- ↳Rebuilt platform access on GKE as a zero-trust WireGuard overlay: databases reachable from the public internet went from 3 of 3 to 0 of 4, with per-environment CI identities and two human access tiers.
- ↳Raised capacity from ~80 to 400+ concurrent users with KEDA-based autoscaling on GKE.
- ↳Cut Next.js build time from 6 min to 2 min by pruning the Docker build context and caching Turborepo outputs.
- ↳Made noisy Datadog alert patterns inspectable in Slack through alert grouping and heatmap views.
- ↳Led response to a brute-forced self-hosted VM: traced a respawning cryptominer to its watchdog and cron/systemd persistence, cut its C2 egress, and closed the entry path with key-only SSH and firewall rules.
- ↳Implemented RobinRelay alert triage workflows with FastAPI, Datadog, Slack, Azure OpenAI, and n8n automation.
- ↳Automated mobile releases with GitLab CI and Expo EAS, including real-device Appium runs on BrowserStack.
- ↳Managed Coolify, Dokploy, on-prem servers, and self-hosted GitLab Appium runners.
Open to the right fully remote DevOps role·Based in Lahore, Pakistan · working US/EU hours
DevOps engineer for production systems that have to hold.

Kubernetes and GKE autoscaling, CI/CD, zero-trust networking, incident response — and the application code underneath.
I read the codebase, not just the infra
I’m Talha — a DevOps engineer who came up as a software engineer first. I work across Kubernetes, GKE, KEDA autoscaling, Dockerized services, GitLab and GitHub CI/CD, observability, incident response, and hardened Linux deployments.
Based in Lahore, Pakistan · working US/EU hours. Open to the right fully remote DevOps role.
From fragile to production-grade
- Read the Adding Pods Made It Worse: ~80 → 400+ Concurrent Users on GKE case study
- Read the Zero-Trust Access for a Regulated Fintech on GKE case study
- Read the React Native CI/CD and BrowserStack E2E QA Pipeline case study
- Read the Datadog Alert Triage Automation for Slack-Based SRE Workflows case study
- Concurrent users on GKE
- ~80 → 400+
Adding Pods Made It Worse: ~80 → 400+ Concurrent Users on GKE
Adding pods made it worse, because Hocuspocus broadcast overhead grows n×(n−1) across pods. The fix was fewer, larger pods sized from load tests, API and Hocuspocus split apart, and KEDA scaling on waiting-room player count instead of CPU — ~80 to 400+ concurrent users on GKE.
Names and live URLs stay private. Happy to walk through any of these systems in detail on a call.
The teams and systems behind these numbers
Employment references
“I'd confidently recommend Talha to any team looking for a reliable, autonomous DevOps or Platform Engineer.”
“I would trust Talha with an ambiguous production issue and expect him to return not just with a fix, but with a clear understanding of the root cause.”
Stack
Kubernetes
GKE
Docker
Terraform
Pulumi
GitHub Actions
GitLab CI
- AWS
Cloudflare
Datadog
Prometheus
Linux
Next.js
TypeScript
Built in public, so you can read every line.
Built a Reference GitOps Platform on AWS EKS with Zero Static Credentials
Designed and published a reference AWS EKS platform using Terraform and GitOps. Eliminated static cloud credentials using IRSA and automated secret delivery with External Secrets Operator, enabling fully declarative infrastructure and application delivery.
No users, no traffic — built to demonstrate the architecture.
Software work that taught me what breaks
Working through these problems in public
Latest writing on DevOps, CI/CD, containers, and platform engineering across Medium and Carbonteq.

Dotfiles as Infrastructure: Managing Developer Environments with Chezmoi
How chezmoi turns shell configs, Git settings, editor setup, and terminal tooling into a reproducible developer environment.
- IaC
- Dotfiles
- Linux
Recent Writing
03 / 05 Articles

Building an Appium E2E QA Pipeline with BrowserStack and GitLab CI
How I automated mobile QA by running Appium tests on BrowserStack, generating reports, emailing results, and uploading artifacts back to GitLab.

Building a Smart GitLab CI/CD Pipeline for React Native: Native Builds, OTA Updates, and Release Control
How I built a GitLab pipeline that decides whether a React Native change needs a full Expo native build or can safely ship as an OTA update.

From Docker to Podman: Understanding Rootless, Daemonless Containers
A practical explanation of how Podman differs from Docker, why rootless containers matter, and how to run a simple app with Podman.
Consulting
I also take a small number of fixed-scope reviews for early-stage teams — a written read on what breaks first under load, ranked by blast radius. The sample report is public and ungated.
Published IEEE research, and where I trained
Formal education and published undergraduate research.
GIKI
Ghulam Ishaq Khan Institute of Engineering Sciences and Technology
Bachelor of Science in Computer Science
3rd Position
FYP Expo @ GIKI
A Novel Genetic Algorithm Timetable Generator
CareCloud Tour
Sponsored Industry Tour
Hiring, or need a read on your infrastructure?
If you are hiring for production infrastructure, send the role — email is fastest, and the résumé is one click. If you are not hiring but your deploy path, CI/CD, or AI-built app needs a serious review, the consulting pages have the scope and the pricing.
Remote setup: EOR or B2B contractor works best from Pakistan. Direct local employment is usually the route that gets stuck.
Or book a 30-minute call →© 2026 TALHA IMTIAZ



