Talha Imtiaz
Talha ImtiazDevOps Engineer

Open to the right fully remote DevOps role·Based in Lahore, Pakistan · working US/EU hours

DevOps engineer for production systems that have to hold.

Talha Imtiaz

Kubernetes and GKE autoscaling, CI/CD, zero-trust networking, incident response — and the application code underneath.

~80 → 400+ Concurrent users on GKE6 min → 2 min Next.js build time3 of 3 → 0 of 4 Databases on the public internet

I read the codebase, not just the infra

I’m Talha a DevOps engineer who came up as a software engineer first. I work across Kubernetes, GKE, KEDA autoscaling, Dockerized services, GitLab and GitHub CI/CD, observability, incident response, and hardened Linux deployments.

Based in Lahore, Pakistan · working US/EU hours. Open to the right fully remote DevOps role.

From fragile to production-grade

Architecture diagram of a multiplayer game on GKE: players reach Next.js frontend pods scaled by KEDA against a custom player-threshold metric API, with the API service and the Hocuspocus collaboration server split into separate single-container pods; kube-prometheus-stack scrapes all three for the metrics that drive scaling.
Kubernetes Scaling
Concurrent users on GKE
~80 → 400+

Adding Pods Made It Worse: ~80 → 400+ Concurrent Users on GKE

Adding pods made it worse, because Hocuspocus broadcast overhead grows n×(n−1) across pods. The fix was fewer, larger pods sized from load tests, API and Hocuspocus split apart, and KEDA scaling on waiting-room player count instead of CPU — ~80 to 400+ concurrent users on GKE.

Read the Full Case Study

Names and live URLs stay private. Happy to walk through any of these systems in detail on a call.

The teams and systems behind these numbers

  • Rebuilt platform access on GKE as a zero-trust WireGuard overlay: databases reachable from the public internet went from 3 of 3 to 0 of 4, with per-environment CI identities and two human access tiers.
  • Raised capacity from ~80 to 400+ concurrent users with KEDA-based autoscaling on GKE.
  • Cut Next.js build time from 6 min to 2 min by pruning the Docker build context and caching Turborepo outputs.
  • Made noisy Datadog alert patterns inspectable in Slack through alert grouping and heatmap views.
  • Led response to a brute-forced self-hosted VM: traced a respawning cryptominer to its watchdog and cron/systemd persistence, cut its C2 egress, and closed the entry path with key-only SSH and firewall rules.
  • Implemented RobinRelay alert triage workflows with FastAPI, Datadog, Slack, Azure OpenAI, and n8n automation.
  • Automated mobile releases with GitLab CI and Expo EAS, including real-device Appium runs on BrowserStack.
  • Managed Coolify, Dokploy, on-prem servers, and self-hosted GitLab Appium runners.
live_deployment

Employment references

I'd confidently recommend Talha to any team looking for a reliable, autonomous DevOps or Platform Engineer.

Saad Abdullah

Saad Abdullah

Cloud, DevOps, & Solutions Architecture · Toptal

Direct Manager

I would trust Talha with an ambiguous production issue and expect him to return not just with a fix, but with a clear understanding of the root cause.

Abdullah Tarar

Abdullah Tarar

Cloud Infrastructure & DevOps Engineer

Senior Teammate

Stack

  • Kubernetes
  • GKE
  • Docker
  • Terraform
  • Pulumi
  • GitHub Actions
  • GitLab CI
  • AWS
  • Cloudflare
  • Datadog
  • Prometheus
  • Linux
  • Next.js
  • TypeScript
Reference architecture · source public

Built in public, so you can read every line.

Platform Engineering

Built a Reference GitOps Platform on AWS EKS with Zero Static Credentials

Designed and published a reference AWS EKS platform using Terraform and GitOps. Eliminated static cloud credentials using IRSA and automated secret delivery with External Secrets Operator, enabling fully declarative infrastructure and application delivery.

No users, no traffic — built to demonstrate the architecture.

AWS EKSTerraformArgoCDExternal Secrets OperatorAWS Secrets Manager

Consulting

I also take a small number of fixed-scope reviews for early-stage teams — a written read on what breaks first under load, ranked by blast radius. The sample report is public and ungated.

Published IEEE research, and where I trained

Formal education and published undergraduate research.

Hiring, or need a read on your infrastructure?

If you are hiring for production infrastructure, send the role — email is fastest, and the résumé is one click. If you are not hiring but your deploy path, CI/CD, or AI-built app needs a serious review, the consulting pages have the scope and the pricing.

Remote setup: EOR or B2B contractor works best from Pakistan. Direct local employment is usually the route that gets stuck.

Or book a 30-minute call →

© 2026 TALHA IMTIAZ