Suraj Kumar

Platform Engineer at Kotak Mahindra Bank

There is no compression algorithm for experience. You can’t learn certain lessons without going through the curve

  • Every day Kubernetes · Terraform · Python · AWS · Bash
  • Regularly Istio · Kong · FastAPI · Azure DevOps
  • Learning Go, slowly
Years doing this
7
Clouds run in production
3
CVE triage, down from 3 days
20 min
Platform tickets a week, down from 15
2
01 Work

Selected work

Three projects. Real details. No template.

CVE Visibility Platform

Kotak Mahindra Bank

The security team got vulnerability reports as PDFs every quarter. By the time anyone read them, the data was already stale. Nobody could answer a simple question: is CVE-2024-XXXXX running in production right now?

I built a Python service that pulls container scan results and joins them against what is actually scheduled in the cluster — not what is in the registry, what is running. A critical CVE in an internal batch job gets deprioritised. A medium CVE in an internet-facing pod gets flagged immediately.

What changed

  • Security team went from 3-day manual triage to about 20 minutes.
  • Auditors get the same dashboard engineers use — no more separate compliance reports nobody trusts.
  • First time the platform team got a thank-you email from security.

Stack: Python, Kubernetes API, container scanning, a lot of Slack threads

Kubernetes Self-Service API

FiftyFive Technologies

We were a 4-person platform team supporting multiple client engagements. Every day: "Can you create a namespace?" "Can you scale this up?" "I need to see the logs for pod-whatever."

I built a FastAPI service that let product teams do this themselves. One API across EKS, AKS and GKE. Quotas and guardrails baked in — if you try to request 100 CPU cores, it says no before anything gets created.

What changed

  • Platform tickets dropped from about 15 a week to 2–3.
  • Developers stopped waiting 4 hours for a namespace.
  • I stopped being a human kubectl proxy.

Stack: Python, FastAPI, Kubernetes API, OIDC, Terraform

Health Monitoring System

Apica · 1,075 days

Nomad and Puppet failures were invisible until something downstream broke. By then you were already in incident mode.

I wrote a Python monitor that polls the scheduler and config management layer directly — the services everything else depends on. Not the apps. The plumbing. Output goes straight into the alerting pipeline we already respond to. No new dashboard nobody watches.

What changed

  • Silent infrastructure failures became alerts people already had a habit of answering.
  • The same visibility exposed idle and oversized resources, which fed a cost optimisation effort.
  • I would tell you the exact saving but I signed an NDA. It was enough that finance noticed.

Stack: Python, Nomad, Puppet, the alerting system we already had

02 Career

Experience

  1. DevOps Engineer 2

    Kotak Mahindra Bank · Gurugram, India

    Oct 2024 — Present current

    • Run Kong API Gateway and Istio service mesh for banking workloads.
    • Blue-green deployments on customer-facing systems. A bad release is a switch flip, not a 2 AM rollback.
    • Built the CVE platform described above.
    • Automated database backups into S3 with lifecycle policies. First time our DR drill did not make the auditor wince.
  2. Lead DevOps Engineer

    FiftyFive Technologies · India

    Jul 2020 — Oct 2024

    • Managed production across Azure, AWS and GCP simultaneously. Would not recommend, but we made it work.
    • Built the Kubernetes self-service API. Still my favourite project.
    • Deployed Kubeflow for ML workflows. ML engineers are a different species — respect.
    • Led a team of 4, ran multiple client projects, somehow did not drop anything.
  3. DevOps Developer

    RoboMQ · India

    Jan 2020 — May 2020

    • First job. First Dockerfiles. First time I deleted a namespace in prod (it was a test cluster, but still).
    • Built DevLogger — a company-wide log system with RabbitMQ. My first production system. It had bugs. It worked anyway.

Client engagements

Apica

1,075 days

Puppet, Nomad, cost optimisation, and the health monitor described above.

NIBE

1,835 days

Azure IoT platform, device provisioning, AKS.

03 Tech

Tech I actually use

Every day

  • Kubernetes
  • Terraform
  • Python
  • AWS
  • Bash

Regularly

  • Istio
  • Kong
  • FastAPI
  • Azure DevOps
  • Docker
  • Git

When needed

  • GCP
  • Azure
  • Go
  • Puppet
  • Nomad
  • Cloud Build
  • CodePipeline

Learning

  • Go — writing more of it, slowly

Opinionated about: YAML is a bad config format and we all know it. We use it anyway.

Tools I built

10 browser-based utilities for DevOps tasks. No backend, no tracking, no AI. Just working code — a visual subnet calculator, a PCAP analyzer, a JWT decoder and the rest.

Open the tools →

Certifications & education

I have them. They helped me learn. But here is the truth: CKA taught me more than CKS, and AZ-400 taught me more than AZ-104. The hands-on ones matter. The associate-level ones are checkboxes.

  • CKA
  • CKAD
  • CKS
  • AZ-400
  • AZ-104
  • RHCSA

B.Tech, Computer Science Engineering — JECRC University, Jaipur, 2020.

05 Contact

Contact

I'm currently open to senior platform and SRE roles. If you're building infrastructure that other engineers ship on, I want to talk.

Location
Gurugram, India
GitHub
@imsurajkr