DevOps Engineer · Ahmedabad, India

Kashish Lakhara

I build infrastructure that holds when things fail.

I keep multiple production Kubernetes clusters running across AWS, GCP, Azure, and air-gapped bare metal, from control-plane bootstrap to the incidents that page at 3am.

About

I treat infrastructure as engineering not guesswork, automate the repetitive work, trace every failure to its root cause, and build so the same problem never has to be solved twice.

Lately that's pulling me deeper into observability: past the dashboards, into why systems actually behave the way they do.

Worth talking about

  • Traced an error-budget-burn alert to a dying disk

    KubeAPIErrorBudgetBurn on a production control plane. smartctl found failing hardware under the etcd fsync latency — swapped the disk live, near-zero downtime.

    read the writeup →
  • Bootstrapped highly available Kubernetes on bare metal

    Ansible and kubeadm from scratch, with HAProxy and Keepalived holding a virtual IP in front of the API servers. Failover validated by pulling nodes on Proxmox.

  • Moved production from on-premise to AWS EKS

    Migrated live workloads, added Karpenter for autoscaling, then spent the following weeks cutting the bill back down.

Everything else I've worked on →

Stack

Kubernetes · kubeadm · EKS · GKE · FluxCD · Terraform · Ansible · Prometheus · Grafana · OpenSearch · Fluent Bit · Longhorn · CrateDB · PostgreSQL · HAProxy · MetalLB · Proxmox · Linux · Bash · Python

Get in touch

Open to collaborating on interesting problems. Reach me at kashishlakhara04@gmail.com, or use the form below.