Continental
Rancher-managed Kubernetes clusters, Helm deployments, Longhorn storage, MinIO, Prometheus, Grafana and Loki. Improved MinIO IOPS by 40% while supporting secure, GDPR-compliant storage.
SHIVAM YADAV // ENGINEERING PROFILE 2026
I build, automate and operate highly available cloud-native systems across Kubernetes, AWS, delivery platforms and observability.
$ whoami
site reliability engineer$ uptime
production: stableA Site Reliability Engineer with four years of hands-on experience designing, automating and maintaining highly available, scalable and secure cloud-native systems.
My work spans Kubernetes, Terraform, AWS, CI/CD, observability and incident management, with a focus on reducing downtime and operational overhead.
SLOs, fault tolerance and operational readiness for production systems.
Terraform, Python, CI/CD and runbooks that remove repetitive work.
Metrics, logs and alerts shaped into signals teams can act on.
Predictable platforms, visible signals and delivery paths that give teams room to do their best work.
Rancher-managed Kubernetes clusters, Helm deployments, Longhorn storage, MinIO, Prometheus, Grafana and Loki. Improved MinIO IOPS by 40% while supporting secure, GDPR-compliant storage.
Leading cross-team response from detection and stakeholder communication through postmortems.
GitHub Copilot, log analysis and incident summaries reduced investigation time by 20–30%.
Responsive Angular frontend and Spring Boot backend deployed securely in an AWS VPC with PostgreSQL.
Roles, systems and responsibilities across cloud infrastructure, identity, delivery and production support.
Incident Commander for business-critical shipping platforms. Rebuilt observability to cut MTTD 35% and alert noise 40%, owned GitLab delivery for 30–50 production deploys per week, and automated AWS operations across 10+ accounts.
Managed Azure AD identity lifecycle, SSO integrations, Conditional Access, MFA and RBAC across the organization.
Automated infrastructure across 100+ environments, defined SLOs, improved fault tolerance and reduced Keycloak upgrade downtime by 85% while maintaining 99.9% uptime.
Deployed a Spring Boot and Angular application on AWS using VPC, CloudWatch, Lambda, S3, SQS, DynamoDB and IAM services.
No arbitrary skill bars. Just the systems and capabilities used to ship and operate.
GLA University, Uttar Pradesh
2018 — 2022
Curiosity is part of the operating system too.
Reliable systems, operational clarity and what keeps production steady.
Small experiments at home to learn, break things safely and understand systems from the inside out.
Taking circuits apart, following signals and learning by making small things work.
Connecting devices, sensors and everyday spaces to the systems behind them.
Books, notes and ideas worth slowing down to understand.
Turning digital designs into physical objects and iterating until the idea fits in your hands.
Plato
READY WHEN YOU ARE
I'm interested in difficult infrastructure problems, distributed systems, reliability and automation.