
Senior Site Reliability, Platform & Cloud Engineer with 8+ years of experience operating and optimizing mission-critical cloud
platforms and production systems at scale, with a proven track record of improving reliability, availability, performance, and
operational efficiency.
Expertise across AWS, Kubernetes, Docker, Terraform, IaC, CI/CD, Linux, Python, Bash, Ansible,GitLab, Datadog, Prometheus,
Grafana, Splunk, and CloudWatch.
Proven ability to reduce MTTR, eliminate operational toil, strengthen production reliability, and automate infrastructure and
operations through observability, incident response, RCA, performance engineering, and disaster recovery.
Strong SRE foundation across SLI/SLO, SLA, error budgets, incident management, resilience, high availability, and production
readiness, with AIOps and AI/ML-assisted operations for intelligent monitoring, anomaly detection, troubleshooting, and
operational automation.