Summary
Overview
Work History
Education
Skills
Timeline
Generic

Kiran Mettu

Summary

Results-driven Site Reliability Engineer with over 7+ years of experience in DevOps, SaaS infrastructure, and network engineering. Specialized in system observability, cloud automation, and Kubernetes operations. Proven ability to maintain highly available, scalable platforms, while actively collaborating with development teams to resolve infrastructure and performance issues. Resourceful Site Reliability Engineer known for high productivity and efficient task completion. Skilled in automation, continuous integration and delivery (CI/CD), and cloud infrastructure management. Excel in problem- solving, collaboration, and adaptability, ensuring seamless operations and system reliability.

Overview

8
8
years of professional experience

Work History

Site Reliability Engineer

JPMorgan Chase
Hyderabad
12.2025 - Current
  • Managed production applications and AWS infrastructure, ensuring availability, reliability, scalability, and performance.
  • Monitored applications and infrastructure using Grafana, Prometheus, Splunk, and AWS CloudWatch.
  • Created and optimized data sources to improve monitoring, query, and dashboard performance.
  • Implemented Z-Score anomaly detection to identify abnormal application and infrastructure behavior proactively.
  • Developed AI agents for automated RCA, analyzing alerts, logs, metrics, and incident data.
  • Created AI-driven Grafana dashboard automation, reducing manual dashboard creation and configuration.
  • Automated repetitive SRE tasks using Python, Shell scripting, APIs, and AI agents.
  • Built dashboards for latency, availability, errors, throughput, CPU, memory, and application health.
  • Supported AWS EC2, Lambda, ECS/Fargate, EKS, Route 53, CloudWatch, S3, and IAM.
  • Supported Kubernetes/EKS and Docker environments and troubleshooting application and container issues.
  • Performed Linux production support, log analysis, troubleshooting, and system monitoring.
  • Created and maintained SOPs, runbooks, RCA documents, and troubleshooting guides.
  • Contributed to observability, reliability, automation, capacity planning, and disaster recovery initiatives.

Site Reliability Engineer

Nokia
Bangalore
06.2022 - 12.2025
  • Manage Datadog observability stack: configure dashboards, alerts, and synthetic tests.
  • Troubleshoot critical incidents and performance issues in Kubernetes clusters.
  • Configure and maintain site-to-site VPNs as per evolving infrastructure requirements.
  • Collaborate with Dev and Infra teams to ensure high availability of SaaS services.
  • Participate in on-call rotation and incident response to minimize service disruption.
  • Automate infrastructure provisioning using Terraform for hybrid cloud environments.
  • Designed, implemented, and managed monitoring and alerting solutions to ensure 99.9%+ uptime and rapid incident response.
  • Created and maintained Service Level Agreements (SLAs) and conducted Root Cause Analysis (RCA) to identify and resolve system failures.

DevOps Engineer

L&T TTS (Client: intel)
04.2021 - 05.2022
  • Designed and implemented CI/CD pipelines using Jenkins to automate build, test, and deployment processes.
  • Automated containerized application deployments using Docker and orchestrated with Kubernetes on AWS EKS. Integrated GitHub with Jenkins and used Maven for build automation and SonarQube for code quality checks.
  • Deployed and managed scalable cloud infrastructure using AWS EC2, S3, and RDS, ensuring high availability.
  • Implemented infrastructure automation with Terraform and AWS CloudFormation to provision and manage resources.
  • Ensured security and compliance by managing IAM roles and automating secret management with AWS Secrets Manager.
  • Optimized workflows and improved collaboration by enforcing version control practices with Git and GitHub. (Client: Intel)

Network Engineer

Value Point (Client: Spar Hypermarket)
Bangalore
03.2019 - 04.2021
  • Provided L1 and L2 support for enterprise networks and handled remote troubleshooting.
  • Configured VLANs, static routes, and IPsec/GRE VPN tunnels on Fortinet and Cisco devices.
  • Monitored network infrastructure using WhatsUp Gold and FortiAnalyzer.
  • Managed DHCP and Mail servers, user access control, and bandwidth utilization policies.

Education

B.Tech -

JNTUH
01-2016

Skills

  • AWS (EC2, S3, VPC, IAM, Route 53, Auto Scaling)
  • GCP
  • Jenkins
  • Git
  • GitHub
  • Gerrit
  • Datadog (alerts, dashboards, log analysis, synthetics)
  • Nagios
  • Terraform
  • Linux
  • Windows
  • VPN
  • VLANs
  • SD-WAN
  • Fortinet Firewalls
  • Shell
  • JIRA
  • UTP

Timeline

Site Reliability Engineer

JPMorgan Chase
12.2025 - Current

Site Reliability Engineer

Nokia
06.2022 - 12.2025

DevOps Engineer

L&T TTS (Client: intel)
04.2021 - 05.2022

Network Engineer

Value Point (Client: Spar Hypermarket)
03.2019 - 04.2021

B.Tech -

JNTUH
Kiran Mettu