Years of professional experience


Senior DevOps and Site Reliability Engineer (SRE) with 5+ years of experience designing, automating, and supporting highly available, cloud-native production environments across AWS and GCP. Strong expertise in Kubernetes, Docker, Terraform, Jenkins, CI/CD, Python, Linux, and Shell Scripting, with hands-on experience in Infrastructure as Code (IaC), Cloud Infrastructure, Production Support, and Infrastructure Automation.
Experienced in Incident Management, Root Cause Analysis (RCA), Monitoring, Observability, SLIs/SLOs, Performance Optimization, and Platform Reliability. Proven ability to build scalable infrastructure, streamline deployments, automate operational workflows, and collaborate with cross-functional engineering teams to deliver secure and resilient cloud platforms.
Actively expanding expertise in AI Infrastructure and LLMOps, with hands-on experience in Prompt Engineering, AI Agents, Agentic AI, Retrieval-Augmented Generation (RAG), Model Context Protocol (MCP), LangGraph, CrewAI, Microsoft AutoGen, OpenAI APIs, and AI-powered Automation Workflows, combining modern DevOps practices with AI-driven platform engineering.
Years of professional experience
Supported enterprise Linux/Unix environments, ensuring operational stability, application availability, and day-to-day production support for business-critical services.
Monitored application and infrastructure health, investigated production issues, and collaborated with engineering teams to resolve incidents and maintain service reliability.
Performed SQL-based data analysis and troubleshooting to support operational processes and resolve application-related issues.
Managed incident, service request, and change management activities, following established ITIL and operational support procedures.
Investigated application and infrastructure issues using logs, system diagnostics, and database queries, contributing to faster issue resolution and improved operational efficiency.
Worked closely with cross-functional teams to maintain production stability, support business operations, and deliver high-quality customer service.
AWS & GCP, CHEF, ANSIBLE
Kubernete, Docker
Terraform
cloud watch(AWS), SPLUNK
ELK
Zabbix, DataDog
SQL Server
MySQL
shell, Python, groovy, yaml