Accomplished Site Reliability Engineer with 7.6 years of experience delivering resilient, high-availability platforms across observability, production support, infrastructure operations, and enterprise application monitoring. Hands-on expertise in AIOps-driven monitoring, DataDog administration, alert engineering, SLO/SLI/SLA governance, incident response, change management, and post-incident analysis within 24/7 mission-critical environments. Proven ability to improve operational stability through alert noise reduction, dashboard standardization, service onboarding, automation with Ansible and Bash, and cross-functional collaboration with development, infrastructure, and business teams. Strong technical foundation in Linux/Unix, Windows, AWS, Azure, Kubernetes, VMware, and enterprise monitoring stacks including Grafana, Kibana, Splunk, Dynatrace, AppDynamics, Zabbix, SiteScope, OMI, and PagerDuty.