
Site Reliability / DevOps Engineer with 5+ years supporting and continuously improving cloud operations for enterprise clients including Thomson Reuters and AT&T, with direct experience across the full incident lifecycle: troubleshooting, deployment and patch management, and monitoring/alerting on AWS and Azure. Sustained 99.8% uptime SLOs and a 99.5% batch job success rate by driving platform stability, release efficiency, and operational maturity across hybrid cloud environments. Hands-on with Kubernetes and containerized workloads, Terraform (Infrastructure as Code), CI/CD pipelines (Jenkins, with working exposure to GitLab CI/CD), VMware/Hyper-V virtualization, and Kafka-adjacent data pipeline operations. Skilled at cross-functional collaboration with DevOps engineers and operations teams to reduce time to detect/resolve and improve customer-facing cloud reliability.
Client: Thomson Reuters
•Drove platform stability and release efficiency by architecting end-to-end Jenkins CI/CD pipelines for Java/Maven applications with Groovy-based Jenkinsfiles, parallelized stages, and automated rollback — cutting deployment cycle time by 40%.
•Maintained 99.8% uptime SLO across 100+ Linux production servers in AWS Cloud Operations by managing deployments and patching, performing vulnerability remediation and compliance validation, and hardening systems — achieving zero critical security incidents.
•Enhanced monitoring, alerting, and early-warning systems using Datadog, Prometheus, Grafana, and CloudWatch with custom dashboards and alert thresholds, improving troubleshooting workflows and reducing time to detect (MTTD).
•Eliminated 50% of recurring operational toil by developing Python and Shell automation for user provisioning, configuration management, and production log monitoring across the fleet.
•Supported daily operations and incident handling in close collaboration with the operations team, leading P1/P2 incident response under ITIL practices — performing RCA and driving systemic fixes that reduced repeat incidents and improved MTTR.
•Collaborated with DevOps engineers and QA in Agile sprints to implement process improvements and platform enhancements, optimizing CI/CD pipeline performance and accelerating release cadence.
•Supported and optimized Kubernetes-based platforms and containerized workloads (EKS/OpenShift) alongside Terraform and AWS CloudFormation, enabling repeatable, production-ready infrastructure provisioning.
Client: AT&T — Enterprise Consolidated Data Warehouse (eCDW)
•Achieved zero-downtime production deployments across hybrid Azure Cloud and on-premises Linux/VMware environments by orchestrating continuous deployments for mission-critical data pipelines serving finance, billing, and customer integration workloads.
•Improved deployment consistency and release governance by owning the complete change lifecycle: authoring migration scripts, executing Linux/Unix deployments, and validating production releases across hybrid, virtualized (VMware/Hyper-V) infrastructure.
•Strengthened cloud security posture by implementing IAM controls, conducting access reviews, and supporting vulnerability remediation across AWS and Azure environments.
•Improved operational visibility for stakeholders by automating weekly and monthly infrastructure health, capacity, compliance, and cloud operations reporting to support proactive planning.