(Qualys is a leading provider of cloud-based security, compliance, and IT asset management solutions trusted by thousands of enterprises globally.)
- Infrastructure at scale: Manage cloud infrastructure at enterprise scale across 15+ Shared Qualys Cloud Platforms and 50+ Private Cloud Platforms, maintaining 99.95%+ service availability across all production environments.
- Designated POC SRE: Serve as designated Point-of-Contact SRE for Qualys's Vulnerability Management (VM), Policy Compliance (PC), and PCI Compliance product — the primary escalation point for Engineering, Support, and Customer teams across upgrades, migrations, and incidents.
- Incident command and P1/P2 response: Led triage and resolution for daily P1/P2 incidents impacting all Qualys customers — coordinated across Engineering, NetOps, SA, DBA, and Support, joined customer calls, and facilitated blameless postmortems with actionable follow-through.
- Application upgrades and change management: Own the end-to-end SRE lifecycle for application upgrade and infrastructure change on the product — from pre-upgrade planning and environment validation through deployment execution, health verification, and post-change monitoring.
- VMSP migration & maintenance (POC, co-led): Co-led migration from legacy PHP-based VM backend processing platform to Kubernetes-native Java application — ensuring seamless cutover with zero disruption across Engineering, NetOps, DBA, and Customer Success teams for one of Qualys's most sensitive, high-throughput backend systems processing all vulnerability and agent scans.
- Oracle-to-S3 data migration (POC, co-led): Co-led the production data storage migration from Oracle Database to Amazon S3 (Massif platform) — a large-scale, customer-impacting transition requiring phased cutover, data integrity validation, and rollback planning.
- Deep application-level troubleshooting: Debugged complex production issues across distributed stack built on Java, PHP, Apache Kafka, Oracle DB, Apache Ignite — diagnosed root causes beyond infrastructure at application and data layer.
- Observability engineering: Reduced alert fatigue by 40% and improved on-call signal quality by architecting observability pipelines with Prometheus, Grafana, ELK Stack, and Oracle metrics.
- SLO/SLI and on-call management: Reduced MTTR by 30% by automating escalation workflows and driving continuous reliability improvements, while defining and tracking SLOs, SLIs, and error budgets and managing PagerDuty on-call rotations.
- Automation and toil reduction: Built and maintained a library of Bash and Python scripts for automated diagnostics, health checks, deployment validation, and operational workflows — reducing recurring manual toil across the team.