Summary
Overview
Work History
Education
Skills
Certification
Timeline
Languages
BusinessAnalyst
Anil Singh

Anil Singh

Senior Site Reliability Engineer
Pune

Summary

Senior Site Reliability Engineer with extensive experience in designing, maintaining, and optimizing highly available, scalable, and secure enterprise SaaS platforms. Served as the primary escalation point for Qualys Vulnerability Management (VM), Policy Compliance (PC), and PCI Compliance products, partnering with Engineering, Customer Support, and cross-functional teams to resolve critical production issues. Achieved 99.95% service availability across shared and private cloud environments through proactive monitoring, incident management, automation, and reliability engineering best practices. Led large-scale infrastructure modernization initiatives, including legacy platform migrations, architecture upgrades, and data storage migrations, improving system performance, scalability, and operational resilience.

Overview

7
7
Certificate
6
6
years of professional experience

Work History

Senior Site Reliability Engineer — Cloud Platform

Qualys
Pune
12.2021 - Current

(Qualys is a leading provider of cloud-based security, compliance, and IT asset management solutions trusted by thousands of enterprises globally.)

  • Infrastructure at scale: Manage cloud infrastructure at enterprise scale across 15+ Shared Qualys Cloud Platforms and 50+ Private Cloud Platforms, maintaining 99.95%+ service availability across all production environments.
  • Designated POC SRE: Serve as designated Point-of-Contact SRE for Qualys's Vulnerability Management (VM), Policy Compliance (PC), and PCI Compliance product — the primary escalation point for Engineering, Support, and Customer teams across upgrades, migrations, and incidents.
  • Incident command and P1/P2 response: Led triage and resolution for daily P1/P2 incidents impacting all Qualys customers — coordinated across Engineering, NetOps, SA, DBA, and Support, joined customer calls, and facilitated blameless postmortems with actionable follow-through.
  • Application upgrades and change management: Own the end-to-end SRE lifecycle for application upgrade and infrastructure change on the product — from pre-upgrade planning and environment validation through deployment execution, health verification, and post-change monitoring.
  • VMSP migration & maintenance (POC, co-led): Co-led migration from legacy PHP-based VM backend processing platform to Kubernetes-native Java application — ensuring seamless cutover with zero disruption across Engineering, NetOps, DBA, and Customer Success teams for one of Qualys's most sensitive, high-throughput backend systems processing all vulnerability and agent scans.
  • Oracle-to-S3 data migration (POC, co-led): Co-led the production data storage migration from Oracle Database to Amazon S3 (Massif platform) — a large-scale, customer-impacting transition requiring phased cutover, data integrity validation, and rollback planning.
  • Deep application-level troubleshooting: Debugged complex production issues across distributed stack built on Java, PHP, Apache Kafka, Oracle DB, Apache Ignite — diagnosed root causes beyond infrastructure at application and data layer.
  • Observability engineering: Reduced alert fatigue by 40% and improved on-call signal quality by architecting observability pipelines with Prometheus, Grafana, ELK Stack, and Oracle metrics.
  • SLO/SLI and on-call management: Reduced MTTR by 30% by automating escalation workflows and driving continuous reliability improvements, while defining and tracking SLOs, SLIs, and error budgets and managing PagerDuty on-call rotations.
  • Automation and toil reduction: Built and maintained a library of Bash and Python scripts for automated diagnostics, health checks, deployment validation, and operational workflows — reducing recurring manual toil across the team.

Cloud Operations Engineer

RChilli Inc.
Remote
10.2020 - 12.2021

(RChilli is an AI-driven resume parsing and recruitment data solutions provider serving global HR technology platforms.)

  • Full-stack operations ownership: Delivered end-to-end operational ownership across SRE, system administration, NetOps, DBA, IT support, and security as part of a small, cross-functional operations team.
  • CI/CD pipeline design and automation: Designed and implemented multiple end-to-end CI/CD pipelines using Jenkins, Ansible, and Git — including fully automated upgrade pipelines that handled version rollouts without manual intervention, reducing deployment errors and improving release efficiency by 35%.
  • SonarQube integration: Integrated SonarQube as a mandatory quality gate within CI/CD pipelines — ensuring every package passed static analysis, security vulnerability scanning, and code quality thresholds before reaching production.
  • Custom performance testing: Built and maintained custom performance test scripts executed against every new package prior to production rollout — validating response times, resource utilisation, and system stability under load before any release reached customers.
  • Project and migration participation: Contributed to infrastructure migrations, service deployments, and application rollouts — providing operational input, environment setup, testing, and go-live support across major initiatives.
  • Compliance certifications: Co-led end-to-end ISO 27001 and SOC 2 compliance certification — including risk assessment, gap analysis, security control implementation, and audit readiness documentation.
  • Multi-cloud infrastructure: Managed multi-cloud infrastructure across AWS, GCP, and Oracle Cloud, optimizing compute, networking, storage, access control, and cost.
  • Security management: Managed firewalls, VPNs, SSL/TLS certificates, and vulnerability assessments with full ownership from implementation through audit evidence collection.
  • Incident response: Owned incident management end-to-end from detection through postmortem, improving MTTR by 25%.
  • Disaster recovery: Planned and executed DR activities including backup strategy design, failover testing, and data integrity validation to ensure business continuity.

Education

Bachelor of Engineering - Computer Science & Engineering

Chandigarh University
Mohali, India
01-2020

Skills

  • Cloud Platforms: AWS, Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), Microsoft Azure
  • Infrastructure at Scale: VM provisioning and management (350 VMs per env), Kubernetes (K8s) cluster operations, pod lifecycle management
  • CI/CD & IaC: Jenkins, Ansible, Puppet, Terraform, Git, Bitbucket Pipelines, Code Deployment & Rollback
  • Scripting & Automation: Bash/Shell scripting, Python, custom automation tooling to reduce toil
  • Observability & Monitoring: Prometheus, Grafana, ELK Stack (Elasticsearch, Logstash, Kibana), Splunk, AppDynamics, Oracle Metrics, Site24x7
  • Incident Management: PagerDuty, On-call Rotations, P1/P2 Incident Response, Blameless Postmortems, RCA Documentation, SLO/SLI
  • Security & Compliance: HashiCorp Vault, SSL/TLS, VPN, Firewall Management, PCI Compliance, ISO 27001, SOC 2
  • OS & Databases: Advanced Linux Admin (RHEL, Ubuntu, CentOS), MySQL, Oracle DB, Google Cloud SQL
  • Application Stack: PHP, Java, Python, Apache Kafka, Oracle DB, Apache Ignite, React Native
  • Collaboration: Jira, Confluence, Runbook Authoring, Cross-team coordination (Eng, NetOps, SA, DBA, Support)

Certification

  • Oracle Autonomous Database Cloud Specialist
  • Oracle Cloud Infrastructure Foundation Associate
  • Google Cloud Networking Fundamentals
  • Fortinet NSE 1 & NSE 2 – Network Security
  • CNSS Network Security Specialist – ICSI, UK
  • HCNA Routing & Switching – Huawei
  • Certificate of Appreciation – Huawei ICT Competition India 2018

Timeline

Senior Site Reliability Engineer — Cloud Platform

Qualys
12.2021 - Current

Cloud Operations Engineer

RChilli Inc.
10.2020 - 12.2021

Bachelor of Engineering - Computer Science & Engineering

Chandigarh University

Languages

English
Proficient
C2
Hindi
Proficient
C2
Anil SinghSenior Site Reliability Engineer