Linux


Site Reliability Engineer with an experience as an on-call Incident Manager, expertise in DevOps tools, such as Linux, Ansible, Docker, Kubernetes, AWS, Spunk, Datadog etc. I am dedicated to ensuring the stability, availability, and performance of complex systems. With 6+ years of experience, I have a proven track record of effectively managing and resolving critical incidents. I thrive in high-pressure environments and excel at coordinating cross-functional teams to minimize downtime and restore services swiftly. Strive to achieve more professional responsibilities in an effective manner with full integrity and zest.
System Reliability: Managed the reliability, availability, and performance of critical production systems, focusing on stability and customer experience.
Production Ownership: Took end-to-end ownership of production services, proactively identifying reliability risks and implementing preventive measures to minimize customer impact.
DevOps Architecture: Collaborated with cross-functional teams to design and document end-to-end DevOps architecture, defining Git/source control, CI/CD, AWS EC2 or Kubernetes deployment, networking, security, monitoring, and logging.
Cloud Reliability: Designed, operated, and improved highly available AWS infrastructure with a focus on reliability, scalability, and resilience.
Infrastructure as Code: Provisioned and managed infrastructure using Terraform and Ansible, enabling consistent and repeatable environments.
CI/CD: Built and improved CI/CD pipelines to automate application delivery, reduce deployment risks, and enable reliable software releases.
Cost Optimization: Monitored and optimized cloud costs through right-sizing, autoscaling, and resource cleanup while maintaining reliability and performance.
Reliability Engineering: Identified single points of failure, service dependencies, bottlenecks, and operational risks, implementing solutions to improve system resilience.
Security: Implemented infrastructure and application security best practices, including IAM, least-privilege access, encryption, vulnerability remediation, and secure configurations.
Incident Management: Led incident response, troubleshooting, service restoration, and post-mortem analysis for critical production incidents.
Monitoring & Observability: Developed and maintained monitoring dashboards, metrics, logs, traces, and alerts to provide visibility into system health and performance.
Alert Management: Improved alert quality by tuning thresholds, reducing noisy alerts, and ensuring alerts were actionable and aligned with service health.
Service Level Objectives: Defined, measured, and reported SLIs, SLOs, SLAs, and error budgets to drive reliability improvements and prioritize engineering efforts.
Toil Reduction: Identified repetitive operational tasks and eliminated them through automation, self-service tooling, and improved engineering processes.
MTTD/MTTR: Improved Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through enhanced monitoring, automation, runbooks, and incident-response processes.
24/7 On-Call: Owned front-line 24/7 on-call rotations and incident response for critical production systems, serving as the first line of defense for production incidents.
Root Cause Analysis: Performed detailed root-cause analysis and implemented permanent corrective actions to prevent recurring production issues.
Documentation: Created and maintained runbooks, architecture documentation, troubleshooting guides, operational procedures, and incident-response documentation.
Cloud Hosting: Managed and supported cloud-hosted environments, ensuring reliable infrastructure, efficient resource utilization, scalability, and high availability.
Server Management: Managed and maintained production Linux servers, including provisioning, configuration, patching, troubleshooting, and performance optimization.
Virtualization: Managed virtual machines including resource allocation, provisioning, migration and performance optimization.
Linux Administration: Performed Linux system administration, including package management, filesystem management, networking, user administration, service management, and system troubleshooting.
Network Troubleshooting: Troubleshot DNS, TCP/IP, routing, firewall, load balancer, connectivity, and network performance issues affecting production services.
Performance Optimization: Analyzed and optimized server performance by troubleshooting CPU, memory, disk I/O, network utilization, processes, and application behaviour.
Production Operations: Provided end-to-end operational support for production infrastructure, from server provisioning and deployment through monitoring, troubleshooting, optimization, and decommissioning.
Log Management: Analyzed system, application, web-server, and security logs to troubleshoot incidents, identify performance issues, and detect abnormal behaviour.
Customer-Facing Support: Provided technical support for critical hosting and infrastructure issues, effectively communicating with customers and internal engineering teams during incidents.
Safe Deployments: Implemented safe deployment and automated rollback strategies to minimize production impact during application and infrastructure changes.
Mentorship: Mentored junior engineers and promoted SRE principles, operational excellence, automation, observability, and reliability best practices across engineering teams.
Travelling
Playing Badminton
Listening Music
Red Hat Certified System Administrator (RHCSA)
Red Hat Certified System Administrator (RHCSA)
Certified Kubernetes Administrator (CKA)
Linux
Listening Music
Travelling
Playing Badminton