Summary
Overview
Work History
Education
Skills
LinkedIn
Credly
Certifications
Hobbies and interests
Timeline
Generic
Harshit Jain

Harshit Jain

Rajsamand, Rajasthan,India

Summary

Site Reliability Engineer with 8+ years of experience in SRE, DevOps, cloud infrastructure, and production operations, specializing in highly available and scalable systems across AWS and Linux environments. Proven track record of maintaining 99.99% uptime and reducing MTTR by 40% through effective incident management, automation, troubleshooting, and root cause analysis. Hands-on experience with Kubernetes, Docker, Terraform, Ansible, Jenkins, CI/CD, Infrastructure as Code, and observability using Prometheus, Grafana, Datadog and Splunk. Experienced in defining SLIs, SLOs, SLAs, and error budgets, designing DevOps architectures, and optimizing system performance, reliability, and cloud costs. Strong in 24/7 on-call operations, cross-functional collaboration, customer coordination, and continuous service improvement.

Overview

8
8
years of professional experience

Work History

Lead Engineer

Nitor Infotech Pvt. Ltd.
Pune, India, (Remote)
01.2024 - 10.2024

System Reliability: Managed the reliability, availability, and performance of critical production systems, focusing on stability and customer experience.

Production Ownership: Took end-to-end ownership of production services, proactively identifying reliability risks and implementing preventive measures to minimize customer impact.

DevOps Architecture: Collaborated with cross-functional teams to design and document end-to-end DevOps architecture, defining Git/source control, CI/CD, AWS EC2 or Kubernetes deployment, networking, security, monitoring, and logging.

Cloud Reliability: Designed, operated, and improved highly available AWS infrastructure with a focus on reliability, scalability, and resilience.

Infrastructure as Code: Provisioned and managed infrastructure using Terraform and Ansible, enabling consistent and repeatable environments.

CI/CD: Built and improved CI/CD pipelines to automate application delivery, reduce deployment risks, and enable reliable software releases.

Cost Optimization: Monitored and optimized cloud costs through right-sizing, autoscaling, and resource cleanup while maintaining reliability and performance.

Reliability Engineering: Identified single points of failure, service dependencies, bottlenecks, and operational risks, implementing solutions to improve system resilience.

Security:Implemented infrastructure and application security best practices, including IAM, least-privilege access, encryption, vulnerability remediation, and secure configurations.

Sr. Associate Technology, L-1 Site Reliability Engineer

Chirpn IT Solutions LLP
Pune, India (Remote) , Worked for US based Cyber Security Client.
11.2022 - 06.2023

Incident Management: Led incident response, troubleshooting, service restoration, and post-mortem analysis for critical production incidents.

Monitoring & Observability: Developed and maintained monitoring dashboards, metrics, logs, traces, and alerts to provide visibility into system health and performance.

Alert Management: Improved alert quality by tuning thresholds, reducing noisy alerts, and ensuring alerts were actionable and aligned with service health.

Service Level Objectives: Defined, measured, and reported SLIs, SLOs, SLAs, and error budgets to drive reliability improvements and prioritize engineering efforts.

Toil Reduction: Identified repetitive operational tasks and eliminated them through automation, self-service tooling, and improved engineering processes.

MTTD/MTTR: Improved Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through enhanced monitoring, automation, runbooks, and incident-response processes.

24/7 On-Call: Owned front-line 24/7 on-call rotations and incident response for critical production systems, serving as the first line of defense for production incidents.

Root Cause Analysis: Performed detailed root-cause analysis and implemented permanent corrective actions to prevent recurring production issues.

Documentation: Created and maintained runbooks, architecture documentation, troubleshooting guides, operational procedures, and incident-response documentation.

Sr. Site Reliability Engineer

GIP Technologies Pvt. Ltd.
Jaipur, India
09.2016 - 10.2022

Cloud Hosting: Managed and supported cloud-hosted environments, ensuring reliable infrastructure, efficient resource utilization, scalability, and high availability.
Server Management: Managed and maintained production Linux servers, including provisioning, configuration, patching, troubleshooting, and performance optimization.

Virtualization: Managed virtual machines including resource allocation, provisioning, migration and performance optimization.

Linux Administration: Performed Linux system administration, including package management, filesystem management, networking, user administration, service management, and system troubleshooting.

Network Troubleshooting: Troubleshot DNS, TCP/IP, routing, firewall, load balancer, connectivity, and network performance issues affecting production services.

Performance Optimization: Analysed and optimized server performance by troubleshooting CPU, memory, disk I/O, network utilization, processes, and application behaviour.

Production Operations: Provided end-to-end operational support for production infrastructure, from server provisioning and deployment through monitoring, troubleshooting, optimization, and decommissioning.

Log Management: Analyzed system, application, web-server, and security logs to troubleshoot incidents, identify performance issues, and detect abnormal behaviour.

Customer-Facing Support: Provided technical support for critical hosting and infrastructure issues, effectively communicating with customers and internal engineering teams during incidents.

Safe Deployments: Implemented safe deployment and automated rollback strategies to minimize production impact during application and infrastructure changes.

Mentorship: Mentored junior engineers and promoted SRE principles, operational excellence, automation, observability, and reliability best practices across engineering teams.

Education

Bachelor of Commerce -

Mohanlal Sukhadia University
Udaipur, Rajasthan, India
2015

12th Standard -

Lakshmipat Singhania School
Rajsamand, Rajasthan, India
2012

Skills

  • Linux
  • AWS Cloud
  • Azure Cloud
  • Terraform
  • Docker
  • Kubernetes
  • Ansible
  • Shell Scripting
  • Python
  • Git
  • CI/CD - Jenkins
  • Prometheus
  • Grafana
  • Kibana
  • DataDog
  • Splunk
  • New Relic
  • PagerDuty
  • Incident Management
  • Fire hydrant
  • Jira Service Management
  • Zendesk

LinkedIn

https://www.linkedin.com/in/harshit-jain-85a312251

Credly

https://www.credly.com/users/harshit-jain.19f4fd84

Certifications

1. RedHat Certified System Administrator (RHCSA),

Certification ID: 230-105-447, May, 2023 - May, 2026.

2. Certified Kubernetes Administrator (CKA),

Certification ID: LF-cu55n9v41j, May, 2023 - May, 2026

3. Introduction to Cyber Security by CISCO, August, 2023

Hobbies and interests

  • Travelling
  • Listening Music
  • Playing Badminton

Timeline

Lead Engineer

Nitor Infotech Pvt. Ltd.
01.2024 - 10.2024

Sr. Associate Technology, L-1 Site Reliability Engineer

Chirpn IT Solutions LLP
11.2022 - 06.2023

Sr. Site Reliability Engineer

GIP Technologies Pvt. Ltd.
09.2016 - 10.2022

Bachelor of Commerce -

Mohanlal Sukhadia University

12th Standard -

Lakshmipat Singhania School
Harshit Jain