Summary
Overview
Work History
Education
Skills
Certification
Timeline
Work Availability
Work Preference
Quote
Software
Languages
Interests
Receptionist
Harshit Jain

Harshit Jain

Site Reliability Engineer
Rajsamand,RJ

Summary

Site Reliability Engineer with an experience as an on-call Incident Manager, expertise in DevOps tools, such as Linux, Ansible, Docker, Kubernetes, AWS, Spunk, Datadog etc. I am dedicated to ensuring the stability, availability, and performance of complex systems. With 6+ years of experience, I have a proven track record of effectively managing and resolving critical incidents. I thrive in high-pressure environments and excel at coordinating cross-functional teams to minimize downtime and restore services swiftly. Strive to achieve more professional responsibilities in an effective manner with full integrity and zest.

Overview

2
2
Languages
2
2
Certificates
3
3
years of post-secondary education
8
8
years of professional experience
2
2
Languages

Work History

Lead Engineer

Nitor Infotech Pvt Ltd
Pune, Maharashtra
01.2024 - 10.2024

System Reliability: Managed the reliability, availability, and performance of critical production systems, focusing on stability and customer experience.

Production Ownership: Took end-to-end ownership of production services, proactively identifying reliability risks and implementing preventive measures to minimize customer impact.

DevOps Architecture: Collaborated with cross-functional teams to design and document end-to-end DevOps architecture, defining Git/source control, CI/CD, AWS EC2 or Kubernetes deployment, networking, security, monitoring, and logging.

Cloud Reliability: Designed, operated, and improved highly available AWS infrastructure with a focus on reliability, scalability, and resilience.

Infrastructure as Code: Provisioned and managed infrastructure using Terraform and Ansible, enabling consistent and repeatable environments.

CI/CD: Built and improved CI/CD pipelines to automate application delivery, reduce deployment risks, and enable reliable software releases.

Cost Optimization: Monitored and optimized cloud costs through right-sizing, autoscaling, and resource cleanup while maintaining reliability and performance.

Reliability Engineering: Identified single points of failure, service dependencies, bottlenecks, and operational risks, implementing solutions to improve system resilience.

Security: Implemented infrastructure and application security best practices, including IAM, least-privilege access, encryption, vulnerability remediation, and secure configurations.

Sr. Associate Technology, L-1 Site Reliability Eng

Chirpn IT Solutions LLP
Pune, Maharashtra
11.2022 - 06.2023

Incident Management: Led incident response, troubleshooting, service restoration, and post-mortem analysis for critical production incidents.

Monitoring & Observability: Developed and maintained monitoring dashboards, metrics, logs, traces, and alerts to provide visibility into system health and performance.

Alert Management: Improved alert quality by tuning thresholds, reducing noisy alerts, and ensuring alerts were actionable and aligned with service health.

Service Level Objectives: Defined, measured, and reported SLIs, SLOs, SLAs, and error budgets to drive reliability improvements and prioritize engineering efforts.

Toil Reduction: Identified repetitive operational tasks and eliminated them through automation, self-service tooling, and improved engineering processes.

MTTD/MTTR: Improved Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through enhanced monitoring, automation, runbooks, and incident-response processes.

24/7 On-Call: Owned front-line 24/7 on-call rotations and incident response for critical production systems, serving as the first line of defense for production incidents.

Root Cause Analysis: Performed detailed root-cause analysis and implemented permanent corrective actions to prevent recurring production issues.

Documentation: Created and maintained runbooks, architecture documentation, troubleshooting guides, operational procedures, and incident-response documentation.

Site Reliability Engineer

GIP Technologies Private Limited
Jaipur, Rajasthan
09.2016 - 10.2022

Cloud Hosting: Managed and supported cloud-hosted environments, ensuring reliable infrastructure, efficient resource utilization, scalability, and high availability.

Server Management: Managed and maintained production Linux servers, including provisioning, configuration, patching, troubleshooting, and performance optimization.

Virtualization: Managed virtual machines including resource allocation, provisioning, migration and performance optimization.

Linux Administration: Performed Linux system administration, including package management, filesystem management, networking, user administration, service management, and system troubleshooting.

Network Troubleshooting: Troubleshot DNS, TCP/IP, routing, firewall, load balancer, connectivity, and network performance issues affecting production services.

Performance Optimization: Analyzed and optimized server performance by troubleshooting CPU, memory, disk I/O, network utilization, processes, and application behaviour.

Production Operations: Provided end-to-end operational support for production infrastructure, from server provisioning and deployment through monitoring, troubleshooting, optimization, and decommissioning.

Log Management: Analyzed system, application, web-server, and security logs to troubleshoot incidents, identify performance issues, and detect abnormal behaviour.

Customer-Facing Support: Provided technical support for critical hosting and infrastructure issues, effectively communicating with customers and internal engineering teams during incidents.

Safe Deployments: Implemented safe deployment and automated rollback strategies to minimize production impact during application and infrastructure changes.

Mentorship: Mentored junior engineers and promoted SRE principles, operational excellence, automation, observability, and reliability best practices across engineering teams.

Education

Bachelor of Commerce -

Mohanlal Sukhadia University
Udaipur, Rajasthan, India
06.2012 - 06.2015

Skills

    Travelling

Playing Badminton

Listening Music

Certification

Red Hat Certified System Administrator (RHCSA)

Timeline

Lead Engineer

Nitor Infotech Pvt Ltd
01.2024 - 10.2024

Red Hat Certified System Administrator (RHCSA)

05-2023

Certified Kubernetes Administrator (CKA)

05-2023

Sr. Associate Technology, L-1 Site Reliability Eng

Chirpn IT Solutions LLP
11.2022 - 06.2023

Site Reliability Engineer

GIP Technologies Private Limited
09.2016 - 10.2022

Bachelor of Commerce -

Mohanlal Sukhadia University
06.2012 - 06.2015

Work Availability

monday
tuesday
wednesday
thursday
friday
saturday
sunday
morning
afternoon
evening
swipe to browse

Work Preference

Work Type

Full Time

Location Preference

HybridRemoteOn-Site

Important To Me

Work-life balanceCompany CultureFlexible work hoursPersonal development programsCareer advancementTeam Building / Company RetreatsPaid sick leaveHealthcare benefits401k matchWork from home optionStock Options / Equity / Profit SharingPaid time off4-day work week

Quote

The best way to make your dreams come true is to wake up.
Paul Valery

Software

Linux

Languages

English
Advanced (C1)
Hindi
Bilingual or Proficient (C2)

Interests

Listening Music

Travelling

Playing Badminton

Harshit JainSite Reliability Engineer