Summary
Overview
Work History
Education
Skills
Certification
Languages
Personal Information
Executive Profile Summary
Timeline
Generic

SAMI Mohammed

Bangalore

Summary

Results-driven Site Reliability Engineer specializing in performance monitoring and incident response. Skilled in Datadog and Ansible automation, enhancing service reliability and operational efficiency. Improved system uptime and incident resolution through effective collaboration with cross-functional teams while managing multiple providers to ensure timely project delivery.

Overview

1
1
Certification
8
8
years of professional experience

Work History

Site Reliability Engineer

UST
12.2025 - Current
  • Leveraged AIOps in production support to improve service reliability and reduce manual intervention time.
  • Utilized OpManager, Datadog, and Dynatrace to enhance infrastructure and application monitoring, improving visibility into system health.
  • Managed Linux-based servers and IIS web applications to ensure optimal performance and uptime.
  • Monitored key metrics including CPU, memory, disk availability, latency, and response time to meet SLA compliance.
  • Conducted incident analysis and root cause investigations to troubleshoot critical outages, minimizing downtime and improving response time.
  • Configured dashboards and threshold alerts to proactively identify incidents and reduce alert noise.
  • Collaborated with cross-functional teams to expedite issue resolution and enhance observability using AIOps.
  • Generated operational reports and dashboards to support service reviews and guide reliability improvement initiatives.
  • Designed dashboards, scorecards, metrics, and reports for business intelligence purposes.

Site Reliability Engineer

Nine Hertz India Pvt. Ltd.
Bangalore
06.2025 - 10.2025
  • Managed observability and production support operations for critical application services, ensuring stability across lower and production environments.
  • Enhanced service reliability through proactive monitoring, failover-aware support practices, and SLA-driven ticket resolution using ServiceNow.
  • Onboarded services into DataDog, enhancing monitoring coverage and visibility across key application and infrastructure components.
  • Built and maintained dashboards for SLO tracking, service performance, and error-budget burn rate, enabling proactive operational decision-making.
  • Configured new alerts and refined thresholds, optimizing monitoring effectiveness and reducing noisy alerts to improve signal-to-noise ratio.
  • Performed alert triage and incident troubleshooting, driving timely resolution of business-impacting issues in globally distributed systems.
  • Developed Ansible playbooks and supported PagerDuty workflows to streamline operations, while publishing stakeholder reports and tracking activities in Asana.
  • Coordinated bridge calls and managed change execution to reduce service disruption during high-priority incidents.
  • Client: Capgemini

Site Reliability Engineer

HCLTech Ltd.
Bangalore
06.2024 - 06.2025
  • Monitored SLA and SLI performance for payment services, providing operational insights and performance updates to enhance decision-making for business stakeholders.
  • Monitored end-to-end digital payment services to ensure high availability, reliability, and seamless transaction processing.
  • Developed and maintained DataDog dashboards, widgets, and monitors to deliver real-time visibility into application health, payment workflows, and service performance.
  • Tracked page-load performance, transaction success rates, and application health metrics to proactively identify service degradation and expedite issue resolution.
  • Defined and implemented alert thresholds and monitoring rules based on service behavior, operational requirements, and performance patterns.
  • Refined alerts and tuning rules to enhance monitoring effectiveness, reducing false positives and improving signal quality for faster response times.
  • Automated monitoring processes and drove continuous improvement initiatives using Ansible and Git for new alert configurations and observability enhancements.
  • Worked closely with engineering teams by raising and managing Jira tickets to escalate issues and drive timely remediation.
  • Client: BCBSRI

Application Production & Infrastructure Administrator

HCLTech Ltd.
08.2023 - 05.2024
  • Provided 24/7 production and infrastructure support for critical Linux and Windows banking platforms in a high-availability enterprise environment.
  • Monitored and responded to platform events, infrastructure alerts, and operational incidents to ensure service continuity and system stability.
  • Managed HP OMI event monitoring to proactively identify and resolve infrastructure and application issues, enhancing system reliability.
  • Administered AWS and Azure cloud services such as EC2, VPC, S3, Route53, CloudWatch, virtual machines, storage, and Kubernetes resources to support secure and reliable operations.
  • Coordinated weekend patching and managed production maintenance tasks, strengthening system security and optimizing performance.
  • Executed production deployments and release activities using BMC BladeLogic, ensuring controlled implementation of changes to maintain service continuity.
  • Managed SSL certificate lifecycle for internal and external customers, overseeing certificate creation, renewal, and deployment to ensure compliance and security.
  • Contributed to security and compliance initiatives by reporting vulnerabilities, monitoring SIEM, coordinating Microsoft patching, and configuring Akamai to strengthen security posture.
  • Client: Commonwealth Bank of Australia (CBA)

Infrastructure & Production Support Engineer

HCLTech Ltd.
08.2022 - 07.2023
  • Managed Linux and Windows production servers, ensuring stable operations through system administration, troubleshooting, and performance optimization.
  • Monitored and optimized system performance, including CPU, memory, disk utilization, and overall server health in production environments.
  • Managed backup operations and infrastructure incidents, ensuring business continuity and service availability through effective node-down recovery activities.
  • Monitored server logs and performed maintenance, supporting server restart and recovery using VMware vCenter and vSphere.
  • Performed LVM administration, package management, and server maintenance activities to support reliable infrastructure operations.
  • Executed incident, change, and major incident management activities, leading bridge calls to maintain service continuity with proactive monitoring using SiteScope, BPM, Dynatrace, and Nagios.
  • Provided front-line and second-level support for global finance applications and middleware platforms, including Apache, Tomcat, and IIS, ensuring optimal performance and reliability.
  • Client: FedEx

Application Monitoring Support Engineer

HCLTech Ltd.
01.2019 - 07.2022
  • Monitored service availability, system logs, scheduled jobs, and platform health across Unix/Linux and Windows environments to maintain operational continuity.
  • Provided L1/L2 production support for SAP PS and internal enterprise applications, ensuring timely resolutions for application incidents, infrastructure alarms, and user-reported issues to maintain service reliability.
  • Executed Linux administration tasks including user and group management, LVM operations, package installations via RPM/YUM, patching, and server maintenance, ensuring optimal system performance.
  • Collaborated with vendors and internal stakeholders during bridge calls, incident investigations, and change-management activities to facilitate timely service restoration.
  • Client: Digital Banking Service of India – Singapore

Education

Bachelor of Engineering - Information Technology

JNTU Anantapur
Anantapur, India

Skills

  • Datadog
  • Application monitoring
  • Performance monitoring
  • AWS cloud services
  • CloudWatch monitoring
  • Kubernetes orchestration
  • Docker containers
  • Ansible automation
  • Jenkins CI/CD pipelines
  • Git
  • Grafana
  • Incident response
  • Log management
  • ITIL practices
  • Linux/Unix systems
  • Cloud security practices
  • ServiceNow ITSM tools
  • Jira project tracking
  • Bitbucket

Certification

  • Microsoft Azure Fundamentals (AZ-900)
  • AWS Certified Cloud Practitioner

Languages

  • English
  • Hindi
  • Telugu

Personal Information

  • Date of Birth: 10/03/90
  • Gender: Male
  • Marital Status: Married

Executive Profile Summary

Accomplished Site Reliability Engineer with 7.6 years of experience delivering resilient, high-availability platforms across observability, production support, infrastructure operations, and enterprise application monitoring. Hands-on expertise in AIOps-driven monitoring, DataDog administration, alert engineering, SLO/SLI/SLA governance, incident response, change management, and post-incident analysis within 24/7 mission-critical environments. Proven ability to improve operational stability through alert noise reduction, dashboard standardization, service onboarding, automation with Ansible and Bash, and cross-functional collaboration with development, infrastructure, and business teams. Strong technical foundation in Linux/Unix, Windows, AWS, Azure, Kubernetes, VMware, and enterprise monitoring stacks including Grafana, Kibana, Splunk, Dynatrace, AppDynamics, Zabbix, SiteScope, OMI, and PagerDuty.

Timeline

Site Reliability Engineer

UST
12.2025 - Current

Site Reliability Engineer

Nine Hertz India Pvt. Ltd.
06.2025 - 10.2025

Site Reliability Engineer

HCLTech Ltd.
06.2024 - 06.2025

Application Production & Infrastructure Administrator

HCLTech Ltd.
08.2023 - 05.2024

Infrastructure & Production Support Engineer

HCLTech Ltd.
08.2022 - 07.2023

Application Monitoring Support Engineer

HCLTech Ltd.
01.2019 - 07.2022

Bachelor of Engineering - Information Technology

JNTU Anantapur
SAMI Mohammed