

Accomplished IT Operations Lead with over 13 years of experience in production and application support for enterprise applications across email marketing, pharma, and finance. Specializes in IT service management, incident resolution, and cloud operations, prioritizing business continuity and disaster recovery. Leads teams to drive continuous service improvements and collaborates with engineering, infrastructure, and security teams to optimize software solution availability and performance.
· Owned IT operations for multiple business-critical enterprise applications.
· Planned, tracked, and managed production support activities using Scrum meetings, ADO Boards, Service now ticket queues.
· Coordinate Major Incident (P1/P2) bridge calls, stakeholder communications, troubleshooting and Root Cause Analysis.
· Improve operational efficiency and reduce MTTR through automation, proactive monitoring, standardized runbooks, and process optimization.
· Drive operational governance through ITIL Incident, Problem, Change, and Release Management processes.
· Support internal audits, technology risk assessments, and operational compliance reviews.
· Produce KPI dashboards covering SLA compliance, incident trends, service availability, and operational performance.
· Coordinate Disaster Recovery planning, testing, and Business Continuity (PiTR) activities.
· Lead cross-functional teams’ collaboration across Infrastructure, Database, Cloud, Security, and Vendor organizations to understand the risk and plan for patching/maintenance, deployments, DR.
· Mentor L1/L2 engineers while driving knowledge sharing, operational readiness, and continuous capability improvement.
· Serve as the primary point of contact between business stakeholders, product teams, infrastructure teams, vendors, and support teams.
· Facilitate service reviews with business stakeholders, vendors, and engineering teams.
· Drive continuous service improvement initiatives through automation, monitoring enhancements, and operational process optimization.
· Recent projects that I led:
1. Created release lifecycle management document
2. Created a Disaster Recovery (DR) and Point in Time Recovery (PiTR) plan
3. Coordinated with different teams for automating
i) Manual application verification after each maintenance activity
ii) SSL cert. renewal for all env. in every 6 months
4. Created an “App Code Archive” using third party Escrow solutions.
5. Coordinated for app migration to azure and helped team to onboard
6. Created few knowledge articles for L1 team for few tasks handover
7. Took care of numerous process revisions to align with organizational standard
8. Restructure APM Configuration items (CIs) in ServiceNow
9. Helped team to set up automation testing env. for regression testing
10. Implemented Privileged Access Management using CyberArk
· Worked on production issues of some in-house ETL Tools like Javelin Incentive Manager (JIM), Javelin Data Manager (JDM), REVO Data Manager (RDM); to provide timely diagnostics, recommendations, workarounds and resolution.
· Worked on few other applications like Javelin Identity Management (IDM), REVO Analytics Workbench (RAW) as well. All REVO branded products are AWS based.
· Guiding product users through features and functionalities.
· Prima facie for some important clients in case of urgent issues and handles escalations.
· Supporting Development team with identifying issues and applying fixes, by debugging application at code level.
· Coordinating change/configure management.
· Analyzing commonly occurring issues and providing insights on required fixes.
· Strong knowledge in ETL, SQL queries and understanding of database.
· Pulling reports from DB or Splunk by creating ad-hoc queries.
· Running profiler and analysing SQL query performance.
· Working with Infra team (Server/Networking/DBA/Cloud etc.) in case of need.
· Helping Product management team for bug fixes and enhancements by explaining the client issues/requirements.
· Helping to resolve any outages by collaborating with various cross functional teams.
· Creating and maintaining the knowledge base or documents for our applications which would help resolve issues faster.
· Understanding various technologies that we work with, developing some level of expertise over a period of time.
· Monitoring Alerts related to Production servers (Windows/IIS) and application.
· Managing Services (Start, Stop and troubleshooting as per necessity) installed on IIS.
· Taking ownership of handling and resolving all the critical incidents.
· User management activities.
· Ensure correct and timely resolution of all incident/problems/client query tickets related to application within SLA.
· Troubleshooting issues and performing root cause analysis based on the logs and data output.
· Designing SQL Queries and preparing Reports as per ad-hoc request by client.
· Performing maintenance activities and application sanity testing.
· Performing day to day operations, administration and maintenance activities of WebLogic application servers.
· Building patches and deploying it to the target servers.
· Writing SQL queries based on ad-hoc request by user.
· Troubleshooting issues and performing root cause analysis based on the logs and data output.
· Migration of object between different servers.
· Maintaining the status for open and closed issues on a weekly basis.
· Monitoring ticket queues and escalate if required.
· Participating in internal and client meetings to discuss the status of the application.
· Opening bridge conference call based on high priority issues.
· Writing and executing shell scripts to automate manual tasks.