

Accomplished Senior Cloud Platform Engineer with over 14 years of IT experience, including 9 years focused on Red Hat OpenStack Platform (RHOSP), OpenShift, Kubernetes, and Ceph Storage. Expert in managing large-scale private cloud environments, executing RHOSP upgrades, and administering OpenShift and Kubernetes clusters. Proficient in implementing Ansible integrations and troubleshooting critical components, ensuring stability in high-demand telecom settings.
Outlined key duties and tasks for the role.
RHOSP Upgrades
· Upgraded RHOSP 16.1.6 to 16.2.5 and 16.2.5 to 17.1.2 across Cloud clusters without downtime.
· Performed minor upgrade from RHOSP 17.1.x to 17.1.y in production environments with zero downtime.
Private Cloud Deployments
· Implemented multiple RHOSP 16.x and 17.x platforms on DL 380 Gen10 and BL 460c G10 bare metal nodes.
• Designed and deployed Ceph storage clusters using HDD, SSD, and NVMe disks.
• Configured and managed OpenStack services such as Keystone, Glance, Nova, Neutron, Horizon, Cinder, and Swift.
• Performed manual testing, applied patches, and handed over setups to clients.
Scaling and Future Planning
• Planned for scaling physical compute and storage resources across multiple data centers.
• Added compute nodes to existing OpenStack clouds and configured Load Balancer as a Service (LBaaS).
Support and Mentorship
• Provided on-call support during Sev1/2 issues and mentored a team of 10 junior cloud engineers.
• Created SOPs for team reference and ensured SLA compliance for infrastructure-related issues.
Managed a variety of major responsibilities related to the administration and maintenance of cloud services.
Administer Red Hat OpenStack Platform (RHOSP) to maintain highly available enterprise cloud services.
Execute Multiple Major & Minor RHOSP upgrade activities, including planning, validation, implementation, and post-upgrade health checks.
Participated in TripleO-based Undercloud and Overcloud deployment and upgrade activities.
Troubleshoot Nova compute services, VM provisioning issues, scheduler failures, and compute node problems.
Resolve Cinder volume provisioning, attachment, and backend storage issues.
Support Glance image management, uploads, and image-related troubleshooting.
Troubleshoot Horizon Dashboard authentication and accessibility issues.
Perform compute live migration and compute evacuation during planned maintenance.
Manage Availability Zones and support compute node onboarding and decommissioning.
Performed SSL Certificate Renewal Activity Across the JAWS Clouds Environment.
Administer OpenShift 4.10 production clusters using the oc CLI.
Manage Kubernetes workloads using kubectl, including namespaces, deployments, pods, and storage resources.
Administer Red Hat Ceph Storage clusters (Pacific and Reef), with previous experience on Nautilus and Octopus.
Monitor Ceph cluster health, OSD lifecycle, MON quorum, and PG recovery.
Troubleshoot OVN networking issues affecting tenant connectivity.
Investigate Pacemaker and RabbitMQ issues impacting controller services.
Execute enterprise Ansible playbooks for operational automation and develop simple playbooks where appropriate.
Perform infrastructure health checks before and after production changes.
Participate in 24×7 production support, incident resolution, and Root Cause Analysis (RCA).
Coordinate with Red Hat support engineers during complex production issues.
Develop SOPs and mentor junior cloud engineers.