Summary
Overview
Work History
Education
Skills
Certification
Timeline
Generic

Raviraj Navada

Bangalore

Summary

Performance Benchmark Engineer with 9+ years of total experience across data center infrastructure, AI, GenAI, HPC systems and virtualization platforms. Currently focused on AI and LLM performance benchmarking, platform sizing, and optimization using NVIDIA NIM, GenAI Perf, MLPerf, and HPC benchmarks. Strong expertise in GPU acceleration, CPU/GPU performance tuning, and analyzing system-level metrics including latency, throughput, power, and thermal behavior to deliver optimized and scalable infrastructure designs.

Overview

10
10
years of professional experience
1
1
Certification

Work History

TME - Performance Benchmark Engineer

Cisco Systems
Bangalore
11.2024 - Current
  • Designed, executed, and optimized AI, GenAI, and HPC performance benchmarks for Cisco UCS platforms, covering training, inference, and compute-intensive workloads.
  • Executed MLPerf Training & Inference benchmarks across vision, NLP, recommendation systems, and LLM workloads.
  • End-to-end platform validation from P0 silicon/board bring-up to production-ready motherboards, collaborating closely with engineering teams and performing thermal, power, and stability validation on server hardware.
  • Test and validate the BIOS tunings parameters specific to Cisco UCS hardware supporting multiple workload (Compute intensive, AI/ML and HPC, Database etc.)
  • Validate the platfotm
  • Currently developing and validating an AI Sizer tool to size Cisco UCS platforms for large-scale LLM inferencing using NVIDIA NIM and GenAI / AI Perf tools.
  • Analyzed GenAI and LLM performance metrics including Time To First Token (TTFT), request latency, end-to-end latency, tokens-per-second throughput, concurrency, and GPU utilization.
  • Performed HPC benchmarking on Cisco UCS platforms using HPL, HPL-STREAM, and STREAM benchmarks with NVIDIA B300 GPU-based systems.
  • Benchmarked NVIDIA GPU platforms (L40S, H100, H200 SXM, NVL, B300) to evaluate compute efficiency, memory bandwidth, and multi-GPU scaling using NCCL.
  • Gained hands-on exposure to AMD MI320 and MI350 GPU platforms, and Intel CPU architectures, focusing on performance characteristics and optimization strategies.
  • Executed CPU-specific benchmarks using SPEC CPU to evaluate processor performance across different workloads and configurations.
  • Monitored, analyzed, and stress-tested power and thermal behavior of Intel processors using Intel PTAT tools to ensure performance stability and thermal compliance.
  • Optimized AI, GenAI, and HPC workloads through CUDA tuning, GPU affinity, NUMA alignment, batch sizing, and concurrency optimization.
  • Validated and optimized storage I/O paths using NVMe and GPUDirect Storage to reduce latency bottlenecks.
  • Tuned network fabrics including RDMA and RoCE to improve distributed training, inference, and HPC workload scalability.
  • Built custom benchmark suites aligned with real-world customer workloads across AI, GenAI, and HPC domains.
  • Analyzed benchmark results to identify performance gaps and delivered actionable sizing, tuning, and optimization recommendations.
  • Developed technical assets, including AI/HPC performance white papers, benchmark methodologies, reference architectures, and executive-level presentations.
  • Collaborated with engineering, product management, sales, and pre-sales teams to align benchmarks with product roadmaps and customer requirements.
  • Worked closely with ecosystem partners including NVIDIA, AMD, and CPU vendors to validate platform compatibility and competitive performance positioning.

Site Reliability Engineer

Deloitte
06.2024 - 11.2024
  • VMware Cloud Foundation lifecycle management, automation, upgrades, reliability engineering, and hybrid cloud operations.
  • Own and provide proactive support by participating in customer upgrades
  • Assess and analyze VCF components before, during and after upgrade activities

UCS Server Virtualization TAC Engineer

Cisco Systems
02.2018 - 06.2024
  • Deep troubleshooting of Cisco UCS compute platforms, Fabric Interconnects, and Intersight-managed environments.
  • Handling break fix cases from customer on UCS product
  • Customer / Escalation management

L2 Network Engineer

Microland Ltd
06.2016 - 01.2018
  • F5 VPN operations, global incident management, network troubleshooting, and junior engineer mentoring.

Education

Bachelor of Engineering - Electronics & Communication

Alva’s Institute of Engineering and Technology
Mangalore

Skills

  • AI performance benchmarking - MLPerf, Gen AI Perf
  • LLM Inferencing and training
  • SPEC CPU, FIO, VDBENCH
  • GPU performance tuning
  • Networking - RDMA, RoCE, InfiniBand, NCCL, Cisco networking
  • Server Virtualization - UCS, VMware
  • OS: Linux, Windows
  • Hardware platform testing
  • Technical Troubleshooting
  • Performance Analysis
  • Technical documentation - CVD's, white papers

Certification

NVIDIA AI infrastructure and operations fundamentals by Coursera
VMware VCP – Data Center Virtualization
Cisco Certified Network Associate (CCNA).

Timeline

TME - Performance Benchmark Engineer

Cisco Systems
11.2024 - Current

Site Reliability Engineer

Deloitte
06.2024 - 11.2024

UCS Server Virtualization TAC Engineer

Cisco Systems
02.2018 - 06.2024

L2 Network Engineer

Microland Ltd
06.2016 - 01.2018

Bachelor of Engineering - Electronics & Communication

Alva’s Institute of Engineering and Technology
Raviraj Navada