Summary
Overview
Work History
Education
Skills
Timeline
Generic

Saran Teja Mallela

Texas

Summary

Healthcare Data Engineer with 3+ years of experience building Azure-native data platforms for clinical, oncology, and EMR/EHR data. Specialized in Azure Data Factory, Azure Databricks, Delta Lake, and HIPAA-compliant data pipelines supporting Real-World Data (RWD) and Real-World Evidence (RWE) analytics. Highly competent Data Engineer with background in designing, testing, and maintaining data management systems. Possess strong skills in database design and data mining, coupled with adeptness at using machine learning to improve business decision making. Previous work resulted in optimizing data retrieval processes and improving system efficiency.

Overview

5
5
years of professional experience

Work History

Data Engineer (Healthcare | Azure)

McKesson
Texas
08.2024 - Current
  • Built Azure-native ingestion pipelines using Azure Data Factory and Databricks to process 150M+ unstructured oncology documents including clinical notes, pathology reports, and lab summaries.
  • Transformed EMR/EHR data (Epic, Cerner) into standardized Clinical Data Models (CDM) enabling downstream RWD/RWE analytics.
  • Implemented Delta Lake Bronze/Silver/Gold architecture on ADLS Gen2 to ensure data lineage, auditability, and reproducibility.
  • Optimized PySpark-based parsing and clinical text normalization workflows, improving processing performance by 40%.
  • Enforced HIPAA-compliant security controls using Azure Key Vault, RBAC, and encrypted storage for PHI datasets.
  • Integrated Azure OpenAI NLP models for clinical entity extraction and oncology terminology mapping, improving data completeness and analytics readiness.
  • Built automated data quality checks (schema validation, clinical code verification) achieving 99%+ accuracy in structured outputs.

Data Engineer (Enterprise Data Platforms)

Accenture
India
05.2021 - 06.2023
  • Designed and built scalable batch and near-real-time data pipelines using Python, SQL, and Apache Spark to ingest large enterprise datasets from multiple source systems.
  • Developed and optimized ETL/ELT workflows supporting analytics, reporting, and downstream consumption across business and analytics teams.
  • Implemented data quality, validation, and reconciliation checks to ensure accuracy, completeness, and consistency of production datasets.
  • Applied data modeling and transformation best practices to deliver analytics-ready tables for BI tools and reporting platforms.
  • Worked in regulated enterprise environments, following security, access control, and audit requirements aligned with organizational governance standards.
  • Collaborated with cross-functional stakeholders to gather requirements, troubleshoot pipeline issues, and support production deployments.

Education

Master’s - Data Science

University of Houston
Houston, TX

Bachelor’s - Computer Science

Velagapudi Ramakrishna Siddhartha Engineering College
India

Skills

  • Azure Data Factory
  • Azure Databricks
  • ADLS Gen2
  • Azure Synapse
  • Python
  • PySpark
  • SQL
  • Delta Lake
  • Lakehouse
  • Data modeling
  • Databricks
  • Data governance
  • Cross-functional collaboration
  • Problem-solving
  • Clinical Data Models (CDM)
  • HIPAA
  • PHI
  • EMR/EHR (Epic, Cerner)
  • ICD-10
  • Schema validation
  • Data reconciliation
  • RBAC
  • Azure Key Vault
  • EXCEL
  • POWERBI
  • ETL development

Timeline

Data Engineer (Healthcare | Azure)

McKesson
08.2024 - Current

Data Engineer (Enterprise Data Platforms)

Accenture
05.2021 - 06.2023

Master’s - Data Science

University of Houston

Bachelor’s - Computer Science

Velagapudi Ramakrishna Siddhartha Engineering College
Saran Teja Mallela