Client: Mars Incorporated (Fortune 500)
Project: Global Consumer Data Platform
- Designed and developed scalable Azure Data Factory and PySpark ETL pipelines processing Freight-to-Freight (F2F), Freight-to-Customer (F2C), Vendor, Labour Cost, Consumption, and Multi-Market datasets supporting analytics across 10 global markets.
- Architected parameterized, metadata-driven ETL frameworks using Azure Data Factory activities such as Lookup, ForEach, Copy Activity, Stored Procedures, and dynamic datasets to improve pipeline scalability and reusability.
- Processed 500GB+ enterprise data daily by optimizing Spark transformations, SQL queries, partitioning strategies, and Delta (Incremental) Loading.
- Built enterprise data warehouse solutions using Unity Catalog, Fact & Dimension modeling, and SCD Type 2 methodologies for downstream analytics and reporting.
- Automated SharePoint and OneDrive data ingestion into ADLS Gen2 using Power Automate and Logic Apps, streamlining data flow and minimizing manual tasks.
- Refactored Pandas-based pipelines to PySpark, enhancing scalability and maintainability of data processing workflows.
- Upgraded legacy codebase from L0 to L3 maturity by implementing modular, reusable design patterns and coding standards.
Client : Procter & Gamble
Project: Consumer Forecast Planning & Trade (CFTP)
- Developed Azure Data Factory and PySpark ETL pipelines supporting Consumer Forecast Planning & Trade (CFTP) business processes.
- Built scalable data ingestion and transformation pipelines integrating enterprise datasets into ADLS Gen2 and Azure SQL for downstream reporting.