Architected a metadata-driven ETL framework using AWS Glue, EMR, Lambda, Step Functions, S3, and Athena for scalable data processing.
Implemented Spark optimizations including partitioning, bucketing, caching, and broadcast joins to enhance performance and minimize shuffle operations.
Reduced ETL job execution time by 40% through Spark optimization techniques, efficient data partitioning, and parallel processing. Lowered AWS costs by 35% by choosing cost-effective compute resources, auto-terminating idle EMR clusters, and leveraging Glue's serverless capabilities.
Optimized Spark jobs on EMR through parallelized computations, avoidance of data skews, and utilization of Adaptive Query Execution (AQE) for dynamic optimization.
Optimized Spark jobs on EMR by parallelizing computations, avoiding data skews, and leveraging Adaptive Query Execution (AQE) for dynamic optimization.
Designed AWS Glue jobs for schema evolution and automated metadata updates in AWS Glue Data Catalog, ensuring data integrity and reducing manual intervention.
Data Engineer
Santarch Infotech LLP
05.2023 - 08.2025
Automated data quality checks and validation using PySpark and AWS Lambda, ensuring clean and reliable datasets for analytics.
Implemented S3 data lake best practices using Parquet and Snappy compression, partition pruning, and optimized storage layouts for Athena queries.
Leveraged AWS Lambda for event-driven processing, selecting Glue or EMR based on data volume and transformation complexity to enhance processing efficiency.
Optimized storage costs through implementation of lifecycle policies and partitioning strategies, improving data management.
Data Engineer
Capgemini
02.2022 - 04.2023
Supported marketing and financial reporting systems using Salesforce Datorama, enhancing data accessibility for stakeholders.
Developed and maintained Excel-based reporting solutions, enabling informed decision-making for business stakeholders.
Troubleshot data pipeline issues, ensuring timely report delivery and maintaining data accuracy.