Summary
Overview
Work History
Education
Skills
Timeline
Generic
POOJA M MALWADE

POOJA M MALWADE

Pimpri-Chinchwad

Summary

Data Engineer with 4.8 years of experience at EXL, experienced in designing, developing, and managing robust data pipelines and architectures using Python, PySpark, Azure Data Factory, Databricks, SQL, Scala, and Apache Spark, Snowflake . Adept at ETL processes, data modeling, and data warehousing solutions, with hands-on expertise across AWS, Azure, and Google Cloud. Skilled in optimizing data workflows for performance, scalability, and reliability while ensuring data security and compliance. Strong problem-solving abilities with a focus on delivering data solutions that support business intelligence and analytics.

Overview

5
5
years of professional experience

Work History

Data Engineer

EXL
Pune
10.2021 - Current

Domain - Banking Confidential (Product-based company – User Analytics Platform)

Project: Event Data Migration & Standardization Platform — AWS to Snowflake Analytics Pipeline : Digital Analytics / Product Analytics Designed and implemented a scalable event-driven data platform to migrate from a legacy third-party event tracking system to an in-house AWS and Snowflake-based solution. The platform enables real-time and batch ingestion of user activity data, standardizes event schemas, ensures data quality, and provides a centralized analytics layer for BI tools like Power BI and Looker.

  • Designed and developed end-to-end ETL pipelines using AWS services and Snowflake.
  • Implemented API-based data ingestion using AWS Lambda.
  • Built transformation pipelines using AWS Glue (PySpark).
  • Standardized multiple event schemas into a unified data model.
  • Implemented data validation frameworks (schema, duplicates, null checks, volume checks).
  • Designed and implemented incremental and backfill loading strategies.
  • Configured Snowpipe for automated data ingestion into Snowflake.
  • Orchestrated workflows using AWS Step Functions.
  • Optimized Glue jobs using partitioning and parallel processing.
  • Implemented metadata-driven pipeline configuration.
  • Monitored pipelines using CloudWatch logs and handled failures.
  • Collaborated with analytics teams to enable reporting and dashboards. Tech Stack: AWS (S3, Lambda, Glue, Step Functions, EventBridge), Snowflake, Snowpipe, Python, PySpark, SQL, Apache Airflow, ETL Design, Data Modeling, Data

Domain-Confidential (Research & Advisory Firm)

Project: Retention Analysis Platform — Automation, Governance & Analytics Domain: Retention Analytics / Customer Lifecycle Management (LCCM) Designed and implemented a scalable Retention Analytics Platform to help the organization proactively identify customer churn risks and improve client retention. The platform consolidates data from multiple enterprise systems, processes large scale engagement data, and generates analytics-ready datasets and engagement metrics to support business decision making and machine learning workflows.

  • Designed and developed end-to-end data pipelines using ADF and Databricks.
  • Implemented data ingestion pipelines from APIs, databases, and file systems.
  • Built scalable transformation logic using PySpark in Databricks.
  • Implemented Medallion Architecture (Bronze, Silver, Gold layers).
  • Performed data cleansing, transformation, and MDM-based identity mapping.
  • Developed engagement metrics and retention indicators.
  • Implemented incremental data loading strategies.
  • Optimized Spark jobs for performance and scalability.
  • Ensured data quality, governance, and security compliance.
  • Supported machine learning workflows with curated datasets.
  • Enabled Power BI dashboards for business insights.
  • Monitored pipelines and handled failures. Tech Stack: Microsoft Azure, Azure Data Factory (ADF), Azure Databricks, Apache Spark, Python, PySpark, ADLS Gen2, Delta Lake, Azure Synapse Analytics, Databricks SQL, Power BI, Git, Jenkins, JIRA, Unity Catalog, Azure Key Vault
  • Developed data pipelines that enabled timely analytics and reporting for informed decision-making.

Education

B.E. ENTC - Engineering

Pune University
Pune
06-2018

Skills

  • Programming languages: Python, SQL, Java, Scala
  • Big data technologies: Hadoop, Spark, Flink, Kafka, Hive
  • Cloud platforms - AWS: S3, Glue, Redshift, Lambda, EC2, SageMaker, Athena, Kinesis
  • Cloud platforms - Azure: Data Factory, Databricks, Synapse, Blob Storage
  • Cloud platforms - GCP: BigQuery, Cloud Storage
  • Databases: MySQL, PostgreSQL, Oracle, MongoDB
  • Data engineering: ETL processes, data migration
  • Data modeling: Star schema, snowflake schema
  • Workflow orchestration: Airflow
  • Data governance and security: GDPR compliance
  • DevOps and CI/CD: Git, Jenkins
  • Visualization and BI: Power BI, Tableau
  • Project management
  • Problem solving
  • Communication

Timeline

Data Engineer

EXL
10.2021 - Current

B.E. ENTC - Engineering

Pune University
POOJA M MALWADE