
4+ years of Experience in Data Engieering and with Python, AWS, Data Migrations, Data bricks, Pyspark, Data Validataion, Data Governance, Django Developement, Data Modeling, ETL, SQL,API Developement, Testing using Docker, CICD, Jenkins, Pyspark, DataBricks, Scala, Spark. SQL, Apache Hadoop, Apache Spark, AWS (S3, Glue, Redshift, Athena), Azure (Data Lake, SQL Database), Apache Kafka, Apache Airflow, Docker, Kubernetes, Terraform, ETL, ELT, Data Warehousing, Pandas, NumPy, Git, Jenkins, Tableau, Power BI, Snowflake, MongoDB, PostgreSQL, MySQL, Hadoop HDFS, Parquet, Avro, JSON, XML, Data Pipeline, Data Governance, BigQuery, DataBricks, REST APIs, Data Lake, Cloud Storage.
Project Name: Standard Insurance MSA
Objectives:
Designed and implemented Scala Spark–based ETL pipelines on Databricks for high-volume insurance claim data. Led data migration from legacy environments to new cloud platforms across Dev, QA, and Prod. Migrated large-scale Parquet datasets into Delta Lake tables ensuring ACID compliance.
Built and optimized ingestion pipelines using Azure Data Factory and Azure Blob Storage. Implemented data validation, reconciliation, and post-migration audit checks.
Optimized Spark jobs using partitioning, caching, and performance tuning. Collaborated with DevOps and QA teams using Azure DevOps, Git, and GitHub.
Ensured data governance, security, and compliance standards within the insurance domain.
Project Name: OneBridge Solution
Objectives:
Designed and implemented end-to-end ETL pipelines using Azure Data Factory, Logic Apps, and Blob Storage to automate data ingestion and preprocessing. Standardized and transformed multi-format data (CSV, Excel, PDF) for integration with downstream systems using Python and master tables. Built data models and enriched datasets for capacity planning, linking hierarchical client data with key metrics like CBM (cubic meters). Orchestrated data storage layers (Bronze, Silver, Gold) in Couchbase NoSQL for structured and processed data. Collaborated with machine learning teams to integrate prediction outputs into the pipelines for real-time decision-making
Project Name: ACA compliance
Responsibilities:
• Data Ingestion: Used AWS Lambda and SharePoint API for file processing.
• Data Storage: Managed data in S3 and RDS/Redshift.
• Data Processing: Built ETL pipelines with AWS Glue and Python/PySpark.
• Querying & Reporting: Optimized Athena queries and visualized data in QuickSight.
• Data Quality & Governance: Implemented quality checks for HIPAA/ACA compliance.
• Cost Optimization: Optimized S3 storage and used Glue/Athena for cost savings.
Python, SQL, data warehousing, ETL, NoSQL, PySpark, Pandas, Django, Flask, REST API, AWS, Azure, Git, Docker, Snowflake, and Databricks, AWS Glue,Apache Hadoop Data Mining Data Visualization Cloud Computing, DataBricks