Summary
Overview
Work History
Education
Skills
Certification
Timeline
Generic

KAMRAN KHAN

BANGALORE

Summary

Results-driven Azure Data Engineer with 3.6 years of experience in designing, developing, and optimizing scalable data pipelines and ETL workflows using Azure Data Factory, Azure Databricks, PySpark, and Azure Data Lake Storage. Seeking new opportunities to leverage my skills and contribute to impactful projects as an Azure Data Engineer.

Experienced in developing end-to-end data solutions, including data extraction, transformation, aggregation, and loading from multiple data sources and file formats such as CSV, JSON, and XML. Proficient in using PySpark and Spark in Databricks for large-scale data processing, transformation and analytics. Experienced in designing and implementing data-driven workflows using Azure Data Factory, including pipeline creation, scheduling triggers and monitoring workflow maintenance. Strong understanding of data warehousing concepts, ETL processes, batch and stream processing, and cloud-based data lake solutions. Experienced in collaborating with clients and business stakeholders to understand requirements and translate them into scalable technical solutions.

Overview

1
1
Certification
4
4
years of professional experience

Work History

Data Engineer

ADUNHILL TECHNOLOGY PRIVATE LIMITED
Bangalore
02.2023 - Current
  • Designed and implemented scalable data engineering solutions for flight, customer, and logistics data using Azure Databricks.
  • Participated in requirement gathering, requirement analysis, solution design, development, testing, and deployment of data engineering workflows.
  • Developed end-to-end ETL pipelines using Azure Data Factory to ingest data from multiple sources and load it into Azure Data Lake Storage Gen2.
  • Created and configured ADF triggers and scheduling mechanisms to automate daily and periodic data processing workflows.
  • Used PySpark in Azure Databricks for data extraction, cleansing, transformation, filtering, and aggregation across multiple datasets.
  • Processed structured and semi-structured data in formats such as CSV, JSON, and Parquet.
  • Implemented data transformation and processing workflows using Spark DataFrames and Spark SQL.
  • Organized data using a Medallion Architecture (Bronze, Silver, and Gold layers) to separate raw, cleansed, and business-ready datasets.
  • Implemented Delta Lake tables for reliable data storage, schema management, and efficient data processing.
  • Developed reusable Databricks notebooks for data cleansing, transformation, validation, and aggregation.
  • Built efficient batch-processing workflows using Azure Databricks and Apache Spark to handle large volumes of data.
  • Performed data quality checks including duplicate detection, null-value handling, data type validation, and source-to-target reconciliation.
  • Monitored Azure Data Factory pipelines and Databricks jobs, investigated failures, and resolved data processing issues.
  • Collaborated with business stakeholders and clients to understand requirements and translate them into scalable technical solutions.
  • Maintained technical documentation for data pipelines, transformation logic, workflows, and data processing processes.

Data Engineer

ADUNHILL TECHNOLOGY PRIVATE LIMITED
Bangalore
02.2023 - Current
  • Designed and implemented scalable ETL pipelines using Azure Data Factory and Azure Databricks for processing taxi-related datasets.
  • Developed end-to-end ETL workflows using Azure Data Factory to extract data from multiple sources and load processed data into the target data warehouse.
  • Designed pipelines to process data from multiple file formats, including CSV, JSON, and XML.
  • Implemented data ingestion and transformation workflows to efficiently process large volumes of structured and semi-structured data.
  • Used Azure Databricks and PySpark for data cleansing, transformation, aggregation, and processing.
  • Designed and maintained data warehouse workflows to store processed data and support downstream reporting and analytics.
  • Developed and maintained Azure Data Factory pipelines, workflows, activities, and triggers for automated data processing.
  • Monitored pipeline executions and performed troubleshooting to identify and resolve data processing and workflow failures.
  • Implemented data protection procedures and validation mechanisms to maintain data integrity and minimize the risk of data loss.
  • Applied data masking and unmasking techniques to protect sensitive information while maintaining usability for authorized processing.
  • Performed data validation and quality checks to ensure consistency and accuracy between source and target datasets.
  • Collaborated with team members to understand business requirements and implement appropriate data processing solutions.

Education

Bachelor of Business Administration -

MGM College
01-2022

Skills

  • Python
  • SQL
  • Unix
  • Databricks
  • Data Factory
  • HDFS
  • Hive
  • Spark
  • Map Reduce
  • Oracle
  • MySQL
  • HBase
  • Azure Data Lake
  • Power BI

Certification

  • Databricks Certified Data Engineer Associate, https://credentials.databricks.com/400c4d95-3d79-496e-a45e-49dc1c263451#acc.oB14tQSb
  • Microsoft certified Fabric Data Engineer Associate, https://learn.microsoft.com/en-us/credentials/certifications/fabric-data-engineer-associate/?practice-assessment-type=certification

Timeline

Data Engineer

ADUNHILL TECHNOLOGY PRIVATE LIMITED
02.2023 - Current

Data Engineer

ADUNHILL TECHNOLOGY PRIVATE LIMITED
02.2023 - Current

Bachelor of Business Administration -

MGM College
KAMRAN KHAN