
Versatile and experienced Data Engineer with over 12+ years of experience in data integration, processing, and warehousing, including 7+ years in cloud-based ecosystems. Specializes in building, scalable data pipelines using both Azure and AWS services, including Azure Data Factory, Azure data Bricks, AWS Glue, and EMR. Skilled in Pyspark, Spark SQL, and ETL development. Delivering performant and reliable solutions in financial and enterprise environments. Adept at enabling analytics through data lake and warehouse solutions with strong data quality, monitoring and orchestration capabilities.
• Architected a metadata-driven ELT pipeline on Databricks/Delta Lake using a custom Python framework, implementing a Data Vault 2.0 model across raw, curated, silver, and gold layers
• Built core Data Vault migration components (custom Hub/Link/Satellite classes) in PySpark, enabling accurate timestamp normalization and event-driven data loading
• Established YAML-based metadata configuration standards governing entity loading, surrogate key resolution, and datamart table definitions across the pipeline
• Resolved Delta Lake constraint violations by redesigning entity load ordering across hub, link, and satellite dependency chains, eliminating recurring pipeline failures
• Designed and implemented a surrogate key null-handling strategy (quarantine-and-drop pattern) to enforce referential integrity across fact and dimension tables, backed by automated unit tests
• Root-caused and resolved data integrity defects, including timestamp misalignment across silver-layer entities that was collapsing rows in historical join logic
• Managed end-to-end pipeline orchestration through Azure Data Factory and Azure DevOps, including parameter propagation across multi-stage pipeline chains and secrets management via Databricks
• Administered Unity Catalog access controls, diagnosing and resolving schema-level governance issues blocking data definition operations
• Maintained Git workflows across a multi-developer codebase, resolving complex rebase conflicts and branch synchronization issues
• Refactored legacy metadata conventions in response to peer code review, replacing hardcoded logic with typed constants and structured configuration formats
• Collaborated in code review cycles, iterating on YAML schema design and data normalization logic based on reviewer feedback
Azure (Data Factory, Data bricks)
AWS (Glue, EMR, Redshift, S3, Athena, Step Functions, Lambda, IAM, SNS)
PySpark
Spark SQL
Azure Data Factory
AWS Glue
Data Vault(Hub,Link and Sat)
Python
SQL
Shell Scripting
Azure Monitor
AWS CloudWatch
IAM
Key Vault
Power BI
Git
Linux
Windows