Summary
Overview
Work History
Education
Skills
Certification
Key Highlights
Timeline
web
Subhayan Ghosh

Subhayan Ghosh

Senior Data Engineer
Bengaluru,KA

Summary

Data Engineering Lead with expertise in scaling cloud data platforms, real-time streaming, and enterprise Data Mesh architecture across Sales and After Sales business domains. Combines deep technical proficiency in PySpark, Databricks, AWS MSK, and Data Vault 2.0 with a strong business focus on infrastructure cost reduction (€220K+ saved throughout), developer velocity, and zero-trust security. Experienced team lead adept at driving Agile execution, establishing automated data-quality frameworks, and mentoring engineering teams to deliver high-availability data products.

Overview

8
8
years of professional experience
1
1
Certification

Work History

Senior Consultant, Data & Analytics

Mercedes-Benz Research and Development India Pvt. Ltd.
Bengaluru
03.2022 - Current
  • Replaced a legacy batch pipeline (4-hour latency) with real-time CDC ingestion using Debezium and AWS MSK (Kafka), achieving sub-minute source-to-Lakehouse latency for critical grounded-vehicle analytics.
  • Migrated 30TB+ of legacy Hive-partitioned tables to Apache Iceberg; implemented hidden partitioning, automated compaction, and snapshot expiration, achieving sub-second query planning and cutting annual compute spend.
  • Augmented streaming CDC pipeline with Grafana observability across MSK and EC2-hosted Debezium connectors, proactively flagging connector lag and ingestion bottlenecks before downstream impact.
  • Reduced pipeline runtimes by up to 40% through Graviton (ARM) cluster migration, join-skew optimization, and DAG simplification, sustaining >60% cluster utilization and delivering €100K in annual infrastructure savings.
  • Standardized platform IaC using Terraform, EKS/Helm, and GitHub Actions; migrated deployment pipelines from Jenkins/ADO to Databricks Asset Bundles (DABs) with OIDC authentication and Black Duck SAST/SCA security gates.
  • Re-engineered the One Touch Retail transaction pipeline (5–10GB/day JSON) using Databricks Lakeflow Pipelines and Auto Loader, eliminating directory-scan overhead and custom compaction to cut processing time by 75% and compute costs by 60%.
  • Designed a multi-domain Data Mesh architecture for 30+ Sales data sources; implemented Data Vault 2.0 at the Silver layer for audited historization and projected into Kimball Star Schemas at the Gold layer for high-performance BI datamarts.
  • Slashed Databricks monthly spend by 50% (€20K to €10K/month) through data model redesign, Delta Z-Ordering, broadcast joins, predicate pushdown, and transitioning from time-based schedules to event-driven triggers.
  • Engineered a configuration-driven PySpark ingestion and data quality framework with automated SCD Type 2 historization, cutting new-source onboarding time by 60%. Integrated a Databricks AI/BI Genie natural-language dashboard over Delta audit tables, reducing developer triage effort by 40%.
  • Centralized data governance by replacing workspace ACLs and IAM policies with Unity Catalog RBAC, cutting user access provisioning time from days to minutes.
  • Technical Leadership: Led Agile delivery for a 6-person engineering team, establishing automated testing standards, driving CI/CD adoption, and mentoring junior data engineers.

Senior Data Engineer

L&T Infotech
Bengaluru
08.2021 - 03.2022
  • Processed multi-format retail banking data (Parquet, CSV, JSON, mainframe) with Spark SQL and DataFrames.
  • Designed a PySpark framework automating data-quality validation, and a Python reporting framework for Autosys job monitoring.

Application Developer

IBM India Pvt. Ltd.
Bengaluru
05.2018 - 08.2021
  • Built and tuned Hive (HQL) and PySpark transformations for a finance risk-evaluation platform; authored SQL stored procedures and monitored Sqoop ingestion from DB2.
  • Developed Java and PL/SQL integration layers connecting third-party platforms to Oracle EBS; optimized SQL using explain plans and materialized views.

Education

B.Tech - Electrical Engineering

West Bengal University of Technology
01.2017

Skills

  • Databricks
  • Delta Lake
  • Apache Iceberg
  • Unity Catalog
  • Delta Live Tables
  • Lakeflow
  • AWS Glue
  • Azure Data Factory
  • Apache Kafka
  • AWS MSK
  • Debezium
  • Azure Event Hub
  • Python
  • PySpark
  • SQL
  • FastAPI
  • Java
  • AWS
  • Azure
  • Terraform
  • Kubernetes
  • Docker
  • GitHub Actions
  • Jenkins
  • Azure DevOps
  • Grafana
  • Data-quality validation frameworks
  • Data Vault 20
  • SCD Type 2
  • RBAC
  • Data lineage
  • PostgreSQL
  • Oracle
  • DB2
  • Apache Hive
  • MongoDB

Certification

  • Databricks Certified Data Engineer Associate — Databricks
  • Databricks Certified Associate Developer for Apache Spark 3.0 — Databricks
  • Data Streaming Engineer — Confluent

Key Highlights

30TB+ data volume, Sales Data Analytics, After Sales analytics., Databricks (Lakeflow, DABs, AI/BI Genie, Unity Catalog), AWS (MSK, S3, Glue, EKS), PySpark, Debezium, Apache Iceberg, Data Vault 2.0, Kimball., €220K+ total savings (€100K infrastructure + €120K Databricks spend reduction), 60% faster onboarding, sub-minute CDC latency., Agile team lead (5–6 engineers), GitHub Actions, Terraform, OIDC, Black Duck SCA/SAST.

Timeline

Senior Consultant, Data & Analytics

Mercedes-Benz Research and Development India Pvt. Ltd.
03.2022 - Current

Senior Data Engineer

L&T Infotech
08.2021 - 03.2022

Application Developer

IBM India Pvt. Ltd.
05.2018 - 08.2021

B.Tech - Electrical Engineering

West Bengal University of Technology
Subhayan GhoshSenior Data Engineer