

Data Engineer with ~3 years of experience in building scalable ETL/ELT Pipelines/DAGS and optimizing workflows using Google Cloud Platform. Skilled in automating workflows, driving efficiency, and aligning with business goals. Proven track record of collaborating with cross-functional teams to deliver high-quality data solutions.
Project: Product Content & Intelligence (Project use case)
1. Led the automation of product tagging and enrichment pipelines using Google Cloud Pub/Sub, BigQuery, Python, Airflow (Cloud Composer), Vertex AI and Gemini models reducing manual effort by over 60% and improving operational efficiency.
2. Improved product attribute prediction accuracy (Color, Shape, Material, Subject, Style, and Life Stage) by 30% through the integration of Vertex AI-powered ML inference pipelines, replacing traditional human-driven tagging processes.
3. Designed and implemented SQL-based data enrichment workflows and Python-driven integrations to publish enriched product metadata to ADS APIs, enabling near real-time updates of product tags on the customer-facing website.
4. Managed cloud infrastructure using Terraform for Airflow (Cloud Composer), Pub/Sub, BigQuery datasets, IAM permissions, and service accounts across development and production environments.
5. Developed standardized schemas, validation frameworks, and data quality checks across multiple Pub/Sub topics, ensuring consistent, scalable, and reliable data exchange between distributed systems.
6. Implemented comprehensive monitoring and observability using Datadog dashboards and alerts, tracking pipeline health, job failures, processing latency, and API success metrics.
7. Collaborated closely with Machine Learning, UFS, and cross-functional engineering teams to deliver production-grade data solutions, maintain schema governance, and create detailed operational documentation and runbooks.
Project: Supply Chain & Logistics Data Platform (Inprogress)
1. Assisting in building end-to-end ELT pipelines for warehouse and logistics data, leveraging Fivetran and Google Cloud Dataflow to ingest data from Mendix and load it into Snowflake and GCP raw data layers.
2. Developing and maintaining modular dbt transformation models to cleanse, standardize, validate, and enrich warehouse and logistics datasets, enabling reliable downstream analytics and reporting.
3. Collaborated with cross-functional teams to design scalable data ingestion and transformation workflows, ensuring data quality and consistency across the platform.
4. Currently contributing to the ORBIT Supply Allocation & Demand Forecasting initiative by working closely with Backend Engineers and Data Science teams to understand forecasting requirements and support the development of the E2E supply forecasting process.
5. Supporting the implementation of forecasting logic and data pipelines that feed downstream planning systems, including SAP IBP, enabling improved demand planning and supply allocation decisions.
Project: GCP Consultant for Cost Optimisation & Data Infrastructure (Pilgrim India)
1. Led end-to-end cloud cost optimization initiatives across BigQuery, GKE, Cloud Composer, and Fivetran workloads, identifying inefficiencies and implementing sustainable cost-reduction strategies.
2. Optimized BigQuery performance through query tuning, partitioning, clustering, workload analysis, and governance controls, improving resource utilization and reducing compute costs.
3. Right-sized Cloud Composer environments by analyzing scheduler, worker, and executor utilization patterns, resulting in improved performance and lower operational expenses.
4. Enhanced Fivetran deployment efficiency by optimizing Kubernetes pod resource allocation, securing external access through controlled ingress configurations, and improving connector reliability and throughput.
5. Developed cost monitoring dashboards, billing analytics reports, and budget guardrails to provide proactive visibility into cloud spending and prevent cost overruns.