Senior Data Engineer with 7+ years of experience designing, developing, and optimizing enterprise-scale data engineering solutions across cloud and big data platforms.
Expertise in AWS, PySpark, Apache Airflow, Snowflake, SQL, and ETL/ELT pipeline development.
Experienced in modernizing legacy data platforms through Teradata-to-Snowflake migrations, implementing automated data quality frameworks, and building scalable data pipelines for analytics and reporting.
Proven ability to optimize performance, automate workflows, and collaborate with cross-functional teams to deliver scalable, high-quality data solutions.
Overview
1
1
Certification
8
8
years of professional experience
Work History
Senior Data Engineer
IBM India
Hyderabad
01.2023 - Current
Enterprise Cloud Data Platform Modernization
Designed and developed scalable ETL pipelines using Apache Airflow (MWAA) for workflow orchestration.
Built PySpark transformation jobs using AWS Glue to process large-scale datasets stored in Amazon S3.
Developed AWS Lambda functions to automate file movement and trigger Glue jobs using S3 events.
Configured AWS AppFlow to integrate Salesforce and SAP data into Amazon S3.
Implemented SodaCL-based data quality validations covering schema, null, duplicate, primary key, uniqueness, volume, and referential integrity checks.
Leveraged Glue Catalog and Crawlers for metadata management and automated schema discovery.
Optimized Glue jobs using partitioning and compression techniques, improving performance and reducing storage costs.
Performed root cause analysis and performance tuning, improving pipeline efficiency by 25%.
Automated end-to-end ingestion, transformation, and reporting workflows, reducing manual effort by 40%.
Maintained GitHub repositories, supported CI/CD deployments, and prepared technical documentation.
Teradata to Snowflake Data Migration Project
IBM India
Hyderabad
01.2023 - Current
Developed Snowflake SQL scripts to create Permanent, Transient, Temporary, External, and Work tables based on business requirements.
Built SQL transformations and data extracts from Snowflake tables to support insurance dashboards and business metrics.
Performed data validation between Teradata and Snowflake by validating record counts, schema, data types, aggregates, and business rules to ensure migration accuracy.
Assisted in migrating Teradata workloads to Snowflake by implementing ELT logic, optimizing SQL queries, and supporting testing activities.
Collaborated with business analysts and reporting teams to validate data quality and deliver accurate reporting datasets.
Utilized GitHub, GitHub Copilot, Windsurf IDE, and Snowflake Cortex AI to improve SQL development, debugging, code optimization, and productivity.
Consultant
Capgemini India
Hyderabad
11.2021 - 01.2023
Enterprise Banking Data Platform
Designed ETL workflows using Oracle Data Integrator (ODI 12c) based on business requirements.
Developed PySpark applications for distributed data processing and transformation.
Created Hive external tables with partitioning, dynamic partitioning, and bucketing.
Processed CSV and Parquet files using Spark SQL and loaded curated datasets into Hive.
Developed PySpark scripts for Slowly Changing Dimensions (SCD), incremental loads, and Oracle-to-Hive transformations.
Built Unix wrapper scripts for automated data loading, merging, and secure SFTP transfers.
Performed SQL optimization, debugging, unit testing, and managed database objects including tables, views, stored procedures, and triggers.
Senior Systems Engineer
Infosys Ltd
11.2018 - 10.2021
AWS Data Engineering Platform
Developed scalable data pipelines for processing enterprise-scale datasets.
Designed Apache Airflow DAGs to automate ETL workflows.
Built Spark SQL transformations on data stored in Amazon S3.
Worked extensively with AWS services including EC2, EMR, Athena, MWAA, CloudWatch, RDS, and S3.
Developed Python scripts for automated data processing and loading into Amazon S3.
Collaborated with business teams to translate requirements into scalable data engineering solutions.
Supported production monitoring, troubleshooting, and issue resolution.
Systems Engineer Trainee
Infosys Ltd
07.2018 - 11.2018
Completed intensive training in Python, DBMS, Java, and Microsoft Business Intelligence (MSBI).
Developed enterprise reports using SSRS.
Integrated data from multiple sources using SSIS.
Created OLAP cubes using SSAS.
Recognized as a High Performer during training.
Education
Bachelor of Technology - Electronics & Communication Engineering
Vaagdevi Engineering College
Warangal
01.2018
Intermediate - PCM
SR Junior College
Warangal
SSC -
Unique High School
Warangal
Skills
Data Engineering & Architecture: Designing and developing scalable data pipelines, cloud-based data solutions, and data processing architectures
AWS Cloud Services: Hands-on experience with Amazon S3, AWS Glue, AppFlow, AirFlow (MWAA), Redshift, Athena, Step Functions, Lambda, CloudWatch, IAM, and Secrets Manager for data processing, orchestration, storage, analytics, and monitoring
Snowflake & Teradata Migration: Experience in migration, including SQL conversion, ETL/ELT migration, data validation, reconciliation, and performance optimization
Apache Spark & PySpark: For distributed data processing, transformation, optimization, and large-scale data workloads
ETL & ELT Pipeline Development: Developing and maintaining data pipelines using AWS Glue, PySpark, SQL, Airflow, and Snowflake
Data Warehousing & Data Lakes: Experience with Snowflake, Teradata, Amazon Redshift, S3-based data lakes, and enterprise data warehouse concepts
SQL & Performance Optimization: Strong expertise in SQL, complex queries, joins, CTEs, subqueries, window functions, query optimization, and performance tuning
Workflow Automation & CI/CD: Experience with Apache Airflow, AWS Step Functions, workflow orchestration, automation, and CI/CD pipelines
Python Programming & Scripting: Experience using Python for data processing, automation, scripting, validation, and data engineering utilities
Big Data Technologies: Experience with Apache Spark, PySpark, Hadoop, HDFS, and Hive for distributed data processing and large-scale data workloads
HDFS Management: Experience working with HDFS for distributed storage, file management, data movement, and large-scale data processing
Team Collaboration: Effective collaboration with cross-functional teams, developers, QA teams, and business stakeholders to deliver data engineering solutions
Problem Solving: Strong analytical and troubleshooting skills with the ability to resolve complex data, pipeline, and production issues
Communication: Strong verbal and written communication skills with experience in technical discussions, stakeholder coordination, and project updates
Adaptability & Ownership: Quick to learn new technologies and take ownership of tasks, resolve challenges, meet deadlines, and drive successful project delivery
Certification
Databricks Certified Associate Spark Developer
AWS Certified Cloud Practitioner
Databricks Lakehouse Fundamentals Accreditation
Microsoft Certified: Azure Fundamentals
AWS Cloud Technical Essentials (Coursera)
Infosys Certified Agile Developer
Timeline
Senior Data Engineer
IBM India
01.2023 - Current
Teradata to Snowflake Data Migration Project
IBM India
01.2023 - Current
Consultant
Capgemini India
11.2021 - 01.2023
Senior Systems Engineer
Infosys Ltd
11.2018 - 10.2021
Systems Engineer Trainee
Infosys Ltd
07.2018 - 11.2018
Bachelor of Technology - Electronics & Communication Engineering