Summary
Overview
Work History
Education
Skills
Certification
Timeline
Generic

Keerthana Nibhanapuri

Hyderabad

Summary

  • Senior Data Engineer with 7+ years of experience designing, developing, and optimizing enterprise-scale data engineering solutions across cloud and big data platforms.
  • Expertise in AWS, PySpark, Apache Airflow, Snowflake, SQL, and ETL/ELT pipeline development.
  • Experienced in modernizing legacy data platforms through Teradata-to-Snowflake migrations, implementing automated data quality frameworks, and building scalable data pipelines for analytics and reporting.
  • Proven ability to optimize performance, automate workflows, and collaborate with cross-functional teams to deliver scalable, high-quality data solutions.

Overview

1
1
Certification
8
8
years of professional experience

Work History

Senior Data Engineer

IBM India
Hyderabad
01.2023 - Current

Enterprise Cloud Data Platform Modernization

  • Designed and developed scalable ETL pipelines using Apache Airflow (MWAA) for workflow orchestration.
  • Built PySpark transformation jobs using AWS Glue to process large-scale datasets stored in Amazon S3.
  • Developed AWS Lambda functions to automate file movement and trigger Glue jobs using S3 events.
  • Configured AWS AppFlow to integrate Salesforce and SAP data into Amazon S3.
  • Implemented SodaCL-based data quality validations covering schema, null, duplicate, primary key, uniqueness, volume, and referential integrity checks.
  • Leveraged Glue Catalog and Crawlers for metadata management and automated schema discovery.
  • Optimized Glue jobs using partitioning and compression techniques, improving performance and reducing storage costs.
  • Performed root cause analysis and performance tuning, improving pipeline efficiency by 25%.
  • Automated end-to-end ingestion, transformation, and reporting workflows, reducing manual effort by 40%.
  • Maintained GitHub repositories, supported CI/CD deployments, and prepared technical documentation.

Teradata to Snowflake Data Migration Project

IBM India
Hyderabad
01.2023 - Current
  • Developed Snowflake SQL scripts to create Permanent, Transient, Temporary, External, and Work tables based on business requirements.
  • Built SQL transformations and data extracts from Snowflake tables to support insurance dashboards and business metrics.
  • Performed data validation between Teradata and Snowflake by validating record counts, schema, data types, aggregates, and business rules to ensure migration accuracy.
  • Assisted in migrating Teradata workloads to Snowflake by implementing ELT logic, optimizing SQL queries, and supporting testing activities.
  • Collaborated with business analysts and reporting teams to validate data quality and deliver accurate reporting datasets.
  • Utilized GitHub, GitHub Copilot, Windsurf IDE, and Snowflake Cortex AI to improve SQL development, debugging, code optimization, and productivity.

Consultant

Capgemini India
Hyderabad
11.2021 - 01.2023

Enterprise Banking Data Platform

  • Designed ETL workflows using Oracle Data Integrator (ODI 12c) based on business requirements.
  • Developed PySpark applications for distributed data processing and transformation.
  • Created Hive external tables with partitioning, dynamic partitioning, and bucketing.
  • Processed CSV and Parquet files using Spark SQL and loaded curated datasets into Hive.
  • Developed PySpark scripts for Slowly Changing Dimensions (SCD), incremental loads, and Oracle-to-Hive transformations.
  • Built Unix wrapper scripts for automated data loading, merging, and secure SFTP transfers.
  • Performed SQL optimization, debugging, unit testing, and managed database objects including tables, views, stored procedures, and triggers.

Senior Systems Engineer

Infosys Ltd
11.2018 - 10.2021

AWS Data Engineering Platform

  • Developed scalable data pipelines for processing enterprise-scale datasets.
  • Designed Apache Airflow DAGs to automate ETL workflows.
  • Built Spark SQL transformations on data stored in Amazon S3.
  • Worked extensively with AWS services including EC2, EMR, Athena, MWAA, CloudWatch, RDS, and S3.
  • Developed Python scripts for automated data processing and loading into Amazon S3.
  • Collaborated with business teams to translate requirements into scalable data engineering solutions.
  • Supported production monitoring, troubleshooting, and issue resolution.

Systems Engineer Trainee

Infosys Ltd
07.2018 - 11.2018
  • Completed intensive training in Python, DBMS, Java, and Microsoft Business Intelligence (MSBI).
  • Developed enterprise reports using SSRS.
  • Integrated data from multiple sources using SSIS.
  • Created OLAP cubes using SSAS.
  • Recognized as a High Performer during training.

Education

Bachelor of Technology - Electronics & Communication Engineering

Vaagdevi Engineering College
Warangal
01.2018

Intermediate - PCM

SR Junior College
Warangal

SSC -

Unique High School
Warangal

Skills

  • Data Engineering & Architecture: Designing and developing scalable data pipelines, cloud-based data solutions, and data processing architectures
  • AWS Cloud Services: Hands-on experience with Amazon S3, AWS Glue, AppFlow, AirFlow (MWAA), Redshift, Athena, Step Functions, Lambda, CloudWatch, IAM, and Secrets Manager for data processing, orchestration, storage, analytics, and monitoring
  • Snowflake & Teradata Migration: Experience in migration, including SQL conversion, ETL/ELT migration, data validation, reconciliation, and performance optimization
  • Apache Spark & PySpark: For distributed data processing, transformation, optimization, and large-scale data workloads
  • ETL & ELT Pipeline Development: Developing and maintaining data pipelines using AWS Glue, PySpark, SQL, Airflow, and Snowflake
  • Data Warehousing & Data Lakes: Experience with Snowflake, Teradata, Amazon Redshift, S3-based data lakes, and enterprise data warehouse concepts
  • SQL & Performance Optimization: Strong expertise in SQL, complex queries, joins, CTEs, subqueries, window functions, query optimization, and performance tuning
  • Workflow Automation & CI/CD: Experience with Apache Airflow, AWS Step Functions, workflow orchestration, automation, and CI/CD pipelines
  • Python Programming & Scripting: Experience using Python for data processing, automation, scripting, validation, and data engineering utilities
  • Big Data Technologies: Experience with Apache Spark, PySpark, Hadoop, HDFS, and Hive for distributed data processing and large-scale data workloads
  • HDFS Management: Experience working with HDFS for distributed storage, file management, data movement, and large-scale data processing
  • Team Collaboration: Effective collaboration with cross-functional teams, developers, QA teams, and business stakeholders to deliver data engineering solutions
  • Problem Solving: Strong analytical and troubleshooting skills with the ability to resolve complex data, pipeline, and production issues
  • Communication: Strong verbal and written communication skills with experience in technical discussions, stakeholder coordination, and project updates
  • Adaptability & Ownership: Quick to learn new technologies and take ownership of tasks, resolve challenges, meet deadlines, and drive successful project delivery

Certification

  • Databricks Certified Associate Spark Developer
  • AWS Certified Cloud Practitioner
  • Databricks Lakehouse Fundamentals Accreditation
  • Microsoft Certified: Azure Fundamentals
  • AWS Cloud Technical Essentials (Coursera)
  • Infosys Certified Agile Developer

Timeline

Senior Data Engineer

IBM India
01.2023 - Current

Teradata to Snowflake Data Migration Project

IBM India
01.2023 - Current

Consultant

Capgemini India
11.2021 - 01.2023

Senior Systems Engineer

Infosys Ltd
11.2018 - 10.2021

Systems Engineer Trainee

Infosys Ltd
07.2018 - 11.2018

Bachelor of Technology - Electronics & Communication Engineering

Vaagdevi Engineering College

Intermediate - PCM

SR Junior College

SSC -

Unique High School
Keerthana Nibhanapuri