Summary
Overview
Work History
Education
Skills
Accomplishments
Certification
Personal Projects
Timeline
Generic

Ajit Kumar Pandit

Bangalore

Summary

Data Science Engineer with 7 years of extensive experience, specializing in architecting scalable ETL processes and leveraging cloud computing on GCP. Proven ability to enhance data pipelines, develop robust machine learning solutions, and collaborate effectively with cross-functional teams. Proficient in Python, Java, and SQL, with a strong focus on problem-solving and delivering impactful, data-driven solutions that drive business innovation.

Overview

8
8
years of professional experience
1
1
Certification

Work History

Data Science Engineer III

Sabre
Bangalore
01.2022 - Current
  • Engineered and maintained scalable data pipelines (ETL) using cloud-native services like Google Cloud Platform (GCP), including BigQuery and Dataflow, to process petabytes of real-time and batch data.
  • Developed production-ready code in Python and Java, following best practices for software development, version control (Git), and continuous integration/continuous deployment (CI/CD).
  • Collaborated with cross-functional teams, including product managers and software engineers, to translate business requirements into technical specifications for data-driven applications.
  • Responsible for migration of all the codes from Bitbucket to GitHub Enterprise.
  • For CI/CD, currently using Jenkins and gradually migrating it to GitHub Actions.
  • Components involved in this project - GCP Services, Oracle DB & Teradata ( data migrated to BigQuery), Apache Beam, Java, Python, SQL, Spock Framework, Groovy, Junit Testing, Terraform, Monorepo (switching gradually from poly repo), Dashboards, GenAI.

Data Engineer

Schlumberger Pvt Ltd
Pune
07.2018 - 01.2022
  • Designed and developed ETL flow for data ingestion and data transformation using Informatica BDM concepts to send the data to downstream consumer on GC and Azure Cloud which makes sure to 100% data availability for consumers respective ponds.
  • Developed Sqoop based ETL pipeline running on Spark Engine which is quite faster than native connectors. Both source and target has configured as sqoop connection. For downstream GC storage is used.
  • Developed and managed custom data ingestion framework using Nifi (ETL Tool) and GCP.
  • Planning and Executing QA and Production release cycle activities.
  • Handling the failure fix or enhancement required in operational flow and responsible for identity and fix of any failure in the data pipeline from source to downstream.
  • GCP Components used in the project - Airflow, Pubsub, Dataflow, Bigquery, Compute Engine, Logging, Monitoring, Dashboard, Cloud Shell, Security Service, IAM services, Triggers/Cloud Function.
  • Attended 1 month workshop conducted by the company for Data Engineering with GCP.

Trainee

IBNC
Pune
05.2017 - 06.2017
  • Hands on Big Data Hadoop and Cloud Computing
  • Exposure of Google Cloud Platform and its services like Storage, Dataflow, Pubsub, Bigquery, Cloud Shell, VMs.

Education

Bachelor of Engineering - Computer Engineering

Army Institute of Technology
Pune
06-2018

Intermediate Certificate - Science

Kendriya Vidyalaya Sangathan
Danapur Cantt
05-2014

Matriculation -

Kendriya Vidyalaya Sangathan
Danapur Cantt
05-2012

Skills

  • ETL processes
  • Cloud computing - GCP
  • Data analysis
  • Data structure understanding
  • Programming languages - C, Java, Python, Groovy, Shell
  • Object-oriented programming
  • SQL, MySQL, Oracle, BigQuey
  • Problem solving
  • Software development
  • Team collaboration
  • Process automation
  • Agile methodologies

Accomplishments

  • Certification of Merit - Scoring 10CGPA in 10th standard
  • Best Outgoing Student (Batch 2014) - Being among the school topper in Class 12th
  • AGIF scholarship - For excellent Academics Performance in Engineering.
  • Runner Up in AIT-UGCON-AMALGAM (03/2018) - 2nd Prize Winner in College Project Competition(Final Year)

Certification

  • ITIL Certification
  • Informatica BDM Certification

Personal Projects

  • Video Trend Analysis Using Youtube Data (07/2017) - PIG is used for sorting the interest of people watching video based on maximum rating, most viewed, etc
  • Sentiment Analysis Using Twitter Data - HDFS is used for storing data and PIG is used for analysis
  • Banking Management System as First Year Project.
  • Hybrid CAT using Naive's Bayes and 2-parameter model - A Platform to test aptitude of candidates by adapting the difficulty level. Paper published in Springer.

Timeline

Data Science Engineer III

Sabre
01.2022 - Current

Data Engineer

Schlumberger Pvt Ltd
07.2018 - 01.2022

Trainee

IBNC
05.2017 - 06.2017

Bachelor of Engineering - Computer Engineering

Army Institute of Technology

Intermediate Certificate - Science

Kendriya Vidyalaya Sangathan

Matriculation -

Kendriya Vidyalaya Sangathan
Ajit Kumar Pandit