Summary
Overview
Work History
Education
Skills
Affiliations
Timeline
Generic

Nitin Raj

Gurugram

Summary

Over 9 years of expertise in data engineering, specializing in ETL flow design and development, data warehousing, and database management. Proven track record in executing full SDLC processes, including requirements gathering, development, testing, and post-production support. Recognized for delivering high-quality solutions that enhance data processing efficiency and accuracy.

Overview

10
10
years of professional experience

Work History

Data Engineer

HCL
Gurugram
01.2024 - Current
  • Leading SAS-to-AWS migration initiatives, transforming legacy flat-file and SAS-based processes into scalable PySpark-based data pipelines.
  • Designed and developed ETL pipelines across landing, raw, curated, and dimensional layers, applying dimensional data modeling techniques to support analytics and reporting.
  • Utilized PySpark for complex transformations, including data cleansing, multi-table joins, aggregations, and analytical computations.
  • Delivered business-ready outputs in Excel and analytical datasets, integrated with Tableau dashboards to support decision-making.
  • Designed and implemented data reconciliation frameworks to validate raw-layer data against source systems, generating audit reports and SNS-based alerts.
  • Developed housekeeping and maintenance scripts for Apache Iceberg tables, supporting storage optimization and metadata cleanup.
  • Optimized multiple Spark jobs to improve performance, scalability, and resource utilization across workloads.
  • Improved development productivity by leveraging GenAI tools such as GitHub Copilot and internal GPT platforms for faster coding and analysis.
  • Actively supporting BAU operations, addressing production issues, enhancements, and performance improvements.
  • Insurance Client: Sunlife Financials
  • Tech Stack: AWS DMS, Amazon S3, AWS Glue (PySpark), Apache Iceberg, Athena, SNS, CI/CD (Dev–SIT–UAT–Prod), Linux, SQL

Data Engineer

HCL
Gurugram
07.2021 - 01.2024
  • Designed and implemented end-to-end AWS-based data ingestion and transformation pipelines, migrating legacy ETL workflows to a cloud-native data lake architecture.
  • Ingested CDC-enabled source data from DB systems into Amazon S3 landing layer using AWS DMS, ensuring near real-time data availability and change tracking.
  • Integrated file-based source feeds delivered via GoAnywhere FTP, landing structured and semi-structured files directly into S3 landing zone.
  • Implemented landing-to-raw data processing using AWS Glue PySpark jobs, converting raw files into standardized Parquet format for optimized storage and performance.
  • Developed pre-processing scripts to cleanse and normalize incoming files (format corrections, delimiter handling, schema alignment) before ingestion into raw datasets.
  • Built and maintained Apache Iceberg tables in the raw layer, enabling ACID transactions, schema evolution, and time travel, with direct query access through Amazon Athena.
  • Designed raw-to-curated transformation pipelines, implementing SCD Type 1 and Type 2 logic to create history-aware and current-state datasets for analytics and reporting.
  • Published curated datasets for downstream BI and ML consumers, enabling analysis via Tableau and model development using Amazon SageMaker.
  • Implemented job completion notifications using Amazon SNS, triggering automated email alerts for pipeline success and failure scenarios.
  • Supported CI/CD-driven promotion of Glue jobs and configurations across Dev, SIT, UAT, and Production environments following controlled release practices.
  • Provided production support post go-live, handling data issues, pipeline failures, and performance tuning until formal stakeholder sign-off.
  • insurance Client : Sunlife Financials
  • Tech Stack: PySpark, PostgreSQL, AWS Glue, Amazon RDS, AWS DMS, Amazon S3, Apache Iceberg, SNS

Data Engineer

Accenture
Gurugram
02.2021 - 06.2021
  • Independently designed and implemented end-to-end PySpark batch data pipelines for customer, loan, repayment, and defaulter datasets (~2M+ records, 118 columns, ~100 GB data).
  • Performed large-scale data cleansing and transformations including null handling, deduplication, schema standardization, datatype casting, and column normalization.
  • Enriched datasets with ingestion and processing timestamps to support data lineage, traceability, and audit requirements.
  • Created external Hive tables on curated datasets to enable SQL-based analytics while preserving raw data immutability.
  • Developed a loan risk scoring framework by combining repayment behavior (20%), defaulter history (45%), and financial health indicators (35%) using configurable scoring logic.
  • Built aggregated and analytical views to support downstream reporting use cases, optimizing for performance and query reuse.
  • Orchestrated Spark batch jobs in a production-like setup using spark-submit, configuring application parameters and submitting jobs to YARN-managed cluster for scheduled execution.
  • Validated Spark applications using YARN ResourceManager and Spark UI, analyzing stages, shuffles, executor behavior, and memory utilization to ensure stable execution.
  • Gained hands-on experience with distributed data processing concepts including partitioning strategies, shuffle optimization, and cluster resource utilization.
  • Delivered a stable, analytics-ready dataset suitable for batch loan risk analysis and reporting.
  • Consumer Finance Domain
  • Tech Stack: PySpark, Apache Spark (YARN), Hadoop HDFS, Hive Metastore, Parquet, CSV, Spark UI, Linux

ETL Developer

COGNIZANT
05.2016 - 01.2021
  • Requirement gathering & impact analysis
  • Preparing necessary design artifacts (LLD, HLD, Mapping document)
  • Coordinate offshore with onshore team, prepare and publish status report
  • Preparing clarification, issue tracker and share with team
  • Based on priority of the task, split and assign to the team
  • Code Development, Unit Testing, Code review, Performance tuning
  • Preparing environment, test plan to perform System integration, performance testing
  • Coordinate with QA team, analyzing on the observations and bug fixing
  • Code debugging, troubleshooting the issue raised by Business
  • DDL preparation for any new/existing database changes
  • Code migration from lower region to higher regions for User acceptance testing
  • Preparing definition for maestro jobs for scheduling new jobs
  • Change ticket creation for production code deployment, Pre and Postproduction code validation
  • Schedule meeting with Application support team to give the KT after each production deployment
  • Preparation and updating the batch documents for different projects
  • Responsible for change management and deployment.
  • Preparation of UTC documents, change management documents and review checklists
  • Working on several process related works i.e. (Mainspring, PMR, C2, BCP)
  • insurance Client :MetLife
  • Tech Stack: Informatica 9.6, DB2, AQT, IBM Maestro, WinSCP, PuTTY

Education

Bachelor of Technology - Electronics and Communications Engineering

College of Engineering Adoor
Adoor, Kerala, India
Adoor, Kerala, India

Skills

Python (PySpark, Pandas) and SQL databases

AWS services (Glue, S3, RDS, DMS, SNS, IAM)

Data warehousing techniques

Generative AI utilization

Project management tools (Jira, Confluence)

Affiliations

Distributed Computing, Distributed Storage, Reading Manga, Travel

Timeline

Data Engineer

HCL
01.2024 - Current

Data Engineer

HCL
07.2021 - 01.2024

Data Engineer

Accenture
02.2021 - 06.2021

ETL Developer

COGNIZANT
05.2016 - 01.2021

Bachelor of Technology - Electronics and Communications Engineering

College of Engineering Adoor
Nitin Raj