
A results driven data engineer with 2 years and 10 months of experience in building and optimizing data pipelines. Proficient in SQL, Python and Data Engineering frameworks for scalable ETL solutions. Experience of building scalable ETL/ELT pipelines supporting structured and semi-structured data. Expertise in data transformation and collaborating with cross-functional teams to deliver data-driven insight.
• Assisted seniors for developing and optimising data pipelines using Databricks notebooks and PySpark, processing over 3 TB/day of data with 99.8% reliability.
• Assisted clients with data and cloud based enterprise software product related questions, feedback and complaints.
• Configured and managed Databricks clusters with appropriate instance types and autoscaling policies to optimize costs
• Excecuted PySpark scripts with multiple tasks, integrating notebooks and scheduled jobs for automated pipeline execution
• Analysed data systems performance, implementing necessary adjustments for optimal operation.
• Assisted in building batch processing pipelines using PySpark on Databricks platform with notebook-based development
• Wrote PySpark scripts for data extraction and transformation, utilizing Databricks widgets for parameterization.
• Performed data profiling and quality analysis using Databricks functions.
• Created basic ETL pipelines for monitoring execution metrics and data quality KPIs.
• Implemented data quality checks using constraints (CHECK, NOT NULL) and monitoring using Databricks.
• Conducted through data quality checks, rectifying inconsistencies and ensuring reliabilty of information.
• Improved data availability and analytics performance by ~30%.
• Crafted custom Extract, Transform, Load processes tailored to specific project needs, enhancing data usability.
• Conducted thorough data quality checks, rectifying inconsistencies and ensuring reliability of information.
• Designed and implemented scalable data pipelines, integrating diverse data sources for streamlined analysis.
• Closely monitored data warehouse functionality and performance, proactively responding to issues.
• Analysed, solved and corrected issues of databases and datasets.
• Python, SQL, PySpark, Databricks, Cluster Configuration, Auto Loader, Unity Catalog, ETL/ELT Architecture, Batch & Near Real-time Pipelines, Data Governance, Data Warehousing. AWS Services - S3, Glue, Athena,Quicksight.
Log File Processing & Error Analytics Platform -
Batch Incremental Data Pipeline -