I’m a Data Engineer with 6+ years of experience building scalable, high-performance data pipelines using Python, PySpark, SQL, and AWS. I help businesses transform raw data into reliable, structured datasets that drive analytics, reporting, and decision-making. I specialize in designing and optimizing ETL pipelines that handle large-scale data efficiently while ensuring data quality, consistency, and performance. What I can help you with: -> Build and optimize ETL/ELT pipelines using PySpark/Scala and Python/Java. -> Design scalable data pipelines on AWS (Glue, EMR, Lambda, S3). -> Process and transform large datasets (batch processing). -> Data cleaning, validation, and quality checks. -> Improve pipeline performance and reduce processing time. -> Data modeling and structuring for analytics and reporting. My experience includes: -> Improving data pipeline efficiency by up to 40%. -> Reducing batch processing time by 20-30%. -> Building reliable data workflows for financial and large-scale systems. -> Working with structured and unstructured datasets. I focus on writing clean, efficient, and production-ready code, and I ensure that every solution is scalable, reliable, and aligned with business needs. If you’re looking for someone who can build or optimize your data pipelines and deliver high-quality results, feel free to reach out-I’d be happy to help.

Deepa Nandi

I’m a Data Engineer with 6+ years of experience building scalable, high-performance data pipelines using Python, PySpark, SQL, and AWS. I help businesses transform raw data into reliable, structured datasets that drive analytics, reporting, and decision-making. I specialize in designing and optimizing ETL pipelines that handle large-scale data efficiently while ensuring data quality, consistency, and performance. What I can help you with: -> Build and optimize ETL/ELT pipelines using PySpark/Scala and Python/Java. -> Design scalable data pipelines on AWS (Glue, EMR, Lambda, S3). -> Process and transform large datasets (batch processing). -> Data cleaning, validation, and quality checks. -> Improve pipeline performance and reduce processing time. -> Data modeling and structuring for analytics and reporting. My experience includes: -> Improving data pipeline efficiency by up to 40%. -> Reducing batch processing time by 20-30%. -> Building reliable data workflows for financial and large-scale systems. -> Working with structured and unstructured datasets. I focus on writing clean, efficient, and production-ready code, and I ensure that every solution is scalable, reliable, and aligned with business needs. If you’re looking for someone who can build or optimize your data pipelines and deliver high-quality results, feel free to reach out-I’d be happy to help.

Available to hire

I’m a Data Engineer with 6+ years of experience building scalable, high-performance data pipelines using Python, PySpark, SQL, and AWS. I help businesses transform raw data into reliable, structured datasets that drive analytics, reporting, and decision-making.

I specialize in designing and optimizing ETL pipelines that handle large-scale data efficiently while ensuring data quality, consistency, and performance.

What I can help you with:
-> Build and optimize ETL/ELT pipelines using PySpark/Scala and Python/Java.
-> Design scalable data pipelines on AWS (Glue, EMR, Lambda, S3).
-> Process and transform large datasets (batch processing).
-> Data cleaning, validation, and quality checks.
-> Improve pipeline performance and reduce processing time.
-> Data modeling and structuring for analytics and reporting.

My experience includes:
-> Improving data pipeline efficiency by up to 40%.
-> Reducing batch processing time by 20-30%.
-> Building reliable data workflows for financial and large-scale systems.
-> Working with structured and unstructured datasets.

I focus on writing clean, efficient, and production-ready code, and I ensure that every solution is scalable, reliable, and aligned with business needs.

If you’re looking for someone who can build or optimize your data pipelines and deliver high-quality results, feel free to reach out-I’d be happy to help.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert

Language

English
Fluent

Work Experience

Associate Data Engineer at J.P. Morgan Chase & Co. (JPMC)
September 1, 2025 - Present
Built and optimized ETL pipelines using Scala, PySpark, AWS Glue, and EMR, improving throughput by 35% and reducing batch runtime by 20%. Designed data ingestion frameworks leveraging Amazon S3 for secure, scalable storage of structured and unstructured financial data. Automated pipeline orchestration using AWS Lambda and Step Functions, reducing manual interventions by 40%. Developed SQL transformations for regulatory and risk datasets, increasing reporting accuracy by 30%. Implemented validation and reconciliation checks to improve data quality, reducing downstream defects by 20%. Collaborated with global teams to translate regulatory requirements into scalable data engineering solutions.
Senior Data Engineer at Accenture
October 1, 2024 - August 1, 2025
Designed and enhanced data pipelines using PySpark and Python, increasing data processing efficiency by 30% and reducing transformation time across diverse data types by 25%. Developed and maintained ETL/ELT processes across 5+ systems, improving data flow efficiency by 40% and cutting processing times by 20%. Worked with Google Cloud Storage for large-scale data storage and retrieval. Collaborated with cross-functional teams to gather requirements and deliver data integration solutions, resulting in a 10% improvement in stakeholder satisfaction with data accessibility. Troubleshot and resolved data pipeline issues, increasing pipeline efficiency by 15%.
Data Engineer at TCS Digital
February 1, 2020 - October 1, 2024
Developed and optimized ETL processes using Hadoop, Spark, Hive, and Python, integrated with Google Cloud Platform (GCP) to process large datasets, improving data processing speed by 30%. Employed data cleaning techniques to resolve issues like duplicate records, missing values, and incorrect data types, improving data integrity by 25%. Deployed Slowly Changing Dimensions (SCD) strategies using Python and GCP, improving data accuracy by 30% and enhancing data consistency, resulting in a 25% improvement in decision-making efficiency for business users. Worked with Google Cloud Platform and other cloud technologies as part of ongoing projects, gaining foundational knowledge of cloud architecture and services.

Education

Bachelor of Computer Science at University of Calcutta
June 1, 2016 - June 1, 2019

Qualifications

Google Cloud Certified - Associate Cloud Engineer
January 11, 2030 - April 1, 2026

Industry Experience

Financial Services, Software & Internet, Professional Services
    End-to-End ETL Pipeline Using PySpark and AWS
    Built a scalable ETL pipeline to process large volumes of raw data from multiple sources into structured datasets for analytics. Extracted data from APIs and files, performed data cleaning, transformation, and validation using PySpark, and loaded the processed data into a cloud storage layer. Ensured data quality, optimized performance, and enabled reliable reporting for business insights.
    End-to-End ETL Pipeline Using PySpark and AWS

    Built a scalable ETL pipeline to process large volumes of raw data from multiple sources into structured datasets for analytics. Extracted data from APIs and files, performed data cleaning, transformation, and validation using PySpark, and loaded the processed data into a cloud storage layer. Ensured data quality, optimized performance, and enabled reliable reporting for business insights.