I’m a Data Engineer with 6+ years of experience building scalable, high-performance data pipelines using Python, PySpark, SQL, and AWS. I help businesses transform raw data into reliable, structured datasets that drive analytics, reporting, and decision-making.
I specialize in designing and optimizing ETL pipelines that handle large-scale data efficiently while ensuring data quality, consistency, and performance.
What I can help you with:
-> Build and optimize ETL/ELT pipelines using PySpark/Scala and Python/Java.
-> Design scalable data pipelines on AWS (Glue, EMR, Lambda, S3).
-> Process and transform large datasets (batch processing).
-> Data cleaning, validation, and quality checks.
-> Improve pipeline performance and reduce processing time.
-> Data modeling and structuring for analytics and reporting.
My experience includes:
-> Improving data pipeline efficiency by up to 40%.
-> Reducing batch processing time by 20-30%.
-> Building reliable data workflows for financial and large-scale systems.
-> Working with structured and unstructured datasets.
I focus on writing clean, efficient, and production-ready code, and I ensure that every solution is scalable, reliable, and aligned with business needs.
If you’re looking for someone who can build or optimize your data pipelines and deliver high-quality results, feel free to reach out-I’d be happy to help.
Skills
Language
Work Experience
Education
Qualifications
Industry Experience
Built a scalable ETL pipeline to process large volumes of raw data from multiple sources into structured datasets for analytics. Extracted data from APIs and files, performed data cleaning, transformation, and validation using PySpark, and loaded the processed data into a cloud storage layer. Ensured data quality, optimized performance, and enabled reliable reporting for business insights.
Hire a Data Scientist
We have the best data scientist experts on Twine. Hire a data scientist in Bengaluru today.