Data Engineer with 6+ years of experience building scalable batch and streaming data pipelines using Spark, Kafka, and PySpark, handling multi-terabyte datasets across enterprise systems. Experienced in AWS-based data platforms (S3, Redshift, Kinesis), real-time stream processing, and ETL/ELT workflows using SQL, Informatica, and dbt. Strong background in orchestration with Apache Airflow, data warehousing with Snowflake/Redshift, and data modeling (star schema, dimensional modeling). Focused on data quality and governance using Great Expectations, along with monitoring and reliability improvements through CloudWatch and Grafana, and deploying containerized pipelines with Docker and Kubernetes.

Manasa Konaganti

Data Engineer with 6+ years of experience building scalable batch and streaming data pipelines using Spark, Kafka, and PySpark, handling multi-terabyte datasets across enterprise systems. Experienced in AWS-based data platforms (S3, Redshift, Kinesis), real-time stream processing, and ETL/ELT workflows using SQL, Informatica, and dbt. Strong background in orchestration with Apache Airflow, data warehousing with Snowflake/Redshift, and data modeling (star schema, dimensional modeling). Focused on data quality and governance using Great Expectations, along with monitoring and reliability improvements through CloudWatch and Grafana, and deploying containerized pipelines with Docker and Kubernetes.

Available to hire

Data Engineer with 6+ years of experience building scalable batch and streaming data pipelines using Spark, Kafka, and PySpark, handling multi-terabyte datasets across enterprise systems. Experienced in AWS-based data platforms (S3, Redshift, Kinesis), real-time stream processing, and ETL/ELT workflows using SQL, Informatica, and dbt.

Strong background in orchestration with Apache Airflow, data warehousing with Snowflake/Redshift, and data modeling (star schema, dimensional modeling). Focused on data quality and governance using Great Expectations, along with monitoring and reliability improvements through CloudWatch and Grafana, and deploying containerized pipelines with Docker and Kubernetes.

See more

Language

Work Experience

Data Engineer at AT&T
November 1, 2023 - Present
Delivered batch and streaming data pipelines using Spark and Kafka to process multi-terabyte datasets daily, reducing latency by 35% for enterprise analytics and reporting. Architected an AWS data platform using S3 and Redshift to store and query large datasets, improving accessibility by 40%. Enabled real-time ingestion pipelines using Kinesis and Spark Structured Streaming to deliver near real-time insights, reducing processing delays by 30%. Transformed raw data into structured datasets using dbt, orchestrated workflows with Apache Airflow, and implemented data quality checks using Great Expectations. Optimized dimensional star schema models for query performance and reduced dashboard load times. Deployed and monitored containerized data workloads using Docker, Kubernetes, CloudWatch, and Grafana, improving reliability and failure detection time.
Data Engineer at Accenture
January 1, 2020 - July 1, 2022
Engineered PySpark/Hadoop pipelines for large-scale enterprise analytics, improving processing performance by 30%. Delivered ETL workflows using Informatica and SQL to integrate data from multiple sources, improving consistency and reducing processing time by 25%. Enabled Snowflake analytics, improving query performance by 35%. Managed Airflow scheduling and dependencies, improving execution efficiency by 20% and reducing failures. Implemented Delta Lake architecture for reliable large-scale processing. Integrated Kafka streaming pipelines to reduce latency by 30%. Optimized SQL (partitioning) to improve retrieval by 25%, and monitored pipelines using Prometheus and Grafana to reduce incident resolution time by 30%. Enforced data governance with lineage and cataloging for compliance and visibility.
Jr Data Engineer at Accenture
January 1, 2019 - December 1, 2019
Delivered ETL pipelines using Python and SQL to process structured and semi-structured data, improving processing efficiency by 20%. Supported data ingestion using Hadoop and Hive to store and retrieve large datasets for analytics. Assisted data transformation for cleaning and structuring datasets, improving pipeline efficiency for reporting needs. Optimized data models to enable faster dashboard creation. Performed data validation and cleansing to reduce inconsistencies by 15%. Automated data processing tasks with Python to reduce manual effort by 25%. Optimized SQL extraction performance by 20%, monitored pipeline execution to identify failures early, and resolved production issues to improve stability and reliability.

Education

Master of Computer Science at New England College
October 1, 2023 - October 1, 2023
Bachelor of Engineering in Computer Science at DVR College of Engineering
July 1, 2019 - July 1, 2019
Master of Computer Science at New England College
October 1, 2023 - October 1, 2023
Bachelor of Engineering in Computer Science at DVR College of Engineering
July 1, 2019 - July 1, 2019

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Professional Services