I am an AWS Data Engineer with 7+ years of experience designing, building, and optimizing cloud-native data pipelines. I specialize in end-to-end data engineering lifecycle in Agile environments, leveraging AWS services (Glue, Redshift, S3, EMR, Lambda, Kinesis, etc.) and Python/Scala to deliver robust, scalable analytics platforms. In my recent roles at KeyBank, Pfizer, and NextEra Energy, I built data lakes, real-time streaming, and data warehouse solutions; I automated infrastructure with Terraform/CloudFormation; I implemented data quality, governance, and lineage; I collaborated with data scientists and product teams to accelerate insights.

Sanjay Bhargav Komma

I am an AWS Data Engineer with 7+ years of experience designing, building, and optimizing cloud-native data pipelines. I specialize in end-to-end data engineering lifecycle in Agile environments, leveraging AWS services (Glue, Redshift, S3, EMR, Lambda, Kinesis, etc.) and Python/Scala to deliver robust, scalable analytics platforms. In my recent roles at KeyBank, Pfizer, and NextEra Energy, I built data lakes, real-time streaming, and data warehouse solutions; I automated infrastructure with Terraform/CloudFormation; I implemented data quality, governance, and lineage; I collaborated with data scientists and product teams to accelerate insights.

Available to hire

I am an AWS Data Engineer with 7+ years of experience designing, building, and optimizing cloud-native data pipelines. I specialize in end-to-end data engineering lifecycle in Agile environments, leveraging AWS services (Glue, Redshift, S3, EMR, Lambda, Kinesis, etc.) and Python/Scala to deliver robust, scalable analytics platforms.

In my recent roles at KeyBank, Pfizer, and NextEra Energy, I built data lakes, real-time streaming, and data warehouse solutions; I automated infrastructure with Terraform/CloudFormation; I implemented data quality, governance, and lineage; I collaborated with data scientists and product teams to accelerate insights.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Language

English
Fluent

Work Experience

Senior AWS Data Engineer at KeyBank
May 1, 2023 - November 4, 2025
Designed and implemented scalable ETL/ELT pipelines using AWS Glue, Lambda, and PySpark to process structured and unstructured data from diverse sources into centralized data lakes. Built data lakes on S3 with efficient partitioning, compression, and lifecycle management. Developed real-time streaming architectures using Kinesis, Kafka, and AWS Lambda for event-driven analytics. Implemented data warehouse solutions with Redshift, integrated with Glue Catalog and Spectrum. Automated infrastructure with Terraform and CloudFormation. Optimized Spark jobs on EMR with partitioning, bucketing, and caching for large datasets. Created CI/CD pipelines with Jenkins and AWS CodePipeline for automated testing, deployment, and monitoring. Integrated data from MySQL, PostgreSQL, Oracle, MongoDB, and DynamoDB; designed event-driven workflows with SQS/SNS; serverless data processing using Lambda, Step Functions, and Glue Jobs. Built data quality and validation frameworks in Python/Pandas for governanc
Senior AWS Data Engineer at Pfizer
April 1, 2023 - April 1, 2023
Designed and developed end-to-end ETL/ELT pipelines on AWS (Glue, Lambda, Step Functions, Redshift) for clinical and R&D datasets. Built HIPAA-compliant data lake on S3 with Raw/Staging/Curated zones, encryption, versioning, and lifecycle management. Optimized PySpark transformations on EMR for high-volume genomic and patient data; applied partitioning, bucketing, and caching to reduce processing times. Implemented data validation and schema enforcement using Glue and DynamicFrames to ensure FDA/GxP compliance. Integrated Glue Data Catalog and Lake Formation for metadata, access control, and lineage. Established CI/CD with Jenkins and CodePipeline; provisioned infrastructure with Terraform and CloudFormation. Configured CloudWatch dashboards and EventBridge rules for monitoring and orchestration. Used Parquet/ORC formats to improve storage and query performance by ~35%. Migrated legacy Informatica/SQL Server workflows to AWS Glue; introduced event-driven, serverless pipelines. Implemen
Big Data Engineer at NextEra Energy
July 1, 2020 - July 1, 2020
Designed and developed robust ETL pipelines for large-scale data, building cloud data warehouses and data lakes on AWS. Implemented real-time streaming with Kafka/Kinesis and Spark Streaming; built batch and micro-batch processing with Spark/Hadoop. Developed dimensional data models (star, snowflake) and automated workflow orchestration. Implemented data validation, quality checks, and security controls; optimized SQL and Spark jobs for performance. Assisted in migrating on-premise data warehouses to cloud-based platforms; built reusable Python/SQL scripts and dashboards for business stakeholders. Ensured governance and auditability; supported CI/CD for ETL and infrastructure deployments.

Education

Master's in Computer Science at Illinois Institute of Technology, Chicago, IL
January 11, 2030 - November 4, 2025

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Financial Services, Healthcare, Professional Services