Hi, I’m Santosh Sai Ram Kottapalli. I am a data engineer with 3+ years of hands-on experience designing, building, and optimizing scalable ETL/ELT pipelines across AWS, Azure, and Databricks, using Python, PySpark, SQL, Spark Structured Streaming, Delta Lake, Snowflake, and Redshift to deliver high-performance, cloud-native data solutions for analytics, ML, and real-time use cases. I excel at building enterprise ingestion and transformation frameworks, orchestrating end-to-end data workflows with metadata-driven pipelines, DAGs, triggers, CI/CD integrations, and automated quality checks. I am passionate about data quality, governance, and observability, and I enjoy collaborating with cross-functional teams to empower data science and analytics throughout the organization.

Santosh Sai Ram Kottapalli

Hi, I’m Santosh Sai Ram Kottapalli. I am a data engineer with 3+ years of hands-on experience designing, building, and optimizing scalable ETL/ELT pipelines across AWS, Azure, and Databricks, using Python, PySpark, SQL, Spark Structured Streaming, Delta Lake, Snowflake, and Redshift to deliver high-performance, cloud-native data solutions for analytics, ML, and real-time use cases. I excel at building enterprise ingestion and transformation frameworks, orchestrating end-to-end data workflows with metadata-driven pipelines, DAGs, triggers, CI/CD integrations, and automated quality checks. I am passionate about data quality, governance, and observability, and I enjoy collaborating with cross-functional teams to empower data science and analytics throughout the organization.

Available to hire

Hi, I’m Santosh Sai Ram Kottapalli. I am a data engineer with 3+ years of hands-on experience designing, building, and optimizing scalable ETL/ELT pipelines across AWS, Azure, and Databricks, using Python, PySpark, SQL, Spark Structured Streaming, Delta Lake, Snowflake, and Redshift to deliver high-performance, cloud-native data solutions for analytics, ML, and real-time use cases.

I excel at building enterprise ingestion and transformation frameworks, orchestrating end-to-end data workflows with metadata-driven pipelines, DAGs, triggers, CI/CD integrations, and automated quality checks. I am passionate about data quality, governance, and observability, and I enjoy collaborating with cross-functional teams to empower data science and analytics throughout the organization.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Work Experience

Data Engineer at CBRE
June 1, 2024 - Present
Designed and developed large-scale ETL/ELT pipelines using Azure Data Factory, Databricks, PySpark, and Scala to process structured, semi-structured, and streaming financial datasets across lakehouse layers (Bronze/Silver/Gold). Implemented medallion architecture with Delta Lake for governed analytics, versioning, and auditability. Built real-time streaming pipelines using Spark Structured Streaming, Event Hubs, and Kafka to support near real-time insights into fraud, compliance, and credit decisioning. Engineered secure ingestion workflows through Apache NiFi, integrating APIs, batch sources, and third-party systems with full metadata capture and lineage tracking. Developed Snowflake dbt models and dynamic SQL transformations, optimizing reporting performance by ~40% for financial analytics. Automated CI/CD deployments using Terraform and GitHub Actions; collaborated with DBAs and application teams to design internal data services and SQL-based interfaces. Collaborated with Data Scien
Data Engineer at Cognizant
January 1, 2022 - June 1, 2023
Developed distributed ETL pipelines using PySpark, AWS Glue, and Databricks to process high-volume batch, streaming, and event-driven datasets across multiple domains. Designed and optimized Redshift and Snowflake data models (partitioning, clustering, materialized views) to improve analytics performance by 30–50%. Built modular dbt transformation layers with automated testing and lineage, and supported enterprise systems by integrating internal APIs and SQL Server sources. Engineered fully automated data workflows using Apache Airflow, AWS Lambda, and CloudWatch, reducing operational troubleshooting time by 25%. Architected scalable AWS S3 data lake structures, incorporating JSON, Avro, and Parquet datasets to enable analytics on Athena and Databricks. Developed Python- and shell-based validation frameworks, improving reconciliation accuracy and reducing manual QA by 60%.
Data Engineer (Intern) at Cognizant
July 1, 2021 - December 1, 2021
Built and supported ETL data pipelines using Python, SQL, and PySpark to ingest, transform, and load structured and semi-structured data into AWS S3 and Databricks environments. Assisted in designing scalable data architectures and integrating multiple data sources including APIs, relational databases, and flat files, ensuring reliable end-to-end data flow. Implemented data validation, reconciliation, and quality checks to improve data accuracy and support downstream analytics and reporting. Collaborated with data engineers, analysts, and stakeholders to translate requirements into optimized SQL queries and transformation logic. Performed testing, debugging, and performance tuning of data pipelines, following SDLC practices and version control.

Education

Master of Science at University of Missouri–Kansas City
August 1, 2023 - May 1, 2025
Bachelor of Technology in Computer Science Engineering at Sathyabama Institute of Science & Technology
June 1, 2019 - June 1, 2023

Qualifications

Add your qualifications or awards here.

Industry Experience

Computers & Electronics, Professional Services, Real Estate & Construction, Software & Internet, Financial Services