Data Engineer with 5+ years of experience building scalable batch and real-time data pipelines, cloud-native data platforms, and enterprise data lake/warehouse solutions across banking, enterprise, and academic environments. Experienced with Python, SQL, Spark (PySpark), Kafka, Hadoop, Airflow, and cloud services on AWS and Azure to deliver reliable ETL/ELT for analytics and reporting. Strong in data modeling (dimensional modeling, SCD), performance tuning, data validation/reconciliation, and operational observability (monitoring, logging, alerting). Adept at orchestrating end-to-end pipelines with governance for sensitive data and collaborating with stakeholders to standardize reporting logic, including migration to centralized frameworks.

Sindhuja Reddy Sama

Data Engineer with 5+ years of experience building scalable batch and real-time data pipelines, cloud-native data platforms, and enterprise data lake/warehouse solutions across banking, enterprise, and academic environments. Experienced with Python, SQL, Spark (PySpark), Kafka, Hadoop, Airflow, and cloud services on AWS and Azure to deliver reliable ETL/ELT for analytics and reporting. Strong in data modeling (dimensional modeling, SCD), performance tuning, data validation/reconciliation, and operational observability (monitoring, logging, alerting). Adept at orchestrating end-to-end pipelines with governance for sensitive data and collaborating with stakeholders to standardize reporting logic, including migration to centralized frameworks.

Available to hire

Data Engineer with 5+ years of experience building scalable batch and real-time data pipelines, cloud-native data platforms, and enterprise data lake/warehouse solutions across banking, enterprise, and academic environments. Experienced with Python, SQL, Spark (PySpark), Kafka, Hadoop, Airflow, and cloud services on AWS and Azure to deliver reliable ETL/ELT for analytics and reporting.

Strong in data modeling (dimensional modeling, SCD), performance tuning, data validation/reconciliation, and operational observability (monitoring, logging, alerting). Adept at orchestrating end-to-end pipelines with governance for sensitive data and collaborating with stakeholders to standardize reporting logic, including migration to centralized frameworks.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert

Language

English
Advanced

Work Experience

Data Engineer at Global Atlantic Financial Company (KKR)
July 1, 2025 - February 1, 2026
Migrated legacy deal-specific actuarial preprocessing logic into the centralized ADLR framework within the Enterprise Data Hub to standardize reporting across business units. Built complex Redshift transformations for policy-level, fund-level, and liability reporting datasets; implemented reconciliation queries to validate financial aggregates and reconcile account/fund/liability metrics across reporting periods. Performed production discrepancy root-cause analysis (e.g., termination edge cases, reporting cutoff variations) and validated data consistency across Policy, Fund, and AVRF extract tables during migration/release cycles. Added ETL data-validation checkpoints to detect inconsistencies prior to downstream actuarial reporting. Enabled scalable onboarding of new deals using metadata-driven configuration mappings rather than hardcoded logic, and handled schema evolution (new settlement file tables and decommissioning deprecated structures). Optimized Redshift queries using joins,
Data Engineer at US Bank
June 1, 2024 - June 1, 2025
Developed and maintained batch and near real-time pipelines using AWS Glue, AWS EMR, and PySpark for large-scale banking and transaction datasets supporting fraud detection and risk analytics. Implemented streaming ingestion workflows with Apache Kafka, AWS Kinesis, and Spark Structured Streaming for low-latency transaction monitoring. Orchestrated reusable ETL workflows with Apache Airflow for reliable scheduling and dependency management across data domains. Designed transformations and optimized data models in Snowflake, Amazon Redshift, and Delta Lake for regulatory reporting and risk analytics. Applied SCD Type 2 modeling to preserve historical customer/account changes. Built data validation, reconciliation, and schema enforcement checks to improve accuracy for audit-ready financial reporting. Managed secure lake storage in AWS S3 with partitioning, IAM-based access controls, and encryption; enforced governance for sensitive data including PII masking and RBAC. Tuned Spark workloa
Data Engineer at Wichita State University Library Technologies
January 1, 2023 - May 1, 2024
Built data ingestion and ETL pipelines using Python and SQL to integrate data from academic systems, library databases, and research repositories. Created automated ETL workflows using PySpark and Apache Airflow to process structured and semi-structured datasets. Ingested data from MongoDB, MySQL, and PostgreSQL and stored curated datasets into an AWS S3-based data lake. Performed transformation and cleansing using Pandas, NumPy, and Spark SQL to standardize JSON/CSV/XML datasets for analytics. Developed complex SQL queries and stored procedures in Oracle and SQL Server for academic reporting and institutional research analytics. Implemented data validation frameworks using Great Expectations to improve data quality across pipelines. Built lightweight REST APIs using Flask to expose curated datasets to internal university applications. Created analytics dashboards in Power BI and Tableau for library usage metrics and academic KPIs. Monitored ETL execution with AWS CloudWatch, troublesh
Data Engineer at Accenture, Hyderabad, India
April 1, 2021 - June 1, 2022
Developed data transformation jobs using Apache Spark (PySpark) and Spark SQL for enterprise analytics and reporting. Built and maintained ETL pipelines using Azure Data Factory to ingest data from Oracle/MySQL into cloud storage. Processed and transformed structured and semi-structured data using Azure Databricks across formats like JSON, Parquet, and CSV. Supported analytical dataset creation using Microsoft Fabric and Azure Synapse services for enterprise reporting and BI. Implemented Spark SQL and Hive-based transformations, wrote optimized SQL/transformation logic for downstream analytics, and assisted in near real-time event-driven ingestion using Apache Kafka and Spark Streaming. Worked with Azure Synapse curated datasets for analytical workloads. Automated ETL scheduling/monitoring using Apache Airflow to improve reliability. Implemented Python-based data validation/cleansing for accuracy and consistency. Collaborated with DevOps to package data applications using Docker and su

Education

Master's Degree (Computer Science) at Wichita State University
August 1, 2022 - May 1, 2024
Bachelor's Degree (Information Technology) at Anurag Group of Institutions
May 1, 2017 - March 1, 2021

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Education, Professional Services