I'm a Senior Data Engineer with 7+ years of experience designing, building, and optimizing large-scale ETL and streaming data pipelines across financial services, insurance, and healthcare. I thrive on delivering robust data platforms with governance and observability, and I enjoy collaborating with data scientists, analysts, and business stakeholders to translate complex data into actionable insights. I specialize in PySpark and AWS Glue, real-time ingestion with Kafka and Kinesis, cloud migrations to Snowflake and Redshift, and automating workflows with Airflow, CI/CD pipelines, and monitoring tools. My focus is on data modernization, quality, and governance to support fraud detection, portfolio monitoring, and enterprise analytics at scale.

Sai Vineeth Neeli

I'm a Senior Data Engineer with 7+ years of experience designing, building, and optimizing large-scale ETL and streaming data pipelines across financial services, insurance, and healthcare. I thrive on delivering robust data platforms with governance and observability, and I enjoy collaborating with data scientists, analysts, and business stakeholders to translate complex data into actionable insights. I specialize in PySpark and AWS Glue, real-time ingestion with Kafka and Kinesis, cloud migrations to Snowflake and Redshift, and automating workflows with Airflow, CI/CD pipelines, and monitoring tools. My focus is on data modernization, quality, and governance to support fraud detection, portfolio monitoring, and enterprise analytics at scale.

Available to hire

I’m a Senior Data Engineer with 7+ years of experience designing, building, and optimizing large-scale ETL and streaming data pipelines across financial services, insurance, and healthcare. I thrive on delivering robust data platforms with governance and observability, and I enjoy collaborating with data scientists, analysts, and business stakeholders to translate complex data into actionable insights.

I specialize in PySpark and AWS Glue, real-time ingestion with Kafka and Kinesis, cloud migrations to Snowflake and Redshift, and automating workflows with Airflow, CI/CD pipelines, and monitoring tools. My focus is on data modernization, quality, and governance to support fraud detection, portfolio monitoring, and enterprise analytics at scale.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Work Experience

Senior Data Engineer at U.S. Bank
April 1, 2023 - October 31, 2025
Led the design and development of PySpark and AWS Glue ETL pipelines handling structured and unstructured data, improving data lake and analytics performance. Built automated Airflow DAGs with Python scripts, reducing manual monitoring by 30% and increasing workflow reliability. Architected streaming fraud detection pipelines using Kafka and Kinesis to process millions of transactions hourly, cutting alert latency by 45%. Migrated Teradata and Oracle to Snowflake, enhancing governance and achieving 50% improvements in query performance and scalability. Implemented cost-optimized infrastructure with S3, EMR, CloudWatch, and EventBridge, lowering compute costs by 25%. Containerized ETL workflows with Docker and Jenkins CI/CD on Kubernetes to accelerate deployments. Established comprehensive metadata and data lineage tracking using AWS Glue Catalog and Apache Atlas. Implemented observability with Grafana, Prometheus, and Datadog to ensure data quality and reliability across distributed pi
Data Engineer at Sentry Insurance
March 1, 2023 - March 1, 2023
Built PySpark and Talend pipelines to ingest structured and unstructured data into Snowflake Data Lakes, improving data models and reporting accuracy by 35%. Automated Airflow DAG orchestration with parameterized workflows, saving 40+ hours monthly on manual oversight. Integrated Kafka streaming for telematics and claims data feeds to enable near real-time fraud scoring, reducing scoring latency by 20%. Tuned Spark jobs and SQL queries on large datasets, delivering up to 40% runtime improvements and lower cluster costs. Led cloud migrations from Sybase and Oracle to Snowflake and Redshift, strengthening governance and scalable analytics. Developed Python validation frameworks and anomaly detection systems to enhance data reliability and actuarial model robustness.
Data Engineer at CommonSpirit Health
December 1, 2020 - December 1, 2020
Engineered PySpark and AWS Glue ETL pipelines for healthcare data transformation, enabling scalable warehousing solutions. Implemented Kafka-based real-time ingestion of EHR and CMS data to support machine learning analytics and distributed computing for clinical insights. Developed HIPAA/PHI validation frameworks in Python to ensure security and compliance across sensitive pipelines. Automated Airflow workflows with recovery and alerting logic, reducing critical pipeline downtime by 25%. Created interactive Power BI dashboards visualizing claims and member utilization metrics to drive leadership decision-making. Collaborated in Agile teams, contributing to code reviews and reusable pipeline modules.

Education

Master's at Central Michigan University
January 11, 2030 - October 31, 2025
Master's, Information Systems at Central Michigan University
January 11, 2030 - October 31, 2025

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Healthcare, Professional Services, Software & Internet