I’m a Generative AI and data engineering professional with about four years of experience building and deploying production AI applications and large-scale data pipelines. I specialize in RAG and agentic AI workflows, real-time voice AI systems, and practical NLP solutions using Python and modern LLM tooling like LangChain/LangGraph, Hugging Face, OpenAI APIs, and Amazon Bedrock. Across my roles, I’ve also focused on building scalable batch and streaming data infrastructure with PySpark, Kafka, Databricks, and governed cloud data platforms like AWS and Snowflake. I care deeply about LLM lifecycle quality—prompt engineering, embeddings/vector search, evaluation with RAGAS, and LLMOps with tracing and monitoring—then delivering it through production-ready services (FastAPI, Docker, CI/CD, and Terraform) and strong data governance practices.

SAI SREE KURAPATI

I’m a Generative AI and data engineering professional with about four years of experience building and deploying production AI applications and large-scale data pipelines. I specialize in RAG and agentic AI workflows, real-time voice AI systems, and practical NLP solutions using Python and modern LLM tooling like LangChain/LangGraph, Hugging Face, OpenAI APIs, and Amazon Bedrock. Across my roles, I’ve also focused on building scalable batch and streaming data infrastructure with PySpark, Kafka, Databricks, and governed cloud data platforms like AWS and Snowflake. I care deeply about LLM lifecycle quality—prompt engineering, embeddings/vector search, evaluation with RAGAS, and LLMOps with tracing and monitoring—then delivering it through production-ready services (FastAPI, Docker, CI/CD, and Terraform) and strong data governance practices.

Available to hire

I’m a Generative AI and data engineering professional with about four years of experience building and deploying production AI applications and large-scale data pipelines. I specialize in RAG and agentic AI workflows, real-time voice AI systems, and practical NLP solutions using Python and modern LLM tooling like LangChain/LangGraph, Hugging Face, OpenAI APIs, and Amazon Bedrock.

Across my roles, I’ve also focused on building scalable batch and streaming data infrastructure with PySpark, Kafka, Databricks, and governed cloud data platforms like AWS and Snowflake. I care deeply about LLM lifecycle quality—prompt engineering, embeddings/vector search, evaluation with RAGAS, and LLMOps with tracing and monitoring—then delivering it through production-ready services (FastAPI, Docker, CI/CD, and Terraform) and strong data governance practices.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Beginner
Beginner
See more

Work Experience

Gen AI / Machine Learning Engineer & Data Engineer at CVS HealthCare — Dallas, Texas
August 1, 2025 - Present
Built transformer-based RAG pipelines integrating S3, Redshift, FAISS, and Pinecone using LangChain, Amazon Bedrock, and OpenAI APIs, improving retrieval precision by 35% via metadata filtering and context-aware search across 1M+ documents. Developed agentic AI workflows with LangGraph and tool calling to orchestrate retrieval, analysis, response generation, and validation, delivering production-ready FastAPI services and reducing manual intervention by 60%. Automated LLMOps workflows for embedding refresh and evaluation using RAGAS, with LangSmith tracing to maintain quality across releases. Implemented scalable inference services in Python/FastAPI with embeddings and vector search for low-latency conversational AI and auditable interaction metadata. Built real-time fraud analytics using PySpark with AWS Kinesis and Kafka for sub-second monitoring over 1M+ daily transactions. Optimized Spark ETL to improve runtime by 40% and migrated 15+ legacy pipelines to AWS Glue/EMR; built metadat
AWS Data Engineer at Truist Bank — Atlanta, Georgia
June 1, 2024 - May 1, 2025
Designed modular Airflow DAGs to orchestrate Glue and EMR PySpark jobs (ingestion, cleansing, enrichment, loading), reducing manual pipeline failures by 60% across 10+ workflows. Orchestrated EMR/Glue/Redshift deployments with Terraform and CloudFormation; developed Glue workflows to ingest customer/transaction/operational data from SAP and 15+ upstream systems into a governed partitioned data lake, cutting provisioning time by 40%. Built Snowflake and Redshift dimensional models supporting risk analytics and regulatory reporting, integrating Tableau with Redshift to reduce reporting time by 50%. Optimized Spark pipelines on EMR/HDFS for high-volume datasets, improving performance by 30% and enabling low-latency access for analytics and modeling. Implemented data quality and governance using Great Expectations and AWS Glue Data Catalog, improving schema validation, lineage tracking, and compliance across 15+ pipelines. Built production-grade Python/PySpark pipelines with automated test
Azure Data Engineer at DXC Technologies — Bangalore, India
January 1, 2022 - July 1, 2023
Optimized Snowflake query performance using clustering/partitioning and extended Medallion outputs to feed downstream AI dashboards and Databricks vector-store semantic search. Implemented CDC pipelines using Debezium, Kafka, and Databricks Delta Lake with MERGE INTO for SCD Type 2 history and idempotent replays, reducing backfill time by 60%. Built incremental ingestion pipelines with Databricks Auto Loader from ADLS Gen2 using checkpointing and schema evolution, reducing latency from hourly to real-time. Created batch and streaming ETL pipelines with Spark and Kafka to process policy data, designing a Medallion (bronze/silver/gold) architecture consolidating data from Hadoop, MySQL, and S3 into ADLS Gen2 for analytics across 25+ geographies. Enforced fine-grained access controls across 100+ notebooks with Unity Catalog and Azure AD, including row/column-level PII masking and CI-enforced data contracts for audit-ready compliance.

Education

Master’s in Business Analytics & Artificial Intelligence (Data Science) at The University of Texas at Dallas
August 1, 2023 - May 1, 2025

Qualifications

Databricks Certified: Generative AI Engineer Associate
January 11, 2030 - September 1, 2026
Databricks Certified: Data Engineer Associate
January 11, 2030 - September 1, 2026
AWS Certified Data Engineer - Associate
January 11, 2030 - September 1, 2026

Industry Experience

Software & Internet, Financial Services, Healthcare, Telecommunications, Professional Services, Education

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Beginner
Beginner
See more

Hire a Data Engineer

We have the best data engineer experts on Twine. Hire a data engineer in Richardson today.