I am a Data Engineer with 3+ years of experience, focused on building end-to-end data pipelines that unlock actionable business insights. I design robust, scalable data solutions and have hands-on experience with Apache Airflow, Databricks, and AWS Glue to ingest data from APIs and relational sources. I enjoy translating complex data into accessible analytics for stakeholders and continuously optimizing pipelines for performance and cost. I excel at managing modern data platforms like Snowflake and DBT, optimizing queries to reduce execution times and costs, and delivering business-ready data for analytics, ML models, and critical reporting. With a strong SQL and NoSQL foundation, I ensure data quality, governance, and reliability across healthcare, financial services, and enterprise applications, while adhering to security and regulatory requirements such as HIPAA.

Vamsee Chilukuri

I am a Data Engineer with 3+ years of experience, focused on building end-to-end data pipelines that unlock actionable business insights. I design robust, scalable data solutions and have hands-on experience with Apache Airflow, Databricks, and AWS Glue to ingest data from APIs and relational sources. I enjoy translating complex data into accessible analytics for stakeholders and continuously optimizing pipelines for performance and cost. I excel at managing modern data platforms like Snowflake and DBT, optimizing queries to reduce execution times and costs, and delivering business-ready data for analytics, ML models, and critical reporting. With a strong SQL and NoSQL foundation, I ensure data quality, governance, and reliability across healthcare, financial services, and enterprise applications, while adhering to security and regulatory requirements such as HIPAA.

Available to hire

I am a Data Engineer with 3+ years of experience, focused on building end-to-end data pipelines that unlock actionable business insights. I design robust, scalable data solutions and have hands-on experience with Apache Airflow, Databricks, and AWS Glue to ingest data from APIs and relational sources. I enjoy translating complex data into accessible analytics for stakeholders and continuously optimizing pipelines for performance and cost.

I excel at managing modern data platforms like Snowflake and DBT, optimizing queries to reduce execution times and costs, and delivering business-ready data for analytics, ML models, and critical reporting. With a strong SQL and NoSQL foundation, I ensure data quality, governance, and reliability across healthcare, financial services, and enterprise applications, while adhering to security and regulatory requirements such as HIPAA.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
See more

Language

English
Fluent
Hindi
Intermediate

Work Experience

Data Engineer at Humana, USA
September 1, 2024 - Present
Built and managed data processing systems on AWS (EMR, S3) and Databricks using PySpark and Python to handle weekly healthcare claims, membership, and clinical data (multi-terabyte volumes). Converted raw data into structured tables for analytics and regulatory reporting. Automated ETL pipelines using Google Cloud Dataflow and Cloud Functions, replacing manual steps and saving 45+ person-hours/year while maintaining 100% data accuracy for compliance and financial reports. Designed and maintained the enterprise data warehouse in Snowflake, improving data models with clustering and materialized views to reduce average query times for 160 users. Implemented IaC with Terraform and GitLab CI/CD to deploy data platforms, cutting setup time from 3 days to under 8 hours while ensuring HIPAA compliance. Built real-time data streaming pipelines with Apache Kafka and Google Cloud Pub/Sub to ingest and process healthcare events supporting fraud detection and operational applications. Created a dat
Data Engineer I at KPIT Technologies, India
February 1, 2022 - July 1, 2023
Built and maintained data pipelines on Microsoft Azure (Data Factory, Blob Storage) and Talend to unify data from multiple sources for client reporting and analytics. Improved pipeline speed by 45% through parallel task execution, reducing data delays. Set up a real-time data processing system using Apache Flink for time-sensitive operations. Implemented data quality checks and validation rules, automating these within daily data flows and reducing data errors by over 90%. Wrote Python scripts to automate routine data validation and cleanup tasks. Built Power BI dashboards to visualize metrics like customer retention and operational trends. Enhanced system stability for distributed workloads by managing Apache Zookeeper for cluster coordination.

Education

Master of Science in Data Science at University of North Carolina, Charlotte
January 11, 2030 - July 2, 2026
Bachelor in Computer Science at SRM University, Amravati, India
January 11, 2030 - July 2, 2026

Qualifications

Python for Data Science
January 11, 2030 - July 2, 2026

Industry Experience

Healthcare, Software & Internet, Professional Services