AI Data Engineer with 6+ years of experience owning production GenAI systems end to end—from data pipeline architecture and structured extraction to monitored deployments. Built and shipped LLM-based enrichment services on Amazon Bedrock, OpenAI APIs, and LangChain, integrating RAG, function calling, and PyTorch model inference with strong data engineering, reliability, and observability practices.

Vaibhav S

AI Data Engineer with 6+ years of experience owning production GenAI systems end to end—from data pipeline architecture and structured extraction to monitored deployments. Built and shipped LLM-based enrichment services on Amazon Bedrock, OpenAI APIs, and LangChain, integrating RAG, function calling, and PyTorch model inference with strong data engineering, reliability, and observability practices.

Available to hire

AI Data Engineer with 6+ years of experience owning production GenAI systems end to end—from data pipeline architecture and structured extraction to monitored deployments.
Built and shipped LLM-based enrichment services on Amazon Bedrock, OpenAI APIs, and LangChain, integrating RAG, function calling, and PyTorch model inference with strong data engineering, reliability, and observability practices.

See more

Experience Level

Language

Work Experience

AI Data Engineering Focus, Data Engineer at Picket Homes, USA
January 1, 2024 - Present
Owned the AI data engineering roadmap for property enrichment, architecting ingestion pipelines on AWS Glue, EMR, and Snowflake to unify listings, transactions, property events, document/image metadata, and user activity for downstream LLM enrichment. Built resilient Python ingestion frameworks (SQLAlchemy, Pandas) with schema validation and error handling to normalize REST/DB/Kinesis/third-party feeds, reducing onboarding time from 2 weeks to under 3 days. Engineered Bronze/Silver/Gold star/medallion models with SCD Type 2 and conformed keys so LLM-extracted fields land in governed, analytics-ready tables. Shipped the core AI enrichment service using Bedrock/OpenAI with LangChain prompt templates, few-shot examples, and structured JSON/function-calling outputs for automated tagging and extraction. Integrated PyTorch embedding/classification models to score document quality and flag anomalies, and implemented retry logic, provider fallback, and output validation for accuracy/cost effic
Data Engineer at CGI, India
January 1, 2018 - December 31, 2021
Delivered Azure and Snowflake data engineering/warehousing solutions for banking and healthcare clients in GDPR/HIPAA/SOX environments, designing dimensional models (star schema, SCD Type 1/2) that reduced report generation time by 35%. Built Python ETL workflows (Pandas, SQLAlchemy) to ingest, cleanse, and transform data from Oracle, SQL Server, flat files, MongoDB, and ADLS Gen2 into Synapse and Snowflake, reducing onboarding time by 40%. Developed Azure Data Factory pipelines with reusable parameterized templates and incremental loads, reducing development effort by 30%. Implemented and optimized PySpark jobs on Azure Databricks (broadcast joins, partition pruning, Delta Lake ACID writes with medallion layering) and authored dbt models with incremental strategies and schema tests, improving processing speed by 40–50% and refresh performance by 25%. Contributed to migration of 30+ on-prem ETL jobs to Azure/Snowflake orchestrated with Airflow, reducing maintenance overhead by 20%.

Education

Master of Science, Computer Science at East Texas A&M University, Commerce, TX
January 1, 2022 - December 31, 2023

Qualifications

AWS Certified Cloud Practitioner
January 11, 2030 - August 3, 2026
Microsoft Certified: Azure Fundamentals (AZ-900)
January 11, 2030 - August 3, 2026

Industry Experience

Financial Services, Healthcare, Real Estate & Construction, Software & Internet, Professional Services