I’m an AI/ML Engineer with 4+ years of experience building production-focused machine learning and generative AI systems for retrieval and conversational experiences. I use Python, PySpark, and XGBoost to turn behavioral data into better ranking and recommendation quality, and I extend that foundation into grounded RAG and agentic workflows. Most recently, I built retrieval pipelines and agent tool-calling experiences using LangChain and LangGraph, with vector search via pgvector, and deployed services using FastAPI on Kubernetes (GKE). I focus heavily on evaluation, reliability, and latency—measuring grounding and retrieval quality, reducing unsupported responses, and optimizing runtime performance so applications stay accurate and dependable in real production workloads.

Vengaiah Chowdary Madineni

I’m an AI/ML Engineer with 4+ years of experience building production-focused machine learning and generative AI systems for retrieval and conversational experiences. I use Python, PySpark, and XGBoost to turn behavioral data into better ranking and recommendation quality, and I extend that foundation into grounded RAG and agentic workflows. Most recently, I built retrieval pipelines and agent tool-calling experiences using LangChain and LangGraph, with vector search via pgvector, and deployed services using FastAPI on Kubernetes (GKE). I focus heavily on evaluation, reliability, and latency—measuring grounding and retrieval quality, reducing unsupported responses, and optimizing runtime performance so applications stay accurate and dependable in real production workloads.

Available to hire

I’m an AI/ML Engineer with 4+ years of experience building production-focused machine learning and generative AI systems for retrieval and conversational experiences. I use Python, PySpark, and XGBoost to turn behavioral data into better ranking and recommendation quality, and I extend that foundation into grounded RAG and agentic workflows.

Most recently, I built retrieval pipelines and agent tool-calling experiences using LangChain and LangGraph, with vector search via pgvector, and deployed services using FastAPI on Kubernetes (GKE). I focus heavily on evaluation, reliability, and latency—measuring grounding and retrieval quality, reducing unsupported responses, and optimizing runtime performance so applications stay accurate and dependable in real production workloads.

See more

Language

English
Advanced

Work Experience

AI Engineer at BigCommerce
September 1, 2024 - Present
Built a LangChain retrieval pipeline using Hugging Face embeddings and pgvector, improving relevant-context recall by 18% for curated merchant queries (products, policies, and availability). Integrated LangGraph tool-calling with BigCommerce catalog, pricing, inventory, and order APIs to generate grounded multi-step answers with validation checks prior to supported transactional actions. Deployed a FastAPI inference service on GKE using Docker with structured logging and robust fallback handling. Conducted evaluation of grounding, retrieval precision, latency, and failure cases using repeatable test sets, reducing unsupported responses by 21% via reranking, prompt refinement, and retrieval tuning. Optimized production query handling using embedding reuse, async requests, and selective context compression, cutting median response latency by 16% while maintaining answer quality.
ML Engineer at TCS
August 1, 2021 - July 31, 2023
Prepared 8.4 million retail records using Python, Pandas, and PySpark to engineer behavioral features for personalized recommendation ranking across customer segments and transaction patterns. Modeled purchase propensity using scikit-learn baselines and XGBoost, comparing ranking quality across cohorts to select robust features for the client recommendation pipeline. Tuned XGBoost ranking parameters in AWS SageMaker, improving top-k precision by 14% over a prior rules-based approach on a held-out validation dataset. Automated training-data ingestion from Amazon S3 and reusable feature-processing jobs, standardizing weekly model refreshes while reducing manual notebook dependencies. Validated feature distributions, recommendation coverage, and model drift with MLflow, reducing model refresh turnaround by 22% while preserving reproducible release comparisons. Deployed the selected recommendation model via REST APIs and maintained responses under 200 milliseconds with integration testing

Education

Master of Science (Applied Statistics and Data Science) at University of Texas at Arlington
August 1, 2023 - December 31, 2024
Bachelor of Technology at Dhanekula Institute of Engineering and Technology
August 1, 2019 - May 31, 2023
Master of Science (Applied Statistics and Data Science) at University of Texas at Arlington
August 1, 2023 - December 1, 2024
Bachelor of Technology at Dhanekula Institute of Engineering and Technology
August 1, 2019 - May 1, 2023
Master of Science, Applied Statistics and Data Science at University of Texas at Arlington
August 1, 2023 - December 1, 2024
Bachelor of Technology at Dhanekula Institute of Engineering and Technology
August 1, 2019 - May 1, 2023

Qualifications

Add your qualifications or awards here.

Industry Experience

Retail, Computers & Electronics, Professional Services, Software & Internet