I'm Vamshi, an AI/ML Engineer focused on building production-scale ML and generative AI systems. I specialize in LLM applications, retrieval-augmented generation (RAG), and agentic AI platforms deployed across enterprise environments. I collaborate across ML research, data engineering, infrastructure, and product teams to deliver secure, compliant, and high-impact AI solutions.

Vamshi

I'm Vamshi, an AI/ML Engineer focused on building production-scale ML and generative AI systems. I specialize in LLM applications, retrieval-augmented generation (RAG), and agentic AI platforms deployed across enterprise environments. I collaborate across ML research, data engineering, infrastructure, and product teams to deliver secure, compliant, and high-impact AI solutions.

Available to hire

I’m Vamshi, an AI/ML Engineer focused on building production-scale ML and generative AI systems. I specialize in LLM applications, retrieval-augmented generation (RAG), and agentic AI platforms deployed across enterprise environments.

I collaborate across ML research, data engineering, infrastructure, and product teams to deliver secure, compliant, and high-impact AI solutions.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert

Language

English
Fluent

Work Experience

AI/ML Engineer at Walmart
October 1, 2024 - Present
Built and deployed an enterprise Retail Operations Knowledge Assistant using LangChain, LangGraph, OpenAI GPT-4, and Pinecone. Implemented hybrid retrieval pipelines (vector + keyword + metadata) to boost response relevance by 35% for policies, manuals, and product knowledge. Developed scalable RAG pipelines using embedding models and SentenceTransformers; designed tool-calling AI agents to interact with internal APIs and support operational workflows. Established end-to-end LLM evaluation and monitoring using RAGAS and LangSmith, and built real-time document ingestion and vector index updates with Apache Spark, Airflow, and Kafka. Deployed high-throughput inference services on Kubernetes with Docker, including autoscaling and rate limiting; observed LLM behavior with Arize Phoenix and Prometheus/Grafana.
Software Engineer at Mindtree
June 1, 2019 - April 1, 2022
Led real-time ML inference services for personalization and ranking, delivering 30-40% faster end-to-end latency. Built backend services using Java (Spring Boot) and Python (FastAPI) to expose PyTorch-based inference endpoints via gRPC and REST. Scaled PyTorch model serving across distributed microservice clusters supporting millions of predictions, contributing to a 20% improvement in recommendation relevance. Designed event-driven data pipelines using Apache Kafka with Celery for asynchronous processing and real-time feature distribution. Integrated image and video processing ML workflows to automate moderation and tagging, reducing manual review workload by 50% and improving safety classification. Containerized with Docker and Kubernetes; contributed to Helm-based deployments and autoscaling. Built observability with Prometheus, Grafana, and OpenTelemetry; maintained production uptime above 92%. Collaborated with ML research, data engineering, and product teams to operationalize mod

Education

Master of Science at Campbellsville University
August 1, 2022 - May 1, 2024

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Retail, Professional Services