Hi, I'm Nagendra Varaprasad Bathula, an AI/ML Engineer with 5+ years of hands-on experience building production-grade AI/ML and LLM-powered systems across large-scale recommendations and search platforms. I’ve consistently improved engagement, latency, and system efficiency through RAG pipelines, generative retrieval, and optimized inference. I thrive on turning data into business impact, partnering with product, data science, and engineering teams to design scalable, reliable ML solutions deployed on cloud-native platforms using Docker, Kubernetes, and CI/CD workflows.

Nagendra Varaprasad Bathula

Hi, I'm Nagendra Varaprasad Bathula, an AI/ML Engineer with 5+ years of hands-on experience building production-grade AI/ML and LLM-powered systems across large-scale recommendations and search platforms. I’ve consistently improved engagement, latency, and system efficiency through RAG pipelines, generative retrieval, and optimized inference. I thrive on turning data into business impact, partnering with product, data science, and engineering teams to design scalable, reliable ML solutions deployed on cloud-native platforms using Docker, Kubernetes, and CI/CD workflows.

Available to hire

Hi, I’m Nagendra Varaprasad Bathula, an AI/ML Engineer with 5+ years of hands-on experience building production-grade AI/ML and LLM-powered systems across large-scale recommendations and search platforms. I’ve consistently improved engagement, latency, and system efficiency through RAG pipelines, generative retrieval, and optimized inference.

I thrive on turning data into business impact, partnering with product, data science, and engineering teams to design scalable, reliable ML solutions deployed on cloud-native platforms using Docker, Kubernetes, and CI/CD workflows.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
See more

Language

English
Fluent

Work Experience

AI/ML Engineer at Pinterest
May 1, 2024 - Present
Led the development of a production-scale AI recommendation platform leveraging transformer-based retrieval, user behavior modeling, and LLM-assisted ranking strategies, increasing user engagement by 34% across personalized content experiences. Architected retrieval-augmented workflows with semantic embeddings, vector search, and contextual ranking models, improving content discovery and reducing churn by 27%. Designed candidate generation pipelines using PyTorch-based transformers trained on billions of interactions, improving retrieval relevance and diversity. Built hybrid retrieval infrastructure with FAISS, ANN indexing, and distributed storage for low-latency semantic search and high-throughput serving. Developed LLM-powered personalization services and scalable model-serving with Triton Inference Server, dynamic batching, and KV caching, reducing end-to-end latency by 31%. Implemented microservices-based AI platform using Python, FastAPI, and gRPC; built production-grade RAG work
AI/ML Engineer at Pinterest CA
May 1, 2024 - Present
Led development of a production-scale AI recommendation platform leveraging transformer-based retrieval, user behavior modeling, and LLM-assisted ranking strategies, increasing user engagement by 34% across personalized content experiences. Architected retrieval-augmented recommendation workflows combining semantic embeddings, vector search, and contextual ranking models, improving content discovery quality while reducing user churn by 27%. Designed and optimized candidate generation pipelines using PyTorch-based transformer architectures trained on billions of user interactions, significantly improving retrieval relevance and recommendation diversity. Built hybrid retrieval infrastructure utilizing FAISS, ANN indexing, embedding models, and distributed storage systems, enabling low-latency semantic search and high-throughput recommendation serving. Developed LLM-powered personalization services that generated contextual content understanding signals and enhanced ranking decisions, imp
AI/ML Engineer at Microsoft India
January 1, 2020 - July 31, 2023
Contributed to the development of large-scale extreme classification and embedding-based retrieval systems using Python, PyTorch, and distributed training frameworks, supporting recommendation and search workloads processing billions of user interactions. Improved recommendation accuracy by 32% through training and optimizing DeepXML-based multi-label classification models, enhancing ranking quality and search relevance across millions of content items. Reduced real-time inference latency by 28% through candidate pruning, optimized batch processing, and ONNX-based model acceleration techniques, enabling low-latency predictions for large-scale production workloads. Designed scalable feature engineering pipelines using Spark, Kafka, and Python, generating behavioral features and dense embeddings from clickstream data to support machine learning models trained on petabyte-scale datasets. Leveraged Python and C++ across performance-sensitive components to improve retrieval efficiency, mode

Education

Master of Science in Data Science at University of Memphis
January 11, 2030 - June 29, 2026
Master of Science in Data Science at University of Memphis
January 11, 2030 - June 29, 2026

Qualifications

Certified Machine Learning Engineer – Associate
January 11, 2030 - June 29, 2026
AWS Solutions Architect Associate
January 11, 2030 - June 29, 2026
NVIDIA DLI Generative AI with LLMs
January 11, 2030 - June 29, 2026
Certified Machine Learning Engineer – Associate
January 11, 2030 - June 29, 2026
AWS Solutions Architect Associate
January 11, 2030 - June 29, 2026
NVIDIA DLI Generative AI with LLMs
January 11, 2030 - June 29, 2026

Industry Experience

Software & Internet, Professional Services, Media & Entertainment

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
See more

Hire a AI Engineer

We have the best ai engineer experts on Twine. Hire a ai engineer in San Francisco today.