I’m an AI/ML Engineer specializing in production-grade AI solutions, LLM applications, RAG systems, ML pipelines, and scalable machine learning infrastructure. I have experience building and deploying AI systems using Python, PyTorch, TensorFlow, LangChain, LangGraph, FastAPI, Docker, Kubernetes, and cloud technologies. I focus on turning complex business problems into reliable, measurable AI solutions, with hands-on experience in LLM serving, model optimization, data pipelines, and MLOps. What sets me apart is my ability to combine strong ML engineering with practical software engineering to build AI products that are scalable, efficient, and production-ready.

Manikanteswar Gandrothula

I’m an AI/ML Engineer specializing in production-grade AI solutions, LLM applications, RAG systems, ML pipelines, and scalable machine learning infrastructure. I have experience building and deploying AI systems using Python, PyTorch, TensorFlow, LangChain, LangGraph, FastAPI, Docker, Kubernetes, and cloud technologies. I focus on turning complex business problems into reliable, measurable AI solutions, with hands-on experience in LLM serving, model optimization, data pipelines, and MLOps. What sets me apart is my ability to combine strong ML engineering with practical software engineering to build AI products that are scalable, efficient, and production-ready.

Available to hire

I’m an AI/ML Engineer specializing in production-grade AI solutions, LLM applications, RAG systems, ML pipelines, and scalable machine learning infrastructure. I have experience building and deploying AI systems using Python, PyTorch, TensorFlow, LangChain, LangGraph, FastAPI, Docker, Kubernetes, and cloud technologies. I focus on turning complex business problems into reliable, measurable AI solutions, with hands-on experience in LLM serving, model optimization, data pipelines, and MLOps. What sets me apart is my ability to combine strong ML engineering with practical software engineering to build AI products that are scalable, efficient, and production-ready.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
See more

Language

Work Experience

AI/ML Engineer at Meta
June 1, 2025 - Present
Owned engineering of Meta's internal LLM serving platform, including reusable inference APIs, model versioning, traffic routing, and deployment workflows for Llama 3 models using PyTorch, vLLM, Kubernetes, Docker, and FastAPI to support 40+ internal applications. Built a shared RAG platform using LangChain/LangGraph with ScaNN and hybrid retrieval for source-grounded responses, reducing lookup time by 45%. Designed distributed fine-tuning pipelines using PyTorch FSDP/DeepSpeed, CUDA, mixed precision, and Hugging Face Transformers, reducing turnaround from 9 days to under 5 days. Implemented multimodal embedding pipelines combining text/images/videos/interaction signals to improve semantic retrieval. Established automated LLM evaluation using Llama Guard, RoBERTa, LangSmith, MLflow, and internal benchmarks for factuality, retrieval quality, latency, and policy compliance (95%+ automated coverage). Built high-throughput feature generation using TorchRec with Spark/Python/SQL for large-sc
Machines Learning Scientist at Sentara Health
September 1, 2024 - May 1, 2025
Developed deep learning models for 30-day readmission, patient deterioration, and length-of-stay prediction using TensorFlow and ClinicalBERT with structured EHR features and longitudinal records, supporting clinical risk stratification across 2M+ encounters. Implemented AI-assisted clinical documentation workflows using GPT-4, Claude 3.5, and Llama 3 to draft encounter summaries and discharge instructions with clinician review. Built clinical NLP pipelines with ClinicalBERT/BioBERT to extract structured clinical entities from physician notes, increasing structured data availability by 32%. Built deployment workflows using Kubernetes, FastAPI, and MLflow to support 15+ production AI services with standardized releases and monitoring. Delivered HIPAA-compliant AI solutions enabling 400+ clinicians to access patient information more efficiently through AI-assisted workflows.
Machine Learning Engineer at Accenture
January 1, 2021 - August 1, 2023
Designed probability of default (PD) and credit risk models using Python, XGBoost, LightGBM, Scikit-learn, SQL, and PySpark for consumer lending portfolios with 8M+ customers. Built a standardized feature engineering framework consolidating credit bureau data, repayment history, transaction behavior, income verification, and customer attributes into governed datasets consumed by multiple risk models. Automated batch model scoring/validation with Airflow, AWS Glue, Spark, and MLflow, processing 1.5M+ loan applications monthly for inference, validation, and regulatory reporting. Productionized ML models as containerized REST services using Docker, Kubernetes, AWS SageMaker/FastAPI/EC2/Lambda with online inference latency below 250 ms integrated into loan origination systems. Implemented MLOps for versioning, drift detection, performance monitoring, and automated retraining using MLflow, Prometheus, Grafana, and CloudWatch, improving reliability. Reduced end-to-end training time by ~7 hou

Education

Master of Science in Data Analytics Engineering at George Mason University
August 1, 2023 - May 1, 2025
Master of Science in Data Analytics Engineering at George Mason University
August 1, 2023 - May 1, 2025

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Healthcare, Professional Services, Software & Internet