AI/ML Engineer with 3 years of experience building and deploying production LLM systems (RAG, fine-tuning, evaluation, and monitoring) plus computer vision solutions across real customer support, recommendation, and safety monitoring use cases. Strong focus on turning prototypes into reliable, measurable production workflows with performance and cost/latency optimization. Experienced in AWS-based deployment, vector databases, semantic search, structured generation with schema validation, and GDPR-safe preprocessing. Built evaluation and observability frameworks to reduce hallucinations, improve retrieval quality, and increase downstream recommendation coverage and relevance.

Lokesh Madem

AI/ML Engineer with 3 years of experience building and deploying production LLM systems (RAG, fine-tuning, evaluation, and monitoring) plus computer vision solutions across real customer support, recommendation, and safety monitoring use cases. Strong focus on turning prototypes into reliable, measurable production workflows with performance and cost/latency optimization. Experienced in AWS-based deployment, vector databases, semantic search, structured generation with schema validation, and GDPR-safe preprocessing. Built evaluation and observability frameworks to reduce hallucinations, improve retrieval quality, and increase downstream recommendation coverage and relevance.

Available to hire

AI/ML Engineer with 3 years of experience building and deploying production LLM systems (RAG, fine-tuning, evaluation, and monitoring) plus computer vision solutions across real customer support, recommendation, and safety monitoring use cases. Strong focus on turning prototypes into reliable, measurable production workflows with performance and cost/latency optimization.

Experienced in AWS-based deployment, vector databases, semantic search, structured generation with schema validation, and GDPR-safe preprocessing. Built evaluation and observability frameworks to reduce hallucinations, improve retrieval quality, and increase downstream recommendation coverage and relevance.

See more

Language

English
Advanced
Hindi
Intermediate
Telugu
Intermediate

Work Experience

Machine Learning Engineer at Tubi
May 1, 2025 - Present
Engineered a production-grade LLM content understanding pipeline for recommendation systems, generating structured, spoiler-safe metadata (summaries, themes, mood, energy, audience fit, recommendation tags). Built schema-constrained generation workflows using prompt engineering, Pydantic validation, and instruction-tuned retries, increasing structured JSON reliability from ~72% to ~99% and reducing re-generation costs. Designed an LLM evaluation framework covering semantic quality, JSON validity, field coverage, spoiler-free rate, latency, and inference cost to benchmark multiple model families for production readiness. Integrated LLM-generated embeddings into two-tower recommendation workflows, improving long-tail recall (~8–12%) and cold-start coverage (~15%) in offline evaluation. Developed a search query understanding agent to convert natural language into structured JSON intents and semantic search representations for hybrid SQL + vector retrieval.
AI Engineer at PayPal
June 1, 2024 - April 1, 2025
Built a production customer-support RAG system using LLMs, all-MiniLM embeddings, and Pinecone vector search to produce grounded, context-aware responses for financial support queries. Optimized multi-format retrieval with recursive chunking and metadata-aware indexing, achieving <150ms top-K retrieval latency and >92% semantic relevance in internal evaluations. Fine-tuned Llama models using LoRA on ~9K curated support dialogues, improving human-judged correctness by ~10% and reducing hallucination rates by ~15% across a 500-query benchmark. Deployed real-time RAG inference services on AWS (Lambda, SageMaker, ECS, Step Functions) for scalable production workloads. Implemented GDPR-compliant PII masking via spaCy NER and regex redaction to support safer embedding and response generation for sensitive data.
AI Engineer at L & T Technology Service
August 1, 2021 - July 1, 2022
Developed a production computer vision system for real-time helmet compliance detection from CCTV streams using YOLO and Faster R-CNN. Trained models on ~23K labeled images using class balancing, augmentation, and preprocessing to improve robustness and violation recall. Created validation and monitoring workflows tracking mAP, precision, recall, and F1-score to accelerate issue detection during iterations. Deployed edge-to-cloud inference with Docker, Kubernetes, AWS, NVIDIA Jetson, and TensorRT to support multi-camera PPE monitoring with automated alerts, sub-50ms latency, and ~99% uptime.

Education

Master of Science in Artificial Intelligence at University of North Texas
August 1, 2022 - May 1, 2024
Bachelor of Technology in Computer Science at Andhra University
July 1, 2018 - May 1, 2022
Master of Science (Artificial Intelligence) at University of North Texas
August 1, 2022 - May 1, 2024
Bachelor of Technology (Computer Science) at Andhra University
July 1, 2018 - May 1, 2022

Qualifications

Neural Network & Deep Learning – Coursera
January 1, 2020 - July 23, 2026
Improving Deep Neural Network: Hyperparameter Tuning, Regularisation & Optimisation - Coursera
January 1, 2020 - July 23, 2026
Python for Everybody Specialisation – Coursera
January 1, 2020 - July 23, 2026
Neural Network & Deep Learning
January 11, 2030 - July 23, 2026
Improving Deep Neural Network: Hyperparameter Tuning, Regularisation & Optimisation
January 11, 2030 - July 23, 2026
Python for Everybody Specialisation
January 11, 2030 - July 23, 2026

Industry Experience

Software & Internet, Financial Services, Professional Services, Computers & Electronics, Education