I’m an AI Engineer focused on Generative AI and LLM systems, especially RAG and hybrid retrieval architectures that work reliably in production. I enjoy building end-to-end AI services—from designing retrieval and orchestration pipelines to implementing async FastAPI backends, monitoring cost/latency, and shipping deployments on AWS. Most recently, I’ve been building a multilingual healthcare RAG platform that ingests and indexes continuously growing medical content and delivers grounded question answering using hybrid retrieval strategies. I also designed a cost-and-latency optimized LLM routing platform and an enterprise hybrid RAG system with reranking and conversational retrieval, so I can improve answer quality while keeping performance and operational overhead under control.

Manibala Balu

I’m an AI Engineer focused on Generative AI and LLM systems, especially RAG and hybrid retrieval architectures that work reliably in production. I enjoy building end-to-end AI services—from designing retrieval and orchestration pipelines to implementing async FastAPI backends, monitoring cost/latency, and shipping deployments on AWS. Most recently, I’ve been building a multilingual healthcare RAG platform that ingests and indexes continuously growing medical content and delivers grounded question answering using hybrid retrieval strategies. I also designed a cost-and-latency optimized LLM routing platform and an enterprise hybrid RAG system with reranking and conversational retrieval, so I can improve answer quality while keeping performance and operational overhead under control.

Available to hire

I’m an AI Engineer focused on Generative AI and LLM systems, especially RAG and hybrid retrieval architectures that work reliably in production. I enjoy building end-to-end AI services—from designing retrieval and orchestration pipelines to implementing async FastAPI backends, monitoring cost/latency, and shipping deployments on AWS.

Most recently, I’ve been building a multilingual healthcare RAG platform that ingests and indexes continuously growing medical content and delivers grounded question answering using hybrid retrieval strategies. I also designed a cost-and-latency optimized LLM routing platform and an enterprise hybrid RAG system with reranking and conversational retrieval, so I can improve answer quality while keeping performance and operational overhead under control.

See more

Language

English
Advanced

Work Experience

AI Engineer — Client Project (Med-ed.ai) at Med-ed.ai
January 1, 2026 - Present
Owned retrieval and backend architecture for a multilingual healthcare RAG platform providing medically grounded question answering across 5,000+ MedlinePlus pages and 20k+ indexed healthcare documents. Built async extraction pipelines to ingest healthcare articles, genetics references, and multimedia for continuous growth. Implemented hybrid retrieval using FAISS vector search, metadata filtering, and sentence-transformer embeddings; improved relevance during internal evaluation. Refactored the backend into modular retrieval, extraction, orchestration, and generation services to reduce coupling and speed debugging. Delivered multilingual conversational APIs with multi-turn memory and context-aware handling. Hardened FastAPI services with async processing, structured logging, retry logic, and input validation. Deployed Dockerized services on AWS Lambda/API Gateway with monitoring for latency, failed requests, and token usage to provide cost/performance visibility.
AI Engineer — Client Project at Med-ed.ai
January 1, 2026 - Present
Own the retrieval and backend architecture for a multilingual healthcare RAG platform providing medically grounded question answering. Built async ingestion/extraction pipelines to continuously ingest healthcare content (articles, genetics references, and multimedia) for semantic retrieval. Implemented hybrid retrieval combining FAISS vector search, metadata filtering, and sentence-transformer embeddings to improve answer relevance. Structured the backend into modular services (extraction, retrieval, orchestration, generation) to reduce coupling and debugging time. Delivered multilingual conversational APIs with multi-turn memory and context-aware interaction. Hardened FastAPI services with async handling, structured logging, retry logic, and input validation. Deployed Dockerized services on AWS Lambda and API Gateway with monitoring for latency, failed requests, and token usage to provide cost/performance visibility in production.
Generative AI Engineer — Cost & Latency Optimized LLM Platform at github.com/mithra06/Cost-Latency-LLM-Platform
September 1, 2025 - January 1, 2026
Designed and built an LLM routing platform that dynamically selects models based on latency, token cost, and query complexity using Python, FastAPI, LangChain, AWS Bedrock, and FAISS. Implemented semantic caching with embeddings and FAISS search to reduce repeat inference requests in testing. Added prompt optimization and context-compression pipelines to lower average token usage while preserving response quality. Built async, batched inference services to improve concurrency and reduce latency under load. Added structured logging, monitoring, retry handling, and health checks to make failure modes debuggable. Set up Docker-based deployment with GitHub Actions CI/CD for automated testing and deployment validation.
Generative AI Engineer — Enterprise Hybrid RAG at github.com/mithra06/Enterprise-Hybrid-RAG
March 1, 2025 - August 1, 2025
Built a hybrid RAG system for document Q&A using LangChain, FastAPI, FAISS, BM25, and cross-encoder reranking to address the precision limitations of vector-only retrieval. Improved retrieval precision by combining embeddings-based retrieval with reranking and validated gains via direct comparisons. Created ingestion pipelines for PDFs/SOPs/support documents using adaptive chunking and metadata enrichment for scalable indexing. Developed conversational retrieval APIs featuring query rewriting, semantic caching, and context compression for multi-turn document Q&A. Tuned chunking, retrieval filtering, and embeddings to reduce retrieval latency across large document collections.
Data Analyst (Volunteer) at Get2Know
June 1, 2023 - May 1, 2024
Built machine learning models for financial analytics and investment forecasting using Python data pipelines. Applied customer segmentation and behavioral clustering to support retention analysis. Partnered with stakeholders to translate financial reporting requirements into scalable analytics workflows. Improved cloud-based data pipelines on GCP, reducing processing time by ~40% for analytics workloads.
Data Analyst — Volunteer at Get2Know
June 1, 2023 - May 1, 2024
Built a machine learning model for financial analytics and investment forecasting using Python data pipelines. Applied customer segmentation and behavioral clustering to support retention analysis. Collaborated with stakeholders to translate financial reporting requirements into scalable analytics workflows. Improved cloud-based data pipelines on GCP, reducing processing time by about 40% for analytics workloads.

Education

Master of Engineering, Computer Science & Engineering (Big Data Analytics) at Anna University, College of Engineering, Chennai, India
June 1, 2015 - June 1, 2017
B.Tech, Information Technology at RMD Engineering College, Tiruvallur, India
March 1, 2011 - April 1, 2015
Master of Engineering, Computer Science & Engineering (Big Data Analytics) at Anna University, College of Engineering, Chennai, India
June 1, 2015 - June 1, 2017
B.Tech, Information Technology at RMD Engineering College, Tiruvallur, India
March 1, 2011 - April 1, 2015

Qualifications

Add your qualifications or awards here.

Industry Experience

Healthcare, Software & Internet, Education, Professional Services