Software Engineer with 6+ years of experience spanning Machine Learning, backend systems, and Generative AI, with strong fundamentals in DSA and scalable system design. Experienced in building end-to-end ML/GenAI pipelines—from model training and fine-tuning to retrieval, orchestration, benchmarking, and production deployment. Has led architecture and delivery across speech and LLM stacks (RAG, prompt-injection/context selection), ML microservices on AWS, and large-scale data/feature pipelines. Also delivered practical ML products such as test-scenario mapping, risk modeling, lead-time prediction, and multi-mandate real-estate chatbots with measurable improvements in quality and performance.

Saurabh Singh

Software Engineer with 6+ years of experience spanning Machine Learning, backend systems, and Generative AI, with strong fundamentals in DSA and scalable system design. Experienced in building end-to-end ML/GenAI pipelines—from model training and fine-tuning to retrieval, orchestration, benchmarking, and production deployment. Has led architecture and delivery across speech and LLM stacks (RAG, prompt-injection/context selection), ML microservices on AWS, and large-scale data/feature pipelines. Also delivered practical ML products such as test-scenario mapping, risk modeling, lead-time prediction, and multi-mandate real-estate chatbots with measurable improvements in quality and performance.

Available to hire

Software Engineer with 6+ years of experience spanning Machine Learning, backend systems, and Generative AI, with strong fundamentals in DSA and scalable system design. Experienced in building end-to-end ML/GenAI pipelines—from model training and fine-tuning to retrieval, orchestration, benchmarking, and production deployment.

Has led architecture and delivery across speech and LLM stacks (RAG, prompt-injection/context selection), ML microservices on AWS, and large-scale data/feature pipelines. Also delivered practical ML products such as test-scenario mapping, risk modeling, lead-time prediction, and multi-mandate real-estate chatbots with measurable improvements in quality and performance.

See more

Experience Level

Expert
Expert
Intermediate
Intermediate
Intermediate

Work Experience

Founding Engineer – Machine Learning at Baeru
December 1, 2025 - Present
Architected core ML infrastructure for a multilingual voice-agent stack covering speech models, LLM inference, retrieval, agent orchestration, and production deployment. Fine-tuned Meta SeamlessM4T v2 Large using PEFT-LoRA with DDP for short-form Kannada speech recognition, improving Sentence Similarity Score from 51% to 84% for the target voice-agent workload. Fine-tuned Qwen2.5-3B using PEFT-LoRA and SFT for domain-specific conversational/agentic tasks. Built dynamic prompt-injection and context-selection using an embedding-based ML model to reduce average LLM context length by 40%+ while retrieving only relevant context. Implemented benchmarking pipelines for LLMs across quantization settings (latency/throughput/GPU memory/quality/serving cost) and benchmarked vector databases for RAG using p95/p99 latency and retrieval metrics; added MMR and re-ranking for better relevance/diversity. Built and operated ML microservices on AWS EC2 for speech processing and LLM-based services.
Senior Software Engineer at Nucleus Software
April 1, 2025 - October 1, 2025
Architected, implemented, and deployed a test scenario–mapping service and knowledge base using FAISS, improving test-case traceability to 94% and optimizing query performance with HNSW to ~5ms on CPU for a 10M record dataset. Led a team of 5 engineers to develop and deploy a pre-delinquency risk model predicting delinquency probability, applying best engineering practices with Spark ML. Implemented a Kafka pub/sub layer to fan out events to multiple ML models (pre-delinquency, credit decisioning, upsell).
Senior Software Engineer at Anarock Technology, Gurgaon, India
May 1, 2022 - January 1, 2025
Led development of an LLM-based context-aware real-estate chatbot using RAG (Elasticsearch Vector DB + LangChain), including web-scraper/data ingestion and LLM tooling (prompt templates, evaluators, guardrails) for reliable answers over 1000 mandates. Developed a supervised XGBoost model for predicting optimal contact times for sales leads and exposed it as a FastAPI microservice, increasing call pick-up rate from 30% to 45% (50% lift) in a 3-month A/B test. Designed and productionized Kubeflow-orchestrated ingestion pipelines across MongoDB aggregation, SQL, and Elasticsearch. Created real-time dashboards for accuracy and drift metrics and KPIs. Deployed MLflow on AWS EC2 with S3 for versioned artifacts and RDS for metadata.
Senior Software Engineer at HSBC Technology India
September 1, 2020 - May 1, 2022
Designed and deployed a multi-class NLP pipeline (TF-IDF + XGBoost) to auto-classify support tickets, achieving 98% classification accuracy. Deployed models as REST endpoints to enable inference for 50k+ tickets/day. Automated monthly retraining, improving average accuracy from 92% to 97% over 6 months.

Education

Bachelor of Engineering (B.E.) at Netaji Subhas Institute of Technology (NSIT), New Delhi
August 1, 2016 - May 1, 2020

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Financial Services, Real Estate & Construction, Computers & Electronics

Experience Level

Expert
Expert
Intermediate
Intermediate
Intermediate

Hire a AI Engineer

We have the best ai engineer experts on Twine. Hire a ai engineer today.