Machine Learning Engineer with 5+ years of experience designing and deploying production-scale AI/ML systems, multimodal applications, and enterprise LLM platforms. Expertise in RAG, RLHF, real-time inference, PyTorch, Hugging Face, and GPU-optimized pipelines, delivering high-performance, low-latency AI solutions. Strong background in MLOps, distributed systems, Kubernetes, AWS, and Azure, building scalable cloud-native architectures that improve model accuracy, reliability, and operational efficiency across fintech, AdTech, and enterprise AI environments.

Sumukh Ramagiri

Machine Learning Engineer with 5+ years of experience designing and deploying production-scale AI/ML systems, multimodal applications, and enterprise LLM platforms. Expertise in RAG, RLHF, real-time inference, PyTorch, Hugging Face, and GPU-optimized pipelines, delivering high-performance, low-latency AI solutions. Strong background in MLOps, distributed systems, Kubernetes, AWS, and Azure, building scalable cloud-native architectures that improve model accuracy, reliability, and operational efficiency across fintech, AdTech, and enterprise AI environments.

Available to hire

Machine Learning Engineer with 5+ years of experience designing and deploying production-scale AI/ML systems, multimodal applications, and enterprise LLM platforms. Expertise in RAG, RLHF, real-time inference, PyTorch, Hugging Face, and GPU-optimized pipelines, delivering high-performance, low-latency AI solutions.

Strong background in MLOps, distributed systems, Kubernetes, AWS, and Azure, building scalable cloud-native architectures that improve model accuracy, reliability, and operational efficiency across fintech, AdTech, and enterprise AI environments.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
See more

Work Experience

Machine Learning Engineer at Scale AI
March 1, 2025 - Present
Developed an AI-driven assistant embedded within admin workflows, reducing merchant task completion time by 35% via automated multi-step execution. Redesigned distributed multimodal preprocessing pipelines (video/audio/text) to increase ingestion throughput by 45% using optimized batching, deduplication, and metadata indexing. Built scalable LLM benchmarking with automated regression testing, prompt/version control, and distributed execution using Ray and containerized GPU inference, reducing evaluation runtime by 38%. Implemented active learning sampling and uncertainty scoring to improve human labeling QA precision by 27%, including model-assisted review, gold-set validation, and automated dispute resolution. Engineered end-to-end enterprise RAG (hybrid retrieval, reranking, orchestration, caching, evaluation) for low-latency workflows. Built RLHF-ready preference pipelines integrating human ranking and reward datasets with dataset versioning, lineage tracking, and reproducible train
Machine Learning Engineer at NVIDIA
January 1, 2024 - February 28, 2025
Optimized TensorRT-LLM pipelines and GPU memory utilization, improving production NIM inference throughput by 47% while reducing p99 latency by 38% through Kubernetes autoscaling. Deployed real-time embedding retrieval, reranking pipelines, and feature-store-backed decision services on GPU clusters, increasing ad personalization CTR by 29% and revenue yield by 18%. Implemented graph-based anomaly detection and calibrated risk models with low-latency microservices, reducing payment fraud losses by 34% while improving chargeback prediction precision by 22%. Architected NIM-based multi-tenant inference services using Triton Inference Server, TensorRT-LLM, Kubernetes, Docker, and gRPC APIs with strict SLO governance. Built real-time personalization pipelines integrating Kafka, Redis feature stores, vector databases, embedding services, and cross-encoder rerankers for millisecond-level recommendation decisioning. Led production MLOps using MLflow, CI/CD, evaluation automation, drift monitor
Software Engineer at Accenture
May 1, 2019 - December 31, 2022
Built high-performance Python backend services integrated with Kafka streaming and rule-based validation frameworks, improving fraud detection accuracy by 34% while reducing false positives by 22%. Optimized REST APIs, added Redis caching, and configured Kubernetes autoscaling to reduce transaction decision latency by 41% and increase approval throughput by 18%. Developed React-based operational dashboards and automated case management workflows with secure API-driven task prioritization, increasing investigator productivity by 29% and reducing manual review volume by 37%. Engineered real-time transaction processing platforms using FastAPI, PyTorch, Kafka, Apache Flink, Redis, and PostgreSQL microservices, ensuring scalable sub-100ms API reliability. Designed CI/CD workflows with GitHub Actions, Docker, and Kubernetes enabling automated builds, container orchestration, blue-green deployments, and high availability for regulated BFSI. Architected secure transaction microservices on Azur

Education

Master of Business Analytics at University of New Haven
January 11, 2030 - August 20, 2026
Master of Business Analytics at University of New Haven
January 11, 2030 - August 26, 2026

Qualifications

AWS Certified Machine Learning Engineer - Associate
January 11, 2030 - August 26, 2026

Industry Experience

Financial Services, Software & Internet