I’m Sai Teja Reddy, a Machine Learning Engineer focused on Speech AI and LLM fine-tuning, building practical systems that reliably deliver low-latency, high-quality results. I work across model training and evaluation, with hands-on experience using PyTorch and Hugging Face Transformers, and I’ve fine-tuned Gemma variants with LoRA/PEFT methods while debugging hallucinations, persona consistency, latency, and failure cases. On the engineering side, I design GPU-aware inference pipelines and distributed ML infrastructure using Kubernetes, Redis, Kafka, and dynamic batching to improve throughput and robustness during rollouts. I’ve also built end-to-end Voice AI and agentic workflows that combine ASR/TTS, retrieval-augmented context, and production-grade reliability checks—drawing from experience with Whisper, FasterWhisper, StyleTTS2, and Qwen-Audio.

SAI TEJA REDDY SHAGA

I’m Sai Teja Reddy, a Machine Learning Engineer focused on Speech AI and LLM fine-tuning, building practical systems that reliably deliver low-latency, high-quality results. I work across model training and evaluation, with hands-on experience using PyTorch and Hugging Face Transformers, and I’ve fine-tuned Gemma variants with LoRA/PEFT methods while debugging hallucinations, persona consistency, latency, and failure cases. On the engineering side, I design GPU-aware inference pipelines and distributed ML infrastructure using Kubernetes, Redis, Kafka, and dynamic batching to improve throughput and robustness during rollouts. I’ve also built end-to-end Voice AI and agentic workflows that combine ASR/TTS, retrieval-augmented context, and production-grade reliability checks—drawing from experience with Whisper, FasterWhisper, StyleTTS2, and Qwen-Audio.

Available to hire

I’m Sai Teja Reddy, a Machine Learning Engineer focused on Speech AI and LLM fine-tuning, building practical systems that reliably deliver low-latency, high-quality results. I work across model training and evaluation, with hands-on experience using PyTorch and Hugging Face Transformers, and I’ve fine-tuned Gemma variants with LoRA/PEFT methods while debugging hallucinations, persona consistency, latency, and failure cases.

On the engineering side, I design GPU-aware inference pipelines and distributed ML infrastructure using Kubernetes, Redis, Kafka, and dynamic batching to improve throughput and robustness during rollouts. I’ve also built end-to-end Voice AI and agentic workflows that combine ASR/TTS, retrieval-augmented context, and production-grade reliability checks—drawing from experience with Whisper, FasterWhisper, StyleTTS2, and Qwen-Audio.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Beginner
See more

Language

English
Advanced

Work Experience

Applied AI Engineer at Machines & Minds
February 1, 2026 - Present
Designed an MCP-style AI integration architecture for agentic crypto workflows, connecting LLM orchestration, retrieval layers, backend APIs, structured tool responses, fallback handling, and frontend components. Built Python-based AI workflows that integrate LLM APIs and retrieval context with wallet/transaction data and reliability checks to reduce hallucinations and improve response quality.
Machine Learning Engineer at Avtar Inc.
January 1, 2025 - January 1, 2026
Researched and prototyped an end-to-end Voice AI pipeline for a personalized avatar product combining ASR, persona-driven LLMs, retrieval, and expressive TTS for real-time conversational experiences. Benchmarked multiple streaming ASR engines (Whisper.cpp, WhisperLive, WhisperX, FasterWhisper), then built a FasterWhisper-based prototype with in-memory audio processing, partial transcription, GPU/INT8 inference, and LiveKit integration. Fine-tuned Gemma 1B–27B using LoRA/LoRA+/DoRA/LoHa/LoKr and experimented with StyleTTS2 fine-tuning and speaker cloning, including StyleTalker-style component replication via trainable projection layers using Qwen-Audio embeddings for multimodal conditioning.
Software Developer (Co-op) at Intuit
January 1, 2024 - May 1, 2024
Built Java/Spring Boot microservices and REST APIs processing 500K+ daily transactions, improving response time by 35%. Wrote JUnit/Mockito tests reaching 82% coverage and automated CI/CD with GitHub Actions and Docker across 15+ releases. Resolved ~200 application and SQL issues per month with ~90% same-day closure; performed root-cause analysis and reduced recurrence by 40%.

Education

B.S. Computer Science (AI/ML Concentration) at University of North Texas
January 11, 2030 - August 29, 2026

Qualifications

Claude Certified Architect
January 11, 2030 - August 29, 2026
Micro1 Certified ML Engineer
January 11, 2030 - August 29, 2026
Supervised Machine Learning (DeepLearning.AI - Andrew Ng)
January 11, 2030 - August 29, 2026

Industry Experience

Software & Internet, Computers & Electronics, Media & Entertainment, Professional Services, Telecommunications

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Beginner
See more

Hire a AI Engineer

We have the best ai engineer experts on Twine. Hire a ai engineer today.