I’m an AI/ML engineer focused on building production-grade Generative AI systems, especially Large Language Model (LLM) applications like fine-tuning, RAG, and inference optimization. Over the last few years, I’ve worked on improving throughput and latency for AI services using techniques like dynamic batching, GPU parallelism, and scalable model serving with vLLM and Triton, while maintaining high uptime for large-scale workloads. I also enjoy the full lifecycle of AI engineering—from distributed training with PyTorch FSDP/DeepSpeed to reliable MLOps automation using tools like Kubernetes, Terraform, Argo Workflows, and MLflow. I build and monitor end-to-end pipelines (streaming data, feature serving, vector search, and evaluation/benchmarking), and I’m particularly driven by turning model quality improvements—like better retrieval relevance and reduced hallucinations—into measurable operational impact.

Srinivasa Vishnu

I’m an AI/ML engineer focused on building production-grade Generative AI systems, especially Large Language Model (LLM) applications like fine-tuning, RAG, and inference optimization. Over the last few years, I’ve worked on improving throughput and latency for AI services using techniques like dynamic batching, GPU parallelism, and scalable model serving with vLLM and Triton, while maintaining high uptime for large-scale workloads. I also enjoy the full lifecycle of AI engineering—from distributed training with PyTorch FSDP/DeepSpeed to reliable MLOps automation using tools like Kubernetes, Terraform, Argo Workflows, and MLflow. I build and monitor end-to-end pipelines (streaming data, feature serving, vector search, and evaluation/benchmarking), and I’m particularly driven by turning model quality improvements—like better retrieval relevance and reduced hallucinations—into measurable operational impact.

Available to hire

I’m an AI/ML engineer focused on building production-grade Generative AI systems, especially Large Language Model (LLM) applications like fine-tuning, RAG, and inference optimization. Over the last few years, I’ve worked on improving throughput and latency for AI services using techniques like dynamic batching, GPU parallelism, and scalable model serving with vLLM and Triton, while maintaining high uptime for large-scale workloads.

I also enjoy the full lifecycle of AI engineering—from distributed training with PyTorch FSDP/DeepSpeed to reliable MLOps automation using tools like Kubernetes, Terraform, Argo Workflows, and MLflow. I build and monitor end-to-end pipelines (streaming data, feature serving, vector search, and evaluation/benchmarking), and I’m particularly driven by turning model quality improvements—like better retrieval relevance and reduced hallucinations—into measurable operational impact.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Language

Javanese
Intermediate
Aragonese
Intermediate

Work Experience

AI/ML Engineer at Meta
March 1, 2025 - Present
Led optimization initiatives improving LLM inference throughput by 65% using token streaming, dynamic batching, and GPU parallelism; reduced end-to-end latency by 45% (220ms to 120ms). Contributed to GPU inference infrastructure with 99.99% uptime, and built scalable serving with vLLM, Triton, and ONNX Runtime, increasing concurrent capacity by 3x. Developed benchmarking to improve GPU utilization by 35% and reduce inference costs by 25%. Implemented distributed training using PyTorch FSDP and DeepSpeed, reducing checkpoint overhead by 55% and training time by 30%. Built RAG systems with LangChain, LlamaIndex, and FAISS improving relevance and reducing hallucinations. Automated deployments using Kubernetes, Terraform, and Argo Workflows; built real-time feature serving using Feast and Redis (under 10ms).
Software Engineer (AI/ML Systems) at Accenture
August 1, 2021 - November 1, 2023
Built ML inference services on AWS SageMaker and Kubernetes supporting 50K+ predictions/sec across recommendation and fraud detection. Developed Kafka and Spark streaming pipelines processing 5TB+ daily, reducing delays by 40%. Implemented MLOps workflows with MLflow, Kubeflow, and Airflow reducing deployment time from days to under 4 hours. Designed GenAI solutions using Azure OpenAI, LangChain, and Pinecone enabling 60% automation adoption and 500K+ monthly conversations. Produced anomaly detection systems improving resolution time by 60%. Productionized CV defect detection with ResNet/Yolo achieving 98.5% accuracy and $1.2M annual savings. Led LoRA fine-tuning initiatives improving task performance by 35%. Developed A/B testing frameworks for 10+ concurrent experiments; authored technical standards used by 25+ engineers; implemented enterprise RAG reaching 94% retrieval accuracy across 2M+ documents.

Education

Master of Science in International Business and Analytics at Hult International Business School
August 1, 2024 - September 1, 2025
Bachelor of Technology in Computer Science Engineering (Artificial Intelligence & Machine Learning) at Bennett University
July 1, 2020 - May 1, 2024
Master of Science in International Business and Analytics at Hult International Business School
August 1, 2024 - September 1, 2025
Bachelor of Technology in Computer Science Engineering (Artificial Intelligence & Machine Learning) at Bennett University
July 1, 2020 - May 1, 2024
Master of Science in International Business and Analytics at Hult International Business School
August 1, 2024 - September 1, 2025
Bachelor of Technology in Computer Science Engineering (Artificial Intelligence & Machine Learning) at Bennett University
July 1, 2020 - May 1, 2024

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Professional Services, Computers & Electronics, Financial Services, Healthcare, Education

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Hire a Software Engineer

We have the best software engineer experts on Twine. Hire a software engineer in San Francisco today.