I am an AI/ML engineer with 5+ years of experience building and deploying production-scale machine learning and generative AI solutions. I specialize in distributed training, RAG systems, and scalable inference, and I enjoy collaborating with product, research, and UX teams to drive measurable business outcomes. I thrive in fast-moving environments and am passionate about responsible AI, observability, and delivering high-value AI capabilities at cloud scale. From enterprise-grade RAG pipelines at Databricks to high-performance AI agents at NVIDIA, I focus on reliability, governance, and user value. I stay curious about new techniques, continuously optimize systems for latency and cost, and lead cross-functional initiatives that translate metrics into revenue growth and improved user engagement.

Samyukth C

I am an AI/ML engineer with 5+ years of experience building and deploying production-scale machine learning and generative AI solutions. I specialize in distributed training, RAG systems, and scalable inference, and I enjoy collaborating with product, research, and UX teams to drive measurable business outcomes. I thrive in fast-moving environments and am passionate about responsible AI, observability, and delivering high-value AI capabilities at cloud scale. From enterprise-grade RAG pipelines at Databricks to high-performance AI agents at NVIDIA, I focus on reliability, governance, and user value. I stay curious about new techniques, continuously optimize systems for latency and cost, and lead cross-functional initiatives that translate metrics into revenue growth and improved user engagement.

Available to hire

I am an AI/ML engineer with 5+ years of experience building and deploying production-scale machine learning and generative AI solutions. I specialize in distributed training, RAG systems, and scalable inference, and I enjoy collaborating with product, research, and UX teams to drive measurable business outcomes. I thrive in fast-moving environments and am passionate about responsible AI, observability, and delivering high-value AI capabilities at cloud scale.

From enterprise-grade RAG pipelines at Databricks to high-performance AI agents at NVIDIA, I focus on reliability, governance, and user value. I stay curious about new techniques, continuously optimize systems for latency and cost, and lead cross-functional initiatives that translate metrics into revenue growth and improved user engagement.

See more

Experience Level

Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate

Work Experience

AI/ML Engineer at Databricks
July 1, 2024 - Present
Developed end-to-end ML pipelines on Databricks using PySpark, MLflow, and Kubernetes, enabling distributed training of large language models across 32+ GPU nodes and reducing training time by 50%. Built an enterprise-scale RAG system by integrating Databricks Lakehouse with FAISS vector search and OpenAI APIs, improving query relevance and supporting 5M+ monthly searches. Optimized inference serving with Ray Serve and Triton for GPU auto-scaling, achieving low 95th percentile latency and real-time predictions at cloud scale. Collaborated with product and UX teams to design an autonomous AI analytics assistant using LangChain and Pinecone, boosting feature adoption by 20%. Implemented responsible AI evaluation pipelines to audit LLM outputs and improve reliability and compliance. Led CI/CD and Kubernetes-based deployment for ML workloads, orchestrating automated MLflow pipelines and increasing deployment frequency. Designed A/B tests and dashboards for personalized recommendations, ach
AI/ML Engineer at NVIDIA
June 1, 2019 - December 1, 2022
Designed a distributed GPU training pipeline using PyTorch DDP and NCCL on A100 clusters, accelerating model convergence by 2.5x and enabling training on 10M+ high-resolution images. Built Triton Inference Server pipelines with TensorRT-optimized models for LLMs, reducing p99 inference latency by 30% across thousands of daily requests. Collaborated with research to integrate an ONNX-optimized multi-modal Transformer, improving end-task accuracy with a double-digit BLEU gain on internal benchmarks. Optimized GPU kernel execution and memory usage (CUDA/FlashAttention) to 5x throughput, significantly reducing latency on edge and cloud. Launched an AI agent powered by LangChain and LlamaIndex to improve user query resolution and reduce support tickets by 20%. Engineered scalable vector retrieval using FAISS and Pinecone for 50M+ documents, increasing retrieval speed by 5x. Conducted extensive A/B testing and error analysis, delivering roughly 10% lift in F1-score and improved user retentio

Education

Master's Degree in Computer Science at University of Texas at Arlington
January 11, 2030 - June 29, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Professional Services, Other

Experience Level

Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate