I’m an AI/ML engineer focused on building production-grade Generative AI systems, especially Large Language Model (LLM) applications like fine-tuning, RAG, and inference optimization. Over the last few years, I’ve worked on improving throughput and latency for AI services using techniques like dynamic batching, GPU parallelism, and scalable model serving with vLLM and Triton, while maintaining high uptime for large-scale workloads.
I also enjoy the full lifecycle of AI engineering—from distributed training with PyTorch FSDP/DeepSpeed to reliable MLOps automation using tools like Kubernetes, Terraform, Argo Workflows, and MLflow. I build and monitor end-to-end pipelines (streaming data, feature serving, vector search, and evaluation/benchmarking), and I’m particularly driven by turning model quality improvements—like better retrieval relevance and reduced hallucinations—into measurable operational impact.
Skills
Experience Level
Language
Work Experience
Education
Qualifications
Industry Experience
Skills
Experience Level
Hire a Software Engineer
We have the best software engineer experts on Twine. Hire a software engineer in San Francisco today.