AI/ML Engineer with 4+ years of experience building and deploying production-grade machine learning systems across NLP, LLMs, and MLOps. Hands-on work includes fine-tuning LLMs (LoRA/QLoRA), engineering RAG pipelines, and improving retrieval and generation quality through evaluation-driven iteration. Skilled in Python, PyTorch, Hugging Face Transformers, and cloud-native ML on AWS and GCP. Experienced in optimizing inference performance (TGI/vLLM/quantization), building robust evaluation and monitoring workflows, and delivering measurable impact across enterprise AI search and IT services.

Chetan patel

AI/ML Engineer with 4+ years of experience building and deploying production-grade machine learning systems across NLP, LLMs, and MLOps. Hands-on work includes fine-tuning LLMs (LoRA/QLoRA), engineering RAG pipelines, and improving retrieval and generation quality through evaluation-driven iteration. Skilled in Python, PyTorch, Hugging Face Transformers, and cloud-native ML on AWS and GCP. Experienced in optimizing inference performance (TGI/vLLM/quantization), building robust evaluation and monitoring workflows, and delivering measurable impact across enterprise AI search and IT services.

Available to hire

AI/ML Engineer with 4+ years of experience building and deploying production-grade machine learning systems across NLP, LLMs, and MLOps. Hands-on work includes fine-tuning LLMs (LoRA/QLoRA), engineering RAG pipelines, and improving retrieval and generation quality through evaluation-driven iteration.

Skilled in Python, PyTorch, Hugging Face Transformers, and cloud-native ML on AWS and GCP. Experienced in optimizing inference performance (TGI/vLLM/quantization), building robust evaluation and monitoring workflows, and delivering measurable impact across enterprise AI search and IT services.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate

Language

Bashkir
Intermediate
Afar
Intermediate
Amharic
Beginner
Javanese
Beginner

Work Experience

AI/ML Engineer (Contract) at Hugging Face
January 1, 2026 - Present
Engineered fine-tuning pipelines for instruction-following LLMs using LoRA and QLoRA, reducing GPU memory usage by ~35% while maintaining benchmark accuracy via internal evaluation suites. Optimized text generation inference (TGI) serving configurations for enterprise clients, cutting token generation latency by ~20% through dynamic batching and INT8 quantization. Built automated model evaluation workflows using LightEval and custom task suites to track quality across iterative checkpoints. Developed reproducible dataset preprocessing pipelines and published curated datasets to the Hugging Face Hub to accelerate cross-team reuse. Triaged and resolved production inference issues on inference endpoints, maintaining 99%+ uptime through proactive monitoring and root-cause analysis.
AI Engineer (Internship) at Perplexity
June 1, 2025 - December 1, 2025
Enhanced a core RAG pipeline using hybrid search (BM25 + dense embeddings), improving answer relevance by ~12% on internal evaluation benchmarks. Built a query intent classifier to route queries to specialized retrieval paths, reducing irrelevant context and improving response coherence. Profiled and optimized the embedding generation step to cut median end-to-end latency by ~80ms without degrading retrieval quality. Designed and executed A/B tests on prompt template variations to improve citation accuracy for factual queries. Integrated a cross-encoder reranker into the source selection layer, improving evidence quality and boosting source credibility for Pro-tier answer generation.
Machine Learning Engineer at Cognizant
July 1, 2021 - July 1, 2024
Designed and deployed end-to-end NLP document classification pipelines for a Fortune 500 healthcare client, automating review workflows and reducing processing time by ~40%. Built scalable ML training infrastructure on AWS SageMaker, reducing model development cycle time from ~2 weeks to ~5 days through parallelized experiment management. Developed a customer churn prediction model using XGBoost with SHAP explainability, achieving ~87% recall and supporting targeted retention campaigns. Migrated batch inference workloads from on-prem to AWS Lambda and S3, reducing infrastructure costs by ~25% while improving deployment reliability. Engineered PySpark feature pipelines on AWS EMR to process 10M+ transactions daily for real-time fraud signal scoring. Mentored junior engineers on MLflow, DVC, and CI/CD; standardized deployment workflows using Docker and FastAPI to reduce staging-to-production handoff from ~3 days to under 4 hours.

Education

Master of Computer Science at Pace University
September 1, 2024 - May 1, 2026
Bachelor of Computer Science and Engineering at Gujarat Technological University
August 1, 2019 - June 1, 2023

Qualifications

Add your qualifications or awards here.

Industry Experience

Healthcare, Financial Services, Software & Internet, Professional Services, Computers & Electronics