I am an AI/ML Engineer with 5 years of experience building production-grade machine learning, generative AI, and NLP systems across financial, healthcare, and enterprise domains. I specialize in scalable retrieval-augmented generation (RAG) based LLM applications, AI agents, and cloud-native ML pipelines using AWS, Kubernetes, and Databricks. I design RAG pipelines, vector database architectures, and semantic search to enable contextual knowledge retrieval from large document repositories. I focus on reducing hallucinations, improving model accuracy, and accelerating ML deployment through MLOps, CI/CD automation, and real-time cloud inference while maintaining AI safety, explainability, and compliance.

Ajay Kumar

I am an AI/ML Engineer with 5 years of experience building production-grade machine learning, generative AI, and NLP systems across financial, healthcare, and enterprise domains. I specialize in scalable retrieval-augmented generation (RAG) based LLM applications, AI agents, and cloud-native ML pipelines using AWS, Kubernetes, and Databricks. I design RAG pipelines, vector database architectures, and semantic search to enable contextual knowledge retrieval from large document repositories. I focus on reducing hallucinations, improving model accuracy, and accelerating ML deployment through MLOps, CI/CD automation, and real-time cloud inference while maintaining AI safety, explainability, and compliance.

Available to hire

I am an AI/ML Engineer with 5 years of experience building production-grade machine learning, generative AI, and NLP systems across financial, healthcare, and enterprise domains. I specialize in scalable retrieval-augmented generation (RAG) based LLM applications, AI agents, and cloud-native ML pipelines using AWS, Kubernetes, and Databricks.

I design RAG pipelines, vector database architectures, and semantic search to enable contextual knowledge retrieval from large document repositories. I focus on reducing hallucinations, improving model accuracy, and accelerating ML deployment through MLOps, CI/CD automation, and real-time cloud inference while maintaining AI safety, explainability, and compliance.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert

Language

English
Fluent

Work Experience

AI/ML Engineer at Capital One
September 1, 2024 - Present
Project: Generative AI Customer Support Assistant. Led the development and deployment of a production RAG-based LLM assistant handling 12,000+ queries/day, reducing irrelevant responses by 54% through improved retrieval grounding. Built embedding/retrieval pipelines with FAISS and OpenSearch, enabling semantic search and hybrid BM25 + vector ranking for enterprise banking queries. Created an end-to-end evaluation framework (Precision/Recall/F1/ROC-AUC/PR) with human-in-the-loop validation to reduce hallucinations. Implemented LangChain-style agent orchestration for multitask customer queries with tool-calling. Established MLOps pipelines on AWS SageMaker (training, tuning, registry, CI/CD) with Lambda/S3/CloudWatch for drift/latency. Deployed secure FastAPI inference services on AWS ECS with authentication, rate limiting, and PII masking to meet GLBA. Optimized Spark/SQL-based feature engineering for LLM personalization and intent classification. Enhanced deep learning models (TensorFl
AI/ML Engineer at Humana
November 1, 2023 - August 1, 2024
Healthcare Risk Stratification: Built LLM-powered clinical analytics pipelines processing 5TB+ of unstructured records, improving automated clinical insights and decision-support by 35%. Architected Retrieval-Augmented Generation (RAG) pipelines using LangChain and LlamaIndex with vector DBs for semantic retrieval across clinical notes, lab reports, and histories. Deployed on AWS (S3/Lambda/SageMaker) with Docker/Kubernetes. Implemented AI safety guardrails and evaluation via MLflow, enabling hallucination detection and monitoring. Established CI/CD with GitHub Actions and observability/drift detection to ensure reliability.
Machine Learning Engineer at LTIMindtree
April 1, 2020 - July 1, 2022
NLP-based ticket classification system: Engineered an NLP pipeline (TF-IDF, Logistic Regression, SVM) achieving 91% routing accuracy; fine-tuned Transformer-based models (BERT) using PyTorch, boosting F1 from 0.74 to 0.88 for multi-class categories. Built scalable preprocessing and feature pipelines for 500K+ historical tickets; deployed AWS S3/Lambda/SageMaker with Docker/Kubernetes, and CI/CD via GitHub Actions. Implemented real-time inference with FastAPI and cluster-based clustering (BERTopic) to identify emerging issue patterns; set up model monitoring and drift detection with MLflow.

Education

Master of Science in Computer Science at Campbellsville University
January 11, 2030 - August 1, 2024

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Healthcare, Software & Internet, Professional Services