Artificial Intelligence Engineer with 3+ years of experience architecting high-throughput, low-latency GenAI and predictive ML pipelines for financial and enterprise systems. Expertise includes optimizing LLM inference (vLLM, TensorRT-LLM) and deploying agentic workflows (LangGraph, CrewAI) alongside advanced RAG architectures (GraphRAG, Pinecone) on Kubernetes and AWS. Track record of delivering production systems with strong reliability and observability, including hybrid semantic-graph retrieval across large-scale datasets, GPU/memory optimization, and end-to-end ML lifecycle management with monitoring and experiment tracking.

Deekshith Ramavath

Artificial Intelligence Engineer with 3+ years of experience architecting high-throughput, low-latency GenAI and predictive ML pipelines for financial and enterprise systems. Expertise includes optimizing LLM inference (vLLM, TensorRT-LLM) and deploying agentic workflows (LangGraph, CrewAI) alongside advanced RAG architectures (GraphRAG, Pinecone) on Kubernetes and AWS. Track record of delivering production systems with strong reliability and observability, including hybrid semantic-graph retrieval across large-scale datasets, GPU/memory optimization, and end-to-end ML lifecycle management with monitoring and experiment tracking.

Available to hire

Artificial Intelligence Engineer with 3+ years of experience architecting high-throughput, low-latency GenAI and predictive ML pipelines for financial and enterprise systems. Expertise includes optimizing LLM inference (vLLM, TensorRT-LLM) and deploying agentic workflows (LangGraph, CrewAI) alongside advanced RAG architectures (GraphRAG, Pinecone) on Kubernetes and AWS.

Track record of delivering production systems with strong reliability and observability, including hybrid semantic-graph retrieval across large-scale datasets, GPU/memory optimization, and end-to-end ML lifecycle management with monitoring and experiment tracking.

See more

Language

English
Advanced

Work Experience

Artificial Intelligence (AI) Engineer at CitiGroup
September 1, 2025 - Present
Engineered an autonomous compliance-checking agent using CrewAI and Llama 3 (70B) to parse 500K+ daily legal documents, reducing manual review time by 85% with 99.2% accuracy (DeepEval). Architected a low-latency LLM serving infrastructure on AWS SageMaker using NVIDIA Triton and TensorRT-LLM with AWQ quantization, reducing latency from 1.2s to 180ms at 10,000+ QPS. Built a hybrid Semantic-GraphRAG pipeline using Neo4j and Milvus to interlink complex financial entities across 50TB of datasets, improving retrieval precision by 40% and generating actionable risk insights. Deployed model lifecycle management with Ray Serve on distributed Kubernetes, integrating LangSmith and Weights & Biases for prompt telemetry; reduced model drift by 22% and maintained 99.9% uptime for AI APIs.
Artificial Intelligence (AI) Engineer at CitiGroup, United States
September 1, 2025 - Present
Engineered an autonomous compliance-checking agent using CrewAI and Llama 3 (70B) to parse 500K+ daily legal documents, reducing manual review time by 85% and achieving 99.2% accuracy (DeepEval). Architected a low-latency LLM serving stack with NVIDIA Triton and TensorRT-LLM on AWS SageMaker, using AWQ quantization to cut inference latency from 1.2s to 180ms at 10,000+ QPS. Designed a hybrid Semantic-GraphRAG pipeline leveraging Neo4j and Milvus to interlink financial entities across 50TB of unstructured transactional data, improving retrieval precision by 40% and producing actionable risk insights. Orchestrated model lifecycle management by deploying Ray Serve on distributed Kubernetes, integrating LangSmith and W&B prompt telemetry to reduce production drift by 22% while maintaining 99.9% uptime for AI APIs.
AI Engineer at Informative Web Solutions
July 1, 2021 - December 1, 2023
Developed a churn prediction engine processing 10M+ daily user telemetry events using Kafka and Spark; trained optimized XGBoost pipelines contributing to a 15% increase in quarterly recurring revenue. Built and fine-tuned BERT-based NLP models with PyTorch and Hugging Face PEFT for automated customer support ticket routing, processing 2M+ monthly inquiries and cutting average resolution time by 60%. Engineered high-throughput defect detection using ViT/CNN models; deployed an ONNX-optimized FastAPI edge pipeline achieving 500+ images/second with 98.7% F1. Automated ML CI/CD with GitHub Actions, Docker, and MLflow on GCP Vertex AI, accelerating releases from bi-weekly to daily. Created an enterprise knowledge extraction system using semantic QA with ChromaDB and LangChain, indexing 1M+ documents to save 10 hours/week in retrieval. Implemented robust ETL pipelines with Airflow to ingest unstructured web data into Databricks Delta Lake, reducing ingestion latency by 35%.
AI Engineer at Informative Web Solutions, India
July 1, 2021 - December 1, 2023
Developed a churn prediction engine processing 10M+ daily user telemetry events with Kafka and Spark; built optimized XGBoost pipelines improving retention strategies and driving a 15% increase in quarterly recurring revenue. Built and fine-tuned BERT-variant NLP models using PyTorch and Hugging Face PEFT for automated customer support ticket routing, processing 2M+ monthly inquiries and reducing average resolution time by 60% while lowering manual triage costs. Engineered high-throughput defect detection using ViT and CNNs, deploying an ONNX-optimized FastAPI pipeline at the edge to process 500+ images/second with 98.7% F1-score. Automated ML deployment CI/CD using GitHub Actions, Docker, and MLflow on GCP Vertex AI, improving deployment frequency from bi-weekly to daily. Implemented an enterprise knowledge extraction system using semantic QA with ChromaDB and LangChain to index 1M+ legacy documents, saving teams ~10 hours/week in retrieval. Built robust ETL pipelines using Airflow to

Education

M.S. in Applied Artificial Intelligence at Stevens Institute of Technology, Hoboken, New Jersey
January 11, 2030 - July 22, 2026
BTech in Electronics and Communication Engineering at Indian Institute of Technology (ISM), Dhanbad, India
January 11, 2030 - July 22, 2026
M.S. in Applied Artificial Intelligence at Stevens Institute of Technology
January 11, 2030 - July 22, 2026
BTech in Electronics and Communication Engineering at Indian Institute of Technology (ISM)
January 11, 2030 - July 22, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Software & Internet, Healthcare, Manufacturing, Professional Services, Computers & Electronics