Available to hire
Artificial Intelligence Engineer with 3+ years of experience architecting high-throughput, low-latency GenAI and predictive ML pipelines for financial and enterprise systems. Expertise includes optimizing LLM inference (vLLM, TensorRT-LLM) and deploying agentic workflows (LangGraph, CrewAI) alongside advanced RAG architectures (GraphRAG, Pinecone) on Kubernetes and AWS.
Track record of delivering production systems with strong reliability and observability, including hybrid semantic-graph retrieval across large-scale datasets, GPU/memory optimization, and end-to-end ML lifecycle management with monitoring and experiment tracking.
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Language
English
Advanced
Work Experience
Artificial Intelligence (AI) Engineer at CitiGroup
September 1, 2025 - PresentEngineered an autonomous compliance-checking agent using CrewAI and Llama 3 (70B) to parse 500K+ daily legal documents, reducing manual review time by 85% with 99.2% accuracy (DeepEval). Architected a low-latency LLM serving infrastructure on AWS SageMaker using NVIDIA Triton and TensorRT-LLM with AWQ quantization, reducing latency from 1.2s to 180ms at 10,000+ QPS. Built a hybrid Semantic-GraphRAG pipeline using Neo4j and Milvus to interlink complex financial entities across 50TB of datasets, improving retrieval precision by 40% and generating actionable risk insights. Deployed model lifecycle management with Ray Serve on distributed Kubernetes, integrating LangSmith and Weights & Biases for prompt telemetry; reduced model drift by 22% and maintained 99.9% uptime for AI APIs.
Artificial Intelligence (AI) Engineer at CitiGroup, United States
September 1, 2025 - PresentEngineered an autonomous compliance-checking agent using CrewAI and Llama 3 (70B) to parse 500K+ daily legal documents, reducing manual review time by 85% and achieving 99.2% accuracy (DeepEval). Architected a low-latency LLM serving stack with NVIDIA Triton and TensorRT-LLM on AWS SageMaker, using AWQ quantization to cut inference latency from 1.2s to 180ms at 10,000+ QPS. Designed a hybrid Semantic-GraphRAG pipeline leveraging Neo4j and Milvus to interlink financial entities across 50TB of unstructured transactional data, improving retrieval precision by 40% and producing actionable risk insights. Orchestrated model lifecycle management by deploying Ray Serve on distributed Kubernetes, integrating LangSmith and W&B prompt telemetry to reduce production drift by 22% while maintaining 99.9% uptime for AI APIs.
AI Engineer at Informative Web Solutions
July 1, 2021 - December 1, 2023Developed a churn prediction engine processing 10M+ daily user telemetry events using Kafka and Spark; trained optimized XGBoost pipelines contributing to a 15% increase in quarterly recurring revenue. Built and fine-tuned BERT-based NLP models with PyTorch and Hugging Face PEFT for automated customer support ticket routing, processing 2M+ monthly inquiries and cutting average resolution time by 60%. Engineered high-throughput defect detection using ViT/CNN models; deployed an ONNX-optimized FastAPI edge pipeline achieving 500+ images/second with 98.7% F1. Automated ML CI/CD with GitHub Actions, Docker, and MLflow on GCP Vertex AI, accelerating releases from bi-weekly to daily. Created an enterprise knowledge extraction system using semantic QA with ChromaDB and LangChain, indexing 1M+ documents to save 10 hours/week in retrieval. Implemented robust ETL pipelines with Airflow to ingest unstructured web data into Databricks Delta Lake, reducing ingestion latency by 35%.
AI Engineer at Informative Web Solutions, India
July 1, 2021 - December 1, 2023Developed a churn prediction engine processing 10M+ daily user telemetry events with Kafka and Spark; built optimized XGBoost pipelines improving retention strategies and driving a 15% increase in quarterly recurring revenue. Built and fine-tuned BERT-variant NLP models using PyTorch and Hugging Face PEFT for automated customer support ticket routing, processing 2M+ monthly inquiries and reducing average resolution time by 60% while lowering manual triage costs. Engineered high-throughput defect detection using ViT and CNNs, deploying an ONNX-optimized FastAPI pipeline at the edge to process 500+ images/second with 98.7% F1-score. Automated ML deployment CI/CD using GitHub Actions, Docker, and MLflow on GCP Vertex AI, improving deployment frequency from bi-weekly to daily. Implemented an enterprise knowledge extraction system using semantic QA with ChromaDB and LangChain to index 1M+ legacy documents, saving teams ~10 hours/week in retrieval. Built robust ETL pipelines using Airflow to
Education
M.S. in Applied Artificial Intelligence at Stevens Institute of Technology, Hoboken, New Jersey
January 11, 2030 - July 22, 2026BTech in Electronics and Communication Engineering at Indian Institute of Technology (ISM), Dhanbad, India
January 11, 2030 - July 22, 2026M.S. in Applied Artificial Intelligence at Stevens Institute of Technology
January 11, 2030 - July 22, 2026BTech in Electronics and Communication Engineering at Indian Institute of Technology (ISM)
January 11, 2030 - July 22, 2026Qualifications
Industry Experience
Financial Services, Software & Internet, Healthcare, Manufacturing, Professional Services, Computers & Electronics
Experience Level
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Hire a AI Engineer
We have the best ai engineer experts on Twine. Hire a ai engineer in Jersey City today.