I am Sumedh Anumala, an Artificial Intelligence and Machine Learning Engineer with 3+ years of experience architecting enterprise-grade Generative AI pipelines and scalable MLOps infrastructures. I specialize in deploying high-throughput, low-latency LLMs and orchestrating complex agentic workflows that drive sub-second inferences and multi-million-dollar operational efficiencies. I partner with cross-functional teams to deliver end-to-end AI services, from data ingestion and feature engineering to model deployment and observability. I enjoy building robust, observable systems with tools like LangSmith, W&B, and MLflow, and I continually explore state-of-the-art techniques in RAG, retrieval, and generative AI to unlock business impact.

Sumedh Anumala

I am Sumedh Anumala, an Artificial Intelligence and Machine Learning Engineer with 3+ years of experience architecting enterprise-grade Generative AI pipelines and scalable MLOps infrastructures. I specialize in deploying high-throughput, low-latency LLMs and orchestrating complex agentic workflows that drive sub-second inferences and multi-million-dollar operational efficiencies. I partner with cross-functional teams to deliver end-to-end AI services, from data ingestion and feature engineering to model deployment and observability. I enjoy building robust, observable systems with tools like LangSmith, W&B, and MLflow, and I continually explore state-of-the-art techniques in RAG, retrieval, and generative AI to unlock business impact.

Available to hire

I am Sumedh Anumala, an Artificial Intelligence and Machine Learning Engineer with 3+ years of experience architecting enterprise-grade Generative AI pipelines and scalable MLOps infrastructures. I specialize in deploying high-throughput, low-latency LLMs and orchestrating complex agentic workflows that drive sub-second inferences and multi-million-dollar operational efficiencies. I partner with cross-functional teams to deliver end-to-end AI services, from data ingestion and feature engineering to model deployment and observability. I enjoy building robust, observable systems with tools like LangSmith, W&B, and MLflow, and I continually explore state-of-the-art techniques in RAG, retrieval, and generative AI to unlock business impact.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Language

Afar
Advanced

Work Experience

Artificial Intelligence (AI) Engineer at PNC Bank
January 1, 2025 - Present
Architected a multi-agent compliance review system utilizing CrewAI and LangGraph, automating the parsing of 2TB+ of unstructured regulatory documents and reducing manual audit cycle times by 40%. Engineered a Graph Retrieval-Augmented Generation (GraphRAG) architecture plumbing Neo4j and LlamaIndex to improve complex semantic query retrieval accuracy by 35% for risk analysts. Deployed heavily quantized (AWQ) Llama-3 models via vLLM and TensorRT-LLM on AWS GPU clusters, driving a 3x increase in inference throughput while reducing quarterly compute infrastructure costs by $150K. Integrated LangSmith and Ragas into the CI/CD pipeline for real-time LLM observability with strict 99.9% factual accuracy in client-facing outputs. Implemented real-time transaction telemetry using Apache Kafka into a PyTorch-based deep learning anomaly detector, slashing fraud detection latency to sub-50ms during peak volumes. Conducted QLoRA fine-tuning on banking datasets, enhancing domain-specific financial
Machine Learning Engineer at Sage SoftTech
June 1, 2021 - August 1, 2023
Developed CNNs in PyTorch for automated document OCR pipeline processing 5M+ scans with 96% text extraction accuracy. Deployed BERT-based transformers for enterprise sentiment analysis and improved customer support ticket routing by 25%. Architected ETL pipelines with Apache Spark and Databricks to cleanse and transform 10TB of data, accelerating feature engineering by 50%. Built an XGBoost-based churn model with feature engineering and hyperparameter tuning, boosting ROC-AUC to 0.89 and preserving ~$77K in annual revenue. Integrated dense vector embeddings into Qdrant and Milvus to enable sub-100ms semantic search. Standardized ML lifecycle with MLflow and Weights & Biases, accelerating cross-team iteration by 30%. Automated end-to-end retraining pipelines with Apache Airflow and data drift detection to maintain accuracy.

Education

Master of Science in Computer Science at Pace University, Manhattan, New York
January 11, 2030 - July 3, 2026
Bachelor of Engineering in Computer Science at Chaitanya Bharathi Institute of Technology (CBIT), India
January 11, 2030 - July 3, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Software & Internet, Professional Services