I’m an AI/ML engineer in the US with 5+ years of experience building scalable machine learning and generative AI systems for enterprise use cases. My work spans transformer-based models, NLP, and production RAG (retrieval-augmented generation) pipelines that turn large document and dataset collections into fast, context-aware answers. I enjoy taking models from experimentation to production—designing data pipelines (Spark, Kafka, SQL), deploying low-latency inference APIs with FastAPI, and running reliable MLOps workflows with MLflow, Docker, Kubernetes, and CI/CD. I’m especially focused on performance optimization (quantization, GPU acceleration, caching) and on integrating AI services smoothly into AWS-based microservices architectures.

Venkata Krishna Ullam

I’m an AI/ML engineer in the US with 5+ years of experience building scalable machine learning and generative AI systems for enterprise use cases. My work spans transformer-based models, NLP, and production RAG (retrieval-augmented generation) pipelines that turn large document and dataset collections into fast, context-aware answers. I enjoy taking models from experimentation to production—designing data pipelines (Spark, Kafka, SQL), deploying low-latency inference APIs with FastAPI, and running reliable MLOps workflows with MLflow, Docker, Kubernetes, and CI/CD. I’m especially focused on performance optimization (quantization, GPU acceleration, caching) and on integrating AI services smoothly into AWS-based microservices architectures.

Available to hire

I’m an AI/ML engineer in the US with 5+ years of experience building scalable machine learning and generative AI systems for enterprise use cases. My work spans transformer-based models, NLP, and production RAG (retrieval-augmented generation) pipelines that turn large document and dataset collections into fast, context-aware answers.

I enjoy taking models from experimentation to production—designing data pipelines (Spark, Kafka, SQL), deploying low-latency inference APIs with FastAPI, and running reliable MLOps workflows with MLflow, Docker, Kubernetes, and CI/CD. I’m especially focused on performance optimization (quantization, GPU acceleration, caching) and on integrating AI services smoothly into AWS-based microservices architectures.

See more

Language

Work Experience

AI/ML Engineer at Cigna
July 1, 2025 - Present
Developed generative AI models using PyTorch and transformer architectures for clinical risk prediction and claims analytics. Engineered RAG pipelines with LangChain, FAISS, and Hugging Face transformers for semantic search and natural language querying across structured clinical datasets. Built and deployed low-latency RESTful inference APIs using FastAPI, Docker, and Kubernetes for integration with enterprise backend systems. Implemented MLflow-driven MLOps workflows with CI/CD for improved model versioning and reproducibility. Optimized inference performance using quantization, GPU acceleration, and caching strategies to reduce latency while maintaining accuracy. Applied prompt engineering and LLM optimization techniques, including few-shot prompting and tool usage, to improve response relevance and contextual accuracy.
ML Engineer at Cognizant
February 1, 2021 - December 1, 2023
Architected distributed machine learning systems using Apache Spark and Python for predictive analytics and recommendation use cases. Built and deployed scalable ML services with Docker, Kubernetes, and FastAPI to support reliable, high-throughput inference. Developed batch and streaming data pipelines with Kafka, Spark, and SQL to enable real-time data processing and downstream analytics. Built NLP models for text classification and sentiment analysis using Scikit-learn and Python, reducing manual processing effort. Performed exploratory data analysis and statistical validation using R to improve feature selection and experimental outcomes. Implemented CI/CD-based deployment pipelines with MLflow to enhance lifecycle management, traceability, and reduce deployment errors. Collaborated cross-functionally to integrate ML solutions into microservices-based architectures.
Junior ML Engineer at Freshworks
June 1, 2019 - January 1, 2021
Prepared and processed structured and unstructured datasets using Python, Pandas, NumPy, and SQL. Developed ML models using Scikit-learn and TensorFlow for automated ticket categorization. Automated data preprocessing, feature engineering, and training pipelines to improve consistency and reduce manual effort. Evaluated model performance using precision, recall, and accuracy, and applied hyperparameter tuning to improve classification accuracy. Integrated ML outputs into Java-based backend systems to support production deployment. Applied NLP preprocessing such as tokenization and vectorization to improve text classification workflows and reduce ticket categorization effort.

Education

Master’s in Data Science at Pace University
January 1, 2024 - December 1, 2025
Bachelor’s Technology in Electronics and Communication Engineering at KL University
January 1, 2016 - May 1, 2020

Qualifications

AWS Certified Machine Learning Engineer
January 11, 2030 - September 1, 2026

Industry Experience

Healthcare, Software & Internet, Professional Services, Financial Services

Hire a AI Engineer

We have the best ai engineer experts on Twine. Hire a ai engineer in Jersey City today.