I’m a Generative AI Engineer with 2+ years of experience building and deploying AI applications end-to-end using Python and modern LLM tooling. I work across LangChain, LangGraph, and retrieval-augmented generation (RAG), integrating vector databases, SQL-backed knowledge sources, and AWS services to deliver grounded, reliable AI experiences. I enjoy engineering scalable AI pipelines—fine-tuning open-source LLMs with LoRA/QLoRA, improving document ranking and semantic retrieval, and shipping production workflows that reduce latency and hallucinations. With a strong foundation in machine learning and NLP (plus hands-on work in computer vision), I’m always focused on measurable business impact through better automation, stronger search relevance, and clear performance monitoring.

Shikhar Jolly

I’m a Generative AI Engineer with 2+ years of experience building and deploying AI applications end-to-end using Python and modern LLM tooling. I work across LangChain, LangGraph, and retrieval-augmented generation (RAG), integrating vector databases, SQL-backed knowledge sources, and AWS services to deliver grounded, reliable AI experiences. I enjoy engineering scalable AI pipelines—fine-tuning open-source LLMs with LoRA/QLoRA, improving document ranking and semantic retrieval, and shipping production workflows that reduce latency and hallucinations. With a strong foundation in machine learning and NLP (plus hands-on work in computer vision), I’m always focused on measurable business impact through better automation, stronger search relevance, and clear performance monitoring.

Available to hire

I’m a Generative AI Engineer with 2+ years of experience building and deploying AI applications end-to-end using Python and modern LLM tooling. I work across LangChain, LangGraph, and retrieval-augmented generation (RAG), integrating vector databases, SQL-backed knowledge sources, and AWS services to deliver grounded, reliable AI experiences.

I enjoy engineering scalable AI pipelines—fine-tuning open-source LLMs with LoRA/QLoRA, improving document ranking and semantic retrieval, and shipping production workflows that reduce latency and hallucinations. With a strong foundation in machine learning and NLP (plus hands-on work in computer vision), I’m always focused on measurable business impact through better automation, stronger search relevance, and clear performance monitoring.

See more

Work Experience

Generative AI Engineer at Prenora Dynamics
August 1, 2025 - Present
Built multi-agent generative AI workflows using LangGraph, LlamaIndex, and Python to orchestrate retrieval, reasoning, and tool execution for enterprise knowledge management across engineering applications. Engineered RAG pipelines integrating Pinecone vector databases, SQL, and proprietary knowledge repositories, improving grounded response accuracy by 34% and reducing hallucinations in production AI assistants. Fine-tuned open-source LLMs with LoRA/QLoRA on AWS SageMaker using TensorFlow, reducing inference costs by 29% and improving domain-specific QA performance. Developed ML models with LightGBM to improve document ranking, semantic retrieval, and prompt routing. Built AI performance dashboards using Tableau and SQL to monitor retrieval quality, latency, token usage, and adoption metrics, reducing average response latency by 24%.
Generative AI Engineer at Reality AI Lab
January 1, 2025 - August 1, 2025
Developed generative AI and NLP applications in Python using LangChain, RoBERTa, NLTK, and Named Entity Recognition to automate document understanding and knowledge extraction, improving entity extraction accuracy by 31% across enterprise datasets. Implemented computer vision pipelines with PyTorch and Mask R-CNN for object detection and image segmentation, reducing manual annotation effort by 42% while improving precision. Optimized predictive AI models using XGBoost, PySpark, SQL, and PostgreSQL to process large-scale structured and unstructured datasets, reducing inference time by 27% for production analytics. Automated cloud-native AI deployments with AWS Lambda and CloudFormation for repeatable, version-controlled releases. Designed Tableau dashboards with SQL/PostgreSQL to track model performance and operational KPIs including inference latency.
Junior AI/ML Engineer at Bizjump Corp
January 1, 2024 - May 31, 2024
Built machine learning models using Python, scikit-learn, Keras, decision trees, random forests, Pandas, and NumPy to automate customer behavior analysis, improving prediction accuracy by 24% through feature engineering and optimization. Performed EDA with Python, SQL, Pandas, and SQL Server, applying PCA, K-Means, and DBSCAN to identify customer segments and reduce dimensionality, decreasing preprocessing time by 30%. Validated NLP pipelines with SpaCy for text preprocessing, tokenization, entity extraction, and feature generation to produce high-quality datasets for downstream training. Deployed and tested experiments in Google Colab and Azure ML, managing training/evaluation/versioning and supporting reproducible development workflows.

Education

Bachelor of Science in Computer Science at Brooklyn College
January 11, 2030 - December 1, 2024
Bachelor of Science in Aerospace Engineering at Rutgers University
January 11, 2030 - December 1, 2021

Qualifications

Python Essentials for MLOps (Duke University)
January 11, 2030 - September 1, 2026
AI Engineer Core Track (Ed Donner)
January 11, 2030 - September 1, 2026

Industry Experience

Software & Internet, Computers & Electronics, Professional Services, Education