I'm Pranay Kumar Galipelly, a Senior Generative AI Engineer focused on designing and deploying LLM-powered systems across finance, cloud, and enterprise domains. I specialize in Retrieval-Augmented Generation (RAG), fine-tuning with PEFT and LoRA, structured prompt engineering, and multi-modal AI applications involving diffusion and CLIP models to deliver grounded, scalable solutions. I build full-stack GenAI apps using Angular/JavaScript frontends and Python backends, implement agentic AI governance and observability, and architect multi-cloud deployments on Azure OpenAI, AWS Bedrock, and GCP Vertex AI. My work emphasizes performance, security, cost-efficiency, and interpretability, with a track record of reducing manual workloads and hallucinations through robust evaluation, human-in-the-loop feedback, and enterprise-grade AI governance.

Pranay Kumar Galipelly

I'm Pranay Kumar Galipelly, a Senior Generative AI Engineer focused on designing and deploying LLM-powered systems across finance, cloud, and enterprise domains. I specialize in Retrieval-Augmented Generation (RAG), fine-tuning with PEFT and LoRA, structured prompt engineering, and multi-modal AI applications involving diffusion and CLIP models to deliver grounded, scalable solutions. I build full-stack GenAI apps using Angular/JavaScript frontends and Python backends, implement agentic AI governance and observability, and architect multi-cloud deployments on Azure OpenAI, AWS Bedrock, and GCP Vertex AI. My work emphasizes performance, security, cost-efficiency, and interpretability, with a track record of reducing manual workloads and hallucinations through robust evaluation, human-in-the-loop feedback, and enterprise-grade AI governance.

Available to hire

I’m Pranay Kumar Galipelly, a Senior Generative AI Engineer focused on designing and deploying LLM-powered systems across finance, cloud, and enterprise domains. I specialize in Retrieval-Augmented Generation (RAG), fine-tuning with PEFT and LoRA, structured prompt engineering, and multi-modal AI applications involving diffusion and CLIP models to deliver grounded, scalable solutions.

I build full-stack GenAI apps using Angular/JavaScript frontends and Python backends, implement agentic AI governance and observability, and architect multi-cloud deployments on Azure OpenAI, AWS Bedrock, and GCP Vertex AI. My work emphasizes performance, security, cost-efficiency, and interpretability, with a track record of reducing manual workloads and hallucinations through robust evaluation, human-in-the-loop feedback, and enterprise-grade AI governance.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert

Language

English
Fluent

Work Experience

Senior Generative AI Engineer | RAG & AI Agents –Support Systems at UnitedHealth Group / Optum
January 1, 2024 - Present
Designed and implemented scalable Retrieval-Augmented Generation (RAG) pipelines using LangChain and LlamaIndex to process healthcare data and improve LLM response grounding. Built LLM-powered applications leveraging GPT-4, Claude, and LLaMA, including fine-tuning with PEFT and LoRA for domain-specific clinical summarization and Q&A. Designed graph-based multi-step agent workflows using LangChain and LangGraph-style orchestration for complex reasoning and automation. Developed modular Python APIs and GenAI components with FastAPI to support scalable enterprise AI applications. Implemented agent development frameworks (ADK-style) to build extensible, tool-integrated AI agent systems. Built LangChain-based AI agents for incident triage, benefits explanation, and prior authorization workflows. Created document ingestion pipelines including chunking strategies, embeddings, and vector search for unstructured healthcare data (HL7/FHIR). Implemented hybrid retrieval mechanisms to improve grou
Generative AI Engineer at Nutanix
June 1, 2023 - December 1, 2023
Built domain-specific RAG pipelines with FAISS, Chroma, and GPT-4 for internal product support. Fine-tuned LLaMA 2 with LoRA using company documentation and internal wikis. Created FastAPI services containerized with Docker and orchestrated with Kubernetes. Implemented few-shot, CoT, and ReAct prompts for summarization and legal automation use cases. Established LLM performance dashboards and latency/cost benchmarks across cloud providers. Optimized GPU workloads with mixed precision and tensor parallelism. Led LLMOps best practices: rollback, CI/CD, A/B testing, and model/version tracking. Engineered LangChain agents capable of structured document Q&A, config analysis, and alert drafting. Automated LLM experimentation workflows with MLflow, enabling rapid evaluation of prompt/model variants. Integrated AI services with enterprise applications built using .NET and REST APIs for backend interoperability. Developed modular GenAI components and orchestration pipelines using LangChain and
AI Solutions Developer – Generative AI Focus at PayPal
December 1, 2022 - May 1, 2023
Fine-tuned GPT-2 to generate summaries, suggestions, and auto-completions for internal knowledge bases and compliance documentation. Built T5-based microservices to auto-summarize support tickets, policy guides, and operational workflows. Applied decoding strategies (top-k, nucleus) to improve clarity, tone, and factual accuracy in customer-facing content. Developed FastAPI services powering explainers, rewriters, and summarizers for support and operations teams. Built full-stack application with Angular frontend and FastAPI backend integrating OpenAI APIs and RAG pipelines. Deployed real-time summarization with GPU acceleration to support instant content delivery in internal tools. Tracked performance metrics like BLEU and ROUGE using MLflow to evaluate quality across prompts. Implemented content validation and hallucination filters to ensure regulatory language compliance. Conducted A/B prompt testing via Jira to optimize response tone and coverage across internal personas. Created r
Data Engineer – ML/ETL at Verizon (via Marlabs)
August 1, 2021 - May 1, 2022
Developed scalable data pipelines to ingest and transform telecom data for machine learning use cases. Built churn prediction models using Scikit-learn with ROC/AUC evaluation. Created classification pipelines for intent detection and sentiment analysis using TF-IDF, logistic regression, and SVM. Designed and deployed neural network prototypes in Keras and TensorFlow for customer behavior modeling. Built ETL jobs to clean and prepare structured/unstructured datasets using SQL, Pandas, and Bash scripts. Automated model deployment workflows on AWS EC2, enabling repeatable experiments. Visualized data trends and KPI dashboards using Matplotlib and Seaborn for cross-functional reporting. Collaborated with product and data science teams to deliver insights for retention and marketing strategies.

Education

Master of Science in Computer and Information Science at Southern Arkansas University
January 11, 2030 - May 1, 2024

Qualifications

Generative AI with LLMs
January 1, 2024 - June 30, 2026
AWS Certified Machine Learning
January 1, 2024 - June 30, 2026

Industry Experience

Healthcare, Financial Services, Telecommunications, Software & Internet, Professional Services