I'm Manikanta Pochamalla, an AI/ML engineer with 6+ years of experience building large-scale AI platforms, LLM applications, retrieval systems, and machine learning infrastructure. I have a proven track record delivering production-grade generative AI, search, and foundation model solutions that improve accuracy, scalability, and operational efficiency. I thrive in distributed systems, cloud-native architectures, and MLOps, and I enjoy driving end-to-end AI products from research and model development through deployment and production optimization.

Manikanta Pochamalla

I'm Manikanta Pochamalla, an AI/ML engineer with 6+ years of experience building large-scale AI platforms, LLM applications, retrieval systems, and machine learning infrastructure. I have a proven track record delivering production-grade generative AI, search, and foundation model solutions that improve accuracy, scalability, and operational efficiency. I thrive in distributed systems, cloud-native architectures, and MLOps, and I enjoy driving end-to-end AI products from research and model development through deployment and production optimization.

Available to hire

I’m Manikanta Pochamalla, an AI/ML engineer with 6+ years of experience building large-scale AI platforms, LLM applications, retrieval systems, and machine learning infrastructure. I have a proven track record delivering production-grade generative AI, search, and foundation model solutions that improve accuracy, scalability, and operational efficiency.

I thrive in distributed systems, cloud-native architectures, and MLOps, and I enjoy driving end-to-end AI products from research and model development through deployment and production optimization.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert

Work Experience

AI/ML Engineer at Perplexity
August 1, 2025 - Present
Led the development of Python-based hybrid retrieval pipelines combining vector search and BM25, serving 3,000+ daily users and improving answer relevance and citation coverage across enterprise knowledge sources. Built RAG-driven query understanding and answer synthesis workflows with PyTorch and transformer models, increasing response accuracy by 32% through optimized retrieval and reranking. Optimized LLM inference and context assembly using vLLM, reducing end-to-end latency by 41% while maintaining grounded responses and source attribution. Engineered semantic retrieval systems with PyTorch, Hugging Face Transformers, embeddings (BGE), Elasticsearch, and vector search to enhance document discovery and contextual relevance. Designed multi-stage ranking pipelines leveraging Cross Encoders, XGBoost, DeepSpeed, and neural reranking models for efficient prioritization of complex queries. Implemented citation-grounded generation workflows with FastAPI, PostgreSQL, Redis, and metadata ser
AI/ML Engineer at NVIDIA
September 1, 2023 - August 31, 2025
Developed Python-based protein foundation model training pipelines using PyTorch, NeMo, and Megatron-LM, enabling large-scale biomolecular learning workflows and improving sequence prediction accuracy by 34%. Built distributed training infrastructure across multi-node clusters using NCCL and Transformer Engine, reducing model training time by 41% while maintaining scalability. Optimized inference workloads using TensorRT-LLM, Triton Inference Server, and CUDA acceleration, reducing end-to-end prediction latency by 37% and lowering infrastructure costs. Designed and fine-tuned transformer-based protein and molecular foundation models using PyTorch, BioNeMo, OpenFold, ESM2, Evo2. Implemented molecular generation and virtual screening workflows using MegaMolBART, DiffDock, Graph Neural Networks, accelerating candidate compound discovery. Engineered scalable data processing pipelines using Pandas, NumPy, Apache Arrow, and Parquet for multi-terabyte datasets. Developed containerized AI micr
Machine Learning Engineer at Accenture
May 1, 2019 - August 31, 2022
Developed Python and PySpark-based ML pipelines on Databricks, processing enterprise-scale datasets and improving model response accuracy by 32% through optimized feature engineering and training workflows. Built automated MLOps workflows using MLflow, Azure DevOps, and CI/CD pipelines, reducing model deployment time by 45% while supporting production inference for 3,000+ business users. Engineered scalable data processing solutions using Apache Spark, Delta Lake, Python, and SQL, enabling reliable feature generation, experiment tracking, and model version management. Designed and deployed containerized microservices using FastAPI, Docker, Kubernetes, and REST APIs, delivering highly available real-time prediction services across enterprise applications. Implemented cloud-native ML infrastructure on AWS using S3, SageMaker, EKS, and Terraform, ensuring secure, scalable model training, deployment, monitoring, and governance. Developed end-to-end data orchestration pipelines using Airflo

Education

Master of Science in Computer Science at Campbellsville University
January 11, 2030 - June 29, 2026

Qualifications

AWS Certified Data Engineer
January 11, 2030 - June 29, 2026
GCP Professional Data Engineer
January 11, 2030 - June 29, 2026
Databricks Certified Data Engineer
January 11, 2030 - June 29, 2026

Industry Experience

Software & Internet, Professional Services, Media & Entertainment, Education, Healthcare