- Increased enterprise document retrieval accuracy by 38% by implementing Retrieval-Augmented Generation pipelines using Gemini APIs, LangChain, Vertex AI, Pinecone, and semantic search models across cloud-based knowledge management platforms. - Reduced manual research effort by 45% after deploying Generative AI solutions supporting contextual summarization, intelligent search, and automated response generation workflows used by analytics and operations teams daily. - Improved recommendation quality by 29% through transformer-based NLP models built using PyTorch and TensorFlow supporting personalization, conversational AI, intelligent ranking, and search optimization across production environments. - Processed over 20 million records daily using distributed PySpark and Kafka pipelines supporting anomaly detection, forecasting, predictive analytics, and near real-time machine learning inference across enterprise cloud infrastructure. - Lowered inference latency by 32% through Kubernetes autoscaling, GPU optimization, model quantization, and deployment improvements supporting stable machine learning performance during high-volume production workloads and concurrent user traffic. - Automated model deployment, retraining, monitoring, and version management workflows using MLflow, Kubeflow, Jenkins, and CI/CD pipelines, reducing manual operational effort by 40% across machine learning lifecycle activities. - Developed FastAPI and Flask microservices deployed on Google Kubernetes Engine supporting low-latency inference, secure API integrations, and distributed machine learning services across enterprise analytics and business applications successfully. - Partnered with data engineers, product managers, architects, and business teams across Agile environments to improve deployment timelines, production reliability, and machine learning integration across multiple cloud-based applications successfully.

Arjun Rao

- Increased enterprise document retrieval accuracy by 38% by implementing Retrieval-Augmented Generation pipelines using Gemini APIs, LangChain, Vertex AI, Pinecone, and semantic search models across cloud-based knowledge management platforms. - Reduced manual research effort by 45% after deploying Generative AI solutions supporting contextual summarization, intelligent search, and automated response generation workflows used by analytics and operations teams daily. - Improved recommendation quality by 29% through transformer-based NLP models built using PyTorch and TensorFlow supporting personalization, conversational AI, intelligent ranking, and search optimization across production environments. - Processed over 20 million records daily using distributed PySpark and Kafka pipelines supporting anomaly detection, forecasting, predictive analytics, and near real-time machine learning inference across enterprise cloud infrastructure. - Lowered inference latency by 32% through Kubernetes autoscaling, GPU optimization, model quantization, and deployment improvements supporting stable machine learning performance during high-volume production workloads and concurrent user traffic. - Automated model deployment, retraining, monitoring, and version management workflows using MLflow, Kubeflow, Jenkins, and CI/CD pipelines, reducing manual operational effort by 40% across machine learning lifecycle activities. - Developed FastAPI and Flask microservices deployed on Google Kubernetes Engine supporting low-latency inference, secure API integrations, and distributed machine learning services across enterprise analytics and business applications successfully. - Partnered with data engineers, product managers, architects, and business teams across Agile environments to improve deployment timelines, production reliability, and machine learning integration across multiple cloud-based applications successfully.

Available to hire
  • Increased enterprise document retrieval accuracy by 38% by implementing Retrieval-Augmented Generation pipelines using Gemini APIs, LangChain, Vertex AI, Pinecone, and semantic search models across cloud-based knowledge management platforms.
  • Reduced manual research effort by 45% after deploying Generative AI solutions supporting contextual summarization, intelligent search, and automated response generation workflows used by analytics and operations teams daily.
  • Improved recommendation quality by 29% through transformer-based NLP models built using PyTorch and TensorFlow supporting personalization, conversational AI, intelligent ranking, and search optimization across production environments.
  • Processed over 20 million records daily using distributed PySpark and Kafka pipelines supporting anomaly detection, forecasting, predictive analytics, and near real-time machine learning inference across enterprise cloud infrastructure.
  • Lowered inference latency by 32% through Kubernetes autoscaling, GPU optimization, model quantization, and deployment improvements supporting stable machine learning performance during high-volume production workloads and concurrent user traffic.
  • Automated model deployment, retraining, monitoring, and version management workflows using MLflow, Kubeflow, Jenkins, and CI/CD pipelines, reducing manual operational effort by 40% across machine learning lifecycle activities.
  • Developed FastAPI and Flask microservices deployed on Google Kubernetes Engine supporting low-latency inference, secure API integrations, and distributed machine learning services across enterprise analytics and business applications successfully.
  • Partnered with data engineers, product managers, architects, and business teams across Agile environments to improve deployment timelines, production reliability, and machine learning integration across multiple cloud-based applications successfully.
See more

Language

Work Experience

AI/ML Engineer at Google
November 1, 2024 - Present
Built enterprise GenAI/RAG pipelines using Gemini APIs, LangChain, Vertex AI, and Pinecone/semantic search to improve document retrieval accuracy by 38%. Deployed GenAI workflows for contextual summarization and intelligent search, reducing manual research effort by 45%. Improved recommendation quality by 29% using transformer-based NLP models. Processed 20M+ records daily with distributed PySpark/Kafka for anomaly detection and forecasting with near real-time inference. Reduced inference latency by 32% via Kubernetes autoscaling, GPU optimization, and quantization. Automated deployment/retraining/monitoring using MLflow, Kubeflow, Jenkins, and CI/CD, reducing manual effort by 40%. Developed low-latency FastAPI/Flask microservices on GKE and collaborated with cross-functional teams to improve delivery and reliability.
AI/ML Engineer at Accenture
April 1, 2020 - August 31, 2023
Improved forecasting accuracy by 27% for customer analytics, operational intelligence, churn prediction, and reporting automation. Reduced manual document processing effort by 41% through NLP solutions for sentiment analysis, entity extraction, and text classification. Built Spark and Airflow ETL pipelines for structured/unstructured data and scalable feature engineering. Accelerated deployment and batch inference by 33% via cloud-native ML integrations across AWS and Google Cloud. Increased model precision by 24% using feature engineering, hyperparameter tuning, preprocessing optimization, and continuous evaluation. Developed Tableau/Power BI dashboards for KPI monitoring and model performance visibility, and supported model governance with monitoring, drift analysis, and retraining workflows.

Education

Master of Science in Information Studies (MSIS) at Trine University
December 1, 2025 - July 24, 2026
B. Tech in Electronics and Communication Engineering at Chaitanya Bharati Institute of Technology
August 1, 2022 - July 24, 2026
Master of Science in Information Studies (MSIS) at Trine University
December 1, 2025 - July 24, 2026
B. Tech in Electronics and Communication Engineering at Chaitanya Bharati Institute of Technology
August 1, 2022 - July 24, 2026
Master of Science in Information Studies (MSIS) at Trine University
December 1, 2025 - July 24, 2026
B. Tech in Electronics and Communication Engineering at Chaitanya Bharati Institute of Technology
August 1, 2022 - July 24, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Professional Services, Software & Internet, Computers & Electronics