I am an innovative and results-driven Lead ML Engineer with 20 years of IT experience, including 9 years specializing in Gen AI, Vertex AI, model enablement & deployment, golden path, LLM tuning & training, and data engineering using cloud-based solutions. I design and implement GenAI-based solutions to drive business impact, collaborating with cross-functional teams to push the boundaries of AI and data science in enterprise settings.

Arun Kumar Mahuri

I am an innovative and results-driven Lead ML Engineer with 20 years of IT experience, including 9 years specializing in Gen AI, Vertex AI, model enablement & deployment, golden path, LLM tuning & training, and data engineering using cloud-based solutions. I design and implement GenAI-based solutions to drive business impact, collaborating with cross-functional teams to push the boundaries of AI and data science in enterprise settings.

Available to hire

I am an innovative and results-driven Lead ML Engineer with 20 years of IT experience, including 9 years specializing in Gen AI, Vertex AI, model enablement & deployment, golden path, LLM tuning & training, and data engineering using cloud-based solutions.

I design and implement GenAI-based solutions to drive business impact, collaborating with cross-functional teams to push the boundaries of AI and data science in enterprise settings.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Work Experience

Automated ETL Pipeline – Audit Data Processing & Cloud Migration at Lloyds Banking Group
May 1, 2022 - Present
Designed and developed automated ETL tools improving efficiency and accuracy. Migrated legacy applications to GCP using Terraform, Docker, GKE, and BigQuery, modernizing data processing infrastructure. Built AI/ML agentic APIs with Spacy and TensorFlow for data mining, predictive analytics, and anomaly detection in audit reports. Automated ingestion of input files into raw and staging layers with refined data pipelines. Implemented containerization and orchestration using Docker, Helm Chart, Istio, and GKE; managed workflows with Cloud Composer (Airflow). Converted SQL Server stored procedures and CTEs into optimized BigQuery SQL models. Established CI/CD pipelines for image build and deployment with Harness and Git.
LLM-based API – GenAI MLOps Enablement at Lloyds Banking Group
May 1, 2022 - Present
Led LLM model enablement, deployment, and inference for Google-managed and open-source models, providing secure API access for data scientists and engineers. Architected GenAI solutions using Vertex AI, Model Garden, Golden Path, Terraform, GKE, Helm, Docker, and Apigee. Trained Google-managed LLMs with SFT and executed custom training jobs. Configured Golden Path for federated deployment across teams. Managed dataset preparation, training workflows, model transfer using VPC bridging. Automated PRs and cloud infrastructure using Terraform and Harness pipelines.
Chat Agent – AI Data Discovery Platform at Lloyds Banking Group
May 1, 2022 - Present
Designed and developed an AI-powered data discovery chat agent leveraging an in-house LLM-based API, RAG pipeline, VectorDB, embeddings, and FastAPI within GCP. Built as a multi-agent application using ADK and VertexAI, deployed on GKE for scalability. Defined high-level architecture to align with enterprise requirements. Implemented a RAG pipeline ingesting data from Collibra, applying semantic chunking, generating embeddings, and storing them in pgVector. Used LLM-based Cortex API for embedding generation and GenAI model integration. Created multiple specialized agents orchestrated by a primary controller agent to enable agentic conversations. Engineered query flow: user input converted to embeddings → semantic search in VectorDB → prompt construction → response generation via LLM → final output to user.
Data Engineer at Goldman Sachs
January 1, 2018 - April 1, 2022
Developed scalable data pipelines for processing and reconciling large equity derivative datasets (Swaps & Synthetics) using PySpark and Airflow. Built and maintained ETL pipelines with AWS Glue, Snowflake, and S3, improving data ingestion and processing efficiency. Optimized SQL-based data models for performance. Automated pipeline scheduling with Autosys and monitoring via Procmon. Collaborated with cross-functional teams to design frameworks for in-memory data handling and real-time processing.
Assistant Consultant at Tata Consultancy Services (TCS)
November 25, 2015 - Present
Assisted in project delivery across multiple engagements, contributing to solution design and implementation activities.
Technology Lead at Infosys Limited
October 3, 2011 - November 19, 2015
Led technology initiatives delivering data & analytics solutions; collaborated with cross-functional teams including architecture and delivery to ensure high-quality, scalable software solutions.
Sr. Application Developer at IBM India Pvt Ltd.
June 22, 2009 - September 27, 2011
Developed enterprise applications, contributing to software engineering activities and delivering robust, scalable solutions.
Software Engineer at Sasken Communications
November 22, 2007 - June 12, 2009
Engineered software solutions across diverse domains, focusing on performance and reliability.
IT Engineer at CMC Ltd.
January 2, 2006 - October 31, 2007
Provided IT engineering support and development contributions across projects.

Education

Bachelor of Engineering in Computer Science and Engineering at BPUT, India
January 11, 2030 - January 6, 2026

Qualifications

Google Certified Professional Data Engineer
January 11, 2030 - January 6, 2026

Industry Experience

Financial Services, Professional Services, Software & Internet