Senior Data Scientist / GenAI/ML Engineer with 11+ years in IT and 8+ years building production ML systems, including 3+ years focused on Generative AI and Agentic AI. Experienced in architecting and deploying enterprise-grade LLM solutions across complex environments, with hands-on expertise in agent orchestration frameworks and evaluation/monitoring for regulated deployments. Deep expertise in Python/Java/Go, Azure and AWS cloud deployments (Azure ML, OpenAI, Functions, Kubernetes; SageMaker, Bedrock, Lambda, OpenSearch), and end-to-end RAG/embeddings/prompt & context engineering. Proven track record delivering measurable business outcomes in finance and healthcare, including fairness/governance, human-in-the-loop workflows, and scalable MLOps/ETL pipelines.

Arun Gilla

Senior Data Scientist / GenAI/ML Engineer with 11+ years in IT and 8+ years building production ML systems, including 3+ years focused on Generative AI and Agentic AI. Experienced in architecting and deploying enterprise-grade LLM solutions across complex environments, with hands-on expertise in agent orchestration frameworks and evaluation/monitoring for regulated deployments. Deep expertise in Python/Java/Go, Azure and AWS cloud deployments (Azure ML, OpenAI, Functions, Kubernetes; SageMaker, Bedrock, Lambda, OpenSearch), and end-to-end RAG/embeddings/prompt & context engineering. Proven track record delivering measurable business outcomes in finance and healthcare, including fairness/governance, human-in-the-loop workflows, and scalable MLOps/ETL pipelines.

Available to hire

Senior Data Scientist / GenAI/ML Engineer with 11+ years in IT and 8+ years building production ML systems, including 3+ years focused on Generative AI and Agentic AI. Experienced in architecting and deploying enterprise-grade LLM solutions across complex environments, with hands-on expertise in agent orchestration frameworks and evaluation/monitoring for regulated deployments.

Deep expertise in Python/Java/Go, Azure and AWS cloud deployments (Azure ML, OpenAI, Functions, Kubernetes; SageMaker, Bedrock, Lambda, OpenSearch), and end-to-end RAG/embeddings/prompt & context engineering. Proven track record delivering measurable business outcomes in finance and healthcare, including fairness/governance, human-in-the-loop workflows, and scalable MLOps/ETL pipelines.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
See more

Language

English
Advanced

Work Experience

Senior Data Scientist at Citizens Financial Group
August 1, 2022 - Present
Built AI/ML-driven customer risk & engagement prioritization in a regulated financial environment, replacing rule-based segmentation with stacked models (XGBoost + TabNet) and embedding-based representation for high-cardinality features. Engineered a PySpark feature pipeline (280+ features), deployed batch scoring at scale with SageMaker Batch Transform, and integrated results into CRM/engagement workflows via S3/Lambda/Snowflake. Implemented governance and monitoring using SageMaker Model Monitor and CloudWatch, with explainability and fairness audits using SHAP; addressed bias impacts across demographic segments and improved targeting outcomes. Led development of an agentic RAG document intelligence platform for loan/credit workflows using LangGraph multi-agent orchestration, Bedrock (Claude 3 Sonnet), AWS Textract, Pydantic structured extraction, and OpenSearch/Titan embeddings. Added audit-ready summaries with traceable citations, enforced human-in-the-loop validation, and ensured
Senior Data Scientist / MLOps Engineer at Citizens Financial Group
August 1, 2022 - Present
Led development of AI/ML customer risk & engagement prioritization within regulated financial services. Built PySpark feature pipelines (280+ features) integrating transactional, credit bureau, digital interaction, call center, and external socio-economic data. Implemented XGBoost and TabNet models and a stacked ensemble with logistic regression meta-learner (AUC improvement to 0.843). Developed embeddings via Gensim Word2Vec for high-cardinality categorical features and standardized them into a reusable internal package. Deployed large-scale batch scoring using AWS SageMaker Batch Transform for 4M+ records/day, with MLflow-based tracking/versioning and audit-ready governance. Established MLOps with SageMaker Model Monitor, CloudWatch, Prometheus/Grafana, PagerDuty, Docker, Kubernetes, and CI/CD (GitHub, CodePipeline, CDK) to support monitoring, drift detection, and retraining. Performed SHAP explainability and fairness audits and addressed bias impacting underbanked/rural segments. A
Senior Data Scientist at HCA Healthcare
June 1, 2019 - July 1, 2022
Established enterprise data science capabilities, modernizing legacy Excel reporting into cloud-based analytics. Built the first cloud data warehouse using AWS Redshift and created Python ETL pipelines for multi-year healthcare ingestion. Developed a Patient 360 model consolidating demographics, visit history, clinical utilization, payer mix, and facility interactions. Implemented record linkage to reduce duplicates (~22%) and recovered fragmented histories (~17,000). Built segmentation using K-Means (RFM cohorts) and automated monthly refresh pipelines to CRM/outreach systems. Developed CLV models (BG/NBD, Gamma-Gamma) to support preventive care prioritization and re-engagement campaigns. Created demand forecasting for resource utilization across multiple locations using Prophet, ARIMA, and moving-average models with external regressors; improved MAPE and mitigated cold-start using similarity-based matching. Incorporated DevOps best practices (Git, CI/CD, IaC concepts) and delivered
Machine Learning Engineer at Comcast
September 1, 2016 - May 1, 2019
Built predictive analytics infrastructure for network service degradation and 30-day churn risk across broadband/video/enterprise services. Developed end-to-end ingestion from OSS/BSS telemetry streams into Teradata staging via Informatica, and engineered 120+ predictive features using PySpark with Hive/Oozie orchestration. Trained Logistic Regression and Gradient Boosted Tree models using rolling 18-month cohorts (AUC-ROC ~0.78) with precision-recall and calibration. Contributed to network reliability early-warning anomaly detection using time-series trend and spike features. Handled feature schema migration and taxonomy changes using crosswalks and feature flags to restore model performance. Migrated batch scoring from single-node Python to distributed PySpark, reducing processing time for ~80,000 daily accounts from 4+ hours to under 25 minutes. Implemented governance, deployment patterns, and stakeholder support for operational risk scoring and dashboards; optimized throughput ~10
Tableau Developer / Data Analyst at eBay
February 1, 2014 - August 1, 2016
Designed and developed Tableau dashboards for weekly GMV, category margin, fulfillment rate, and return/fraud loss metrics. Built/optimized Teradata SQL layers integrating Oracle and buyer engagement data marts; resolved performance bottlenecks using pre-aggregated summary tables. Implemented Informatica-driven nightly ETL to refresh summaries, reducing extract/query runtime from 90 minutes to 22 minutes. Standardized canonical metrics across business units (canonical KPI layer) and configured Tableau Server projects, user groups, and row-level security for regional isolation. Improved Tableau Server performance under peak concurrency by switching to scheduled extracts and tuning backgrounder workloads. Increased adoption from 30 to 200+ named users, and reduced manual Excel reporting effort by ~15–20 hours/week through automation and reusable onboarding frameworks.

Education

Master’s in Information Technology at University Of North Texas
January 1, 2013 - December 1, 2013
Master’s in Information Technology at University Of North Texas
January 11, 2030 - December 1, 2013

Qualifications

Databricks Data Engineer Associate
January 11, 2030 - July 8, 2026
Databricks Machine Learning Engineer Associate
January 11, 2030 - July 8, 2026
Databricks Generative AI Engineer Associate
January 11, 2030 - July 8, 2026
Databricks Associate ML Engineer
January 11, 2030 - July 23, 2026
Databricks Associate GenAI Engineer
January 11, 2030 - July 23, 2026
Databricks Data Engineer
January 11, 2030 - July 23, 2026

Industry Experience

Financial Services, Healthcare, Telecommunications, Software & Internet