Data Engineer with 8+ years designing and running ETL/ELT pipelines and data platforms for healthcare, banking, and telecom organizations. Recent focus has been connecting core data engineering with AI/ML delivery—building retrieval, feature, and ingestion pipelines that put LLM and ML tools into production. Comfortable owning the full data lifecycle, from raw ingestion through pipelines that feed a model, dashboard, or LLM-based tool. Delivered measurable results in reducing manual review effort, improving batch reliability, and partnering with compliance, data science, and business stakeholders in regulated environments (HIPAA, financial governance).

Tejaswini Divi

Data Engineer with 8+ years designing and running ETL/ELT pipelines and data platforms for healthcare, banking, and telecom organizations. Recent focus has been connecting core data engineering with AI/ML delivery—building retrieval, feature, and ingestion pipelines that put LLM and ML tools into production. Comfortable owning the full data lifecycle, from raw ingestion through pipelines that feed a model, dashboard, or LLM-based tool. Delivered measurable results in reducing manual review effort, improving batch reliability, and partnering with compliance, data science, and business stakeholders in regulated environments (HIPAA, financial governance).

Available to hire

Data Engineer with 8+ years designing and running ETL/ELT pipelines and data platforms for healthcare, banking, and telecom organizations. Recent focus has been connecting core data engineering with AI/ML delivery—building retrieval, feature, and ingestion pipelines that put LLM and ML tools into production.

Comfortable owning the full data lifecycle, from raw ingestion through pipelines that feed a model, dashboard, or LLM-based tool. Delivered measurable results in reducing manual review effort, improving batch reliability, and partnering with compliance, data science, and business stakeholders in regulated environments (HIPAA, financial governance).

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Beginner
Beginner
See more

Language

English
Advanced

Work Experience

Data Engineer – AI/ML at UHG
January 1, 2024 - Present
Built LLM-powered retrieval pipelines and underlying data infrastructure for AI-assisted claims processing and clinical documentation tools. Engineered RAG pipelines using LangChain with Pinecone and Azure OpenAI/AWS Bedrock for clinical-note summarization and eligibility-check use cases, tuned chunking for clinical formatting. Owned feature pipelines for readmission risk and fraud anomaly models including deduplication, schema validation, and refresh scheduling. Deployed services as containerized REST APIs (Flask/FastAPI) and set up MLflow-based retraining pipelines with CI/CD to reduce manual rollback issues. Implemented HIPAA-aligned controls including PHI masking before indexing, access controls, and audit logging; added validation on RAG outputs and retrieval-quality drift monitoring. Optimized SQL transformations to reduce preprocessing time and orchestrated daily loads with Airflow so downstream models and the retrieval index used same-day data. Built extracts/dashboards in Inco
Data Engineer at Bank of America
April 1, 2021 - November 30, 2023
Built data and preprocessing pipelines behind a computer-vision-based document intelligence initiative for mortgage operations. Developed image preprocessing for scanned documents (deskewing, binarization, noise removal, layout detection with OpenCV) feeding OCR/NLP extraction models; reduced manual review queues by ~40%. Merged structured borrower/origination data from PostgreSQL with unstructured vision-derived outputs in S3 to provide consistent features for downstream risk scoring. Benchmarked OCR/CV architectures on sample sets and produced error-analysis to support model selection. Created geospatial fraud-detection features using GeoPandas and Shapely integrated via PySpark spatial joins. Implemented MLflow and GitHub Actions for bi-weekly retraining, metric reporting, and Slack alerts when AUC dropped beyond threshold; included documentation/lineage for model risk management. Automated data quality checks and supported reporting via Power BI and Incorta. Fed time-series feature
Senior Data Engineer / Data Engineer at AT&T
November 1, 2019 - February 28, 2021
Designed ingestion and transformation pipelines for subscriber analytics, churn prediction, and network performance monitoring. Architected Hadoop-based pipelines (HDFS, Hive, Pig, Sqoop) pulling structured/semi-structured OSS/BSS and network data; delivered the backbone for downstream churn and network-KPI models. Built feature engineering pipelines in collaboration with data science teams, and led distributed feature workflows on Databricks to process CDR/telemetry/subscriber datasets at scale, improving throughput versus Hadoop-only setup. Deployed batch and near real-time scoring pipelines on AWS EC2 in a hybrid on-prem/cloud environment for near real-time ops actions. Built supporting NLP pipelines for service tickets/transcripts and implemented data quality/validation frameworks to prevent malformed CDR batches from reaching models. Established governance and documentation standards for reproducible, auditable analytics workflows.
Data Engineer at Freshworks
July 1, 2017 - June 30, 2019
Built backend data pipelines and reporting infrastructure for enterprise analytics applications used by finance and retail clients. Replaced Excel-based reporting with automated, queryable SQL-backed pipelines. Developed Python backend modules and REST APIs to move data between source systems and SQL databases; implemented SQL extraction/transformation logic and automated reporting scripts. Packaged supervised ML models (feature engineering, cross-validation, packaging) to automate client manual data-tagging. Deployed pipeline components in Docker, maintained CI/CD, documented API contracts, and supported requirements sessions to translate business needs into schema/query logic. Conducted code reviews, optimized nightly batch queries, and added unit tests to reduce recurring production issues.

Education

Bachelor of Technology in Computer Science at KL University
June 1, 2013 - May 31, 2017

Qualifications

Add your qualifications or awards here.

Industry Experience

Healthcare, Financial Services, Telecommunications, Software & Internet

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Beginner
Beginner
See more