Data Architecture & Machine Learning Engineer with 4+ years of experience designing schema architectures, structured metadata models, and data-modeling frameworks that power large-scale retrieval and analytics systems. Skilled in translating business and compliance requirements into deterministic, well-documented data models including semantic tagging, taxonomy design, and relationship mapping across ingestion and retrieval pipelines. Experienced evaluating and selecting scalable platforms (vector databases, cloud data platforms, orchestration frameworks) and building GenAI/RAG systems using tool-calling and agent orchestration. Proven track record partnering with data engineering and compliance stakeholders to improve retrieval precision, groundedness, and operational reliability while ensuring auditability and security across the ML lifecycle.

DIVAKAR SARAGADAM

Data Architecture & Machine Learning Engineer with 4+ years of experience designing schema architectures, structured metadata models, and data-modeling frameworks that power large-scale retrieval and analytics systems. Skilled in translating business and compliance requirements into deterministic, well-documented data models including semantic tagging, taxonomy design, and relationship mapping across ingestion and retrieval pipelines. Experienced evaluating and selecting scalable platforms (vector databases, cloud data platforms, orchestration frameworks) and building GenAI/RAG systems using tool-calling and agent orchestration. Proven track record partnering with data engineering and compliance stakeholders to improve retrieval precision, groundedness, and operational reliability while ensuring auditability and security across the ML lifecycle.

Available to hire

Data Architecture & Machine Learning Engineer with 4+ years of experience designing schema architectures, structured metadata models, and data-modeling frameworks that power large-scale retrieval and analytics systems. Skilled in translating business and compliance requirements into deterministic, well-documented data models including semantic tagging, taxonomy design, and relationship mapping across ingestion and retrieval pipelines.

Experienced evaluating and selecting scalable platforms (vector databases, cloud data platforms, orchestration frameworks) and building GenAI/RAG systems using tool-calling and agent orchestration. Proven track record partnering with data engineering and compliance stakeholders to improve retrieval precision, groundedness, and operational reliability while ensuring auditability and security across the ML lifecycle.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
See more

Work Experience

GenAI / Data Architecture Engineer at Onbe
September 1, 2025 - Present
Architected a structured metadata and schema-tagging framework for a Retrieval-Augmented Generation (RAG) platform supporting 150,000+ users, reducing document lookup time from 20–30 minutes to under 8 seconds. Defined a two-phase architecture separating offline schema/indexing from runtime retrieval, enforcing deterministic data boundaries. Collaborated with data engineers and compliance subject-matter experts to build ingestion pipelines processing 10,000+ regulatory documents with version-aware schema updates. Designed semantic tagging and chunking taxonomy (600–800 tokens) with structured metadata attributes, improving retrieval precision. Built and maintained a pgvector-based vector store with relationship-aware metadata filtering in a VPC-isolated environment, improving grounded response accuracy. Evaluated and standardized GenAI tooling for multi-agent orchestration (LangGraph/LangChain/MCP). Implemented secure backend services (FastAPI, OAuth2/JWT/RBAC, IAM) and operational
GenAI / Data Architecture Engineer
September 1, 2025 - Present
Architected a structured metadata and schema-tagging framework for a Retrieval-Augmented Generation (RAG) platform serving 150,000+ users, reducing document lookup time from 20–30 minutes to under 8 seconds (~60% efficiency gain). Designed a two-phase architecture separating offline schema/indexing from runtime retrieval to enforce deterministic data boundaries and reduce compliance escalations (~50%). Built ingestion pipelines for 10,000+ regulatory documents with version-aware schema updates using Cloud Dataflow/Apache Beam and Pub/Sub/BigQuery. Developed semantic tagging and chunking taxonomy (600–800 token units) with structured metadata attributes to improve retrieval precision (~32%). Implemented a vector-based knowledge store in PostgreSQL/pgvector with relationship-aware metadata filtering in a VPC-isolated environment, improving grounded response accuracy (~38%). Evaluated and standardized GenAI orchestration tools (LangGraph, LangChain, MCP) and implemented secure backend
Data Modeling / ML Engineer at Tata Consultancy Services Limited, Mumbai, India
June 1, 2023 - July 31, 2024
Designed a structured data model for a credit-risk early warning system across millions of accounts, improving delinquency detection precision and reducing false positives. Built governed, reproducible feature/schema pipelines supporting regulatory compliance and auditability. Performed structural data analysis (Python/Pandas/NumPy/SQL) for outlier detection, missing-value handling, and schema-consistency checks. Implemented schema-aware ingestion pipelines using Azure Data Factory processing 3–5 million records daily with incremental loading and schema evolution into ADLS Gen2. Orchestrated event-driven workflows with Azure Functions to trigger schema validation and scoring pipelines. Built distributed modeling workflows with Azure Databricks/PySpark, documented model logic, and exposed scoring services via Azure API Management. Implemented schema-drift monitoring to trigger controlled retraining workflows and reduced job failures through monitoring (Azure Monitor/Log Analytics/ELK)
Data Modeling / ML Engineer at Tata Consultancy Services Limited
June 1, 2023 - July 1, 2024
Designed structured credit-risk early warning data models for millions of accounts, improving delinquency detection precision (~11%) and reducing false positives (~17%). Built governed and reproducible feature/schema pipelines to meet regulatory compliance, data transparency, and auditability requirements. Performed structural data analysis using Python/Pandas/NumPy/SQL for outliers, missing values, and schema-consistency checks. Implemented schema-aware ingestion pipelines using Azure Data Factory processing 3–5 million records daily into ADLS Gen2 with incremental loading and schema evolution handling. Orchestrated event-driven workflows using Azure Functions to trigger schema validation and scoring pipelines. Developed distributed data-modeling workflows with Azure Databricks/PySpark to structure behavioral relationship attributes and improve processing efficiency (~23%). Implemented schema-drift monitoring to trigger controlled retraining while maintaining structural parity and a
Data Modeling / ML Engineer at LanceSoft Engineering, Hyderabad, India
March 1, 2021 - May 31, 2023
Structured daily data architectures processing multi-terabyte datasets to ensure reliable, schema-consistent access for 2,000+ users. Built and scaled an AWS data lake on S3 to modernize legacy schema/ETL pipelines and improve reporting availability. Designed streaming/batch schema-validated ingestion pipelines using Kinesis, Lambda, and Glue for near-real-time visibility into airline operations. Developed distributed data-processing workflows with PySpark/EMR for cleaning, normalization, and relationship partitioning to improve query performance and reduce batch processing time. Modeled predictive delay/disruption features using scikit-learn, XGBoost, and TensorFlow, defining relationship variables such as delay propagation risk. Documented and versioned data models via SageMaker Pipelines for reproducible and auditable deployment. Integrated structured outputs into Redshift/Snowflake for operational insights and implemented CI/CD with Git/Jenkins, Docker, and Kubernetes (EKS), improv
Data Modeling / ML Engineer at LanceSoft Engineering
March 1, 2021 - May 1, 2023
Structured daily data architectures processing multi-terabyte datasets for reliable, schema-consistent access by 2,000+ analytics and operations users. Built and scaled an AWS data lake on Amazon S3, modernizing legacy schema/ETL and improving reporting availability (~35%). Designed streaming and batch schema-validated ingestion pipelines using Kinesis, Lambda, and Glue to enable near real-time airline operations visibility. Created distributed data-processing workflows in PySpark/EMR for cleaning, normalization, and relationship partitioning to improve query performance and reduce batch processing time. Modeled predictive delay/disruption features (scikit-learn, XGBoost, TensorFlow) with structured relationship variables such as delay-propagation risk. Documented and versioned data models using SageMaker Pipelines for reproducible, auditable deployments; integrated structured outputs into Redshift/Snowflake for operational insights. Implemented CI/CD with Git/Jenkins, containerized se

Education

Master of Science, Computer Science at Rivier University
September 1, 2024 - May 1, 2026
Master of Science, Computer Science at Rivier University
September 1, 2024 - May 1, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Transportation & Logistics