I am a Data and Generative AI Engineer with 5+ years of experience across AWS, GCP, and Azure, building scalable ETL, data lake, and data warehouse architectures for financial services and retail. I designed enterprise-grade RAG pipelines using Amazon Bedrock to enable secure LLM-driven semantic search and contextual document retrieval for advisor research and compliance operations. I architected scalable data engineering frameworks leveraging AWS Glue, Cloud Dataflow, and Azure Data Factory to standardize ingestion, transformation, and curation of structured and unstructured enterprise datasets. I have built unified Customer 360 platforms in BigQuery integrating transactional and behavioral data to power advanced segmentation models and marketing analytics. I developed distributed processing solutions using Spark and cloud-native services to optimize batch processing, improve data quality, and enable reliable downstream analytics and machine learning workloads. I focus on secure data access controls, robust monitoring, and collaboration with data scientists and business stakeholders to translate requirements into production-ready data models and AI-enabled analytics solutions, while modernizing legacy systems to support Generative AI at scale.

Charitha Veeramachaneni

I am a Data and Generative AI Engineer with 5+ years of experience across AWS, GCP, and Azure, building scalable ETL, data lake, and data warehouse architectures for financial services and retail. I designed enterprise-grade RAG pipelines using Amazon Bedrock to enable secure LLM-driven semantic search and contextual document retrieval for advisor research and compliance operations. I architected scalable data engineering frameworks leveraging AWS Glue, Cloud Dataflow, and Azure Data Factory to standardize ingestion, transformation, and curation of structured and unstructured enterprise datasets. I have built unified Customer 360 platforms in BigQuery integrating transactional and behavioral data to power advanced segmentation models and marketing analytics. I developed distributed processing solutions using Spark and cloud-native services to optimize batch processing, improve data quality, and enable reliable downstream analytics and machine learning workloads. I focus on secure data access controls, robust monitoring, and collaboration with data scientists and business stakeholders to translate requirements into production-ready data models and AI-enabled analytics solutions, while modernizing legacy systems to support Generative AI at scale.

Available to hire

I am a Data and Generative AI Engineer with 5+ years of experience across AWS, GCP, and Azure, building scalable ETL, data lake, and data warehouse architectures for financial services and retail. I designed enterprise-grade RAG pipelines using Amazon Bedrock to enable secure LLM-driven semantic search and contextual document retrieval for advisor research and compliance operations. I architected scalable data engineering frameworks leveraging AWS Glue, Cloud Dataflow, and Azure Data Factory to standardize ingestion, transformation, and curation of structured and unstructured enterprise datasets.

I have built unified Customer 360 platforms in BigQuery integrating transactional and behavioral data to power advanced segmentation models and marketing analytics. I developed distributed processing solutions using Spark and cloud-native services to optimize batch processing, improve data quality, and enable reliable downstream analytics and machine learning workloads. I focus on secure data access controls, robust monitoring, and collaboration with data scientists and business stakeholders to translate requirements into production-ready data models and AI-enabled analytics solutions, while modernizing legacy systems to support Generative AI at scale.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert

Work Experience

Data and Generative AI Engineer at Charles Schwab
July 1, 2024 - Present
Led secure AWS Financial RAG Assistant enabling semantic search across research and client documents. Built end-to-end RAG pipelines on AWS using Glue to ingest 18M documents from 9 enterprise systems into S3, processing 6TB of structured and unstructured content for advisor and compliance access. Implemented document chunking with Lambda generating 140M text chunks and 512-token windows, improving contextual retrieval accuracy. Built vector embeddings with Bedrock to create and store 140M embeddings, enabling semantic search queries with average latency under 2 seconds for 120 advisors. Engineered metadata-driven ingestion across 25 pipelines orchestrated with Glue, reducing onboarding time from 3 weeks to 6 days. Optimized storage across 6TB of data in S3, reducing retrieval overhead by 27%. Developed access-controlled APIs via Lambda for 15K daily queries with RBAC across 8 advisor groups. Implemented monitoring with CloudWatch across 60+ workflows, achieving 99.9% SLA. Built automa
Data Engineer at The Home Depot
July 1, 2023 - May 1, 2024
Built scalable GCP Customer 360 platform enabling segmentation, personalization, recommendations & campaign analytics. Designed end-to-end ETL pipelines using Cloud Dataflow to ingest 120M daily transactions from 18 systems into BigQuery, standardizing identifiers and enabling a unified Customer 360 view across 2 channels. Built a centralized data warehouse in BigQuery processing 1.5B historical records with partitioning across 36 months, reducing complex join query time from 22 minutes to 6 minutes for marketing analytics teams. Orchestrated 25+ pipelines with Cloud Composer across 4 environments, ensuring refresh cycles under 4 hours for segmentation and attribution workflows. Implemented identity resolution across 220M profiles across 5 sources, improving match accuracy by 31%. Engineered 450+ behavioral features for recommendation models and targeted promotions across 3 business units. Developed incremental ingestion frameworks processing 10K+ files per day, reducing redundant load
Data Engineer at GVK Biosciences
May 1, 2020 - July 1, 2022
Built Azure-based genomics and drug discovery data platform enabling scalable ETL, feature engineering, and ML-ready research analytics pipelines. Ingested 15TB of genomic FASTQ and VCF files monthly from 8 sequencing systems, orchestrating 25 workflows that standardized metadata for downstream analytics. Engineered distributed transformations in Azure Databricks using Spark to process 2B genomic records across 5 clusters, reducing batch runtime to 6 hours. Built curated layers in Azure Data Lake storing 25TB of structured and semi-structured data across 40+ domains to improve query efficiency. Automated data validation rules across 150+ checks in Azure Databricks, improving data reliability by 28%. Created compound screening ingestion pipelines to integrate 12M assay results for SAR modeling. Optimized Spark logic to process 500M molecular features daily for scalable feature pipelines. Implemented orchestration monitoring with Azure Monitor across 70+ jobs in 3 environments, achieving

Education

Add your educational history here.

Qualifications

Master's in Computer and Information Science
January 11, 2030 - May 1, 2024
AWS Certified Solutions Architect – Professional
January 11, 2030 - June 29, 2026
AWS Certified Machine Learning – Specialty
January 11, 2030 - June 29, 2026
Google Professional Data Engineer
January 11, 2030 - June 29, 2026
Microsoft Certified: Azure Data Engineer Associate
January 11, 2030 - June 29, 2026

Industry Experience

Financial Services, Retail

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert

Hire a Data Scientist

We have the best data scientist experts on Twine. Hire a data scientist today.