I am a Senior Data Engineer with over 5 years of experience specializing in building scalable cloud data platforms and Generative AI (RAG) initiatives. My expertise spans both the Azure and AWS ecosystems, including hands-on experience with Databricks, Snowflake, PySpark, and dbt. At Citibank and ADP, I have successfully architected ETL pipelines for multi-terabyte datasets and implemented real-time streaming solutions for millions of daily events. I stand out for my ability to deliver significant operational improvements, such as reducing query runtimes by 35% and automating compliance workflows through AI. I am dedicated to providing high-performance, AI-ready infrastructure that drives business intelligence and enterprise analytics.

TEJA B

I am a Senior Data Engineer with over 5 years of experience specializing in building scalable cloud data platforms and Generative AI (RAG) initiatives. My expertise spans both the Azure and AWS ecosystems, including hands-on experience with Databricks, Snowflake, PySpark, and dbt. At Citibank and ADP, I have successfully architected ETL pipelines for multi-terabyte datasets and implemented real-time streaming solutions for millions of daily events. I stand out for my ability to deliver significant operational improvements, such as reducing query runtimes by 35% and automating compliance workflows through AI. I am dedicated to providing high-performance, AI-ready infrastructure that drives business intelligence and enterprise analytics.

Available to hire

I am a Senior Data Engineer with over 5 years of experience specializing in building scalable cloud data platforms and Generative AI (RAG) initiatives. My expertise spans both the Azure and AWS ecosystems, including hands-on experience with Databricks, Snowflake, PySpark, and dbt.
At Citibank and ADP, I have successfully architected ETL pipelines for multi-terabyte datasets and implemented real-time streaming solutions for millions of daily events. I stand out for my ability to deliver significant operational improvements, such as reducing query runtimes by 35% and automating compliance workflows through AI. I am dedicated to providing high-performance, AI-ready infrastructure that drives business intelligence and enterprise analytics.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Beginner
Beginner
See more

Work Experience

Data Engineer at Citibank Toronto
April 1, 2024 - Present
Developed cloud-native ETL pipelines using Azure Databricks, PySpark, and Azure Data Factory to process 4TB+ of monthly banking transaction data supporting fraud detection analytics. Architected a Generative AI (RAG) orchestration layer in Databricks using Azure OpenAI to automate compliance audits of unstructured financial documents, reducing review cycles by 40%. Designed Delta Lake architecture with partitioning and Z-order optimization reducing Azure Synapse SQL runtimes by 35% for BI. Implemented real-time event streaming with Azure Event Hubs and Spark Streaming ingesting 500K+ daily transaction events. Collaborated with BI and risk teams to design SQL-based data models in Azure Synapse improving compliance reporting turnaround by 40%. Automated DataOps deployments using Azure DevOps, Jenkins CI/CD, and Git reducing manual release effort by 45%. Provisioned Azure data platform infrastructure via Terraform (Databricks clusters, ADLS, Function Apps, Docker, Kubernetes). Integrated
Data Pipeline Engineer at ADP Bengaluru
January 1, 2022 - August 31, 2023
Designed scalable ETL pipelines using AWS Glue, PySpark, and Apache Airflow to ingest payroll and workforce datasets into a centralized analytics platform. Engineered high-performance feature stores in Snowflake supporting predictive ML models for customer retention, contributing to a 17% accuracy improvement. Built a 50TB+ workforce data lake on Amazon S3 with IAM access controls and data governance policies. Engineered Snowflake and Amazon Redshift dimensional modeling pipelines improving BI query performance by 32%. Standardized workforce metrics using dbt transformation models in collaboration with analytics and HR stakeholders. Implemented event-driven ingestion using Kafka, MongoDB, and AWS streaming services for near real-time payroll events. Automated orchestration using Step Functions, Lambda, and Airflow reducing manual monitoring effort. Supported ML feature pipelines through SageMaker and Informatica IDMC workflows improving retention insights by 17%. Developed Tableau dash
Data Platform Engineer at Algocode Hyderabad
August 1, 2020 - December 31, 2021
Implemented distributed big data processing pipelines using Apache Spark and Scala for high-volume transactional dataset transformations. Built real-time event streaming pipelines with Kafka, AWS Kinesis, and Spark Streaming processing 10M+ application events per day. Orchestrated ETL workflows using Apache Airflow to schedule, monitor, and recover pipelines reducing failures by 40%. Designed scalable ingestion frameworks loading application data into Amazon S3 and Azure Data Lake Storage for analytics workloads. Developed optimized Snowflake tables and SQL transformations improving reporting performance by 22%.

Education

Postgraduate Diploma in Data Analytics and Business Decision Making at Canada Durham College
January 1, 2023 - January 1, 2024
Bachelor of Technology in Computer Science and Engineering at Andhra University
January 1, 2018 - January 1, 2022

Qualifications

Microsoft Certified: Azure Data Engineer Associate (DP-203)
January 11, 2030 - August 20, 2026
AWS Data Warehouse Specialization - Coursera
January 11, 2030 - August 20, 2026
Microsoft Certified: Azure Fundamentals (AZ-900)
January 11, 2030 - August 20, 2026
Big Data Analysis with Apache Spark - Coursera
January 11, 2030 - August 20, 2026

Industry Experience

Financial Services, Software & Internet, Education, Professional Services

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Beginner
Beginner
See more