Data engineering professional with 6+ years of experience building scalable batch and streaming data platforms on Azure and AWS. Expertise includes Databricks cluster/workspace optimization, ETL/ELT pipelines (ADF, Glue, dbt), data warehousing (Snowflake, Synapse), and governance with Azure Purview and secure credential management (Key Vault/Secrets Manager).

Karthik Banka

Data engineering professional with 6+ years of experience building scalable batch and streaming data platforms on Azure and AWS. Expertise includes Databricks cluster/workspace optimization, ETL/ELT pipelines (ADF, Glue, dbt), data warehousing (Snowflake, Synapse), and governance with Azure Purview and secure credential management (Key Vault/Secrets Manager).

Available to hire

Data engineering professional with 6+ years of experience building scalable batch and streaming data platforms on Azure and AWS. Expertise includes Databricks cluster/workspace optimization, ETL/ELT pipelines (ADF, Glue, dbt), data warehousing (Snowflake, Synapse), and governance with Azure Purview and secure credential management (Key Vault/Secrets Manager).

See more

Language

Work Experience

Data Engineer at Bayer Corp
February 1, 2024 - Present
Architected an intelligent data quality framework in Databricks with ML-based anomaly detection for real-time drift detection, schema validation, and inconsistency identification across 10TB+ datasets. Built an Azure OpenAI + Azure AI Search RAG pipeline to index pipeline metadata, lineage graphs, and anomaly reports for natural-language querying, reducing analyst escalations by ~30%. Designed fault-tolerant batch and real-time streaming pipelines using Azure Event Hubs, Kafka, and Databricks Structured Streaming to meet SLA targets. Integrated Azure OpenAI into dbt workflows to auto-generate and validate SQL transformations from natural language, reducing model development time by 35%. Engineered Spark workload profiling to recommend partition sizes and AQE tuning, reducing job latency by 20%. Implemented dynamic Airflow DAG building and integrated dbt runs/tests into Airflow orchestration. Led modernization of legacy systems to a modern big data architecture, reducing processing time
Data Engineer at SADA Systems
May 1, 2023 - January 31, 2024
Implemented a large-scale Azure Data Lake + Databricks data lake for IoT sensor data using partitioning strategies and lifecycle management to optimize storage costs and query performance. Developed dbt models for Snowflake, including unit tests and documentation to improve reliability and governance. Built real-time streaming analytics using Azure Event Hubs, Azure Stream Analytics, and Power BI to reduce reporting latency. Designed serverless event-driven pipelines with Azure Functions and Logic Apps. Implemented real-time Kafka pipelines for application logs/IoT ingestion. Automated ML pipelines using Azure ML and MLflow for training/versioning/deployment. Performed complex Snowflake transformations with SQL, UDFs, and stored procedures. Built custom Airflow operators to trigger Databricks/AWS Glue/ADF pipelines. Led migration of on-prem Hadoop clusters to Azure HDInsight, reducing infrastructure costs by 20%. Mentored a team of 4 junior engineers on AWS/Azure best practices.
Data Engineer at Enphase Energy
March 1, 2020 - July 31, 2022
Migrated an on-prem data warehouse to Azure Synapse and AWS Redshift using Azure Data Factory and AWS Glue, applying partitioning and indexing to improve query performance and reduce costs. Built CI/CD pipelines for Databricks notebooks, ADF pipelines, and AWS Lambda using Azure DevOps and GitHub. Implemented data quality checks in Databricks and ADF using schema validation, profiling, and business-rule validations. Optimized Spark and EMR workloads via autoscaling, instance selection, and code refactoring, reducing processing times and improving resource utilization. Loaded star/snowflake-style dimensional datasets into Synapse DW and Redshift. Implemented data masking and row-level security in Azure SQL Database and AWS RDS (including dynamic data masking and security roles). Enabled Delta Lake (ACID transactions and time travel) for reliable historical change handling.
Data Engineer Intern at Riverbed Technologies
March 1, 2019 - February 29, 2020
Developed Power BI dashboards and reports using Azure SQL Database data. Wrote complex SQL transformations using window functions, CTEs, and indexed views to optimize performance. Collaborated with data scientists to prepare ML datasets using Databricks notebooks (PySpark). Supported a data quality framework with validation checks and schema enforcement using Databricks and ADF. Participated in migration of on-prem SQL Server databases to Azure SQL Database. Developed Spark script to flatten deeply nested JSON. Assisted with security measures including encryption and RBAC in Azure SQL Database. Supported automated CICD pipeline for DAG deployment.

Education

Add your educational history here.

Qualifications

DP-203 Microsoft Certified Azure Data Engineer Associate
January 11, 2030 - July 23, 2026

Industry Experience

Energy & Utilities, Healthcare, Life Sciences, Professional Services