Hi, I’m Deepak, a results-driven Data Engineer with 4+ years of experience building large-scale data ecosystems across telecom, banking, and healthcare. I design and optimize petabyte-scale pipelines, real-time streaming architectures, and cloud-native ETL/ELT solutions across AWS, Azure, and GCP. I’m proficient with Spark, Kafka, Flink, Databricks, Delta Lake, data warehousing (Snowflake, Redshift, BigQuery, Synapse), and orchestration (Airflow, Prefect, Dagster).\n\nI focus on governance, quality, and compliance (GDPR, HIPAA, SOX, TRAI) while enabling self-serve analytics and Data Mesh. I collaborate with data scientists to drive AI/ML insights for fraud detection, network optimization, and predictive analytics. I’ve delivered cost savings, reduced latency, and democratized data access across multi-cloud environments.

Hi, I’m Deepak, a results-driven Data Engineer with 4+ years of experience building large-scale data ecosystems across telecom, banking, and healthcare. I design and optimize petabyte-scale pipelines, real-time streaming architectures, and cloud-native ETL/ELT solutions across AWS, Azure, and GCP. I’m proficient with Spark, Kafka, Flink, Databricks, Delta Lake, data warehousing (Snowflake, Redshift, BigQuery, Synapse), and orchestration (Airflow, Prefect, Dagster).\n\nI focus on governance, quality, and compliance (GDPR, HIPAA, SOX, TRAI) while enabling self-serve analytics and Data Mesh. I collaborate with data scientists to drive AI/ML insights for fraud detection, network optimization, and predictive analytics. I’ve delivered cost savings, reduced latency, and democratized data access across multi-cloud environments.

Available to hire

Hi, I’m Deepak, a results-driven Data Engineer with 4+ years of experience building large-scale data ecosystems across telecom, banking, and healthcare. I design and optimize petabyte-scale pipelines, real-time streaming architectures, and cloud-native ETL/ELT solutions across AWS, Azure, and GCP. I’m proficient with Spark, Kafka, Flink, Databricks, Delta Lake, data warehousing (Snowflake, Redshift, BigQuery, Synapse), and orchestration (Airflow, Prefect, Dagster).\n\nI focus on governance, quality, and compliance (GDPR, HIPAA, SOX, TRAI) while enabling self-serve analytics and Data Mesh. I collaborate with data scientists to drive AI/ML insights for fraud detection, network optimization, and predictive analytics. I’ve delivered cost savings, reduced latency, and democratized data access across multi-cloud environments.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
See more

Work Experience

Data Engineer at Charter
March 1, 2024 - Present
Engineered petabyte-scale data pipelines for telecom CDRs, network logs, and IoT feeds using Apache Spark, Databricks, and Delta Lake, achieving 99.99% uptime and optimized latency. Designed real-time streaming architectures with Kafka, Flink, and Confluent Cloud for fraud detection, outage monitoring, and churn prediction. Built cloud-native ETL/ELT pipelines with AWS Glue, Azure Data Factory, and Google Cloud Dataflow, automating ingestion from OSS/BSS systems, CRM platforms, and 5G telemetry sources. Optimized data models/storage across Snowflake, BigQuery, Redshift, and Iceberg for analytics. Implemented data quality/governance with Great Expectations, Collibra, and Apache Atlas to meet GDPR/CCPA/TRAI. Applied AI-driven network analytics with MLflow/TFX/PyTorch Lightning to predict congestion and improve QoE. Used dbt for KPI transformations and orchestrated pipelines across multi-cloud with Airflow/Dagster/Prefect. Architected Data Mesh and built real-time dashboards with Power BI
Data Engineer at Silicon Valley Bank
March 1, 2023 - February 1, 2024
Implemented real-time data pipelines using Apache Kafka and Spring Boot microservices, reducing transaction latency by ~45% in fraud workflows. Designed and deployed AWS-based ETL with S3, Redshift, Lambda, and Glue, improving data accessibility and reducing batch runtimes by ~30%. Optimized SQL queries across Oracle and PostgreSQL, cutting report generation time by ~35%. Automated CI/CD with Jenkins/Maven/Git; Dockerized services with Kubernetes to boost scalability and reduce downtime by ~40%. Strengthened data validation with Hibernate/JPA for FINRA/SOX compliance. Built REST APIs for secure data exchange with partners; migrated legacy ETL to AWS Glue/EMR for serverless execution (~20%). Configured Athena for self-service analytics and performed root-cause analysis to improve SLA. Collaborated with data scientists to curate Redshift datasets for predictive risk modeling, improving forecasting accuracy by ~18%. Implemented monitoring with CloudWatch/CloudTrail and automated infra pro
Data Engineer at PulseArc
July 1, 2018 - December 1, 2020
Assisted in developing Java- and Python-based ETL workflows to process patient, billing, and insurance data across hospital systems. Contributed to designing Star/Snowflake data warehouse schemas, improving reporting efficiency. Supported REST APIs for secure EHR/EMR data exchange with third-party applications. Migrated datasets to Azure Synapse Analytics and Blob Storage, gaining hands-on cloud data warehousing experience. Worked with Apache Kafka to stream HL7/FHIR messages for near real-time updates. Aided in automating data ingestion with Azure Data Factory pipelines and employed Azure Key Vault for secrets management and RBAC to support HIPAA compliance. Enhanced error handling, retries, and logging in ETL jobs. Prepared structured datasets for Power BI/Tableau dashboards and performed SQL tuning to resolve performance issues. Contributed to Azure-based data lake design for self-service BI.

Education

Masters in Computer Technology at Eastern Illinois University
January 11, 2030 - January 8, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Telecommunications, Financial Services, Healthcare, Software & Internet, Professional Services