I’m a data engineer with 6+ years of experience building reliable data pipelines, cloud data platforms, and distributed processing solutions across AWS and data platforms like Databricks. I focus on end-to-end ETL/ELT and data quality—ingesting, transforming, validating, and curating enterprise datasets for analytics and operational reporting. I work hands-on with Python, SQL, PySpark/Spark, dbt, Airflow, Snowflake, Kafka, and AWS services to deliver batch and near-real-time streaming systems. I also enjoy partnering with engineers, analysts, and stakeholders to translate data requirements into production-ready pipelines, monitoring workflows for health and latency, and improving performance and retrievability of downstream data products.

shradhanjali pradhan

I’m a data engineer with 6+ years of experience building reliable data pipelines, cloud data platforms, and distributed processing solutions across AWS and data platforms like Databricks. I focus on end-to-end ETL/ELT and data quality—ingesting, transforming, validating, and curating enterprise datasets for analytics and operational reporting. I work hands-on with Python, SQL, PySpark/Spark, dbt, Airflow, Snowflake, Kafka, and AWS services to deliver batch and near-real-time streaming systems. I also enjoy partnering with engineers, analysts, and stakeholders to translate data requirements into production-ready pipelines, monitoring workflows for health and latency, and improving performance and retrievability of downstream data products.

Available to hire

I’m a data engineer with 6+ years of experience building reliable data pipelines, cloud data platforms, and distributed processing solutions across AWS and data platforms like Databricks. I focus on end-to-end ETL/ELT and data quality—ingesting, transforming, validating, and curating enterprise datasets for analytics and operational reporting.

I work hands-on with Python, SQL, PySpark/Spark, dbt, Airflow, Snowflake, Kafka, and AWS services to deliver batch and near-real-time streaming systems. I also enjoy partnering with engineers, analysts, and stakeholders to translate data requirements into production-ready pipelines, monitoring workflows for health and latency, and improving performance and retrievability of downstream data products.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Work Experience

Data Engineer at Softtech Technology Group Inc
February 1, 2025 - Present
Architected end-to-end Python and SQL data pipelines to ingest, transform, validate, and curate enterprise datasets, establishing reliable foundations for analytics and downstream applications. Developed scalable PySpark and Databricks pipelines processing 2M+ enterprise documents, standardizing structured/unstructured data into curated datasets for reporting and consumption. Built incremental ETL/ELT workflows with reusable transformation logic, data validation, schema standardization, and error handling to improve consistency across processing jobs. Designed distributed data processing workflows with Databricks and PySpark, optimizing transformations, partitioning, and resource utilization for high-volume datasets. Implemented production data-quality and monitoring workflows using AWS CloudWatch and Prometheus, tracking pipeline health, processing latency, failures, and data issues while reducing detection time from 4 hours to 20 minutes. Built AI-ready pipelines using Python, LangCh
Data Engineer at DE Holdings
January 1, 2025 - February 1, 2025
Built scalable ETL pipelines using Python, PySpark, and Databricks to transform raw data into accurate Delta Lake tables for analytics and reporting. Developed dbt models and automated data tests/documentation, standardizing business metrics and reducing reporting inconsistencies by 25%. Built near-real-time pipelines using AWS Kinesis, Lambda, and S3 to process time-sensitive data and reduce operational decision latency by 30%. Improved PostgreSQL query performance through indexing, partitioning, and query optimization. Integrated external systems via REST APIs to automate ingestion and reduce manual data transfers. Used AWS Athena and RDS for querying operational datasets supporting ad-hoc analysis and downstream reporting. Implemented data validation and monitoring workflows with Apache Airflow and Python, achieving 99.5% pipeline reliability and reducing production data incidents by 35%. Partnered with BI/analytics/operations teams to define data requirements and resolve data issue
Data Engineer at D E Holdings
January 1, 2025 - February 28, 2025
Built scalable ETL pipelines with Python, PySpark, and Databricks to transform raw data into accurate Delta Lake tables for analytics and reporting. Developed dbt models and automated data tests and documentation, standardizing business metrics and reducing reporting inconsistencies by 25%. Implemented near real-time pipelines using AWS Kinesis, Lambda, and S3 to process time-sensitive data and reduce operational decision latency by 30%. Improved PostgreSQL query performance through indexing, partitioning, and query optimization, reducing processing time for frequently accessed datasets. Integrated external systems via REST APIs, automating data ingestion and reducing manual transfers between platforms. Used AWS Athena and RDS to query and manage operational datasets for ad-hoc analysis and downstream reporting. Built data validation and monitoring workflows with Apache Airflow and Python, achieving 99.5% pipeline reliability and reducing production data incidents by 35%.
Data Engineer at Data Glacier
June 1, 2024 - August 31, 2024
Developed Python and SQL ETL pipelines to ingest and transform more than 1M healthcare records into Delta Lake tables for downstream analytics and reporting. Orchestrated data workflows using Databricks and Apache Airflow, with AWS CloudWatch monitoring for pipeline failures and SLA tracking. Designed and optimized PostgreSQL tables and queries using indexing and query tuning to improve data retrieval for operational and clinical reporting. Developed REST APIs for standardized data ingestion from external healthcare systems, reducing manual reporting effort by 15%. Collaborated with stakeholders to define data requirements, document workflows, and resolve data issues across ingestion and transformation processes.
Data Engineer at Steradian Technologies
June 1, 2021 - July 31, 2023
Built batch ETL pipelines using PySpark, AWS Glue, and dbt to process high-volume healthcare IoT data for analytics and regulatory reporting. Developed streaming pipelines with Apache Kafka and AWS Lambda to support early-warning systems and reduce clinical response times by 40%. Designed dimensional data models in Snowflake using clustering and partitioning to improve query performance for analytics and executive reporting. Implemented data validation and quality checks using Python and AWS Glue, reducing downstream data-quality issues and improving overall pipeline reliability. Built reusable data pipelines for machine learning workflows, automating feature preparation and data ingestion for clinical anomaly detection models. Worked with data scientists and analytics teams to troubleshoot pipeline issues, improve data availability, and productionize data workflows.
Data Analyst at Remote care
August 1, 2020 - June 30, 2022
Analyzed healthcare operations data using SQL and Python to identify appointment no-show patterns, contributing to a 12% improvement in attendance across priority clinics. Built PySpark and PostgreSQL data pipelines to prepare features for A/B testing and machine learning workflows, reducing feature refresh cycles from days to hours. Developed predictive models using scikit-learn to support patient retention and service optimization initiatives. Built Tableau dashboards for operational KPIs and predictive insights used by program managers for weekly decision-making. Automated model versioning and retraining workflows using Git-based controls, improving reproducibility and reducing post-release issues.

Education

Applied Data Analytics at Boston University
January 1, 2024 - January 1, 2025
Applied Data Analytics at Boston University
January 11, 2030 - January 1, 2025

Qualifications

Add your qualifications or awards here.

Industry Experience

Healthcare, Software & Internet