Data Engineering professional with 6+ years of experience building end-to-end big data solutions across ingest, storage, processing, querying, and analytics. Strong hands-on expertise with Hadoop ecosystem and modern cloud platforms (AWS, GCP, and Azure), including ETL pipelines, data warehousing, and performance optimization. Experienced in designing and developing data architectures and reusable ETL frameworks, integrating multiple data sources (Oracle, MySQL, PostgreSQL, SQL Server, APIs, files), and ensuring data reliability through testing, validation, and monitoring. Skilled in building streaming and batch data workflows using tools like Airflow, Spark, Hive, Kafka, and scheduling systems (Oozie/ZooKeeper), with additional analytics support using BI/reporting tools (Power BI, Tableau, SSRS).

Muskan Khadka

Data Engineering professional with 6+ years of experience building end-to-end big data solutions across ingest, storage, processing, querying, and analytics. Strong hands-on expertise with Hadoop ecosystem and modern cloud platforms (AWS, GCP, and Azure), including ETL pipelines, data warehousing, and performance optimization. Experienced in designing and developing data architectures and reusable ETL frameworks, integrating multiple data sources (Oracle, MySQL, PostgreSQL, SQL Server, APIs, files), and ensuring data reliability through testing, validation, and monitoring. Skilled in building streaming and batch data workflows using tools like Airflow, Spark, Hive, Kafka, and scheduling systems (Oozie/ZooKeeper), with additional analytics support using BI/reporting tools (Power BI, Tableau, SSRS).

Available to hire

Data Engineering professional with 6+ years of experience building end-to-end big data solutions across ingest, storage, processing, querying, and analytics. Strong hands-on expertise with Hadoop ecosystem and modern cloud platforms (AWS, GCP, and Azure), including ETL pipelines, data warehousing, and performance optimization.

Experienced in designing and developing data architectures and reusable ETL frameworks, integrating multiple data sources (Oracle, MySQL, PostgreSQL, SQL Server, APIs, files), and ensuring data reliability through testing, validation, and monitoring. Skilled in building streaming and batch data workflows using tools like Airflow, Spark, Hive, Kafka, and scheduling systems (Oozie/ZooKeeper), with additional analytics support using BI/reporting tools (Power BI, Tableau, SSRS).

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Language

Work Experience

Data Engineer at CVS Health
August 1, 2024 - Present
Hands-on development and support for new and existing data applications. Supported migration of projects from Teradata to Google Cloud Platform (GCP), including converting scripts (Teradata to BigQuery/SQL), mapping/session/workflow logic to load into BigQuery tables, and optimizing query performance by reducing overall execution time. Built Airflow DAG to execute scripts and load data into tables. Prepared unit test cases and documentation for data loads, maintained comparison sheets for Teradata vs GCP tables, and implemented pre-ingestion data checks before loading into the RAW layer. Developed ED M job workflows to load files from source into RAW and to move data RAW-to-persistent intake, error, stage, and target layers, including same-environment QA testing. Managed scheduling and monitoring, performed post-run validation by comparing with the old environment data, and provided production issue support.
Data Engineer at Selective Insurance, Charlotte, NC
August 1, 2023 - July 1, 2024
Developed and supported data applications focused on data warehousing, data integration, and data engineering projects. Advanced SQL skills for query and database performance optimization. Worked closely with leads to create/review low-level implementation designs. Built and enhanced ETL pipelines using SQL and Python and tools such as Informatica PowerCenter, DataBricks, and IICS. Wrote SQL queries and used joins to access data from Oracle and PostgreSQL. Supported design, development, testing, deployment, and documentation of data engineering deliverables. Performance tuned SQL and ETL pipelines across data lake and data warehouse environments. Worked with Azure DevOps for version control and CI/CD pipelines, developed Informatica PowerCenter mappings and loads into Azure, and built IICS mappings to push data to Azure. Monitored pipeline health and created contingency plans, and produced dashboards/reports using Power BI and Azure ML as needed.
Data Engineer at CVS Health
August 1, 2023 - Present
Hands-on development and support for new and existing data applications. Migrated a project from Teradata to Google Cloud Platform (GCP), including converting scripts (Teradata to BigQuery/SQL), Informatica mappings/workflow logic to BigQuery for loading data, and optimizing query execution time to reduce overall runtime. Built Airflow DAGs/operators to execute scripts and load data into tables. Prepared unit test cases and excel documentation to track table comparisons between Teradata and GCP. Added pre-ingestion data checks and created EDW job flows to load data into RAW, then RAW to persistent/performance layers (intake/error/stage/target). Performed data validation post-load by comparing against the old environment. Provided production issue support and maintained scheduling/run sheets. Used Git/Git Bash to push code to final production.
Data Engineer at CVS Health, Hartford, CT
July 1, 2022 - June 1, 2023
Maintained data infrastructure across multiple projects on Google Cloud Platform using Terraform (Infrastructure as Code). Performed data aggregation from varied sources including databases, APIs, files, and streams; linked related data points across datasets to improve data coherence. Implemented data standardization processes including cleaning, transforming, and harmonizing data. Conducted quality checks to identify and correct errors. Developed ETL pipelines to support data warehouse integration and built regulatory and financial reports using advanced SQL (Snowflake). Loaded and transformed large structured and semi-structured datasets; analyzed them by running Hive queries. Created dashboards/visualizations/reports in support of decision-making. Migrated database from SQLite to MySQL to PostgreSQL ensuring data integrity. Developed and deployed transformations using PySpark in the data processing cluster, and leveraged Spark SQL APIs for extraction and loading and SQL queries. Wo
Data Engineer at Selective Insurance
July 1, 2022 - July 1, 2024
Developed and supported data warehousing, data integration, and data engineering projects. Strong SQL knowledge with query and database performance optimization. Partnered with the lead to create/review low-level implementation designs. Built/used ETL tools and pipelines (Informatica PowerCenter, DataBricks, IIC S). Worked in Azure cloud services and wrote SQL queries using joins to access Oracle and PostgreSQL. Supported design, development, testing, deployment, and documentation of data engineering components. Monitored pipeline health and implemented contingency plans. Used Azure DevOps for version control and CI/CD. Developed PowerCenter mappings and pushed data to Azure cloud. Delivered reporting using Power BI and Azure ML. Worked closely with Product Owner, Product Managers, Scrum Master, and team members; created detailed documentation for downstream integrations and supported production issues.
Data Engineer at CVS Health
June 1, 2022 - June 1, 2023
Maintained data infrastructure across multiple projects on Google Cloud using Terraform (Infrastructure as Code). Conducted data aggregation and standardization: cleaning, transforming, and harmonizing data for compatibility. Performed data quality checks (identify/correct errors). Developed ETL pipelines into and out of the data warehouse and built major regulatory and financial reports in Snowflake using advanced SQL queries. Loaded and transformed large structured and semi-structured datasets; analyzed them by running Hive queries. Created dashboards/visualizations/reports for stakeholders. Migrated an SQLite database to MySQL and then PostgreSQL with full data integrity. Developed and deployed processing logic using PySpark in the data cluster; used Spark SQL APIs. Built Airflow workflows on GCP for scheduled updates, including BigQuery-authorized views. Designed Informatica (SSIS) packages to extract/transform/load data from heterogeneous sources to data warehouses/data marts. Mig
Senior Big Data Engineer at Expedia
April 1, 2022 - June 1, 2022
Performed ETL from multiple sources (e.g., Kafka, NIFI, Teradata, DB2) using Hadoop/Spark. Moved data from Teradata to a Hadoop cluster using TDCH/FastExport and Apache NiFi. Developed Python/PySpark/Bash scripts to transform and load data across on-prem and cloud platforms. Built Azure Data Factory pipelines and automated mediations using PowerShell/JSON templates. Worked with Informatica PowerCenter tool components (Designer, Repository Manager, Workflow Manager, Workflow Monitor) including Docker/Kubernetes setup for orchestration. Handled imports from multiple data sources and transformations using Hive/MapReduce; loaded data into HDFS and extracted from MySQL into HDFS via Sqoop. Ingested data into Azure services (Data Lake, Storage, Azure SQL, Azure DW) and created tables/views in Snowflake as per business requirements. Worked with NIFI workflows to pick up data from REST APIs/server and send to Kafka broker; used Spark for intraday/real-time processing. Developed web application
Senior Big Data Engineer
June 1, 2021 - December 1, 2021
Worked on data movement and ETL tasks between HDFS and AWS S3, with extensive S3 bucket work. Created data ingestion modules using Glue for loading data into various layers in S3 and reporting using Athena and QuickSight. Designed data models for use in analytics and serverless applications on AWS Lambda for complex analysis, end-to-end traceability, lineage, and key business element definitions. Conducted User Acceptance Testing (UAT) and documented test cases for business lines. Hands-on with AWS databases such as RDS (Aurora), Redshift, DynamoDB, and ElasticCache (Memcached/Redis). Converted Hadoop jobs to run in EMR by configuring the cluster based on data size. Used AWS Glue Data Catalog with crawlers to bring data from S3 and query it via AWS Athena; generated reports using QuickSight. Created Kinesis data streams/firehose/analytics to capture and process streaming data and output to S3/DynamoDB/Redshift. Performed data cleaning using Ab Initio components (joins, dedup, denormali
Senior Big Data Engineer at Homestead Insurance / (Hartford, MA area)
June 1, 2021 - March 1, 2022
Worked on data ingestion and pipelines between HDFS and AWS S3, including extensive S3 bucket processing. Created data ingestion modules using AWS Glue to load data across layers in S3; reported using Athena and QuickSight. Designed data models for AWS Lambda-based analytical applications requiring end-to-end traceability, lineage, and key business element definitions. Conducted User Acceptance Testing (UAT) and documented test cases. Built pipelines using AWS services (RDS/Aurora, Redshift, DynamoDB, Elasticache/Memcached/Redis). Converted Hadoop jobs to run on EMR by configuring cluster based on data size. Built Glue crawlers for S3 data and executed SQL queries in Athena. Implemented Kinesis streams/firehose/analytics to capture streaming data and output to S3/DynamoDB/Redshift. Built ETL and cleaning using Ab Initio components (join/dedup/denormalize/normalize/refromatting/filter-by-expression/rollup). Wrote SQL scripts for data mismatch checks and loaded history from Teradata SQL
Senior Big Data Engineer / Data Engineer
January 1, 2019 - March 1, 2021
Performed ETL from multiple sources such as Kafka, NIFi, Teradata, and DB2 using Hadoop Spark. Moved data from Teradata to a Hadoop cluster using TDC H/Fast export and Apache NiFi. Developed Python/PySpark/Bash scripts to transform and load data across on-prem and cloud. Built Azure Data Factory pipelines, integration runtimes, and data ingestion components; developed remediation/automation systems using PowerShell scripts and JSON templates. Worked with Informatica PowerCenter tools (Designer, Repository Manager, Workflow Manager, Workflow Monitor) and implemented Docker/Kubernetes for orchestration purposes. Handled importing data from multiple sources, performed transformations using Hive/MapReduce, loaded data into HDFS, and extracted data from MySQL into HDFS via Sqoop. Built data ingestion workflows to Azure services such as Azure Data Lake, Azure Storage, Azure SQL, and Azure DW; created tables/views on Snowflake as per business needs; implemented NiFi flows to pick up data from
Hadoop Engineer / Data Engineer
June 1, 2018 - December 1, 2019
Involved in functional, integration, regression, smoke, and performance testing of Hadoop MapReduce code developed in Python, Pig, and Hive. Built SQL scripts for data matching and loaded history data from Teradata SQL to Snowflake. Created Azure Data Factory and managed policies; used Azure Blob Storage for storage/backup. Built Azure Notebook functions using Python/Scala/Spark. Implemented Metadata tables and end-user views in Snowflake for Tableau refresh. Created performance dashboards in Tableau/Excel/PowerPoint. Implemented defect tracking in Jira and supported data processing improvements. Developed Spark code and Spark SQL/streaming for faster testing and data processing. Evaluated traffic/performance of daily deals ads and suggested improvements to BI components (reports/stored procedures).

Education

Bachelor's in Science at Doon University
January 1, 2018 - August 26, 2026
Master's in Analytics and Modelling at Valparaiso University
January 1, 2021 - August 26, 2026
Bachelor's in Science at Doon University
January 11, 2030 - January 1, 2018
Master's in Analytics and Modelling at Vaparaisao University
January 1, 2021 - August 26, 2026

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Computers & Electronics, Financial Services, Healthcare, Professional Services

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Hire a Data Engineer

We have the best data engineer experts on Twine. Hire a data engineer today.