Data Engineer/AIML Engineer with 9+ years of experience across DataOps, data engineering, ETL/ELT, and data pipeline automation, building scalable batch and incremental pipelines for enterprise analytics and ML-driven use cases. Strong hands-on expertise in Python, Spark, Hadoop ecosystem tools, Airflow orchestration, and cloud platforms with a focus on reliability, performance, and data quality. Experienced in designing dimensional models (Kimball/star schemas), developing and optimizing Snowflake/Spark/SQL transformations, and managing production support with root-cause analysis, restartable workflows, and governance controls. Collaborates with business and technical teams to translate requirements into robust data architectures, improving processing efficiency and enabling secure, curated data for fraud/risk/compliance and analytics stakeholders.

Pranav M

Data Engineer/AIML Engineer with 9+ years of experience across DataOps, data engineering, ETL/ELT, and data pipeline automation, building scalable batch and incremental pipelines for enterprise analytics and ML-driven use cases. Strong hands-on expertise in Python, Spark, Hadoop ecosystem tools, Airflow orchestration, and cloud platforms with a focus on reliability, performance, and data quality. Experienced in designing dimensional models (Kimball/star schemas), developing and optimizing Snowflake/Spark/SQL transformations, and managing production support with root-cause analysis, restartable workflows, and governance controls. Collaborates with business and technical teams to translate requirements into robust data architectures, improving processing efficiency and enabling secure, curated data for fraud/risk/compliance and analytics stakeholders.

Available to hire

Data Engineer/AIML Engineer with 9+ years of experience across DataOps, data engineering, ETL/ELT, and data pipeline automation, building scalable batch and incremental pipelines for enterprise analytics and ML-driven use cases. Strong hands-on expertise in Python, Spark, Hadoop ecosystem tools, Airflow orchestration, and cloud platforms with a focus on reliability, performance, and data quality.

Experienced in designing dimensional models (Kimball/star schemas), developing and optimizing Snowflake/Spark/SQL transformations, and managing production support with root-cause analysis, restartable workflows, and governance controls. Collaborates with business and technical teams to translate requirements into robust data architectures, improving processing efficiency and enabling secure, curated data for fraud/risk/compliance and analytics stakeholders.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
See more

Language

English
Advanced
Telugu
Advanced
Hindi
Intermediate

Work Experience

Sr Data Engineer/AIML Engineer at Royal Bank of Canada, One liberty plaza, NY
February 1, 2025 - Present
Designed, developed, and supported scalable ETL/ELT pipelines for banking transaction, fraud, risk, compliance, and analytics workloads. Built enterprise pipelines using Python, PySpark, Airflow, AWS, Snowflake, Oracle, and SQL to ingest, transform, validate, reconcile, and deliver curated datasets. Developed Snowflake SQL/dbt transformations to standardize data across banking systems and modernized source-to-target integration by resolving schema/identifier/business-rule differences while maintaining source traceability. Implemented cloud-native batch and incremental workflows using S3, AWS Glue, EMR, Spark, and Airflow for large-scale processing, including idempotent and restartable designs to prevent duplicates during reruns. Optimized Spark/Airflow/SQL/Snowflake performance for 5TB+ daily processing through partitioning, join strategy improvements, and workflow dependency tuning. Added comprehensive data quality controls (null/duplicate checks, schema validation, reconciliation to
Sr Data Engineer at Royal Bank of Canada
February 1, 2025 - Present
Designed, developed, and supported scalable ETL/ELT pipelines for banking transaction, fraud, risk, and compliance workloads. Built curated financial datasets using Python, PySpark, Airflow, AWS, Snowflake, Oracle, and SQL; implemented complex Snowflake SQL and dbt transformations to standardize data across banking systems. Developed cloud-native ingestion workflows using S3, AWS Glue, EMR, Spark, Airflow, and Snowflake for large-scale batch processing. Created reusable Airflow DAG orchestration patterns with scheduling, retries, failure handling, backfills, monitoring, and restartable processing. Implemented idempotent batch and incremental loads to prevent duplicates on reruns. Built Kimball dimensional models (STAR schemas, fact/dimension tables, SCDs) and ensured source traceability in source-to-target mappings. Added comprehensive data-quality controls (null/duplicate validation, schema checks, referential integrity, record-count reconciliation, freshness). Optimized Spark/Airflow
Data Engineer at John Deere, Moline, Illinois
August 1, 2023 - January 1, 2025
Designed and deployed AWS-based multi-tier solutions emphasizing high availability, fault tolerance, and autoscaling using CloudFormation and services such as EC2, Route53, S3, RDS, DynamoDB, SNS, SQS, and IAM. Built data pipelines and ingestion workflows via API Gateway, Lambda, and nested JSON processing into Redshift, along with scalable Spark/Scala processing for RDBMS and streaming sources. Optimized Spark jobs and migrated MapReduce to Spark SQL; implemented Spark Streaming workflows to process newly arrived S3 files and persist results to HDFS. Used AWS Glue (catalog and crawlers) for metadata management and SQL querying on S3-backed datasets. Migrated legacy schedulers to Apache Airflow by analyzing schedules/dependencies and creating deployment docs, rerun procedures, and operational runbooks. Developed reusable Airflow automation patterns and test coverage (unit, integration, regression, end-to-end) using SQL validation and reconciliation logic. Built Tableau reports/dashboa
Data Engineer at John Deere
August 1, 2023 - January 1, 2025
Designed and deployed multi-tier applications using AWS services (EC2, Route53, S3, RDS, DynamoDB, SNS, SQS, IAM) with high availability, fault tolerance, and autoscaling via CloudFormation. Built AWS data pipelines using API Gateway, Lambda, and Python to transform nested JSON and load results into Redshift. Developed Spark applications in Scala and optimized jobs by converting MapReduce logic to Spark SQL. Implemented Spark Streaming workflows to process new S3 files, transform/aggregate data, and persist results to HDFS. Used AWS Glue (catalog + crawlers) for schema discovery and SQL operations. Developed Python/Airflow workflow management (including time sensors) and migrated legacy schedulers to Airflow with documentation, rerun procedures, and operational runbooks. Executed module/integration/regression/end-to-end tests with SQL-based validation and reconciliation; supported UAT and production releases with defect management and RCA. Built Tableau dashboards for business reportin
Big Data Engineer at Big Lots, Columbus OH
February 1, 2020 - October 1, 2021
Extracted and processed large-scale customer behavior, sales/revenue, supply chain, and logistics datasets. Transferred data to AWS S3 using Apache NiFi dataflows; validated and cleaned data using Python before storage. Built transformations in PySpark and implemented data-quality checks, metadata management, data lineage, schema validation, and reconciliation across pipelines. Designed and maintained distributed database solutions using CockroachDB, including schema design and SQL performance optimization, and loaded transformed data into AWS Redshift for analytics. Scheduled Hadoop jobs with Oozie and maintained reusable Airflow DAG templates/operators to improve consistency and code quality. Led a team of three engineers to deliver a complex ingestion/processing pipeline for a new data source, reducing time to insights by 50%. Performed analytical querying in HDFS using Hive; converted Hive queries into PySpark transformations (RDD/DataFrame API). Monitored pipelines with Grafana a
ETL Developer at Big lots (ETL Developer) / (as listed in resume)
February 1, 2020 - October 1, 2021
Developed and maintained Informatica PowerCenter (Designer/Workflow Manager/Monitor/Repository) ETL workflows and handled ETL standards using QMC. Built streaming and near-real-time pipelines using Kafka and Spark Streaming. Used Sqoop for data transfer between relational databases and Hadoop. Developed real-time ingestion with Kafka and processing with Spark Streaming. Designed STAR schemas and analytical data models for reporting. Extracted and loaded data from web servers and Teradata using Sqoop/Flume and streaming components. Implemented MapReduce/Python modules for predictive/ML analytics on Hadoop, including distributed random forest via Python streaming. Created multiple MapReduce programs for extraction/transformation/aggregation across formats (XML/JSON/CSV and compressed files). Built complex Informatica mappings (joiner, expression, aggregate, lookup, filter, router, update strategy) and supported data loads using Unix shell scripts and SQL loader. Worked on SFTP setup and
Big Data Engineer at Big Lots
February 1, 2020 - October 1, 2021
Extracted customer behavior, sales/revenue, and supply-chain/logistics data from HDFS and transferred data to AWS S3 using Apache NiFi. Validated and cleaned data using Python before storing in S3. Processed and transformed data with PySpark; implemented data quality checks, metadata management, schema validation, and reconciliation across pipelines. Designed and supported distributed databases using CockroachDB with schema design and SQL performance tuning. Loaded curated datasets into AWS Redshift for analysis. Scheduled Hadoop workflows using Oozie. Built reusable Airflow DAG templates/operators and led a team of three engineers to deliver an ingestion/processing pipeline that reduced time to insights by 50%. Queried and transformed data using Hive/HiveQL and converted Hive logic into PySpark transformations. Monitored pipelines using Grafana and supported distributed application coordination via Zookeeper.
ETL Developer at Company not specified (Alibaba Cloud deployment mentioned in resume)
February 1, 2019 - October 1, 2021
Developed and maintained ETL pipelines using Informatica PowerCenter (Designer, Workflow Manager, Workflow Monitor, Repository Manager), including complex mappings and enterprise data transformations. Used Kafka for live streaming ingestion and analytics, and implemented real-time pipelines using Kafka and Spark Streaming. Transferred data via Sqoop and Flume between relational systems and Hadoop, and built MapReduce/PySpark modules for machine learning/predictive analytics. Designed star schemas (fact/dimension tables) and analytical data models for reporting. Built MapReduce programs for extraction/transformation/aggregation across XML/JSON/CSV and compressed file formats. Integrated heterogeneous sources including Oracle and flat files, and performed joins, expressions, lookups, routers, filters, and update strategies to load targets. Managed database deployments on Alibaba Cloud with focus on security, scalability, and availability. Also supported Informatica workflow migrations,
Data Analyst - Python at Voice Gate Technologies India Pvt.
September 1, 2017 - November 1, 2019
Performed exploratory data analysis and quantitative analysis using Python libraries such as NumPy, pandas, SciPy, and Matplotlib. Built SQL queries for data validation and accuracy checks aligned to business requirements. Produced high-level analysis reports and dashboards using Excel and Tableau, including identifying billing patterns, outliers, and data quality limitations. Worked with Tableau filters/sorting and created Excel summary reports (pivot tables and charts). Developed backend Python business logic and used pandas/Lambda (map/filter/reduce) patterns, plus time-series analysis using pandas API. Created regression test frameworks for new code and supported data modeling/requirements gathering. Used multiple data source types (CSV, Excel, HTML, SQL) and handled data writing back to files and databases.
Data Analyst (Python) at Voice Gate Technologies India Pvt. Ltd.
September 1, 2017 - November 1, 2019
Performed exploratory data analysis and quantitative analysis using Python libraries (NumPy, pandas, SciPy, Matplotlib) and SQL. Developed complex SQL queries and scripts for data validation, aggregation, and requirement-driven extraction. Built reports and dashboards using Excel and Tableau; supported analysis by identifying billing patterns, outliers, and data quality limitations. Implemented data validation using standard SQL and created Excel summary outputs (pivot tables/charts). Worked with CSV/Excel/HTML and wrote analytical outputs back to files/databases. Used pandas DataFrame operations and functional concepts (map/reduce/filter patterns); implemented Lambda-based transformations. Also created regression test frameworks for new code and handled business logic via backend Python. Supported time-series analysis with pandas APIs.

Education

degree at Aurora Degree & PG college
June 1, 2014 - September 1, 2017
Degree & PG college at Aurora Degree & PG college
June 1, 2014 - September 1, 2017

Qualifications

Add your qualifications or awards here.

Industry Experience

Financial Services, Software & Internet, Professional Services

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
See more