I'm Srinivas Sai, a data engineer and Python developer based in Toronto with 6+ years delivering data-intensive applications across Hadoop ecosystems, cloud platforms, and data warehouses. I enjoy turning complex datasets into reliable pipelines and actionable insights, collaborating with cross‑functional teams to enable data-driven decision making. I thrive on building scalable architectures for streaming and batch workloads on AWS, Azure, and GCP, with a focus on data quality, automation, and CI/CD. In my spare time, I contribute to open source and mentor teammates on best practices in data engineering.

Srinivas Sai

I'm Srinivas Sai, a data engineer and Python developer based in Toronto with 6+ years delivering data-intensive applications across Hadoop ecosystems, cloud platforms, and data warehouses. I enjoy turning complex datasets into reliable pipelines and actionable insights, collaborating with cross‑functional teams to enable data-driven decision making. I thrive on building scalable architectures for streaming and batch workloads on AWS, Azure, and GCP, with a focus on data quality, automation, and CI/CD. In my spare time, I contribute to open source and mentor teammates on best practices in data engineering.

Available to hire

I’m Srinivas Sai, a data engineer and Python developer based in Toronto with 6+ years delivering data-intensive applications across Hadoop ecosystems, cloud platforms, and data warehouses. I enjoy turning complex datasets into reliable pipelines and actionable insights, collaborating with cross‑functional teams to enable data-driven decision making.

I thrive on building scalable architectures for streaming and batch workloads on AWS, Azure, and GCP, with a focus on data quality, automation, and CI/CD. In my spare time, I contribute to open source and mentor teammates on best practices in data engineering.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
See more

Work Experience

Sr. Data Engineer at Manulife
March 1, 2024 - November 25, 2025
Identified, designed, and implemented internal process improvements to automate manual tasks, optimize data delivery, and redesign infrastructure for scalability. Assisted BI users by connecting Redshift with Power BI, Excel, Spotfire, and Python. Led data enrichment and ETL efforts: PySpark pipelines cleansing and enriching raw data, including merging Snowflake data with MongoDB details. Ingested streaming data from REST APIs into AWS EMR via Kinesis; created Snowflake schemas to support complex analytics and built a pipeline to consolidate similar products using Word2Vec, Spark, Snowflake, and Airflow. Implemented real‑time processing with Kafka and Spark Streaming, loading into DataFrames and saving as Parquet in HDFS. Migrated data from Teradata to AWS and moved reports from OBIEE to Power BI. Automated Hive queries and scripts with Airflow and shell scripting. Built ETL workflows using Azure Data Factory, Azure Synapse Analytics, and Logic Apps. Leveraged SQL for data modeling a
Data Engineer at Point Click Care
February 1, 2024 - February 1, 2024
Responsible for reliability improvements to increase efficiency, eliminate downtime, and maintain performance at scale across platforms. Conducted data mining, cleaning, modeling, validation, and visualization. Built ETL pipelines from data lake to various databases; authored Python and Scala code in Azure Databricks. Automated scripts and workflows with Airflow and shell scripting for daily production tasks. Created and managed data pipelines using Azure Data Factory and Azure Databricks, enabling real‑time processing with Spark Streaming and Kafka. Created user accounts and data architecture artifacts and designed Snowflake schemas for analytical queries. Migrated data from AWS Redshift to partitioned S3 datasets. Developed multiple Kafka producers/consumers; used pub-sub to trigger orchestration. Installed and configured Apache Airflow for S3 and Snowflake data warehouse; built DAGs. Implemented MapReduce programs in Hortonworks for data cleaning and preprocessing. Performed text
Data Analyst at Wipro
August 1, 2021 - August 1, 2021
Designed and developed a web app BI for performance analytics. Converted Avro/Parquet to optimize data processing. Built real-time streaming with Spark, Kafka, Scala, and Hive for streaming ETL and ML. Wrote shell scripts to orchestrate Hadoop tasks; created Spark apps with PySpark and Spark SQL for multi-format data; prepared Python notebooks for automated weekly/monthly/quarterly reporting ETL. Implemented AWS fully managed Kafka streaming to connect APIs to Spark clusters, Redshift, Glue, and Lambda/Python. Used TensorFlow and Scikit-learn for ML; migrated Hive UDFs to Spark SQL. Orchestrated Airflow in a hybrid cloud. Wrote shell FTP scripts for moving data to S3. Analyzed datasets to determine aggregation and reporting. Used Oozie to manage Hadoop jobs; wrote HiveQL to Spark transformations. Installed/configured Hive, Pig, Sqoop, and Oozie; built an API to write XML from a database; used XML/XSL for dynamic web content. Created Airflow DAGs for onboarding datasets and change manag

Education

Bachelors in Mechanical Engineering at Cmr technical campus
January 11, 2030 - March 1, 2019

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Financial Services, Professional Services