Hi, I’m Alex Zhao, a data science and machine learning engineer based in Toronto with experience spanning medical imaging, finance, and energy. I enjoy turning complex data into actionable insights, designing scalable pipelines, and deploying models that actually improve decision-making. Beyond coding, I value clear communication and collaborating with cross-functional teams to translate business needs into robust data solutions. I’m excited to continue building end-to-end systems that scale and solve real-world problems.

Alex Zhao

Hi, I’m Alex Zhao, a data science and machine learning engineer based in Toronto with experience spanning medical imaging, finance, and energy. I enjoy turning complex data into actionable insights, designing scalable pipelines, and deploying models that actually improve decision-making. Beyond coding, I value clear communication and collaborating with cross-functional teams to translate business needs into robust data solutions. I’m excited to continue building end-to-end systems that scale and solve real-world problems.

Available to hire

Hi, I’m Alex Zhao, a data science and machine learning engineer based in Toronto with experience spanning medical imaging, finance, and energy. I enjoy turning complex data into actionable insights, designing scalable pipelines, and deploying models that actually improve decision-making.

Beyond coding, I value clear communication and collaborating with cross-functional teams to translate business needs into robust data solutions. I’m excited to continue building end-to-end systems that scale and solve real-world problems.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Intermediate

Work Experience

Machine Learning Engineer Intern at SonoScape
May 1, 2025 - October 1, 2025
Designed scalable Python ETL pipelines for 3D medical imaging segmentation, including resampling, normalization, cropping, and augmentation to ensure consistency across multi-source volumetric datasets. Built an end-to-end PyTorch data pipeline processing ~2TB of imaging data, automating preprocessing to cut data preparation and experiment turnaround by 90%. Implemented automated QA and filtering to remove corrupted scans, and published quality metrics for cross-team validation. Benchmarked MedSAM, SAM3D-ViT, and YOLO-based baselines; selected the best memory–performance trade-off and performed domain-specific fine-tuning to achieve Dice 0.91 from 0.80. Used high-confidence pseudo-labeling to reduce labeling; profiling showed inference around 0.75s per 3D volume, with a lightweight variant achieving ~0.3s and enabling 8GB-GPU feasibility.
Data Scientist Intern at Wuhan Financial Holdings (Group) Co., Ltd.
May 1, 2024 - August 1, 2024
Developed an XGBoost loan default model on 50K+ borrower records, achieving 0.89 AUC and strong ranking performance. Created 50+ predictive features from credit bureau data, transaction history, and borrower demographics using Pandas to improve risk assessment. Visualized portfolio risk distribution and model performance, reducing manual risk review by 20% and cutting decision turnaround time by 15% for high-risk account actions.
Data Scientist Intern at Hua Zhong Electric Power Technology Development Co.
June 1, 2023 - August 1, 2023
Built and deployed a Graph Neural Network-based electricity demand forecasting model using PyTorch and SQL on 10M+ grid sensor records, delivering automated regional demand predictions and improving planning accuracy by 8% R². Designed data integration pipelines to combine weather signals and grid topology with demand history, enabling geographically aware forecasting. Delivered a lightweight Streamlit dashboard to visualize forecasts, trends, and model performance, reducing manual planning effort by 30% and cutting decision turnaround by 25%. Built ETL and validation workflows with data quality checks, and benchmarked against baseline forecasting models.
Database DevOps Intern at Changjiang Securities
September 1, 2020 - August 1, 2021
Built containerized Oracle-to-HDFS and Hive migration workflows on an Alibaba distributed Hadoop stack, onboarding new nodes and migrating multi-TB datasets into partitioned Hive tables, cutting end-to-end runtime by 20%. Developed PySpark ETL pipelines over 200,000+ client accounts, producing curated Hive datasets for segmentation, product recommendation analysis, and retention analysis, and improving downstream query responsiveness via partitioning and file layout tuning. Strengthened reliability with structured logging, checkpointed reruns, and automated recovery scripts, reducing manual intervention time by 30% during peak migrations. Implemented operational monitoring with Prometheus and Grafana, tracking job runtime, HDFS I/O throughput, Hive query latency, executor memory, and failure rates to speed up incident triage and keep daily batches stable.

Education

Master of Engineering (Electrical & Computer Engineering) at University of Toronto
September 1, 2023 - May 1, 2025
Bachelor of Computing (Honours) in Computing & Mathematics Analysis at Queen's University
September 1, 2019 - May 1, 2023

Qualifications

Add your qualifications or awards here.

Industry Experience

Healthcare, Financial Services, Energy & Utilities, Software & Internet