I am Monica Chen, a results-driven Data Engineer with 2+ years of experience building scalable data pipelines and analytical systems. I specialise in distributed data processing, scalable ETL workflows, and data-intensive application development, with a strong applied AI background and hands-on experience with LLMs. I thrive in Agile environments and enjoy turning data insights into actionable business decisions. Across projects in finance, education, and research, I combine data engineering with software development to deliver end-to-end solutions. I am comfortable collaborating across cross-functional teams to ship production-grade systems on cloud platforms such as AWS and GCP, using PySpark, Hadoop, SQL, Python, and modern visualization tools.

Monica (Yiyang) Chen

I am Monica Chen, a results-driven Data Engineer with 2+ years of experience building scalable data pipelines and analytical systems. I specialise in distributed data processing, scalable ETL workflows, and data-intensive application development, with a strong applied AI background and hands-on experience with LLMs. I thrive in Agile environments and enjoy turning data insights into actionable business decisions. Across projects in finance, education, and research, I combine data engineering with software development to deliver end-to-end solutions. I am comfortable collaborating across cross-functional teams to ship production-grade systems on cloud platforms such as AWS and GCP, using PySpark, Hadoop, SQL, Python, and modern visualization tools.

Available to hire

I am Monica Chen, a results-driven Data Engineer with 2+ years of experience building scalable data pipelines and analytical systems. I specialise in distributed data processing, scalable ETL workflows, and data-intensive application development, with a strong applied AI background and hands-on experience with LLMs. I thrive in Agile environments and enjoy turning data insights into actionable business decisions.

Across projects in finance, education, and research, I combine data engineering with software development to deliver end-to-end solutions. I am comfortable collaborating across cross-functional teams to ship production-grade systems on cloud platforms such as AWS and GCP, using PySpark, Hadoop, SQL, Python, and modern visualization tools.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
See more

Language

English
Fluent
Chinese
Fluent

Work Experience

Full Stack Developer at ArtCert
November 1, 2025 - November 1, 2025
Development of an artwork authentication and provenance platform, ArtCert, establishing data engineering foundations, secure data workflows, and scalable backend architecture. Implemented robust Flask REST services, JWT security, and Redis-backed caching to enhance data access performance. Designed and optimised relational MySQL schema supporting high-integrity transactional workloads and efficient analytics. Containerised backend services with Docker and automated CI/CD pipelines for reliable deployment on Google Cloud Platform. Collaborated in an Agile environment to validate requirements, optimise data structures, and ensure scalable system performance.
Project Manager, Backend & Frontend Engineer at Flavours of Melbourne Platform
November 1, 2025 - November 1, 2025
Collected and cleaned 2000+ venue data from Google API; built automated functions for column mapping, pricing, and business hours standardization. Developed a full-stack interactive web application using RShiny, Leaflet, and Tableau Embedding API. Built a real-time recommendation engine based on budget, meal type, distance, and geolocation; routing optimization via OSRM improved itinerary accuracy by 35% over static mapping.
Product Owner at ArtCert - Artwork Authentication Platform
November 1, 2025 - November 1, 2025
Designed and developed a secure, transparent artwork authentication platform for artists and collectors. Implemented structured metadata storage and verification workflows to ensure trust and provenance. Built a React–Flask–MariaDB full-stack architecture, supporting user authentication, submission portals, and admin dashboards. Optimized deployment scripts and local development environments to maintain system stability during production rollout.
Data Analyst at GF Securities Co., Ltd.
February 1, 2025 - February 1, 2025
Supported data-driven decision-making by building reliable data pipelines, preparing ML-ready datasets, and delivering analytical insights through SQL, Python, and automated ETL workflows. Extracted, transformed and validated business data to support analytics and ML-ready dataset preparation. Built automated ETL workflows to improve data refresh efficiency and enable reliable, repeatable reporting. Performed EDA to identify key customer patterns and operational insights. Developed clear data visualisations using Power BI to communicate trends and support stakeholder decision-making. Collaborated with modeling and risk teams to assemble, clean, and quality-check datasets for predictive modeling use cases.
Full Stack Developer at WEHI
November 1, 2024 - November 1, 2024
Optimised the Student Organiser Platform by building data pipelines and automated testing workflows to improve system reliability and support data-driven development. Built synthetic test datasets using SQL to expand test scenarios and significantly improve test coverage. Developed and optimised Flask backend services to streamline data ingestion, processing, and API delivery. Containerised components with Docker and implemented CI/CD pipelines to automate testing and deployment for reliable releases. Deployed backend services to AWS for scalable cloud deployment. Collaborated with cross-functional teams to diagnose bottlenecks and deliver scalable, production-ready improvements.
Full-Stack Engineer Intern at Walter and Eliza Hall Institute of Medical Research
November 1, 2024 - November 1, 2024
Led the development of the institute's internal internship management web application. Used MySQL to optimize database structures and data retrieval pipelines. Designed and implemented synthetic public datasets to improve testing coverage and performance.
Machine Learning Engineer at Climate Change Fact-Checking System
May 1, 2024 - May 1, 2024
Developed end-to-end ML pipelines for climate misinformation detection using deep learning, NLP and computer vision, including scalable preprocessing, feature engineering, and robust model evaluation. Processed 1M+ text samples and 5K+ visual data, building automated NLP (cleaning, feature extraction) and CV (resizing, augmentation) pipelines. Used Transformer and DNN models (TensorFlow, PyTorch) to identify misinformation patterns with improved classification precision. Implemented multimodal feature extraction (OpenCV) and generated analytical reports with performance metrics to support model interpretability. Built reproducible ML workflows with clear train/test separation and production-ready validation practices.
Project Manager at Climate Change Fact-Check System
May 1, 2024 - May 1, 2024
Specialized in Transformer, CNN (TensorFlow, PyTorch) deep learning architectures, along with pre-trained Inception (OpenCV) models with transfer learning. Led model selection and optimization for a large-scale fact verification task using over 1 million Twitter comments and over 5,000 pictures. Achieved F-score 0.3967 in evidence retrieval and 0.4416 in claim classification on the development set, a 24% improvement over the baseline.
Data Engineer at Fujian Big Data Exchange Platform
February 1, 2024 - February 1, 2024
Designed and deployed ML models to predict bid-winning amounts and award timelines, transitioning from hypothesis through modelling, validation, and business delivery. Trained regression and time-series models using Python (scikit-learn) and PySpark MLlib to forecast bid outcomes. Cleaned, transformed, and engineered features across large, multi-source datasets with Spark DataFrames to improve data quality for downstream modelling. Optimised model performance via feature engineering and tuning to support data-driven bidding strategy. Communicated insights to stakeholders and built Tableau dashboards to visualise key features of large-scale data. Built automated preprocessing workflows on a Hadoop-based environment to enable scalable ETL and cloud-aligned data engineering practices.
Data Analyst Intern at Fujian Big Data Exchange Platform
February 1, 2024 - February 1, 2024
Managed and analyzed backend data using Java and Hadoop frameworks. Processed and analyzed large-scale bidding transaction data to support business decision-making. Contributed to negotiation strategy formulation and handled bidding documentation and invoice management.

Education

Master of Information Technology (Artificial Intelligence) at The University of Melbourne
March 1, 2024 - November 1, 2024
Bachelor of Data Science at The University of Melbourne
March 1, 2021 - November 1, 2023
Master of Information Technology at University of Melbourne
March 1, 2024 - December 1, 2025
Bachelor of Science at University of Melbourne
March 1, 2021 - December 1, 2023

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Education, Healthcare, Professional Services, Media & Entertainment, Computers & Electronics, Financial Services