I’m a data scientist and machine learning engineer with a strong math foundation (B.S. Mathematics and an M.S. in Applied Mathematics from the University of Houston). I focus on the principles behind model performance—empirical risk minimization, capacity control, and the generalization gap—and use them to design and evaluate systems across classification, clustering, forecasting, and modern deep learning. In my recent work at C++ Alliance, I build and maintain a production knowledge base over the Boost and C++ ecosystem using Pinecone, combining careful ingestion (chunking, metadata, namespaces) with hybrid dense/sparse retrieval and reranking to keep answers grounded in source material. I also develop tooling that makes this knowledge accessible to developers and agents, and I’ve built analytics pipelines over WG21 (ISO C++ standards) papers for topic classification, author activity, deduplication, and forecasting.

John Zhihao Zhao

I’m a data scientist and machine learning engineer with a strong math foundation (B.S. Mathematics and an M.S. in Applied Mathematics from the University of Houston). I focus on the principles behind model performance—empirical risk minimization, capacity control, and the generalization gap—and use them to design and evaluate systems across classification, clustering, forecasting, and modern deep learning. In my recent work at C++ Alliance, I build and maintain a production knowledge base over the Boost and C++ ecosystem using Pinecone, combining careful ingestion (chunking, metadata, namespaces) with hybrid dense/sparse retrieval and reranking to keep answers grounded in source material. I also develop tooling that makes this knowledge accessible to developers and agents, and I’ve built analytics pipelines over WG21 (ISO C++ standards) papers for topic classification, author activity, deduplication, and forecasting.

Available to hire

I’m a data scientist and machine learning engineer with a strong math foundation (B.S. Mathematics and an M.S. in Applied Mathematics from the University of Houston). I focus on the principles behind model performance—empirical risk minimization, capacity control, and the generalization gap—and use them to design and evaluate systems across classification, clustering, forecasting, and modern deep learning.

In my recent work at C++ Alliance, I build and maintain a production knowledge base over the Boost and C++ ecosystem using Pinecone, combining careful ingestion (chunking, metadata, namespaces) with hybrid dense/sparse retrieval and reranking to keep answers grounded in source material. I also develop tooling that makes this knowledge accessible to developers and agents, and I’ve built analytics pipelines over WG21 (ISO C++ standards) papers for topic classification, author activity, deduplication, and forecasting.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
Intermediate
Intermediate
Intermediate
Intermediate
See more

Language

English
Fluent
Chinese
Advanced

Work Experience

Data Scientist / Machine Learning Engineer at C++ Alliance (CppAlliance)
October 1, 2025 - Present
Built and operated a Pinecone knowledge base for the Boost and C++ ecosystem, including ingestion over public corpora (Boost mailing lists and documentation, C++ reference materials, Clang GitHub history, Reddit r/cpp, and WG21 paper corpora). Designed chunking and metadata strategies, namespaces, and hybrid dense/sparse retrieval with reranking so responses remain grounded in the corpus. Owned upsert/sync operations via cppa_pinecone_sync in boost-data-collector, including preprocessing orchestration and failure/retry tracking. Contributed to a production Pinecone read-only MCP TypeScript server exposing hybrid search, semantic reranking, namespace discovery, and typed error surfaces; improved robustness with benchmarking, test coverage gates, and end-to-end verification. Also developed WG21 author-data-mining analytics for topic dictionary classification, author activity analysis, paper deduplication, and topic-level time-series forecasting.
Data Scientist / Machine Learning Engineer at C++ Alliance
October 1, 2025 - Present
Built and operate a Pinecone knowledge base for the Boost and C++ ecosystem, including ingestion pipelines over public C++ sources (documentation, mailing lists, reference materials, and WG21 corpora). Designed chunking, metadata, namespaces, and hybrid dense/sparse retrieval with reranking to keep outputs grounded in the corpus, and owned ingestion/sync orchestration with failure/retry tracking via internal sync tooling. Contributed to and hardened a production Pinecone read-only MCP TypeScript server exposing hybrid search, reranking, namespace discovery, and typed tool surfaces for agent/IDE workflows. Also supported related boost-data-collector services, raising production readiness via test coverage, API documentation generation, Docker hardening, readiness/health endpoints, structured logging, and security governance.
Senior Data Scientist at Capgemini Engineering
January 1, 2022 - September 30, 2025
Led end-to-end data science delivery for enterprise clients, covering problem framing, feature engineering, and model development for classification, clustering, and forecasting. Emphasized evaluation design (cross-validation, calibration, and ROC/AUC-style metrics) rather than relying on training accuracy alone, and set up production monitoring. Designed MLOps workflows with MLflow and containerized inference, and mentored junior engineers on reproducible experimentation and evaluation. From 2023 onward, prototyped NLP and retrieval workflows (embeddings, semantic search, early RAG patterns) for knowledge-heavy domains, bridging industrial analytics into developer-tooling contexts.
Senior Data Scientist at Capgemini Engineering (Remote · EU project delivery)
January 1, 2022 - September 30, 2025
Led end-to-end data science delivery for enterprise clients: problem framing, feature engineering, model selection across classification, clustering, and forecasting, and production monitoring. Emphasized evaluation design (cross-validation, calibration, ROC/AUC style assessment) rather than optimizing for training accuracy alone. Designed MLOps workflows (MLflow, containerized inference, CI/CD) and mentored engineers on reproducible experimentation and trustworthy metrics. From 2023, prototyped NLP and early retrieval/RAG workflows (embeddings, semantic search) for knowledge-heavy domains, bridging industrial analytics experience into developer-tooling contexts.
Data Scientist / Machine Learning Engineer at Luxoft (DXC Technology)
May 1, 2018 - December 31, 2021
Built forecasting, classification, and clustering pipelines for manufacturing and supply-chain clients using sensor time-series and quality/demand metrics. Implemented boosted additive ensembles with capacity control to manage the bias–variance tradeoff, and used bagging/stacking with model disagreement as a signal of predictive uncertainty. Applied information-theoretic feature selection (mutual information/dependence ranking) for high-dimensional operational factors. Developed Python microservices and batch jobs to ingest data, train models, and expose REST APIs for BI/MES integration; partnered on predictive maintenance/anomaly detection proofs of concept and produced deployment runbooks for client handoff.
Data Scientist / Machine Learning Engineer at Luxoft (DXC Technology) (Remote · European client engagements)
May 1, 2018 - December 31, 2021
Built forecasting, classification, and clustering pipelines for manufacturing and supply-chain domains, including sensor time series and demand-quality related metrics. Developed boosted additive ensembles with explicit bias–variance control (e.g., limited tree depth/shrinkage) and used bagged/stacked disagreement as a predictive uncertainty signal. Applied information-theoretic feature selection (mutual information/dependence ranking) for high-dimensional operational factors. Implemented Python microservices and batch jobs to ingest operational data, train models, and serve REST APIs for BI/MES integrations; supported predictive maintenance/anomaly detection proof-of-concepts and documented deployment runbooks for client handoff.

Education

M.S. Applied Mathematics at University of Houston
January 1, 2014 - January 1, 2016
B.S. Mathematics at University of Houston
January 1, 2009 - January 1, 2013
B.S. in Mathematics at University of Houston
January 1, 2009 - January 1, 2013
M.S. in Applied Mathematics at University of Houston
January 1, 2014 - January 1, 2016

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet, Professional Services, Manufacturing, Computers & Electronics, Education