I’m Lingyu Li, a computing graduate student focused on building practical AI systems—especially LLM-driven agent workflows, post-training pipelines, and large-scale inference. I’ve worked across the full lifecycle, from designing routing and skill orchestration logic to constructing training/evaluation data loops that continuously improve accuracy and reliability in production-like customer service and dialogue settings. I enjoy turning complex requirements into robust engineering solutions: optimizing distributed training on GPU clusters, distilling and aligning models for better tool use and decision-making, and developing data pipelines for real-world trajectory and behavioral logs. Alongside hands-on ML work, I also value strong modeling foundations (including math modeling competitions) and like shipping measurable improvements such as higher dispatch accuracy, better recall at constrained interception rates, and lower training cost/latency.

Lingyu Li

I’m Lingyu Li, a computing graduate student focused on building practical AI systems—especially LLM-driven agent workflows, post-training pipelines, and large-scale inference. I’ve worked across the full lifecycle, from designing routing and skill orchestration logic to constructing training/evaluation data loops that continuously improve accuracy and reliability in production-like customer service and dialogue settings. I enjoy turning complex requirements into robust engineering solutions: optimizing distributed training on GPU clusters, distilling and aligning models for better tool use and decision-making, and developing data pipelines for real-world trajectory and behavioral logs. Alongside hands-on ML work, I also value strong modeling foundations (including math modeling competitions) and like shipping measurable improvements such as higher dispatch accuracy, better recall at constrained interception rates, and lower training cost/latency.

Available to hire

I’m Lingyu Li, a computing graduate student focused on building practical AI systems—especially LLM-driven agent workflows, post-training pipelines, and large-scale inference. I’ve worked across the full lifecycle, from designing routing and skill orchestration logic to constructing training/evaluation data loops that continuously improve accuracy and reliability in production-like customer service and dialogue settings.

I enjoy turning complex requirements into robust engineering solutions: optimizing distributed training on GPU clusters, distilling and aligning models for better tool use and decision-making, and developing data pipelines for real-world trajectory and behavioral logs. Alongside hands-on ML work, I also value strong modeling foundations (including math modeling competitions) and like shipping measurable improvements such as higher dispatch accuracy, better recall at constrained interception rates, and lower training cost/latency.

See more

Work Experience

Algorithm Engineer Intern at Shopee
May 1, 2026 - August 31, 2026
Designed a first-layer router for a multi-agent customer service system using Qwen3.5-4B to identify and route CS/BD intents; built an intent taxonomy plus hard-case mining/feedback loop, improving routing accuracy to 98%. Drove agentization of customer-service workflows across eight regions by decomposing tree-based flows into atomic units and redesigning them as a hierarchical skill architecture for scenario-aware skill selection/orchestration. Improved skill-dispatch accuracy from 33% to 91% by enhancing skill atomization, function-calling governance, system prompt design, and fallback optimization. Built a post-training data and evaluation pipeline processing 10,000+ real customer-service trajectories daily, enabling closed-loop bad-case mining and regression evaluation with improved reasoning data formats. Led Qwen3.6-27B post-training on a B300 cluster with Hindsight SEED and Few Teacher Steps to improve tool use and decision-making (90%+ tool-calling accuracy) and increased onli
Machine Learning Engineer Intern at Garena
January 1, 2026 - April 30, 2026
Built post-training infrastructure for 35B-scale models on an H200 cluster and optimized the training pipeline for long-context full-parameter fine-tuning. Addressed memory/throughput bottlenecks using FlashAttention-3, DeepSpeed ZeRO-3, and gradient checkpointing, improving memory utilization by 25%+ and training throughput by 40%+. Created an instruction-tuning data pipeline across Qwen and Gemma families (0.8B–35B) including multi-turn cleaning, quality filtering, hard-case feedback, and dynamic domain/general mixing to mitigate catastrophic forgetting. Designed CoT distillation data construction for complex scenarios; achieved 0.62+ on an in-house benchmark and approached Claude 3.7 under the same evaluation protocol. Built a multi-turn dialogue reward model and LLM-as-judge evaluation framework spanning 10+ dimensions, introducing conditional rewards and key-metric gating for long-tail scenarios. Optimized multi-turn policies with GRPO using group-relative rewards and SPIF augme
Algorithm Intern at Trip.com Group
September 1, 2025 - December 31, 2025
Constructed a robust training dataset from user environment and behavioral logs, extracted key risk features, and modeled long-sequence behaviors using Transformer/LSTM/TiSASRec to detect account takeover risk, achieving AUC=0.84 and substantially improving recall at a 0.2% interception rate. Enhanced DPO alignment pipeline to improve output controllability/stability, raising instruction-following accuracy by ~20% over SFT-only baselines while reducing hallucinations and inconsistent reasoning. Fine-tuned models to generate natural-language risk analysis reports (risk level + anomaly features + evidence chain), reaching F1=0.86 on internal test sets to support production deployment.
Backend Engineer Intern at Chengdu BigQuant Technology Co., Ltd.
May 1, 2025 - August 31, 2025
Implemented ReAct-style agent reasoning pipelines supporting multi-turn dialogue management, strategy generation, and document-grounded question answering, improving internal user satisfaction by 20%. Developed a scalable long-context memory system using Redis-based session storage, automatic summarization, and token-budget optimization, reducing LLM inference costs by 15%. Built a high-throughput document understanding pipeline for financial research PDFs, converting complex layouts into structured Markdown with metadata extraction support. Designed a distributed GPU-powered asynchronous processing architecture with Redis Queue and FastAPI, reducing PDF parsing latency by 30% and maintaining 95%+ task success rates under concurrent workloads.

Education

M.S. in Computing at National University of Singapore
July 1, 2024 - January 1, 2027
B.E. in Environmental Engineering, Minor in Economics at Harbin Institute of Technology, Shenzhen
September 1, 2020 - June 30, 2024

Qualifications

Add your qualifications or awards here.

Industry Experience

Software & Internet