Zhuoyun Du 杜卓耘

Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization

Shan He, Runze Wang, Zhuoyun Du, Huiyu Bai, Zouying Cao, Yucheng, and Bo Zheng

Findings of the Association for Computational Linguistics (ACL Findings), 2026

MemCoRL: Alternating Co-Optimization of Memory Retrieval and Utilization via Collaborative Reinforcement Learning

Yuewen Liu, Peng Xu, Zhuoyun Du, Muxi Diao, Anyi Zhang, Yang Li, and Yutong Zhang

Association for Computational Linguistics (ACL), 2026

Other Papers

* denotes equal contribution.

Experience

Talent Program Intern, ByteDance (Mar. 2026 - Jul. 2026)
  • Built the team's first data-flywheel rollout system for multi-turn, task-oriented dialogue, supporting downstream multi-turn dialogue RL and agentic RL.
  • Designed a self-play simulation and evaluation environment with high-fidelity, persona-conditioned user-simulation agents for rollout generation and evaluation signals.
  • Conducted post-training and preference alignment of large-scale LLMs (e.g., Qwen3.5-122B/397B) using SFT and RLHF/DPO, alongside autonomous agent-driven evaluation of dialogue policies.
Research Intern, Future Living Lab (now part of Token Foundry), Alibaba Group (Jan. 2025 - Mar. 2026)
  • Proposed Interlat (code), a latent-space communication framework that replaces natural-language message passing with temporally aligned hidden-state exchange.
  • Designed a token-latent curriculum and learned compression scheme, achieving up to 24x communication speedup at comparable task performance.
  • Gained hands-on experience with distributed training on 256+ GPUs; co-authored Online-PVLM, MemCoRL, and Learning to Evolve.
M.Eng. Research, Zhejiang University, State Key Lab of CAD&CG (Sept. 2024 - present)
  • Led EvoPatient (code), an open-source doctor-patient coevolution system for virtual standardized patients, accepted to ACL 2025.
  • Drove the project end-to-end, including research formulation, system implementation, medical-case curation, expert evaluation, and paper writing.
  • Improved patient-answer ability from 0.763 to 0.860 after 200 training cases.
Research Assistant, THUNLP, Tsinghua University (Dec. 2023 - Aug. 2024)
  • Worked with Prof. Zhiyuan Liu and Prof. Chen Qian on ChatDev-based multi-agent collaboration, studying role organization, information exchange, and coordination.
  • Led Croto (code), designing cross-team orchestration with parallel teams, hierarchy partitioning, and greedy aggregation for multi-agent software development.
  • Contributed to MACNet (code) and iAgent (code); related OpenBMB/ChatDev projects have accumulated over 33k GitHub stars.

Awards and Honors

  • 2025: National Scholarship; Outstanding Graduate Student Scholarship.
  • 2024: First-Class Scholarship for Outstanding Graduates.
  • Undergraduate: three undergraduate scholarships.

Academic Service

Miscellaneous

Beyond research, I enjoy soccer, fencing, snowboarding, billiards, ballroom dancing, piano, photography, and physics. I occasionally write notes on my personal blog.