Zhen Wan

PhD Student · Kyoto University · Kurohashi–Chu Lab · Research Intern at NVIDIA Research

prof_pic.jpg

Yoshida-Honmachi, Sakyo-ku

Kyoto 606-8501, Japan

zhenwan.nlp [at] gmail.com

I am a PhD student in Informatics (Intelligence Science & Technology) at Kyoto University, advised by Prof. Sadao Kurohashi and Prof. Chenhui Chu. I am also a research intern at NVIDIA Research, working with Dr. Chao-Han Huck Yang, Dr. Rafael Valle, and Dr. Boris Ginsburg on voice agentic and omni-perception models.

My research is on voice agentic systems and multimodal LLMs — specifically, teaching omni-models when to trust themselves versus when to consult external perception tools, and how to evaluate the intelligence of speech LLMs. Recent work includes:

  • Speech-Hands (ACL 2026, Oral) — a self-reflection voice agentic framework for ASR and audio reasoning.
  • SpeechIQ (ACL 2025 Main) — an agentic intelligence quotient for voice-understanding LLMs across cognitive levels.
  • NeKo (ACL 2025, Best Industrial Paper Honor Mention) — cross-modality post-recognition error correction with mixture-of-experts.

I am supported by a JSPS DC2 fellowship (acceptance rate 19.1%). Before Kyoto, I received a B.E. in Energy Engineering from Zhejiang University. Outside research, I have been a research assistant on LLM-jp and the SIP medical LLM project, and previously interned at Alibaba DAMO Academy on Chinese legal LLMs.

I regularly review for ACL, EMNLP, NAACL and ARR. Feel free to reach out — I am happy to chat about voice agents, multimodal LLMs, low-resource Japanese NLP, or anything in between.

news

Apr 21, 2026 Speech-Hands is accepted as an Oral at ACL 2026 Main Conference 🎤 [project page] [code]
Sep 20, 2025 Three papers accepted at EMNLP 2025CoVoGER (Main), Causal Tree Extraction (Main), and Cross-lingual Japanese Medical LLMs (Findings).
Aug 01, 2025 NeKo received an ACL 2025 Best Industrial Paper Honor Mention 🎉, and SpeechIQ was presented at the ACL 2025 Main Conference.
Mar 01, 2025 Started as a Research Intern at NVIDIA Research (remote, Japan), working on Omni models and voice agentic systems.
Apr 01, 2024 Awarded the JSPS Research Fellowship for Young Scientists (DC2)Multi-document & Multi-modal Information Extraction.

selected publications

  1. ACL ’26
    Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
    Zhen Wan, Chao-Han Huck Yang, Jinchuan Tian, and 14 more authors
    In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), 2026
    Oral presentation
  2. ACL ’25
    SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
    Zhen Wan, Chao-Han Huck Yang, Yahan Yu, and 8 more authors
    In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Main), 2025
  3. ACL ’25
    NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
    Yen-Ting Lin, Zhehuai Chen, Piotr Zelasko, and 11 more authors
    In ACL 2025 Industrial Track, 2025
    Best Industrial Paper Honor Mention
  4. ACL ’24
    Reformulating Domain Adaptation of Large Language Models as Adapt-Retrieve-Revise: A Case Study on Chinese Legal Domain
    Zhen Wan, Yating Zhang, Yexiang Wang, and 2 more authors
    In Findings of the Association for Computational Linguistics (ACL 2024), 2024
  5. EMNLP ’23
    GPT-RE: In-context Learning for Relation Extraction using Large Language Models
    Zhen Wan, Fei Cheng, Zhuoyuan Mao, and 4 more authors
    In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023), 2023