Zhen Wan
PhD Student · Kyoto University · Kurohashi–Chu Lab · Research Intern at NVIDIA Research
Yoshida-Honmachi, Sakyo-ku
Kyoto 606-8501, Japan
zhenwan.nlp [at] gmail.com
I am a PhD student in Informatics (Intelligence Science & Technology) at Kyoto University, advised by Prof. Sadao Kurohashi and Prof. Chenhui Chu. I am also a research intern at NVIDIA Research, working with Dr. Chao-Han Huck Yang, Dr. Rafael Valle, and Dr. Boris Ginsburg on voice agentic and omni-perception models.
My research is on voice agentic systems and multimodal LLMs — specifically, teaching omni-models when to trust themselves versus when to consult external perception tools, and how to evaluate the intelligence of speech LLMs. Recent work includes:
- Speech-Hands (ACL 2026, Oral) — a self-reflection voice agentic framework for ASR and audio reasoning.
- SpeechIQ (ACL 2025 Main) — an agentic intelligence quotient for voice-understanding LLMs across cognitive levels.
- NeKo (ACL 2025, Best Industrial Paper Honor Mention) — cross-modality post-recognition error correction with mixture-of-experts.
I am supported by a JSPS DC2 fellowship (acceptance rate 19.1%). Before Kyoto, I received a B.E. in Energy Engineering from Zhejiang University. Outside research, I have been a research assistant on LLM-jp and the SIP medical LLM project, and previously interned at Alibaba DAMO Academy on Chinese legal LLMs.
I regularly review for ACL, EMNLP, NAACL and ARR. Feel free to reach out — I am happy to chat about voice agents, multimodal LLMs, low-resource Japanese NLP, or anything in between.
news
| Apr 21, 2026 | Speech-Hands is accepted as an Oral at ACL 2026 Main Conference 🎤 [project page] [code] |
|---|---|
| Sep 20, 2025 | Three papers accepted at EMNLP 2025 — CoVoGER (Main), Causal Tree Extraction (Main), and Cross-lingual Japanese Medical LLMs (Findings). |
| Aug 01, 2025 | NeKo received an ACL 2025 Best Industrial Paper Honor Mention 🎉, and SpeechIQ was presented at the ACL 2025 Main Conference. |
| Mar 01, 2025 | Started as a Research Intern at NVIDIA Research (remote, Japan), working on Omni models and voice agentic systems. |
| Apr 01, 2024 | Awarded the JSPS Research Fellowship for Young Scientists (DC2) — Multi-document & Multi-modal Information Extraction. |
selected publications
- ACL ’25SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language ModelsIn Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Main), 2025
- ACL ’25NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language ModelIn ACL 2025 Industrial Track, 2025Best Industrial Paper Honor Mention
- ACL ’24Reformulating Domain Adaptation of Large Language Models as Adapt-Retrieve-Revise: A Case Study on Chinese Legal DomainIn Findings of the Association for Computational Linguistics (ACL 2024), 2024
- EMNLP ’23GPT-RE: In-context Learning for Relation Extraction using Large Language ModelsIn Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023), 2023