I am an undergraduate student pursuing a B.Sc. in Computational Data Science at The Chinese University of Hong Kong (CUHK) (GPA: 3.704/4.00), with an expected graduation in May 2027. My research interests focus on:
- MLLM reasoning beyond text verbalization: pushing multimodal reasoning when visual evidence is not faithfully captured by a textual description.
- Post-training & evaluation: benchmark design and diagnostic evaluations; training signals beyond imitation.
- Safety & robustness of multimodal agents: reliability under uncertainty and distribution shift; mitigating harmful or non-grounded behaviors in multimodal decision-making.
I am actively involved in multiple research projects spanning multimodal large language models (MLLMs), spatial reasoning, and vision-language systems. I have experience working on post-training methods for MLLMs, building evaluation benchmarks, and developing distributed systems for LLM reinforcement learning.
🔥 News
- 2026.03: 📄 Multiple papers submitted to NeurIPS 2026 (I-WebGenBench, Unify-Agent, OpenSearch-VL, Human Cognitive Benchmarks, Modality Interference).
- 2026.03: 📄 “ReMAP-PET” and “PaperVoyager” submitted to ARR May 2026.
- 2026.01: 🚀 Started reviewer service for Pattern Recognition.
- 2025.11: 💼 Completed AI Software Engineer Internship at Huawei Technologies (2012 Labs).
- 2025.06: 🔬 Started research on probing generalization boundaries of multimodal reasoning at CUHK.
- 2025.02: 📄 “Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs” published on arXiv.
- 2024.07: 🏆 Achieved Dean’s List recognition (Top 10%) for exceptional academic performance.
- 2024.05: 🏅 Meritorious Winner in the MCM Mathematical Modeling Contest (COMAP).
📝 Publications
Accepted

IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation
Hao Wei, Wenjin Qi, Dasen Dai, Minqing Zhang, Wu Yuan
MICCAI 2026
Submitted / Under Review

Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
Jen-Tse Huang, Dasen Dai, Jen-Yuan Huang, Youliang Yuan, Xiaoyuan Liu, Wenxuan Wang, Wenxiang Jiao, Pinjia He, Zhaopeng Tu, Haodong Duan
NeurIPS 2026 under review arXiv
- Built a comprehensive benchmark suite based on the Kit of Factor-Referenced Cognitive Tests to evaluate MLLMs’ spatial intelligence, identifying significant cognitive gaps between human and machine vision.
- Engineered an automated, scalable data generation pipeline to batch-produce spatial reasoning tasks with fine-grained difficulty control, creating an RL-ready training corpus.

FMVP: Masked Flow Matching for Adversarial Video Purification
Duoxun Tang, Xueyi Zhang, Chak Hin Wang, Xi Xiao, Dasen Dai, Xinhang Jiang, Wentao Shi, Rui Li, Qing Li
NeurIPS 2026 under review arXiv
- Proposed a novel video purification framework integrating Conditional Flow Matching (CFM) with a masking strategy to physically disrupt adversarial patterns.
- Designed a Frequency-Gated Loss (FGL) to suppress high-frequency adversarial noise while preserving low-frequency semantic fidelity.

I-WebGenBench: Evaluating Interactivity in LLM-Generated Scientific Web Applications
Dasen Dai, Biao Wu, Meng Fang, Shuoqi Li, Wenhao Wang
NeurIPS 2026 under review

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
Shuang Chen, Kaituo Feng, Hangting Chen, Wenxuan Huang, Dasen Dai, Quanxin Shou, Yunlong Lin, Xiangyu Yue, Shenghua Gao, Tianyu Pang
NeurIPS 2026 under review arXiv

PaperVoyager: Building Interactive Web with Large Multimodal Models
Dasen Dai*, Biao Wu*, Meng Fang, Wenhao Wang
ARR May 2026 under review arXiv

Dasen Dai*, Yanteng Zhang*, Shuoqi Li*, Yuxiang Wei, Hongjie Yu, Qingxin Zhang, Qizhen Lan, Jagath C. Rajapakse, Vince D. Calhoun
ARR May 2026 under review

VidDoS: Universal DoS Attack on Video-based LLMs
Duoxun Tang, Dasen Dai, Jiyao Wang, Xiao Yang, Jianyu Wang, Siqi Cai
- Mitigating Modality Interference for Unified Reasoning and Perception in Multimodal Large Language Models, Shuang Chen, Yimeng Ye, Dasen Dai, Yicheng Xiao, Wenxuan Huang, Kaituo Feng, Kaixuan Fan, Manyuan Zhang, Yucheng Zhou, Hanwen Du, Haoxiao Wang, Ziqian Bi, Youhua Li, Tianyu Shi. NeurIPS 2026 under review.
- Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis, Shuang Chen, Quanxin Shou, Hangting Chen, Yucheng Zhou, Kaituo Feng, Wenbo Hu, Yi-Fan Zhang, Yunlong Lin, Wenxuan Huang, Mingyang Song, Dasen Dai, Bolin Jiang, Manyuan Zhang, Shi-Xue Zhang, Zhengkai Jiang, Lucas Wang, Zhao Zhong, Yu Cheng, Nanyun Peng. NeurIPS 2026 under review.
(* denotes equal contribution)
🎖 Honors and Awards
- 2025 Dean’s List (Top 10%), Faculty of Engineering, CUHK
- 2024 Dr Shu-chia Yang GOAL Programme Memorial Scholarship, United College, CUHK
- 2024 The Alumni Association of United College of the CUHK Ltd Prize, United College, CUHK
- 2024 ELITE Stream Scholarship, Faculty of Engineering, CUHK
- 2024 Dean’s List (Top 10%), Faculty of Engineering, CUHK
- 2024 Talent Development Scholarship, HKSAR Government
- 2024 Meritorious Winner, MCM Mathematical Modeling Contest, COMAP
📖 Education
- 2024.09 - 2027.05 (Expected), B.Sc. in Computational Data Science (CDAS), The Chinese University of Hong Kong, New Territories, Hong Kong. GPA: 3.704/4.00
🔬 Research Experience
- 2025.06 - 2026.04, Probing Generalization Boundaries of Multimodal Reasoning, Prof. Xiangyu Yue, CUHK, Hong Kong
- Engineered a robust evaluation framework to quantify MLLM generalization via image-grounded logic puzzles, incorporating multi-level difficulty scaling and OOD puzzle-family shifts.
- Developed a “solver-in-the-loop” pipeline by coupling LLMs with symbolic solvers to synthesize 2k+ verified Chain-of-Thought traces, facilitating the transition of Qwen3-VL-32B from instruction-following to advanced reasoning.
- Systematically investigated easy-to-hard extrapolation and cross-task transferability, providing empirical insights into the scaling laws of multimodal logical reasoning.
- 2024.09 - 2025.06, Benchmarking Spatial Reasoning Abilities of MLLMs, Dr. Jen-Tse Huang, JHU, Baltimore
- Spearheaded the development of a comprehensive benchmark suite based on the Kit of Factor-Referenced Cognitive Tests to evaluate MLLMs’ spatial intelligence.
- Engineered an automated, scalable data generation pipeline to batch-produce spatial reasoning tasks with fine-grained difficulty control, creating an RL-ready training corpus.
- Conducted rigorous error analysis to categorize prevailing failure modes in spatial tasks, uncovering fundamental limitations of current MLLM architectures in geometric and topological reasoning.
💼 Industrial Experience
- 2025.11 - 2026.01, AI Software Engineer Intern, Huawei Technologies (2012 Labs), Distributed & Parallel Software Lab, Shenzhen, China
- Architected SampleTable, a high-throughput distributed data middleware for LLM Reinforcement Learning, enabling asynchronous collaboration between VLLM inference and policy training.
- Designed a strong-typed columnar storage engine with fine-grained state management to orchestrate complex multi-agent workflows, resolving pipeline blocking issues in large-scale training.
- Optimized system performance leveraging OpenEuler Yuanrong serverless infrastructure; implemented data-function affinity scheduling and zero-copy shared memory access, significantly reducing I/O latency.
💻 Technical Skills
- Programming Languages: Python, C/C++, R
- Frameworks: PyTorch, Verl, VLLM, Ray
- Languages: Chinese (Native), English (Fluent)
🌐 Service
- Journal Reviewer: Pattern Recognition (2026 – Present)