I am an undergraduate student pursuing a B.Sc. in Computational Data Science at The Chinese University of Hong Kong (CUHK) (GPA: 3.704/4.00), with an expected graduation in May 2027. My research interests focus on:

  • MLLM reasoning beyond text verbalization: pushing multimodal reasoning when visual evidence is not faithfully captured by a textual description.
  • Post-training & evaluation: benchmark design and diagnostic evaluations; training signals beyond imitation.
  • Safety & robustness of multimodal agents: reliability under uncertainty and distribution shift; mitigating harmful or non-grounded behaviors in multimodal decision-making.

I am actively involved in multiple research projects spanning multimodal large language models (MLLMs), spatial reasoning, and vision-language systems. I have experience working on post-training methods for MLLMs, building evaluation benchmarks, and developing distributed systems for LLM reinforcement learning.

🔥 News

  • 2026.03:  📄 Multiple papers submitted to NeurIPS 2026 (I-WebGenBench, Unify-Agent, OpenSearch-VL, Human Cognitive Benchmarks, Modality Interference).
  • 2026.03:  📄 “ReMAP-PET” and “PaperVoyager” submitted to ARR May 2026.
  • 2026.01:  🚀 Started reviewer service for Pattern Recognition.
  • 2025.11:  💼 Completed AI Software Engineer Internship at Huawei Technologies (2012 Labs).
  • 2025.06:  🔬 Started research on probing generalization boundaries of multimodal reasoning at CUHK.
  • 2025.02:  📄 “Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs” published on arXiv.
  • 2024.07:  🏆 Achieved Dean’s List recognition (Top 10%) for exceptional academic performance.
  • 2024.05:  🏅 Meritorious Winner in the MCM Mathematical Modeling Contest (COMAP).

📝 Publications

Accepted

MICCAI 2026
sym

IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation

Hao Wei, Wenjin Qi, Dasen Dai, Minqing Zhang, Wu Yuan

MICCAI 2026

Submitted / Under Review

NeurIPS 2026
sym

Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs

Jen-Tse Huang, Dasen Dai, Jen-Yuan Huang, Youliang Yuan, Xiaoyuan Liu, Wenxuan Wang, Wenxiang Jiao, Pinjia He, Zhaopeng Tu, Haodong Duan

NeurIPS 2026 under review   arXiv

  • Built a comprehensive benchmark suite based on the Kit of Factor-Referenced Cognitive Tests to evaluate MLLMs’ spatial intelligence, identifying significant cognitive gaps between human and machine vision.
  • Engineered an automated, scalable data generation pipeline to batch-produce spatial reasoning tasks with fine-grained difficulty control, creating an RL-ready training corpus.
NeurIPS 2026
sym

FMVP: Masked Flow Matching for Adversarial Video Purification

Duoxun Tang, Xueyi Zhang, Chak Hin Wang, Xi Xiao, Dasen Dai, Xinhang Jiang, Wentao Shi, Rui Li, Qing Li

NeurIPS 2026 under review   arXiv

  • Proposed a novel video purification framework integrating Conditional Flow Matching (CFM) with a masking strategy to physically disrupt adversarial patterns.
  • Designed a Frequency-Gated Loss (FGL) to suppress high-frequency adversarial noise while preserving low-frequency semantic fidelity.
NeurIPS 2026
sym

I-WebGenBench: Evaluating Interactivity in LLM-Generated Scientific Web Applications

Dasen Dai, Biao Wu, Meng Fang, Shuoqi Li, Wenhao Wang

NeurIPS 2026 under review

NeurIPS 2026
sym

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

Shuang Chen, Kaituo Feng, Hangting Chen, Wenxuan Huang, Dasen Dai, Quanxin Shou, Yunlong Lin, Xiangyu Yue, Shenghua Gao, Tianyu Pang

NeurIPS 2026 under review   arXiv

ARR 2026
sym

PaperVoyager: Building Interactive Web with Large Multimodal Models

Dasen Dai*, Biao Wu*, Meng Fang, Wenhao Wang

ARR May 2026 under review   arXiv

ARR 2026
sym

ReMAP-PET: Beyond Visual Understanding – Learning Region-Guided Metabolic Alignment Semantics from Brain PET

Dasen Dai*, Yanteng Zhang*, Shuoqi Li*, Yuxiang Wei, Hongjie Yu, Qingxin Zhang, Qizhen Lan, Jagath C. Rajapakse, Vince D. Calhoun

ARR May 2026 under review

ECCV 2026
sym

VidDoS: Universal DoS Attack on Video-based LLMs

Duoxun Tang, Dasen Dai, Jiyao Wang, Xiao Yang, Jianyu Wang, Siqi Cai

arXiv

(* denotes equal contribution)

🎖 Honors and Awards

  • 2025 Dean’s List (Top 10%), Faculty of Engineering, CUHK
  • 2024 Dr Shu-chia Yang GOAL Programme Memorial Scholarship, United College, CUHK
  • 2024 The Alumni Association of United College of the CUHK Ltd Prize, United College, CUHK
  • 2024 ELITE Stream Scholarship, Faculty of Engineering, CUHK
  • 2024 Dean’s List (Top 10%), Faculty of Engineering, CUHK
  • 2024 Talent Development Scholarship, HKSAR Government
  • 2024 Meritorious Winner, MCM Mathematical Modeling Contest, COMAP

📖 Education

  • 2024.09 - 2027.05 (Expected), B.Sc. in Computational Data Science (CDAS), The Chinese University of Hong Kong, New Territories, Hong Kong. GPA: 3.704/4.00

🔬 Research Experience

  • 2025.06 - 2026.04, Probing Generalization Boundaries of Multimodal Reasoning, Prof. Xiangyu Yue, CUHK, Hong Kong
    • Engineered a robust evaluation framework to quantify MLLM generalization via image-grounded logic puzzles, incorporating multi-level difficulty scaling and OOD puzzle-family shifts.
    • Developed a “solver-in-the-loop” pipeline by coupling LLMs with symbolic solvers to synthesize 2k+ verified Chain-of-Thought traces, facilitating the transition of Qwen3-VL-32B from instruction-following to advanced reasoning.
    • Systematically investigated easy-to-hard extrapolation and cross-task transferability, providing empirical insights into the scaling laws of multimodal logical reasoning.
  • 2024.09 - 2025.06, Benchmarking Spatial Reasoning Abilities of MLLMs, Dr. Jen-Tse Huang, JHU, Baltimore
    • Spearheaded the development of a comprehensive benchmark suite based on the Kit of Factor-Referenced Cognitive Tests to evaluate MLLMs’ spatial intelligence.
    • Engineered an automated, scalable data generation pipeline to batch-produce spatial reasoning tasks with fine-grained difficulty control, creating an RL-ready training corpus.
    • Conducted rigorous error analysis to categorize prevailing failure modes in spatial tasks, uncovering fundamental limitations of current MLLM architectures in geometric and topological reasoning.

💼 Industrial Experience

  • 2025.11 - 2026.01, AI Software Engineer Intern, Huawei Technologies (2012 Labs), Distributed & Parallel Software Lab, Shenzhen, China
    • Architected SampleTable, a high-throughput distributed data middleware for LLM Reinforcement Learning, enabling asynchronous collaboration between VLLM inference and policy training.
    • Designed a strong-typed columnar storage engine with fine-grained state management to orchestrate complex multi-agent workflows, resolving pipeline blocking issues in large-scale training.
    • Optimized system performance leveraging OpenEuler Yuanrong serverless infrastructure; implemented data-function affinity scheduling and zero-copy shared memory access, significantly reducing I/O latency.

💻 Technical Skills

  • Programming Languages: Python, C/C++, R
  • Frameworks: PyTorch, Verl, VLLM, Ray
  • Languages: Chinese (Native), English (Fluent)

🌐 Service

  • Journal Reviewer: Pattern Recognition (2026 – Present)