← Search

Wenwen Tong

4 accepted papers

2026

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

ICML 2026poster

Recent Omni-MLLMs are driving a paradigm shift in multimodal emotion recognition from label-only prediction toward *Multimodal Emotion Reasoning* (MER), where models output both emotions and textual explanations grounded in visual, acoustic, and linguistic signals. However, we show that current emot…

Cited by 0SourceScholar
2026

SenseSearch: Empowering Vision-Language Models with High-Resolution Agentic Search-Reasoning via Reinforcement Learning

CVPR 2026

Vision-Language Models (VLMs) are limited by static knowledge and insufficient fine-grained visual analysis, hindering their performance on knowledge-intensive and visually complex tasks. While recent research has explored VLMs that employ external tools like search or cropping to enhance model perf

Cited by 0SourcecodeScholar
2026

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning

ICML 2026poster

Multimodal large language models (MLLMs) have achieved remarkable success in general perception, yet complex multi-step visual reasoning remains a persistent challenge. Although recent agentic approaches incorporate tool use, they often neglect critical execution feedback. Consequently, they suffer …

Cited by 0SourceScholar