← Search

Mingrui Chen

4 accepted papers

2026

Think 360deg: Beyond Depth: Evaluating the Width-centric Reasoning Capability of MLLMs

CVPR 2026

In this paper, we present a holistic multimodal benchmark that evaluates the reasoning capabilities of MLLMs with an explicit focus on reasoning width, a complementary dimension to the more commonly studied reasoning depth. Specifically, reasoning depth measures the model's ability to carry out long

Cited by 0SourceScholar
2025

Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens

ICCV 2025poster

The Vision Transformer (ViT) has gained prominence for its superior relational modeling prowess. However, its global attention mechanism's quadratic complexity poses substantial computational burdens. A common remedy spatially groups tokens for self-attention, reducing computational requirements. No…

Cited by 0SourcePDFScholar
2024

RMT: Retentive Networks Meet Vision Transformers

CVPR 2024poster

Vision Transformer (ViT) has gained increasing attention in the computer vision community in recent years. However the core component of ViT Self-Attention lacks explicit spatial priors and bears a quadratic computational complexity thereby constraining the applicability of ViT. To alleviate these i…

2023

Occ^2Net: Robust Image Matching Based on 3D Occupancy Estimation for Occluded Regions

ICCV 2023poster

Image matching is a fundamental and critical task in various visual applications, such as Simultaneous Localization and Mapping (SLAM) and image retrieval, which require accurate pose estimation. However, most existing methods ignore the occlusion relations between objects caused by camera motion an…

Cited by 4PDFScholar