← Search

Ming Dai

9 accepted papers

2026

DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation

CVPR 2026

Referring video object segmentation (RVOS) aims to segment objects within a video according to natural language expressions. Unlike earlier works focusing on static single-object scenarios, recent studies address more complex motion scenes. Previous methods typically adopt a query-based, logically m

Cited by 0SourceScholar
2026

Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMs

CVPR 2026

Recently, multimodal large language models (MLLMs) have achieved remarkable success in general multimodal tasks. Increasing attention has been given to leveraging MLLMs for fine-grained visual understanding, such as region-level captioning and pixel-level grounding. However, most existing approaches

Cited by 0SourceScholar
2026

Q Cache: Visual Attention Is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model

AAAI 2026technical

Multimodal large language models (MLLMs) are plagued by exorbitant inference costs attributable to the profusion of visual tokens within the vision encoder. The redundant visual tokens engenders a substantial computational load and key-value (KV) cache footprint bottleneck. Existing approaches focus

Cited by 0SourcePDFScholar
2026

VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

ICML 2026poster

Reasoning Video Object Segmentation (RVOS) demands a sophisticated integration of temporal dynamics, spatial details, and linguistic reasoning to achieve precise pixel-level localization. Existing methods are limited to reasoning over fixed initial inputs and lack the capacity to actively acquire fu…

Cited by 0SourceScholar
2025

DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy

ICCV 2025poster

Referring Image Segmentation (RIS) is a challenging task that aims to segment objects in an image based on natural language expressions. While prior studies have predominantly concentrated on improving vision-language interactions and achieving fine-grained localization, a systematic analysis of the…

2025

Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints

AAAI 2025technical

Multi-task visual grounding involves the simultaneous execution of localization and segmentation in images based on textual expressions. The majority of advanced methods predominantly focus on transformer-based multimodal fusion, aiming to extract robust multimodal representations. However, ambiguit…

2025

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination

ICCV 2025poster

Recent advances in visual grounding have largely shifted away from traditional proposal-based two-stage frameworks due to their inefficiency and high computational complexity, favoring end-to-end direct reference paradigms. However, these methods rely exclusively on the referred target for supervisi…

2025

ST3: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming

AAAI 2025technical

Multimodal large language models (MLLMs) enhance their perceptual capabilities by integrating visual and textual information. However, processing the massive number of visual tokens incurs a significant computational cost. Existing analysis of the MLLM attention mechanisms remains shallow, leading t…

Cited by 2SourcePDFScholar
2019

Restricted Orientation Dubins Path With Application to Sailboats

RA-L 2019

This letter develops a geometrical construction of the shortest Dubins path in a discontinuous orientation-restricted environment. The method proposed here builds the shortest path from one pose to the other while avoiding a no-go zone in terms of orientation, and being constrained to move forward.

Cited by 12SourceScholar