← Search

Jinrong Zhang

4 accepted papers

2026

Breaking the Regional Perception Bottleneck of Multimodal Large Language Models via External Reasoning Framework

CVPR 2026

High-quality pixel-level responses remain a major bottleneck for multimodal large language models (MLLMs) in regional perception. Existing approaches generally attach regression decoders to MLLM features, achieving strong grounding performance but compromising end-to-end design and increasing traini

Cited by 0SourceScholar
2025

Cluster-Refined Optimal Transport for Unsupervised Action Segmentation

ICASSP 2025accepted

Action segmentation in untrimmed videos is essential for comprehensive video understanding. Despite significant progress in unsupervised methods, capturing both long-range dependencies and short-duration actions simultaneously remains a challenging task. To address this challenge, this paper introdu…

Cited by 0SourceScholar
2025

DTOS: Dynamic Time Object Sensing with Large Multimodal Model

CVPR 2025poster

Existing multimodal large language models (MLLMs) face significant challenges in Referring Video Object Segmentation(RVOS). We identify three critical challenges: (C1) insufficient quantitative representation of textual numerical data, (C2) repetitive and degraded response templates for spatiotempor…

2025

Just a Few Glances: Open-Set Visual Perception with Image Prompt Paradigm

AAAI 2025technical

To break through the limitations of pre-training models on fixed categories, Open-Set Object Detection (OSOD) and Open-Set Segmentation (OSS) have attracted a surge of interest from researchers. Inspired by large language models, mainstream OSOD and OSS methods generally utilize text as a prompt, ac…

Cited by 0SourcePDFScholar