← Search

Yunzhi Zhuge

15 accepted papers

2026

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

ICLR 2026poster

Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventional video LLMs evaluated in offline settings. The shift to an online, streaming paradigm introduces significant challen…

Cited by 0SourceScholar
2026

Reinforcing Video Object Segmentation to Think before it Segments

CVPR 2026

Video reasoning segmentation (VRS) endeavors to delineate referred objects in videos guided by implicit instructions that encapsulate human intent and temporal logic. Previous approaches leverage large vision language models (LVLMs) to encode object semantics into \SEG tokens for mask prediction. Ho

Cited by 0SourceScholar
2026

Spatial-Frequency Spiking Neural Network for Underwater Object Detection

AAAI 2026technical

Underwater object detection presents significant challenges due to the unique visual degradations in underwater environments, such as low contrast, poor visibility, and blurry object boundaries. While ANNs have achieved impressive detection accuracy, their high computational cost and power consumpti

Cited by 0SourcePDFScholar
2025

Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding

AAAI 2025technical

Injecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance o…

2025

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

NeurIPS 2025poster

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particul…

Cited by 0SourceScholar
2025

Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

ICLR 2025poster

Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long video sequences, supporting multi-turn dialogues, and adapti…

2025

The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

CVPR 2025poster

Existing methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To overcome these challenges, we propose VRS-HQ, an end-to-end video reasoning segme…

2025

Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation

AAAI 2025technical

Recently, deep learning based methods have revolutionized remote sensing image segmentation. However, these methods usually rely on a predefined semantic class set, thus needing additional image annotation and model training when adapting to new classes. More importantly, they are unable to segment…

2024

Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters

CVPR 2024poster

Continual learning can empower vision-language models to continuously acquire new knowledge without the need for access to the entire historical dataset. However mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout lifelong learning and (…

2024

DME: Unveiling the Bias for Better Generalized Monocular Depth Estimation

AAAI 2024technical

This paper aims to design monocular depth estimation models with better generalization abilities. To this end, we have conducted quantitative analysis and discovered two important insights. First, the Simulation Correlation phenomenon, commonly seen in long-tailed classification problems, also exist…

2024

LLMs Can Evolve Continually on Modality for $\mathbb{X}$-Modal Reasoning

NeurIPS 2024poster

Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily on extensive modal-specific pretraining and joint-modal tuning, leading to significant computational burdens when expand…

2024

SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning

ECCV 2024poster

"Parameter-efficient transfer learning (PETL) has emerged as a flourishing research field for adapting large pre-trained models to downstream tasks, greatly reducing trainable parameters while grappling with memory challenges during fine-tuning. To address it, memory-efficient series (METL) avoid ba…

2023

CTVIS: Consistent Training for Online Video Instance Segmentation

ICCV 2023poster

The discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/nega…

Cited by 46PDFcodeScholar
2019

Joint Learning of Saliency Detection and Weakly Supervised Semantic Segmentation

ICCV 2019poster

Existing weakly supervised semantic segmentation (WSSS) methods usually utilize the results of pre-trained saliency detection (SD) models without explicitly modelling the connections between the two tasks, which is not the most efficient configuration. Here we propose a unified multi-task learning f…

Cited by 246PDFcodeScholar
2019

Multi-Source Weak Supervision for Saliency Detection

CVPR 2019poster

The high cost of pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source usually does not contain enough information to train a well-performing model. To this end, we propose a unified framework to train saliency…

Cited by 227PDFcodeScholar