← Search

Zhaobo Qi

8 accepted papers

2026

ActiveScope: Actively Seeking and Correcting Perception for MLLMs

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in vision-language understanding, yet they still struggle with fine-grained perception in high-resolution images. While existing training-free methods typically rely on attention-based localization or coarse-to-fine s…

Cited by 0SourceScholar
2026

PTNET: A PROPOSAL-CENTRIC TRANSFORMER NET- WORK FOR 3D OBJECT DETECTION

ICLR 2026poster

3D object detection from LiDAR point cloud data is important for autonomous driving systems. Recent two-stage 3D object detectors struggle to achieve satisfactory performance due to limitations in proposal quality, stemming from the degradation of geometric detail information in the generated propos…

Cited by 0SourceScholar
2025

Enhancing Pre-trained Representation Classifiability can Boost its Interpretability

ICLR 2025spotlight

The visual representation of a pre-trained model prioritizes the classifiability on downstream tasks, while the widespread applications for pre-trained visual models have posed new requirements for representation interpretability. However, it remains unclear whether the pre-trained representations c…

2025

Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval

ICLR 2025poster

With the explosive growth of video data, finding videos that meet detailed requirements in large datasets has become a challenge. To address this, the composed video retrieval task has been introduced, enabling users to retrieve videos using complex queries that involve both visual and textual infor…

2025

Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos

ICLR 2025poster

In this paper, we address the challenge of procedure planning in instructional videos, aiming to generate coherent and task-aligned action sequences from start and end visual observations. Previous work has mainly relied on text-level supervision to bridge the gap between observed states and unobser…

Cited by 0SourcePDFScholar
2025

Procedure Knowledge Decoupled Distillation Strategy for Procedure Planning in Instructional Videos

AAAI 2025technical

Procedure planning in instructional videos, producing a structured and plannable action sequence facilitating the transition from the start to the goal states, has achieved significant progress. The dominant single-branch non-autoregressive planning paradigm guides action sequence generation through…

2025

Video Language Model Pretraining with Spatio-temporal Masking

CVPR 2025poster

The development of self-supervised video-language models based on mask learning has significantly advanced downstream video tasks. These models leverage masked reconstruction to facilitate joint learning of visual and linguistic information. However, recent study reveals that reconstructing image fe…

Cited by 0SourcePDFScholar
2024

Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in Video

AAAI 2024technical

Temporal Sentence Grounding in Video (TSGV) is troubled by dataset bias issue, which is caused by the uneven temporal distribution of the target moments for samples with similar semantic components in input videos or query texts. Existing methods resort to utilizing prior knowledge about bias to art…