← Search

Zhiyu Pan

12 accepted papers

2026

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modaliti…

Cited by 0SourcecodeScholar
2026

Semi-Supervised High Dynamic Range Image Reconstructing via Bi-Level Uncertain Area Masking

AAAI 2026technical

Reconstructing high dynamic range (HDR) images from low dynamic range (LDR) bursts plays an essential role in the computational photography. Impressive progress has been achieved by learning-based algorithms which require LDR-HDR image pairs. However, these pairs are hard to obtain, which motivates

Cited by 0SourcePDFScholar
2026

Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs

ICLR 2026oral

Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self-improving techniques that refine reasoning paths for subsequent fine-tuning. However, extending these language-based self-improving approaches to vision language models (V…

Cited by 0SourcecodeScholar
2025

Exploring Contextual Attribute Density in Referring Expression Counting

CVPR 2025poster

Referring expression counting (REC) algorithms are for more flexible and interactive counting ability across varied fine-grained text expressions. However, the requirement for fine-grained attribute understanding poses challenges for prior arts, as they struggle to accurately align attribute informa…

2025

SRefiner: Soft-Braid Attention for Multi-Agent Trajectory Refinement

ICCV 2025poster

Accurate prediction of multi-agent future trajectories is crucial for autonomous driving systems to make safe and efficient decisions. Trajectory refinement has emerged as a key strategy to enhance prediction accuracy. However, existing refinement methods often overlook the topological relationships…

2024

Camera-LiDAR Cross-modality Gait Recognition

ECCV 2024poster

"Gait recognition is a crucial biometric identification technique. Camera-based gait recognition has been widely applied in both research and industrial fields. LiDAR-based gait recognition has also begun to evolve most recently, due to the provision of 3D structural information. However, in certain…

2024

Self-Supervised Class-Agnostic Motion Prediction with Spatial and Temporal Consistency Regularizations

CVPR 2024poster

The perception of motion behavior in a dynamic environment holds significant importance for autonomous driving systems wherein class-agnostic motion prediction methods directly predict the motion of the entire point cloud. While most existing methods rely on fully-supervised learning the manual labe…

2024

Semi-supervised Class-Agnostic Motion Prediction with Pseudo Label Regeneration and BEVMix

AAAI 2024technical

Class-agnostic motion prediction methods aim to comprehend motion within open-world scenarios, holding significance for autonomous driving systems. However, training a high-performance model in a fully-supervised manner always requires substantial amounts of manually annotated data, which can be bot…

2023

Find Beauty in the Rare: Contrastive Composition Feature Clustering for Nontrivial Cropping Box Regression

AAAI 2023technical

Automatic image cropping algorithms aim to recompose images like human-being photographers by generating the cropping boxes with improved composition quality. Cropping box regression approaches learn the beauty of composition from annotated cropping boxes. However, the bias of annotations leads to q…

Cited by 6SourcePDFScholar
2023

When Epipolar Constraint Meets Non-Local Operators in Multi-View Stereo

ICCV 2023poster

Learning-based multi-view stereo (MVS) method heavily relies on feature matching, which requires distinctive and descriptive representations. An effective solution is to apply non-local feature aggregation, e.g., Transformer. Albeit useful, these techniques introduce heavy computation overheads for…

Cited by 33PDFcodeScholar
2021

TransView: Inside, Outside, and Across the Cropping View Boundaries

ICCV 2021poster

We show that relation modeling between visual elements matters in cropping view recommendation. Cropping view recommendation addresses the problem of image recomposition conditioned on the composition quality and the ranking of views (cropped sub-regions). This task is challenging because the visual…

Cited by 23PDFScholar