← Search

Zhiheng Fu

12 accepted papers

2026

Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval

CVPR 2026

Composed Image Retrieval (CIR) has attracted significant attention due to its flexible multimodal query method, yet its development is severely constrained by the Noisy Triplet Correspondence (NTC) problem. Most existing robust learning methods rely on the "small loss hypothesis", but the unique sem

Cited by 0SourcecodeScholar
2026

ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval

CVPR 2026

The Composed Image Retrieval (CIR) task provides a flexible retrieval paradigm via a reference image and modification text, but it heavily relies on expensive and error-prone triplet annotations. This paper systematically investigates the Noisy Triplet Correspondence (NTC) problem introduced by anno

Cited by 0SourcecodeScholar
2026

HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval

AAAI 2026technical

Composed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal query composed of a reference image and modification text. Although this task has demonstrated promising applications in personalized search and recomme

Cited by 0SourcePDFScholar
2026

INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

AAAI 2026technical

Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting of reference images and modification texts. Although substantial progress has been made in recent years, existing methods assume that all samples are

Cited by 0SourcePDFScholar
2026

MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement Learning

CVPR 2026

Offline Multi-Agent Reinforcement Learning (MARL) is critical for coordinating multiple agents in costly and unsafe environments, yet existing methods struggle with high sensitivity to reward functions and weak generalization to new goals, limiting its practical impact. Inspired by single-agent Offl

Cited by 0SourceScholar
2026

ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval

AAAI 2026technical

With the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing attention from researchers. Unlike unimodal video retrieval methods, the CVR task takes a multi-modal query consisting of a reference video and a piece o

Cited by 0SourcePDFScholar
2025

Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval

IJCAI 2025

Current text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieve

Cited by 0SourcePDFScholar
2025

ENCODER: Entity Mining and Modification Relation Binding for Composed Image Retrieval

AAAI 2025technical

The objective of Composed Image Retrieval (CIR) is to identify a target image that meets the requirement based on a multimodal query (including the reference image and the modification text) provided by the user. Despite the notable success of existing approaches, they fail to adequately address the…

Cited by 2SourcePDFScholar
2025

PAIR: Complementarity-guided Disentanglement for Composed Image Retrieval

ICASSP 2025accepted

Composed Image Retrieval (CIR) is a novel image retrieval paradigm that aims at searching for the target images via the multimodal query including a reference image and a modification text. Although existing works have made significant progress, they overlook the inter-modal coherence and incoherenc…

Cited by 0SourceScholar
2024

AEDNet: Adaptive Embedding and Multiview-Aware Disentanglement for Point Cloud Completion

ECCV 2024poster

"Point cloud completion involves inferring missing parts of 3D objects from incomplete point cloud data. It requires a model that understands the global structure of the object and reconstructs local details. To this end, we propose a global perception and local attention network, termed AEDNet, for…

Cited by 1SourcePDFScholar
2023

VAPCNet: Viewpoint-Aware 3D Point Cloud Completion

ICCV 2023poster

Most existing learning-based 3D point cloud completion methods ignore the fact that the completion process is highly coupled with the viewpoint of a partial scan. However, the various viewpoints of incompletely scanned objects in real-world applications are normally unknown and directly estimating t…

Cited by 12PDFcodeScholar
2022

SLFNet: A Stereo and LiDAR Fusion Network for Depth Completion

RA-L 2022

Acquiring dense and precise depth information in real time is highly demanded for robotic perception and automatic driving. Motivated by the complementary nature of stereo images and LiDAR point clouds, we propose an efficient stereo-LiDAR fusion network (SLFNet) to predict a dense depth map of a sc

Cited by 14SourceScholar