← Search

Zhihai He

20 accepted papers

2026

CausalLens: Sensitivity-Guided Multi-Head Causal Intervention for Hallucination Mitigation in Large Vision-Language Models

CVPR 2026

Recent Large Vision-Language Models (LVLMs) have shown impressive capabilities in multimodal understanding and generation. Despite this progress, they remain prone to *hallucination*, where model outputs conflict with the visual input due to an over-reliance on textual priors. Existing inference-tim

Cited by 0SourceScholar
2026

OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text

ICLR 2026poster

Composed video retrieval presents a complex challenge: retrieving a target video based on a source video and a textual modification instruction. This task demands fine-grained reasoning over multimodal transformations. However, existing benchmarks predominantly focus on vision–text alignment, largel…

Cited by 0SourceScholar
2026

OmniPortrait: Fine-Grained Personalized Portrait Synthesis via Pivotal Optimization

ICLR 2026poster

Image identity customization aims to synthesize realistic and diverse portraits of a specified identity, given a reference image and a text prompt. This task presents two key challenges: (1) generating realistic portraits that preserve fine-grained facial details of the reference identity, and (2) m…

Cited by 0SourceScholar
2026

TRAINING-FREE TEST-TIME ADAPTATION WITH BROWNIAN DISTANCE COVARIANCE IN VISION-LANGUAGE MODELS

ICASSP 2026poster

Vision-language models suffer performance degradation under domain shift, limiting real-world applicability. Existing test-time adaptation methods are computationally intensive, rely on back-propagation, and often focus on single modalities. To address these issues, we propose Training-free Test-Tim…

Cited by 0SourcePDFScholar
2025

Conditional Latent Coding with Learnable Synthesized Reference for Deep Image Compression

AAAI 2025technical

In this paper, we study how to synthesize a dynamic reference from an external dictionary to perform conditional coding of the input image in the latent domain and how to learn the conditional latent synthesis and coding modules in an end-to-end manner. Our approach begins by constructing a universa…

2025

Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations

AAAI 2025technical

We introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO address…

Cited by 1SourcePDFScholar
2024

Concept-Guided Prompt Learning for Generalization in Vision-Language Models

AAAI 2024technical

Contrastive Language-Image Pretraining (CLIP) model has exhibited remarkable efficacy in establishing cross-modal connections between texts and images, yielding impressive performance across a broad spectrum of downstream applications through fine-tuning. However, for generalization tasks, the curre…

2024

Cross-Constrained Progressive Inference for 3D Hand Pose Estimation with Dynamic Observer-Decision-Adjuster Networks

AAAI 2024technical

Generalization is very important for pose estimation, especially for 3D pose estimation where small changes in the 2D images could trigger structural changes in the 3D space. To achieve generalization, the system needs to have the capability of detecting estimation errors by double-checking the pro…

Cited by 0SourcePDFScholar
2024

Learning Inference-Time Drift Sensor-Actuator for Domain Generalization

ICASSP 2024accepted

In machine learning tasks, models trained in the source domain often suffer from performance degradation in the target domain due to domain drift or distribution shift. In this paper, we explore the concept of sensor-actuator design in adaptive control to address this domain drift problem and develo…

Cited by 0SourceScholar
2023

Neuro-Modulated Hebbian Learning for Fully Test-Time Adaptation

CVPR 2023poster

Fully test-time adaptation aims to adapt the network model based on sequential analysis of input samples during the inference stage to address the cross-domain performance degradation problem of deep neural networks. We take inspiration from the biological plausibility learning where the neuron resp…

2023

Self-Correctable and Adaptable Inference for Generalizable Human Pose Estimation

CVPR 2023poster

A central challenge in human pose estimation, as well as in many other machine learning and prediction tasks, is the generalization problem. The learned network does not have the capability to characterize the prediction error, generate feedback information from the test sample, and correct the pred…

Cited by 19SourcePDFScholar
2022

Coded Residual Transform for Generalizable Deep Metric Learning

NeurIPS 2022accept

A fundamental challenge in deep metric learning is the generalization capability of the feature embedding network model since the embedding network learned on training classes need to be evaluated on new test classes. To address this challenge, in this paper, we introduce a new method called coded…

Cited by 4SourcePDFScholar
2022

Self-Constrained Inference Optimization on Structural Groups for Human Pose Estimation

ECCV 2022poster

"We observe that human poses exhibit strong group-wise structural correlation and spatial coupling between keypoints due to the biological constraints of different body parts. This group-wise structural correlation can be explored to improve the accuracy and robustness of human pose estimation. In t…

Cited by 19SourcePDFScholar
2021

Relative Order Analysis and Optimization for Unsupervised Deep Metric Learning

CVPR 2021poster

In unsupervised learning of image features without labels, especially on datasets with fine-grained object classes, it is often very difficult to tell if a given image belongs to one specific object class or another, even for human eyes. However, we can reliably tell if image C is more similar to im…

Cited by 14PDFcodeScholar
2020

Unsupervised Deep Metric Learning with Transformed Attention Consistency and Contrastive Clustering Loss

ECCV 2020poster

Existing approaches for unsupervised metric learning focus on exploring self-supervision information within the input image itself. We observe that, when analyzing images, human eyes often compare images against each other instead of examining images individually. In addition, they often pay attenti…

Cited by 26SourcePDFScholar
2017

Rate-coverage analysis and optimization for joint audio-video multimedia retrieval

ICASSP 2017accepted

In this work, we consider the problem of automatic content retrieval (ACR) using joint audio-video fingerprints. We focus on how to balance the query accuracy and the size of fingerprint, and how to allocate the fingerprint bits to video and audio frames to maximize the query accuracy. By introducin…

Cited by 0SourceScholar