← Search

Yuhui Zheng

11 accepted papers

2026

Learning Brain Representation with Hierarchical Visual Embeddings

ICLR 2026poster

Decoding visual representations from brain signals has attracted significant attention in both neuroscience and artificial intelligence. However, the degree to which brain signals truly encode visual information remains unclear. Current visual decoding approaches explore various brain–image alignmen…

Cited by 0SourceScholar
2026

Progressive Guessing to Fixed Point: Rethinking Human Motion Prediction with Deep Equilibrium Models

CVPR 2026

Many recent human motion prediction methods adopt a multi-stage refinement framework, where each stage produces an initial guess of future poses for the next stage. These guesses are progressively refined towards the target prediction through a sequence of spatial-temporal reasoning stages.However,

Cited by 0SourceScholar
2026

RAC-DMVC: Reliability-Aware Contrastive Deep Multi-View Clustering Under Multi-Source Noise

AAAI 2026technical

Multi-view clustering (MVC), which aims to separate the multi-view data into distinct clusters in an unsupervised manner, is a fundamental yet challenging task. To enhance its applicability in real-world scenarios, this paper addresses a more challenging task: MVC under multi-source noises, includin

Cited by 0SourcePDFScholar
2025

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions

EMNLP 2025

While densely annotated image captions significantly facilitate the learning of robust vision-language alignment, methodologies for systematically optimizing human annotation efforts remain underexplored. We introduce Chain-of-Talkers (CoTalk), an AI-in-the-loop methodology designed to maximize the

Cited by 0SourcePDFScholar
2024

A Transformer-Based Adaptive Prototype Matching Network for Few-Shot Semantic Segmentation

IJCAI 2024poster

Few-shot semantic segmentation (FSS) aims to generate a model for segmenting novel classes using a limited number of annotated samples. Previous FSS methods have shown sensitivity to background noise due to inherent bias, attention bias, and spatial-aware bias. In this study, we propose a Transforme…

Cited by 0SourcePDFScholar
2024

Contrastive Transformer Cross-Modal Hashing for Video-Text Retrieval

IJCAI 2024poster

As video-based social networks continue to grow exponentially, there is a rising interest in video retrieval using natural language. Cross-modal hashing, which learns compact hash code for encoding multi-modal data, has proven to be widely effective in large-scale cross-modal retrieval, e.g., image-…

Cited by 0SourcePDFScholar
2024

Contrastive Transformer Masked Image Hashing for Degraded Image Retrieval

IJCAI 2024poster

Hashing utilizes hash code as a compact image representation, offering excellent performance in large-scale image retrieval due to its computational and storage advantages. However, the prevalence of degraded images on social media platforms, resulting from imperfections in the image capture process…

Cited by 0SourcePDFScholar
2024

Distributed Manifold Hashing for Image Set Classification and Retrieval

AAAI 2024technical

Conventional image set methods typically learn from image sets stored in one location. However, in real-world applications, image sets are often distributed or collected across different positions. Learning from such distributed image sets presents a challenge that has not been studied thus far. Mor…

Cited by 1SourcePDFScholar
2024

Generalizable Fourier Augmentation for Unsupervised Video Object Segmentation

AAAI 2024technical

The performance of existing unsupervised video object segmentation methods typically suffers from severe performance degradation on test videos when tested in out-of-distribution scenarios. The primary reason is that the test data in real- world may not follow the independent and identically distrib…

Cited by 6SourcePDFScholar
2024

Glance, Focus and Refinement Network for Remote Sensing Change Detection

ICASSP 2024accepted

Existing change detection (CD) methods often directly fuse the multi-level features from bi-temporal remote sensing images without discriminatively considering each pixel's importance. Despite the demonstrated success, unselectively mixing the features degrades the model's performance to effectively…

Cited by 0SourceScholar