← Search

Yupeng Hu

21 accepted papers

2026

Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval

CVPR 2026

Composed Image Retrieval (CIR) has attracted significant attention due to its flexible multimodal query method, yet its development is severely constrained by the Noisy Triplet Correspondence (NTC) problem. Most existing robust learning methods rely on the "small loss hypothesis", but the unique sem

Cited by 0SourcecodeScholar
2026

Circuit-Think: A Multimodal Reasoning Framework for Automated Circuit-to-Netlist Translation with Trajectory-Guided Reinforcement Learning

AAAI 2026technical

Vision Language Models (VLMs) have shown strong performance in multimodal understanding, offering promise for the circuit-to-netlist translation task. However, the diverse component symbols and complex connections in circuit images challenge VLMs in understanding physical layouts and reasoning for e

Cited by 0SourcePDFScholar
2026

ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval

CVPR 2026

The Composed Image Retrieval (CIR) task provides a flexible retrieval paradigm via a reference image and modification text, but it heavily relies on expensive and error-prone triplet annotations. This paper systematically investigates the Noisy Triplet Correspondence (NTC) problem introduced by anno

Cited by 0SourcecodeScholar
2026

D2MoRA: Diversity-Regulated Asymmetric MoE-LoRA Decomposition for Efficient Multi-Task Adaptation

AAAI 2026technical

Low-Rank Adaptation (LoRA) has emerged as a powerful parameter-efficient fine-tuning method for adapting large language models to downstream tasks. Recent studies have leveraged Mixture-of-Experts (MoE) mechanism to effectively integrate multiple LoRA modules, facilitating efficient parameter adapta

Cited by 0SourcePDFScholar
2026

Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action Localization

AAAI 2026technical

Current Zero-Shot Temporal Action Localization (ZSTAL) methods, whether training-based or training-free ones, still predominantly rely on a single, unified query to localize an entire action. This unified representation is fundamentally ill-suited for complex real-world activities, as it fails to ca

Cited by 0SourcePDFScholar
2026

HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval

AAAI 2026technical

Composed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal query composed of a reference image and modification text. Although this task has demonstrated promising applications in personalized search and recomme

Cited by 0SourcePDFScholar
2026

INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

AAAI 2026technical

Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting of reference images and modification texts. Although substantial progress has been made in recent years, existing methods assume that all samples are

Cited by 0SourcePDFScholar
2026

MELT: IMPROVE COMPOSED IMAGE RETRIEVAL VIA THE MODIFICATION FREQUENTATION-RARITY BALANCE NETWORK

ICASSP 2026poster

Composed Image Retrieval (CIR) uses a reference image and a modification text as a query to retrieve a target image satisfying the requirement of ``modifying the reference image according to the text instructions''. However, existing CIR methods face two limitations: (1) frequency bias leading to ``…

Cited by 0SourcePDFScholar
2026

ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval

AAAI 2026technical

With the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing attention from researchers. Unlike unimodal video retrieval methods, the CVR task takes a multi-modal query consisting of a reference video and a piece o

Cited by 0SourcePDFScholar
2026

When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?

AAAI 2026technical

Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, which simulates an “Audio-Visual Confusion” scene by modifying the corresponding sound of an object in the video, e.g., mute

Cited by 0SourcePDFScholar
2025

Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection

ICCV 2025poster

Open vocabulary Human-Object Interaction (HOI) detection is a challenging task that detects all <human, verb, object> triplets of interest in an image, even those that are not pre-defined in the training set. Existing approaches typically rely on output features generated by large Vision-Language Mo…

2025

CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG

ACL 2025long

Multimodal Retrieval-Augmented Generation (MMRAG) has been introduced to enhance Multimodal Large Language Models by incorporating externally retrieved multimodal knowledge, but it introduces two challenges: Parametric-Retrieved Knowledge Inconsistency (PRKI), where discrepancies between parametric…

2025

Content-aware Balanced Spectrum Encoding in Masked Modeling for Time Series Classification

AAAI 2025technical

Due to the superior ability of global dependency, transformer and its variants have become the primary choice in Masked Time-series Modeling (MTM) towards time-series classification task. In this paper, we experimentally analyze that existing transformer-based MTM methods encounter with two under-ex…

Cited by 0SourcePDFScholar
2025

CurMIM: Curriculum Masked Image Modeling

ICASSP 2025accepted

Masked Image Modeling (MIM), following “mask-andreconstruct” scheme, is a promising self-supervised method to learn scalable visual representation. Studies indicate that selecting an effective mask strategy is vital for MIM. However, existing approaches often rely on static pre-defined priors, which…

Cited by 0SourceScholar
2025

ENCODER: Entity Mining and Modification Relation Binding for Composed Image Retrieval

AAAI 2025technical

The objective of Composed Image Retrieval (CIR) is to identify a target image that meets the requirement based on a multimodal query (including the reference image and the modification text) provided by the user. Despite the notable success of existing approaches, they fail to adequately address the…

Cited by 2SourcePDFScholar
2025

MEDIAN: Adaptive Intermediate-grained Aggregation Network for Composed Image Retrieval

ICASSP 2025accepted

The Composed Image Retrieval (CIR) task aims to retrieve a target image that meets the requirements based on a given multimodal query (includes a reference image and modification text). Most existing works align multimodal semantics at both local and global granularity. However, they have failed to…

Cited by 0SourceScholar
2025

PAIR: Complementarity-guided Disentanglement for Composed Image Retrieval

ICASSP 2025accepted

Composed Image Retrieval (CIR) is a novel image retrieval paradigm that aims at searching for the target images via the multimodal query including a reference image and a modification text. Although existing works have made significant progress, they overlook the inter-modal coherence and incoherenc…

Cited by 0SourceScholar
2025

Towards Stable and Storage-efficient Dataset Distillation: Matching Convexified Trajectory

CVPR 2025poster

The rapid evolution of deep learning and large language models has led to an exponential growth in the demand for training data, prompting the development of Dataset Distillation methods to address the challenges of managing large datasets. Among these, Matching Training Trajectories (MTT) has been…

Cited by 2SourcePDFScholar
2025

Unveiling the Pruning Risks on Privacy Vulnerabilities of Deep Neural Networks

ICASSP 2025accepted

Large-scale deep neural networks (DNNs), such as large language models, have gained immense popularity due to their outstanding performance across various tasks. However, their application in resource-constrained scenarios faces significant challenges due to the high computational costs and memory u…

Cited by 0SourceScholar
2024

Breaking Barriers of System Heterogeneity: Straggler-Tolerant Multimodal Federated Learning via Knowledge Distillation

IJCAI 2024poster

Internet of Things (IoT) devices possess valuable yet private multimodal data, calling for a decentralized machine learning scheme. Though several multimodal federated learning (MFL) methods have been proposed, most of them merely overlook the system heterogeneity across IoT devices, resulting in th…

Cited by 2SourcePDFScholar
2024

Exploiting the Social-Like Prior in Transformer for Visual Reasoning

AAAI 2024technical

Benefiting from instrumental global dependency modeling of self-attention (SA), transformer-based approaches have become the pivotal choices for numerous downstream visual reasoning tasks, such as visual question answering (VQA) and referring expression comprehension (REC). However, some studies hav…

Cited by 4SourcePDFScholar