← Search

Yinwei Wei

6 accepted papers

2026

HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval

AAAI 2026technical

Composed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal query composed of a reference image and modification text. Although this task has demonstrated promising applications in personalized search and recomme

Cited by 0SourcePDFScholar
2026

INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

AAAI 2026technical

Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting of reference images and modification texts. Although substantial progress has been made in recent years, existing methods assume that all samples are

Cited by 0SourcePDFScholar
2026

VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

ICML 2026poster

In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-videos paradigms have emerged, adopting a localize–clip–answer pipeline in which th…

Cited by 2SourceScholar
2025

Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization

NeurIPS 2025poster

Direct Preference Optimization (DPO) has emerged as an effective approach for mitigating hallucination in Multimodal Large Language Models (MLLMs). Although existing methods have achieved significant progress by utilizing vision-oriented contrastive objectives for enhancing MLLMs' attention to visua…

Cited by 0SourceScholar
2025

Modality-Independent Graph Neural Networks with Global Transformers for Multimodal Recommendation

AAAI 2025technical

Multimodal recommendation systems can learn users' preferences from existing user-item interactions as well as the semantics of multimodal data associated with items. Many existing methods model this through a multimodal user-item graph, approaching multimodal recommendation as a graph learning task…

2024

Mrtnet: Multi-Resolution Temporal Network for Video Sentence Grounding

ICASSP 2024accepted

Video sentence grounding locates a specific moment in a video based on a text query. Existing methods focus on single temporal resolution, ignoring multi-scale temporal consistency. We introduce MRTNet, a multi-resolution grounding network with four key components: a feature encoder, a Multi-Resolut…

Cited by 0SourceScholar