← Search

Lin Yuanbo Wu

10 accepted papers

2026

Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective

AAAI 2026technical

Multimodal Misinformation Detection (MMD) refers to the task of detecting social media posts involving misinformation, where the post often contains text and image modalities. However, by observing the MMD posts, we hold that the text modality may be much more informative than the image modality bec

Cited by 0SourcePDFScholar
2026

Geometry-Aware Noisy Correspondence Mitigation for Cross-Modal Text-Based Person Retrieval

AAAI 2026technical

Text-Based Person Retrieval (TBPR) aims to accurately retrieve target individuals from large-scale image databases using only textual descriptions. Existing methods typically assume a ground-truth correspondence between text and images (i.e., strongly correlated). However, in real-world scenarios, t

Cited by 0SourcePDFScholar
2025

Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding

CVPR 2025highlight

Large Vision-Language Models (LVLMs) have obtained impressive performance in visual content understanding and multi-modal reasoning. Unfortunately, these large models suffer from serious hallucination problems and tend to generate fabricated responses. Recently, several Contrastive Decoding (CD) str…

2025

Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models

ICCV 2025poster

Although Large Vision-Language Models (LVLMs) have achieved impressive results, their high computational costs pose a significant barrier to wide application. To enhance inference efficiency, most existing approaches can be categorized as parameter-dependent or token-dependent strategies to reduce c…

2025

Towards Robust Category-level Articulation Pose Estimation via Integrated Differentiable Rendering

ICASSP 2025accepted

Accurate object pose estimation is crucial for embodied intelligence tasks such as manipulation, grasping, and human-robot interaction. However, due to the inherent characteristics of articulated objects, such as kinematic constraints and self-occlusion, pose estimation for articulated objects has r…

Cited by 1SourceScholar
2025

Unlocking Generalization Power in LiDAR Point Cloud Registration

CVPR 2025highlight

In real-world environments, a LiDAR point cloud registration method with robust generalization capabilities (across varying distances and datasets) is crucial for ensuring safety in autonomous driving and other LiDAR-based applications. However, current methods fall short in achieving this level of…

2025

Utterance-level Emotion Recognition in Conversation with Conversation-level Supervision

AAAI 2025technical

Emotion Recognition in Conversations (ERC) involves automatically identifying the emotion of each utterance in conversations. The emotion of an utterance is contingent to the conversation context, and thus, annotating each utterance in ERC entails repetitive screening the whole conversation from ann…

Cited by 0SourcePDFScholar
2024

De novo Protein Design Using Geometric Vector Field Networks

ICLR 2024spotlight

Advances like protein diffusion have marked revolutionary progress in $\textit{de novo}$ protein design, a central topic in life science. These methods typically depend on protein structure encoders to model residue backbone frames, where atoms do not exist. Most prior encoders rely on atom-wise fea…

2024

EfficientCAPER: An End-to-End Framework for Fast and Robust Category-Level Articulated Object Pose Estimation

NeurIPS 2024poster

Human life is populated with articulated objects. Pose estimation for category-level articulated objects is a significant challenge due to their inherent complexity and diverse kinematic structures. Current methods for this task usually meet the problems of insufficient consideration of kinematic co…

Cited by 0SourcePDFScholar
2023

CTVIS: Consistent Training for Online Video Instance Segmentation

ICCV 2023poster

The discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/nega…

Cited by 46PDFcodeScholar