← Search

Xiaoyong Wei

8 accepted papers

2026

Compositional Transformation Reasoning for Composed Video Retrieval

CVPR 2026

Composed Video Retrieval aims to retrieve a target video given a reference video and a textual modification describing the desired change. The core challenge lies in modeling compositional multimodal transformations, i.e., how entities, actions, and scenes evolve across video and language modalities

Cited by 0SourcecodeScholar
2025

Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond

EMNLP 2025

The increasing demand for domain-specific evaluation of large language models (LLMs) has led to the development of numerous benchmarks. These efforts often adhere to the principle of data scaling, relying on large corpora or extensive question-answer (QA) sets to ensure broad coverage. However, the

2025

GLProtein: Global-and-Local Structure Aware Protein Representation Learning

EMNLP 2025

Proteins are central to biological systems, participating as building blocks across all forms of life. Despite advancements in understanding protein functions through protein sequence analysis, there remains potential for further exploration in integrating protein structural information. We argue th

Cited by 0SourcePDFScholar
2025

Removal of Hallucination on Hallucination: Debate-Augmented RAG

ACL 2025long

Retrieval-Augmented Generation (RAG) enhances factual accuracy by integrating external knowledge, yet it introduces a critical issue: erroneous or biased retrieval can mislead generation, compounding hallucinations, a phenomenon we term Hallucination on Hallucination. To address this, we propose Deb…

2025

Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues

CVPR 2025poster

Understanding human behavior and the environmental information in the egocentric video is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-person view. Leveraging the corresponding exocentric video to provide global context has…

2024

Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for Dialogue

ACL 2024long

Tuning language models for dialogue generation has been a prevalent paradigm for building capable dialogue agents. Yet, traditional tuning narrowly views dialogue generation as resembling other language generation tasks, ignoring the role disparities between two speakers and the multi-round interact…

2023

Rethinking Multimodal Entity and Relation Extraction from a Translation Point of View

ACL 2023long

We revisit the multimodal entity and relation extraction from a translation point of view. Special attention is paid on the misalignment issue in text-image datasets which may mislead the learning. We are motivated by the fact that the cross-modal misalignment is a similar problem of cross-lingual d…

2022

M5Product: Self-Harmonized Contrastive Learning for E-Commercial Multi-Modal Pretraining

CVPR 2022poster

Despite the potential of multi-modal pre-training to learn highly discriminative feature representations from complementary data modalities, current progress is being slowed by the lack of large-scale modality-diverse datasets. By leveraging the natural suitability of E-commerce, where different mod…

Cited by 44PDFcodeScholar