← Search

Xuemeng Song

17 accepted papers

2026

D2MoRA: Diversity-Regulated Asymmetric MoE-LoRA Decomposition for Efficient Multi-Task Adaptation

AAAI 2026technical

Low-Rank Adaptation (LoRA) has emerged as a powerful parameter-efficient fine-tuning method for adapting large language models to downstream tasks. Recent studies have leveraged Mixture-of-Experts (MoE) mechanism to effectively integrate multiple LoRA modules, facilitating efficient parameter adapta

Cited by 0SourcePDFScholar
2026

MELT: IMPROVE COMPOSED IMAGE RETRIEVAL VIA THE MODIFICATION FREQUENTATION-RARITY BALANCE NETWORK

ICASSP 2026poster

Composed Image Retrieval (CIR) uses a reference image and a modification text as a query to retrieve a target image satisfying the requirement of ``modifying the reference image according to the text instructions''. However, existing CIR methods face two limitations: (1) frequency bias leading to ``…

Cited by 0SourcePDFScholar
2026

TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs

AAAI 2026technical

Video large language models have achieved remarkable performance in tasks such as video question answering, however, their temporal understanding remains suboptimal. To address this limitation, we curate a dedicated instruction fine-tuning dataset that focuses on enhancing temporal comprehension acr

Cited by 0SourcePDFScholar
2026

VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

ICML 2026poster

In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-videos paradigms have emerged, adopting a localize–clip–answer pipeline in which th…

Cited by 2SourceScholar
2025

Content-aware Balanced Spectrum Encoding in Masked Modeling for Time Series Classification

AAAI 2025technical

Due to the superior ability of global dependency, transformer and its variants have become the primary choice in Masked Time-series Modeling (MTM) towards time-series classification task. In this paper, we experimentally analyze that existing transformer-based MTM methods encounter with two under-ex…

Cited by 0SourcePDFScholar
2025

Language-Assisted Debiasing and Smoothing for Foundation Model-Based Semi-Supervised Learning

CVPR 2025poster

Recent studies have focused on introducing pre-trained foundation models into semi-supervised learning (SSL) tasks. Nevertheless, these foundation models can exhibit biases toward different classes and tend to generate imbalanced pseudo-labels for SSL. Thus, efforts have been made to introduce the l…

2025

MEDIAN: Adaptive Intermediate-grained Aggregation Network for Composed Image Retrieval

ICASSP 2025accepted

The Composed Image Retrieval (CIR) task aims to retrieve a target image that meets the requirements based on a given multimodal query (includes a reference image and modification text). Most existing works align multimodal semantics at both local and global granularity. However, they have failed to…

Cited by 0SourceScholar
2025

Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization

NeurIPS 2025poster

Direct Preference Optimization (DPO) has emerged as an effective approach for mitigating hallucination in Multimodal Large Language Models (MLLMs). Although existing methods have achieved significant progress by utilizing vision-oriented contrastive objectives for enhancing MLLMs' attention to visua…

Cited by 0SourceScholar
2025

PAIR: Complementarity-guided Disentanglement for Composed Image Retrieval

ICASSP 2025accepted

Composed Image Retrieval (CIR) is a novel image retrieval paradigm that aims at searching for the target images via the multimodal query including a reference image and a modification text. Although existing works have made significant progress, they overlook the inter-modal coherence and incoherenc…

Cited by 0SourceScholar
2024

Exploiting the Social-Like Prior in Transformer for Visual Reasoning

AAAI 2024technical

Benefiting from instrumental global dependency modeling of self-attention (SA), transformer-based approaches have become the pivotal choices for numerous downstream visual reasoning tasks, such as visual question answering (VQA) and referring expression comprehension (REC). However, some studies hav…

Cited by 4SourcePDFScholar
2024

Multi-Factor Adaptive Vision Selection for Egocentric Video Question Answering

ICML 2024poster

The challenge of interpreting the world from a human perspective in Artificial Intelligence (AI) is particularly evident in egocentric video question answering, which grapples with issues like small object recognition, noise suppression, and spatial-temporal reasoning. To address these challenges, w…

2024

VK-G2T: Vision and Context Knowledge Enhanced Gloss2text

ICASSP 2024accepted

Existing sign language translation methods follow a two-stage pipeline: first converting the sign language video to a gloss sequence (i.e., Sign2Gloss) and then translating the generated gloss sequence into a spoken language sentence (i.e., Gloss2Text). While previous studies have focused on boostin…

Cited by 0SourceScholar
2023

Multi-source Semantic Graph-based Multimodal Sarcasm Explanation Generation

ACL 2023long

Multimodal Sarcasm Explanation (MuSE) is a new yet challenging task, which aims to generate a natural language sentence for a multimodal social post (an image as well as its caption) to explain why it contains sarcasm. Although the existing pioneer study has achieved great success with the BART back…

2023

Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm Detection

AAAI 2023technical

Sarcasm is a sophisticated linguistic phenomenon that is prevalent on today's social media platforms. Multi-modal sarcasm detection aims to identify whether a given sample with multi-modal information (i.e., text and image) is sarcastic. This task's key lies in capturing both inter- and intra-modal…

2022

MERIt: Meta-Path Guided Contrastive Learning for Logical Reasoning

ACL 2022findings

Logical reasoning is of vital importance to natural language understanding. Previous studies either employ graph-based models to incorporate prior knowledge about logical relations, or introduce symbolic logic into neural models through data augmentation. These methods, however, heavily depend on an…

2022

Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph Generation

CVPR 2022poster

Scene Graph Generation, which generally follows a regular encoder-decoder pipeline, aims to first encode the visual contents within the given image and then parse them into a compact summary graph. Existing SGG approaches generally not only neglect the insufficient modality fusion between vision and…

Cited by 127PDFcodeScholar
2020

Auxiliary Template-Enhanced Generative Compatibility Modeling

IJCAI 2020poster

In recent years, there has been a growing interest in the fashion analysis (e.g., clothing matching) due to the huge economic value of the fashion industry. The essential problem is to model the compatibility between the complementary fashion items, such as the top and bottom in clothing matching. T…

Cited by 0SourcePDFScholar