← Search

Xinyue Liu

21 accepted papers

2026

TAPE: Task-Adaptive Prototype Evolution in Audio-Language Models for Fully Few-shot Class-incremental Audio Classification

CVPR 2026

Fully Few-shot Class-incremental Audio Classification (FFCAC) is challenging since the training samples are limited both in the incremental sessions and in the base session. Existing few-shot learning methods suffer from catastrophic forgetting and overfitting when applied to FFCAC.Pre-trained Audio

Cited by 0SourcecodeScholar
2026

Think-Then-Generate: Structural Chain-of-Thought Reasoning for Consistent 3D Generation

CVPR 2026

Recently, generating 3D assets using visual priors from pretrained diffusion models has shown remarkable results. However, due to the inherent lack of 3D geometric priors in 2D diffusion, the synthesized results often suffer from spatial hallucination and multi-view inconsistency. To address this li

Cited by 0SourcecodeScholar
2026

Towards Fine-Grained Attribution: Instance-Aware Preference Optimization for Aligning Diffusion Models

CVPR 2026

Direct Preference Optimization has achieved remarkable success in aligning diffusion models with human feedback. However, existing methods heavily rely on image-level preferences, which suffer from sparse rewards in the spatial dimension. This creates a fundamental misalignment: while an image may b

Cited by 0SourceScholar
2025

Conditional Semantic Textual Similarity via Conditional Contrastive Learning

COLING 2025main

Conditional semantic textual similarity (C-STS) assesses the similarity between pairs of sentence representations under different conditions. The current method encounters the over-estimation issue of positive and negative samples. Specifically, the similarity within positive samples is excessively…

2025

Full Network Capacity Framework for Sample-Efficient Deep Reinforcement Learning

UAI 2025

In deep reinforcement learning (DRL), the presence of dormant neurons leads to a significant reduction in network capacity, which results in sub-optimal performance and limited sample efficiency. Existing training techniques, especially those relying on periodic resetting (PR), exacerbate this issue

2025

Hawkes based Representation Learning for Reasoning over Scale-free Community-structured Temporal Knowledge Graphs

COLING 2025main

Temporal knowledge graph (TKG) reasoning has become a hot topic due to its great value in many practical tasks. The key to TKG reasoning is modeling the structural information and evolutional patterns of the TKGs. While great efforts have been devoted to TKG reasoning, the structural and evolutional…

2025

Online Contrastive Continual Learning with Hard Negative Samples

ICASSP 2025accepted

Online continual learning (OCL) is a strict setting of continual learning (CL), where the OCL agent faces a never-ending data stream and encounters each new sample only once. An OCL agent suffers more serious catastrophic forgetting (i.e., forgetting previous knowledge of old classes) than a CL agen…

Cited by 0SourceScholar
2025

Text-Guided Fine-grained Counterfactual Inference for Short Video Fake News Detection

AAAI 2025technical

Detecting fake news in short videos is crucial for combating misinformation. Existing methods utilize topic modeling and co-attention mechanism, overlooking the modality heterogeneity and resulting in suboptimal performance. To address this issue, we introduce Text-Guided Fine-grained Counterfactual…

Cited by 0SourcePDFScholar
2025

Unlocking the Potential of Reverse Distillation for Anomaly Detection

AAAI 2025technical

Knowledge Distillation (KD) is a promising approach for unsupervised Anomaly Detection (AD). However, the student network's over-generalization often diminishes the crucial representation differences between teacher and student in anomalous regions, leading to detection failures. To address this pro…

2024

A Goal Interaction Graph Planning Framework for Conversational Recommendation

AAAI 2024technical

Multi-goal conversational recommender system (MG-CRS) which is more in line with realistic scenarios has attracted a lot of attention. MG-CRS can dynamically capture the demands of users in conversation, continuously engage their interests, and make recommendations. The key of accomplishing these ta…

2024

Improve Meta-learning for Few-Shot Text Classification with All You Can Acquire from the Tasks

EMNLP 2024finding

Meta-learning has emerged as a prominent technology for few-shot text classification and has achieved promising performance. However, existing methods often encounter difficulties in drawing accurate class prototypes from support set samples, primarily due to probable large intra-class differences a…

2024

RENN: A Rule Embedding Enhanced Neural Network Framework for Temporal Knowledge Graph Completion

COLING 2024main

Temporal knowledge graph completion is a critical task within the knowledge graph domain. Existing approaches encompass deep neural network-based methods for temporal knowledge graph embedding and rule-based logical symbolic reasoning. However, the former may not adequately account for structural de…

Cited by 2SourcePDFScholar
2024

Temporal Knowledge Graph Reasoning with Dynamic Hypergraph Embedding

COLING 2024main

Reasoning over the Temporal Knowledge Graph (TKG) that predicts facts in the future has received much attention. Most previous works attempt to model temporal dynamics with knowledge graphs and graph convolution networks. However, these methods lack the consideration of high-order interactions betwe…

Cited by 4SourcePDFScholar
2024

Unveiling Opinion Evolution via Prompting and Diffusion for Short Video Fake News Detection

ACL 2024findings

Short video fake news detection is crucial for combating the spread of misinformation. Current detection methods tend to aggregate features from individual modalities into multimodal features, overlooking the implicit opinions and the evolving nature of opinions across modalities. In this paper, we…

Cited by 3SourcePDFScholar
2024

Video-Context Aligned Transformer for Video Question Answering

AAAI 2024technical

Video question answering involves understanding video content to generate accurate answers to questions. Recent studies have successfully modeled video features and achieved diverse multimodal interaction, yielding impressive outcomes. However, they have overlooked the fact that the video contains r…

Cited by 3SourcePDFScholar
2023

Knowledge-Aware Graph Convolutional Network with Utterance-Specific Window Search for Emotion Recognition In Conversations

ICASSP 2023accepted

Emotion recognition in conversation (ERC) enables a deeper understanding of emotion for each utterance within a conversation. Recent progress on ERC has proved that using Graph Neural Networks (GNN) to model conversational context is effective for identifying emotions. However, existing GNN-based ap…

Cited by 0SourceScholar
2023

T2IAT: Measuring Valence and Stereotypical Biases in Text-to-Image Generation

ACL 2023findings

*Warning: This paper contains several contents that may be toxic, harmful, or offensive.*In the last few years, text-to-image generative models have gained remarkable success in generating images with unprecedented quality accompanied by a breakthrough of inference speed. Despite their rapid progres…

Cited by 32SourcePDFScholar
2022

Multimodal Context Carryover

EMNLP 2022industry

Multi-modality support has become an integral part of creating a seamless user experience with modern voice assistants with smart displays. Users refer to images, video thumbnails, or the accompanying text descriptions on the screen through voice communication with AI powered devices. This raises th…

Cited by 3SourcePDFScholar
2020

Multi-task Learning of Spoken Language Understanding by Integrating N-Best Hypotheses with Hierarchical Attention

COLING 2020industry

Currently, in spoken language understanding (SLU) systems, the automatic speech recognition (ASR) module produces multiple interpretations (or hypotheses) for the input audio signal and the natural language understanding (NLU) module takes the one with the highest confidence score for domain or inte…

Cited by 8SourcePDFScholar