← Search

Xueyi Li

8 accepted papers

2026

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

ICLR 2026poster

Recent advances in large language models (LLMs) have attracted significant interest in extending their capabilities to multimodal scenarios, particularly for speech-to-speech conversational systems. However, existing multimodal models handling interleaved audio and text rely on autoregressive (AR) m…

Cited by 0SourcecodeScholar
2026

Improving Scientific Formula Verbalization in Large Speech Language Models for Accessible Learning

IJCAI 2026

Online learning systems provide accessible learning opportunities for blind or low-vision students. To support access to complex scientific materials, the speech models used in these systems need to deliver accurate scientific formula verbalization. While recent large speech language models (LSLMs)

Cited by 0Scholar
2025

Cognitive Fluctuations Enhanced Attention Network for Knowledge Tracing

AAAI 2025technical

Knowledge tracing (KT) involves using the historical records of student-learning interactions to anticipate their performance on forthcoming questions. Central to this process is the modeling of human cognition to gain deeper insights into how knowledge is acquired and retained. Human cognition is c…

Cited by 0SourcePDFScholar
2025

Rethinking and Improving Student Learning and Forgetting Processes for Attention based Knowledge Tracing Models

AAAI 2025technical

Knowledge tracing (KT) models students' knowledge states and predicts their future performance based on their historical interaction data. However, attention based KT models struggle to accurately capture diverse forgetting behaviors in ever-growing interaction sequences. First, existing models us…

Cited by 0SourcePDFScholar
2024

Enhancing Length Generalization for Attention Based Knowledge Tracing Models with Linear Biases

IJCAI 2024poster

Knowledge tracing (KT) is the task of predicting students' future performance based on their historical learning interaction data. With the rapid advancement of attention mechanisms, many attention based KT models are developed. However, existing attention based KT models exhibit performance drops a…

Cited by 8SourcePDFScholar
2021

Group-Wise Semantic Mining for Weakly Supervised Semantic Segmentation

AAAI 2021technical

Acquiring sufficient ground-truth supervision to train deep vi- sual models has been a bottleneck over the years due to the data-hungry nature of deep learning. This is exacerbated in some structured prediction tasks, such as semantic segmen- tation, which requires pixel-level annotations. This work…

2021

Local Representation is Not Enough: Soft Point-Wise Transformer for Descriptor and Detector of Local Features

IJCAI 2021poster

Significant progress has been witnessed for the descriptor and detector of local features, but there still exist several challenging and intractable limitations, such as insufficient localization accuracy and non-discriminative description, especially in repetitive- or blank-texture regions, which h…

Cited by 12SourcePDFScholar
2021

Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object Segmentation

CVPR 2021poster

This paper addresses the task of unsupervised video multi-object segmentation. Current approaches follow a two-stage paradigm: 1) detect object proposals using pre-trained Mask R-CNN, and 2) conduct generic feature matching for temporal association using re-identification techniques. However, the ge…

Cited by 59PDFScholar