← Search

Xiaojie Wang

55 accepted papers

2026

Accelerating Langevin Monte Carlo via Efficient Stochastic Runge-Kutta Methods beyond Log-Concavity

ICML 2026poster

Sampling from a high-dimensional probability distribution is a fundamental algorithmic task arising in wide-ranging applications across multiple disciplines, including scientific computing, computational statistics and machine learning. Langevin Monte Carlo (LMC) algorithms are among the most widely…

Cited by 0SourceScholar
2026

An Information-Theoretic Parameter-Free Bayesian Framework for Probing Labeled Dependency Trees from Attention Score

ICLR 2026poster

Figuring out how neural language models comprehend syntax acts as a key to revealing how they understand languages. We systematically analyzed methods of extracting syntax from models, namely _probing_, and found limitations yet widely exist in previous probing practice. We proposed a method capab…

Cited by 0SourcecodeScholar
2026

ChainGPT: Dual-Reasoning Model with Recurrent Depth and Multi-Rank State Updates

ICLR 2026poster

Large language models, constrained by the fixed-depth Transformer architecture, struggle to solve complex reasoning tasks in an end-to-end manner. Existing approaches, such as Chain of Thought, improve reasoning depth to some extent but rely heavily on natural language generation, with computational…

Cited by 0SourceScholar
2026

CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics Priority

CVPR 2026

Cross-modal alignment aims to learn semantically consistent latent representations across diverse modalities. Prevailing methods rely on a text-guided aggregation paradigm to achieve fine-grained alignment, while they suffer from redundant patch-word correlations and high computational costs. To add

Cited by 0SourceScholar
2026

EntroKV: Entropy-Guided Dynamic Budget Allocation for KV-Cache Compression

ICML 2026spotlight

The prohibitive memory footprint of the Key-Value (KV) cache imposes a critical bottleneck for efficient long-context LLM serving. Current compression techniques typically rely on static or uniform budget allocation, overlooking the significant heterogeneity in information density across attention h…

Cited by 0SourceScholar
2026

Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models

ICLR 2026poster

Vision-Language Models (VLMs) have demonstrated remarkable performance across a variety of real-world tasks. However, existing VLMs typically process visual information by serializing images, a method that diverges significantly from the parallel nature of human vision. Moreover, their opaque intern…

Cited by 0SourcecodeScholar
2026

Self-Refining Vision Language Model for Robotic Failure Detection and Reasoning

ICLR 2026poster

Reasoning about failures is crucial for building reliable and trustworthy robotic systems. Prior approaches either treat failure reasoning as a closed-set classification problem or assume access to ample human annotations. Failures in the real world are typically subtle, combinatorial, and difficult…

Cited by 0SourceScholar
2026

Semantic Impact–Driven Visual Scheduling in Vision-Language Models

ICML 2026poster

Vision-Language Models (VLMs) suffer from high inference latency due to long visual sequences. To enable efficient, on-demand utilization of visual information, we argue that visual necessity should be assessed by its semantic impact on the output distribution, rather than inferred from intermediate…

Cited by 0SourceScholar
2026

Small Object, Great Challenge: A Benchmark for Small Object Visual Grounding

CVPR 2026

The task of visual grounding (i.e., VG) aims to locate or segment objects in images based on referring expressions. Existing research on VG primarily focuses on large objects. However, these images often contain objects at various scales. Although large objects are usually the visual focus, small ob

Cited by 0SourcecodeScholar
2025

A Non-Contrastive Learning Framework for Sequential Recommendation with Preference-Preserving Profile Generation

ICLR 2025poster

Contrastive Learning (CL) proves to be effective for learning generalizable user representations in Sequential Recommendation (SR), but it suffers from high computational costs due to its reliance on negative samples. To overcome this limitation, we propose the first Non-Contrastive Learning (NCL) f…

Cited by 0SourcePDFScholar
2025

A Systematic Exploration of Knowledge Graph Alignment with Large Language Models in Retrieval Augmented Generation

AAAI 2025technical

Retrieval Augmented Generation (RAG) with Knowledge Graphs (KGs) is an effective way to enhance Large Language Models (LLMs). Due to the natural discrepancy between structured KGs and sequential LLMs, KGs must be linearized to text before being inputted into LLMs, leading to the problem of KG Alignm…

2025

A Weighted Cross-entropy Loss for Mitigating LLM Hallucinations in Cross-lingual Continual Pretraining

ICASSP 2025accepted

Recently, due to the explosive advances of large language models (LLMs) on English, cross-lingual continual pretraining has been widely applied in obtaining Chinese LLMs. However, previous studies showed that these LLMs have suffered severe hallucinations, mainly caused by noisy tokens. To this aim,…

Cited by 0SourceScholar
2025

CoTD-PO: Chain-of-Thought Distillation with Preference Optimization

EMNLP 2025

Chain-of-Thought (CoT) distillation has emerged as a promising paradigm to enhance the reasoning ability of small language models by imitating the reasoning and outputs of larger teacher models. However, existing approaches suffer from a critical limitation: a distribution mismatch between teacher-g

2025

Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents

EMNLP 2025

Large Language Models (LLMs) based agent systems have made great strides in real-world applications beyond traditional NLP tasks. This paper proposes a new LLM-based Multi-Agent System (LLM-MAS) benchmark, Collab-Overcooked, built on the popular Overcooked-AI game with more applicable and challengin

2025

Concept Conductor: Orchestrating Multiple Personalized Concepts in Text-to-Image Synthesis

AAAI 2025technical

The customization of text-to-image models has seen significant advancements, yet generating multiple personalized concepts remains a challenging task. Current methods struggle with attribute leakage and layout confusion when handling multiple concepts, leading to reduced concept fidelity and semanti…

2025

Controlled Low-Rank Adaptation with Subspace Regularization for Continued Training on Large Language Models

ACL 2025long

Large language models (LLMs) exhibit remarkable capabilities in natural language processing but face catastrophic forgetting when learning new tasks, where adaptation to a new domain leads to a substantial decline in performance on previous tasks. In this paper, we propose Controlled LoRA (CLoRA), a…

2025

Data with High and Consistent Preference Difference Are Better for Reward Model

AAAI 2025technical

Reinforcement Learning from Human Feedback (RLHF) is a commonly used alignment method for Large Language Models (LLMs). This method relies on a reward model trained on a preference dataset to provide scalar rewards. However, the human-annotated preference data is often sparse, noisy, and costly to o…

2025

Enhancing Tactile Sensing in Robotics Using Null-Space Diffusion Model with EIT-based Sensors

IROS 2025

Robotic tactile sensors based on Electrical Impedance Tomography (EIT) have gained great attention in robotic sensing applications due to their features such as no internal wiring, "all-in-one" structure, and continuous sensing capabilities. However, the effectiveness of EIT-based tactile sensors is

Cited by 0SourceScholar
2025

Multimodal Aspect-Based Sentiment Analysis under Conditional Relation

COLING 2025main

Multimodal Aspect-Based Sentiment Analysis (MABSA) aims to extract aspect terms from text-image pairs and identify their sentiments. Previous methods are based on the premise that the image contains the objects referred by the aspects within the text. However, this condition cannot always be met, re…

2025

Non-asymptotic Error Bounds in $\mathcal{W}_2$-Distance with Sqrt(d) Dimension Dependence and First Order Convergence for Langevin Monte Carlo beyond Log-Concavity

ICML 2025poster

Generating samples from a high dimensional probability distribution is a fundamental task with wide-ranging applications in the area of scientific computing, statistics and machine learning. This article revisits the popular Langevin Monte Carlo (LMC) sampling algorithms and provides a non-asymptoti…

Cited by 0SourcePDFScholar
2024

Correcting Non-Uniform Sensitivity in EIT Tactile Sensing via Jacobian Vector Approximation

RA-L 2024

Electrical impedance tomography (EIT)-based tactile sensors enable promising capabilities for safe human-robot interaction through large-area distributed force sensing. However, their practical realization is hampered by non-uniform sensitivity distribution which varies at different locations. This

Cited by 9SourceScholar
2024

Dual-Stage Multi-Task Syntax-Oriented Pre-Training for Syntactically Controlled Paraphrase Generation

ACL 2024findings

Syntactically Controlled Paraphrase Generation (SCPG), which aims at generating sentences having syntactic structures resembling given exemplars, is attracting more research efforts in recent years. We took an empirical survey on previous SCPG datasets and methods and found three tacitly approved wh…

2024

Enhancing Tactile Sensing in Robotics: Dual-Modal Force and Shape Perception with EIT-based Sensors and MM-CNN

ICRA 2024poster

Electrical Impedance Tomography (EIT)-based tactile sensors offer durability, scalability, and cost-effective manufacturing. However, simultaneously reconstructing force and shape from boundary measurements remains challenging due to EIT’s inherent location dependencies and image artifacts. This stu…

Cited by 2SourceScholar
2024

KG-Adapter: Enabling Knowledge Graph Integration in Large Language Models through Parameter-Efficient Fine-Tuning

ACL 2024findings

Although large language models (LLMs) show remarkable capabilities and generalizability across various tasks, they are criticized for lack of expertise. One promising solution is to combine knowledge graphs (KGs) with LLMs, and recent studies focus on integrating KGs into LLMs through prompt-based m…

2024

Phased Instruction Fine-Tuning for Large Language Models

ACL 2024findings

Instruction Fine-Tuning, a method enhancing pre-trained language models’ capabilities from mere next-word prediction to complex instruction following, often employs a one-off training approach on diverse instruction dataset. However, this method may not effectively enhance models’ adherence to instr…

2024

Pseudo-Domain Adversarial Networks with Electrical Impedance Tomography for Electrode Offset Error

IROS 2024poster

This paper propose a novel transfer learning approach, Pseudo-Domain Adversarial Network (PDAN), to tackle the issue of electrode displacement in Electrical Impedance Tomography (EIT). Electrode displacement, caused by human movement or improper operation, significantly affects the accuracy of EIT b…

Cited by 0SourceScholar
2024

Visual Prompt Tuning for Weakly Supervised Phrase Grounding

ICASSP 2024accepted

Previous works on the task of weakly supervised phrase grounding (WSG) rely heavily on object detectors providing RoIs for the localization. However, such methods cannot be applied effectively to real-world scenarios largely because that the detectors are trained with limited categories. In this pap…

Cited by 0SourceScholar
2023

An Asynchronous Updating Reinforcement Learning Framework for Task-Oriented Dialog System

ICASSP 2023accepted

Reinforcement learning has been applied to train the dialog systems in many works. Previous approaches divide the dialog system into multiple modules including DST (dialog state tracking) and DP (dialog policy), and train these modules simultaneously. However, different modules influence each other…

Cited by 0SourceScholar
2023

Explicit Alignment and Many-to-many Entailment Based Reasoning for Conversational Machine Reading

EMNLP 2023long findings

Conversational Machine Reading (CMR) requires answering a user's initial question through multi-turn dialogue interactions based on a given document. Although there exist many effective methods, they largely neglected the alignment between the $\textit{document}$ and the $\textit{user-provided infor…

Cited by 0SourceScholar
2023

Multimodal Recommendation Dialog with Subjective Preference: A New Challenge and Benchmark

ACL 2023findings

Existing multimodal task-oriented dialog data fails to demonstrate the diverse expressions of user subjective preferences and recommendation acts in the real-life shopping scenario. This paper introduces a new dataset SURE (Multimodal Recommendation Dialog with Subjective Preference), which contains…

2023

SPRING: Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph

AAAI 2023technical

Existing multimodal conversation agents have shown impressive abilities to locate absolute positions or retrieve attributes in simple scenarios, but they fail to perform well when complex relative positions and information alignments are involved, which poses a bottleneck in response quality. In thi…

2023

USSA: A Unified Table Filling Scheme for Structured Sentiment Analysis

ACL 2023long

Most previous studies on Structured Sentiment Analysis (SSA) have cast it as a problem of bi-lexical dependency parsing, which cannot address issues of overlap and discontinuity simultaneously. In this paper, we propose a niche-targeting and effective solution. Our approach involves creating a novel…

Cited by 9SourcePDFScholar
2022

A Simple Model for Distantly Supervised Relation Extraction

COLING 2022main

Distantly supervised relation extraction is challenging due to the noise within data. Recent methods focus on exploiting bag representations based on deep neural networks with complex de-noising scheme to achieve remarkable performance. In this paper, we propose a simple but effective BERT-based Gra…

2022

A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots

ACL 2022findings

A slot value might be provided segment by segment over multiple-turn interactions in a dialog, especially for some important information such as phone numbers and names. It is a common phenomenon in daily life, but little attention has been paid to it in previous work. To fill the gap, this paper de…

2022

COM-MRC: A COntext-Masked Machine Reading Comprehension Framework for Aspect Sentiment Triplet Extraction

EMNLP 2022main

Aspect Sentiment Triplet Extraction (ASTE) aims to extract sentiment triplets from sentences, which was recently formalized as an effective machine reading comprehension (MRC) based framework. However, when facing multiple aspect terms, the MRC-based methods could fail due to the interference from o…

2022

Co-VQA : Answering by Interactive Sub Question Sequence

ACL 2022findings

Most existing approaches to Visual Question Answering (VQA) answer questions directly, however, people usually decompose a complex question into a sequence of simple sub questions and finally obtain the answer to the original question after answering the sub question sequence(SQS). By simulating the…

Cited by 22SourcePDFScholar
2022

Enhanced Multi-Channel Graph Convolutional Network for Aspect Sentiment Triplet Extraction

ACL 2022long

Aspect Sentiment Triplet Extraction (ASTE) is an emerging sentiment analysis task. Most of the existing studies focus on devising a new tagging scheme that enables the model to extract the sentiment triplets in an end-to-end fashion. However, these methods ignore the relations between words for ASTE…

2022

Learn to Adapt for Generalized Zero-Shot Text Classification

ACL 2022long

Generalized zero-shot text classification aims to classify textual instances from both previously seen classes and incrementally emerging unseen classes. Most existing methods generalize poorly since the learned parameters are only optimal for seen classes rather than for both classes, and the param…

2022

Towards Unifying Reference Expression Generation and Comprehension

EMNLP 2022main

Reference Expression Generation (REG) and Comprehension (REC) are two highly correlated tasks. Modeling REG and REC simultaneously for utilizing the relation between them is a promising way to improve both. However, the problem of distinct inputs, as well as building connections between them in a si…

2021

Converse, Focus and Guess – Towards Multi-Document Driven Dialogue

AAAI 2021technical

We propose a novel task, Multi-Document Driven Dialogue (MD3), in which an agent can guess the target document that the user is interested in by leading a dialogue. To benchmark progress, we introduce a new dataset of GuessMovie, which contains 16,881 documents, each describing a movie, and associat…

2021

DialogueTRM: Exploring Multi-Modal Emotional Dynamics in a Conversation

EMNLP 2021finding

Emotion dynamics formulates principles explaining the emotional fluctuation during conversations. Recent studies explore the emotion dynamics from the self and inter-personal dependencies, however, ignoring the temporal and spatial dependencies in the situation of multi-modal conversations. To addre…

2021

Dual Graph Convolutional Networks for Aspect-based Sentiment Analysis

ACL 2021long

Aspect-based sentiment analysis is a fine-grained sentiment classification task. Recently, graph neural networks over dependency trees have been explored to explicitly model connections between aspects and opinion words. However, the improvement is limited due to the inaccuracy of the dependency par…

2021

Enhancing Visual Dialog Questioner with Entity-based Strategy Learning and Augmented Guesser

EMNLP 2021finding

Considering the importance of building a good Visual Dialog (VD) Questioner, many researchers study the topic under a Q-Bot-A-Bot image-guessing game setting, where the Questioner needs to raise a series of questions to collect information of an undisclosed image. Despite progress has been made in S…

2021

Grouped-Attention for Content-Selection and Content-Plan Generation

EMNLP 2021finding

Content-planning is an essential part of data-to-text generation to determine the order of data mentioned in generated texts. Recent neural data-to-text generation models employ Pointer Networks to explicitly learn content-plan given a set of attributes as input. They use LSTM to encode the input, w…

Cited by 1SourcePDFScholar
2021

MIEHDR CNN: Main Image Enhancement based Ghost-Free High Dynamic Range Imaging using Dual-Lens Systems

AAAI 2021technical

We study the High Dynamic Range (HDR) imaging problem using two Low Dynamic Range (LDR) images that are shot from dual-lens systems in a single shot time with different exposures. In most of the related HDR imaging methods, the problem is usually solved by Multiple Images Merging, i.e. the final HDR…

Cited by 8SourcePDFScholar
2021

Multi-stage Pre-training over Simplified Multimodal Pre-training Models

ACL 2021long

Multimodal pre-training models, such as LXMERT, have achieved excellent results in downstream tasks. However, current pre-trained models require large amounts of training data and have huge model sizes, which make them impossible to apply in low-resource situations. How to obtain similar or even bet…

2021

Task-Oriented Clustering for Dialogues

EMNLP 2021finding

A reliable clustering algorithm for task-oriented dialogues can help developer analysis and define dialogue tasks efficiently. It is challenging to directly apply prior normal text clustering algorithms for task-oriented dialogues, due to the inherent differences between them, such as coreference, o…

2021

Topic-Aware Contrastive Learning for Abstractive Dialogue Summarization

EMNLP 2021finding

Unlike well-structured text, such as news reports and encyclopedia articles, dialogue content often comes from two or more interlocutors, exchanging information with each other. In such a scenario, the topic of a conversation can vary upon progression and the key information for a certain topic is o…

2020

Multi-scale Two-way Deep Neural Network for Stock Trend Prediction

IJCAI 2020poster

Stock Trend Prediction(STP) has drawn wide attention from various fields, especially Artificial Intelligence. Most previous studies are single-scale oriented which results in information loss from a multi-scale perspective. In fact, multi-scale behavior is vital for making intelligent investment dec…

2019

Doubly Robust Joint Learning for Recommendation on Data Missing Not at Random

ICML 2019oral

In recommender systems, usually the ratings of a user to most items are missing and a critical problem is that the missing ratings are often missing not at random (MNAR) in reality. It is widely acknowledged that MNAR ratings make it difficult to accurately predict the ratings and unbiasedly estimat…

Cited by 282SourcePDFScholar
2018

KDGAN: Knowledge Distillation with Generative Adversarial Networks

NeurIPS 2018poster

Knowledge distillation (KD) aims to train a lightweight classifier suitable to provide accurate inference with constrained resources in multi-label learning. Instead of directly consuming feature-label pairs, the classifier is trained by a teacher, i.e., a high-capacity model whose training may be r…