← Search

Jing Xu

77 accepted papers

2026

Flow Caching for Autoregressive Video Generation

ICLR 2026poster

Autoregressive models, often built on Transformer architectures, represent a powerful paradigm for generating ultra-long videos by synthesizing content in sequential chunks. However, this sequential generation process is notoriously slow. While caching strategies have proven effective for accelerati…

Cited by 0SourcecodeScholar
2026

Hybrid Reinforcement: when reward is sparse, better to be dense

ICLR 2026poster

Post-training for reasoning in large language models has increasingly relied on verifiable rewards: deterministic checkers that provide $0$–$1$ correctness signals. While reliable, such binary feedback is brittle—many tasks admit partially correct or alternative answers that verifiers under-credit,…

Cited by 0SourceScholar
2026

LAMER-SSL: LAYER-AWARE MIXTURE OF LORA EXPERTS FOR CONTINUAL MULTILINGUAL EXPANSION OF SELF-SUPERVISED MODELS WITHOUT FORGETTING

ICASSP 2026poster

Despite their impressive performance, self-supervised speech models often struggle to generalize to new languages and tend to forget previously acquired knowledge during continual training. To address this, we propose Lamer-SSL, a parameter-efficient framework that integrates a Layer-Aware MixturE o…

Cited by 0SourcePDFScholar
2026

Mesh-Pro: Asynchronous Advantage-guided Ranking Preference Optimization for Artist-style Quadrilateral Mesh Generation

CVPR 2026

Reinforcement learning (RL) has demonstrated remarkable success in text and image generation, yet its potential in 3D generation remains largely unexplored. Existing attempts typically rely on offline direct preference optimization (DPO) method, which suffers from low training efficiency and limited

Cited by 0SourceScholar
2026

Motion-Aware Caching for Efficient Autoregressive Video Generation

ICML 2026poster

Autoregressive video generation paradigms offer theoretical promise for long video synthesis, yet their practical deployment is hindered by the computational burden of sequential iterative denoising. While cache reuse strategies can accelerate generation by skipping redundant denoising steps, existi…

Cited by 0SourceScholar
2026

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer

CVPR 2026

Recent advances in diffusion models have greatly improved pose-driven character animation. However, existing methods are limited to spatially aligned reference-pose pairs with matched skeletal structures. Handling reference-pose misalignment remains unsolved. To address this, we present One-to-All A

Cited by 0SourcecodeScholar
2026

QuadGPT: Native Quadrilateral Mesh Generation with Autoregressive Models

ICLR 2026poster

The generation of quadrilateral-dominant meshes is a cornerstone of professional 3D content creation. However, existing generative models generate quad meshes by first generating triangle meshes and then merging triangles into quadrilaterals with some specific rules, which typically produces quad m…

Cited by 0SourceScholar
2026

RESTRAIN: From Spurious Votes to Signals — Self-Training RL with Self-Penalization

ICLR 2026poster

Reinforcement learning with human-annotated data has boosted chain-of-thought reasoning in large reasoning models, but these gains come at high costs in labeled data while faltering on harder tasks. A natural next step is experience-driven learning, where models improve without curated labels by ada…

Cited by 0SourceScholar
2026

Scaling Knowledge Editing in LLMs to 100,000 Facts with Neural KV Database

ICLR 2026poster

Efficiently editing knowledge stored in Large Language Models (LLMs) enables model updates without large-scale training. One promising solution is Locate-and-Edit (L\&E), allowing simultaneous modifications of a massive number of factual knowledge. However, such editing may compromise the general ab…

Cited by 0SourcecodeScholar
2026

Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

CVPR 2026

Humans inhabit a physical 4D world, where spatial geometry and semantic content evolve over time, forming a dynamic reality. While current Multimodal Large Language Models (MLLMs) demonstrate strong capabilities in understanding static visual inputs, it remains unclear whether they can effectively "

Cited by 0SourcecodeScholar
2025

Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference Optimization

NeurIPS 2025poster

We introduce Auto-Connect, a novel approach for automatic rigging that explicitly preserves skeletal connectivity through a connectivity-preserving tokenization scheme. Unlike previous methods that predict bone positions represented as two joints or first predict points before determining connectivi…

Cited by 0SourceScholar
2025

Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object Detection

ICCV 2025poster

At the core of Camouflaged Object Detection (COD) lies segmenting objects from their highly similar surroundings. Previous efforts navigate this challenge primarily through image-level modeling or annotation-based optimization. Despite advancing considerably, this commonplace practice hardly taps va…

2025

CoPRA: Bridging Cross-domain Pretrained Sequence Models with Complex Structures for Protein-RNA Binding Affinity Prediction

AAAI 2025technical

Accurately measuring protein-RNA binding affinity is crucial in many biological processes and drug design. Previous computational methods for protein-RNA binding affinity prediction rely on either sequence or structure features, unable to capture the binding mechanisms comprehensively. The recent em…

2025

Detecting Hallucination in Large Language Models Through Deep Internal Representation Analysis

IJCAI 2025

Large language models (LLMs) have shown exceptional performance across various domains. However, LLMs are prone to hallucinate facts and generate non-factual responses, which can undermine their reliability in real-world applications. Current hallucination detection methods suffer from external reso

2025

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

CVPR 2025poster

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging…

Cited by 23SourcePDFScholar
2025

Efficient and Privacy-Preserving Soft Prompt Transfer for LLMs

ICML 2025poster

Prompting has become a dominant paradigm for adapting large language models (LLMs). While discrete (textual) prompts are widely used for their interpretability, soft (parameter) prompts have recently gained traction in APIs. This is because they can encode information from more training samples whil…

Cited by 0SourcePDFScholar
2025

Following Length Constraints in Instructions

EMNLP 2025

Aligned instruction following models can better fulfill user requests than their unaligned counterparts. However, it has been shown that there is a length bias in evaluation of such models, and that training algorithms tend to exploit this bias by learning longer responses. In this work we show how

2025

Integrating Potential Pronunciations for Enhanced Mispronunciation Detection and Diagnosis Ability in LLMs

ICASSP 2025accepted

Large Language Models (LLMs) have exhibited significant potentials across various tasks. However, how to leverage the power of LLMs in the mispronunciation detection and diagnosis (MDD) task is still under-explored. In this paper, we propose a PP-ATP model, which integrates potential pronunciations…

Cited by 0SourceScholar
2025

LLM Agents Can Be Choice-Supportive Biased Evaluators: An Empirical Study

AAAI 2025technical

With Large Language Model (LLM) agents taking on more evaluation responsibilities in decision-making, it is essential to recognize their possible biases to guarantee fair and trustworthy AI-supported decisions. This study is the first to thoroughly examine the choice-supportive bias in LLM agents, a…

Cited by 0SourcePDFScholar
2025

Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning

NeurIPS 2025spotlight

Existing pretrained models for 3D mesh generation often suffer from data biases and produce low-quality results, while global reinforcement learning (RL) methods rely on object-level rewards that struggle to capture local structure details. To address these challenges, we present $\textbf{Mesh-RFT}$…

Cited by 0SourceScholar
2025

Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

EMNLP 2025

Large Language Models (LLMs) are rapidly surpassing human knowledge in many domains. While improving these models traditionally relies on costly human data, recent self-rewarding mechanisms have shown that LLMs can improve by judging their own responses instead of relying on human labelers. However,

Cited by 0SourcePDFScholar
2025

R.I.P.: Better Models by Survival of the Fittest Prompts

ICML 2025poster

Training data quality is one of the most important drivers of final model quality. In this work, we introduce a method for evaluating data integrity based on the assumption that low-quality input prompts result in high variance and low quality responses. This is achieved by measuring the rejected re…

Cited by 1SourcePDFScholar
2025

Self-Consistency Preference Optimization

ICML 2025poster

Self-alignment, whereby models learn to improve themselves without human annotation, is a rapidly growing research area. However, existing techniques often fail to improve complex reasoning tasks due to the difficulty of assigning correct rewards. An orthogonal approach that is known to improve corr…

Cited by 9SourcePDFScholar
2025

Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution

CVPR 2025poster

Existing diffusion-based video super-resolution (VSR) methods are susceptible to introducing complex degradations and noticeable artifacts into high-resolution videos due to their inherent randomness. In this paper, we propose a noise-robust real-world VSR framework by incorporating self-supervised…

Cited by 0SourcePDFScholar
2025

Shift the Lens: Environment-Aware Unsupervised Camouflaged Object Detection

CVPR 2025poster

Camouflaged Object Detection (COD) seeks to distinguish objects from their highly similar backgrounds. Existing work has essentially focused on isolating camouflaged objects from the environment, demonstrating ever-improving performance but at the cost of extensive annotations and complex optimizati…

2025

Understanding Nonlinear Implicit Bias via Region Counts in Input Space

ICML 2025poster

One explanation for the strong generalization ability of neural networks is implicit bias. Yet, the definition and mechanism of implicit bias in non-linear contexts remains little understood. In this work, we propose to characterize implicit bias by the count of connected regions in the input space…

Cited by 0SourcePDFScholar
2024

Augmenting Reasoning Capabilities of LLMs with Graph Structures in Knowledge Base Question Answering

EMNLP 2024finding

Recently, significant progress has been made in employing Large Language Models (LLMs) for semantic parsing to address Knowledge Base Question Answering (KBQA) tasks. Previous work utilize LLMs to generate query statements on Knowledge Bases (KBs) for retrieving answers. However, LLMs often generate…

2024

CADTalk: An Algorithm and Benchmark for Semantic Commenting of CAD Programs

CVPR 2024highlight

CAD programs are a popular way to compactly encode shapes as a sequence of operations that are easy to parametrically modify. However without sufficient semantic comments and structure such programs can be challenging to understand let alone modify. We introduce the problem of semantic commenting CA…

2024

Chain-of-Verification Reduces Hallucination in Large Language Models

ACL 2024findings

Generation of plausible yet incorrect factual information, termed hallucination, is an unsolved issue in large language models. We study the ability of language models to deliberate on the responses they give in order to correct their mistakes. We develop the Chain-of-Verification (CoVe) method wher…

Cited by 390SourcePDFScholar
2024

Con4m: Context-aware Consistency Learning Framework for Segmented Time Series Classification

NeurIPS 2024poster

Time Series Classification (TSC) encompasses two settings: classifying entire sequences or classifying segmented subsequences. The raw time series for segmented TSC usually contain Multiple classes with Varying Duration of each class (MVD). Therefore, the characteristics of MVD pose unique challenge…

2024

Enhancing Generalizable 6D Pose Tracking of an In-Hand Object With Tactile Sensing

RA-L 2024

When manipulating an object to accomplish complex tasks, humans rely on both vision and touch to keep track of the object's 6D pose. However, most existing object pose tracking systems in robotics rely exclusively on visual signals, which hinder a robot's ability to manipulate objects effectively. T

Cited by 25SourcecodeScholar
2024

Functionally Constrained Algorithm Solves Convex Simple Bilevel Problem

NeurIPS 2024poster

This paper studies simple bilevel problems, where a convex upper-level function is minimized over the optimal solutions of a convex lower-level problem. We first show the fundamental difficulty of simple bilevel problems, that the approximate optimal value of such problems is not obtainable by first…

Cited by 0SourcePDFScholar
2024

Global and Local Hierarchical Prompt Tuning Framework for Multi-level Implicit Discourse Relation Recognition

COLING 2024main

Multi-level implicit discourse relation recognition (MIDRR) is a challenging task to recognize the hierarchical discourse relations between the arguments with the absence of connectives. Recent methods tend to incorporate the static hierarchical structure containing all senses (defined as global hie…

Cited by 1SourcePDFScholar
2024

MultiSum: A Multi-Facet Approach for Extractive Social Summarization Utilizing Semantic and Sociological Relationships

AAAI 2024technical

Social summarization aims to provide summaries for a large number of social texts (called posts) about a single topic. To extract a summary, both the representation of post and summary selection method are crucial. Previous methods introduce social relation to enhance post embedding to mitigate th…

Cited by 2SourcePDFScholar
2024

Self-Rewarding Language Models

ICML 2024poster

We posit that to achieve superhuman agents, future models require superhuman feedback in order to provide an adequate training signal. Current approaches commonly train reward models from human preferences, which may then be bottlenecked by human performance level, and secondly these reward models r…

Cited by 0SourcePDFScholar
2024

Separation and Fusion: A Novel Multiple Token Linking Model for Event Argument Extraction

NAACL 2024long

In event argument extraction (EAE), a promising approach involves jointly encoding text and argument roles, and performing multiple token linking operations. This approach further falls into two categories. One extracts arguments within a single event, while the other attempts to extract arguments f…

2024

When Life Gives You Lemons, Make Cherryade: Converting Feedback from Bad Responses into Good Labels

NAACL 2024long

Deployed dialogue agents have the potential to integrate human feedback to continuously improve themselves. However, humans may not always provide explicit signals when the chatbot makes mistakes during interactions. In this work, we propose Juicer, a framework to make use of both binary and free-fo…

Cited by 19SourcePDFScholar
2023

A Closer Look at Few-shot Classification Again

ICML 2023poster

Few-shot classification consists of a training phase where a model is learned on a relatively large dataset and an adaptation phase where the learned model is adapted to previously-unseen tasks with limited labeled samples. In this paper, we empirically prove that the training algorithm and the adap…

2023

Continual Dialogue State Tracking via Example-Guided Question Answering

EMNLP 2023long main

Dialogue systems are frequently updated to accommodate new services, but naively updating them by continually training with data for new services in diminishing performance on previously learnt services. Motivated by the insight that dialogue state tracking (DST), a crucial component of dialogue sys…

Cited by 0SourcecodeScholar
2023

Dialogue State Distillation Network with Inter-slot Contrastive Learning for Dialogue State Tracking

AAAI 2023technical

In task-oriented dialogue systems, Dialogue State Tracking (DST) aims to extract users' intentions from the dialogue history. Currently, most existing approaches suffer from error propagation and are unable to dynamically select relevant information when utilizing previous dialogue states. Moreover,…

Cited by 7SourcePDFScholar
2023

Fast Actuating Multi-Helically Heated Twisted and Coiled Polymer Actuator

RA-L 2023

The twisted and coiled polymer actuator (TCPA), promisingly used in wearable robots, soft exoskeletons and prosthesis, retains the advantages of convenience, high energy density, scalable stroke and hysteresis-free. However, as a thermal actuator, its dynamic response is still rather low, not only b

Cited by 5SourceScholar
2023

Infusing Hierarchical Guidance into Prompt Tuning: A Parameter-Efficient Framework for Multi-level Implicit Discourse Relation Recognition

ACL 2023long

Multi-level implicit discourse relation recognition (MIDRR) aims at identifying hierarchical discourse relations among arguments. Previous methods achieve the promotion through fine-tuning PLMs. However, due to the data scarcity and the task gap, the pre-trained feature space cannot be accurately tu…

2023

Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedback

ACL 2023long

Frozen models trained to mimic static datasets can never improve their performance. Models that can employ internet-retrieval for up-to-date information and obtain feedback from humans during deployment provide the promise of both adapting to new information, and improving their performance. In this…

Cited by 41SourcePDFScholar
2023

Part-Guided 3D RL for Sim2Real Articulated Object Manipulation

RA-L 2023

Manipulating unseen articulated objects through visual feedback is a critical but challenging task for real robots. Existing learning-based solutions mainly focus on visual affordance learning or other pre-trained visual models to guide manipulation policies, which face challenges for novel instance

Cited by 15SourcecodeScholar
2023

Sim2Real2: Actively Building Explicit Physics Model for Precise Articulated Object Manipulation

ICRA 2023poster

Accurately manipulating articulated objects is a challenging yet important task for real robot applications. In this paper, we present a novel framework called Sim2Real2 to enable the robot to manipulate an unseen articulated object to the desired state precisely in the real world with no human demo…

Cited by 14SourcecodeScholar
2023

The CRINGE Loss: Learning what language not to model

ACL 2023long

Standard language model training employs gold human documents or human-human interaction data, and treats all training data as positive examples. Growing evidence shows that even with very large amounts of positive training data, issues remain that can be alleviated with relatively small amounts of…

Cited by 35SourcePDFScholar
2023

Towards Data-Algorithm Dependent Generalization: a Case Study on Overparameterized Linear Regression

NeurIPS 2023poster

One of the major open problems in machine learning is to characterize generalization in the overparameterized regime, where most traditional generalization bounds become inconsistent even for overparameterized linear regression. In many scenarios, this failure can be attributed to obscuring the cruc…

Cited by 2SourcePDFScholar
2023

Training Models to Generate, Recognize, and Reframe Unhelpful Thoughts

ACL 2023long

Many cognitive approaches to well-being, such as recognizing and reframing unhelpful thoughts, have received considerable empirical support over the past decades, yet still lack truly widespread adoption in self-help format. A barrier to that adoption is a lack of adequately specific and diverse ded…

2023

TransTouch: Learning Transparent Objects Depth Sensing Through Sparse Touches

IROS 2023poster

Transparent objects are common in daily life. However, depth sensing for transparent objects remains a challenging problem. While learning-based methods can leverage shape priors to improve the sensing quality, the labor-intensive data collection in real world and the sim-to-real domain gap restrict…

Cited by 3SourceScholar
2022

A Multi-turn Machine Reading Comprehension Framework with Rethink Mechanism for Emotion-Cause Pair Extraction

COLING 2022main

Emotion-cause pair extraction (ECPE) is an emerging task in emotion cause analysis, which extracts potential emotion-cause pairs from an emotional document. Most recent studies use end-to-end methods to tackle the ECPE task. However, these methods either suffer from a label sparsity problem or fail…

2022

Alleviating the Sample Selection Bias in Few-shot Learning by Removing Projection to the Centroid

NeurIPS 2022accept

Few-shot learning (FSL) targets at generalization of vision models towards unseen tasks without sufficient annotations. Despite the emergence of a number of few-shot learning methods, the sample selection bias problem, i.e., the sensitivity to the limited amount of support data, has not been well un…

2022

Bidirectional Sim-to-Real Transfer for GelSight Tactile Sensors With CycleGAN

RA-L 2022

GelSight optical tactile sensors have high-resolution and low-cost advantages and have witnessed growing adoption in various contact-rich robotic applications. Sim2Real for GelSight sensors can reduce the time cost and sensor damage during data collection and is crucial for learning-based tactile pe

Cited by 46SourcecodeScholar
2022

ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of Mind

ICLR 2022poster

Being able to predict the mental states of others is a key factor to effective social interaction. It is also crucial for distributed multi-agent systems, where agents are required to communicate and cooperate. In this paper, we introduce such an important social-cognitive skill, i.e. Theory of Mind…

2021

Bot-Adversarial Dialogue for Safe Conversational Agents

NAACL 2021long

Conversational agents trained on large unlabeled corpora of human interactions will learn patterns and mimic behaviors therein, which include offensive or otherwise toxic behavior. We introduce a new human-and-model-in-the-loop framework for evaluating the toxicity of such models, and compare a vari…

2021

Joint Communications with FH-MIMO Radar Systems: An Extended Signaling Strategy

ICASSP 2021accepted

In this paper, we investigate the signaling strategy of communications embedding in frequency-hopping (FH) multiple input multiple output (MIMO) radar. Previous work that embeds communication symbols into the emission of MIMO radar with orthogonal FH waveforms via phase modulation, and waveform orth…

Cited by 5SourceScholar
2021

MiniSeg: An Extremely Minimum Network for Efficient COVID-19 Segmentation

AAAI 2021technical

The rapid spread of the new pandemic, i.e., COVID-19, has severely threatened global health. Deep-learning-based computer-aided screening, e.g., COVID-19 infected CT area segmentation, has attracted much attention. However, the publicly available COVID-19 training data are limited, easily causing ov…

2021

Modeling of Planar Hydraulically Amplified Self-Healing Electrostatic Actuators

RA-L 2021

With the advantages of high actuation strain and specific power and ability of self-healing after dielectric breakdown, the planar hydraulically amplified self-healing electrostatic (pHASEL) actuators are promising for extensive emerging applications of soft robots. However, the relationship between

Cited by 0SourceScholar
2021

Modularized Interaction Network for Named Entity Recognition

ACL 2021long

Although the existing Named Entity Recognition (NER) models have achieved promising performance, they suffer from certain drawbacks. The sequence labeling-based NER models do not perform well in recognizing long entities as they focus only on word-level information, while the segment-based NER model…

Cited by 40SourcePDFScholar
2020

Learning Multi-Agent Coordination for Enhancing Target Coverage in Directional Sensor Networks

NeurIPS 2020poster

Maximum target coverage by adjusting the orientation of distributed sensors is an important problem in directional sensor networks (DSNs). This problem is challenging as the targets usually move randomly but the coverage range of sensors is limited in angle and distance. Thus, it is required to coor…

2019

S4G: Amodal Single-view Single-Shot SE(3) Grasp Detection in Cluttered Scenes

CoRL 2019

Grasping is among the most fundamental and long-lasting problems in robotics study. This paper studies the problem of 6-DoF(degree of freedom) grasping by a parallel gripper in a cluttered scene captured using a commodity depth sensor from a single viewpoint. We address the problem in a learning-bas

2018

Attention-Aware Compositional Network for Person Re-Identification

CVPR 2018poster

Person re-identification (ReID) is to identify pedestrians observed from different camera views based on visual appearance. It is a challenging task due to large pose variations, complex background clutters and severe occlusions. Recently, human pose estimation by predicting joint locations was larg…

Cited by 565SourcePDFScholar