← Search

Jun Xu

64 accepted papers

2026

Distributed Model Predictive Control for Energy and Comfort Optimization in Large Buildings Using Piecewise Affine Approximation

RA-L 2026

The control of large buildings encounters challenges in computational efficiency due to their size and nonlinear components. To address these issues, this paper proposes a Piecewise Affine (PWA)-based distributed scheme for Model Predictive Control (MPC) that optimizes energy and comfort through PWA

Cited by 0SourceScholar
2026

Mechanistic Detection and Mitigation of Hallucination in Large Reasoning Models

ICLR 2026poster

Large Reasoning Models (LRMs) have shown impressive capabilities in multi-step reasoning tasks. However, alongside these successes, a more deceptive form of model error has emerged—**Reasoning Hallucination**—where logically coherent but factually incorrect reasoning traces lead to persuasive yet fa…

Cited by 0SourcecodeScholar
2026

StyliTruth : Unlocking Stylized yet Truthful LLM Generation via Disentangled Steering

ICLR 2026poster

Generating stylized large language model (LLM) responses via representation editing is a promising way for fine-grained output control. However, there exists an inherent trade-off: imposing a distinctive style often degrades truthfulness. Existing representation editing methods, by naively injecting…

Cited by 0SourceScholar
2026

Thinker: Training LLMs in Hierarchical Thinking for Deep Search via Multi-Turn Interaction

AAAI 2026technical

Efficient retrieval of external knowledge bases and web pages is crucial for enhancing the reasoning abilities of LLMs. Previous works on training LLMs to leverage external retrievers for solving complex problems have predominantly employed end-to-end reinforcement learning. However, these approache

Cited by 0SourcePDFScholar
2026

Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm

ICML 2026poster

Continual Pre-Training (CPT) is essential for enabling Language Models (LMs) to integrate new factual knowledge without erasing old. While classical CPT techniques like data replay have become the standard paradigm, the mechanisms underlying how LMs acquire and retain facts over time, termed as cont…

Cited by 0SourceScholar
2025

AdaO2B: Adaptive Online to Batch Conversion for Out-of-Distribution Generalization

AAAI 2025technical

Online to batch conversion involves constructing a new batch learner by utilizing a series of models generated by an existing online learning algorithm, for achieving generalization guarantees under i.i.d assumption. However, when applied to real-world streaming applications such as streaming recomm…

Cited by 0SourcePDFScholar
2025

Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning

EMNLP 2025

In legal practice, judges apply the trichotomous dogmatics of criminal law, sequentially assessingthe elements of the offense, unlawfulness, and culpability to determine whether an individual’s conduct constitutes a crime. Although current legal large language models (LLMs) show promising accuracy i

2025

Bridging Jensen Gap for Max-Min Group Fairness Optimization in Recommendation

ICLR 2025poster

Group max-min fairness (MMF) is commonly used in fairness-aware recommender systems (RS) as an optimization objective, as it aims to protect marginalized item groups and ensures a fair competition platform. However, our theoretical analysis indicates that integrating MMF constraint violates the assu…

2025

From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions

ICLR 2025oral

Tool learning enables Large Language Models (LLMs) to interact with external environments by invoking tools, serving as an effective strategy to mitigate the limitations inherent in their pre-training data. In this process, tool documentation plays a crucial role by providing usage instructions for…

2025

GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents

NeurIPS 2025poster

Recent Graphical User Interface (GUI) agents replicate the R1-Zero paradigm, coupling online Reinforcement Learning (RL) with explicit chain-of-thought reasoning prior to object grounding and thereby achieving substantial performance gains. In this paper, we first conduct extensive analysis experime…

Cited by 0SourcecodeScholar
2025

Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis

AAAI 2025technical

High-quality, large-scale instructions are crucial for aligning large language models (LLMs), however, there is a severe shortage of instruction in the field of natural language understanding (NLU). Previous works on constructing NLU instructions mainly focus on information extraction (IE), neglect…

Cited by 0SourcePDFScholar
2025

Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learning

EMNLP 2025

Legal mathematical reasoning is essential for applying large language models (LLMs) in high-stakes legal contexts, where outputs must be both mathematically accurate and procedurally compliant. However, existing legal LLMs lack structured numerical reasoning, and open-domain models, though capable o

2025

Length-Induced Embedding Collapse in PLM-based Models

ACL 2025long

Text embeddings from PLM-based models enable a wide range of applications, yet their performance often degrades on longer texts. In this paper, we introduce a phenomenon we call Length Collapse, where embeddings of longer texts tend to cluster together. This clustering results in a distributional in…

2025

MAPS: Motivation-Aware Personalized Search via LLM-Driven Consultation Alignment

ACL 2025long

Personalized product search aims to retrieve and rank items that match users’ preferences and search intent. Despite their effectiveness, existing approaches typically assume that users’ query fully captures their real motivation. However, our analysis of a real-world e-commerce platform reveals tha…

2025

Media Source Matters More Than Content: Unveiling Political Bias in LLM-Generated Citations

EMNLP 2025

Unlike traditional search engines that present ranked lists of webpages, generative search engines rely solely on in-line citations as the key gateway to original real-world webpages, making it crucial to examine whether LLM-generated citations have biases—particularly for politically sensitive quer

2025

ParseCaps: An Interpretable Parsing Capsule Network for Medical Image Diagnosis

AAAI 2025technical

Deep learning has excelled in medical image classification, but its clinical application is limited by poor interpretability. Capsule networks, known for encoding hierarchical relationships and spatial features, show potential in addressing this issue. Nevertheless, traditional capsule networks…

Cited by 1SourcePDFScholar
2025

Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents

ICLR 2025poster

Previous studies have found that PLM-based retrieval models exhibit a preference for LLM-generated content, assigning higher relevance scores to these documents even when their semantic quality is comparable to human-written ones. This phenomenon, known as source bias, threatens the sustainable deve…

2025

Q-PRM: Adaptive Query Rewriting for Retrieval-Augmented Generation via Step-level Process Supervision

EMNLP 2025

Query rewriting plays a pivotal role in Retrieval-Augmented Generation (RAG) by refining real-world queries of varying complexity. Existing approaches typically rely on outcome-supervised training or heuristic rules to guide the rewriting process. However, these paradigms often struggle to handle qu

Cited by 0SourcePDFScholar
2025

ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability

ICLR 2025spotlight

Retrieval-Augmented Generation (RAG) models are designed to incorporate external knowledge, reducing hallucinations caused by insufficient parametric (internal) knowledge. However, even with accurate and relevant retrieved content, RAG models can still produce hallucinations by generating outputs th…

Cited by 8SourcePDFScholar
2025

Similarity = Value? Consultation Value-Assessment and Alignment for Personalized Search

EMNLP 2025

Personalized search systems in e-commerce platforms increasingly involve user interactions with AI assistants, where users consult about products, usage scenarios, and more. Leveraging consultation to personalize search services is trending. Existing methods typically rely on semantic similarity to

2025

Towards Effective and Efficient Continual Pre-training of Large Language Models

ACL 2025long

Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. In this paper, we comprehensively study its key designs to balance the new abilities while retaining the original abilities, and present an effective CPT method that can greatly imp…

2025

Trigger3:Refining Query Correction via Adaptive Model Selector

AAAI 2025technical

In search scenarios, user experience can be hindered by erroneous queries due to typos, voice errors, or knowledge gaps. Therefore, query correction is crucial for search engines. Current correction models, usually small models trained on specific data, often struggle with queries beyond their train…

2025

Uplift-RAG: Uplift-Driven Knowledge Preference Alignment for Retrieval-Augmented Generation

EMNLP 2025

Retrieval-augmented generation (RAG) has proven effective in enhancing the knowledge coverage of large language models (LLMs) and mitigating hallucinations by incorporating external retrieved documents. However, documents deemed relevant by the retriever are not necessarily helpful for answer genera

2024

A Study of Implicit Ranking Unfairness in Large Language Models

EMNLP 2024finding

Recently, Large Language Models (LLMs) have demonstrated a superior ability to serve as ranking models. However, concerns have arisen as LLMs will exhibit discriminatory ranking behaviors based on users’ sensitive attributes (gender). Worse still, in this paper, we identify a subtler form of discrim…

2024

ChatUIE: Exploring Chat-based Unified Information Extraction Using Large Language Models

COLING 2024main

Recent advancements in large language models have shown impressive performance in general chat. However, their domain-specific capabilities, particularly in information extraction, have certain limitations. Extracting structured information from natural language that deviates from known schemas or i…

2024

Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration

ACL 2024findings

The proliferation of Large Language Models (LLMs) has led to an influx of AI-generated content (AIGC) on the internet, transforming the corpus of Information Retrieval (IR) systems from solely human-written to a coexistence with LLM-generated content. The impact of this surge in AIGC on IR systems r…

2024

Effective In-Context Example Selection through Data Compression

ACL 2024findings

In-context learning has been extensively validated in large language models. However, the mechanism and selection strategy for in-context example selection, which is a crucial ingredient in this approach, lacks systematic and in-depth research. In this paper, we propose a data compression approach t…

Cited by 1SourcePDFScholar
2024

Efficient Knowledge Infusion via KG-LLM Alignment

ACL 2024findings

To tackle the problem of domain-specific knowledge scarcity within large language models (LLMs), knowledge graph-retrievalaugmented method has been proven to be an effective and efficient technique for knowledge infusion. However, existing approaches face two primary challenges: knowledge mismatch b…

2024

Logic Rules as Explanations for Legal Case Retrieval

COLING 2024main

In this paper, we address the issue of using logic rules to explain the results from legal case retrieval. The task is critical to legal case retrieval because the users (e.g., lawyers or judges) are highly specialized and require the system to provide logic, faithful, and interpretable explanations…

2024

OrthCaps: An Orthogonal CapsNet with Sparse Attention Routing and Pruning

CVPR 2024poster

Redundancy is a persistent challenge in Capsule Networks (CapsNet) leading to high computational costs and parameter counts. Although previous studies have introduced pruning after the initial capsule layer dynamic routing's fully connected nature and non-orthogonal weight matrices reintroduce redun…

Cited by 8SourcePDFScholar
2024

Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation

ICML 2024poster

Fine-tuning pretrained large models to downstream tasks is an important problem, which however suffers from huge memory overhead due to large-scale parameters. This work strives to reduce memory overhead in fine-tuning from perspectives of activation function and layer normalization. To this end, we…

2024

Smooth Start: A Unified Approach for Gradual Transition from Cold to Old in Recommender Systems

ICASSP 2024accepted

In recommender systems, the cold-start problem poses a significant challenge, especially as users transition from being new to more engaged. Existing solutions often lack the granularity to accommodate this evolving user engagement, resulting in suboptimal performance for intermediate and older user…

Cited by 0SourceScholar
2023

A Dynamic Multi-Scale Voxel Flow Network for Video Prediction

CVPR 2023highlight

The performance of video prediction has been greatly boosted by advanced deep neural networks. However, most of the current methods suffer from large model sizes and require extra inputs, e.g., semantic/depth maps, for promising performance. For efficiency consideration, in this paper, we propose a…

2023

Conditional counterfactual causal effect for individual attribution

UAI 2023poster

Identifying the causes of an event, also termed as causal attribution, is a commonly encountered task in many application problems. Available methods, mostly in Bayesian or causal inference literature, suffer from two main drawbacks: 1) cannot attribute for individuals, and 2) attributing one singl…

Cited by 9SourcePDFScholar
2023

Joint Semantic and Strategy Matching for Persuasive Dialogue

EMNLP 2023long findings

Persuasive dialogue aims to persuade users to achieve some targets by conversations. While previous persuasion models have achieved notable successes, they mostly base themselves on utterance semantic matching, and an important aspect has been ignored, that is, the strategy of the conversations, for…

Cited by 0SourceScholar
2023

Reward Imputation with Sketching for Contextual Batched Bandits

NeurIPS 2023poster

Contextual batched bandit (CBB) is a setting where a batch of rewards is observed from the environment at the end of each episode, but the rewards of the non-executed actions are unobserved, resulting in partial-information feedback. Existing approaches for CBB often ignore the rewards of the non-ex…

Cited by 0SourcePDFScholar
2022

"Restore Globally, Refine Locally: A Mask-Guided Scheme to Accelerate Super-Resolution Networks"

ECCV 2022poster

"Single image super-resolution (SR) has been boosted by deep convolutional neural networks with growing model complexity and computational costs. To deploy existing SR networks onto edge devices, it is necessary to accelerate them for large image (4K) processing. The different areas in an image ofte…

2022

Extracting Trigger-sharing Events via an Event Matrix

EMNLP 2022finding

A growing interest emerges in event extraction which aims to extract multiple events with triggers and arguments. Previous methods mitigate the problem of multiple events extraction by predicting the arguments conditioned on the event trigger and event type, assuming that these arguments belong to a…

2022

Incremental Cross-View Mutual Distillation for Self-Supervised Medical CT Synthesis

CVPR 2022poster

Due to the constraints of the imaging device and high cost in operation time, computer tomography (CT) scans are usually acquired with low within-slice resolution. Improving the inter-slice resolution is beneficial to the disease diagnosis for both human experts and computer-aided systems. To this e…

Cited by 25PDFScholar
2022

Optimal Partial Transport Based Sentence Selection for Long-form Document Matching

COLING 2022main

One typical approach to long-form document matching is first conducting alignment between cross-document sentence pairs, and then aggregating all of the sentence-level matching signals. However, this approach could be problematic because the alignment between documents is partial — despite two docum…

2022

Where to Go for the Holidays: Towards Mixed-Type Dialogs for Clarification of User Goals

ACL 2022long

Most dialog systems posit that users have figured out clear and specific goals before starting an interaction. For example, users have determined the departure, the destination, and the travel time for booking a flight. However, in many scenarios, limited by experience and knowledge, users may know…

2021

A Joint Model for Dropped Pronoun Recovery and Conversational Discourse Parsing in Chinese Conversational Speech

ACL 2021long

In this paper, we present a neural model for joint dropped pronoun recovery (DPR) and conversational discourse parsing (CDP) in Chinese conversational speech. We show that DPR and CDP are closely related, and a joint model benefits both tasks. We refer to our model as DiscProReco, and it first encod…

2021

Discovering Dialog Structure Graph for Coherent Dialog Generation

ACL 2021long

Learning discrete dialog structure graph from human-human dialogs yields basic insights into the structure of conversation, and also provides background knowledge to facilitate dialog generation. However, this problem is less studied in open-domain dialogue. In this paper, we conduct unsupervised di…

2021

Regret Bounds for Online Kernel Selection in Continuous Kernel Space

AAAI 2021technical

Regret bounds of online kernel selection in a finite kernel set have been well studied, having at least an order O( √ NT) of magnitude after T rounds, where N is the number of candidate kernels. But it is still an unsolved problem to achieve sublinear regret bounds of online kernel selection in a co…

Cited by 5SourcePDFScholar
2021

Temporal Modulation Network for Controllable Space-Time Video Super-Resolution

CVPR 2021poster

Space-time video super-resolution (STVSR) aims to increase the spatial and temporal resolutions of low-resolution and low-frame-rate videos. Recently, deformable convolution based methods have achieved promising STVSR performance, but they could only infer the intermediate frame pre-defined in the t…

Cited by 113PDFcodeScholar
2020

Enhancing Dialog Coherence with Event Graph Grounded Content Planning

IJCAI 2020poster

How to generate informative, coherent and sustainable open-domain conversations is a non-trivial task. Previous work on knowledge grounded conversation generation focus on improving dialog informativeness with little attention on dialog coherence. In this paper, to enhance multi-turn dialog coherenc…

2020

ICNet: Intra-saliency Correlation Network for Co-Saliency Detection

NeurIPS 2020poster

Intra-saliency and inter-saliency cues have been extensively studied for co-saliency detection (Co-SOD). Model-based methods produce coarse Co-SOD results due to hand-crafted intra- and inter-saliency features. Current data-driven models exploit inter-saliency cues, but undervalue the potential powe…

2020

Learning to Learn Kernels with Variational Random Features

ICML 2020poster

We introduce kernels with random Fourier features in the meta-learning framework for few-shot learning. We propose meta variational random features (MetaVRF) to learn adaptive kernels for the base-learner, which is developed in a latent variable model by treating the random feature basis as the late…

Cited by 34SourcePDFScholar
2020

Learning to Learn with Variational Information Bottleneck for Domain Generalization

ECCV 2020poster

Domain generalization models learn to generalize to previously unseen domains, but suffer from prediction uncertainty and domain shift. In this paper, we address both problems. We introduce a probabilistic meta-learning model for domain generalization, in which classifier parameters shared across do…

Cited by 196SourcePDFScholar
2019

RANet: Ranking Attention Network for Fast Video Object Segmentation

ICCV 2019poster

Despite online learning (OL) techniques have boosted the performance of semi-supervised video object segmentation (VOS) methods, the huge time costs of OL greatly restricts their practicality. Matching based and propagation based methods run at a faster speed by avoiding OL techniques. However, they…

Cited by 274PDFcodeScholar
2018

A Hybrid l1-l0 Layer Decomposition Model for Tone Mapping

CVPR 2018poster

Tone mapping aims to reproduce a standard dynamic range image from a high dynamic range image with visual information preserved. State-of-the-art tone mapping algorithms mostly decompose an image into a base layer and a detail layer, and process them accordingly. These methods may have problems of h…

Cited by 175SourcePDFScholar
2018

A PID Controller Approach for Stochastic Optimization of Deep Networks

CVPR 2018poster

Deep neural networks have demonstrated their power in many computer vision applications. State-of-the-art deep architectures such as VGG, ResNet, and DenseNet are mostly optimized by the SGD-Momentum algorithm, which updates the weights by considering their past and current gradients. Nonetheless, S…

2018

Multivariate Time Series Imputation with Generative Adversarial Networks

NeurIPS 2018poster

Multivariate time series usually contain a large number of missing values, which hinders the application of advanced analysis methods on multivariate time series data. Conventional approaches to addressing the challenge of missing values, including mean/zero imputation, case deletion, and matrix fac…

Cited by 673SourcePDFScholar
2017

Multi-Channel Weighted Nuclear Norm Minimization for Real Color Image Denoising

ICCV 2017poster

Most of the existing denoising algorithms are developed for grayscale images. It is not trivial to extend them for color image denoising since the noise statistics in R, G, and B channels can be very different for real noisy images. In this paper, we propose a multi-channel (MC) optimization model f…

Cited by 359PDFScholar
2015

Patch Group Based Nonlocal Self-Similarity Prior Learning for Image Denoising

ICCV 2015poster

Patch based image modeling has achieved a great success in low level vision such as image denoising. In particular, the use of image nonlocal self-similarity (NSS) prior, which refers to the fact that a local patch often has many nonlocal similar patches to it across the image, has significantly enh…

Cited by 472PDFScholar