← Search

Haotian Xu

32 accepted papers

2026

CODEPMP: SCALABLE PREFERENCE MODEL PRETRAINING FOR LARGE LANGUAGE MODEL REASONING

ICASSP 2026poster

Large language models (LLMs) have made significant progress in natural language understanding and generation, driven by scalable pretraining and advanced finetuning. However, enhancing reasoning abilities in LLMs, particularly via reinforcement learning from human feedback (RLHF), remains challengin…

Cited by 0SourcePDFScholar
2026

IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction

ICLR 2026poster

Recent advances in deep-research agents have shown promise for autonomous knowledge construction through dynamic reasoning over external sources. However, existing approaches rely on a mono-contextual paradigm that accumulates all information in a single, expanding context window, leading to context…

Cited by 0SourcecodeScholar
2026

LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs

AAAI 2026technical

High-quality long-context data is essential for training large language models (LLMs) capable of processing extensive documents, yet existing synthesis approaches using relevance-based aggregation face challenges of computational efficiency. We present LiteLong, a resource-efficient method for synth

Cited by 0SourcePDFScholar
2026

Online Change Point Detection for Multivariate Inhomogeneous Poisson Processes Time Series

ICML 2026poster

We study online change point detection for multivariate inhomogeneous Poisson point process time series. This setting arises commonly in applications such as earthquake seismology, climate monitoring, and epidemic surveillance, yet remains underexplored in the machine learning and statistics literat…

Cited by 0SourceScholar
2026

Reinforcing Structured Chain-of-Thought for Video Understanding

CVPR 2026

Multi-modal Large Language Models (MLLMs) show promise in video understanding. However, their reasoning often suffers from thinking drift and weak temporal comprehension, even when enhanced by Reinforcement Learning (RL) techniques like Group Relative Policy Optimization (GRPO). Moreover, existing R

Cited by 0SourceScholar
2026

Resting Neurons, Active Insights: Robustify Activation Sparsity for Large Language Models

ICML 2026poster

Activation sparsity offers a compelling route to accelerate large language model (LLM) inference by selectively suppressing hidden activations, yet existing approaches exhibit severe accuracy degradation at high sparsity. We show that this failure stems from representational instability: *activation…

Cited by 0SourceScholar
2026

SFedPO: Streaming Federated Learning with a Prediction Oracle under Temporal Shifts

ICML 2026poster

Federated Learning (FL) enables decentralized clients to collaboratively train a global model without sharing raw data. However, most existing FL frameworks assume that clients train on static local datasets collected in advance or that the data follows a fixed underlying distribution, which limits …

Cited by 0SourceScholar
2026

SafeCompass: Dynamic Chain-of-Thought Steering via Inference-Time Safety Signals

ICML 2026poster

Large reasoning models (LRMs) achieve strong performance by explicitly generating chain-of-thought (CoT) reasoning, but this reasoning process can be manipulated by adversarial prompts. Inference-time CoT interventions offer a simple and lightweight approach to improving safety, yet existing methods…

Cited by 0SourceScholar
2026

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling

ICML 2026poster

Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees. Existing safety defenses are largely incompatible with speculative inference: they either introduce additional computation or disrupt the draft–verify mechanism, negating acceleration be…

Cited by 0SourceScholar
2026

Search Self-Play: Pushing the Frontier of Agent Capability without Supervision

ICLR 2026poster

Reinforcement learning with verifiable rewards (RLVR) has become the mainstream technique for training LLM agents. However, RLVR highly depends on well-crafted task queries and corresponding ground-truth answers to provide accurate rewards, which requires significant human effort and hinders the sca…

Cited by 0SourcecodeScholar
2026

Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs

ICLR 2026poster

Direct Preference Optimization (DPO) has emerged as a cornerstone of reinforcement learning from human feedback (RLHF) due to its simplicity and efficiency. However, existing DPO-based methods typically treat all preference pairs equally, overlooking substantial variations in data quality and learni…

Cited by 0SourceScholar
2025

Agentic RL Scaling Law: Spontaneous Code Execution for Mathematical Problem Solving

NeurIPS 2025poster

Large Language Models (LLMs) often struggle with mathematical reasoning tasks requiring precise, verifiable computation. While Reinforcement Learning (RL) from outcome-based rewards enhances text-based reasoning, understanding how agents autonomously learn to leverage external tools like code execu…

Cited by 0SourcecodeScholar
2025

Beyond Single-Task: Robust Multi-Task Length Generalization for LLMs

NeurIPS 2025poster

Length generalization—the ability to solve problems longer than those seen during training—remains a critical challenge for large language models (LLMs). Previous work modifies positional encodings (PEs) and data formats to improve length generalization on specific symbolic tasks such as addition an…

Cited by 0SourceScholar
2025

DriftRemover: Hybrid Energy Optimizations for Anomaly Images Synthesis and Segmentation

IJCAI 2025

This paper tackles the challenge of anomaly image synthesis and segmentation to generate various anomaly images and their segmentation labels to mitigate the issue of data scarcity. Existing approaches employ the precise mask to guide the generation, relying on additional mask generators, leading to

2025

JPDS-NN: Reinforcement Learning-Based Dynamic Task Allocation for Agricultural Vehicle Routing Optimization

IROS 2025

The Entrance Dependent Vehicle Routing Problem (EDVRP) is a variant of the Vehicle Routing Problem (VRP) where the scale of cities influences routing outcomes, necessitating consideration of their entrances. This paper addresses EDVRP in agriculture, focusing on multi-parameter vehicle planning for

Cited by 1SourceScholar
2025

Navi2Gaze: Leveraging Foundation Models for Navigation and Target Gazing

IROS 2025

Task-aware navigation continues to be a challenging area of research, especially in scenarios involving open vocabulary. Previous studies primarily focus on finding suitable locations for task completion, often overlooking the importance of the robot’s pose. However, the robot’s orientation is cruci

Cited by 6SourcecodeScholar
2025

SilentStriker: Toward Stealthy Bit-Flip Attacks on Large Language Models

NeurIPS 2025poster

The rapid adoption of large language models (LLMs) in critical domains has spurred extensive research into their security issues. While input manipulation attacks (e.g., prompt injection) have been well-studied, Bit-Flip Attacks (BFAs)—which exploit hardware vulnerabilities to corrupt model paramete…

Cited by 0SourceScholar
2025

ZeroDiff: Solidified Visual-semantic Correlation in Zero-Shot Learning

ICLR 2025poster

Zero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-semantic correlations from seen classes. However, most current generative approaches heavily rely on having a sufficient…

2024

DexCatch: Learning to Catch Arbitrary Objects with Dexterous Hands

CoRL 2024poster

Achieving human-like dexterous manipulation remains a crucial area of research in robotics. Current research focuses on improving the success rate of pick-and-place tasks. Compared with pick-and-place, throwing-catching behavior has the potential to increase the speed of transporting objects to thei…

Cited by 4SourceScholar
2024

EMONA: Event-level Moral Opinions in News Articles

NAACL 2024long

Most previous research on moral frames has focused on social media short texts, little work has explored moral sentiment within news articles. In news articles, authors often express their opinions or political stance through moral judgment towards events, specifically whether the event is right or…

2024

FedFa: A Fully Asynchronous Training Paradigm for Federated Learning

IJCAI 2024poster

Federated learning has been identified as an efficient decentralized training paradigm for scaling the machine learning model training on a large number of devices while guaranteeing the data privacy of the trainers. FedAvg has become a foundational parameter update strategy for federated learning,…

Cited by 5SourcePDFScholar
2024

ItD: Large Language Models Can Teach Themselves Induction through Deduction

ACL 2024long

Although Large Language Models (LLMs) are showing impressive performance on a wide range of Natural Language Processing tasks, researchers have found that they still have limited ability to conduct induction. Recent works mainly adopt “post processes” paradigms to improve the performance of LLMs on…

2024

Latent 3D Graph Diffusion

ICLR 2024poster

Generating 3D graphs of symmetry-group equivariance is of intriguing potential in broad applications from machine vision to molecular discovery. Emerging approaches adopt diffusion generative models (DGMs) with proper re-engineering to capture 3D graph distributions. In this paper, we raise an ortho…

2023

A Policy Optimization Method Towards Optimal-time Stability

CoRL 2023poster

In current model-free reinforcement learning (RL) algorithms, stability criteria based on sampling methods are commonly utilized to guide policy optimization. However, these criteria only guarantee the infinite-time convergence of the system's state to an equilibrium point, which leads to sub-optima…

Cited by 3SourceScholar
2023

Change point detection and inference in multivariate non-parametric models under mixing conditions

NeurIPS 2023poster

This paper addresses the problem of localizing and inferring multiple change points, in non-parametric multivariate time series settings. Specifically, we consider a multivariate time series with potentially short-range dependence, whose underlying distributions have Hölder smooth densities and can…

Cited by 15SourcePDFScholar
2023

Efficient Exploration Using Extra Safety Budget in Constrained Policy Optimization

IROS 2023poster

Reinforcement learning (RL) has achieved promising results on most robotic control tasks. Safety of learning-based controllers is an essential notion of ensuring the effectiveness of the controllers. Current methods adopt whole consistency constraints during the training, thus resulting in inefficie…

Cited by 2SourceScholar
2022

Crossroads, Buildings and Neighborhoods: A Dataset for Fine-grained Location Recognition

NAACL 2022long

General domain Named Entity Recognition (NER) datasets like CoNLL-2003 mostly annotate coarse-grained location entities such as a country or a city. But many applications require identifying fine-grained locations from texts and mapping them precisely to geographic sites, e.g., a crossroad, an apart…

2022

Generating Disentangled Arguments with Prompts: A Simple Event Extraction Framework That Works

ICASSP 2022accepted

Event Extraction bridges the gap between text and event signals. Based on the assumption of trigger-argument dependency, existing approaches have achieved state-of-the-art performance with expert-designed templates or complicated decoding constraints. In this paper, for the first time we introduce t…

Cited by 0SourceScholar