← Search

Jianru Xue

25 accepted papers

2026

Efficient Test-Time Scaling via Hierarchical Search and Self-Verification for Discrete Diffusion Language Models

ICML 2026poster

Inference-time compute has re-emerged as a practical way to improve LLM reasoning. Most test-time scaling (TTS) algorithms rely on autoregressive decoding, which is ill-suited to discrete diffusion language models (dLLMs) due to their parallel decoding over the entire sequence. As a result, developi…

Cited by 0SourceScholar
2026

Event-Driven Sleep-Wake Scheduling for Heterogeneous Robots under LTL Constraints

RSS 2026poster

In large-scale heterogeneous robot systems (HRS), scheduling efficiency in terms of throughput and makespan relies on exploiting parallel execution, while human-issued safety and precedence instructions impose rigid temporal-logic constraints that create severe combinatorial complexity and challenge…

Cited by 0SourceScholar
2026

Guided Distillation and Risk Adaptive Evolution for Multi-Robot Navigation

AAAI 2026technical

Recent advancements in multi-robot navigation have explored methods that combine Large Language Models (LLMs) for tasks like scene understanding or high-level decision-making. However, these approaches face challenges with high inference latency and potential hallucinations. To address these challen

Cited by 0SourcePDFScholar
2026

MetaDAT: Generalizable Trajectory Prediction Via Meta Pre-Training and Data-Adaptive Test-Time Updating

ICRA 2026poster

Existing trajectory prediction methods exhibit significant performance degradation under distribution shifts during test time. Although test-time training techniques have been explored to enable adaptation, current approaches rely on an offline pre-trained predictor that lacks online learning flexib…

2026

Redundant Queries in DETR-Based 3D Detection Methods: Unnecessary and Prunable

AAAI 2026technical

Query-based models are extensively used in 3D object detection tasks, with a wide range of pre-trained checkpoints readily available online. However, despite their popularity, these models often require an excessive number of object queries, far surpassing the actual number of objects to detect. The

Cited by 0SourcePDFScholar
2025

Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis

ICCV 2025poster

Egocentricly comprehending the causes and effects of car accidents is crucial for the safety of self-driving cars, and synthesizing causal-entity reflected accident videos can facilitate the capability test to respond to unaffordable accidents in reality. However, incorporating causal relations as s…

Cited by 0SourcePDFScholar
2025

Causal-Planner: Causal Interaction Disentangling with Episodic Memory Gating for Autonomous Planning

IROS 2025

Autonomous vehicle trajectory planning faces significant challenges in dynamic traffic environments due to the complex and mixed causal relationships between critical scene elements (e.g., pedestrians, vehicles, road markings) and safe decision-making. To identify the causal factors influencing plan

Cited by 0SourcecodeScholar
2024

Abductive Ego-View Accident Video Understanding for Safe Driving Perception

CVPR 2024highlight

We present MM-AU a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11727 in-the-wild ego-view accident videos each with temporally aligned text descriptions. We annotate over 2.23 million object boxes and 58650 pairs of video-based accident reasons covering 58 accident cat…

Cited by 12SourcePDFScholar
2024

TICMapNet: A Tightly Coupled Temporal Fusion Pipeline for Vectorized HD Map Learning

RA-L 2024

High-Definition (HD) map construction is essential for autonomous driving to accurately understand the surrounding environment. Most existing methods rely on single-frame inputs to predict local map, which often fail to effectively capture the temporal correlations between frames. This limitation re

Cited by 7SourceScholar
2023

CalibDepth: Unifying Depth Map Representation for Iterative LiDAR-Camera Online Calibration

ICRA 2023poster

LiDAR-Camera online calibration is of great significance for building a stable autonomous driving perception system. For online calibration, a key challenge lies in constructing a unified and robust representation between multi-modal sensor data. Most methods extract features manually or implicitly…

Cited by 26SourcecodeScholar
2023

FEND: A Future Enhanced Distribution-Aware Contrastive Learning Framework for Long-Tail Trajectory Prediction

CVPR 2023poster

Predicting the future trajectories of the traffic agents is a gordian technique in autonomous driving. However, trajectory prediction suffers from data imbalance in the prevalent datasets, and the tailed data is often more complicated and safety-critical. In this paper, we focus on dealing with the…

2022

Sparse Semantic Map-Based Monocular Localization in Traffic Scenes Using Learned 2D-3D Point-Line Correspondences

RA-L 2022

Vision-based localization in a prior map is of crucial importance for autonomous vehicles. Given a query image, the goal is to estimate the camera pose corresponding to the prior map, and the key is the registration problem of camera images within the map. While autonomous vehicles drive on the road

Cited by 9SourceScholar
2020

Navigation Command Matching for Vision-based Autonomous Driving

ICRA 2020poster

Learning an optimal policy for autonomous driving task to confront with complex environment is a long- studied challenge. Imitative reinforcement learning is accepted as a promising approach to learn a robust driving policy through expert demonstrations and interactions with environments. However, t…

Cited by 9SourceScholar
2020

Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition

CVPR 2020poster

Skeleton-based human action recognition has attracted great interest thanks to the easy accessibility of the human skeleton data. Recently, there is a trend of using very deep feedforward neural networks to model the 3D coordinates of joints without considering the computational efficiency. In this…

Cited by 635PDFcodeScholar
2019

BLVD: Building A Large-scale 5D Semantics Benchmark for Autonomous Driving

ICRA 2019poster

In autonomous driving community, numerous benchmarks have been established to assist the tasks of 3D/2D object detection, stereo vision, semantic/instance segmentation. However, the more meaningful dynamic evolution of the surrounding objects of ego-vehicle is rarely exploited, and lacks a large-sca…

Cited by 74SourcecodeScholar
2019

Precise Correntropy-based 3D Object Modelling With Geometrical Traffic Prior

IROS 2019poster

Robust 3D perception using LiDAR is of prime importance for robotics, and its fundamental core lies in precise object modelling resisting to noise and outliers. In this paper, a precise 3D object modelling algorithm is designed especially for the intelligent vehicles. The proposed algorithm is advan…

Cited by 1SourceScholar
2019

SR-LSTM: State Refinement for LSTM Towards Pedestrian Trajectory Prediction

CVPR 2019poster

In crowd scenarios, reliable trajectory prediction of pedestrians requires insightful understanding of their social behaviors. These behaviors have been well investigated by plenty of studies, while it is hard to be fully expressed by hand-craft rules. Recent studies based on LSTM networks have show…

Cited by 636PDFScholar
2018

Adding Attentiveness to the Neurons in Recurrent Neural Networks

ECCV 2018poster

Recurrent neural networks (RNNs) are capable of modeling the temporal dynamics of complex sequential information. However, the structures of existing RNN neurons mainly focus on controlling the contributions of current and historical information but do not explore the different importance levels of…

Cited by 105SourcePDFScholar
2017

ER3: A Unified Framework for Event Retrieval, Recognition and Recounting

CVPR 2017poster

We develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alignment procedure identifies and removes the feature redundancies across frames an…

Cited by 28PDFScholar
2017

View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition From Skeleton Data

ICCV 2017poster

Skeleton-based human action recognition has recently attracted increasing attention due to the popularity of 3D skeleton data. One main challenge lies in the large view variations in captured human actions. We propose a novel view adaptation scheme to automatically regulate observation viewpoints du…

Cited by 683PDFcodeScholar