← Search

Xinyi Liu

26 accepted papers

2026

AeroGS: Scale-Aware Gaussian Splatting for Pose-Free Dynamic UAV Scene Reconstruction

CVPR 2026

Monocular UAV videos pose a fundamental challenge for 3D reconstruction: dynamic scene modeling requires accurate camera poses, yet recovering poses from long UAV trajectories often fails in texture-sparse regions and in the presence of moving objects. Existing approaches typically handle either pos

Cited by 0SourceScholar
2026

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning

ICML 2026poster

Reinforcement Learning (RL) has become pivotal for improving model capabilities yet suffers from rollout efficiency bottlenecks due to the long-tail response length distribution. While existing works mitigate the impact of long tails via prompt-level tail scheduling, we focus on the root source of i…

Cited by 0SourceScholar
2026

FreeAdapt: Unleashing Diffusion Priors for Ultra-High-Definition Image Restoration

ICLR 2026poster

Latent Diffusion Models (LDMs) have recently shown great potential for image restoration owing to their powerful generative priors. However, directly applying them to ultra-high-definition image restoration (UHD-IR) often results in severe global inconsistencies and loss of fine-grained details, pri…

Cited by 0SourceScholar
2026

OncoCoT: A Temporal-causal Chain-of-Thought Dataset for Oncologic Decision-Making

AAAI 2026technical

Long Chain-of-Thought (CoT) reasoning has shown great promise in complex reasoning tasks, but its application to medical decision-making presents unique challenges. Unlike structured tasks relying on static verification frameworks, medical decision-making requires dynamic validation through longitud

Cited by 0SourcePDFScholar
2026

PiLoT: Neural Pixel-to-3D Registration for UAV-based Ego and Target Geo-localization

CVPR 2026

We present PiLoT, a unified framework that tackles UAV-based ego and target geo-localization. Conventional approaches rely on decoupled pipelines that fuse GNSS and Visual-Inertial Odometry (VIO) for ego-pose estimation, and active sensors like laser rangefinders for target localization. However, th

Cited by 0SourcecodeScholar
2026

ShieldRAG: Safeguarding Retrieval-Augmented Generation from Untrusted Knowledge Bases

AAAI 2026technical

Open knowledge bases (e.g., websites) are widely adopted in Retrieval-Augmented Generation (RAG) systems to provide supplementary knowledge (e.g., latest information). However, such sources inevitably contain biased or harmful content, and incorporating these untrusted contents into the RAG process

Cited by 0SourcePDFScholar
2026

SkySplat: Generalizable 3D Gaussian Splatting from Multi-Temporal Sparse Satellite Images

AAAI 2026technical

Three-dimensional scene reconstruction from sparse-view satellite images is a long-standing and challenging task. While 3D Gaussian Splatting (3DGS) and its variants have recently attracted attention for its high efficiency, existing methods remain unsuitable for satellite images due to incompatibil

Cited by 0SourcePDFScholar
2025

A Systematic Survey of Claim Verification: Corpora, Systems, and Case Studies

EMNLP 2025

Automated Claim Verification (CV)—the task of assessing a claim’s veracity against explicitly provided evidence—is a critical tool in the fight against growing misinformation. This survey offers a comprehensive analysis of 198 studies published between January 2022 and March 2025, synthesizing recen

Cited by 0SourcePDFScholar
2025

Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction

EMNLP 2025

LLM-as-a-judge has become a promising paradigm for using large language models (LLMs) to evaluate natural language generation (NLG), but the uncertainty of its evaluation remains underexplored. This lack of reliability may limit its deployment in many applications. This work presents the first frame

Cited by 0SourcePDFScholar
2025

CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for Guidance

ICCV 2025poster

Semi-dense feature matching methods have shown strong performance in challenging scenarios. However, the existing pipeline relies on a global search across the entire feature map to establish coarse matches, limiting further improvements in accuracy and efficiency. Motivated by this limitation, we p…

2025

Exploring Fine-Grained Human Motion Video Captioning

COLING 2025main

Detailed descriptions of human motion are crucial for effective fitness training, which highlights the importance of research in fine-grained human motion video captioning. Existing video captioning models often fail to capture the nuanced semantics of videos, resulting in the generated descriptions…

2025

MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals

ICASSP 2025accepted

Video-based physiology, exemplified by remote photoplethysmography (rPPG), extracts physiological signals such as pulse and respiration by analyzing subtle changes in video recordings. This non-contact, real-time monitoring method holds great potential for home settings. Despite the valuable contrib…

Cited by 0SourceScholar
2025

NetMoE: Accelerating MoE Training through Dynamic Sample Placement

ICLR 2025spotlight

Mixture of Experts (MoE) is a widely used technique to expand model sizes for better model quality while maintaining the computation cost constant. In a nutshell, an MoE model consists of multiple experts in each model layer and routes the training tokens to only a fixed number of experts rather tha…

Cited by 1SourcePDFScholar
2025

PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy

ACL 2025long

This paper introduces PreP-OCR, a two-stage pipeline that combines document image restoration with semantic-aware post-OCR correction to enhance both visual clarity and textual consistency, thereby improving text extraction from degraded historical documents.First, we synthesize document-image pairs…

2025

The Role of Model Confidence on Bias Effects in Measured Uncertainties for Vision-Language Models

EMNLP 2025

With the growing adoption of Large Language Models (LLMs) for open-ended tasks, accurately assessing epistemic uncertainty, which reflects a model’s lack of knowledge, has become crucial to ensuring reliable outcomes. However, quantifying epistemic uncertainty in such tasks is challenging due to the

2025

UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis

ACL 2025finding

Recent advancements in Large Vision-Language Models are accelerating the development of Graphical User Interface (GUI) agents that utilize human-like vision perception capabilities to enhance productivity on digital devices. Compared to approaches predicated on GUI metadata, which are platform-depen…

Cited by 0SourcePDFScholar
2025

V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and Prediction

ICCV 2025poster

Vehicle-to-everything (V2X) technologies offer a promising paradigm to mitigate the limitations of constrained observability in single-vehicle systems. Prior work primarily focuses on single-frame cooperative perception, which fuses agents' information across different spatial locations but ignores…

2024

Accurate and Efficient Loop Closure Detection With Deep Binary Image Descriptor and Augmented Point Cloud Registration

IROS 2024poster

Loop Closure Detection (LCD) is an essential component of Simultaneous Localization and Mapping (SLAM), helping to correct drift errors, facilitate map merging, or both by identifying previously observed scenes. Despite its importance, traditional LCD algorithms based on single sensor such as camera…

Cited by 0SourceScholar
2024

Guided Online Distillation: Promoting Safe Reinforcement Learning by Offline Demonstration

ICRA 2024poster

Safe Reinforcement Learning (RL) aims to find a policy that achieves high rewards while satisfying cost constraints. When learning from scratch, safe RL agents tend to be overly conservative, which impedes exploration and restrains the overall performance. In many realistic tasks, e.g. autonomous dr…

Cited by 10SourceScholar
2024

RANSAC Back to SOTA: A Two-Stage Consensus Filtering for Real-Time 3D Registration

RA-L 2024

Correspondence-based point cloud registration (PCR) plays a key role in robotics and computer vision. However, challenges like sensor noises, object occlusions, and descriptor limitations inevitably result in numerous outliers. RANSAC family is the most popular outlier removal solution. However, the

Cited by 19SourcecodeScholar
2024

SGCalib: A Two-stage Camera-LiDAR Calibration Method Using Semantic Information and Geometric Features

ICRA 2024poster

Extrinsic calibration is an essential prerequisite for the applications of camera-LiDAR fusion. Existing methods either suffer from the complex offline setting of man-made targets or tend to produce suboptimal and unrobust results. In this paper, we propose an online two-stage calibration method tha…

Cited by 4SourceScholar
2022

ELSR: Efficient Line Segment Reconstruction With Planes and Points Guidance

CVPR 2022poster

Three-dimensional (3D) line segments are helpful for scene reconstruction. Most of the existing 3D-line-segment-reconstruction algorithms deal with two views or dozens of small-size images; while in practice there are usually hundreds or thousands of large-size images. In this paper, we propose an e…

Cited by 23PDFScholar
2022

Predicting Sagittal-Plane Swing Hip Kinematics in Response to Trips

RA-L 2022

State-of-the-art wearable lower-limb robot controllers typically use established baseline human kinematics during common mobility tasks. Unfortunately due to the variability in human response during perturbations, these lower-limb controllers are unable to effectively assist with perturbation recove

Cited by 2SourceScholar