← Search

Rui Song

74 accepted papers

2026

An Intention-Guided Reinforcement Learning Approach With Dirichlet Energy Constraint for Heterogeneous Multi-Robot Cooperation

RA-L 2026

Multi-robot systems have demonstrated significant potential in accomplishing complex tasks, such as cooperative pursuit, search-and-rescue operations. The emergence of heterogeneous robots with diverse capabilities and characteristics shows superior adaptability compared with homogeneous teams. Howe

Cited by 0SourceScholar
2026

AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese Corpora

AAAI 2026technical

Comprehension of ancient texts plays an important role in archaeology and understanding of Chinese history and civilization. The rapid development of large language models needs benchmarks that can evaluate their comprehension of ancient characters. Existing Chinese benchmarks are mostly targeted at

Cited by 0SourcePDFScholar
2026

CaLoRA-Stereo: Robust Stereo Endoscopic Depth Estimation Network Via Camera-Aware LoRA and Dual-View Geometry

ICRA 2026poster

Stereo depth estimation has drawn widespread attention from the robotics and vision community due to its broad applications such as 3D reconstruction. Recently, stereo matching foundation models have made significant progress by being trained on the large-scale datasets containing natural images. Ho…

Cited by 0Scholar
2026

EnerGS: Energy-Based Gaussian Splatting under Partial Geometric Observability

ICML 2026poster

3D Gaussian Splatting (3DGS) has been widely adopted for scene reconstruction, where training inherently constitutes a highly coupled and non-convex optimization problem. Recent works commonly incorporate geometric priors, such as LiDAR measurements, either for initialization or as training constrai…

Cited by 0SourceScholar
2026

Exploring 6D Object Pose Estimation with Deformation

CVPR 2026

We present DeSOPE, a large-scale dataset for 6DoF deformed objects. Most 6D object pose methods assume rigid or articulated objects, an assumption that fails in practice as objects deviate from their canonical shapes due to wear, impact, or deformation. To model this, we introduce the DeSOPE dataset

Cited by 0SourcecodeScholar
2026

Low-Latency Neural LiDAR Compression with 2D Context Models

ICLR 2026poster

Context modeling is fundamental to LiDAR point cloud compression. Existing methods rely on computationally intensive 3D contexts, such as voxel and octree, which struggle to balance the compression efficiency and coding speed. In this work, we propose a neural LiDAR compressor based on 2D context mo…

Cited by 0SourcecodeScholar
2026

OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial Perspective

CVPR 2026

Semantic Scene Completion (SSC) is essential for 3D perception in mobile robotics, as it enables holistic scene understanding by jointly estimating dense volumetric occupancy and per-voxel semantics. Although SSC has been widely studied in terrestrial domains such as autonomous driving, aerial setti

Cited by 0SourcecodeScholar
2026

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

CVPR 2026

End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining. However, we find that current VLMs struggle to understand fine-gra

Cited by 0SourcecodeScholar
2026

TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments

ICML 2026poster

Robots in dynamic, human-centric environments must follow language instructions while maintaining real-time reactive control. Vision-language-action (VLA) models offer a promising framework, but they assume temporally aligned reasoning and control, despite semantic inference being inherently delayed…

Cited by 0SourceScholar
2026

Towards Global Sparse and Partial Point Set Registration with Pose-Robust Completion for Computer-Assisted Orthopedic Surgery

ICRA 2026poster

In computer-assisted orthopedic surgery (CAOS), accurately registering sparse and partial intraoperative point sets with a complete preoperative model remains highly challenging due to limited overlap, extreme sparsity, and point localisation noise. In this paper, we propose a novel end-to-end compl…

Cited by 0Scholar
2025

A Crab-Inspired Soft Gripper with Single-Finger Dexterous Grasping Capabilities

IROS 2025

Soft grippers conform to the shape and surface properties of the objects to be grasped, effectively avoiding damage to soft and fragile items. Despite the variety of existing soft gripper designs, their structures lack sufficient flexibility for effectively grasping slender objects or operating in n

Cited by 0SourceScholar
2025

A Deep Learning Modeling Method for the Identification of Robot Curtain Wall Assembly State

RA-L 2025

The precise identification of the robot curtain wall assembly state is crucial for improving construction efficiency. Traditional methods remain sensitive to noise in high-dimensional sensor data and require extensive datasets, resulting in limited generalization. A method for identifying the robot

Cited by 0SourceScholar
2025

CoDa-4DGS: Dynamic Gaussian Splatting with Context and Deformation Awareness for Autonomous Driving

ICCV 2025poster

Dynamic scene rendering opens new avenues in autonomous driving by enabling closed-loop simulations with photorealistic data, which is crucial for validating end-to-end algorithms. However, the complex and highly dynamic nature of traffic environments presents significant challenges in accurately re…

Cited by 0SourcePDFScholar
2025

Complex Robotic Manipulation via Hindsight Goal Diffusion and Graph-based Experience Replay

IROS 2025

Goal-conditioned reinforcement learning (GCRL) is an effective method for multi-goal robotic manipulation tasks. Many studies based on hindsight experience replay (HER) and hindsight goal generation (HGG) have achieved the autonomous acquisition of robotic manipulation in reward-sparse environments

Cited by 0SourceScholar
2025

DGJA: Dependency Graph-enhanced Joint Attention Structure for Multimodal Sarcasm Detection

ICASSP 2025accepted

Multimodal sarcasm detection (MSD) leverages multimodal data, including both images and text, to detect whether the input content contains sarcastic information. Despite recent advances, existing MSD approaches often overlook the imbalance in sarcastic content between text and image modalities, wher…

Cited by 0SourceScholar
2025

Directed Spatial Consistency-Based Partial-to-Partial Point Cloud Registration with Deep Graph Matching

IROS 2025

3D point cloud registration is an essential problem in computer vision, robotics, surgical navigation and augmented reality. Accurate registration of partially overlapped intraoperative point clouds (e.g., femoral reconstruction) remains critical yet challenging in orthopedic navigation due to incom

Cited by 0SourcecodeScholar
2025

Feature-aligned Fisheye Object Detection Network for Autonomous Driving

IROS 2025

Fisheye cameras, renowned for their panoramic field of view (FOV) of 360°, are crucial for surround-view perception in autonomous driving. However, research on object perception in fisheye images lags behind that of standard images. To address this gap, we propose a feature-aligned fisheye object de

Cited by 0SourceScholar
2025

IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals

NeurIPS 2025poster

Semantic Scene Completion (SSC) has emerged as a pivotal approach for jointly learning scene geometry and semantics, enabling downstream applications such as navigation in mobile robotics. The recent generalization to Panoptic Scene Completion (PSC) advances the SSC domain by integrating instance-le…

Cited by 0SourceScholar
2025

Registration After Completion: Towards Sparse and Partial Point Set Registration for Computer-Assisted Orthopedic Surgery

IROS 2025

In computer-assisted orthopedic surgery (CAOS), accurate point set registration is essential for enhancing surgical accuracy. However, the sparse and low-overlap nature of intraoperative point sets presents significant challenges for reliable registration. To deal with these challenges, we propose a

Cited by 0SourceScholar
2025

Revisiting 3D Curve to Surface Registration using Tangent and Normal Vectors for Computer-Assisted Orthopedic Surgery

IROS 2025

In this paper, we present a novel curve-to-surface registration method, termed Bi-directional Hybrid Mixture Model Registration based on Dual-constrained Tangent and Normal Vectors (BiHMM-DTN), where two different tangent vectors at the intraoperative point are simultaneously used with the normal ve

Cited by 1SourcecodeScholar
2025

Robust and Accurate Multi-View 2D/3D Image Registration with Differentiable X-Ray Rendering and Dual Cross-View Constraints

ICRA 2025

Robust and accurate 2D/3D registration, which aligns preoperative models with intraoperative images of the same anatomy, is crucial for successful interventional navigation. To mitigate the challenge of a limited field of view in single-image intraoperative scenarios, multi-view 2D/3D registration i

Cited by 1SourceScholar
2025

SCFlow2: Plug-and-Play Object Pose Refiner with Shape-Constraint Scene Flow

CVPR 2025poster

We introduce SCFlow2, a plug-and-play refinement framework for 6D object pose estimation. Most recent 6D object pose methods rely on refinement to get accurate results. However, most existing refinements either suffer from noises in establishing correspondences, or rely on retraining for novel objec…

Cited by 0SourcePDFScholar
2025

TrafficLoc: Localizing Traffic Surveillance Cameras in 3D Scenes

ICCV 2025poster

We tackle the problem of localizing traffic cameras within a 3D reference map and propose a novel image-to-point cloud registration (I2P) method, TrafficLoc, in a coarse-to-fine matching fashion. To overcome the lack of large-scale real-world intersection datasets, we first introduce Carla Intersect…

2025

Unsupervised Liver Deformation Correction Network Using Optimal Transport for Image-Guided Liver Surgery

IROS 2025

In this paper, we propose a novel unsupervised intraoperative liver deformation correction method, called Learning Coherent point drift Network (LCNet), for image-guided liver surgery (IGLS). We first estimate the correspondences between the preoperative and intraoperative point sets in the optimal

Cited by 0SourceScholar
2024

Bidirectional Partial-to-Full Non-Rigid Point Set Registration with Non-Overlapping Filtering

IROS 2024poster

In this paper, we introduce Bidirectional Non-Overlapping Filtering Network (Bi-NOFNet), which registers the partial intraoperative point set with full preoperative point set for computer-assisted interventions (CAI). Our contributions are three-folds. First, Bi-NOFNet adopts customised feature extr…

Cited by 0SourceScholar
2024

Collaborative Semantic Occupancy Prediction with Hybrid Feature Fusion in Connected Automated Vehicles

CVPR 2024poster

Collaborative perception in automated vehicles leverages the exchange of information between agents aiming to elevate perception results. Previous camera-based collaborative 3D perception methods typically employ 3D bounding boxes or bird's eye views as representations of the environment. However th…

Cited by 19SourcePDFScholar
2024

Data-Centric Explainable Debiasing for Improving Fairness in Pre-trained Language Models

ACL 2024findings

Human-like social bias of pre-trained language models (PLMs) on downstream tasks have attracted increasing attention. The potential flaws in the training data are the main factor that causes unfairness in PLMs. Existing data-centric debiasing strategies mainly leverage explicit bias words (defined a…

2024

DeepBHMR: Learning Bidirectional Hybrid Mixture Models for Generalized Rigid Point Set Registration

IROS 2024

In this paper, we introduce a novel normal-assisted learning-based rigid registration approach, i.e., Deep Bi-directional Hybrid Mixture Registration (DeepBHMR). Our approach utilises helpful normal vectors explicitly in both correspondence and transformation stages and formulates the optimization o

Cited by 3SourcecodeScholar
2024

Effect Size Estimation for Duration Recommendation in Online Experiments: Leveraging Hierarchical Models and Objective Utility Approaches

AAAI 2024technical

The selection of the assumed effect size (AES) critically determines the duration of an experiment, and hence its accuracy and efficiency. Traditionally, experimenters determine AES based on domain knowledge. However, this method becomes impractical for online experimentation services managing numer…

Cited by 1SourcePDFScholar
2024

FlingFlow: LLM-Driven Dynamic Strategies for Efficient Cloth Flattening

RA-L 2024

The proficiency of robots in cloth manipulation is crucial for their potential widespread deployment in household service contexts, with the task of unfolding cloth being particularly indispensable. Unlike rigid objects, cloth has a high-dimensional state space, which poses significant challenges fo

Cited by 7SourceScholar
2024

History-Aware Planning for Risk-free Autonomous Navigation on Unknown Uneven Terrain

ICRA 2024poster

It is challenging for the mobile robot to achieve autonomous and mapless navigation in the unknown environment with uneven terrain. In this study, we present a layered and systematic pipeline. At the local level, we maintain a tree structure that is dynamically extended with the navigation. This str…

Cited by 1SourcecodeScholar
2024

OBHMR: Robust Partial-to-full Generalized Point Set Registration with Overlap-guided Bidirectional Hybrid Mixture Model

IROS 2024poster

In this paper, we introduce a novel overlap-based bidirectional point set registration approach, i.e., Overlap-guided Bidirectional Hybrid Mixture Registration (OBHMR), which incorporates geometric information (i.e., normal vectors) in both the correspondence and transformation stages and formulates…

Cited by 1SourcecodeScholar
2024

Online Posterior Sampling with a Diffusion Prior

NeurIPS 2024poster

Posterior sampling in contextual bandits with a Gaussian prior can be implemented exactly or approximately using the Laplace approximation. The Gaussian prior is computationally efficient but it cannot describe complex distributions. In this work, we propose approximate posterior sampling algorithms…

Cited by 0SourcePDFScholar
2024

Self-supervised Preference Optimization: Enhance Your Language Model with Preference Degree Awareness

EMNLP 2024finding

Recently, there has been significant interest in replacing the reward model in Reinforcement Learning with Human Feedback (RLHF) methods for Large Language Models (LLMs), such as Direct Preference Optimization (DPO) and its variants. These approaches commonly use a binary cross-entropy mechanism on…

2024

TACIT: A Target-Agnostic Feature Disentanglement Framework for Cross-Domain Text Classification

AAAI 2024technical

Cross-domain text classification aims to transfer models from label-rich source domains to label-poor target domains, giving it a wide range of practical applications. Many approaches promote cross-domain generalization by capturing domaininvariant features. However, these methods rely on unlabeled…

2024

TUMTraf V2X Cooperative Perception Dataset

CVPR 2024poster

Cooperative perception offers several benefits for enhancing the capabilities of autonomous vehicles and improving road safety. Using roadside sensors in addition to onboard sensors increases reliability and extends the sensor range. External sensors offer higher situational awareness for automated…

2023

A Reinforcement Learning Framework for Dynamic Mediation Analysis

ICML 2023poster

Mediation analysis learns the causal effect transmitted via mediator variables between treatments and outcomes, and receives increasing attention in various scientific domains to elucidate causal relations. Most existing works focus on point-exposure studies where each subject only receives one trea…

2023

An Instrumental Variable Approach to Confounded Off-Policy Evaluation

ICML 2023poster

Off-policy evaluation (OPE) aims to estimate the return of a target policy using some pre-collected observational data generated by a potentially different behavior policy. In many cases, there exist unmeasured variables that confound the action-reward or action-next-state relationships, rendering m…

Cited by 21SourcePDFScholar
2023

Design and Development of a Rapidly Deployable Low-Cost Tensegrity In-Pipe Robot

IROS 2023poster

Existing in-pipe robots have insufficient adaptability when dealing with accidents in unfamiliar pipe environments. Developing a pipe robot that can be designed and manufactured quickly is one solution. The tensegrity structure is a self-stressing spatial structure formed by the interaction of rigid…

Cited by 0SourceScholar
2023

Enhancing Ontology Translation Through Cross-Lingual Agreement

ICASSP 2023accepted

Ontology serves as the foundation for the underlying representation of knowledge. In order to achieve the sharing of knowledge across languages, ontologies that are typically only represented in English must be translated into different languages. Building a domain-specific translation system is nec…

Cited by 0SourceScholar
2023

FCC: Feature Clusters Compression for Long-Tailed Visual Recognition

CVPR 2023poster

Deep Neural Networks (DNNs) are rather restrictive in long-tailed data, since they commonly exhibit an under-representation for minority classes. Various remedies have been proposed to tackle this problem from different perspectives, but they ignore the impact of the density of Backbone Features (BF…

2023

Fast Recognition of Snap-Fit for Industrial Robot Using a Recurrent Neural Network

RA-L 2023

Snap-fit recognition is an essential capability for industrial robots in manufacturing. The goal is to protect fragile parts by quickly detecting snap-fit signals in the assembly. In this letter, we propose a fast recognition method of snap-fit for industrial robots. A snap-fit dataset generation st

Cited by 11SourceScholar
2023

Human-Robot Deformation Manipulation Skill Transfer: Sequential Fabric Unfolding Method For Robots

RA-L 2023

Deformable object manipulation has been considered a challenging task for robots for its complex dynamics and the infinite dimensional configuration space. Fabric unfolding manipulation takes on critical significance in the textile industry and household services. Accordingly, enabling robots to pos

Cited by 5SourceScholar
2023

On Heterogeneous Treatment Effects in Heterogeneous Causal Graphs

ICML 2023poster

Heterogeneity and comorbidity are two interwoven challenges associated with various healthcare problems that greatly hampered research on developing effective treatment and understanding of the underlying neurobiological mechanism. Very few studies have been conducted to investigate heterogeneous ca…

2023

Pseudo Flow Consistency for Self-Supervised 6D Object Pose Estimation

ICCV 2023poster

Most self-supervised 6D object pose estimation methods can only work with additional depth information or rely on the accurate annotation of 2D segmentation masks, limiting their application range. In this paper, we propose a 6D object pose estimation method that can be trained with pure RGB images…

Cited by 12PDFcodeScholar
2023

Rigidity-Aware Detection for 6D Object Pose Estimation

CVPR 2023poster

Most recent 6D object pose estimation methods first use object detection to obtain 2D bounding boxes before actually regressing the pose. However, the general object detection methods they use are ill-suited to handle cluttered scenes, thus producing poor initialization to the subsequent pose networ…

2023

Shape-Constraint Recurrent Flow for 6D Object Pose Estimation

CVPR 2023poster

Most recent 6D object pose estimation methods rely on 2D optical flow networks to refine their results. However, these optical flow methods typically do not consider any 3D shape information of the targets during matching, making them suffer in 6D object pose estimation. In this work, we propose a s…

2023

V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative Perception

CVPR 2023highlight

Modern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has demonstrated that the Vehicle-to-Vehicle (V2V) cooperative perception…

2022

A Simple Contrastive Learning Framework for Interactive Argument Pair Identification via Argument-Context Extraction

EMNLP 2022main

Interactive argument pair identification is an emerging research task for argument mining, aiming to identify whether two arguments are interactively related. It is pointed out that the context of the argument is essential to improve identification performance. However, current context-based methods…

2022

A Tensegrity-Based Inchworm-Like Robot for Crawling in Pipes With Varying Diameters

RA-L 2022

Most current in-pipe robots are usually designed for pipes of a specific size. In this letter, we propose a novel inchworm-like in-pipe robot based on the concept of tensegrity for moving in pipes with varying diameters. Firstly, a tensegrity-based robotic module capable of two kinds of shape change

Cited by 34SourceScholar
2022

An In-pipe Crawling Robot based on Tensegrity Structures

IROS 2022poster

This paper presents a novel concept to develop robots capable of crawling in tubular environments, inspired by the movement of earthworms and the biological musculoskeletal systems in nature. A tensegrity structures-based robotic module with shape changeability actuated by only one linear actuator i…

Cited by 3SourceScholar
2022

Federated Multi-Task Attention for Cross-Individual Human Activity Recognition

IJCAI 2022poster

Federated Learning (FL) is an emerging privacy-aware machine learning technique that applies successfully to the collaborative learning of global models for Human Activity Recognition (HAR). As of now, the applications of FL for HAR assume that the data associated with diverse individuals follow the…

2022

Locally Aggregated Feature Attribution on Natural Language Model Understanding

NAACL 2022long

With the growing popularity of deep-learning models, model understanding becomes more important. Much effort has been devoted to demystify deep neural networks for better explainability. Some feature attribution methods have shown promising results in computer vision, especially the gradient-based m…

2022

Low-drift LiDAR-only Odometry and Mapping for UGVs in Environments with Non-level Roads

IROS 2022poster

This study focuses on localization and mapping for UGVs when they are deployed in environments with non-level roads. In these scenarios, the vehicles need to travel through flat but not necessarily level grounds, i.e., ascent or descent, which may cause drifts of the robot pose and distortion of the…

Cited by 2SourceScholar
2022

OctAttention: Octree-Based Large-Scale Contexts Model for Point Cloud Compression

AAAI 2022technical

In point cloud compression, sufficient contexts are significant for modeling the point cloud distribution. However, the contexts gathered by the previous voxel-based methods decrease when handling sparse point clouds. To address this problem, we propose a multiple-contexts deep learning framework ca…

2021

ANOCE: Analysis of Causal Effects with Multiple Mediators via Constrained Structural Learning

ICLR 2021poster

In the era of causal revolution, identifying the causal effect of an exposure on the outcome of interest is an important problem in many areas, such as epidemics, medicine, genetics, and economics. Under a general causal graph, the exposure may have a direct effect on the outcome and also an indirec…

Cited by 13SourcePDFScholar
2021

Deep Jump Learning for Off-Policy Evaluation in Continuous Treatment Settings

NeurIPS 2021poster

We consider off-policy evaluation (OPE) in continuous treatment settings, such as personalized dose-finding. In OPE, one aims to estimate the mean outcome under a new treatment decision rule using historical data generated by a different decision rule. Most existing works on OPE focus on discrete tr…

2020

Does the Markov Decision Process Fit the Data: Testing for the Markov Property in Sequential Decision Making

ICML 2020poster

The Markov assumption (MA) is fundamental to the empirical validity of reinforcement learning. In this paper, we propose a novel Forward-Backward Learning procedure to test MA in sequential decision making. The proposed test does not assume any parametric form on the joint distribution of the observ…

Cited by 52SourcePDFScholar
2020

On Validation and Planning of An Optimal Decision Rule with Application in Healthcare Studies

ICML 2020poster

In the current era of personalized recommendation, one major interest is to develop an optimal individualized decision rule that assigns individuals with the best treatment option according to their covariates. Estimation of optimal decision rules (ODR) has been extensively investigated recently, ho…

Cited by 1SourcePDFScholar