← Search

Yue Wang

330 accepted papers

2026

$\pi$-BA: Probabilistic Neural Bundle Adjustment With Iterative Cycle Optimization for Driving Scene Reconstruction

RA-L 2026

Urban scene reconstruction under noisy camera poses remains a critical challenge for autonomous driving. While recent neural dense Bundle Adjustment (BA) methods have shown promising results in specific settings, their performance often degrades in real-world urban scenarios due to noisy corresponde

Cited by 0SourceScholar
2026

AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis

ICRA 2026poster

The collection of large-scale and diverse robot demonstrations remains a major bottleneck for imitation learning, as real-world data acquisition is costly and simulators offer limited diversity and fidelity with pronounced sim-to-real gaps. While generative models present an attractive solution, exi…

2026

AnyAmber: A Generalist for Versatile Anonymous Bearing and Range Based Position Tracking

RSS 2026poster

Position tracking based on bearing measurements and Ultra-wideband (UWB) ranging is widely used in robotic navigation tasks. However, due to variations in the number of robots, anchor configurations, UWB tag layouts, and the presence or absence of anonymous visual observations, existing methods typi…

Cited by 0SourceScholar
2026

BIOARC: Discovering Optimal Neural Architectures for Biological Foundation Models

ICML 2026poster

Foundation models have revolutionized AI, yet biological applications often repurpose general architectures without accounting for the intrinsic structural and functional properties of distinct modalities, such as genomic and proteomic sequences. Consequently, these architectures lack the inductive …

Cited by 0SourceScholar
2026

Beyond Buffer Limits: Energy-Based Data Reassembly for Continual Learning

ICML 2026poster

Continual learning (CL) aims to acquire new knowledge from a non-stationary data stream while retaining performance on previously learned tasks. Memory-based replay methods mitigate catastrophic forgetting by storing and revisiting past samples, but their effectiveness is fundamentally constrained b…

Cited by 0SourceScholar
2026

Bio-Inspired Liquid Crystal Elastomer Suction Actuator for Intelligent Robotic Grasping

ICRA 2026poster

Grasping operations constitute a fundamental mechanism for robotic interaction with the environment and task execution, playing a critical role in logistics, unmanned systems, and complex terrain exploration. Conventional rigid grasping devices are often bulky and exhibit limited adaptability and co…

Cited by 0Scholar
2026

CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers

ICML 2026poster

Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness. However, since human reviews often cover only a subset of salient issues and sometimes contain mistakes, they are unreliable as gold references. To a…

Cited by 0SourceScholar
2026

D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping

ICLR 2026poster

Simulation provides a cost-effective and flexible platform for data generation and policy learning to develop robotic systems. However, bridging the gap between simulation and real-world dynamics remains a significant challenge, especially in physical parameter identification. In this work, we intro…

Cited by 0SourcecodeScholar
2026

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

ICLR 2026poster

Reinforcement learning (RL) with large language models shows promise in complex reasoning. However, its progress is hindered by the lack of large-scale training data that is sufficiently challenging, contamination-free and verifiable. To this end, we introduce DeepMath-103K, a large-scale mathematic…

Cited by 0SourcecodeScholar
2026

Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving

ICLR 2026poster

End-to-End (E2E) solutions have emerged as a mainstream approach for autonomous driving systems, with Vision-Language-Action (VLA) models representing a new paradigm that leverages pre-trained multimodal knowledge from Vision-Language Models (VLMs) to interpret and interact with complex real-world e…

Cited by 0SourcecodeScholar
2026

Disentangling Length Bias in Preference Learning via Response-Conditioned Modeling

ICLR 2026poster

Reinforcement Learning from Human Feedback (RLHF) has achieved considerable success in aligning large language models (LLMs) by modeling human preferences with a learnable reward model and employing a reinforcement learning algorithm to maximize the reward model's scores. However, these reward model…

Cited by 0SourceScholar
2026

Distilling Geometry Priors for 3D-Consistent Video Generation

ICML 2026poster

While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures arise because standard denoising objectives lack explicit ince…

Cited by 0SourceScholar
2026

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment

ICML 2026poster

Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complicated training process. To address this, we propose a training-free inference method to sample directly from the optimal RL policy. The transition probab…

Cited by 0SourceScholar
2026

Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in Clutter (I)

ICRA 2026poster

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation…

Cited by 0codeScholar
2026

Electrospun TPU/LCE Composite Fibers for High-Performance Biomimetic Tendon Actuation

ICRA 2026poster

Traditional rigid actuators in soft robotics, particularly for bionic hands, suffer from structural complexity, bulkiness, and limited biomimetic motion. To address these limitations, we developed an electrospun composite fiber membrane composed of thermoplastic polyurethane (TPU) and liquid crystal…

Cited by 0Scholar
2026

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation

ICRA 2026poster

From loco-motion to dextrous manipulation, humanoid robots have made remarkable strides in demonstrating complex full-body capabilities. However, the majority of current robot learning datasets and benchmarks mainly focus on stationary robot arms, and the few existing humanoid datasets are either co…

2026

InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields

CVPR 2026

Existing depth estimation methods are fundamentally limited to predicting depth on discrete image grids. Such representations restrict their scalability to arbitrary output resolutions and hinder the geometric detail recovery. This paper introduces InfiniDepth, which represents depth as neural impli

Cited by 0SourcecodeScholar
2026

MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning

ICML 2026poster

Multi-objective discrete optimization problems, such as molecular design, pose significant challenges due to their vast and unstructured combinatorial spaces. Traditional evolutionary algorithms often get trapped in local optima, while expert knowledge can provide crucial guidance for accelerating c…

Cited by 0SourceScholar
2026

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

ICML 2026poster

Robust reinforcement learning (RL) under the average-reward criterion is essential for long-term decision-making, particularly when the environment may differ from its training dynamics. However, most existing studies focus on model-based settings and provide only asymptotic guarantees, hindering th…

Cited by 0SourceScholar
2026

ORVIT: Near-Optimal Online Distributionally Robust Reinforcement Learning

AAAI 2026technical

Reinforcement learning (RL) faces significant challenges in real-world deployments due to the sim-to-real gap, where policies trained in simulators often underperform in practice due to mismatches between training and deployment conditions. Distributionally robust RL addresses this issue by optimizi

Cited by 0SourcePDFScholar
2026

Pantheon360: Taming Digital Twin Generation via 3D-Aware 360deg Video Diffusion

CVPR 2026

Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency--constraints that remain challenging for perspective video generators due to their limited field of view (FoV). Their narrow FoV forces long or multi-view tra

Cited by 0SourceScholar
2026

PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies

RSS 2026poster

A significant challenge for robot learning research is our ability to accurately measure and compare the performance of robot policies. Benchmarking in robotics is historically challenging due to the stochasticity, reproducibility, and time-consuming nature of real-world rollouts. This challenge is …

Cited by 26SourceScholar
2026

Predictive Local Planning with Multi-Step Reward and Q-Value Forecasting

ICRA 2026poster

Planning in dynamic environments often relies on explicit future observation prediction or value-based estimation, both of which can be brittle or hard to generalize in uncertain settings. We propose a novel model-based reinforcement learning framework that performs trajectory rollout and optimizati…

Cited by 0Scholar
2026

ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation

ICLR 2026poster

Unified multimodal models (UMMs) have shown remarkable advances in jointly understanding and generating text and images. However, prevailing evaluations treat these abilities in isolation, such that tasks with multimodal inputs and outputs are scored primarily through unimodal reasoning: textual ben…

Cited by 0SourcecodeScholar
2026

Resource Efficient Sleep Staging via Multi-Level Masking and Prompt Learning

AAAI 2026technical

Automatic sleep staging plays a vital role in assessing sleep quality and diagnosing sleep disorders. Most existing methods rely heavily on long and continuous EEG recordings, which poses significant challenges for data acquisition in resource-constrained systems, such as wearable or home-based moni

Cited by 0SourcePDFScholar
2026

SFCLTA: Spectral Fusion Contrastive Learning with Topology-Adaptive Graph Augmentation

ICML 2026poster

Graph Neural Networks (GNNs) have achieved remarkable successes in graph analysis due to the Message-Passing (MP) mechanism, yet they struggle with heterophilic graphs where connected nodes often have distinct labels or dissimilar attributes. Graph Contrastive Learning (GCL) serves as a promising ap…

Cited by 0SourceScholar
2026

SSHPool: The Separated Subgraph-based Hierarchical Pooling

AAAI 2026technical

In this paper, we develop a novel local graph pooling method, namely the Separated Subgraph-based Hierarchical Pooling (SSHPool), for graph classification. We commence by assigning the nodes of a sample graph into different clusters, resulting in a family of separated subgraphs. We individually empl

Cited by 0SourcePDFScholar
2026

Sample-Efficient Distributionally Robust Multi-Agent Reinforcement Learning via Online Interaction

ICLR 2026poster

Well-trained multi-agent systems can fail when deployed in real-world environments due to model mismatches between the training and deployment environments, caused by environment uncertainties including noise or adversarial attacks. Distributionally Robust Markov Games (DRMGs) enhance system resilie…

Cited by 0SourceScholar
2026

SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment

ICRA 2026poster

Developing efficient and accurate visuomotor policies poses a central challenge in robotic imitation learning. While recent rectified flow approaches have advanced visuomotor policy learning, they suffer from a key limitation: After iterative distillation, generated actions may deviate from the grou…

2026

Shape Sensing and Tip Tracking Via Reciprocating Magnet in the Soft Continuum Robot

ICRA 2026poster

Soft continuum robots, attributable to inherently compliant trunks and shape manipulability, have been widely deployed in complex scenarios requiring safe human-robot interaction. However, their nonlinear deformations and hyperredundant degrees of freedom pose substantial challenges for full-body sh…

Cited by 0Scholar
2026

Sym-Servo: Disambiguate Symmetric Object Pose by End-To-End Optimal Visual Servo

ICRA 2026poster

Controlling symmetric objects is an indispensable but challenging task in robotic manipulation. Mainstream perception-action frameworks rely on accurate 6D pose estimation to guide the controller. However, the majority of existing 6D pose estimation methods for symmetric objects are designed to outp…

Cited by 0Scholar
2026

Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?

ICLR 2026poster

Spatial embodied intelligence often operates under partial observability, where agents must act to acquire missing information rather than passively consume complete observations. In such settings, progress depends on actively selecting informative actions that reduce uncertainty and support the con…

Cited by 0SourcecodeScholar
2026

Toward Embodiment Equivariant Vision-Language-Action Policy

ICRA 2026poster

Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches emphasize model size, dataset scale and diversity while paying le…

2026

UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as Reasoning

ICLR 2026poster

GUI grounding, which maps natural-language instructions to actionable UI elements, is a core capability of GUI agents. Prior work largely treats instructions as a static proxy for user intent, overlooking the impact of instruction diversity on grounding performance. Through a careful investigation o…

Cited by 0SourcecodeScholar
2026

UnIRe: Unsupervised Instance Decomposition for Dynamic Urban Scene Reconstruction

ICRA 2026poster

Reconstructing and decomposing dynamic urban scenes is crucial for autonomous driving, urban planning, and scene editing. However, existing methods fail to perform instance-aware decomposition without manual annotations, which is crucial for instance-level scene editing. We propose UnIRe, a 3D Gauss…

2026

UniST-Pred: A Robust Unified Framework for Spatio-Temporal Traffic Forecasting in Transportation Networks Under Disruptions

IJCAI 2026

Spatio-temporal traffic forecasting is a core component of intelligent transportation systems, supporting various downstream tasks such as signal control and network-level traffic management. In real-world deployments, forecasting models must operate under structural and observational uncertainties,

Cited by 0Scholar
2026

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

ICML 2026poster

Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelmed by global context. Recent "Thinking-with-Images" methods alleviate this by iteratively zooming into regions of interes…

Cited by 0SourceScholar
2026

Ψ0Ψ0\Psi_0: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation

RSS 2026poster

We introduce Ψ₀ (Psi-Zero), an open foundation model to address challenging humanoid loco-manipulation tasks. While existing approaches often attempt to address this fundamental problem by co-training on large and diverse human and humanoid data, we argue that this strategy is suboptimal due to the …

Cited by 0SourceScholar
2025

A Reduction Framework for Distributionally Robust Reinforcement Learning under Average Reward

ICML 2025poster

Robust reinforcement learning (RL) under the average reward criterion, which seeks to optimize long-term system performance in uncertain environments, remains a largely unexplored area. To address this challenge, we propose a reduction-based framework that transforms robust average reward optimizati…

Cited by 0SourcePDFScholar
2025

A Selective Learning Method for Temporal Graph Continual Learning

ICML 2025poster

Node classification is a key task in temporal graph learning (TGL). Real-life temporal graphs often introduce new node classes over time, but existing TGL methods assume a fixed set of classes. This assumption brings limitations, as updating models with full data is costly, while focusing only on ne…

Cited by 0SourcePDFScholar
2025

AKBR: Learning Adaptive Kernel-based Representations for Graph Classification

IJCAI 2025

In this paper, we propose a new model to learn Adaptive Kernel-based Representations (AKBR) for graph classification. Unlike state-of-the-art R-convolution graph kernels that are defined by merely counting any pair of isomorphic substructures between graphs and cannot provide an end-to-end learning

2025

Adaptive Neural Uncalibrated Visual Servo with Zero-shot Transfer of Extrinsics and Scenes

IROS 2025

Deploying visual servo controller to novel scenes with uncertain parameters requires additional manual effort for calibration. Traditional methods tackle this problem by online estimating the Jacobian matrix. However, they struggle in challenging scenes due to intrinsic limitations. For instance, im

Cited by 0SourceScholar
2025

Adaptive Wavelet-Positional Encoding for High-Frequency Information Learning in Implicit Neural Representation

AAAI 2025technical

Implicit Neural Representation (INR) has shown great potential in constructing the complex nature signal as a continuous implicit function. However, the representation results are incomplete since different components of the signal correspond to different frequencies and neural network inherently te…

Cited by 0SourcePDFScholar
2025

An End-to-End Simple Clustering Hierarchical Pooling Operation for Graph Learning Based on Top-K Node Selection

IJCAI 2025

Graph Neural Networks (GNNs) are powerful tools for graph learning, but one of the important challenges is how to effectively extract representations for graph-level tasks. In this paper, we propose an end-to-end Simple Clustering Hierarchical Pooling (SCHPool) operation, which is based on Top-K nod

2025

An Intelligent Skeleton Based on Liquid Metal for Biohybrid Actuator Powered by Muscle

IROS 2025

Biological machines that use biological cells and soft materials in combination to obtain a sense of the environment driven by bioenergy and generate driving force are called biohybrid actuators. With the development of tissue engineering and organoid technology, researchers have applied biohybrid a

Cited by 0SourceScholar
2025

Aria-UI: Visual Grounding for GUI Instructions

ACL 2025finding

Digital agents for automating tasks across different platforms by directly manipulating the GUIs are increasingly important. For these agents, grounding from language instructions to target elements remains a significant challenge due to reliance on HTML or AXTree inputs. In this paper, we introduce…

Cited by 0SourcePDFScholar
2025

BEV-DWPVO: BEV-Based Differentiable Weighted Procrustes for Low Scale-Drift Monocular Visual Odometry on Ground

RA-L 2025

Monocular Visual Odometry (MVO) provides a cost-effective, real-time positioning solution for autonomous vehicles. However, MVO systems face the common issue of lacking inherent scale information from monocular cameras. Traditional methods have good interpretability but can only obtain relative scal

Cited by 2SourceScholar
2025

CNSv2: Probabilistic Correspondence Encoded Neural Image Servo

ICRA 2025

Visual servo based on traditional image matching methods often requires accurate keypoint correspondence for high precision control. However, keypoint detection or matching tends to fail in challenging scenarios with inconsistent illuminations or textureless objects, resulting significant performanc

Cited by 2SourceScholar
2025

COME: Dual Structure-Semantic Learning with Collaborative MoE for Universal Lesion Detection Across Heterogeneous Ultrasound Datasets

ICCV 2025poster

Conventional single-dataset training often fails with new data distributions, especially in ultrasound (US) image analysis due to limited data, acoustic shadows, and speckle noise.Therefore, constructing a universal framework for multi-heterogeneous US datasets is imperative. However, a key challeng…

2025

Capsizing-Guided Trajectory Optimization for Autonomous Navigation with Rough Terrain

IROS 2025

It is a challenging task for ground robots to autonomously navigate in harsh environments due to the presence of non-trivial obstacles and uneven terrain. This requires trajectory planning that balances safety and efficiency. The primary challenge is to generate a feasible trajectory that prevents r

Cited by 0SourceScholar
2025

CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-Scale Reinforcement Learning in Autonomous Driving

CVPR 2025poster

Trajectory planning is vital for autonomous driving, ensuring safe and efficient navigation in complex environments. While recent learning-based methods, particularly reinforcement learning (RL), have shown promise in specific scenarios, RL planners struggle with training inefficiencies and managing…

2025

Contradicted in Reliable, Replicated in Unreliable: Dual-Source Reference for Fake News Early Detection

AAAI 2025technical

Early detection of fake news is crucial to mitigate its negative impact. Current research in fake news detection often utilizes the difference between real and fake news regarding the support degree from reliable sources. However, it has overlooked their different semantic outlier degrees among unre…

Cited by 0SourcePDFScholar
2025

DORec: Decomposed Object Reconstruction and Segmentation Utilizing 2D Self-Supervised Features

RA-L 2025

Recovering 3D geometry and textures of individual objects is crucial for many robotics applications, such as manipulation, pose estimation, and autonomous driving. However, decomposing a target object from a complex background is challenging. Most existing approaches rely on costly manual labels to

Cited by 1SourceScholar
2025

Deformation Configuration Estimation for Soft Continuum Robot Utilizing Seq2Seq Learning

RA-L 2025

Inspired by biological tentacles, soft continuum robots exhibit the potential for navigating through narrow spaces and operating in complex environments, offering extensive application possibilities. However, owing to their inherent compliance, soft continuum robots may undergo unpredictable deforma

Cited by 1SourceScholar
2025

Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning

IROS 2025

Grasp-based manipulation tasks are fundamental to robots interacting with their environments, yet gripper state ambiguity significantly reduces the robustness of imitation learning policies for these tasks. Data-driven solutions face the challenge of high real-world data costs, while simulation data

Cited by 2SourceScholar
2025

DreamDrive: Generative 4D Scene Modeling from Street View Images

ICRA 2025

Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based methods create 3D scenes from driving logs and synthesize geometry-consistent driving videos through neural rendering, b

Cited by 24SourceScholar
2025

ENAHPool: The Edge-Node Attention-based Hierarchical Pooling for Graph Neural Networks

ICML 2025poster

Graph Neural Networks (GNNs) have emerged as powerful tools for graph learning, and one key challenge arising in GNNs is the development of effective pooling operations for learning meaningful graph representations. In this paper, we propose a novel Edge-Node Attention-based Hierarchical Pooling (EN…

Cited by 0SourcePDFScholar
2025

EVolSplat: Efficient Volume-based Gaussian Splatting for Urban View Synthesis

CVPR 2025poster

Novel view synthesis of urban scenes is essential for autonomous driving-related applications. Existing NeRF and 3DGS-based methods show promising results in achieving photorealistic renderings but require slow, per-scene optimization. We introduce EVolSplat, an efficient 3D Gaussian Splatting model…

Cited by 0SourcePDFScholar
2025

Exploring the Over-smoothing Problem of Graph Neural Networks for Graph Classification: An Entropy-based Viewpoint

IJCAI 2025

The over-smoothing has emerged as a major challenge in the development of Graph Neural Networks (GNNs). While existing state-of-the-art methods effectively mitigate the diminishing distance between nodes and improve the performance of node classification, they tend to be elusive for graph-level task

2025

Extrapolated Urban View Synthesis Benchmark

ICCV 2025poster

Photorealistic simulators are essential for the training and evaluation of vision-centric autonomous vehicles (AVs). At their core is Novel View Synthesis (NVS), a crucial capability that generates diverse unseen viewpoints to accommodate the broad and continuous pose distribution of AVs. Recent adv…

2025

Fantastic Copyrighted Beasts and How (Not) to Generate Them

ICLR 2025poster

Recent studies show that image and video generation models can be prompted to reproduce copyrighted content from their training data, raising serious legal con- cerns about copyright infringement. Copyrighted characters (e.g., Mario, Batman) present a significant challenge: at least one lawsuit has…

Cited by 12SourcePDFScholar
2025

First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training

NeurIPS 2025poster

Improving Multi-modal Large Language Models (MLLMs) in the post-training stage typically relies on supervised fine-tuning (SFT) or reinforcement learning (RL), which require expensive and manually annotated multi-modal data--an ultimately unsustainable resource. This limitation has motivated a growi…

Cited by 0SourcecodeScholar
2025

From Uncertain to Safe: Conformal Adaptation of Diffusion Models for Safe PDE Control

ICML 2025poster

The application of deep learning for partial differential equation (PDE)-constrained control is gaining increasing attention. However, existing methods rarely consider safety requirements crucial in real-world applications. To address this limitation, we propose Safe Diffusion Models for PDE Control…

2025

GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs

ICML 2025poster

Large language model (LLM) unlearning has demonstrated its essential role in removing privacy and copyright-related responses, crucial for their legal and safe applications. However, the pursuit of complete unlearning often comes with substantial costs due to its compromises in their general functio…

Cited by 0SourcePDFScholar
2025

Global Static Pruning via Adaptive Sample Complexity Awareness

ICASSP 2025accepted

Dynamic pruning leverage the feature information of each input sample to dynamically adjust the network structure, generating multiple subnetworks suitable for different sample complexity. However, it inevitably introduces higher computational complexity and increased memory consumption. In addition…

Cited by 0SourceScholar
2025

Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions

CVPR 2025poster

Grounding 3D object affordance is a task that locates objects in 3D space where they can be manipulated, which links perception and action for embodied intelligence. For example, for an intelligent robot, it is necessary to accurately ground the affordance of an object and grasp it according to huma…

Cited by 0SourcePDFScholar
2025

HA-SCN: Learning Hierarchical Aligned Subtree Convolutional Networks for Graph Classification

IJCAI 2025

In this paper, we propose a Hierarchical Aligned Subtree Convolutional Network (HA-SCN) for graph classification. Our idea is to transform graphs of arbitrary sizes into fixed-sized aligned graphs and construct a normalized K-layer m-ary subtree for each node in the aligned graphs. By sliding convol

2025

HR-Extreme: A High-Resolution Dataset for Extreme Weather Forecasting

ICLR 2025poster

The application of large deep learning models in weather forecasting has led to significant advancements in the field, including higher-resolution forecasting and extended prediction periods exemplified by models such as Pangu and Fuxi. Despite these successes, previous research has largely been cha…

2025

High-Precision and High-Efficiency Trajectory Tracking for Excavators Based on Closed-Loop Dynamics

IROS 2025

The complex nonlinear dynamics of hydraulic excavators, such as time delays and control coupling, pose significant challenges to achieving high-precision trajectory tracking. Traditional control methods often fall short in such applications due to their inability to effectively handle these nonlinea

Cited by 0SourcecodeScholar
2025

How Sememic Components Can Benefit Link Prediction for Lexico-Semantic Knowledge Graphs?

EMNLP 2025

Link Prediction (LP) aims to predict missing triple information within a Knowledge Graph (KG). Existing LP methods have sought to improve the performance by integrating structural and textual information. However, for lexico-semantic KGs designed to document fine-grained sense distinctions, these ty

2025

How to Re-enable PDE Loss for Physical Systems Modeling Under Partial Observation

AAAI 2025technical

In science and engineering, machine learning techniques are increasingly successful in physical systems modeling (predicting future states of physical systems). Effectively integrating PDE loss as a constraint of system transition can improve the model's prediction by overcoming generalization issue…

2025

Human-guided robotic-assistance handheld continuum medical robot system

IROS 2025

Nowadays, laparoscopic surgery procedures face a trade-off between expensive, complex robotic systems and manual instruments with limited functionality. Fully robotic solutions offer precision but lack portability and intuitive control, while manual tools rely solely on the surgeon’s dexterity, limi

Cited by 0SourceScholar
2025

Hybrid Offline Passive Grammatical Inference and Online Planning for Non-Markovian Tasks

ICASSP 2025accepted

Planning in non-Markovian environments often requires inferring task structures, such as reward machines, through interactions with the environment. Traditional active grammatical inference methods, like Angluin’s L<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/19…

Cited by 0SourceScholar
2025

InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models

ICCV 2025poster

We present InfiniCube, a scalable and controllable method to generate unbounded and dynamic 3D driving scenes with high fidelity.Previous methods for scene generation are constrained either by their applicability to indoor scenes or by their lack of controllability.In contrast, we take advantage of…

Cited by 0SourcePDFScholar
2025

JointSwinUNETR: an Efficient Feature-enhanced Architecture for Small Intestine Cine MRI Segmentation

ICASSP 2025accepted

The Cine MRI of the small intestine is a dynamic magnetic resonance imaging technique used to observe and evaluate small intestine motility. It captures sequential images of the organ in motion over time through rapid imaging. The Transformer architecture is highly effective at capturing long-range…

Cited by 0SourceScholar
2025

LI-GS: Gaussian Splatting With LiDAR Incorporated for Accurate Large-Scale Reconstruction

RA-L 2025

Large-scale 3D reconstruction is critical in the field of robotics, and the potential of 3D Gaussian Splatting (3DGS) for achieving accurate object-level reconstruction has been demonstrated. However, ensuring geometric accuracy in outdoor and unbounded scenes remains a significant challenge. This s

Cited by 32SourceScholar
2025

LTRS: Improving Word Sense Disambiguation via Learning to Rank Senses

COLING 2025main

Word Sense Disambiguation (WSD) is a fundamental task critical for accurate semantic understanding. Conventional training strategies usually only consider predefined senses for target words and learn each of them from relatively limited instances, neglecting the influence of similar ones. To address…

Cited by 0SourcePDFScholar
2025

Language-Image Models with 3D Understanding

ICLR 2025poster

Multi-modal large language models (MLLMs) have shown incredible capabilities in a variety of 2D vision and language tasks. We extend MLLMs’ perceptual capabilities to ground and reason about images in 3-dimensional space. To that end, we first develop a large-scale pretraining dataset for 2D and 3D…

Cited by 15SourcePDFScholar
2025

Learning Temporally Consistent Video Depth from Video Diffusion Priors

CVPR 2025poster

This work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal in fostering temporal consistency. Therefore, we reformulate…

2025

Learning an Implicit Physics Model for Image-based Fluid Simulation

ICCV 2025poster

Humans possess an exceptional ability to imagine 4D scenes, encompassing both motion and 3D geometry, from a single still image. This ability is rooted in our accumulated observations of similar scenes and an intuitive understanding of physics. In this paper, we aim to replicate this capacity in neu…

2025

LiLoc: Lifelong Localization Using Adaptive Submap Joining and Egocentric Factor Graph

ICRA 2025

This paper proposes a versatile graph-based lifelong localization framework using LiDAR, LiLoc, which enhances its timeliness by maintaining a single central session while improves the accuracy through multi-modal factors between the central and subsidiary sessions. First, an adaptive submap joining

Cited by 3SourcecodeScholar
2025

LoGS: Visual Localization via Gaussian Splatting with Fewer Training Images

ICRA 2025

Visual localization involves estimating a query image's 6-DoF (degrees of freedom) camera pose, which is a fundamental component in various computer vision and robotic tasks. This paper presents LoGS, a vision-based localization pipeline utilizing the 3D Gaussian Splatting (GS) technique as scene re

Cited by 8SourcecodeScholar
2025

LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation models

ICLR 2025spotlight

Emerging 3D geometric foundation models, such as DUSt3R, offer a promising approach for in-the-wild 3D vision tasks. However, due to the high-dimensional nature of the problem space and scarcity of high-quality 3D data, these pre-trained models still struggle to generalize to many challenging circum…

2025

MSCI: Addressing CLIP&#039;s Inherent Limitations for Compositional Zero-Shot Learning

IJCAI 2025

Compositional Zero-Shot Learning (CZSL) aims to recognize unseen state-object combinations by leveraging known combinations. Existing studies basically rely on the cross-modal alignment capabilities of CLIP but tend to overlook its limitations in capturing fine-grained local features, which arise fr

2025

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation

CoRL 2025poster

Vision-Language Models (VLMs) have revolutionized artificial intelligence and robotics due to their commonsense reasoning capabilities. In robotic manipulation, VLMs are used primarily as high-level planners, but recent work has also studied their lower-level reasoning ability, which refers to makin…

Cited by 0SourceScholar
2025

Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions

NeurIPS 2025poster

The synthesis of realistic Martian landscape videos, essential for mission rehearsal and robotic simulation, presents unique challenges. These primarily stem from the scarcity of high-quality Martian data and the significant domain gap relative to terrestrial imagery. To address these challenges, we…

Cited by 0SourceScholar
2025

Microfluidics-Based Analysis of Controlled Mixing and Bubble Formation in Soda Solutions for Education

IROS 2025

This study describes a microfluidics experiment with ready classroom applications, designed to enhance students' understanding of fluid dynamics, controlled mixing, and bubble formation. The materials employed are safe and readily accessible, such as vinegar and baking soda, combined with PDMS micro

Cited by 0SourceScholar
2025

Model-Based Closed-Loop Control Algorithm for Stochastic Partial Differential Equation Control

IJCAI 2025

Neural operators have demonstrated promise in modeling and controlling systems governed by Partial Differential Equations (PDEs). Beyond PDEs, Stochastic Partial Differential Equations (SPDEs) play a critical role in modeling systems influenced by randomness, with applications in finance, physics, a

Cited by 0SourcePDFScholar
2025

Model-Free Offline Reinforcement Learning with Enhanced Robustness

ICLR 2025poster

Offline reinforcement learning (RL) has gained considerable attention for its ability to learn policies from pre-collected data without real-time interaction, which makes it particularly useful for high-risk applications. However, due to its reliance on offline datasets, existing works inevitably in…

Cited by 0SourcePDFScholar
2025

Mr. Virgil: Learning Multi-robot Visual-range Relative Localization

IROS 2025

Ultra-wideband (UWB)-vision fusion localization has achieved extensive applications in the domain of multiagent relative localization. The challenging matching problem between robots and visual detection renders existing methods highly dependent on identity-encoded hardware or delicate tuning algori

Cited by 0SourcecodeScholar
2025

Ms. NAMI: Multimodal Semantic Navigation on Relative Metric Intention Graph

ICRA 2025

Embodied navigation in unknown environments presents the significant challenge of integrating tasks with multimodal goals into a unified framework. In this paper, we propose the Multimodal Semantic Navigation on Relative Metric Intention Graph (Ms. NAMI), a framework that integrates various navigati

Cited by 0SourceScholar
2025

Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning

ICLR 2025poster

Vision foundation models, particularly the ViT family, have revolutionized image understanding by providing rich semantic features. However, despite their success in 2D comprehension, their abilities on grasping 3D spatial relationships are still unclear. In this work, we evaluate and enhance the 3D…

2025

Muscle-on-a-Chip: A Self-Healing Actuator Platform in Robotic Systems

IROS 2025

The regulation of muscle function is very important for tissue engineering and sports science. This paper presents a simple microfluidic chip platform and its control method to investigate the regulation of muscle function. By employing C2C12 cells as the model system for skeletal muscle research, t

Cited by 0SourceScholar
2025

N-ForGOT: Towards Not-forgetting and Generalization of Open Temporal Graph Learning

ICLR 2025poster

Temporal Graph Neural Networks (TGNNs) lay emphasis on capturing node interactions over time but often overlook evolution in node classes and dynamic data distributions triggered by the continuous emergence of new class labels, known as the open-set problem. This problem poses challenges for existin…

Cited by 0SourcePDFScholar
2025

Natural Humanoid Robot Locomotion with Generative Motion Prior

IROS 2025

Natural and lifelike locomotion remains a fundamental challenge for humanoid robots to interact with human society. However, previous methods either neglect motion naturalness or rely on unstable and ambiguous style rewards. In this paper, we propose a novel Generative Motion Prior (GMP) that provid

Cited by 10SourceScholar
2025

Neural Eulerian Scene Flow Fields

ICLR 2025poster

We reframe scene flow as the task of estimating a continuous space-time ordinary differential equation (ODE) that describes motion for an entire observation sequence, represented with a neural prior. Our method, EulerFlow, optimizes this neural prior estimate against several multi-observation recons…

Cited by 1SourcePDFScholar
2025

OmniRe: Omni Urban Scene Reconstruction

ICLR 2025spotlight

We introduce OmniRe, a comprehensive system for efficiently creating high-fidelity digital twins of dynamic real-world scenes from on-device logs. Recent methods using neural fields or Gaussian Splatting primarily focus on vehicles, hindering a holistic framework for all dynamic foregrounds demanded…

2025

Out-of-Distribution Detection with Prototypical Outlier Proxy

AAAI 2025technical

Out-of-distribution (OOD) detection is a crucial task for deploying deep learning models in the wild. One of the major challenges is that well-trained deep models tend to perform over-confidence on unseen test data. Recent research attempts to leverage real or synthetic outliers to mitigate the issu…

2025

PanopticSplatting: End-to-End Panoptic Gaussian Splatting

IROS 2025

Open-vocabulary panoptic reconstruction is a challenging task for simultaneous scene reconstruction and understanding. Recently, methods have been proposed for 3D scene understanding based on Gaussian splatting. However, these methods are multi-staged, suffering from the accumulated errors and the d

Cited by 2SourceScholar
2025

Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning

ICML 2025poster

Transfer reinforcement learning aims to derive a near-optimal policy for a target environment with limited data by leveraging abundant data from related source domains. However, it faces two key challenges: the lack of performance guarantees for the transferred policy, which can lead to undesired ac…

Cited by 0SourcePDFScholar
2025

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

ICLR 2025oral

Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. While Vision-Language Models (VLMs) have shown great promise in reasoning and task planning for embodied agents, their abil…

2025

RISED: Accurate and Efficient RGB-Colorized Mapping Using Image Selection and Point Cloud Densification

ICRA 2025

Recent advances in robotics have underscored the critical role of colorized point clouds in enhancing environmental perception accuracy. However, conventional multisensor fusion Simultaneous Localization and Mapping (SLAM) systems typically employ all available images indiscriminately for point clou

Cited by 1SourceScholar
2025

Reinforcement Learning for Adaptive Planner Parameter Tuning: A Perspective on Hierarchical Architecture

ICRA 2025

Automatic parameter tuning methods for planning algorithms, which integrate pipeline approaches with learning-based techniques, are regarded as promising due to their stability and capability to handle highly constrained environments. While existing parameter tuning methods have demonstrated conside

Cited by 2SourceScholar
2025

RoboVerse: A Unified Platform, Benchmark and Dataset for Scalable and Generalizable Robot Learning

RSS 2025poster

Data scaling and standardized evaluation benchmarks have driven remarkable advances in natural language processing and computer vision. However, in robotics, scaling up data and establishing evaluation protocols pose significant challenges. Directly collecting real-world data is inefficient and reso…

Cited by 0PDFScholar
2025

Robot Learning from Any Images

CoRL 2025poster

We introduce RoLA, a framework that transforms any in‑the‑wild image into an interactive, physics‑enabled robotic environment. Unlike previous methods, RoLA operates directly on a single image without requiring additional hardware or digital assets. Our framework democratizes robotic data generatio…

Cited by 0SourcecodeScholar
2025

SMART: Advancing Scalable Map Priors for Driving Topology Reasoning

ICRA 2025

Topology reasoning is crucial for autonomous driving as it enables comprehensive understanding of connec-tivity and relationships between lanes and traffic elements. While recent approaches have shown success in perceiving driving topology using vehicle-mounted sensors, their scalability is hindered

Cited by 9SourcecodeScholar
2025

SOLAR: Serendipity Optimized Language Model Aligned for Recommendation

EMNLP 2025

Recently, Large Language Models (LLMs) have shown strong potential in recommendation tasks due to their broad world knowledge and reasoning capabilities. However, applying them to serendipity-oriented recommendation remains challenging, mainly due to a domain gap of LLMs in modeling personalized use

2025

STORM: Spatio-TempOral Reconstruction Model For Large-Scale Outdoor Scenes

ICLR 2025poster

We present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often rely on per-scene optimization, dense observations across space and time, and strong motion supervision, resulting in le…

2025

Seeing the Wind from a Falling Leaf

NeurIPS 2025poster

A longstanding goal in computer vision is to model motions from videos, while the representations behind motions, i.e. the invisible physical interactions that cause objects to deform and move, remain largely unexplored. In this paper, we study how to recover the invisible forces from visual observa…

Cited by 0SourcecodeScholar
2025

Self-Sensing Liquid Crystal Elastomer Actuator with Magnetic-Thermal Synergy

IROS 2025

Fueled by the rapid evolution of robotics, the demand for intelligent and lightweight robotic systems continues to grow across industries. However, conventional designs often separate sensing and actuation, resulting in structural complexity and diminished reliability. While integrated sensor-actuat

Cited by 0SourceScholar
2025

Sparse Hierarchical LiDAR Bundle Adjustment for Online Collaborative Localization and Mapping

RA-L 2025

This letter presents a sparse hierarchical LiDAR bundle adjustment method for online multi-robot collaborative simultaneous localization and mapping (C-SLAM). The motivation behind this work is that the pose graph cannot directly reflect map inconsistencies. As a result, the map divergence across mu

Cited by 1SourceScholar
2025

Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting

CVPR 2025poster

The task of occupancy forecasting (OCF) involves utilizing past and present perception data to predict future occupancy states of autonomous vehicle surrounding environments, which is critical for downstream tasks such as obstacle avoidance and path planning. Existing 3D OCF approaches struggle to p…

2025

StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following

ACL 2025finding

Multi-turn instruction following capability constitutes a core competency of large language models (LLMs) in real-world applications. Existing evaluation benchmarks predominantly focus on fine-grained constraint satisfaction and domain-specific capability assessment, yet overlook the crucial structu…

2025

Symbolic Representation for Any-to-Any Generative Tasks

CVPR 2025poster

We propose a symbolic generative task description language and a corresponding inference engine that can represent arbitrary multimodal tasks as structured symbolic flows. Unlike conventional generative models, which rely on large-scale training and implicit neural representations to learn cross-mod…

2025

Thoughts Are All Over the Place: On the Underthinking of Long Reasoning Models

NeurIPS 2025spotlight

Long reasoning models (LRMs) such as OpenAI's o1 and DeepSeek's R1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where LRMs frequently switch between dif…

Cited by 0SourcecodeScholar
2025

Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training

NeurIPS 2025poster

Mixture-of-Experts (MoE) architectures within Large Reasoning Models (LRMs) have achieved impressive reasoning capabilities by selectively activating experts to facilitate structured cognitive processes. Despite notable advances, existing reasoning models often suffer from cognitive inefficiencies l…

Cited by 0SourceScholar
2025

UGNA-VPR: A Novel Training Paradigm for Visual Place Recognition Based on Uncertainty-Guided NeRF Augmentation

RA-L 2025

Visual place recognition (VPR) is crucial for robots to identify previously visited locations, playing an important role in autonomous navigation in both indoor and outdoor environments. However, most existing VPR datasets are limited to single-viewpoint scenarios, leading to reduced recognition acc

Cited by 1SourcecodeScholar
2025

UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

ACL 2025long

Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored. We introduce a benchmark to evaluate whether video-large language models (Video-LLMs) can naturally process continuous first-person v…

Cited by 0SourcePDFScholar
2025

Wavelet Diffusion Neural Operator

ICLR 2025poster

Simulating and controlling physical systems described by partial differential equations (PDEs) are crucial tasks across science and engineering. Recently, diffusion generative models have emerged as a competitive class of methods for these tasks due to their ability to capture long-term dependencies…

2025

𝒜3: Automatic Alignment Framework for Attributed Text Generation

ACL 2025long

Attributed text generation aims to enhance the reliability of content generated from large language models by providing citations for each claim, which thereby enables users to easily verify the correctness of the responses.However, the scarcity of high-quality training samples presents a significan…

2024

A Facile one-step injection novel composite sensor for robot tactile assistance

IROS 2024poster

Tactile information is the research hotspot of wearable flexible sensors due to its importance and complexity. With the innovation of wearable technology and robotics in healthcare, researchers are increasingly integrating wearable flexible sensors on the front end of robots to reproduce the hand ta…

Cited by 0SourceScholar
2024

A Fast Motion and Foothold Planning Framework for Legged Robots on Discrete Terrain

IROS 2024poster

Legged robot proved their capability to cross complex terrain in recent research, yet the autonomy of robots on discrete terrain still needs to be enhanced since it requires a full stack framework. This paper introduces a real-time motion and foothold planning framework tailored for legged robots na…

Cited by 0SourceScholar
2024

A Unified Principle of Pessimism for Offline Reinforcement Learning under Model Mismatch

NeurIPS 2024poster

In this paper, we address the challenges of offline reinforcement learning (RL) under model mismatch, where the agent aims to optimize its performance through an offline dataset that may not accurately represent the deployment environment. We identify two primary challenges under the setting: inaccu…

Cited by 0SourcePDFScholar
2024

Adapting for Calibration Disturbances: A Neural Uncalibrated Visual Servoing Policy

ICRA 2024poster

Visual servoing (VS) is a widely used technique in industries where there are hundreds of robots, but it requires accurate camera calibration including camera intrinsic and extrinsic parameters. However, it is labour-intensive to calibrate robots one-by-one in practical use. In this paper, we propos…

Cited by 1SourceScholar
2024

Advancing Virtual Reality Interaction: A Ring-Shaped Controller and Pose Tracking

ICRA 2024poster

Ensuring robust tracking of controllers’ movement is critical for human-robot interaction in virtual reality (VR) scenarios. This paper proposes a robust tracking algorithm based on a novel wearable ring-shaped controller equipped with an inertial measurement unit (IMU) and a light-emitting diode (L…

Cited by 0SourceScholar
2024

Augmenting Lane Perception and Topology Understanding with Standard Definition Navigation Maps

ICRA 2024poster

Autonomous driving has traditionally relied heavily on costly and labor-intensive High Definition (HD) maps, hindering scalability. In contrast, Standard Definition (SD) maps are more affordable and have worldwide coverage, offering a scalable alternative. In this work, we systematically explore the…

Cited by 34SourcecodeScholar
2024

BEV-ODOM: Reducing Scale Drift in Monocular Visual Odometry with BEV Representation

IROS 2024poster

Monocular visual odometry (MVO) is vital in autonomous navigation and robotics, providing a cost-effective and flexible motion tracking solution, but the inherent scale ambiguity in monocular setups often leads to cumulative errors over time. In this paper, we present BEV-ODOM, a novel MVO framework…

Cited by 1SourceScholar
2024

BEV2PR: BEV-Enhanced Visual Place Recognition with Structural Cues

IROS 2024

In this paper, we propose a new image-based visual place recognition (VPR) framework by exploiting the structural cues in bird’s-eye view (BEV) from a single monocular camera. The motivation arises from two key observations about place recognition methods based on both appearance and structure: 1) F

Cited by 4SourcecodeScholar
2024

Demonstration Data-Driven Parameter Adjustment for Trajectory Planning in Highly Constrained Environments

RA-L 2024

Trajectory planning in highly constrained environments is crucial for robotic navigation. Classical algorithms are widely used for their interpretability, generalization, and system robustness. However, these algorithms often require parameter retuning when adapting to new scenarios. To address this

Cited by 2SourceScholar
2024

Denoising Vision Transformers

ECCV 2024oral

"We study a crucial yet often overlooked issue inherent to Vision Transformers (ViTs): feature maps of these models exhibit grid-like artifacts (“Original features” in fig:teaser), which hurt the performance of ViTs in downstream dense prediction tasks such as semantic segmentation, depth prediction…

2024

DiffPhyCon: A Generative Approach to Control Complex Physical Systems

NeurIPS 2024poster

Controlling the evolution of complex physical systems is a fundamental task across science and engineering. Classical techniques suffer from limited applicability or huge computational costs. On the other hand, recent deep learning and reinforcement learning-based approaches often struggle to optim…

2024

Disambiguate Words like Composing Them: A Morphology-Informed Approach to Enhance Chinese Word Sense Disambiguation

ACL 2024long

In parataxis languages like Chinese, word meanings are highly correlated with morphological knowledge, which can help to disambiguate word senses. However, in-depth exploration of morphological knowledge in previous word sense disambiguation (WSD) methods is still lacking due to the absence of publi…

2024

DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features

NeurIPS 2024poster

We propose DistillNeRF, a self-supervised learning framework addressing the challenge of understanding 3D environments from limited 2D observations in outdoor autonomous driving scenes. Our method is a generalizable feedforward model that predicts a rich neural scene representation from sparse, sing…

2024

Driving Everywhere with Large Language Model Policy Adaptation

CVPR 2024poster

Adapting driving behavior to new environments customs and laws is a long-standing problem in autonomous driving precluding the widespread deployment of autonomous vehicles (AVs). In this paper we present LLaDA a simple yet powerful tool that enables human drivers and autonomous vehicles alike to dri…

2024

Dropout Mixture Low-Rank Adaptation for Visual Parameters-Efficient Fine-Tuning

ECCV 2024poster

"Parameter-efficient fine-tuning methods adjust a small subset of parameters in large models, achieving performance comparable to or even surpassing that of models fine-tuned with the full parameter set, and significantly reducing the time and computational costs associated with the fine-tuning proc…

2024

EDA: Evolving and Distinct Anchors for Multimodal Motion Prediction

AAAI 2024technical

Motion prediction is a crucial task in autonomous driving, and one of its major challenges lands in the multimodality of future behaviors. Many successful works have utilized mixture models which require identification of positive mixture components, and correspondingly fall into two main lines: pre…

2024

Efficient Global Trajectory Planning for Multi-robot System with Affinely Deformable Formation

IROS 2024

Global trajectory planning is crucial for long-range formation navigation tasks of multi-robot systems in efficiency improvement and energy saving, whose main challenges are the joint space constraints of the whole team and the long-range deployment. To overcome the above difficulties, we reformulat

Cited by 1SourceScholar
2024

EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-Supervision

ICLR 2024poster

We present EmerNeRF, a simple yet powerful approach for learning spatial-temporal representations of dynamic driving scenes. Grounded in neural fields, EmerNeRF simultaneously captures scene geometry, appearance, motion, and semantics via self-bootstrapping. EmerNeRF hinges upon two core components:…

2024

Enhancing Closed-Loop Performance in Learning-Based Vehicle Motion Planning by Integrating Rule-Based Insights

RA-L 2024

This letter introduces an innovative vehicle motion planning method that leverages the integration of rule-based insights to significantly improve closed-loop performance within a learning-based framework. We first employ rule-based methods to heuristically search and generate a diverse set of traje

Cited by 2SourceScholar
2024

Explicit Interaction for Fusion-Based Place Recognition

IROS 2024poster

Fusion-based place recognition is an emerging technique jointly utilizing multi-modal perception data, to recognize previously visited places in GPS-denied scenarios for robots and autonomous vehicles. Recent fusion-based place recognition methods combine multi-modal features in implicit manners. Wh…

Cited by 2SourcecodeScholar
2024

Fast Cross-Modality Knowledge Transfer via a Contextual Autoencoder Transformation

ICASSP 2024accepted

Cross-modality knowledge transfer aims to apply knowledge learned in the source modality to the target modality. It is more challenging than the general knowledge transfer task because of the aggravated modality shift problem due to introducing heterogeneous data. This paper proposes a novel fast cr…

Cited by 0SourceScholar
2024

HC-GAE: The Hierarchical Cluster-based Graph Auto-Encoder for Graph Representation Learning

NeurIPS 2024poster

Graph Auto-Encoders (GAEs) are powerful tools for graph representation learning. In this paper, we develop a novel Hierarchical Cluster-based GAE (HC-GAE), that can learn effective structural characteristics for graph data analysis. To this end, during the encoding process, we commence by utilizing…

Cited by 1SourcePDFScholar
2024

HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting

CVPR 2024poster

Holistic understanding of urban scenes based on RGB images is a challenging yet important problem. It encompasses understanding both the geometry and appearance to enable novel view synthesis parsing semantic labels and tracking moving objects. Despite considerable progress existing approaches often…

2024

Harmonic Retrieval for Non-Circular Coherent Signals via Double Decoupled Atomic Norm Minimization

ICASSP 2024accepted

This paper studies super-resolution harmonic retrieval for strictly non-circular coherent signals. We develop gridless sparse representations of both their covariance and pseudo-covariance matrices over a common matrix-form atom set. This enables the decoupled atomic norm minimization (D-ANM) techni…

Cited by 0SourceScholar
2024

Large Spatial Model: End-to-end Unposed Images to Semantic 3D

NeurIPS 2024poster

Reconstructing and understanding 3D structures from a limited number of images is a classical problem in computer vision. Traditional approaches typically decompose this task into multiple subtasks, involving several stages of complex mappings between different data representations. For example, den…

2024

Learning 3D-aware GANs from Unposed Images with Template Feature Field

ECCV 2024oral

"Collecting accurate camera poses of training images has been shown to well serve the learning of 3D-aware generative adversarial networks (GANs) yet can be quite expensive in practice. This work targets learning 3D-aware GANs from unposed images, for which we propose to perform on-the-fly pose esti…

Cited by 1SourcePDFScholar
2024

Learning Hierarchical Graph-Based Policy for Goal-Reaching in Unknown Environments

RA-L 2024

Goal-reaching in unknown environments is one of the essential tasks in robot applications. Large-scale perception and long-horizon decision-making are the keys to solving this task as the operation scope expands or complexity rises. Existing navigation methods may suffer from degraded performance in

Cited by 6SourceScholar
2024

Learning the Inverse Kinematics of Magnetic Continuum Robot for Teleoperated Navigation

IROS 2024

Magnetic continuum robots are subject to external magnetic fields and deformed remotely, simplifying the robot’s transmission mechanism and providing it with significant potential for miniaturization and operational flexibility. However, modeling magnetic field distribution generated by permanent ma

Cited by 2SourceScholar
2024

Let Occ Flow: Self-Supervised 3D Occupancy Flow Prediction

CoRL 2024poster

Accurate perception of the dynamic environment is a fundamental task for autonomous driving and robot systems. This paper introduces Let Occ Flow, the first self-supervised work for joint 3D occupancy and occupancy flow prediction using only camera inputs, eliminating the need for 3D annotations. Ut…

Cited by 10SourceScholar
2024

Memorize What Matters: Emergent Scene Decomposition from Multitraverse

NeurIPS 2024spotlight

Humans naturally retain memories of permanent elements, while ephemeral moments often slip through the cracks of memory. This selective retention is crucial for robotic perception, localization, and mapping. To endow robots with this capability, we introduce 3D Gaussian Mapping (3DGM), a self-superv…

2024

Morpheme Sense Disambiguation: A New Task Aiming for Understanding the Language at Character Level

COLING 2024main

Morphemes serve as a strong linguistic feature to capture lexical semantics, with higher coverage than words and more natural than sememes. However, due to the lack of morpheme-informed resources and the expense of manual annotation, morpheme-enhanced methods remain largely unexplored in Computation…

2024

NGEL-SLAM: Neural Implicit Representation-based Global Consistent Low-Latency SLAM System

ICRA 2024poster

Neural implicit representations have emerged as a promising solution for providing dense geometry in Simultaneous Localization and Mapping (SLAM). However, existing methods in this direction fall short in terms of global consistency and low latency. This paper presents NGEL-SLAM to tackle the above…

Cited by 29SourceScholar
2024

Non-Asymptotic Analysis for Single-Loop (Natural) Actor-Critic with Compatible Function Approximation

ICML 2024poster

Actor-critic (AC) is a powerful method for learning an optimal policy in reinforcement learning, where the critic uses algorithms, e.g., temporal difference (TD) learning with function approximation, to evaluate the current policy and the actor updates the policy along an approximate gradient direct…

Cited by 12SourcePDFScholar
2024

OTVIC: A Dataset with Online Transmission for Vehicle-to-Infrastructure Cooperative 3D Object Detection

IROS 2024poster

Vehicle-to-infrastructure cooperative 3D object detection (VIC3D) is a task that leverages both vehicle and roadside sensors to jointly perceive the surrounding environment. However, considering the high speed of vehicles, the real-time requirements, and the limitations of communication bandwidth, r…

Cited by 1SourceScholar
2024

Online Trajectory Deformation and Tracking for Self-entanglement-free Differential-Driven Robots

ICRA 2024poster

This paper introduces an optimisation-based trajectory deformation and tracking algorithm for tethered differential-driven mobile robots. The motivation of this work is to generate self-entanglement-free (SEF) commands for a tethered differential-driven robot to track a path. Whilst existing path pl…

Cited by 0SourceScholar
2024

Optimal Non-Redundant Manipulator Surface Coverage with Rank-Deficient Manipulability Constraints

RSS 2024poster

A generalised solver for the manipulator non-revisiting coverage path planning (NCPP) problem is proposed in this paper. Nonlinear manipulator kinematics and the imposition of task-specific constraints dictate that applying conventional coverage path planning (CPP) solutions based on 2D template mat…

Cited by 0SourcePDFScholar
2024

PARA-Drive: Parallelized Architecture for Real-time Autonomous Driving

CVPR 2024poster

Recent works have proposed end-to-end autonomous vehicle (AV) architectures comprised of differentiable modules achieving state-of-the-art driving performance. While they provide advantages over the traditional perception-prediction-planning pipeline (e.g. removing information bottlenecks between co…

Cited by 37SourcePDFScholar
2024

PEP: Policy-Embedded Trajectory Planning for Autonomous Driving

RA-L 2024

Autonomous driving demands proficient trajectory planning to ensure safety and comfort. This letter introduces Policy-Embedded Planner (PEP), a novel framework that enhances closed-loop performance of imitation learning (IL) based planners by embedding a neural policy for sequential ego pose generat

Cited by 8SourceScholar
2024

PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic Reconstruction

IROS 2024

Panoptic reconstruction is a challenging task in 3D scene understanding. However, most existing methods heavily rely on pre-trained semantic segmentation models and known 3D object bounding boxes for 3D panoptic segmentation, which is not available for in-the-wild scenes. In this paper, we propose a

Cited by 8SourceScholar
2024

Parallelized Spatiotemporal Slot Binding for Videos

ICML 2024poster

While modern best practices advocate for scalable architectures that support long-range interactions, object-centric models are yet to fully embrace these architectures. In particular, existing object-centric models for handling sequential inputs, due to their reliance on RNN-based implementation, s…

Cited by 0SourcePDFScholar
2024

PreSight: Enhancing Autonomous Vehicle Perception with City-Scale NeRF Priors

ECCV 2024poster

"Autonomous vehicles rely extensively on perception systems to navigate and interpret their surroundings. Despite significant advancements in these systems recently, challenges persist under conditions like occlusion, extreme lighting, or in unfamiliar urban areas. Unlike these systems, humans do no…

2024

Q-SLAM: Quadric Representations for Monocular SLAM

CoRL 2024poster

In this paper, we reimagine volumetric representations through the lens of quadrics. We posit that rigid scene components can be effectively decomposed into quadric surfaces. Leveraging this assumption, we reshape the volumetric representations with million of cubes by several quadric planes, which…

Cited by 6SourceScholar
2024

QBMK: Quantum-based Matching Kernels for Un-attributed Graphs

ICML 2024spotlight

In this work, we develop a new Quantum-based Matching Kernel (QBMK) for un-attributed graphs, by computing the kernel-based similarity between the quantum Shannon entropies of aligned vertices through the Continuous-time Quantum Walk (CTQW). The theoretical analysis reveals that the proposed QBMK ke…

Cited by 0SourcePDFScholar
2024

RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation

CoRL 2024poster

This work proposes a retrieve-and-transfer framework for zero-shot robotic manipulation, dubbed RAM, featuring generalizability across various objects, environments, and embodiments. Unlike existing approaches that learn manipulation from expensive in-domain demonstrations, RAM capitalizes on a retr…

Cited by 26SourcecodeScholar
2024

RGBD-based Image Goal Navigation with Pose Drift: A Topo-metric Graph based Approach

ICRA 2024poster

Image-goal navigation in unknown environments with sensor error is of considerable difficulty for autonomous robots. In this paper, we propose a drift-resisting topo-metric graph to map the environment and localize the robot using only relative poses. The error-sharing mechanism under this represent…

Cited by 1SourceScholar
2024

SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving

IROS 2024

Monocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D scene understanding is semantic scene completion (SSC), which jointly completes semantic information and geometric details

Cited by 90SourcecodeScholar
2024

Scale Disparity of Instances in Interactive Point Cloud Segmentation

IROS 2024poster

Interactive point cloud segmentation has become a pivotal task for understanding 3D scenes, enabling users to guide segmentation models with simple interactions such as clicks, therefore significantly reducing the effort required to tailor models to diverse scenarios and new categories. However, in…

Cited by 2SourceScholar
2024

Semantics-aware Motion Retargeting with Vision-Language Models

CVPR 2024poster

Capturing and preserving motion semantics is essential to motion retargeting between animation characters. However most of the previous works neglect the semantic information or rely on human-designed joint-level representations. Here we present a novel Semantics-aware Motion reTargeting (SMT) metho…

Cited by 5SourcePDFScholar
2024

Social Physics Informed Diffusion Model for Crowd Simulation

AAAI 2024technical

Crowd simulation holds crucial applications in various domains, such as urban planning, architectural design, and traffic arrangement. In recent years, physics-informed machine learning methods have achieved state-of-the-art performance in crowd simulation but fail to model the heterogeneity and mul…

2024

Soft Hybrid Actuated Hierarchical Bronchoscope Robot for Deep Lung Examination

RA-L 2024

Lungdiseases are becoming one of the world's most serious health issues. Soft bronchoscope robots can achieve safe and controllable lung navigation, which will be crucial for the future early examination of lung diseases. However, due to the single driving method, the large size, and insufficient fl

Cited by 10SourceScholar
2024

Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving

CoRL 2024poster

The autonomous driving industry is increasingly adopting end-to-end learning from sensory inputs to minimize human biases in system design. Traditional end-to-end driving models, however, suffer from long-tail events due to rare or unseen inputs within their training distributions. To address this,…

Cited by 13SourceScholar
2024

Towards More Realistic Chinese Spell Checking with New Benchmark and Specialized Expert Model

COLING 2024main

Large Language Models (LLMs) hold considerable promise for artificial general intelligence, given their intrinsic abilities to accomplish a wide range of open-domain tasks either independently or in tandem with specialized expert models. However, despite these capabilities, the performance of LLMs h…

2024

Tree-based Representation of Locally Shortest Paths for 2D k-Shortest Non-homotopic Path Planning

ICRA 2024poster

A novel algorithm to solve the 2D k-shortest non-homotopic path planning (k-SNPP) task is proposed in this paper. The task is of practical significance as a sub-module for higherlevel planning and scheduling tasks, and is gaining increasing attention and focus in recent years. There have existed alg…

Cited by 2SourceScholar
2024

VIVO: A Visual-Inertial-Velocity Odometry with Online Calibration in Challenging Condition

IROS 2024poster

State estimation is a central component of autonomous navigation. To date, many methods presented have a disruptive potential for application, such as visual-inertial odometry (VIO), wheel and leg odometry (for short, body odometry). However, most of them are prone to fail in some challenging condit…

Cited by 0SourceScholar
2024

Vertebrae-based Global X-ray to CT Registration for Thoracic Surgeries

IROS 2024poster

X-ray to CT registration is an essential technique to provide on-site guidance for clinicians and medical robots by aligning preoperative information with intraoperative images. Current methods focus on local registration with small capture ranges and necessitate a manual initial alignment before pr…

Cited by 0SourcecodeScholar
2024

ν-DBA: Neural Implicit Dense Bundle Adjustment Enables Image-Only Driving Scene Reconstruction

IROS 2024poster

The joint optimization of the sensor trajectory and 3D map is a crucial characteristic of bundle adjustment (BA), essential for autonomous driving. This paper presents ν-DBA, a novel framework implementing geometric dense bundle adjustment (DBA) using 3D neural implicit surfaces for map parametrizat…

Cited by 0SourceScholar
2023

A Hyper-Network Based End-to-End Visual Servoing With Arbitrary Desired Poses

RA-L 2023

Recently, several works achieve end-to-end visual servoing (VS) for robotic manipulation by replacing traditional controller with differentiable neural networks, but lose the ability to servo arbitrary desired poses. This letter proposes a differentiable architecture for arbitrary pose servoing: a h

Cited by 8SourceScholar
2023

A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter

ICRA 2023poster

We focus on the task of language-conditioned grasping in clutter, in which a robot is supposed to grasp the target object based on a language instruction. Previous works separately conduct visual grounding to localize the target object, and generate a grasp for that object. However, these works requ…

Cited by 49SourcecodeScholar
2023

A Robust and Constrained Multi-Agent Reinforcement Learning Electric Vehicle Rebalancing Method in AMoD Systems

IROS 2023poster

Electric vehicles (EVs) play critical roles in autonomous mobility-on-demand (AMoD) systems, but their unique charging patterns increase the model uncertainties in AMoD systems (e.g. state transition probability). Since there usually exists a mismatch between the training and test/true environments,…

Cited by 35SourceScholar
2023

A Two-Stage Based Social Preference Recognition in Multi-Agent Autonomous Driving System

IROS 2023poster

Multi-Agent Reinforcement Learning (MARL) has become a promising solution for constructing a multi-agent autonomous driving system (MADS) in complex and dense scenarios. But most methods consider agents acting selfishly, which leads to conflict behaviors. Some existing works incorporate the concept…

Cited by 2SourceScholar
2023

An Efficient Multi-solution Solver for the Inverse Kinematics of 3-Section Constant-Curvature Robots

RSS 2023poster

Piecewise constant curvature is a popular kinematics framework for continuum robots. Computing the model parameters from the desired end pose, known as the inverse kinematics problem, is fundamental in manipulation, tracking and planning tasks. In this paper, we propose an efficient multi-solution s…

Cited by 7SourcePDFScholar