← Search

Jun Ma

97 accepted papers

2026

A Reconfigured Wheel-Legged Robot for Enhanced Steering and Adaptability

RA-L 2026

Wheel-legged robots integrate leg agility on rough terrain with wheel efficiency on flat ground. However, most existing designs do not fully capitalize on the benefits of both legged and wheeled structures, which limits overall system flexibility and efficiency. We present FLORES, a novel wheel-legg

Cited by 1SourcecodeScholar
2026

ApexNav: An Adaptive Exploration Strategy for Zero-Shot Object Navigation with Target-Centric Semantic Fusion

ICRA 2026poster

Navigating unknown environments to find a target object is a significant challenge. While semantic information is crucial for navigation, relying solely on it for decision-making may not always be efficient, especially in environments with weak semantic cues. Additionally, many methods are susceptib…

2026

Beyond Distribution Estimation: Simplex Anchored Structural Inference Towards Universal Semi-supervised Learning

ICML 2026poster

Semi-supervised learning (SSL) faces significant challenges in realistic scenarios where labeled data is extremely scarce and unlabeled data follows unknown, arbitrary distributions. We formalize this critical yet under-explored paradigm as Universal Semi-supervised Learning (UniSSL). Existing metho…

Cited by 0SourceScholar
2026

CoPlanner: An Interactive Motion Planner with Contingency-Aware Diffusion for Autonomous Driving

ICRA 2026poster

Accurate trajectory prediction and motion planning are crucial for autonomous driving systems to navigate safely in complex, interactive environments characterized by multimodal uncertainties. However, current generation-then-evaluation frameworks typically construct multiple plausible trajectory hy…

2026

CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving

CVPR 2026

The pursuit of autonomous agents with predictive cognitive world models is hindered by a fundamental flaw in current vision-language models (VLMs): they lack cognitive inertia. Operating on isolated snapshots, these models cannot form a temporally coherent world view, leading to erratic decision jit

Cited by 0SourceScholar
2026

DRM-Net: Explicit Residual Modelling with Subaquatic Multi-Scale Context Fusion for Underwater Image Enhancement

AAAI 2026technical

Clear and high-quality underwater images are essential for marine applications, including autonomous navigation, ecological monitoring, and infrastructure inspection. However, underwater images typically suffer from severe colour distortion, low contrast, and diminished structural visibility due to

Cited by 0SourcePDFScholar
2026

Domain Adaptation Guided Infrared and Visible Image Fusion

AAAI 2026technical

Infrared and Visible Image Fusion (IVIF) integrates complementary information from distinct modalities to enhance image quality. However, the effectiveness declines under unseen conditions such as novel weather or scenes, due to domain shifts primarily from variations of data distribution in the vis

Cited by 0SourcePDFScholar
2026

Embracing Bulky Objects with Humanoid Robots: Whole-Body Manipulation with Reinforcement Learning

ICRA 2026poster

Whole-body manipulation (WBM) for humanoid robots presents a promising approach for executing embracing tasks involving bulky objects, where traditional grasping relying on end-effectors only remains limited in such scenarios due to inherent stability and payload constraints. This paper introduces a…

2026

GDP: Enhancing End-To-End Autonomous Driving with Goal-Driven Planner

ICRA 2026poster

End-to-end (E2E) autonomous driving has emerged as a promising paradigm with the pervasive power of model architectures and the availability of large-scale driving datasets. Despite tremendous efforts in recent research, most E2E driving frameworks rely on rather general driving commands, such as "G…

Cited by 0Scholar
2026

GUIDE: A Diffusion-Based Autonomous Robot Exploration Framework Using Global Graph Inference

ICRA 2026poster

Autonomous exploration in structured and complex indoor environments remains a challenging task, as existing methods often struggle to appropriately model unobserved space and plan globally efficient paths. To address these limitations, we propose GUIDE, a novel exploration framework that synergisti…

2026

HATIR: Heat-Aware Diffusion for Turbulent Infrared Video Super-Resolution

AAAI 2026technical

Infrared video has been of great interest in visual tasks under challenging environments, but often suffers from severe atmospheric turbulence and compression degradation. Existing video super-resolution (VSR) methods either neglect the inherent modality gap between infrared and visible images or fa

Cited by 0SourcePDFScholar
2026

IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs

ICLR 2026poster

Existing evaluation frameworks for Multimodal Large Language Models (MLLMs) primarily focus on image reasoning or general video understanding tasks, largely overlooking the significant role of image context in video comprehension. To bridge this gap, we propose \textbf{IV-Bench}, the first comprehen…

Cited by 0SourcecodeScholar
2026

Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning

AAAI 2026technical

Zero-shot stance detection (ZSSD) seeks to determine the stance of text toward previously unseen targets, a task critical for analyzing dynamic and polarized online discourse with limited labeled data. While large language models (LLMs) offer zero-shot capabilities, prompting-based approaches often

Cited by 0SourcePDFScholar
2026

LISA: Language-guided Interference-aware Spatial-Frequency Attention for Driver Gaze Estimation

IJCAI 2026

Driver gaze estimation serves as a fundamental metric for evaluating driver attentiveness in modern monitoring systems. Beyond being vulnerable to sudden lighting changes and sensor noise, spatial-domain models struggle to disentangle authentic gaze cues from irrelevant visual attributes. In this pa

Cited by 0Scholar
2026

ManiVID-3D: Generalizable View-Invariant Reinforcement Learning for Robotic Manipulation via Disentangled 3D Representations

RA-L 2026

Deploying visual reinforcement learning (RL) policies in real-world manipulation is often hindered by camera viewpoint changes. A policy trained from a fixed front-facing camera may fail when the camera is shifted-an unavoidable situation in real-world settings where sensor placement is hard to mana

Cited by 5SourceScholar
2026

Occlusion-Aware Consistent Model Predictive Control for Robot Navigation in Occluded Obstacle-Dense Environments

ICRA 2026poster

Ensuring safety and motion consistency for robot navigation in occluded, obstacle-dense environments is a critical challenge. In this context, this study presents an occlusion-aware Consistent Model Predictive Control (CMPC) strategy. To account for the occluded obstacles, it incorporates adjustable…

2026

One Flow Fits All! A Scale-Aware Generative Framework for Diverse Data

IJCAI 2026

Real-world systems increasingly require coherent reasoning and generation over diverse data modalities simultaneously. Current generative frameworks rely on complex, multi-stage training, resulting in low efficiency due to iterative inference and high computational cost. They also struggle with unif

Cited by 0Scholar
2026

Online Trajectory Optimization for Arbitrary-Shaped Mobile Robots via Polynomial Separating Hypersurfaces

RA-L 2026

An emerging class of trajectory optimization methods enforces collision avoidance by jointly optimizing the robot's configuration and a separating hyperplane. However, as linear separators only apply to convex sets, these methods require convex approximations of both the robot and obstacles, which b

Cited by 1SourceScholar
2026

Rethinking the Practicality of Vision-Language-Action Model: A Comprehensive Benchmark and an Improved Baseline

ICRA 2026poster

Vision-Language-Action (VLA) models have emerged as a generalist robotic agent. However, existing VLAs are hindered by excessive parameter scales, prohibitive pre-training requirements, and limited applicability to diverse embodiments. To improve the practicality of VLAs, we propose a comprehensive …

2026

SEG-Parking: Towards Safe, Efficient, and Generalizable Autonomous Parking Via End-To-End Offline Reinforcement Learning

ICRA 2026poster

Autonomous parking is a critical component for achieving safe and efficient urban autonomous driving. However, unstructured environments and dynamic interactions pose significant challenges to autonomous parking tasks. To address this problem, we propose SEG-Parking, a novel end-to-end offline reinf…

2026

SVP: Improving Vision-Language-Action Models with Dual Stochastic Visual Prompting

ICRA 2026poster

Vision-Language-Action (VLA) models, such as OpenVLA, hold the promise of generalist robots, yet their performance is often impaired by distracted attention, which we identify as a manifestation of shortcut learning. We posit that the solution lies not in architectural modifications, but in a new tr…

Cited by 0Scholar
2026

Samples Are Not Equal: A Sample Selection Approach for Deep Clustering

ICLR 2026poster

Deep clustering has recently achieved remarkable progress across various domains. However, existing clustering methods typically treat all samples equally, neglecting the inherent differences in their feature patterns and learning states. Such redundant learning often drives models to overemphasize…

Cited by 0SourcecodeScholar
2026

Semantic-LiDAR-Inertial-Wheel Odometry Fusion for Robust Localization in Large-Scale Dynamic Environments

ICRA 2026poster

Reliable, drift-free global localization presents significant challenges yet remains crucial for autonomous navigation in large-scale dynamic environments. In this paper, we introduce a tightly-coupled Semantic-LiDAR-Inertial-Wheel Odometry fusion framework, which is specifically designed to provide…

2026

Semi-SMD: Semi-Supervised Metric Depth Estimation Via Surrounding Cameras for Autonomous Driving

ICRA 2026poster

In this paper, we introduce Semi-SMD, a novel metric depth estimation framework tailored for surrounding cameras equipment in autonomous driving. In this work, the input data consists of adjacent surrounding frames and camera parameters. We propose a unified spatial-temporal-semantic fusion module t…

2026

Toward Real-World High-Precision Image Matting and Segmentation

AAAI 2026technical

High-precision scene parsing tasks, including image matting and dichotomous segmentation, aim to accurately predict masks with extremely fine details (such as hair). Most existing methods focus on salient, single foreground objects. While interactive methods allow for target adjustment, their class-

Cited by 0SourcePDFScholar
2026

Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset

CVPR 2026

Infrared image super-resolution (IISR) under real-world conditions is a practically significant yet rarely addressed task. Pioneering works are often trained and evaluated on simulated datasets or neglect the intrinsic differences between infrared and visible imaging. In practice, however, real infr

Cited by 0SourcecodeScholar
2026

Uncertainty-Aware Spatial-Frequency Registration and Fusion for Infrared and Visible Images

IJCAI 2026

Infrared and Visible Image Fusion (IVIF) has shown promise in visual tasks under challenging environments, but fusion under unregistered conditions faces inherent misalignments. Current studies to solve them either predict the deformation parameters coarse-to-fine (i.e., coarse registration and fine

Cited by 0Scholar
2026

VLM-E2E: Enhancing End-To-End Autonomous Driving with Multimodal Driver Attention Fusion

ICRA 2026poster

Human drivers adeptly navigate complex scenarios by utilizing rich attentional semantics, but the current autonomous systems struggle to replicate this ability, as they often lose critical semantic information when converting 2D observations into 3D space. In this sense, it hinders their effective d…

2025

A Simple Data Augmentation for Feature Distribution Skewed Federated Learning

CVPR 2025poster

Federated Learning (FL) facilitates collaborative learning among multiple clients in a distributed manner and ensures the security of privacy. However, its performance inevitably degrades with non-Independent and Identically Distributed (non-IID) data. In this paper, we focus on the feature distribu…

2025

Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector

NeurIPS 2025poster

LLM-as-a-Judge has emerged as a promising tool for automatically evaluating generated outputs, but its reliability is often undermined by potential biases in judgment. Existing efforts to mitigate these biases face key limitations: in-context learning-based methods fail to address rooted biases due…

Cited by 0SourceScholar
2025

ApexNAV: An Adaptive Exploration Strategy for Zero-Shot Object Navigation With Target-Centric Semantic Fusion

RA-L 2025

Navigating unknown environments to find a target object is a significant challenge. While semantic information is crucial for navigation, relying solely on it for decision-making may not always be efficient, especially in environments with weak semantic cues. Additionally, many methods are susceptib

Cited by 24SourceScholar
2025

DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution

CVPR 2025poster

Infrared imaging is essential for autonomous driving and robotic operations as a supportive modality due to its reliable performance in challenging environments. Despite its popularity, the limitations of infrared cameras, such as low spatial resolution and complex degradations, consistently challen…

2025

DuLoc: Life-Long Dual-Layer Localization in Changing and Dynamic Expansive Scenarios

IROS 2025

LiDAR-based localization serves as a critical component in autonomous systems, yet existing approaches face persistent challenges in balancing repeatability, accuracy, and environmental adaptability. Traditional point cloud registration methods relying solely on offline maps often exhibit limited ro

Cited by 0SourceScholar
2025

FERMI: Flexible Radio Mapping with a Hybrid Propagation Model and Scalable Autonomous Data Collection

RSS 2025poster

Communication is fundamental for multi-robot collaboration, with accurate radio mapping playing a crucial role in predicting signal strength between robots. However, modeling radio signal propagation in large and occluded environments is challenging due to complex interactions between signals and ob…

Cited by 0PDFScholar
2025

FRTree Planner: Robot Navigation in Cluttered and Unknown Environments With Tree of Free Regions

RA-L 2025

In this work, we present FRTree planner, a novel robot navigation framework that leverages a tree structure of free regions, specifically designed for navigation in cluttered and unknown environments with narrow passages. The framework continuously incorporates real-time perceptive information to id

Cited by 6SourceScholar
2025

FisheyeDepth: A Real Scale Self-Supervised Depth Estimation Model for Fisheye Camera

ICRA 2025

Accurate depth estimation is crucial for 3D scene comprehension in robotics and autonomous vehicles. Fisheye cameras, known for their wide field of view, have inherent geometric benefits. However, their use in depth estimation is restricted by a scarcity of ground truth data and image distortions. W

Cited by 7SourcecodeScholar
2025

G2-SDF: Geometry-Guided Neural Signed Distance Fields for Scalable and Detailed Reconstruction

RA-L 2025

Effcient reconstruction methods, particularly capable of providing detailed information on obstacle distances across diverse environments, are crucial for effective robot motion planning. In this context, neural Signed Distance Fields (SDFs) offer a powerful solution by learning implicit representat

Cited by 1SourceScholar
2025

GDTS: Goal-Guided Diffusion Model with Tree Sampling for Multi-Modal Pedestrian Trajectory Prediction

IROS 2025

Accurate prediction of pedestrian trajectories is crucial for improving the safety of autonomous driving. However, this task is generally nontrivial due to the inherent stochasticity of human motion, which naturally requires the predictor to generate multi-modal prediction. Previous works leverage v

Cited by 2SourceScholar
2025

GS-LIVM: Real-Time Photo-Realistic LiDAR-Inertial-Visual Mapping with Gaussian Splatting

ICCV 2025poster

In this paper, we introduce GS-LIVM, a real-time photo-realistic LiDAR-Inertial-Visual mapping framework with Gaussian Splatting tailored for outdoor scenes. Compared to existing methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), our approach enables real-time photo-rea…

2025

Interactive Navigation for Legged Manipulators with Learned Arm-Pushing Controller

IROS 2025

Interactive navigation is crucial in scenarios where proactively interacting with objects can yield shorter paths, thus significantly improving traversal efficiency. Existing methods primarily focus on using the robot body to relocate obstacles during navigation. However, they prove ineffective in n

Cited by 5SourcecodeScholar
2025

Local Reactive Control for Mobile Manipulators With Whole-Body Safety in Complex Environments

RA-L 2025

Mobile manipulators typically encounter significant challenges in navigating narrow, cluttered environments due to their high-dimensional state spaces and complex kinematics. While reactive methods excel in dynamic settings, they struggle to efficiently incorporate complex, coupled constraints acros

Cited by 5SourceScholar
2025

MF-BERT: A Siamese Pre-training Framework for Motion Forecasting

ICASSP 2025accepted

Accurately predicting the future motions of traffic agents is essential for autonomous systems. Despite the significant success of existing motion forecasting methods based on supervised learning, they still exhibit two main limitations. First, when annotated data for a scene is limited, these metho…

Cited by 0SourceScholar
2025

Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models

IROS 2025

Diffusion models have been widely employed in the field of 3D manipulation due to their efficient capability to learn distributions, allowing for precise prediction of action trajectories. However, diffusion models typically rely on large parameter UNet backbones as policy networks, which can be cha

Cited by 22SourcecodeScholar
2025

MedRAX: Medical Reasoning Agent for Chest X-ray

ICML 2025poster

Chest X-rays (CXRs) play an integral role in driving critical decisions in disease management and patient care. While recent innovations have led to specialized models for various CXR interpretation tasks, these solutions often operate in isolation, limiting their practical utility in clinical pract…

2025

MorphoDiff: Cellular Morphology Painting with Diffusion Models

ICLR 2025spotlight

Understanding cellular responses to external stimuli is critical for parsing biological mechanisms and advancing therapeutic development. High-content image-based assays provide a cost-effective approach to examine cellular phenotypes induced by diverse interventions, which offers valuable insights…

Cited by 2SourcePDFScholar
2025

OTIAS: OcTree Implicit Adaptive Sampling for Multispectral and Hyperspectral Image Fusion

AAAI 2025technical

Implicit Neural Representation (INR) methods have demonstrated great potential in arbitrary-scale super-resolution tasks. This success is primarily due to their ability to continuously represent images using coordinates. In the task of remote sensing image fusion, INR methods have also shown promisi…

2025

On the Surprising Robustness of Sequential Convex Optimization for Contact-Implicit Motion Planning

RSS 2025poster

Contact-implicit motion planning—embedding contact sequencing as implicit complementarity constraints—holds the promise of leveraging continuous optimization to discover new contact patterns online. Nevertheless, the resulting optimization, being an instance of Mathematical Programming with Compleme…

Cited by 0PDFScholar
2025

PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding

IROS 2025

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in

Cited by 60SourceScholar
2025

PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly

NeurIPS 2025poster

While vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, remains severely limited. To close this gap, we introduce PhyBlock, a progressiv…

Cited by 0SourceScholar
2025

RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation

IROS 2025

This paper introduces RoboDexVLM, an innovative framework for robot task planning and grasp detection tailored for a collaborative manipulator equipped with a dexterous hand. Previous methods focus on simplified and limited manipulation tasks, which often neglect the complexities associated with gra

Cited by 15SourcecodeScholar
2025

Robot Navigation in Unknown and Cluttered Workspace with Dynamical System Modulation in Starshaped Roadmap

ICRA 2025

Compared to conventional decomposition methods that use ellipses or polygons to represent free space, starshaped representation can better capture the natural distribution of sensor data, thereby exploiting a larger portion of traversable space. This paper introduces a novel motion planning and cont

Cited by 3SourcecodeScholar
2025

SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering

AAAI 2025technical

The general capabilities of large language models (LLMs) make them the infrastructure for various AI applications, but updating their inner knowledge requires significant resources. Recent model editing is a promising technique for efficiently updating a small amount of knowledge of LLMs and has att…

2025

Safe Motion Planning for Multi-Vehicle Autonomous Driving in Uncertain Environment

RA-L 2025

In the field of motion planning for autonomous driving systems, ensuring the safety of multi-vehicle navigation is one of the crucial topics. An unavoidable problem in practice is that the noise-induced uncertainties in real-world applications highly degrade the safety of multi-vehicle navigation. I

Cited by 2SourceScholar
2025

Scene-Aware Explainable Multimodal Trajectory Prediction

ICRA 2025

Advancements in intelligent technologies have significantly improved navigation in complex traffic environments by enhancing environment perception and trajectory prediction for automated vehicles. However, current research often overlooks the joint reasoning of scenario agents and lacks explainabil

Cited by 2SourcecodeScholar
2025

Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning Distillation

ICASSP 2025accepted

The rapid advancement of large language models (LLMs) has significantly enhanced their reasoning abilities, enabling increasingly complex tasks. However, these capabilities often diminish in smaller, more computationally efficient models like GPT-2. Recent research shows that reasoning distillation…

Cited by 0SourceScholar
2025

TSCLIP: Robust CLIP Fine-Tuning for Worldwide Cross-Regional Traffic Sign Recognition

ICRA 2025

Traffic sign is a critical map feature for navigation and traffic control. Nevertheless, current methods for traffic sign recognition rely on traditional deep learning models, which typically suffer from significant performance degradation considering the variations in data distribution across diffe

Cited by 6SourcecodeScholar
2025

Towards Verifiable Text Generation with Generative Agent

AAAI 2025technical

Text generation with citations makes it easy to verify the factuality of Large Language Models’ (LLMs) generations. Existing one-step generation studies expose distinct shortages in answer refinement and in-context demonstration matching. In light of these challenges, we propose R2-MGA, a Retrieval…

Cited by 0SourcePDFScholar
2025

Zero-shot Stance Detection with Logically Consistent Data Augmentation

ICASSP 2025accepted

Zero-shot stance detection (ZSSD) is a challenging task that requires classifying stances towards unseen targets without large, well-curated training datasets. Existing data augmentation methods for ZSSD often suffer from semantic inconsistencies, hindering their effectiveness. To address these limi…

Cited by 0SourceScholar
2024

A Generic Trajectory Planning Method for Constrained All-Wheel-Steering Robots

IROS 2024poster

This paper presents a generic trajectory planning method for wheeled robots with fixed steering axes while the steering angle of each wheel is constrained. In the existing literatures, All-Wheel-Steering (AWS) robots, incorporating modes such as rotation-free translation maneuvers, in-situ rotationa…

Cited by 0SourcecodeScholar
2024

Arm-Constrained Curriculum Learning for Loco-Manipulation of a Wheel-Legged Robot

IROS 2024poster

Incorporating a robotic manipulator into a wheellegged robot enhances its agility and expands its potential for practical applications. However, the presence of potential instability and uncertainties presents additional challenges for control objectives. In this paper, we introduce an arm-constrain…

Cited by 5SourcecodeScholar
2024

Chance-Aware Lane Change with High-Level Model Predictive Control Through Curriculum Reinforcement Learning

ICRA 2024poster

Lane change in dense traffic typically requires the recognition of an appropriate opportunity for maneuvers, which remains a challenging problem in self-driving. In this work, we propose a chance-aware lane-change strategy with high-level model predictive control (MPC) through curriculum reinforceme…

Cited by 7SourceScholar
2024

Collision-Free Trajectory Optimization in Cluttered Environments Using Sums-of-Squares Programming

RA-L 2024

In this work, we propose a trajectory optimization approach for robot navigation in cluttered 3D environments. We represent the robot's geometry as a semialgebraic set defined by polynomial inequalities such that robots with general shapes can be suitably characterized. We exploit the collision-free

Cited by 12SourcecodeScholar
2024

Confucius: Iterative Tool Learning from Introspection Feedback by Easy-to-Difficult Curriculum

AAAI 2024technical

Augmenting large language models (LLMs) with external tools has emerged as a promising approach to extending the capability of LLMs. Although there are some works that employ open-source LLMs for the tool-learning task, most of them are trained in a controlled environment in which LLMs only learn to…

2024

Every Dataset Counts: Scaling up Monocular 3D Object Detection with Joint Datasets Training

IROS 2024poster

Monocular 3D object detection is essential for autonomous driving. However, current monocular 3D detection algorithms rely on expensive 3D labels from LiDAR scans, making it difficult to use in new datasets and unfamiliar environments. This study explores training a monocular 3D object detection mod…

Cited by 4SourceScholar
2024

Geometry-Aware Safety-Critical Local Reactive Controller for Robot Navigation in Unknown and Cluttered Environments

RA-L 2024

This work proposes a safety-critical local reactive controller that enables the robot to navigate in unknown and cluttered environments. In particular, the trajectory tracking task is formulated as a constrained polynomial optimization problem. Then, safety constraints are imposed on the control var

Cited by 13SourceScholar
2024

Improving Factual Error Correction by Learning to Inject Factual Errors

AAAI 2024technical

Factual error correction (FEC) aims to revise factual errors in false claims with minimal editing, making them faithful to the provided evidence. This task is crucial for alleviating the hallucination problem encountered by large language models. Given the lack of paired data (i.e., false claims and…

2024

MCGMapper: Light-Weight Incremental Structure from Motion and Visual Localization with Planar Markers and Camera Groups

IROS 2024poster

Structure from Motion (SfM) and visual localization in indoor texture-less scenes and industrial scenarios present prevalent yet challenging research topics. Existing SfM methods designed for natural scenes typically yield low accuracy or map-building failures due to insufficient robust feature extr…

Cited by 1SourcecodeScholar
2024

PMET: Precise Model Editing in a Transformer

AAAI 2024technical

Model editing techniques modify a minor proportion of knowledge in Large Language Models (LLMs) at a relatively low cost, which have demonstrated notable success. Existing methods assume Transformer Layer (TL) hidden states are values of key-value memories of the Feed-Forward Network (FFN). They usu…

2024

Parallel Optimization with Hard Safety Constraints for Cooperative Planning of Connected Autonomous Vehicles

ICRA 2024poster

The development of connected autonomous vehicles (CAVs) facilitates the enhancement of traffic efficiency in complicated scenarios. Difficulties remain unsolved in developing an effective and efficient coordination strategy for CAVs. In this paper, we formulate the cooperative autonomous driving tas…

Cited by 4SourceScholar
2024

Reward-Driven Automated Curriculum Learning for Interaction-Aware Self-Driving at Unsignalized Intersections

IROS 2024poster

In this work, we present a reward-driven automated curriculum reinforcement learning approach for interaction-aware self-driving at unsignalized intersections, taking into account the uncertainties associated with surrounding vehicles (SVs). These uncertainties encompass the uncertainty of SVs’ driv…

Cited by 6SourceScholar
2024

S3E: A Multi-Robot Multimodal Dataset for Collaborative SLAM

RA-L 2024

The burgeoning demand for collaborative robotic systems to execute complex tasks collectively has intensified the research community's focus on advancing simultaneous localization and mapping (SLAM) in a cooperative context. Despite this interest, the scalability and diversity of existing datasets f

Cited by 42SourcecodeScholar
2024

SCPNet: Unsupervised Cross-modal Homography Estimation via Intra-modal Self-supervised Learning

ECCV 2024poster

"We propose a novel unsupervised cross-modal homography estimation framework based on intra-modal Self-supervised learning, Correlation, and consistent feature map Projection, namely SCPNet. The concept of intra-modal self-supervised learning is first presented to facilitate the unsupervised cross-m…

2024

Towards Proactive Interactions for In-Vehicle Conversational Assistants Utilizing Large Language Models

IJCAI 2024poster

Research demonstrates that the proactivity of in-vehicle conversational assistants (IVCAs) can help to reduce distractions and enhance driving safety, better meeting users' cognitive needs. However, existing IVCAs struggle with user intent recognition and context awareness, which leads to suboptimal…

2023

PivotFEC: Enhancing Few-shot Factual Error Correction with a Pivot Task Approach using Large Language Models

EMNLP 2023long findings

Factual Error Correction (FEC) aims to rectify false claims by making minimal revisions to align them more accurately with supporting evidence. However, the lack of datasets containing false claims and their corresponding corrections has impeded progress in this field. Existing distantly supervised…

Cited by 0SourceScholar
2023

Reinforcement Learning for Robot Navigation with Adaptive Forward Simulation Time (AFST) in a Semi-Markov Model

IROS 2023poster

Deep reinforcement learning (DRL) algorithms have proven effective in robot navigation, especially in unknown environments, by directly mapping perception inputs into robot control commands. However, most existing methods ignore the local minimum problem in navigation and thereby cannot handle compl…

Cited by 0SourcecodeScholar
2023

Why Is the Winner the Best?

CVPR 2023poster

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and…

Cited by 29SourcePDFScholar
2022

Few-shot Named Entity Recognition with Entity-level Prototypical Network Enhanced by Dispersedly Distributed Prototypes

COLING 2022main

Few-shot named entity recognition (NER) enables us to build a NER system for a new domain using very few labeled examples. However, existing prototypical networks for this task suffer from roughly estimated label dependency and closely distributed prototypes, thus often causing misclassifications. T…

Cited by 36SourcePDFScholar
2022

Incremental Few-Shot Object Detection for Robotics

ICRA 2022poster

Incremental few-shot learning is highly expected for practical robotics applications. On one hand, robot is desired to learn new tasks quickly and flexibly using only few annotated training samples; on the other hand, such new additional tasks should be learned in a continuous and incremental manner…

Cited by 15SourceScholar
2022

Learning to Socially Navigate in Pedestrian-rich Environments with Interaction Capacity

ICRA 2022poster

Existing navigation policies for autonomous robots tend to focus on collision avoidance while ignoring human-robot interactions in social life. For instance, robots can pass along the corridor safer and easier if pedestrians notice them. Sounds have been considered as an efficient way to attract the…

Cited by 18SourceScholar
2021

Crowd-Aware Robot Navigation for Pedestrians with Multiple Collision Avoidance Strategies via Map-based Deep Reinforcement Learning

IROS 2021poster

It is challenging for a mobile robot to navigate through human crowds. Existing approaches usually assume that pedestrians follow a predefined collision avoidance strategy, like social force model (SFM) or optimal reciprocal collision avoidance (ORCA). However, their performances commonly need to be…

Cited by 41SourceScholar
2021

DRQN-based 3D Obstacle Avoidance with a Limited Field of View

IROS 2021poster

In this paper, we propose a map-based end-to-end DRL approach for three-dimensional (3D) obstacle avoidance in a partially observed environment, which is applied to achieve autonomous navigation for an indoor mobile robot using a depth camera with a narrow field of view. We first train a neural netw…

Cited by 10SourceScholar
2021

EfficientTTS: An Efficient and High-Quality Text-to-Speech Architecture

ICML 2021spotlight

In this work, we address the Text-to-Speech (TTS) task by proposing a non-autoregressive architecture called EfficientTTS. Unlike the dominant non-autoregressive TTS models, which are trained with the need of external aligners, EfficientTTS optimizes all its parameters with a stable, end-to-end trai…

2021

End-to-End Conversational Search for Online Shopping with Utterance Transfer

EMNLP 2021main

Successful conversational search systems can present natural, adaptive and interactive shopping experience for online shopping customers. However, building such systems from scratch faces real word challenges from both imperfect product schema/knowledge and lack of training dialog data. In this work…

2021

Improving Neural Text Normalization with Partial Parameter Generator and Pointer-Generator Network

ICASSP 2021accepted

Text Normalization (TN) is an essential part in conversational systems like text-to-speech synthesis (TTS) and automatic speech recognition (ASR). It is a process of transforming non-standard words (NSW) into a representation of how the words are to be spoken. Existing approaches to TN are mainly ru…

Cited by 0SourceScholar
2021

SEQ-CPC : Sequential Contrastive Predictive Coding for Automatic Speech Recognition

ICASSP 2021accepted

Inspired by the contrastive predictive coding (CPC), we propose a feature representation scheme for automatic speech recognition (ASR), which encodes sequential dependency information from raw audio signals. Following the original CPC, for a given frame, mutual information (MI) lower bound is maximi…

Cited by 0SourceScholar
2021

Unsupervised Learning for Multi-Style Speech Synthesis with Limited Data

ICASSP 2021accepted

Existing multi-style speech synthesis methods require either style labels or large amounts of unlabeled training data, making data acquisition difficult. In this paper, we present an unsupervised multi-style speech synthesis method that can be trained with limited data. We leverage instance discrimi…

Cited by 0SourceScholar
2020

Auxiliary Template-Enhanced Generative Compatibility Modeling

IJCAI 2020poster

In recent years, there has been a growing interest in the fashion analysis (e.g., clothing matching) due to the huge economic value of the fashion industry. The essential problem is to model the compatibility between the complementary fashion items, such as the top and bottom in clothing matching. T…

Cited by 0SourcePDFScholar
2020

Flow-TTS: A Non-Autoregressive Network for Text to Speech Based on Flow

ICASSP 2020accepted

In this work, we propose Flow-TTS, a non-autoregressive end-to-end neural TTS model based on generative flow. Unlike other non-autoregressive models, Flow-TTS can achieve high-quality speech generation by using a single feed-forward network. To our knowledge, Flow-TTS is the first TTS model utilizin…

Cited by 0SourceScholar
2020

Grasping Detection Network with Uncertainty Estimation for Confidence-Driven Semi-Supervised Domain Adaptation

IROS 2020poster

Data-efficient domain adaptation with only a few labelled data is desired for many robotic applications, e.g., in grasping detection, the inference skill learned from a grasping dataset is not universal enough to directly apply on various other daily/industrial applications. This paper presents an a…

Cited by 32SourceScholar
2020

Learning-Based Controller Optimization for Repetitive Robotic Tasks

IROS 2020poster

Dynamic control for robotic automation tasks is traditionally designed and optimized with a model-based approach, and the performance relies heavily upon accurate system modeling. However, modeling the true dynamics of increasingly complex robotic systems is an extremely challenging task and it ofte…

Cited by 2SourceScholar
2020

Span-based Joint Entity and Relation Extraction with Attention-based Span-specific and Contextual Semantic Representations

COLING 2020main

Span-based joint extraction models have shown their efficiency on entity recognition and relation extraction. These models regard text spans as candidate entities and span tuples as candidate relation tuples. Span semantic representations are shared in both entity recognition and relation extraction…

Cited by 88SourcePDFScholar