← Search

Xiao Liang

36 accepted papers

2026

Anatomical Region-Guided Contrastive Decoding: A Plug-and-Play Strategy for Mitigating Hallucinations in Medical VLMs

AAAI 2026technical

Medical Vision-Language Models (MedVLMs) show immense promise in clinical applicability. However, their reliability is hindered by hallucinations, where models often fail to derive answers from visual evidence, instead relying on learned textual priors. Existing mitigation strategies for MedVLMs hav

Cited by 0SourcePDFScholar
2026

Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVR

ICLR 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a key paradigm for post-training Large Language Models (LLMs), particularly for complex reasoning tasks. However, vanilla RLVR training has been shown to improve Pass@1 performance at the expense of policy entropy, leading…

Cited by 0SourcecodeScholar
2026

Feedback Matters: Augmenting Autonomous Dissection with Visual and Topological Feedback

ICRA 2026poster

Autonomous surgical systems must adapt to highly dynamic environments where tissue properties and visual cues evolve rapidly. Central to such adaptability is feedback: the ability to sense, interpret, and respond to changes during execution. While feedback mechanisms have been explored in surgical r…

2026

LapSurgie: Humanoid Robots Performing Surgery Via Teleoperated Handheld Laparoscopy

ICRA 2026poster

Robotic laparoscopic surgery has gained increasing attention in recent years for its potential to deliver more efficient and precise minimally invasive procedures. However, adoption of surgical robotic platforms remains largely confined to high-resource medical centers, exacerbating healthcare dispa…

2026

PEMTRS: Perception-Enhanced Memory With Temporal Region Selection for Vision-Based Multirotor Navigation

RA-L 2026

End-to-end deep learning methods for autonomous aerial vehicles like multirotors have been successfully applied to high-agility flight. However, existing end-to-end navigation methods for multirotors typically rely solely on the robot's current observation and state, neglecting explicit modeling of

Cited by 0SourceScholar
2026

ReMatch: Boosting Representation through Matching for Multimodal Retrieval

CVPR 2026

We present ReMatch, a framework that leverages the generative strength of MLLMs for multimodal retrieval. Previous approaches treated an MLLM as a simple encoder, ignoring its generative nature, and under-utilising its compositional reasoning and world knowledge. We train the embedding MLLM end-to-e

Cited by 0SourcecodeScholar
2026

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

ICLR 2026poster

Recent advancements in long chain-of-thought (CoT) reasoning, particularly through the Group Relative Policy Optimization algorithm used by DeepSeek-R1, have led to significant interest in the potential of Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Models (LLMs). While…

Cited by 0SourceScholar
2026

Towards Autonomous Tape Handling for Robotic Wound Redressing

ICRA 2026poster

Chronic wounds, such as diabetic, pressure, and venous ulcers, affect over 6.5 million patients in the United States alone and generate an annual cost exceeding 25 billion. Despite this burden, chronic wound care remains a routine yet manual process performed exclusively by trained clinicians due to…

2026

UniGeoRS: A Unified Benchmark for Tri-view Geo-Localization

CVPR 2026

Cross-view geo-localization (CVGL) aims to estimate an image's geographic location by matching it with geo-referenced images from different viewpoints, supporting applications such as autonomous driving, UAV navigation, and visual surveillance. However, due to the high cost of image collection, curr

Cited by 0SourceScholar
2025

AutoPeel: Adhesion-Aware Safe Peeling Trajectory Optimization for Robotic Wound Care

ICRA 2025

Chronic wounds, including diabetic ulcers, pressure ulcers, and ulcers secondary to venous hypertension, affects more than 6.5 million patients and a yearly cost of more than $25 billion in the United States alone. Chronic wound treatment is currently a manual process, and we envision a future where

Cited by 2SourceScholar
2025

Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective

ACL 2025long

Large Language Models (LLMs) have made notable progress in mathematical reasoning, yet they often rely on single-paradigm reasoning that limits their effectiveness across diverse tasks. In this paper, we introduce Chain-of-Reasoning (CoR), a novel unified framework that integrates multiple reasoning…

2025

GuiLoMo: Allocating Experts and Ranks for LoRA-MoE via Bilevel Optimization with GuidedSelection Vectors

EMNLP 2025

Parameter-efficient fine-tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA), offer an efficient way to adapt large language models with reduced computational costs. However, their performance is limited by the small number of trainable parameters. Recent work combines LoRA with the Mixtu

2025

Integrative Decoding: Improving Factuality via Implicit Self-consistency

ICLR 2025poster

Self-consistency-based approaches, which involve repeatedly sampling multiple outputs and selecting the most consistent one as the final response, prove to be remarkably effective in improving the factual accuracy of large language models. Nonetheless, existing methods usually have strict constraint…

Cited by 4SourcePDFScholar
2025

LAMPS: A Novel Robot Generalization Framework for Learning Adaptive Multi-Periodic Skills

IROS 2025

Learning from Demonstrations (LfD) methods are applied to transfer human skills to robots from expert demonstrations, enabling them to perform complex tasks. However, existing methods often struggle to handle such long-horizon human skills as cleaning or wiping stains on the surface, which involve m

Cited by 0SourcecodeScholar
2025

MEDiC: Autonomous Surgical Robotic Assistance to Maximizing Exposure for Dissection and Cautery

ICRA 2025

Surgical automation has the capability to improve the consistency of patient outcomes and broaden access to advanced surgical care in underprivileged communities. Shared autonomy, where the robot automates routine subtasks while the surgeon retains partial teleoperative control, offers great potenti

Cited by 6SourceScholar
2025

Online Anti-Swing Trajectory Refinement for Variable-Length Cable-Suspended Aerial Transportation Robot

IROS 2025

Aerial robots have demonstrated significant potential in suspended cargo transportation, especially in industries such as logistics and food delivery. Due to the underactuated and nonlinear dynamics of the cable-suspended system, directly tracking a given trajectory with a multicopter without modify

Cited by 0SourceScholar
2025

Planning and Compliant Control for Laparoscopic Ultrasound Scanning System

RA-L 2025

Laparoscopic ultrasound (LUS) serves as a critical technology in intraoperative surgeries, particularly for guiding complex procedures in liver diseases. However, the development of robotic systems for LUS examination remains hindered by challenges such as high costs and the absence of force feedbac

Cited by 0SourceScholar
2025

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning

NeurIPS 2025poster

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for training large language models (LLMs) on complex reasoning tasks, such as mathematical problem solving. A prerequisite for the scalability of RLVR is a high-quality problem set with precise and verifiable answers. However…

Cited by 0SourcecodeScholar
2025

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models

ICCV 2025poster

The rapid advancements in Vision Language Models (VLMs) have prompted the development of multi-modal medical assistant systems. Despite this progress, current models still have inherent probabilistic uncertainties, often producing erroneous or unverified responses--an issue with serious implications…

2024

Achieving Autonomous Cloth Manipulation with Optimal Control via Differentiable Physics-Aware Regularization and Safety Constraints

ICRA 2024poster

Cloth manipulation is a category of deformable object manipulation of great interest to the robotics community, from applications of automated laundry-folding and home organizing to textiles and flexible manufacturing. Despite the desire for automated cloth manipulation, the thin-shell dynamics and…

Cited by 3SourceScholar
2024

Chunk, Align, Select: A Simple Long-sequence Processing Method for Transformers

ACL 2024long

Although dominant in natural language processing, transformer-based models still struggle with long-sequence processing, due to the computational costs of their self-attention operations, which increase exponentially as the length of the input sequence grows. To address this challenge, we propose a…

2024

CoSTA: End-to-End Comprehensive Space-Time Entanglement for Spatio-Temporal Video Grounding

AAAI 2024technical

This paper studies the spatio-temporal video grounding task, which aims to localize a spatio-temporal tube in an untrimmed video based on the given text description of an event. Existing one-stage approaches suffer from insufficient space-time interaction in two aspects: i) less precise prediction o…

Cited by 1SourcePDFScholar
2024

DE-TGN: Uncertainty-Aware Human Motion Forecasting Using Deep Ensembles

RA-L 2024

Ensuring the safety of human workers in a collaborative environment with robots is of utmost importance. Although accurate pose prediction models can help prevent collisions between human workers and robots, they are still susceptible to critical errors. In this study, we propose a novel approach ca

Cited by 19SourceScholar
2024

Improving Disturbance Estimation and Suppression via Learning Among Systems With Mismatched Dynamics

RA-L 2024

Iterative learning control (ILC) is a method for reducing system tracking or estimation errors over multiple iterations by using information from past iterations. The disturbance observer (DOB) is used to estimate and mitigate disturbances within the system, while the system is being affected by the

Cited by 4SourceScholar
2024

JIGGLE: An Active Sensing Framework for Boundary Parameters Estimation in Deformable Surgical Environments

RSS 2024poster

Surgical automation can improve the accessibility and consistency of life-saving procedures. Most surgeries require separating layers of tissue to access the surgical site, and suturing to re-attach incisions. These tasks involve deformable manipula- tion to safely identify and alter tissue attachme…

Cited by 6SourcePDFScholar
2024

Real-to-Sim Deformable Object Manipulation: Optimizing Physics Models with Residual Mappings for Robotic Surgery

ICRA 2024poster

Accurate deformable object manipulation (DOM) is essential for achieving autonomy in robotic surgery, where soft tissues are being displaced, stretched, and dissected. Many DOM methods can be powered by simulation, which ensures realistic deformation by adhering to the governing physical constraints…

Cited by 9SourceScholar
2024

TMFN: A Target-oriented Multi-grained Fusion Network for End-to-end Aspect-based Multimodal Sentiment Analysis

COLING 2024main

End-to-end multimodal aspect-based sentiment analysis (MABSA) combines multimodal aspect terms extraction (MATE) with multimodal aspect sentiment classification (MASC), aiming to simultaneously extract aspect words and classify the sentiment polarity of each aspect. However, existing MABSA methods h…

Cited by 3SourcePDFScholar
2024

Task Oriented In-Domain Data Augmentation

EMNLP 2024main

Large Language Models (LLMs) have shown superior performance in various applications and fields. To achieve better performance on specialized domains such as law and advertisement, LLMs are often continue pre-trained on in-domain data. However, existing approaches suffer from two major issues. First…

2024

TransFusion: A Practical and Effective Transformer-Based Diffusion Model for 3D Human Motion Prediction

RA-L 2024

Predicting human motion plays a crucial role in ensuring a safe and effective human-robot close collaboration in intelligent remanufacturing systems of the future. Existing works can be categorized into two groups: those focusing on accuracy, predicting a single future motion, and those generating d

Cited by 46SourcecodeScholar
2024

Unleashing Region Understanding in Intermediate Layers for MLLM-based Referring Expression Generation

NeurIPS 2024poster

The Multi-modal Large Language Model (MLLM) based Referring Expression Generation (REG) task has gained increasing popularity, which aims to generate an unambiguous text description that applies to exactly one object or region in the image by leveraging foundation models. We empirically found that t…

2022

Uncertainty-Assisted Image-Processing for Human-Robot Close Collaboration

RA-L 2022

The safety of human workers has been the main concern in human-robot close collaboration. Along with rapidly developed artificial intelligence techniques, deep learning models using two-dimensional images have become feasible solutions for human motion detection. These models serve as “sensors” in t

Cited by 24SourceScholar
2021

Large Scale Image Completion via Co-Modulated Generative Adversarial Networks

ICLR 2021spotlight

Numerous task-specific variants of conditional generative adversarial networks have been developed for image completion. Yet, a serious limitation remains that all existing algorithms tend to fail when handling large-scale missing regions. To overcome this challenge, we propose a generic new approac…

2018

Joint List Polar Decoder with Successive Cancellation and Sphere Decoding

ICASSP 2018accepted

For polar codes, both successive cancellation list (SCL) decoding and list sphere decoding (LSD) aim to balance performance and complexity. The same list structure but different decoding schedules of SCL and LSD can lead to a combination of both schemes. In this paper, an efficient joint list decode…

Cited by 0SourceScholar