← Search

Xinyu Liu

65 accepted papers

2026

A Cable-Driven Soft Robotic Hand with an In-Hand RGB-D Camera for Dexterous Grasping and Manipulation

ICRA 2026poster

The aspiration to replicate the capabilities of the human hand has driven innovations in the design of soft robotic hands. Despite these advancements, many existing designs of soft hands still lack effective in-hand vision and the ability for each finger to achieve active multidegree-of-freedom moti…

Cited by 1Scholar
2026

A VAGP-Based Adaptive Kalman Filter for Force Estimation of Robot

RA-L 2026

In human-robot interaction, external force measurement is fundamental to achieving robot compliance control. Parameter identification based on robot dynamics enables external force detection without expensive sensors. However, the unmodeled dynamic errors inherent in real robots pose a challenge to

Cited by 0SourceScholar
2026

Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors

ICLR 2026poster

Context compression presents a promising approach for accelerating large language model (LLM) inference by compressing long contexts into compact representations.Current context compression methods predominantly rely on autoencoding tasks to train context-agnostic compression tokens to compress cont…

Cited by 0SourcecodeScholar
2026

BioFormer: Rethinking Cross-Subject Generalization via Spectral Structural Alignment in Biomedical Time-Series

ICML 2026poster

Cross-subject generalization in biomedical time-series (BTS) refers to training on data from some subjects and testing on unseen subjects. The key challenge is to suppress subject-specific variability in BTS representations. Most existing methods implicitly suppress the variability through model bui…

Cited by 0SourceScholar
2026

Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities

IJCAI 2026

Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbalance. Existing methods generate features for modality missing from available ones, but differences in expression mechanisms and sentiment dynamics ac

Cited by 0Scholar
2026

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

AAAI 2026technical

Human-defined creativity is highly abstract, posing a challenge for multimodal large language models (MLLMs) to comprehend and assess creativity that aligns with human judgments. The absence of an existing benchmark further exacerbates this dilemma. To this end, we propose CreBench, which consists o

Cited by 0SourcePDFScholar
2026

GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language Models

ICML 2026poster

Geometric shapes play important roles in both physical world and human cognition. While multimodal large language models (MLLMs) have made significant advancements in visual understanding, their abilities to recognize geometric shapes and their spatial relationships, which we term geometric percepti…

Cited by 0SourceScholar
2026

Learning Latent Imaging Biomarkers for Interpretable Microvascular Invasion Prediction in Hepatocellular Carcinoma

AAAI 2026technical

Microvascular invasion (MVI) is a critical prognostic factor that significantly impacts postoperative outcomes in hepatocellular carcinoma (HCC). As the current gold standard for the diagnosis of MVI is based on the postoperative histopathological examination of whole slide images, accurate preopera

Cited by 0SourcePDFScholar
2026

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

ICML 2026poster

While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models (LLMs), the absence of many folklore lemmas in Mathlib remains a persistent barrier that limits Lean's usability as an everyday tool for mathematicians like …

Cited by 0SourceScholar
2026

Offline Two-Player Zero-Sum Markov Games with KL Regularization

ICML 2026poster

We study the problem of learning Nash equilibria in offline two-player zero-sum Markov games. While existing approaches often rely on explicit pessimism to address distribution shift, we show that KL regularization alone suffices to stabilize learning and guarantee convergence. We first introduce Re…

Cited by 0SourceScholar
2026

PHOTONS: Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views

AAAI 2026technical

We present PHOTONS (Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views), a real-time framework for novel view synthesis without requiring camera calibration. Our method reconstructs consistent 3D Gaussian point clouds and synthesizes 2K photo-realistic novel vie

Cited by 0SourcePDFScholar
2026

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers

ICML 2026poster

The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures. A critical design choice is the placement of normalization layers, leading to a fundamental trade-off: the ''PreNorm'' architecture ensures training stability at the cost of potential perform…

Cited by 0SourceScholar
2026

Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs

ICLR 2026poster

Although unified MLLMs aim to unify generation and understanding, they are considered to exhibit an internal gap, with understanding outperforming generation. Through large‑scale evaluation across multiple MLLMs and tasks, we confirm the widespread non‑unification of MLLMs, and demonstrate that it i…

Cited by 0SourceScholar
2026

Twin-DP3: View-Invariant 3D Diffusion Policy With Digital Twin

RA-L 2026

Learning visuomotor policies with imitation learning from 3D observations is a primary research direction in robotic manipulation, as 3D data inherently captures spatial features critical for action. While many existing methods rely on multiview point cloud fusion, recent studies like 3D Diffusion P

Cited by 0SourceScholar
2025

CoMamba: Real-time Cooperative Perception Unlocked with State-Space Models

IROS 2025

Cooperative perception systems play a vital role in enhancing the safety and efficiency of vehicular autonomy. Although recent studies have highlighted the efficacy of vehicle-to-everything (V2X) communication techniques in autonomous driving, a significant challenge persists: how to efficiently int

Cited by 7SourceScholar
2025

ContextAware: A Multi-Agent Framework for Detecting Harmful Image-Based Comments on Social Media

IJCAI 2025

Detecting hidden stigmatization in social media poses significant challenges due to semantic misalignments between textual and visual modalities, as well as the subtlety of implicit stigmatization. Traditional approaches often fail to capture these complexities in real-world, multimodal content. To

2025

ExAct: A Video-Language Benchmark for Expert Action Analysis

NeurIPS 2025poster

We present ExAct, a new video-language benchmark for expert-level understanding of skilled physical human activities. Our new benchmark contains 3,521 expert-curated video question-answer pairs spanning 11 physical activities in 6 domains: Sports, Bike Repair, Cooking, Health, Music, and Dance. ExAc…

Cited by 0SourcecodeScholar
2025

Finite Sample Analysis of Linear Temporal Difference Learning with Arbitrary Features

NeurIPS 2025poster

Linear TD($\lambda$) is one of the most fundamental reinforcement learning algorithms for policy evaluation. Previously, convergence rates are typically established under the assumption of linearly independent features, which does not hold in many practical scenarios. This paper instead establishes…

Cited by 0SourceScholar
2025

GaussianReg: Rapid 2D/3D Registration for Emergency Surgery via Explicit 3D Modeling with Gaussian Primitives

ICCV 2025poster

Intraoperative 2D/3D registration, which aligns preoperative CT scans with intraoperative X-ray images, is critical for surgical navigation. However, existing methods require extensive preoperative training (several hours), making them unsuitable for emergency surgeries where minutes significantly i…

2025

IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method

EMNLP 2025

High-order numerical methods enhance Transformer performance in tasks like NLP and CV, but introduce a performance-efficiency trade-off due to increased computational overhead. Our analysis reveals that conventional efficiency techniques, such as distillation, can be detrimental to the performance o

2025

InfoBridge: Balanced Multimodal Integration through Conditional Dependency Modeling

ICCV 2025poster

Developing systems that interpret diverse real-world signals remains a fundamental challenge in multimodal learning. Current approaches face significant obstacles from inherent modal heterogeneity. While existing methods attempt to enhance fusion through cross-modal alignment or interaction mechanis…

2025

LiON: Learning Point-Wise Abstaining Penalty for LiDAR Outlier DetectioN Using Diverse Synthetic Data

AAAI 2025technical

LiDAR-based semantic scene understanding is an important module in the modern autonomous driving perception stack. However, identifying outlier points in a LiDAR point cloud is challenging as LiDAR point clouds lack semantically-rich information. While former SOTA methods adopt heuristic architectur…

2025

Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set

ICML 2025poster

$Q$-learning is one of the most fundamental reinforcement learning algorithms. It is widely believed that $Q$-learning with linear function approximation (i.e., linear $Q$-learning) suffers from possible divergence until the recent work Meyn (2024) which establishes the ultimate almost sure boundedn…

Cited by 0SourcePDFScholar
2025

MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation

ICCV 2025poster

High-resolution (HR) medical videos are vital for accurate diagnosis, yet are hard to acquire due to hardware limitations and physiological constraints. Clinically, the collected low-resolution (LR) medical videos present unique challenges for video super-resolution (VSR) models, including camera sh…

2025

New Network Protocol for Supermedia-Enhanced Telerobotics

IROS 2025

The growing complexity of robotic teleoperation systems necessitates the integration of multiple feedback modalities, including video, audio, force, tactile, and temperature feedback. The concept of supermedia is utilized to describe the aggregation of these feedback streams. By integrating multiple

Cited by 0SourceScholar
2025

Position IDs Matter: An Enhanced Position Layout for Efficient Context Compression in Large Language Models

EMNLP 2025

Using special tokens (e.g., gist, memory, or compressed tokens) to compress context information is a common practice for large language models (LLMs). However, existing approaches often neglect that position encodings inherently induce local inductive biases in models, causing the compression proces

Cited by 0SourcePDFScholar
2025

SMR-Net: Semantic-Guided Mutually Reinforcing Network for Cross-Modal Image Fusion and Salient Object Detection

AAAI 2025technical

This paper introduces a lightweight Semantic-guided Mutually Reinforcing network (SMR-Net) for the tasks of cross-modal image fusion and salient object detection (SOD). The core concept of SMR-Net is to leverage semantics for directing the mutual reinforcing between image fusion and SOD. Specificall…

Cited by 0SourcePDFScholar
2025

TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making

EMNLP 2025

Using effective generalization capabilities of vision language models (VLMs) in context-specific dynamic tasks for embodied artificial intelligence remains a significant challenge. Although supervised fine-tuned models can better align with the real physical world, they still exhibit sluggish respon

Cited by 0SourcePDFScholar
2025

Track Any Anomalous Object:A Granular Video Anomaly Detection Pipeline

CVPR 2025poster

Video anomaly detection (VAD) is crucial in scenarios such as surveillance and autonomous driving, where timely detection of unexpected activities is essential. Albeit existing methods have primarily focused on detecting anomalous objects in videos--either by identifying anomalous frames or objects-…

Cited by 0SourcePDFScholar
2025

U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation

AAAI 2025technical

U-Net has become a cornerstone in various visual applications such as image segmentation and diffusion probability models. While numerous innovative designs and improvements have been introduced by incorporating transformers or MLPs, the networks are still limited to linearly modeling patterns as we…

2025

V2X-DG: Domain Generalization for Vehicle-to-Everything Cooperative Perception

ICRA 2025

LiDAR-based Vehicle-to-Everything (V2X) cooperative perception has demonstrated its impact on the safety and effectiveness of autonomous driving. Since current cooperative perception algorithms are trained and tested on the same dataset, the generalization ability of cooperative perception systems r

Cited by 3SourceScholar
2025

V2X-DGW: Domain Generalization for Multi-Agent Perception Under Adverse Weather Conditions

ICRA 2025

Current LiDAR-based Vehicle-to-Everything (V2X) multi-agent perception systems have shown the significant success on 3D object detection. While these models perform well in the trained clean weather, they struggle in unseen adverse weather conditions with the domain gap. In this paper, we propose a

Cited by 19SourcecodeScholar
2024

AdvGPS: Adversarial GPS for Multi-Agent Perception Attack

ICRA 2024poster

The multi-agent perception system collects visual data from sensors located on various agents and leverages their relative poses determined by GPS signals to effectively fuse information, mitigating the limitations of single-agent sensing, such as occlusion. However, the precision of GPS signals can…

Cited by 6SourcecodeScholar
2024

Breaking Data Silos: Cross-Domain Learning for Multi-Agent Perception from Independent Private Sources

ICRA 2024poster

The diverse agents in multi-agent perception systems may be from different companies. Each company might use the identical classic neural network architecture based encoder for feature extraction. However, the data source to train the various agents is independent and private in each company, leadin…

Cited by 7SourcecodeScholar
2024

CLIFF: Continual Latent Diffusion for Open-Vocabulary Object Detection

ECCV 2024oral

"Open-vocabulary object detection (OVD) utilizes image-level cues to expand the linguistic space of region proposals, thereby facilitating the detection of diverse novel classes. Recent works adapt CLIP embedding by minimizing the object-image and object-text discrepancy combinatorially in a discrim…

2024

Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAM

NeurIPS 2024poster

As the vision foundation models like the Segment Anything Model (SAM) demonstrate potent universality, they also present challenges in giving ambiguous and uncertain predictions. Significant variations in the model output and granularity can occur with simply subtle changes in the prompt, contradict…

2024

Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-Context Models

EMNLP 2024main

Numerous recent works target to extend effective context length for language models and various methods, tasks and benchmarks exist to measure model’s effective memory length. However, through thorough investigations, we find limitations for currently existing evaluations on model’s memory. We provi…

2024

GTP-4o: Modality-prompted Heterogeneous Graph Learning for Omni-modal Biomedical Representation

ECCV 2024poster

"Recent advances in learning multi-modal representation have witnessed the success in biomedical domains. While established techniques enable handling multi-modal information, the challenges are posed when extended to various clinical modalities and practical modality-missing setting due to the inhe…

2024

Light the Night: A Multi-Condition Diffusion Framework for Unpaired Low-Light Enhancement in Autonomous Driving

CVPR 2024poster

Vision-centric perception systems for autonomous driving have gained considerable attention recently due to their cost-effectiveness and scalability especially compared to LiDAR-based systems. However these systems often struggle in low-light conditions potentially compromising their performance and…

Cited by 24SourcePDFScholar
2024

S2R-ViT for Multi-Agent Cooperative Perception: Bridging the Gap from Simulation to Reality

ICRA 2024poster

Due to the lack of enough real multi-agent data and time-consuming of labeling, existing multi-agent cooperative perception algorithms usually select the simulated sensor data for training and validating. However, the perception performance is degraded when these simulation-trained models are deploy…

Cited by 21SourceScholar
2024

TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes

ECCV 2024poster

"3D dense captioning stands as a cornerstone in achieving a comprehensive understanding of 3D scenes through natural language. It has recently witnessed remarkable achievements, particularly in indoor settings. However, the exploration of 3D dense captioning in outdoor scenes is hindered by two majo…

2023

A Small-Scale Untethered Tensegrity Robot With High Velocity and Multi Locomotion Modes

RA-L 2023

Tensegrity mobile robots are well-appraised for their high stiffness-to-mass ratio and superior structural compliance. However, traditional untethered tensegrity mobile robots usually have low velocity due to the large actuation force required by coupling effects among stiff struts and soft cables.

Cited by 15SourceScholar
2023

ADAPT: Action-aware Driving Caption Transformer

ICRA 2023poster

End-to-end autonomous driving has great potential in the transportation industry. However, the lack of transparency and interpretability of the automatic decision-making process hinders its industrial adoption in practice. There have been some early attempts to use attention maps or cost volume for…

Cited by 90SourcecodeScholar
2023

Delving Into Shape-Aware Zero-Shot Semantic Segmentation

CVPR 2023poster

Thanks to the impressive progress of large-scale vision-language pretraining, recent recognition models can classify arbitrary objects in a zero-shot and open-set manner, with a surprisingly high accuracy. However, translating this success to semantic segmentation is not trivial, because this dense…

2023

ESPVR: Entity Spans Position Visual Regions for Multimodal Named Entity Recognition

EMNLP 2023long findings

Multimodal Named Entity Recognition (MNER) uses visual information to improve the performance of text-only Named Entity Recognition (NER). However, existing methods for acquiring local visual information suffer from certain limitations: (1) using an attention-based method to extract visual regions r…

Cited by 0SourceScholar
2023

EfficientViT: Memory Efficient Vision Transformer With Cascaded Group Attention

CVPR 2023poster

Vision transformers have shown great success due to their high model capabilities. However, their remarkable performance is accompanied by heavy computation costs, which makes them unsuitable for real-time applications. In this paper, we propose a family of high-speed vision transformers named Effic…

2023

Robotic Barrier Construction through Weaved, Inflatable Tubes

IROS 2023poster

In this article, we present a mechanism and related path planning algorithm to construct light-duty barriers out of extruded, inflated tubes weaved around existing environmental features. Our extruded tubes are based on everted vine-robots and in this context, we present a new method to steer their…

Cited by 1SourceScholar
2022

Generalizing to New Domains by Mapping Natural Language to Lifted LTL

ICRA 2022poster

Recent work on using natural language to specify commands to robots has grounded that language to LTL. However, mapping natural language task specifications to LTL task specifications using language models require probability distributions over finite vocabulary. Existing state-of-the-art methods ha…

Cited by 16SourceScholar
2022

Maintaining Reasoning Consistency in Compositional Visual Question Answering

CVPR 2022poster

A compositional question refers to a question that contains multiple visual concepts (e.g., objects, attributes, and relationships) and requires compositional reasoning to answer. Existing VQA models can answer a compositional question well, but cannot work well in terms of reasoning consistency in…

Cited by 29PDFcodeScholar
2022

SCAN: Cross Domain Object Detection with Semantic Conditioned Adaptation

AAAI 2022technical

The domain gap severely limits the transferability and scalability of object detectors trained in a specific domain when applied to a novel one. Most existing works bridge the domain gap by minimizing the domain discrepancy in the category space and aligning category-agnostic global features. Though…

2022

Towards Robust Adaptive Object Detection Under Noisy Annotations

CVPR 2022poster

Domain Adaptive Object Detection (DAOD) models a joint distribution of images and labels from an annotated source domain and learns a domain-invariant transformation to estimate the target labels with the given target domain images. Existing methods assume that the source domain labels are completel…

Cited by 39PDFcodeScholar
2021

A Soft Robotic Gripper with Anti-Freezing Ionic Hydrogel-Based Sensors for Learning-Based Object Recognition

ICRA 2021poster

Soft robotic grippers possess high structural compliance and adaptability for grasping objects with unknown and irregular shapes and sizes. To enable more dexterous manipulation, soft sensors with similar mechanical properties to common elastomer materials are desired to be integrated into soft grip…

Cited by 22SourceScholar
2021

Cross Scene Video Foreground Segmentation Via Co-Occurrence Probability Oriented Supervised and Unsupervised Model Interaction

ICASSP 2021accepted

Using only one deep model for cross scene video foreground segmentation is still very challenging because existing methods are scene-dependent, which restricts the consistent segmentation. In this paper, we propose a cross scene video foreground segmentation framework to extend the generalization ca…

Cited by 0SourceScholar
2021

Force-Controlled Mechanical Stimulation and Single-Neuron Fluorescence Imaging of Drosophila Larvae

RA-L 2021

Studying the neural response of a Drosophila larva to touch stimulation could decipher neural basis of the creature's danger-escaping behaviors. This letter reports force-controlled robotic mechanical stimulation and single-neuron fluorescence imaging of Drosophila larvae. A force control architectu

Cited by 8SourceScholar
2020

An SEM-Based Nanomanipulation System for Multi-Physical Characterization of Single InGaN/GaN Nanowires

IROS 2020poster

Functional nanomaterials possess exceptional multi-physical (e.g., mechanical, electrical and optical) properties compared with their bulk counterparts. To facilitate both synthesis and device applications of these nanomaterials, it is highly desired to characterize their multi-physical properties w…

Cited by 14SourceScholar
2018

Dex-Net 3.0: Computing Robust Vacuum Suction Grasp Targets in Point Clouds Using a New Analytic Model and Deep Learning

ICRA 2018poster

Vacuum-based end effectors are widely used in industry and are often preferred over parallel-jaw and multifinger grippers due to their ability to lift objects with a single point of contact. Suction grasp planners often target planar surfaces on point clouds near the estimated centroid of an object.…

Cited by 690SourcecodeScholar
2017

Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics

RSS 2017poster

To reduce data collection time for deep learning of robust robotic grasp plans, we explore training from a synthetic dataset of 6.7 million point clouds, grasps, and robust analytic grasp metrics generated from thousands of 3D models from Dex-Net 1.0 in randomized poses on a table. We use the resul…

Cited by 1474SourcePDFScholar
2017

Regulating surface traction of a soft robot through electrostatic adhesion control

IROS 2017poster

This paper reports the electrostatic regulation of surface traction of a quadruped soft robot to improve its locomotion efficiency. The soft robot, containing five pneumatic channel networks (PneuNets) in different parts of its body, is actuated to achieve undulated locomotion. Electrostatic adhesio…

Cited by 17SourceScholar
2016

A model compensation-prediction scheme for control of micromanipulation systems with a single feedback loop

ICRA 2016

Many micromanipulation systems employ sensorless actuators and possess unknown modeling errors, feedback measurement noise, and time delays. Conventional modelbased control schemes ignore some of these characteristics, and thus sacrifice the control performance of the system. This paper presents a n

Cited by 6SourceScholar
2015

An automated robotic system for high-speed microinjection of Caenorhabditis elegans

ICRA 2015poster

The tiny nematode worm Caenorhabditis elegans has long been a popular model organism for genetic, developmental, and biochemical studies in which worm microinjection plays a critical role. This paper presents an automated robotic system for high-speed injection of C. elegans with an efficiency more…

Cited by 12SourceScholar
2015

Switched fuzzy-PD control of contact forces in robotic micromanipulation of Drosophila larvae

ICRA 2015poster

Force sensing and control are of paramount importance in robotic micromanipulation. A contact force regulator capable of accurately applying mechanical stimuli to a live Drosophila larva could greatly facilitate mechanobiology research on Drosophila and may eventually lead to novel discoveries in me…

Cited by 5SourceScholar