← Search

Xiao Chen

56 accepted papers

2026

ADAFLOW: EFFICIENT LONG VIDEO EDITING VIA ADAPTIVE ATTENTION SLIMMING AND KEYFRAME SELECTION

ICASSP 2026poster

Despite great progress, text-driven long video editing is still notoriously challenging mainly due to excessive memory overhead. Although recent efforts have simplified this task into a two-step process of keyframe translation and interpolation generation, the token-wise keyframe translation still p…

Cited by 0SourcePDFScholar
2026

AutoMat: Physics-Guided Agentic Reasoning for Solving Ill-Posed Inverse Microscopy Problems

ICML 2026poster

Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can explain similar contrast, and purely feed-forward models cannot verify physical validity. We present **AutoMat**, a failure-aware agentic *controller* that performs …

Cited by 0SourceScholar
2026

DynamicTree: Interactive Real Tree Animation via Sparse Voxel Spectrum

CVPR 2026

Generating dynamic and interactive 3D trees has wide applications in virtual reality, games, and world simulation. However, existing methods still face various challenges in generating structurally consistent and realistic 4D motion for complex real trees. In this paper, we propose DynamicTree, the

Cited by 0SourcecodeScholar
2026

FullPart: Generating each 3D Part at Full Resolution

ICLR 2026poster

Part-based 3D generation holds great potential for various applications. Previous part generators that represent parts using implicit vector-set tokens often suffer from insufficient geometric details. Another line of work adopts an explicit voxel representation but shares a global voxel grid among…

Cited by 0SourcecodeScholar
2026

GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable perceptual and reasoning abilities. However, they struggle to perceive fine-grained geometric structures, constraining their ability of geometric understanding and visual reasoning. To address this, we propose GeoTikzBrid

Cited by 0SourcecodeScholar
2026

Learning a Shape-Adaptive Assist-As-Needed Rehabilitation Policy from Therapist-Informed Input

ICRA 2026poster

Therapist-in-the-loop robotic rehabilitation has shown the promise to enhance rehabilitation outcomes by integrating the strengths of therapists and robotic systems. However, its broader adoption is limited due to insufficient interaction and limited adaptation capability. This article proposes a no…

2026

Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions

CVPR 2026

Video data is more cost-effective than motion capture data for learning 3D character controllers, yet using it to generate realistic and physically plausible motions remains challenging. Previous approaches typically rely on off-the-shelf motion reconstruction techniques to extract 3D kinematic traj

Cited by 0SourcecodeScholar
2026

MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNorm

AAAI 2026technical

Test-time adaptation (TTA) has proven effective in mitigating performance drops under single-domain distribution shifts by updating model parameters during inference. However, real-world deployments often involve mixed distribution shifts---where test samples are affected by diverse and potentially

Cited by 0SourcePDFScholar
2026

Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism

CVPR 2026

Long video understanding is a key challenge that plagues the advancement of Multimodal Large language Models (MLLMs). In this paper, we study this problem from the perspective of visual memory mechanism, and proposed a novel and training-free approach, termed Flexible Memory (FlexMem). In principle,

Cited by 0SourcecodeScholar
2026

UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes

CVPR 2026

We present UniTEX, a novel two-stage 3D texture generation framework to create high-quality, consistent textures for 3D assets. Existing approaches predominantly rely on UV-based models in the second stage to refine textures after reprojecting the generated multi-view images onto the 3D shapes, whic

Cited by 0SourcecodeScholar
2025

C2KD: Cross-layer and Cross-head Knowledge Distillation for Small Language Model-based Recommendation

ACL 2025finding

Sequential recommenders predict users’ next interactions based on historical behavior and are essential in modern recommendation systems. While Large Language Models (LLMs) show promise, their size and high inference costs limit deployment on resource-constrained devices. Small Language Models (SLMs…

Cited by 0SourcePDFScholar
2025

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

CVPR 2025poster

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging…

Cited by 23SourcePDFScholar
2025

FIRM: Flexible Interactive Reflection ReMoval

AAAI 2025technical

Removing reflection from a single image is challenging due to the absence of general reflection priors. Although existing methods incorporate extensive user guidance for satisfactory performance, they often lack the flexibility to adapt user guidance in different modalities, and dense user interacti…

2025

From One to More: Contextual Part Latents for 3D Generation

ICCV 2025poster

To generate 3D objects, early research focused on multi-view-driven approaches relying solely on 2D renderings. Recently, the 3D native latent diffusion paradigm has demonstrated superior performance in 3D generation, because it fully leverages the geometric information provided in ground truth 3D d…

2025

GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scene

ICCV 2025poster

Generalizable active mapping in complex unknown environments remains a critical challenge for mobile robots. Existing methods, constrained by limited training data and conservative exploration strategies, struggle to generalize across scenes with diverse layouts and complex connectivity. To enable s…

Cited by 0SourcePDFScholar
2025

LSU-NET: Lightweight Automatic Organs Segmentation Network for Medical Images

ICASSP 2025accepted

UNet and its variants have widespread applications in medical image segmentation. However, the substantial number of parameters and computational complexity of these models make them less suitable for use in clinical settings with limited computational resources. To address this limitation, we propo…

Cited by 0SourceScholar
2025

Label-Efficient Data Augmentation with Video Diffusion Models for Guidewire Segmentation in Cardiac Fluoroscopy

AAAI 2025technical

The accurate segmentation of guidewires in interventional cardiac fluoroscopy videos is crucial for computer-aided navigation tasks. Although deep learning methods have demonstrated high accuracy and robustness in wire segmentation, they require substantial annotated datasets for generalizability, u…

Cited by 0SourcePDFScholar
2025

Learning Humanoid Standing-up Control across Diverse Postures

RSS 2025poster

Standing-up control is crucial for humanoid robots, with the potential for integration into current locomotion and loco-manipulation systems. Existing approaches are either limited to simulations that neglect hardware constraints or rely on predefined ground-specific motion trajectories, failing to…

Cited by 6PDFScholar
2025

Model-Mediated Teleoperation with 3D Dynamic Environment Tracking (MMT-DET): A Comparative Study of Task Performance with Time-Domain Passivity Control

IROS 2025

Teleoperation with haptic feedback allows users to interact with remote environments while retaining a sense of touch. However, the stability and transparency of these systems are compromised under communication network delay. This paper presents an augmented Model-Mediated Teleoperation with 3D obj

Cited by 0SourceScholar
2025

UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation

CVPR 2025poster

Multimodal Large Language Models (MLLMs) have emerged to tackle the challenges of Visual Question Answering (VQA), sparking a new research focus on conducting objective evaluations of these models. Existing evaluation mechanisms face limitations due to the significant human workload required to desi…

Cited by 0SourcePDFScholar
2025

Where Does This Data Come From? Enhanced Source Inference Attacks in Federated Learning

IJCAI 2025

Federated learning (FL) enables collaborative model training without exposing raw data, offering a privacy-aware alternative to centralized learning. However, FL remains vulnerable to various privacy attacks that exploit shared model updates, including membership inference, property inference, and g

Cited by 0SourcePDFScholar
2024

DDGS-CT: Direction-Disentangled Gaussian Splatting for Realistic Volume Rendering

NeurIPS 2024poster

Digitally reconstructed radiographs (DRRs) are simulated 2D X-ray images generated from 3D CT volumes, widely used in preoperative settings but limited in intraoperative applications due to computational bottlenecks. Physics-based Monte Carlo simulations provide accurate representations but are extr…

Cited by 6SourcePDFScholar
2024

EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

CVPR 2024poster

In the realm of computer vision and robotics embodied agents are expected to explore their environment and carry out human instructions. This necessitates the ability to fully understand 3D scenes given their first-person observations and contextualize them into language for interaction. However tra…

2024

Enhancing the Tracking Performance of Passivity-based High-Frequency Robot Cloud Control

ICRA 2024poster

This paper addresses the migration of high-frequency robot controllers to remote computing services, which are connected via a communication channel prone to delays and packet loss. The stability of the networked system is guaranteed by ensuring passivity of each subcomponent in the interconnection,…

Cited by 0SourceScholar
2024

GenNBV: Generalizable Next-Best-View Policy for Active 3D Reconstruction

CVPR 2024poster

While recent advances in neural radiance field enable realistic digitization for large-scale scenes the image-capturing process is still time-consuming and labor-intensive. Previous works attempt to automate this process using the Next-Best-View (NBV) policy for active 3D reconstruction. However the…

2024

Online Transfer and Adaptation of Tactile Skill: A Teleoperation Framework

CoRL 2024poster

This paper presents a teleoperation framework designed for online learning and adaptation of tactile skills, which provides an intuitive interface without need for physical access to execution robot. The proposed tele-teaching approach utilizes periodical Dynamical Movement Primitives (DMP) and Recu…

Cited by 1SourceScholar
2024

SACNet: A Scattered Attention-Based Network With Feature Compensator for Visual Localization

RA-L 2024

Visual localization, an integral component of a vast array of computer applications, has been effectively resolved by scene coordinate regression (SCoRe) methods. However, due to the limited receptive field of convolutional neural networks (CNNs), current SCoRe methods have difficulty in distinguish

Cited by 4SourceScholar
2024

SPTESleepNet: Automatic Sleep Staging Model Based On Strip Patch Embeddings And Transformer Encoder

ICASSP 2024accepted

Although the research on automatic sleep staging has made great progress, there is still a certain distance from its clinical application. For a machine scoring system, in order to work in an interactive and collaborative manner with practitioners, two barriers need to be addressed, which are high a…

Cited by 0SourceScholar
2023

A Force-Sensitive Exoskeleton for Teleoperation: An Application in Elderly Care Robotics

ICRA 2023poster

With the increasing demand for new healthcare solutions and technologies, such as those resulting from the COVID-19 crisis, and the growing elderly population, exoskeletons for teleoperation are a promising solution for many future medical applications. In this context, we propose two force- sensiti…

Cited by 30SourceScholar
2023

A Passivity-based Approach on Relocating High-Frequency Robot Controller to the Edge Cloud

ICRA 2023poster

As robots become more and more intelligent, the complexity of the algorithms behind them is increasing. Since these algorithms require high computation power from the onboard robot controller, the weight of the robot and energy consumption increases. A promising solution to tackle this issue is to r…

Cited by 5SourceScholar
2023

Cooperative Assist-as-Needed Control for Robotic Rehabilitation: A Two-Player Game Approach

RA-L 2023

This letter presents an adaptive optimal control strategy for developing assist-as-needed robotic rehabilitation. The primary goal is to encourage patient participation and increase the effectiveness of training sessions by minimizing robot intervention while following a predefined path. To achieve

Cited by 25SourceScholar
2023

Identification of a Generalized Base Inertial Parameter Set of Robotic Manipulators Considering Mounting Configurations

ICRA 2023poster

Identifying the inertial parameters of real robotic manipulators is a fundamental step towards realistic modeling and better controller performances, which is crucial for safe human-robot interaction. Our work introduces a novel framework for identifying a generalized set of base inertial parameters…

Cited by 5SourceScholar
2023

Improving End-to-End Speech Processing by Efficient Text Data Utilization with Latent Synthesis

EMNLP 2023long findings

Training a high performance end-to-end speech (E2E) processing model requires an enormous amount of labeled speech data, especially in the era of data-centric artificial intelligence. However, labeled speech data are usually scarcer and more expensive for collection, compared to textual data. We pro…

Cited by 0SourceScholar
2022

M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus

NeurIPS 2022accept

The lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Multi-singer Mandarin singing collection with elaborately annotated Musical scores a…

Cited by 89SourcePDFScholar
2022

On the Communication Channel in Bilateral Teleoperation: An Experimental Study for Ethernet, WiFi, LTE and 5G

IROS 2022poster

Teleoperated robots are believed to play an important role for future applications in industry, medicine and other domains. Examples for this are remote assembly and maintenance, surgery, diagnosis or deep-sea and space exploration. Such applications are made possible by state-of-the-art tactile man…

Cited by 9SourceScholar
2022

Origami Robot Self-folding by Magnetic Induction

IROS 2022poster

Inspired by the traditional art of paper folding, origami, autonomous production of 3D structures from 2D sheets can be achieved by the implementation of self-folding techniques. One technique to achieve such transformation is the usage of thermo-responsive smart materials such as self-folding polym…

Cited by 6SourceScholar
2022

Robust Landmark-Based Stent Tracking in X-Ray Fluoroscopy

ECCV 2022poster

"In clinical procedures of angioplasty (i.e., open clogged coronary arteries), devices such as balloons and stents need to be placed and expanded in arteries under the guidance of X-ray fluoroscopy. Due to the limitation of X-ray dose, the resulting images are often noisy. To check the correct place…

Cited by 7SourcePDFScholar
2022

Tactile Robotic Telemedicine for Safe Remote Diagnostics in Times of Corona: System Design, Feasibility and Usability Study

RA-L 2022

The current crisis surrounding the COVID-19 pandemic demonstrates the amount of responsibility and the workload on our healthcare system and, above all, on the medical staff around the world. In this work, we propose a promising approach to overcome this problem using robot-assisted telediagnostics,

Cited by 15SourceScholar
2022

bert2BERT: Towards Reusable Pretrained Language Models

ACL 2022long

In recent years, researchers tend to pre-train ever-larger language models to explore the upper limit of deep models. However, large language model pre-training costs intensive computational resources, and most of the models are trained from scratch without reusing the existing pre-trained models, w…

Cited by 86SourcePDFScholar
2021

A Streaming End-to-End Framework For Spoken Language Understanding

IJCAI 2021poster

End-to-end spoken language understanding (SLU) recently attracted increasing interest. Compared to the conventional tandem-based approach that combines speech recognition and language understanding as separate modules, the new approach extracts users' intentions directly from the speech signals, res…

Cited by 12SourcePDFScholar
2021

AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language Models

ACL 2021long

Pre-trained language models (PLMs) have achieved great success in natural language processing. Most of PLMs follow the default setting of architecture hyper-parameters (e.g., the hidden dimension is a quarter of the intermediate dimension in feed-forward sub-networks) in BERT. Few studies have been…

2021

DyLex: Incorporating Dynamic Lexicons into BERT for Sequence Labeling

EMNLP 2021main

Incorporating lexical knowledge into deep learning models has been proved to be very effective for sequence labeling tasks. However, previous works commonly have difficulty dealing with large-scale dynamic lexicons which often cause excessive matching noise and problems of frequent updates. In this…

2021

Fcl-Taco2: Towards Fast, Controllable and Lightweight Text-to-Speech Synthesis

ICASSP 2021accepted

Sequence-to-sequence (seq2seq) learning has greatly improved text-to-speech (TTS) synthesis performance, but effective implementation on resource-restricted devices remains challenging as seq2seq models are usually computationally expensive and memory intensive. To achieve fast inference speed and s…

Cited by 0SourceScholar
2021

GhostBERT: Generate More Features with Cheap Operations for BERT

ACL 2021long

Transformer-based pre-trained language models like BERT, though powerful in many tasks, are expensive in both memory and computation, due to their large number of parameters. Previous works show that some parameters in these models can be pruned away without severe accuracy drop. However, these redu…

Cited by 26SourcePDFScholar
2021

Multi-Speaker Emotional Speech Synthesis with Fine-Grained Prosody Modeling

ICASSP 2021accepted

We present an end-to-end system for multi-speaker emotional speech synthesis. In particular, our system learns emotion classes from just two speakers then generalizes these classes to other speakers from whom no emotional data was seen. We address the problem by integrating disentangled, fine-graine…

Cited by 0SourceScholar
2021

Vset: A Multimodal Transformer for Visual Speech Enhancement

ICASSP 2021accepted

The transformer architecture has shown great capability in learning long-term dependency and works well in multiple domains. However, transformer has been less considered in audio-visual speech enhancement (AVSE) research, partly due to the convention that treats speech enhancement as a short-time s…

Cited by 0SourceScholar
2020

DynaBERT: Dynamic BERT with Adaptive Width and Depth

NeurIPS 2020spotlight

The pre-trained language models like BERT, though powerful in many natural language processing tasks, are both computation and memory expensive. To alleviate this problem, one approach is to compress them for specific tasks before deployment. However, recent works on BERT compression usually compres…

2020

FOAL: Fast Online Adaptive Learning for Cardiac Motion Estimation

CVPR 2020poster

Motion estimation of cardiac MRI videos is crucial for the evaluation of human heart anatomy and function. Recent researches show promising results with deep learning-based methods. In clinical deployment, however, they suffer dramatic performance drops due to mismatched distributions between traini…

Cited by 64PDFScholar
2018

Learning Efficient Single-stage Pedestrian Detectors by Asymptotic Localization Fitting

ECCV 2018poster

Though Faster R-CNN based two-stage detectors have witnessed significant boost in pedestrian detection accuracy, it is still slow for practical applications. One solution is to simplify this working flow as a single-stage detector. However, current single-stage detectors (e.g. SSD) have not presente…