← Search

Kun Xu

46 accepted papers

2026

A Dual-Adhesion-Enhanced Soft Gripper with Microwedge Adhesives and SMA-Driven Microspines

ICRA 2026poster

,软握把因其适应性和安全性而备受推崇,但 它们固有的柔软性常常导致在重物下抓握失败 很多。大多数增强附着力的握把依赖单一附着 针对光滑或粗糙表面量身定制的策略。蜥蜴, 但在非结构化环境中,有效导航时,通过以下方式 基于 地表状况。灵感来自混合粘附策略 壁虎和变色龙,本研究展示了一种仿生的软抓握器 它集成了微楔干胶和SMA驱动的微棘。 微楔胶提供可控的附着力,保证平滑 而SMA驱动的微棘则延伸用于粗糙表面 粘附和回放以避免干扰。优化模型为 开发目的是确定最优链路维度,提升抓取能力 性能方面,力和半径。实验结果 各种表面验证了其有效

Cited by 0SourceScholar
2026

A Mole-Inspired Scratch-Digging Robot for Granular Media Traversal

RA-L 2026

This letter proposes a mole-inspired scratch-digging robot to investigate four-limb subsurface locomotion in granular media. The robot integrated a conical head, a rigid torso, a hybrid crank-rocker and crank-slider forelimb that reproduces scratch-digging strokes, and a two-degree-of-freedom (DOF)

Cited by 0SourceScholar
2026

BDRP: A Binary Divisive Recursive Planner for Path Planning

RA-L 2026

Narrow passage scenarios pose significant challenges for path planning, especially for tasks requiring real-time performance. Traditional asymptotically converging sampling-based planners (SBPs) often exhibit poor initial path quality and slow convergence, limiting their ability to efficiently const

Cited by 0SourceScholar
2026

CAMEL: Confidence-Gated Reflection for Reward Modeling

ICML 2026poster

Reward models play a fundamental role in aligning large language models with human preferences. Existing methods predominantly follow two paradigms: scalar discriminative preference models, which are efficient but lack interpretability, and generative judging models, which offer richer reasoning at …

Cited by 0SourceScholar
2026

Continuous-Time Optical Flow Estimation from Asynchronous Event-Frame Streams for Embedded Systems

ICRA 2026poster

Bioinspired event cameras, with their high temporal resolution, low power consumption, and inherent motion responsiveness, have been widely adopted for fundamental vision tasks in robotics, notably optical flow estimation. Recent studies have shown that incorporating complementary frame data can sig…

Cited by 0Scholar
2026

Generating Attribute-Aware Human Motions from Textual Prompt

AAAI 2026technical

Text-driven human motion generation has recently attracted considerable attention, allowing models to generate human motions based on textual descriptions. However, current methods neglect the influence of human attributes—such as age, gender, weight, and height—which are key factors shaping human m

Cited by 0SourcePDFScholar
2025

A Dual-Adhesion-Enhanced Soft Gripper With Microwedge Adhesives and SMA-Driven Microspines

RA-L 2025

Soft grippers are highly valued for their adaptability and safety, but their inherent softness often leads to grasping failure under heavy loads. Most adhesion-enhanced grippers rely on single-adhesion strategies tailored for either smooth or rough surfaces. Lizards, however, effectively navigate in

Cited by 0SourceScholar
2025

A Mole-inspired Incisor-Burrowing Robotic Platform for Planetary Exploration

IROS 2025

Planetary exploration requires efficient methods for subsurface sampling, especially in extreme energy limitations. Traditional drilling methods are often energy intensive and require large platforms, limiting their applicability. Bio-inspired burrowing techniques, inspired by animals like moles, of

Cited by 0SourceScholar
2025

Following the Autoregressive Nature of LLM Embeddings via Compression and Alignment

EMNLP 2025

A new trend uses LLMs as dense text encoders via contrastive learning. However, since LLM embeddings predict the probability distribution of the next token, they are inherently generative and distributive, conflicting with contrastive learning, which requires embeddings to capture full-text semantic

2025

Granularity-Adaptive Spatial Evidence Tokenization for Video Question Answering

AAAI 2025technical

Video question answering plays a vital role in computer vision, and recent advances in large language models have further propelled the development of this field. However, existing video question answering techniques often face limitations in grasping fine-grained video content in spatial dimensions…

Cited by 0SourcePDFScholar
2025

Pyramidal Flow Matching for Efficient Video Generative Modeling

ICLR 2025poster

Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational de…

2024

Bionic Bird Claw Design for Grabbing and Perching Inspired by Tendon-Locking Mechanism

RA-L 2024

This letter proposes a novel bionic bird claw with a digit-locking mechanism inspired by a tendon-locking mechanism (TLM). First, the biological mechanisms of both the bird digit-locked TLM (tendon locking) and the leg automatic digit flexion mechanism (ADFM) (quick digit flexion) of birds are intro

Cited by 6SourceScholar
2024

Dynamic Interaction Control in Legged Mobile Manipulators: A Decoupled Approach

ICRA 2024poster

Legged mobile manipulators are receiving much more attention. Mobile platforms can infinitely expand the workspace of robotic arms, providing more possibilities for robot application scenarios. Compared with wheeled mobile manipulators, legged mobile manipulators have higher requirements for coopera…

Cited by 1SourceScholar
2024

Enhancing VIO Robustness Under Sudden Lighting Variation: A Learning-Based IMU Dead-Reckoning for UAV Localization

RA-L 2024

Visual Inertial Odometry (VIO) is commonly used for real-time Unmanned Aerial Vehicle (UAV) localization. However, the performance of VIO significantly deteriorates when UAV encounters sudden lighting variation in the environment, which poses a significant risk during flight. To address this issue w

Cited by 13SourceScholar
2024

Harder Task Needs More Experts: Dynamic Routing in MoE Models

ACL 2024long

In this paper, we introduce a novel dynamic expert selection framework for Mixture of Experts (MoE) models, aiming to enhance computational efficiency and model performance by adjusting the number of activated experts based on input difficulty. Unlike existing MoE approaches that rely on fixed TopK…

2024

Limited Information Aggregation for Collaborative Driving in Multi-Agent Autonomous Vehicles

RA-L 2024

Multi-agent reinforcement learning (MARL) methods have emerged as a promising solution for multi-agent collaborative driving in the intersection and roundabout scenarios. However, these methods need large amounts of training data obtained from the interaction with the driving simulator, and learning

Cited by 13SourceScholar
2024

Probing Multimodal Large Language Models for Global and Local Semantic Representations

COLING 2024main

The advancement of Multimodal Large Language Models (MLLMs) has greatly accelerated the development of applications in understanding integrated texts and images. Recent works leverage image-caption datasets to train MLLMs, achieving state-of-the-art performance on image-to-text tasks. However, there…

2024

RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance

NeurIPS 2024poster

Customizing diffusion models to generate identity-preserving images from user-provided reference images is an intriguing new problem. The prevalent approaches typically require training on extensive domain-specific images to achieve identity preservation, which lacks flexibility across different use…

2024

Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization

ICLR 2024poster

Recently, the remarkable advance of the Large Language Model (LLM) has inspired researchers to transfer its extraordinary reasoning capability to both vision and language data. However, the prevailing approaches primarily regard the visual input as a prompt and focus exclusively on optimizing the te…

2024

Unlocking Versatile Locomotion: A Novel Quadrupedal Robot with 4-DoFs Legs for Roller Skating

ICRA 2024poster

Roller skating with passive wheels on a quadrupedal robot is more efficient than traditional walking. However, the typical mammalian quadruped robot with 3-DoFs legs can only perform one dynamic roller skating gait and has difficulty achieving turning motion. To address this limitation, we designed…

Cited by 0SourceScholar
2024

Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization

ICML 2024oral

In light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos. Compared to static images, video poses unique challenges for effective large-scale pre-training due to the modeling of its…

2022

DRG-SLAM: A Semantic RGB-D SLAM using Geometric Features for Indoor Dynamic Scene

IROS 2022poster

Visual SLAM methods based on point features have achieved acceptable results in texture-rich static scenes, but they often suffer from a deficiency of texture and the existence of dynamic objects in real indoor scenes, which limits the application of these methods. In this paper, we have presented D…

Cited by 20SourceScholar
2022

Learning a Grammar Inducer from Massive Uncurated Instructional Videos

EMNLP 2022main

Video-aided grammar induction aims to leverage video information for finding more accurate syntactic grammars for accompanying text. While previous work focuses on building systems for inducing grammars on text that are well-aligned with video content, we investigate the scenario, in which text and…

2022

Neural Color Operators for Sequential Image Retouching

ECCV 2022poster

"We propose a novel image retouching method by modeling the retouching process as performing a sequence of newly introduced trainable neural color operators. The neural color operator mimics the behavior of traditional color operators and learns pixelwise color transformation while its strength is c…

2022

Variational Graph Autoencoding as Cheap Supervision for AMR Coreference Resolution

ACL 2022long

Coreference resolution over semantic graphs like AMRs aims to group the graph nodes that represent the same entity. This is a crucial step for making document-level formal semantic representations. With annotated data on AMR coreference resolution, deep learning approaches have recently shown great…

2022

Zero-shot Cross-lingual Conversational Semantic Role Labeling

NAACL 2022findings

While conversational semantic role labeling (CSRL) has shown its usefulness on Chinese conversational tasks, it is still under-explored in non-Chinese languages due to the lack of multilingual CSRL annotations for the parser training. To avoid expensive data collection and error-propagation of trans…

2021

CSAGN: Conversational Structure Aware Graph Network for Conversational Semantic Role Labeling

EMNLP 2021main

Conversational semantic role labeling (CSRL) is believed to be a crucial step towards dialogue understanding. However, it remains a major challenge for existing CSRL parser to handle conversational structural information. In this paper, we present a simple and effective architecture for CSRL which a…

2021

Domain-Adaptive Pretraining Methods for Dialogue Understanding

ACL 2021short

Language models like BERT and SpanBERT pretrained on open-domain data have obtained impressive gains on various NLP tasks. In this paper, we probe the effectiveness of domain-adaptive pretraining objectives on downstream tasks. In particular, three objectives, including a novel objective focusing on…

Cited by 25SourcePDFScholar
2021

DyLex: Incorporating Dynamic Lexicons into BERT for Sequence Labeling

EMNLP 2021main

Incorporating lexical knowledge into deep learning models has been proved to be very effective for sequence labeling tasks. However, previous works commonly have difficulty dealing with large-scale dynamic lexicons which often cause excessive matching noise and problems of frequent updates. In this…

2021

Exophoric Pronoun Resolution in Dialogues with Topic Regularization

EMNLP 2021main

Resolving pronouns to their referents has long been studied as a fundamental natural language understanding problem. Previous works on pronoun coreference resolution (PCR) mostly focus on resolving pronouns to mentions in text while ignoring the exophoric scenario. Exophoric pronouns are common in d…

2021

Hierarchical Layout-Aware Graph Convolutional Network for Unified Aesthetics Assessment

CVPR 2021poster

Learning computational models of image aesthetics can have a substantial impact on visual art and graphic design. Although automatic image aesthetics assessment is a challenging topic by its subjective nature, psychological studies have confirmed a strong correlation between image layouts and percei…

Cited by 98PDFcodeScholar
2021

Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation

CVPR 2021poster

Weakly supervised phrase grounding aims at learning region-phrase correspondences using only image-sentence pairs. A major challenge thus lies in the missing links between image regions and sentence phrases during training. To address this challenge, we leverage a generic object detector at training…

Cited by 84PDFcodeScholar
2021

Instance-adaptive training with noise-robust losses against noisy labels

EMNLP 2021main

In order to alleviate the huge demand for annotated datasets for different tasks, many recent natural language processing datasets have adopted automated pipelines for fast-tracking usable data. However, model training with such datasets poses a challenge because popular optimization objectives are…

Cited by 10SourcePDFScholar
2021

On the Generation of Medical Dialogs for COVID-19

ACL 2021short

Under the pandemic of COVID-19, people experiencing COVID19-related symptoms have a pressing need to consult doctors. Because of the shortage of medical professionals, many people cannot receive online consultations timely. To address this problem, we aim to develop a medical dialog system that can…

2021

RAST: Domain-Robust Dialogue Rewriting as Sequence Tagging

EMNLP 2021main

The task of dialogue rewriting aims to reconstruct the latest dialogue utterance by copying the missing content from the dialogue context. Until now, the existing models for this task suffer from the robustness issue, i.e., performances drop dramatically when testing on a different dataset. We addre…

2021

Rethinking and Reweighting the Univariate Losses for Multi-Label Ranking: Consistency and Generalization

NeurIPS 2021poster

The (partial) ranking loss is a commonly used evaluation measure for multi-label classification, which is usually optimized with convex surrogates for computational efficiency. Prior theoretical efforts on multi-label ranking mainly focus on (Fisher) consistency analyses. However, there is a gap bet…

Cited by 14SourcePDFScholar
2021

Self-Supervised Neural Networks for Spectral Snapshot Compressive Imaging

ICCV 2021poster

We consider using untrained neural networks to solve the reconstruction problem of snapshot compressive imaging (SCI), which uses a two-dimensional (2D) detector to capture a high-dimensional (usually 3D) data-cube in a compressed manner. Various SCI systems have been built in recent years to captur…

Cited by 126PDFcodeScholar
2021

Variational (Gradient) Estimate of the Score Function in Energy-based Latent Variable Models

ICML 2021spotlight

This paper presents new estimates of the score function and its gradient with respect to the model parameters in a general energy-based latent variable model (EBLVM). The score function and its gradient can be expressed as combinations of expectation and covariance terms over the (generally intracta…

2021

Video-aided Unsupervised Grammar Induction

NAACL 2021long

We investigate video-aided grammar induction, which learns a constituency parser from both unlabeled text and its corresponding video. Existing methods of multi-modal grammar induction focus on grammar induction from text-image pairs, with promising results showing that the information from static i…

2020

Bi-level Score Matching for Learning Energy-based Latent Variable Models

NeurIPS 2020poster

Score matching (SM) provides a compelling approach to learn energy-based models (EBMs) by avoiding the calculation of partition function. However, it remains largely open to learn energy-based latent variable models (EBLVMs), except some special cases. This paper presents a bi-level score matching (…

2020

Boosting Adversarial Training with Hypersphere Embedding

NeurIPS 2020poster

Adversarial training (AT) is one of the most effective defenses against adversarial attacks for deep learning models. In this work, we advocate incorporating the hypersphere embedding (HE) mechanism into the AT procedure by regularizing the features onto compact manifolds, which constitutes a lightw…

2020

Efficient Learning of Generative Models via Finite-Difference Score Matching

NeurIPS 2020poster

Several machine learning applications involve the optimization of higher-order derivatives (e.g., gradients of gradients) during training, which can be expensive with respect to memory and computation even with automatic differentiation. As a typical example in generative modeling, score matching~(S…

2020

Rethinking Softmax Cross-Entropy Loss for Adversarial Robustness

ICLR 2020poster

Previous work shows that adversarially robust generalization requires larger sample complexity, and the same dataset, e.g., CIFAR-10, which enables good standard accuracy may not suffice to train robust models. Since collecting new training data could be costly, we focus on better utilizing the give…

Cited by 214SourcecodeScholar
2020

Understanding and Stabilizing GANs’ Training Dynamics Using Control Theory

ICML 2020poster

Generative adversarial networks (GANs) are effective in generating realistic images but the training is often unstable. There are existing efforts that model the training dynamics of GANs in the parameter space but the analysis cannot directly motivate practically effective stabilizing methods. To t…

Cited by 36SourcePDFScholar
2019

Improving Adversarial Robustness via Promoting Ensemble Diversity

ICML 2019oral

Though deep neural networks have achieved significant progress on various tasks, often enhanced by model ensemble, existing high-performance models can be vulnerable to adversarial attacks. Many efforts have been devoted to enhancing the robustness of individual networks and then constructing a stra…