← Search

Tao Zhang

148 accepted papers

2026

A Dual-Adhesion-Enhanced Soft Gripper with Microwedge Adhesives and SMA-Driven Microspines

ICRA 2026poster

,软握把因其适应性和安全性而备受推崇,但 它们固有的柔软性常常导致在重物下抓握失败 很多。大多数增强附着力的握把依赖单一附着 针对光滑或粗糙表面量身定制的策略。蜥蜴, 但在非结构化环境中,有效导航时,通过以下方式 基于 地表状况。灵感来自混合粘附策略 壁虎和变色龙,本研究展示了一种仿生的软抓握器 它集成了微楔干胶和SMA驱动的微棘。 微楔胶提供可控的附着力,保证平滑 而SMA驱动的微棘则延伸用于粗糙表面 粘附和回放以避免干扰。优化模型为 开发目的是确定最优链路维度,提升抓取能力 性能方面,力和半径。实验结果 各种表面验证了其有效

Cited by 0SourceScholar
2026

A Mole-Inspired Scratch-Digging Robot for Granular Media Traversal

RA-L 2026

This letter proposes a mole-inspired scratch-digging robot to investigate four-limb subsurface locomotion in granular media. The robot integrated a conical head, a rigid torso, a hybrid crank-rocker and crank-slider forelimb that reproduces scratch-digging strokes, and a two-degree-of-freedom (DOF)

Cited by 0SourceScholar
2026

A Transendoscopic Telerobotic System Using Heterogeneous Flexible Manipulators for Bimanual Endoscopic Submucosal Dissection

ICRA 2026poster

Endoscopic submucosal dissection (ESD) is an effective technique to resect early cancers in the gastrointestinal (GI) tract. Bimanual telerobotic manipulation is an approach to performing ESD intuitively and efficiently, which requires two robotic instruments with flexibility, stiffness, dexterity a…

Cited by 0Scholar
2026

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators

ICLR 2026poster

Large Language Models (LLMs) have shown impressive performance across diverse domains, with code generation emerging as a particularly prominent application. However, existing benchmarks designed to evaluate code generation exhibit several critical limitations. First, most rely on manual annotations…

Cited by 0SourcecodeScholar
2026

Breaking Safety Paradox with Feasible Dual Policy Iteration

ICLR 2026poster

Achieving zero constraint violations in safe reinforcement learning poses a significant challenge. We discover a key obstacle called the safety paradox, where improving policy safety reduces the frequency of constraint-violating samples, thereby impairing feasibility function estimation and ultimate…

Cited by 0SourceScholar
2026

Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning

ICML 2026poster

Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is effective, it raises concerns regarding privacy and storage. Thus, generative replay has emerged as a viable alternative, synthesizing old data usin…

Cited by 0SourceScholar
2026

Differentially Private Subspace Fine-Tuning for Large Language Models

AAAI 2026technical

Fine-tuning large language models on downstream tasks is crucial for realizing their cross-domain potential but often relies on sensitive data, raising privacy concerns. Differential privacy (DP) offers rigorous privacy guarantees and has been widely adopted in fine-tuning; however, naively injecti

Cited by 0SourcePDFScholar
2026

EMR-Diff: Edge-aware Multimodal Residual Diffusion Model for Hyperspectral Image Super-resolution

CVPR 2026

Hardware constraints make it challenging to simultaneously acquire hyperspectral images (HSIs) with both high spatial and high spectral resolutions. A promising solution is to fuse low-resolution HSI (LR-HSI) with high-resolution multispectral images (HR-MSI) to generate high-resolution HSI (HR-HSI)

Cited by 0SourcecodeScholar
2026

EasyMimic: A Low-Cost Framework for Robot Imitation Learning from Human Videos

ICRA 2026poster

Robot imitation learning is often hindered by the high cost of collecting large-scale, real-world data. This challenge is especially significant for low-cost robots designed for home use, as they must be both user-friendly and affordable. To address this, we propose the EasyMimic framework, a lowcos…

2026

Eliminate Distance Differences Induced by Backdoor Attacks: Layer-Selective Training and Clipping to Mask Backdoor Models

CVPR 2026

Federated learning (FL) enables a central server to collaboratively train a global model with multiple clients while preserving data privacy. However, the distributed nature of FL makes the paradigm vulnerable to backdoor attacks, as proved by numerous recent studies. Although existing studies impro

Cited by 0SourceScholar
2026

Enhancing Unregistered Hyperspectral Image Super-Resolution via Unmixing-based Abundance Fusion Learning

CVPR 2026

Unregistered hyperspectral image (HSI) super-resolution (SR) typically aims to enhance a low-resolution HSI using an unregistered high-resolution reference image. In this paper, we propose an unmixing-based fusion framework that decouples spatial-spectral information to simultaneously mitigate the i

Cited by 0SourcecodeScholar
2026

Federated Manifold Learning (FML): Tackling Domain Heterogeneity with Structural Knowledge Transfer

ICML 2026poster

Federated Learning (FL) faces significant challenges due to domain heterogeneity, where data from different clients exhibit substantial statistical shifts that hinder the generalization of the global model. Although existing methods attempt to mitigate this by exchanging class prototypes, they fall …

Cited by 0SourceScholar
2026

GenCP: Towards Generative Modeling Paradigm of Coupled physics

ICLR 2026poster

Real-world physical systems are inherently complex, often involving the coupling of multiple physics, making their simulation both highly valuable and challenging. Many mainstream approaches face challenges when dealing with decoupled data. Besides, they also suffer from low efficiency and fidelity…

Cited by 0SourcecodeScholar
2026

Grasp Any Region: Prompting MLLM to Understand the Dense World

ICLR 2026poster

While Multimodal Large Language Models (MLLMs) excel at holistic understanding, they struggle with the dense world, i.e., complex scenes requiring fine-grained analysis of intricate details and object inter-relationships. Region-level MLLMs have been a promising step. However, previous attempts are…

Cited by 0SourcecodeScholar
2026

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

CVPR 2026

In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making them impractical for real-time or resource-constrained applications.Visual token pruning is a promising strategy for reducing the cost of MLLM inferen

Cited by 0SourcecodeScholar
2026

IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting

CVPR 2026

Recent advances in multimodal large language models (MLLMs) have led to impressive progress across various benchmarks. However, their capability in understanding infrared images remains unexplored. To address this gap, we introduce **IF-Bench**, the first high-quality benchmark designed for evaluati

Cited by 0SourcecodeScholar
2026

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

ICASSP 2026poster

Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instructions. To address this, we adopt a reinforcement learning (RL) based post-training strategy for MLLMs in multi-image grou…

Cited by 0SourcePDFScholar
2026

MMhops-R1: Multimodal Multi-hop Reasoning

AAAI 2026technical

The ability to perform multi-modal multi-hop reasoning by iteratively integrating information across various modalities and external knowledge is critical for addressing complex real-world challenges. However, existing Multi-modal Large Language Models (MLLMs) are predominantly limited to single-ste

Cited by 0SourcePDFScholar
2026

Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation

ICML 2026poster

Reverse Chain-of-Thought Generation (RCG) synthesizes reasoning traces from query-answer pairs, but runs the risk of producing post-hoc rationalizations: when models can see the answer during generation, the answer serves as a cognitive anchor that shapes the entire explanation. We formalize this ph…

Cited by 0SourceScholar
2026

MoM: Linear Sequence Modeling with Mixture-of-Memories

ICLR 2026poster

Linear sequence modeling methods, such as linear attention, state space modeling, and linear RNNs, offer significant efficiency improvements by reducing the complexity of training and inference. However, these methods typically compress the entire input sequence into a single fixed-size memory state…

Cited by 0SourcecodeScholar
2026

Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence

ICML 2026poster

Most video reasoning models only generate textual reasoning traces without indicating when and where key evidence appears. Recent models such as OpenAI-o3 have sparked wide interest in evidence-centered reasoning for images, yet extending this ability to videos is more challenging due to the need fo…

Cited by 43SourceScholar
2026

PHOTONS: Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views

AAAI 2026technical

We present PHOTONS (Pose-Free Human-Centric Photo-Realistic Real-Time Novel View Synthesis from Sparse Views), a real-time framework for novel view synthesis without requiring camera calibration. Our method reconstructs consistent 3D Gaussian point clouds and synthesizes 2K photo-realistic novel vie

Cited by 0SourcePDFScholar
2026

Physically-Guided Optical Inversion Enable Non-Contact Side-Channel Attack on Isolated Screens

ICLR 2026poster

Noncontact exfiltration of electronic screen content poses a security challenge, with side-channel incursions as the principal vector. We introduce an optical projection side-channel paradigm that confronts two core instabilities: (i) the near-singular Jacobian spectrum of projection mapping breache…

Cited by 0SourceScholar
2026

S3LAM: Surfel Splatting SLAM for Geometrically Accurate Tracking and Mapping

ICRA 2026poster

We propose S3LAM, a novel RGB-D SLAM system that leverages 2D surfel splatting to achieve geometrically accurate scene representations for simultaneous tracking and mapping. Unlike existing 3DGS-based SLAM approaches that rely on 3D Gaussian ellipsoids, we utilize 2D Gaussian surfels as primitives f…

Cited by 0codeScholar
2026

SAMTok: Representing Any Mask with Two Words

CVPR 2026

Pixel-wise capabilities are essential for building interactive intelligent systems. However, pixel-wise multi-modal LLMs (MLLMs) remain difficult to scale due to complex region-level encoders, specialized segmentation decoders, and incompatible training objectives. To address these challenges, we pr

Cited by 0SourcecodeScholar
2026

SceneDirector: Bridging Explicit Geometry and Generative Priors for Unified Driving Scene Editing

ICML 2026poster

Validating autonomous driving systems requires diverse scenarios, yet real-world data collection is biased and costly. Editing existing driving logs offers a scalable solution, but simultaneously editing objects and ego-trajectory—termed unified editing—remains challenging. Current methods face an i…

Cited by 0SourceScholar
2026

Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Method

ICLR 2026poster

Models like OpenAI-o3 pioneer visual grounded reasoning by dynamically ref- erencing visual regions, just like human “thinking with images”. However, no benchmark exists to evaluate these capabilities holistically. To bridge this gap, we propose TreeBench (Traceable Evidence Evaluation Benchmark), a…

Cited by 58SourcecodeScholar
2026

VFScale: Intrinsic Reasoning through Verifier-Free Test-time Scalable Diffusion Model

ICLR 2026poster

Inspired by human SYSTEM 2 thinking, LLMs excel at complex reasoning tasks via extended Chain-of-Thought. However, similar test-time scaling for diffusion models to tackle complex reasoning remains largely unexplored. From existing work, two primary challenges emerge in this setting: (i) the depende…

Cited by 0SourcecodeScholar
2026

Wavefront-Constrained Passive Obscured Object Detection

AAAI 2026technical

Accurately localizing and segmenting obscured objects from faint light patterns beyond the field of view is highly challenging due to multiple scattering and medium-induced perturbations. Most existing methods, based on real-valued modeling or local convolutional operations, are inadequate for captu

Cited by 0SourcePDFScholar
2026

Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs

CVPR 2026

User interface to code (UI2Code) aims to generate executable code that can faithfully reconstruct a given input UI. Prior work focuses largely on web pages and mobile screens, leaving app widgets underexplored. Unlike web or mobile UIs with rich hierarchical context, widgets are compact, context-fre

Cited by 0SourcecodeScholar
2025

4DRC-OC: Online Calibration of 4D Millimeter Wave Radar-Camera With Depth Map Assistance

RA-L 2025

The online calibration of 4D millimeter-wave radar and camera is crucial for advancing perception and SLAM technologies in complex environments. It eliminates the reliance on manual labeling, offering real-time and convenience. However, the sparse nature of 4D radar point clouds presents challenges

Cited by 2SourceScholar
2025

A Dual-Adhesion-Enhanced Soft Gripper With Microwedge Adhesives and SMA-Driven Microspines

RA-L 2025

Soft grippers are highly valued for their adaptability and safety, but their inherent softness often leads to grasping failure under heavy loads. Most adhesion-enhanced grippers rely on single-adhesion strategies tailored for either smooth or rough surfaces. Lizards, however, effectively navigate in

Cited by 0SourceScholar
2025

A Mole-inspired Incisor-Burrowing Robotic Platform for Planetary Exploration

IROS 2025

Planetary exploration requires efficient methods for subsurface sampling, especially in extreme energy limitations. Traditional drilling methods are often energy intensive and require large platforms, limiting their applicability. Bio-inspired burrowing techniques, inspired by animals like moles, of

Cited by 0SourceScholar
2025

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

ICCV 2025poster

Recent advancements in multimodal large language models (MLLM) have shown a strong ability in visual perception, reasoning abilities, and vision-language understanding. However, the visual matching ability of MLLMs is rarely studied, despite finding the visual correspondence of objects is essential…

2025

CFBench: A Comprehensive Constraints-Following Benchmark for LLMs

ACL 2025long

The adeptness of Large Language Models (LLMs) in comprehending and following natural language instructions is critical for their deployment in sophisticated real-world applications. Existing evaluations mainly focus on fragmented constraints or narrow scenarios, but they overlook the comprehensivene…

2025

CL-DiffPhyCon: Closed-loop Diffusion Control of Complex Physical Systems

ICLR 2025poster

The control problems of complex physical systems have broad applications in science and engineering. Previous studies have shown that generative control methods based on diffusion models offer significant advantages for solving these problems. However, existing generative control approaches face ch…

2025

Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization

NeurIPS 2025poster

Preference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face…

Cited by 0SourcecodeScholar
2025

ECERC: Evidence-Cause Attention Network for Multi-Modal Emotion Recognition in Conversation

ACL 2025long

Multi-modal Emotion Recognition in Conversation (MMERC) aims to identify speakers’ emotional states using multi-modal conversational data, significant for various domains. MMERC requires addressing emotional causes: contextual factors that influence emotions, alongside emotional evidence directly ex…

Cited by 0SourcePDFScholar
2025

Effective Techniques for Scaling Audio Encoder Pretraining

ICASSP 2025accepted

This work presents advancements in audio pretraining objectives designed to generate semantically rich embeddings, capable of addressing a wide range of audio-related tasks. Despite significant progress in the field, current methods often emphasize full fine-tuning in downstream applications, which…

Cited by 0SourceScholar
2025

EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues

NAACL 2025long

Role-playing agents (RPAs) powered by large language models (LLMs) have been widely utilized in dialogue systems for their capability to deliver personalized interactions. Current evaluations of RPAs mainly focus on personality fidelity, tone imitation, and knowledge consistency, while overlooking e…

Cited by 0SourcePDFScholar
2025

From Uncertain to Safe: Conformal Adaptation of Diffusion Models for Safe PDE Control

ICML 2025poster

The application of deep learning for partial differential equation (PDE)-constrained control is gaining increasing attention. However, existing methods rarely consider safety requirements crucial in real-world applications. To address this limitation, we propose Safe Diffusion Models for PDE Control…

2025

GLiM: Integrating Graph Transformer and LLM for Document-Level Biomedical Relation Extraction with Incomplete Labeling

ACL 2025finding

Document-level relation extraction (DocRE) identifies relations between entities across an entire document. However, as the number and complexity of entities and entity-pair relations grow, the problem space expands quadratically, causing incomplete annotations and frequent false negatives, especial…

2025

GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models

ACL 2025long

Large Language Models (LLMs) are prone to generating content that exhibits gender biases, raising significant ethical concerns. Alignment, the process of fine-tuning LLMs to better align with desired behaviors, is recognized as an effective approach to mitigate gender biases. Although proprietary LL…

2025

High-dimension Prototype is a Better Incremental Object Detection Learner

ICLR 2025poster

Incremental object detection (IOD), surpassing simple classification, requires the simultaneous overcoming of catastrophic forgetting in both recognition and localization tasks, primarily due to the significantly higher feature space complexity. Integrating Knowledge Distillation (KD) would mitigate…

Cited by 0SourcePDFScholar
2025

JPDS-NN: Reinforcement Learning-Based Dynamic Task Allocation for Agricultural Vehicle Routing Optimization

IROS 2025

The Entrance Dependent Vehicle Routing Problem (EDVRP) is a variant of the Vehicle Routing Problem (VRP) where the scale of cities influences routing outcomes, necessitating consideration of their entrances. This paper addresses EDVRP in agriculture, focusing on multi-parameter vehicle planning for

Cited by 1SourceScholar
2025

M2PDE: Compositional Generative Multiphysics and Multi-component PDE Simulation

ICML 2025poster

Multiphysics simulation, which models the interactions between multiple physical processes, and multi-component simulation of complex structures are critical in fields like nuclear and aerospace engineering. Previous studies use numerical solvers or ML-based surrogate models for these simulations. H…

2025

MAITFuse: Multi-Dimension Adaptive Interaction Transform Network For Infrared-visible Image Fusion

ICASSP 2025accepted

In recent years, Transformers have achieved significant success in image fusion. These methods utilize self-attention mechanism across different spatial or channel dimensions and have demonstrated impressive performance. However, existing methods only optimize along a single dimension and struggle t…

Cited by 0SourceScholar
2025

MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding

ICCV 2025poster

Accurate driving behavior recognition and reasoning are critical for autonomous driving video understanding. However, existing methods often tend to dig out the shallow causal, fail to address spurious correlations across modalities, and ignore the ego-vehicle level causality modeling. To overcome t…

2025

MMCD: Memory-Based Multimodal Change Detection

ICASSP 2025accepted

Single-modal change detection methods based on optical or Synthetic Aperture Radar (SAR) images face challenges such as degradation due to adverse weather or noise interference. In contrast, multimodal change detection struggles with significant domain gaps between different modalities. Inspired by…

Cited by 0SourceScholar
2025

Minimally Invasive Endotracheal Inside-Out Flexible Needle Driving System Towards Microendoscope-Guided Robotic Tracheostomy

ICRA 2025

Open tracheostomy (OT) is considered the traditional way and golden standard for treating airway obstruction patients. However, OT has many unavoidable drawbacks, including strict performing scenarios, significant scarring, and the risk of surgeon infection. Percutaneous dilation tracheostomy (PDT)

Cited by 0SourceScholar
2025

Minimizing Disparities between Real and Pseudo Queries for Unsupervised Visual Grounding

ICASSP 2025accepted

Visual grounding involves the identification and localization of image regions given textual descriptions. To reduce the manual labeling effort on region-text pairs, unsupervised visual grounding aims to generate pseudo bounding box and query pairs for training grounding models. However, there exist…

Cited by 0SourceScholar
2025

Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format Alignment

ICLR 2025poster

Adapting large language models (LLMs) to specialized domains typically requires domain-specific corpora for continual pre-training to facilitate knowledge memorization and related instructions for fine-tuning to apply this knowledge. However, this method may lead to inefficient knowledge memorizatio…

Cited by 1SourcePDFScholar
2025

Multi-Relational Geometric Regularization Framework for Multi-Modal Emotion Recognition in Conversation

ICASSP 2025accepted

Existing studies on multi-modal emotion recognition in conversation (MMERC) mainly focus on multi-modal fusion and context modeling for emotion representation, facing limitations in uncovering the intrinsic structure of emotion-related data. The existing geometric consistency regularization (GCR) te…

Cited by 0SourceScholar
2025

Navi2Gaze: Leveraging Foundation Models for Navigation and Target Gazing

IROS 2025

Task-aware navigation continues to be a challenging area of research, especially in scenarios involving open vocabulary. Previous studies primarily focus on finding suitable locations for task completion, often overlooking the importance of the robot’s pose. However, the robot’s orientation is cruci

Cited by 6SourcecodeScholar
2025

Noise Calibration and Spatial-Frequency Interactive Network for STEM Image Enhancement

CVPR 2025poster

Scanning Transmission Electron Microscopy (STEM) enables the observation of atomic arrangements at sub-angstrom resolution, allowing for atomically resolved analysis of the physical and chemical properties of materials. However, due to the effects of noise, electron beam damage, sample thickness, et…

2025

On Path to Multimodal Generalist: General-Level and General-Bench

ICML 2025oral

The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of language-based LLMs. Unlike their specialist predecessors, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially limited to understanding multiple mod…

Cited by 0SourcePDFScholar
2025

Point Cloud Mamba: Point Cloud Learning via State Space Model

AAAI 2025technical

Recently, state space models have exhibited strong global modeling capabilities and linear computational complexity in contrast to transformers. This research focuses on applying such architecture to more efficiently and effectively model point cloud data globally with linear computational complexit…

2025

RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement

NAACL 2025long

Existing large language models (LLMs) show exceptional problem-solving capabilities but might struggle with complex reasoning tasks. Despite the successes of chain-of-thought and tree-based search methods, they mainly depend on the internal knowledge of LLMs to search over intermediate reasoning ste…

2025

Robust Graph Based Social Recommendation Through Contrastive Multi-View Learning

AAAI 2025technical

Social recommendation leverages the social connections between users to mitigate the issue of data sparsity and enhance recommendation quality. Although existing related works show their effectiveness, there remain two critical questions: i) The patterns of preference interactions among users are va…

Cited by 0SourcePDFScholar
2025

Robust State Estimation for Legged Robots With Dual Beta Kalman Filter

RA-L 2025

Existing state estimation algorithms for legged robots that rely on proprioceptive sensors often overlook foot slippage and leg deformation in the physical world, leading to large estimation errors. To address this limitation, we propose a comprehensive measurement model that accounts for both foot

Cited by 6SourceScholar
2025

Sketch-based Point Cloud Generation with Diffusion Model and Pre-training Enhancement

ICASSP 2025accepted

Diffusion models, known for their success in various generative tasks like image generation and super-resolution, are applied in this study for point cloud generation, a field that has not been extensively explored due to the complexity of point clouds. We propose a novel method using a diffusion mo…

Cited by 0SourceScholar
2025

SysBench: Can LLMs Follow System Message?

ICLR 2025poster

Large Language Models (LLMs) have become instrumental across various applications, with the customization of these models to specific scenarios becoming increasingly critical. System message, a fundamental component of LLMs, is consist of carefully crafted instructions that guide the behavior of mod…

Cited by 0SourcePDFScholar
2025

THESAURUS: Contrastive Graph Clustering by Swapping Fused Gromov-Wasserstein Couplings

AAAI 2025technical

Graph node clustering is a fundamental unsupervised task. Existing methods typically train an encoder through self-supervised learning and then apply K-means to the encoder output. Some methods use this clustering result directly as the final assignment, while others initialize centroids based on th…

Cited by 0SourcePDFScholar
2025

Three-Dimension Tip Force Perception and Axial Contact Location Identification for Flexible Endoscopy Using Tissue-Compliant Soft Distal Attachment Cap Sensors

ICRA 2025

In endoluminal surgeries, inserting a flexible endo-scope is one of the fundamental procedures. During this process, vision remains the primary feedback, while the perception of tactile magnitude and location is insufficient. This limitation can hinder the clinician's efficiency when navigating the

Cited by 0SourceScholar
2025

Transferable Latent-To-Latent Locomotion Policy for Efficient and Versatile Motion Control of Diverse Legged Robots

IROS 2025

Reinforcement learning (RL) has demonstrated remarkable capability in acquiring robot skills, but learning each new skill still requires substantial data collection for training. The pretrain-and-finetune paradigm offers a promising approach for efficiently adapting to new robot entities and tasks.

Cited by 2SourceScholar
2025

UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic Segmentation

ICLR 2025poster

Pre-training techniques significantly enhance the performance of semantic segmentation tasks with limited training data. However, the efficacy under a large domain gap between pre-training (e.g. RGB) and fine-tuning (e.g. infrared) remains underexplored. In this study, we first benchmark the infrare…

2025

VCADNet: Vision-based Circular Accessible Depth Prediction for UGV Perception

IROS 2025

Circular accessible depth (CAD) provides a lightweight and robust traversability representation for autonomous navigation of unmanned ground vehicles (UGV). Aiming at the limitations of existing LiDAR-based methods in detecting low-thickness targets and executing semantic reasoning, we propose VCADN

Cited by 0SourceScholar
2025

Wavelet Diffusion Neural Operator

ICLR 2025poster

Simulating and controlling physical systems described by partial differential equations (PDEs) are crucial tasks across science and engineering. Recently, diffusion generative models have emerged as a competitive class of methods for these tasks due to their ability to capture long-term dependencies…

2025

WebQuality: A Large-scale Multi-modal Web Page Quality Assessment Dataset with Multiple Scoring Dimensions

NAACL 2025long

The assessment of web page quality plays a critical role in a range of downstream applications, yet there is a notable absence of datasets for the evaluation of web page quality. This research presents the pioneering task of web page quality assessment and introduces the first comprehensive, multi-m…

2024

A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete Labeling

AAAI 2024technical

The goal of document-level relation extraction (RE) is to identify relations between entities that span multiple sentences. Recently, incomplete labeling in document-level RE has received increasing attention, and some studies have used methods such as positive-unlabeled learning to tackle this issu…

2024

A Two-Stage Reinforcement Learning Approach for Robot Navigation in Long-range Indoor Dense Crowd Environments

IROS 2024poster

Safe and efficient mobility is vital for mobile robots navigating long-range indoor crowd environments, such as supermarkets, restaurants, and railway stations. Traditional path planning methods are challenged because of the high dynamics of pedestrians and constrained feasible regions. Existing lon…

Cited by 4SourceScholar
2024

Compositional Generative Inverse Design

ICLR 2024spotlight

Inverse design, where we seek to design input variables in order to optimize an underlying objective function, is an important problem that arises across fields such as mechanical engineering to aerospace engineering. Inverse design is typically formulated as an optimization problem, with recent wor…

2024

ControlCap: Controllable Captioning via No-Fuss Lexicon

ICASSP 2024accepted

Controllable captioning has received much attention in recent years. Although substantial progress has been made, existing methods still face challenges such as high training costs, intricate control signals and limited control capabilities. To address these issues, we propose a straightforward and…

Cited by 0SourceScholar
2024

Decoupling Representation and Knowledge for Few-Shot Intent Classification and Slot Filling

AAAI 2024technical

Few-shot intent classification and slot filling are important but challenging tasks due to the scarcity of finely labeled data. Therefore, current works first train a model on source domains with sufficiently labeled data, and then transfer the model to target domains where only rarely labeled data…

Cited by 0SourcePDFScholar
2024

Design, Analysis, and Tests of an Impulsive Aquatic Jumping Mechanism Inspired by Frog Limbs

RA-L 2024

The Southeast Asia frog (Euphlyctis hexadactylus), a natural jumping expert in water, demonstrates exceptional jumping capability by contracting its leg muscles and coordinating its flippers to propel itself out of water, significantly aiding its hunting activities. This bionic behavior serves as a

Cited by 2SourceScholar
2024

DexCatch: Learning to Catch Arbitrary Objects with Dexterous Hands

CoRL 2024poster

Achieving human-like dexterous manipulation remains a crucial area of research in robotics. Current research focuses on improving the success rate of pick-and-place tasks. Compared with pick-and-place, throwing-catching behavior has the potential to increase the speed of transporting objects to thei…

Cited by 4SourceScholar
2024

DiffPhyCon: A Generative Approach to Control Complex Physical Systems

NeurIPS 2024poster

Controlling the evolution of complex physical systems is a fundamental task across science and engineering. Classical techniques suffer from limited applicability or huge computational costs. On the other hand, recent deep learning and reinforcement learning-based approaches often struggle to optim…

2024

Exploring Object-Centered External Knowledge for Fine-Grained Video Paragraph Captioning

ICASSP 2024accepted

Video paragraph captioning task aims to generate a detailed, fluent and relevant paragraph for a given video. Prior studies often focus on isolating visual objects (potential main components in a sentence) from the overall video content. They rarely explore the latent semantic relations between obje…

Cited by 0SourceScholar
2024

FusedNet: End-to-End Mobile Robot Relocalization in Dynamic Large-Scale Scene

RA-L 2024

To improve robot relocalization accuracy in both static and dynamic environments, we introduce a novel network, FusedNet, which incorporates a cross-attention to fuse global and local image features for end-to-end relocalization. This approach relies solely on a monocular camera sensor that is fixed

Cited by 5SourceScholar
2024

Mobile Robot Oriented Large-Scale Indoor Dataset for Dynamic Scene Understanding

ICRA 2024poster

Most existing robotic datasets capture static scene data and thus are limited in evaluating robots’ dynamic performance. To address this, we present a mobile robot oriented large-scale indoor dataset, denoted as THUD (Tsinghua University Dynamic) robotic dataset, for training and evaluating their dy…

Cited by 9SourceScholar
2024

MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music

IJCAI 2024poster

The rapidly evolving multimodal Large Language Models (LLMs) urgently require new benchmarks to uniformly evaluate their performance on understanding and textually describing music. However, due to semantic gaps between Music Information Retrieval (MIR) algorithms and human understanding, discrepanc…

2024

OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

NeurIPS 2024poster

Current universal segmentation methods demonstrate strong capabilities in pixel-level image and video understanding. However, they lack reasoning abilities and cannot be controlled via text instructions. In contrast, large vision-language multimodal models exhibit powerful vision-based conversation…

Cited by 47SourcePDFScholar
2024

Stronger, Lighter, Better: Towards Life-Long Attribute Value Extraction for E-Commerce Products

ACL 2024findings

Attribute value extraction involves identifying the value spans of predetermined attributes in product texts. This area of research has traditionally operated under a closed-world assumption, focusing on products from a static set of categories and their associated attributes. However, products in e…

Cited by 0SourcePDFScholar
2024

Your Career Path Matters in Person-Job Fit

AAAI 2024technical

We are again confronted with one of the most vexing aspects of the advancement of technology: automation and AI technology cause the devaluation of human labor, resulting in unemployment. With this background, automatic person-job fit systems are promising solutions to promote the employment rate. T…

2024

i-Octree: A Fast, Lightweight, and Dynamic Octree for Proximity Search

ICRA 2024poster

Establishing the correspondences between newly acquired points and historically accumulated data (i.e., the map) through nearest neighbor search is crucial in numerous robotic applications. However, static tree data structures are inadequate to handle large and dynamically growing maps in real-time.…

Cited by 5SourcecodeScholar
2023

A Policy Optimization Method Towards Optimal-time Stability

CoRL 2023poster

In current model-free reinforcement learning (RL) algorithms, stability criteria based on sampling methods are commonly utilized to guide policy optimization. However, these criteria only guarantee the infinite-time convergence of the system's state to an equilibrium point, which leads to sub-optima…

Cited by 3SourceScholar
2023

Binary Image Fingerprint: Stable Structure Identifier for 3D LiDAR Place Recognition

RA-L 2023

Place recognition is considered as an effective strategy to reduce robot drift errors. In this work, a place recognition method that uses binary features to match loop closure frames is proposed for 3D LiDAR. The method extracts important information from structural feature matrix by image compressi

Cited by 5SourceScholar
2023

Boosting Fine-Grained Sketch-Based Image Retrieval with Self-Supervised Learning

ICASSP 2023accepted

Fine-grained sketch-based image retrieval (FG-SBIR) aims at aligning images and sketches at the instance level. It is a challenging task as there are significant differences between sketch and image. Existing methods usually produce less desired performance due to the lack of large-scale fine-graine…

Cited by 0SourceScholar
2023

Characteristics of Permanent Magnet Coupling Based Wireless Manipulation via Simulation

IROS 2023poster

Characteristics of wireless manipulation based on permanent magnet coupling, including anchoring distance, panning torque, and translational force, are assessed in this paper. The study focuses on a typical scenario where a slave robot embedded with a small permanent magnet can be remotely controlle…

Cited by 0SourceScholar
2023

CoF-CoT: Enhancing Large Language Models with Coarse-to-Fine Chain-of-Thought Prompting for Multi-domain NLU Tasks

EMNLP 2023short main

While Chain-of-Thought prompting is popular in reasoning tasks, its application to Large Language Models (LLMs) in Natural Language Understanding (NLU) is under-explored. Motivated by multi-step reasoning of LLMs, we propose Coarse-to-Fine Chain-of-Thought (CoF-CoT) approach that breaks down NLU tas…

Cited by 0SourcecodeScholar
2023

DVIS: Decoupled Video Instance Segmentation Framework

ICCV 2023poster

Video instance segmentation (VIS) is a critical task with diverse applications, including autonomous driving and video editing. Existing methods often underperform on complex and long videos in real world, primarily due to two factors. Firstly, offline methods are limited by the tightly-coupled mode…

Cited by 59PDFcodeScholar
2023

Design and Implementation of a Miniature Jellyfish-Inspired Robot

RA-L 2023

With the development of the global marine industry, the demand to investigate underwater environments using robots is increasing. Inspired by the unique structure and excellent swimming ability of <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Aurel

Cited by 16SourceScholar
2023

Efficient Exploration Using Extra Safety Budget in Constrained Policy Optimization

IROS 2023poster

Reinforcement learning (RL) has achieved promising results on most robotic control tasks. Safety of learning-based controllers is an essential notion of ensuring the effectiveness of the controllers. Current methods adopt whole consistency constraints during the training, thus resulting in inefficie…

Cited by 2SourceScholar
2023

Enhancing Cross-lingual Transfer via Phonemic Transcription Integration

ACL 2023findings

Previous cross-lingual transfer methods are restricted to orthographic representation learning via textual scripts. This limitation hampers cross-lingual transfer and is biased towards languages sharing similar well-known scripts. To alleviate the gap between languages from different writing scripts…

2023

Motion-Aware Video Paragraph Captioning via Exploring Object-Centered Internal Knowledge

ICASSP 2023accepted

Video paragraph captioning task aims at generating a fine-grained, coherent and relevant paragraph for a video. Different from the images where objects are static, the temporal states of objects are changing in videos. The dynamic information could be contributed to understanding the whole video con…

Cited by 0SourceScholar
2023

Sparse Frame Grouping Network with Action Centered for Untrimmed Video Paragraph Captioning

EMNLP 2023long findings

Generating paragraph captions for untrimmed videos without event annotations is challenging, especially when aiming to enhance precision and minimize repetition at the same time. To address this challenge, we propose a module called Sparse Frame Grouping (SFG). It dynamically groups event informatio…

Cited by 0SourceScholar
2023

Video Captioning via Relation-Aware Graph Learning

ICASSP 2023accepted

Recent neural models for video captioning usually employed an encoder-decoder framework. However, most approaches either neglected the spatial and temporal interactions between objects in a video or implicitly modelled the interactions, resulting in less desired performance. In this paper, we propos…

Cited by 0SourceScholar
2022

A Unified Positive-Unlabeled Learning Framework for Document-Level Relation Extraction with Different Levels of Labeling

EMNLP 2022main

Document-level relation extraction (RE) aims to identify relations between entities across multiple sentences. Most previous methods focused on document-level RE under full supervision. However, in real-world scenario, it is expensive and difficult to completely label all relations in a document bec…

2022

CREAM: Weakly Supervised Object Localization via Class RE-Activation Mapping

CVPR 2022poster

Weakly Supervised Object Localization (WSOL) aims to localize objects with image-level supervision. Existing works mainly rely on Class Activation Mapping (CAM) derived from a classification model. However, CAM-based methods usually focus on the most discriminative parts of an object (i.e., incomple…

Cited by 45PDFcodeScholar
2022

Collision-Free Trajectory Planning for a 6-DoF Free-Floating Space Robot via Hierarchical Decoupling Optimization

RA-L 2022

Collision-free trajectory planning is a critical technique for space robot mission. In this letter, we developed a model-free Hierarchical Decoupling Optimization (HDO) algorithm to realize 6D-pose multi-target trajectory planning for the free-floating space robot. In order to reduce the complexity

Cited by 31SourceScholar
2022

Design and Analysis of a Long-range Magnetic Actuated and Guided Endoscope for Uniport VATS

ICRA 2022poster

This paper presents a long-range magnetic actuated and guided endoscope for uniport video-assisted thoracic surgery (VATS). In VATS, the incision is quite narrow and part of the chest wall may be very thick. So, the magnetic endoscope system is required to produce sufficient attractive force at a co…

Cited by 5SourceScholar
2022

E2EC: An End-to-End Contour-Based Method for High-Quality High-Speed Instance Segmentation

CVPR 2022poster

Contour-based instance segmentation methods have developed rapidly recently but feature rough and handcrafted front-end contour initialization, which restricts the model performance, and an empirical and fixed backend predicted-label vertex pairing, which contributes to the learning difficulty. In t…

Cited by 113PDFcodeScholar
2022

Federated Learning Challenges and Opportunities: An Outlook

ICASSP 2022accepted

Federated learning (FL) has been developed as a promising framework to leverage the resources of edge devices, enhance customers’ privacy, comply with regulations, and reduce development costs. Although many methods and applications have been developed for FL, several critical challenges for practic…

Cited by 0SourceScholar
2022

Self-Aware Personalized Federated Learning

NeurIPS 2022accept

In the context of personalized federated learning (FL), the critical challenge is to balance local model improvement and global model tuning when the personal and global objectives may not be exactly aligned. Inspired by Bayesian hierarchical models, we develop a self-aware personalized FL method wh…

Cited by 28SourcePDFScholar
2022

Specialised Video Quality Model For Enhanced User Generated Content (UGC) With Special Effects

ICASSP 2022accepted

User Generated Content (UGC) refers to media generated by users for end-consumers that represent most of the media exchange on social media. UGC is subject to acquisition and transmission limitations that disable access to the pristine, i.e., perfect source content. Evaluating their quality, especia…

Cited by 0SourceScholar
2022

TCCNet: Temporally Consistent Context-Free Network for Semi-supervised Video Polyp Segmentation

IJCAI 2022poster

Automatic video polyp segmentation (VPS) is highly valued for the early diagnosis of colorectal cancer. However, existing methods are limited in three respects: 1) most of them work on static images, while ignoring the temporal information in consecutive video frames; 2) all of them are fully superv…

2022

Unified Speculation, Detection, and Verification Keyword Spotting

ICASSP 2022accepted

Accurate and timely recognition of the trigger keyword is vital for a good customer experience on smart devices. In the traditional keyword spotting task, there is typically a trade-off needed between accuracy and latency, where higher accuracy can be achieved by waiting for more context. In this pa…

Cited by 0SourceScholar
2021

A Multi-Target Trajectory Planning of a 6-DoF Free-Floating Space Robot via Reinforcement Learning

IROS 2021poster

Space robots have played an essential role in space junk removal. Compared with traditional model-based methods, model-free reinforcement learning methods are promising in tackling space capture missions, which is challenging due to the dynamic singular problem and measuring errors of dynamics param…

Cited by 27SourceScholar
2021

Generative Partial Visual-Tactile Fused Object Clustering

AAAI 2021technical

Visual-tactile fused sensing for object clustering has achieved significant progresses recently, since the involvement of tactile modality can effectively improve clustering performance. However, the missing data (i.e., partial data) issues always happen due to occlusion and noises during the data c…

Cited by 17SourcePDFScholar
2021

M3D-VTON: A Monocular-to-3D Virtual Try-On Network

ICCV 2021poster

Virtual 3D try-on can provide an intuitive and realistic view for online shopping and has a huge potential commercial value. However, existing 3D virtual try-on methods mainly rely on annotated 3D human shapes and garment templates, which hinders their applications in practical scenarios. 2D virtual…

Cited by 77PDFcodeScholar
2021

Multi-layer VI-GNSS Global Positioning Framework with Numerical Solution aided MAP Initialization

IROS 2021poster

Motivated by the goal of achieving long-term drift-free camera pose estimation in complex scenarios, we propose a global positioning framework fusing visual, inertial and Global Navigation Satellite System (GNSS) measurements in multiple layers. Different from previous loosely- and tightly-coupled m…

Cited by 5SourceScholar
2021

PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity Recognition

EMNLP 2021main

Cross-domain Named Entity Recognition (NER) transfers the NER knowledge from high-resource domains to the low-resource target domain. Due to limited labeled resources and domain shift, cross-domain NER is a challenging task. To address these challenges, we propose a progressive domain adaptation Kno…

Cited by 26SourcePDFScholar
2020

Evaluation of Joint Auditory Attention Decoding and Adaptive Binaural Beamforming Approach for Hearing Devices with Attention Switching

ICASSP 2020accepted

Beamforming is a common technique used to improve speech intelligibility and listening comfort of hearing aids users in a noisy environment. Traditional hearing aids beamforming algorithms require the a priori knowledge of the auditory of the listener, which may not be available in real applications…

Cited by 0SourceScholar
2020

MZET: Memory Augmented Zero-Shot Fine-grained Named Entity Typing

COLING 2020main

Named entity typing (NET) is a classification task of assigning an entity mention in the context with given semantic types. However, with the growing size and granularity of the entity types, few previous researches concern with newly emerged entity types. In this paper, we propose MZET, a novel mem…

Cited by 36SourcePDFScholar
2020

Regression Before Classification for Temporal Action Detection

ICASSP 2020accepted

Action classification combined with location regression is a widely-utilized mechanism in existing temporal action detection methods. However, there exists an inconsistency problem between locations and categories of action instances in this mechanism. More specifically, while the location of the pr…

Cited by 0SourceScholar
2019

A Joint Auditory Attention Decoding and Adaptive Binaural Beamforming Algorithm for Hearing Devices

ICASSP 2019accepted

Traditional adaptive binaural beamforming algorithms for hearing devices often assume that the target talker is known or can be derived from the listener's look direction. When this assumption is violated, the traditional beamforming algorithms often produce distorted target speech and less than opt…

Cited by 0SourceScholar
2019

Boundary Information Matters More: Accurate Temporal Action Detection with Temporal Boundary Network

ICASSP 2019accepted

Temporal action detection in untrimmed videos is an important yet challenging task. How to locate complex actions accurately is still an open question due to the ambiguous boundaries between action instances and the background. Recently a newly proposed work exploits Structured Segment Networks (SSN…

Cited by 0SourceScholar
2019

Hyperspectral Image Super-Resolution With Optimized RGB Guidance

CVPR 2019poster

To overcome the limitations of existing hyperspectral cameras on spatial/temporal resolution, fusing a low resolution hyperspectral image (HSI) with a high resolution RGB (or multispectral) image into a high resolution HSI has been prevalent. Previous methods for this fusion task usually emplo…

Cited by 108PDFcodeScholar
2019

Mechanical Framework Design with Experimental Verification of a Wearable Exoskeleton Chair

ICRA 2019poster

In this study, a human-chair model was developed as the basis for a wearable chair design. A prototype chair, HUST-EC, was fabricated and evaluated. Employing the optimization under an inner point penalty function, an optimized simulation of the operating mode with the lowest chair height was implem…

Cited by 15SourceScholar
2018

Evaluation of the Penalized Inequality Constrained Minimum Variance Beamformer for Hearing Aids

ICASSP 2018accepted

Beamforming is a common technique used to improve speech intelligibility and listening comfort of hearing aids users in a noisy environment. Traditional beamforming algorithms such as linearly constrained minimum variance (LCMV) beamformer cannot effectively suppress multiple interferences when the…

Cited by 0SourceScholar
2018

Improved Noise Characterization for Relative Impulse Response Estimation

ICASSP 2018accepted

Relative Impulse Responses (ReIRs) have several applications in speech enhancement, noise suppression and source localization for multi-channel speech processing in reverberant environments. Noise is usually assumed to be white Gaussian during the estimation of the ReIR between two microphones. We s…

Cited by 0SourceScholar
2018

Joint Camera Spectral Sensitivity Selection and Hyperspectral Image Recovery

ECCV 2018poster

Hyperspectral image (HSI) recovery from a single RGB image has attracted much attention, whose performance has recently been shown to be sensitive to the camera spectral sensitivity (CSS). In this paper, we present an efficient convolutional neural network (CNN) based method, which can jointly selec…

Cited by 70SourcePDFScholar
2018

Late Reverberation Suppression Using Recurrent Neural Networks with Long Short-Term Memory

ICASSP 2018accepted

Human speech is usually distorted by room reverberation. These corruptions degrade speech quality and intelligibility, especially under a long reverberation time, and they also pose a serious problem for many speech-related applications such as automatic speech recognition. In this paper, we propose…

Cited by 0SourceScholar
2017

Comparison of two binaural beamforming approaches for hearing aids

ICASSP 2017accepted

Beamforming algorithms in binaural hearing aids are crucial to improve speech understanding in background noise for hearing impaired persons. In this study, we compare and evaluate the performance of two recently proposed minimum variance (MV) beamforming approaches for binaural hearing aids. The bi…

Cited by 0SourceScholar
2016

Dynamic relative impulse response estimation using structured sparse Bayesian learning

ICASSP 2016accepted

In this paper we present a novel Hierarchical Bayesian approach to estimate Relative Impulse Response (ReIR) using short, noisy and reverberant microphone recordings. The information contained in ReIRs between two microphones is useful for a wide range of multichannel speech processing applications…

Cited by 10SourceScholar
2015

Incorporating spatial information in binaural beamforming for noise suppression in hearing aids

ICASSP 2015accepted

In this paper, we propose a beamforming algorithm for binaural hearing aids with enhanced noise suppression capability. The enhancement is based on incorporating a priori spatial information into the conventional multichannel Wiener filtering (MWF) approach for noise suppression. We develop a low co…

Cited by 11SourceScholar