← Search

Weidong Chen

43 accepted papers

2026

CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design

ICLR 2026poster

Graphic design plays a vital role in visual communication across advertising, marketing, and multimedia entertainment. Prior work has explored automated graphic design generation using diffusion models, aiming to streamline creative workflows and democratize design capabilities. However, complex gra…

Cited by 0SourcecodeScholar
2026

Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation

ICLR 2026poster

Current speech language models require their core component, the speech codec, to discretize continuous speech signals into tokens that not only capture high-level cues for autoregressive modeling but also preserve sufficient acoustic details for perceptual quality. To address this need, we propose…

Cited by 0SourceScholar
2026

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

ICASSP 2026poster

Mainstream Automatic Speech Recognition (ASR) systems excel at transcribing lexical content, but largely fail to recognize nonverbal vocalizations (NVs) embedded in speech, such as sighs, laughs, and coughs. This capability is important for a comprehensive understanding of human communication, as NV…

Cited by 0SourcePDFScholar
2025

A 4D Radar Camera Extrinsic Calibration Tool Based on 3D Uncertainty Perspective N Points

IROS 2025

4D imaging radar is a type of low-cost millimeter-wave radar(costing merely 10-20% of lidar systems) capable of providing range, azimuth, elevation, and Doppler velocity information. Accurate extrinsic calibration between millimeter-wave radar and camera systems is critical for robust multimodal per

Cited by 2SourceScholar
2025

CGS-SLAM: Compact 3D Gaussian Splatting for Dense Visual SLAM

IROS 2025

Recent work has shown that 3D Gaussian-based SLAM enables high-quality reconstruction, accurate pose estimation, and real-time rendering of scenes. However, these approaches are built on a tremendous number of redundant 3D Gaussian ellipsoids, leading to high memory and storage costs and slow traini

Cited by 61SourceScholar
2025

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

ICASSP 2025accepted

Controlling text-to-speech (TTS) systems to synthesize speech with the prosodic characteristics expected by users has attracted much attention. To achieve controllability, current studies focus on two main directions: (1) using reference speech as prosody prompt to guide speech synthesis, and (2) us…

Cited by 0SourceScholar
2025

Graph Mixture of Experts and Memory-augmented Routers for Multivariate Time Series Anomaly Detection

AAAI 2025technical

Multivariate time series (MTS) anomaly detection is a critical task that involves identifying abnormal patterns or events in data that consist of multiple interrelated time series. In order to better model the complex interdependence between entities and the various inherent characteristics of each…

2025

MNE-SLAM: Multi-Agent Neural SLAM for Mobile Robots

CVPR 2025poster

Neural implicit scene representations have recently shown promising results in dense visual SLAM. However, existing implicit SLAM algorithms are constrained to single-agent scenarios, and fall difficulty in large indoor scenes and long sequences. Existing multi-agent SLAM frameworks cannot meet the…

2025

SN-LiDAR: Semantic Neural Fields for Novel Space-time View LiDAR Synthesis

IROS 2025

Recent research has begun exploring novel view synthesis (NVS) for LiDAR point clouds, aiming to generate realistic LiDAR scans from unseen viewpoints. However, most existing approaches do not reconstruct semantic labels, which are crucial for many downstream applications such as autonomous driving

Cited by 1SourcecodeScholar
2024

An Interpretable and Generalizable Speech Detector Based on a CNN-LSTM Framework

ICASSP 2024accepted

Speech brain-computer interface (speech BCI) aims to reconstruct speech from recorded brain signals. Real-time speech BCI relies on speech detection, which is greatly impacted by the selection of speech-related neural frequency features. However, most studies did not investigate this aspect when des…

Cited by 0SourceScholar
2024

Bootstrapping Large Language Models for Radiology Report Generation

AAAI 2024technical

Radiology report generation (RRG) aims to automatically generate a free-text description from a specific clinical radiograph, e.g., chest X-Ray images. Existing approaches tend to perform RRG with specific models trained on the public yet limited data from scratch, where they often lead to inferior…

2024

Improving Radiology Report Generation with D2-Net: When Diffusion Meets Discriminator

ICASSP 2024accepted

Radiology report generation (RRG) aims to automatically provide observations and insight into a patient’s condition based on radiology images, which is able to greatly reduce the workload of physicians on the premise of ensuring the quality of medical treatment. Existing works leverage the Transform…

Cited by 0SourceScholar
2024

PLGSLAM: Progressive Neural Scene Represenation with Local to Global Bundle Adjustment

CVPR 2024poster

Neural implicit scene representations have recently shown encouraging results in dense visual SLAM. However existing methods produce low-quality scene reconstruction and low-accuracy localization performance when scaling up to large indoor scenes and long sequences. These limitations are mainly due…

Cited by 67SourcePDFScholar
2024

Prompting Few-shot Multi-hop Question Generation via Comprehending Type-aware Semantics

NAACL 2024findings

Given several documents, multi-hop question generation (MQG) is a task aims to generate complicated questions that require reasoning over multiple pieces of these documents to find the answer. To perform this task, existing studies focus on designing advanced architectures to locate essential keywor…

2023

DST: Deformable Speech Transformer for Emotion Recognition

ICASSP 2023accepted

Enabled by multi-head self-attention, Transformer has exhibited remarkable results in speech emotion recognition (SER). Compared to the original full attention mechanism, window-based attention is more effective in learning fine-grained features while greatly reducing model redundancy. However, emot…

Cited by 0SourceScholar
2023

DWFormer: Dynamic Window Transformer for Speech Emotion Recognition

ICASSP 2023accepted

Speech emotion recognition is crucial to human-computer interaction. The temporal regions that represent different emotions scatter in different parts of the speech locally. Moreover, the temporal scales of important information may vary over a large range within and across speech segments. Although…

Cited by 0SourceScholar
2023

End-to-end Aspect-based Sentiment Analysis with Combinatory Categorial Grammar

ACL 2023findings

End-to-end Aspect-based Sentiment Analysis (EASA) is a natural language processing (NLP) task that involves extracting aspect terms and identifying the sentiments for them, which provides a fine-grained level of text analysis and thus requires a deep understanding of the running text. Many previous…

2023

Improving Image Captioning via Predicting Structured Concepts

EMNLP 2023long main

Having the difficulty of solving the semantic gap between images and texts for the image captioning task, conventional studies in this area paid some attention to treating semantic concepts as a bridge between the two modalities and improved captioning performance accordingly. Although promising res…

Cited by 0SourceScholar
2023

Text Style Transfer with Contrastive Transfer Pattern Mining

ACL 2023long

Text style transfer (TST) is an important task in natural language generation, which aims to alter the stylistic attributes (e.g., sentiment) of a sentence and keep its semantic meaning unchanged. Most existing studies mainly focus on the transformation between styles, yet ignore that this transform…

2022

Fixed and Sliding FBG Sensors-Based Triaxial Tip Force Sensing for Cable-Driven Continuum Robots

ICRA 2022poster

Tip force sensing for cable-driven continuum robots are vital to provide the force information for safe and reliable human-robot interaction. However, traditional triaxial force sensors usually have a complicated structure occupying its inner lumen, without enough space for additional instrumental t…

Cited by 6SourceScholar
2022

Key-Sparse Transformer for Multimodal Speech Emotion Recognition

ICASSP 2022accepted

Speech emotion recognition is a challenging research topic that plays a critical role in human-computer interaction. Multimodal inputs further improve the performance as more emotional information is used. However, existing studies learn all the information in the sample while only a small portion o…

Cited by 0SourceScholar
2022

ROLL: Long-Term Robust LiDAR-based Localization With Temporary Mapping in Changing Environments

IROS 2022poster

Long-term scene changes pose challenges to localization systems using a pre-built map. This paper presents a LiDAR-based system that provides robust localization against those challenges. Our method starts with activation of a mapping process temporarily when global matching towards the pre-built ma…

Cited by 18SourcecodeScholar
2021

Hybrid Vision/Force Control for Interaction with the Bottle-like Object

ICRA 2021poster

This study proposes a hybrid vision/force control scheme for interaction with the inner surface of the bottle-like object. Based on the geometry of the object, a new generalized constraint called the bottleneck (BN) constraint is proposed, which ensures the tool passes through a fixed 3-D region and…

Cited by 1SourceScholar
2021

LSSED: A Large-Scale Dataset and Benchmark for Speech Emotion Recognition

ICASSP 2021accepted

Speech emotion recognition is a vital contributor to the next generation of human-computer interaction (HCI). However, current existing small-scale databases have limited the development of related research. In this paper, we present LSSED, a challenging large-scale english speech emotion dataset, w…

Cited by 0SourceScholar
2021

Soft Manipulator Fault Detection and Identification Using ANC-based LSTM

IROS 2021poster

Timely fault detection and identification (FDI) of soft manipulators are critical in the design of surgical systems to improve reliability. However, due to the intrinsic compliance of soft manipulators, their end effectors vibrate during the dynamic control process, which introduces noise into the m…

Cited by 7SourcecodeScholar
2021

Toward State-Unsaturation Guaranteed Fault Detection Method in Visual Servoing of Soft Robot Manipulators

IROS 2021poster

This paper puts forward a novel sensor-less fault detection method with only task errors feedback and applies it to visual servoing tasks of soft robot manipulators. The method is developed by introducing a suitably designed endogenous accessory signal (EAS). On the one hand, EAS transforms the chan…

Cited by 4SourceScholar
2021

Towards Collision Detection, Localization and Force Estimation for a Soft Cable-driven Robot Manipulator

ICRA 2021poster

Soft robots have been applied widely to various constrained scenarios due to the advantages over traditional rigid manipulators such as softness, deformability and adaptability to constrained surroundings. To make full use of this merit, this paper proposes a method that integrates collision detecti…

Cited by 1SourceScholar
2020

Hierarchical Quadtree Feature Optical Flow Tracking Based Sparse Pose-Graph Visual-Inertial SLAM

ICRA 2020poster

Accurate, robust and real-time localization under constrained-resources is a critical problem to be solved. In this paper, we present a new sparse pose-graph visual-inertial SLAM (SPVIS). Unlike the existing methods that are costly to deal with a large number of redundant features and 3D map points,…

Cited by 9SourceScholar
2020

Long-Term Localization With Time Series Map Prediction for Mobile Robots in Dynamic Environments

IROS 2020poster

In many applications of mobile robot, the environment is constantly changing. How to use historical information to analysis environmental changes and generate a map corresponding with current environment is important to achieve high-precision localization. Inspired by predictive mechanism of brain,…

Cited by 13SourceScholar
2019

Local Pose optimization with an Attention-based Neural Network

IROS 2019poster

In this paper, we propose a novel pose optimizer which can be inserted into either supervised or unsupervised end-to-end visual odometry for the purpose of local pose optimization. The pose optimizer is an analogue of the pose graph optimization used in traditional VSLAM algorithms. Local pose optim…

Cited by 2SourceScholar
2019

Long-Term Visual Inertial SLAM based on Time Series Map Prediction

IROS 2019poster

With the advance in the field of mobile robots, autonomous robots are required for long-term deployment in dynamic and complex environments. However, the performance of Visual Inertial SLAM systems in long-term operation is not satisfactory, and most long-term SLAM systems assumes periodic changes i…

Cited by 16SourceScholar
2019

Retrieval-based Localization Based on Domain-invariant Feature Learning under Changing Environments

IROS 2019poster

Visual localization is a crucial problem in mobile robotics and autonomous driving. One solution is to retrieve images with known pose from a database for the localization of query images. However, in environments with drastically varying conditions (e.g. illumination changes, seasons, occlusion, dy…

Cited by 30SourcecodeScholar
2019

Unsupervised Learning of Monocular Depth and Ego-Motion Using Multiple Masks

ICRA 2019poster

A new unsupervised learning method of depth and ego-motion using multiple masks from monocular video is proposed in this paper. The depth estimation network and the ego-motion estimation network are trained according to the constraints of depth and ego-motion without truth values. The main contribut…

Cited by 38SourcecodeScholar
2018

A Failure-Tolerant Approach to Synchronous Formation Control of Mobile Robots Under Communication Delays

ICRA 2018poster

Robot malfunction is inevitable in practical applications of the robot formation control due to uncontrolled crashing, system malfunction or communication loss. In this paper, we study the synchronous formation control problem in the presence of robot malfunctions. Our main idea is to improve the ne…

Cited by 7SourceScholar
2018

Tagging Like Humans: Diverse and Distinct Image Annotation

CVPR 2018poster

In this work we propose a new automatic image annotation model, dubbed diverse and distinct image annotation (D2IA). The generative model D2IA is inspired by the ensemble of human annotations, which create semantically relevant, yet distinct and diverse tags. In D2IA, we generate a relevant and dist…

2017

A unified leader-follower scheme for mobile robots with uncalibrated on-board camera

ICRA 2017poster

This paper studies the problem of image-based leader-follower formation control for mobile robots, where the controller is designed independently of the leader's motion. An adaptive control scheme, which is suitable for both omnidirectional and perspective cameras, is proposed. The proposed approach…

Cited by 14SourceScholar
2015

A gradient-based self-healing algorithm for mobile robot formation

IROS 2015poster

In this paper, we investigate the self-healing problem of mobile robot formation after some robots have been damaged, and present a gradient-based algorithm which enables mobile robots to restore the topology of the formation through local interactions among neighboring robots. Firstly, in order to…

Cited by 10SourceScholar