← Search

Sicheng Yang

23 accepted papers

2026

CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation

CVPR 2026

Co-speech gesture generation has significantly advanced human-computer interaction, yet speaker movements remain constrained due to the omission of text-driven non-spontaneous gestures (e.g., bowing while talking). Existing methods face two key challenges: 1) the semantic prior gap due to the lack o

Cited by 0SourceScholar
2026

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

AAAI 2026technical

The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs)

Cited by 0SourcePDFScholar
2026

VAEVQ: Enhancing Discrete Visual Tokenization Through Variational Modeling

AAAI 2026technical

Vector quantization (VQ) transforms continuous image features into discrete representations, providing compressed, tokenized inputs for generative models. However, VQ-based frameworks suffer from several issues, such as non-smooth latent spaces, weak alignment between representations before and aft

Cited by 0SourcePDFScholar
2025

Robotic Hand Tool Use with Contact-Based Demonstration: The Case of Cucumber Peeling

IROS 2025

Robotic hand tool use has garnered significant attention from robotics researchers, because it enhances dexterity beyond the limitations imposed by manipulators with fixed tool configurations and human-involved manual tool changes. Despite extensive research, current methodologies predominantly focu

Cited by 0SourceScholar
2025

VQ-Seg: Vector-Quantized Token Perturbation for Semi-Supervised Medical Image Segmentation

NeurIPS 2025poster

Consistency learning with feature perturbation is a widely used strategy in semi-supervised medical image segmentation. However, many existing perturbation methods rely on dropout, and thus require a careful manual tuning of the dropout rate, which is a sensitive hyperparameter and often difficult t…

Cited by 0SourcecodeScholar
2024

A High-Performance Anthropomorphic Robotic Arm for Household Applications

IROS 2024poster

Anthropomorphic robotic arms, mimicking the structure and function of human arms, show great potential for helping people in various tedious and repetitive household tasks. However, such arms mostly consist of multiple serial links controlled independently by actuators at joints with high reduction…

Cited by 0SourceScholar
2024

Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control

AAAI 2024technical

This study aims to improve the generation of 3D gestures by utilizing multimodal information from human speech. Previous studies have focused on incorporating additional modalities to enhance the quality of generated gestures. However, these methods perform poorly when certain modalities are missing…

Cited by 14SourcePDFScholar
2024

Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model

CVPR 2024poster

Co-speech gestures if presented in the lively form of videos can achieve superior visual effects in human-machine interaction. While previous works mostly generate structural human skeletons resulting in the omission of appearance information we focus on the direct generation of audio-driven co-spee…

2024

Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion Models

ICASSP 2024accepted

Audio-driven co-speech human gesture generation has made remarkable advancements recently. However, most previous works only focus on single person audio-driven gesture generation. We aim at solving the problem of conversational co-speech gesture generation that considers multiple participants in a…

Cited by 0SourceScholar
2024

FreeTalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness

ICASSP 2024accepted

Current talking avatars mostly generate co-speech gestures based on audio and text of the utterance, without considering the non-speaking motion of the speaker. Furthermore, previous works on co-speech gesture generation have designed network structures based on individual gesture datasets, which re…

Cited by 0SourceScholar
2024

MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models

NeurIPS 2024poster

Gesture synthesis is a vital realm of human-computer interaction, with wide-ranging applications across various fields like film, robotics, and virtual reality. Recent advancements have utilized the diffusion model to improve gesture synthesis. However, the high computational complexity of these t…

2024

TRX-Hand5: An Anthropomorphic Hand with Integrated Tactile Feedback for Grasping and Manipulation in Human Environments

IROS 2024poster

Objects of daily life are designed to suit the human hand. Without major modifications to these objects and our environments, robots will need end-effectors with human hand-like configuration and dexterity to efficiently operate on them. Tight integration of tactile and proprioceptive sensors are al…

Cited by 0SourceScholar
2024

Thermoformed electronic skins for conformal tactile sensor arrays

ICRA 2024poster

Robots and prostheses are increasingly designed with curvilinear surfaces for functional, aesthetic, aerodynamic, and safety reasons. Electronic skins (e-skins) capable of sensing contact location and pressure across complex, non-developable surfaces are essential for empowering next-generation robo…

Cited by 2SourceScholar
2024

VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception

ICML 2024poster

This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the "Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision,…

2023

A Unified Trajectory Generation Algorithm for Dynamic Dexterous Manipulation

IROS 2023poster

This paper proposes a novel efficient multi-phase trajectory generation algorithm for dynamic dexterous manipulation tasks, such as throwing, catching, dynamic regrasping, and dynamic handover, which can be decomposed into multiple manipulation primitives, including sticking, rolling, approaching, s…

Cited by 1SourceScholar
2023

DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models

IJCAI 2023poster

The art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the rhythm and semantics of the gesture to the corresponding spee…

2023

GTN-Bailando: Genre Consistent long-Term 3D Dance Generation Based on Pre-Trained Genre Token Network

ICASSP 2023accepted

Music-driven 3D dance generation has become an intensive research topic in recent years with great potential for real-world applications. Most existing methods lack the consideration of genre, which results in genre inconsistency in the generated dance movements. In addition, the correlation between…

Cited by 0SourceScholar
2023

QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation

CVPR 2023highlight

Speech-driven gesture generation is highly challenging due to the random jitters of human motion. In addition, there is an inherent asynchronous relationship between human speech and gestures. To tackle these challenges, we introduce a novel quantization-based and phase-guided motion matching framew…

2023

Wavsyncswap: End-To-End Portrait-Customized Audio-Driven Talking Face Generation

ICASSP 2023accepted

Audio-driven talking face with portrait customization enhances the flexibility of avatar applications for different scenarios, such as on-line meetings, mixed reality, and data generation. Among the existing methods, audio-driven talking face and face swapping are typically viewed as separate tasks…

Cited by 0SourceScholar
2022

Real-time Inertial Parameter Identification of Floating-Base Robots Through Iterative Primitive Shape Division

ICRA 2022poster

Dynamic models play a key role in robot motion generation and control and the identification of inertial parameters is a critical component for obtaining an accurate dynamic model of a robot. This paper presents a novel iterative primitive shape division method for the inertia parameter identificati…

Cited by 1SourceScholar
2020

Gain Scheduled Controller Design for Balancing an Autonomous Bicycle

IROS 2020poster

In this paper, the gain scheduling technique is applied to design a balance controller for an autonomous bicycle with an inertia wheel. Previously, two different balance controllers are needed depending on whether the bicycle is stationary or dynamic. The switch between the two different controllers…

Cited by 16SourceScholar
2020

Nonlinear Balance Control of an Unmanned Bicycle: Design and Experiments

IROS 2020poster

In this paper, nonlinear control techniques are exploited to balance an unmanned bicycle with enlarged stability domain. We consider two cases. For the first case when the autonomous bicycle is balanced by the flywheel, the steering angle is set to zero, and the torque of the flywheel is used as the…

Cited by 26SourceScholar
2019

Development of a Continuous Vertical-pulling Automatic Doffing Robot for the Ring Spinning

IROS 2019poster

Doffing robot is an important part of the spinning process in the textile production. This paper analyzes the doffing process of spinning machines and points out the requirements of the structure and functions of the doffer. The locking two-finger gripper, the three-dimensional circulating operation…

Cited by 1SourceScholar