← Search

Xiang Yin

35 accepted papers

2026

Bridging Perception and Planning: Towards End-To-End Planning for Signal Temporal Logic Tasks

ICRA 2026poster

We investigate the task and motion planning problem for Signal Temporal Logic (STL) specifications in robotics. Existing STL methods rely on pre-defined maps or mobility representations, which are ineffective in unstruc- tured real-world environments. We propose the Structured- MoE STL Planner (S-MS…

2026

InfinityHuman: Towards Long-Term Audio-Driven Human Animation

CVPR 2026

Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration videos with consistent appearance and natural hand motions. Existing methods extend videos using overlapping motion frames

Cited by 0SourcecodeScholar
2025

Adaptive Visual Servoing Control Barrier Function of Robotic Manipulators with Uncalibrated Camera

IROS 2025

This paper investigates the problem of safe visual servoing control of manipulators using an uncalibrated eye-in-hand camera based on control barrier functions (CBFs). Traditional CBFs are defined in the workspace, corresponding to the global coordinates of the base frame. However, when the camera’s

Cited by 0SourceScholar
2025

Argumentative Large Language Models for Explainable and Contestable Claim Verification

AAAI 2025technical

The profusion of knowledge encoded in large language models (LLMs) and their ability to apply this knowledge zero-shot in a range of settings makes them promising candidates for use in decision-making. However, they are currently limited by their inability to provide outputs which can be faithfully…

2025

Automated Manipulation of Magnetic Microswarms for Temporal Logic Cargo Delivery Tasks in Complex Environments

IROS 2025

Micromanipulation using magnetic microswarms has garnered significant attention in recent years due to their potential in microscale cargo delivery tasks. While existing studies have demonstrated the capabilities of microswarms in basic manipulation tasks, they often lack the autonomy required to ha

Cited by 0SourceScholar
2025

Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis

ACL 2025finding

Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to insufficient emotional perception and redundant discrete speech coding…

2025

MaxAuc: A Max-Plus-Based Auction Approach for Multi-Robot Allocations for Time-Ordered Temporal Logic Tasks

IROS 2025

In this paper, we investigate a multi-robot task allocation problem where a team of heterogeneous robots operates in a discrete workspace to achieve a set of tasks expressed by linear temporal logic formulas. In contrast to existing works, we further consider inter-task-time-order constraints, which

Cited by 0SourceScholar
2025

Online Synthesis of Control Barrier Functions with Local Occupancy Grid Maps for Safe Navigation in Unknown Environments

IROS 2025

Control Barrier Functions (CBFs) have emerged as an effective and non-invasive safety filter for ensuring the safety of autonomous systems in dynamic environments with formal guarantees. However, most existing works on CBF synthesis focus on fully known settings. Synthesizing CBFs online based on pe

Cited by 0SourceScholar
2024

Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context Modeling

AAAI 2024technical

Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not thoroughly investigated the emotional expressiveness problem…

2024

Explaining Arguments’ Strength: Unveiling the Role of Attacks and Supports

IJCAI 2024poster

Quantitatively explaining the strength of arguments under gradual semantics has recently received increasing attention. Specifically, several works in the literature provide quantitative explanations by computing the attribution scores of arguments. These works disregard the importance of attacks an…

Cited by 8SourcePDFScholar
2024

FedST: Federated Style Transfer Learning for Non-IID Image Segmentation

AAAI 2024technical

Federated learning collaboratively trains machine learning models among different clients while keeping data privacy and has become the mainstream for breaking data silos. However, the non-independently and identically distribution (i.e., Non-IID) characteristic of different image domains among diff…

2024

Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

ICLR 2024poster

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspe…

2024

MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes

NeurIPS 2024poster

Talking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically…

2024

NNgTL: Neural Network Guided Optimal Temporal Logic Task Planning for Mobile Robots

ICRA 2024poster

In this work, we investigate task planning for mobile robots under linear temporal logic (LTL) specifications. This problem is particularly challenging when robots navigate in continuous workspaces due to the high computational complexity involved. Sampling-based methods have emerged as a promising…

Cited by 6SourcecodeScholar
2024

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

ICLR 2024spotlight

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talkin…

2024

Sleep When Everything Looks Fine: Self-Triggered Monitoring for Signal Temporal Logic Tasks

RA-L 2024

Online monitoring is a widely used technique in assessing if the performance of the system satisfies some desired requirements during run-time operation. Existing works on online monitoring usually assume that the monitor can acquire system information periodically at each time instant, which may be

Cited by 8SourcecodeScholar
2024

Synthesis of Temporally-Robust Policies for Signal Temporal Logic Tasks using Reinforcement Learning

ICRA 2024poster

This paper investigates the problem of designing control policies that satisfy high-level specifications described by signal temporal logic (STL) in unknown, stochastic environments. While many existing works concentrate on optimizing the spatial robustness of a system, our work takes a step further…

Cited by 5SourcecodeScholar
2023

AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation

ACL 2023long

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from distinct degradation in noisy environments and fail to translate visual speech (i.e.,…

2023

CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-Training

ACL 2023long

Improving text representation has attracted much attention to achieve expressive text-to-speech (TTS). However, existing works only implicitly learn the prosody with masked token reconstruction tasks, which leads to low training efficiency and difficulty in prosody modeling. We propose CLAPSpeech, a…

2023

Explaining Random Forests Using Bipolar Argumentation and Markov Networks

AAAI 2023technical

Random forests are decision tree ensembles that can be used to solve a variety of machine learning problems. However, as the number of trees and their individual size can be large, their decision making process is often incomprehensible. We show that their decision process can be naturally represen…

2023

LiteG2P: A Fast, Light and High Accuracy Model for Grapheme-to-Phoneme Conversion

ICASSP 2023accepted

As a key component of automated speech recognition (ASR) and the front-end in text-to-speech (TTS), grapheme-to-phoneme (G2P) plays the role of converting letters to their corresponding pronunciations. Existing methods are either slow or poor in performance, and are limited in application scenarios,…

Cited by 0SourceScholar
2023

Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

ICML 2023poster

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio pairs, and the complexity of modeling long continuous audio…

2023

Security-Aware Reinforcement Learning under Linear Temporal Logic Specifications

ICRA 2023poster

In this paper, we investigate the problem of reinforcement learning under linear temporal logic (LTL) specifications for Markov decision processes (MDPs) with security constraints. We consider an outside passive intruder (observer) that can observe the external output behavior of the system through…

Cited by 4SourceScholar
2023

UniLG: A Unified Structure-aware Framework for Lyrics Generation

ACL 2023long

As a special task of natural language generation, conditional lyrics generation needs to consider the structure of generated lyrics and the relationship between lyrics and music. Due to various forms of conditions, a lyrics generation system is expected to generate lyrics conditioned on different si…

2023

Unsupervised Video Domain Adaptation for Action Recognition: A Disentanglement Perspective

NeurIPS 2023poster

Unsupervised video domain adaptation is a practical yet challenging task. In this work, for the first time, we tackle it from a disentanglement view. Our key idea is to handle the spatial and temporal domain divergence separately through disentanglement. Specifically, we consider the generation of c…

2023

Virtual Try-On with Pose-Garment Keypoints Guided Inpainting

ICCV 2023poster

Virtual try-on is an important technology supporting online apparel shopping, which provides consumers with a virtual experience to fit garments without physically wearing them. Recently, the image-based virtual try-on has received growing research attention. However, the synthetic results of existi…

Cited by 32PDFcodeScholar
2022

Towards Using Clothes Style Transfer for Scenario-Aware Person Video Generation

ICASSP 2022accepted

Clothes style transfer for person video generation is a challenging task, due to drastic variations of intra-person appearance and video scenarios. To tackle this problem, most recent AdaIN-based architectures are proposed to extract clothes and scenario features for generation. However, these appro…

Cited by 0SourceScholar
2021

A Chapter-Wise Understanding System for Text-To-Speech in Chinese Novels

ICASSP 2021accepted

In TTS-based audiobook production, multi-role dubbing and emotional expressions can significantly improve the naturalness of audiobooks. However, it requires manual annotation of original novels with explicit speaker and emotion tags in sentence level, which is extremely time-consuming and costly. I…

Cited by 0SourceScholar
2021

PPG-Based Singing Voice Conversion with Adversarial Representation Learning

ICASSP 2021accepted

Singing voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion works, we propose a novel model to steadily convert songs while keeping their naturalness and intonation. We build an end-to…

Cited by 0SourceScholar
2020

A Hybrid Text Normalization System Using Multi-Head Self-Attention For Mandarin

ICASSP 2020accepted

In this paper, we propose a hybrid text normalization system using multi-head self-attention. The system combines the advantages of a rule-based model and a neural model for text preprocessing tasks. Previous studies in Mandarin text normalization usually use a set of hand-written rules, which are h…

Cited by 0SourceScholar
2020

A Unified Sequence-to-Sequence Front-End Model for Mandarin Text-to-Speech Synthesis

ICASSP 2020accepted

In Mandarin text-to-speech (TTS) system, the front-end text processing module significantly influences the intelligibility and naturalness of synthesized speech. Building a typical pipeline-based front-end which consists of multiple individual components requires extensive efforts. In this paper, we…

Cited by 0SourceScholar
2016

Modeling spectral envelopes using deep conditional restricted Boltzmann machines for statistical parametric speech synthesis

ICASSP 2016accepted

This paper proposes a spectral modeling method using a deep conditional restricted Boltzmann machine (DCRBM) for statistical parametric speech synthesis. In this method, a DCRBM, which combines a deep neural network (DNN) with a conditional restricted Boltzmann machine (CRBM), is utilized to describ…

Cited by 0SourceScholar