← Search

Yi Yuan

23 accepted papers

2026

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMs

AAAI 2026technical

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of e

Cited by 0SourcePDFScholar
2025

FlowSep: Language-Queried Sound Separation with Rectified Flow Matching

ICASSP 2025accepted

Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds and minimize interference from other sources. However, these…

Cited by 0SourceScholar
2025

General Incomplete Time Series Analysis via Patch Dropping Without Imputation

IJCAI 2025

Missing values in multivariate time series data present significant challenges to effective analysis. Existing methods for multivariate time series analysis either ignore missing data, sacrificing performance, or follow the impute-then-analyze paradigm, which suffers from redundant training and erro

2025

MMNet: Missing-Aware and Memory-Enhanced Network for Multivariate Time Series Imputation

IJCAI 2025

Multivariate time series (MTS) data in real-world scenarios are often incomplete, which hinders effective data analysis. Therefore, MTS imputation has been widely studied to facilitate various MTS tasks. Existing imputation methods primarily initialize missing values with zeros in order to perform e

2025

Sound-VECaps: Improving Audio Generation with Visually Enhanced Captions

ICASSP 2025accepted

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work…

Cited by 0SourceScholar
2025

Uncertainty-Aware Iterative Preference Optimization for Enhanced LLM Reasoning

ACL 2025long

Direct Preference Optimization (DPO) has recently emerged as an efficient and effective method for aligning large language models with human preferences. However, constructing high-quality preference datasets remains challenging, often necessitating expensive manual or powerful LM annotations. Addit…

2024

Retrieval-Augmented Text-to-Audio Generation

ICASSP 2024accepted

Despite recent progress in text-to-audio (TTA) generation, we show that the state-of-the-art models, such as AudioLDM, trained on datasets with an imbalanced class distribution, such as AudioCaps, are biased in their generation performance. Specifically, they excel in generating common audio classes…

Cited by 0SourceScholar
2023

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

ICML 2023poster

Text-to-audio (TTA) systems have recently gained attention for their ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study, we propose AudioLDM, a TTA system that is built on a lat…

2023

Enhancing Robustness and Imperceptibility of Blind Watermarking with Improved Message Processor

ICASSP 2023accepted

The current state-of-the-art(SOTA) blind watermark embedding method MBRS based on deep learning is less robust to Crop, and additional diffusion layers need to be added for optimization. However, the diffusion layer will make the model less robust to noise other than Crop. Therefore, MBRS which need…

Cited by 0SourceScholar
2023

SwiftAvatar: Efficient Auto-Creation of Parameterized Stylized Character on Arbitrary Avatar Engines

AAAI 2023technical

The creation of a parameterized stylized character involves careful selection of numerous parameters, also known as the "avatar vectors" that can be interpreted by the avatar engine. Existing unsupervised avatar vector estimation methods that auto-create avatars for users, however, often fail to wor…

2022

A Unified Framework for Real Time Motion Completion

AAAI 2022technical

Motion completion, as a challenging and fundamental problem, is of great significance in film and game applications. For different motion completion application scenarios (in-betweening, in-filling, and blending), most previous methods deal with the completion problems with case-by-case methodology…

Cited by 23SourcePDFScholar
2022

Learning Implicit Body Representations from Double Diffusion Based Neural Radiance Fields

IJCAI 2022poster

In this paper, we present a novel double diffusion based neural radiance field, dubbed DD-NeRF, to reconstruct human body geometry and render the human body appearance in novel views from a sparse set of images. We first propose a double diffusion mechanism to achieve expressive representations of i…

Cited by 10SourcePDFScholar
2021

Automatic Translation of Music-to-Dance for In-Game Characters

IJCAI 2021poster

Music-to-dance translation is an emerging and powerful feature in recent role-playing games. Previous works of this topic consider music-to-dance as a supervised motion generation problem based on time-series data. However, these methods require a large amount of training data pairs and may suffer f…

2021

HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation

AAAI 2021technical

Self-supervised learning shows great potential in monocular depth estimation, using image sequences as the only source of supervision. Although people try to use the high-resolution image for depth estimation, the accuracy of prediction has not been significantly improved. In this work…

2021

In-game Residential Home Planning via Visual Context-aware Global Relation Learning

AAAI 2021technical

In this paper, we propose an effective global relation learning algorithm to recommend an appropriate location of a building unit for in-game customization of residential home complex. Given a construction layout, we propose a visual context-aware graph generation network that learns the implicit gl…

Cited by 5SourcePDFScholar
2021

One-shot Face Reenactment Using Appearance Adaptive Normalization

AAAI 2021technical

The paper proposes a novel generative adversarial network for one-shot face reenactment, which can animate a single face image to a different pose-and-expression (provided by a driving image) while keeping its original appearance. The core of our network is a novel mechanism called appearance adapti…

Cited by 31SourcePDFScholar
2021

RFNet: Recurrent Forward Network for Dense Point Cloud Completion

ICCV 2021poster

Point cloud completion is an interesting and challenging task in 3D vision, aiming to recover complete shapes from sparse and incomplete point clouds. Existing learning-based methods often require vast computation cost to achieve excellent performance, which limits their practical applications. In t…

Cited by 48PDFScholar
2021

Structure-aware Person Image Generation with Pose Decomposition and Semantic Correlation

AAAI 2021technical

In this paper we tackle the problem of pose guided person image generation, which aims to transfer a person image from the source pose to a novel target pose while maintaining the source appearance. Given the inefficiency of standard CNNs in handling large spatial transformation, we propose a struct…

Cited by 23SourcePDFScholar
2020

Ladybird: Quasi-Monte Carlo Sampling for Deep Implicit Field Based 3D Reconstruction with Symmetry

ECCV 2020poster

Deep implicit field regression methods are effective for 3D reconstruction from single-view images. However, the impact of different sampling patterns on the reconstruction quality is not well-understood. In this work, we first study the effect of point set discrepancy on the network training. Based…

Cited by 0SourcePDFScholar
2020

Towards High-Fidelity 3D Face Reconstruction From In-the-Wild Images Using Graph Convolutional Networks

CVPR 2020poster

3D Morphable Model (3DMM) based methods have achieved great success in recovering 3D face shapes from single-view images. However, the facial textures recovered by such methods lack the fidelity as exhibited in the input images. Recent works demonstrate high-quality facial texture recovering with ge…

Cited by 152PDFcodeScholar
2019

Face-to-Parameter Translation for Game Character Auto-Creation

ICCV 2019poster

Character customization system is an important component in Role-Playing Games (RPGs), where players are allowed to edit the facial appearance of their in-game characters with their own preferences rather than using default templates. This paper proposes a method for automatically creating in-game c…

Cited by 75PDFScholar