← Search

Shan Yang

20 accepted papers

2026

Beyond Test-Time Training: Learning to Reason via Hardware-Efficient Optimal Control

ICML 2026poster

Associative memory has long underpinned the design of sequential models. Beyond recall, humans reason by *projecting future states and selecting goal-directed actions*, a capability that modern language models increasingly require but do not natively encode. While prior work uses reinforcement learn…

Cited by 0SourceScholar
2026

TMD-Bench: A Multi-Level Evaluation Paradigm for Music–Dance Co-Generation

ICML 2026poster

Unified audio--visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio--video synthesis to music–dance co-generation, the task becomes substantially harder: musical rhythm, phra…

Cited by 0SourceScholar
2026

UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation

AAAI 2026technical

Cued Speech (CS) enhances lipreading via hand coding, offering visual phonemic cues that support precise speech perception for the hearing-impaired. The task of CS Video-to-Speech generation (CSV2S) aims to convert CS videos into intelligible speech signals. Most existing research focuses on CS Reco

Cited by 0SourcePDFScholar
2025

DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions

ICASSP 2025accepted

Controlling text-to-speech (TTS) systems to synthesize speech with the prosodic characteristics expected by users has attracted much attention. To achieve controllability, current studies focus on two main directions: (1) using reference speech as prosody prompt to guide speech synthesis, and (2) us…

Cited by 0SourceScholar
2025

ThinkAnswer Loss: Balancing Semantic Similarity and Exact Matching for LLM Reasoning Enhancement

EMNLP 2025

Knowledge distillation for large language models often uses Chain-of-Thought (CoT) and answer pairs, but existing methods struggle with appropriate supervision signals. Uniform constraints (e.g., cross-entropy) on CoT can enforce literal, verbose reasoning and suppress expressive diversity, while so

Cited by 0SourcePDFScholar
2024

EffLoc: Lightweight Vision Transformer for Efficient 6-DOF Camera Relocalization

ICRA 2024poster

Camera relocalization is pivotal in computer vision, with applications in AR, drones, robotics, and autonomous driving. It estimates 3D camera position and orientation (6-DoF) from images. Unlike traditional methods like SLAM, recent strides use deep learning for direct end-to-end pose estimation. W…

Cited by 4SourceScholar
2024

Heuristic-Driven, Type-Specific Embedding in Parallel Spaces for Enhancing Knowledge Graph Reasoning

ICASSP 2024accepted

Knowledge Graph Reasoning aims to derive new insights from existing Knowledge Graphs (KGs) and address any missing or incomplete data. Existing models primarily rely on explicit information while neglecting the implicit constraints imposed by entity types on relations types. For example, when the en…

Cited by 0SourceScholar
2024

ICAR: Image-Based Complementary Auto Reasoning

AAAI 2024technical

Scene-aware Complementary Item Retrieval (CIR) is a challenging task which requires to generate a set of compatible items across domains. Due to the subjectivity, it is difficult to set up a rigorous standard for both data collection and learning objectives. To address this challenging task, we prop…

Cited by 1SourcePDFScholar
2024

Unleashing the Power of Large Language Models in Zero-shot Relation Extraction via Self-Prompting

EMNLP 2024finding

Recent research in zero-shot Relation Extraction (RE) has focused on using Large Language Models (LLMs) due to their impressive zero-shot capabilities. However, current methods often perform suboptimally, mainly due to a lack of detailed, context-specific prompts needed for understanding various sen…

Cited by 0SourcePDFScholar
2024

ViLA: Efficient Video-Language Alignment for Video Question Answering

ECCV 2024poster

"We propose an efficient Video-Language Alignment (ViLA) network. Our ViLA model addresses both efficient frame sampling and effective cross-modal alignment in a unified way. In our ViLA network, we design a new learnable text-guided Frame-Prompter together with a cross-modal distillation (QFormer-D…

2023

UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice Synthesis

AAAI 2023technical

Text-to-speech (TTS) and singing voice synthesis (SVS) aim at generating high-quality speaking and singing voice according to textual input and music scores, respectively. Unifying TTS and SVS into a single system is crucial to the applications requiring both of them. Existing methods usually suffer…

Cited by 10SourcePDFScholar
2022

Referee: Towards Reference-Free Cross-Speaker Style Transfer with Low-Quality Data for Expressive Speech Synthesis

ICASSP 2022accepted

Cross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker’s voice. Most previous CSST approaches rely on expensive high-quality data carrying desired speaking style during training and require a re…

Cited by 0SourceScholar
2022

VCVTS: Multi-Speaker Video-to-Speech Synthesis Via Cross-Modal Knowledge Transfer from Voice Conversion

ICASSP 2022accepted

Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speech, while allowing flexible control of speaker identity, all in a single system. This paper proposes a novel multi-speake…

Cited by 0SourceScholar
2021

AI Choreographer: Music Conditioned 3D Dance Generation With AIST++

ICCV 2021poster

We present AIST++, a new multi-modal dataset of 3D dance motion and music, along with FACT, a Full-Attention Cross-modal Transformer network for generating 3D dance motion conditioned on music. The proposed AIST++ dataset contains 1.1M frames of 3D dance motion in 1408 sequences, covering 10 dance g…

Cited by 576PDFcodeScholar
2021

Attention Bottlenecks for Multimodal Fusion

NeurIPS 2021poster

Humans perceive the world by concurrently processing and fusing high-dimensional inputs from multiple modalities such as vision and audio. Machine perception models, in stark contrast, are typically modality-specific and optimised for unimodal benchmarks. A common approach for building multimodal m…

2021

Entity Concept-enhanced Few-shot Relation Extraction

ACL 2021short

Few-shot relation extraction (FSRE) is of great importance in long-tail distribution problem, especially in special domain with low-resource data. Most existing FSRE algorithms fail to accurately classify the relations merely based on the information of the sentences together with the recognized ent…

2021

GLAVNet: Global-Local Audio-Visual Cues for Fine-Grained Material Recognition

CVPR 2021poster

In this paper, we aim to recognize materials with combined use of auditory and visual perception. To this end, we construct a new dataset named GLAudio that consists of both the geometry of the object being struck and the sound captured from either modal sound synthesis (for virtual objects) or real…

Cited by 8PDFScholar
2019

Enhancing Hybrid Self-attention Structure with Relative-position-aware Bias for Speech Synthesis

ICASSP 2019accepted

Compared with the conventional "front-end"-"back-end"- "vocoder" structure, based on the attention mechanism, end-to-end speech synthesis systems directly train and synthesize from text sequence to the acoustic feature sequence as a whole. Recently, a more calculation efficient end-to-end architectu…

Cited by 0SourceScholar