← Search

Yuanyuan Wang

23 accepted papers

2026

DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models

AAAI 2026technical

Extending pre-trained text Large Language Models (LLMs)’s speech understanding or generation abilities by introducing various effective speech tokens has attracted great attention in the speech research community. However, building a unified speech understanding and generation model still faces the

Cited by 0SourcePDFScholar
2026

Geometry-Aware Stereo Matching via Monocular Disparity Distribution Prior and Gradient Enhancement

AAAI 2026technical

Stereo matching recovers 3D scene information based on the correlation between corresponding pixels. Despite impressive progress, existing methods lack sufficient correlation priors in ill-posed regions such as occlusions, detailed and reflective regions. In this paper, we propose Geometry Aware Ste

Cited by 0SourcePDFScholar
2026

Physics from Video: Identifiability of Time-Invariant Second-Order ODEs under Minimal Trajectory Conditions

ICML 2026poster

Bridging the gap between visual realism and physical understanding is a core challenge for video-based world models. We study the structural identifiability of continuous-time physical laws from raw pixels, focusing on whether an encoder-only pipeline can uniquely recover the parameters of second-or…

Cited by 0SourceScholar
2025

ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling

ICML 2025poster

Recent advancements in audio language models have underscored the pivotal role of audio tokenization, which converts audio signals into discrete tokens, thereby facilitating the application of language model architectures to the audio domain. In this study, we introduce ALMTokenizer, a novel low-bit…

Cited by 0SourcePDFScholar
2025

AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions

ICASSP 2025accepted

Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by incorporating additional frame-level conditions or control net…

Cited by 0SourceScholar
2025

CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCR

CVPR 2025poster

Scene Text Retrieval (STR) seeks to identify all images containing a given query string. Existing methods typically rely on an explicit Optical Character Recognition (OCR) process of text spotting or localization, which is susceptible to complex pipelines and accumulated errors. To settle this, we r…

Cited by 0SourcePDFScholar
2025

HarmonySeg: Tubular Structure Segmentation with Deep-Shallow Feature Fusion and Growth-Suppression Balanced Loss

ICCV 2025poster

Accurate segmentation of tubular structures in medical images, such as vessels and airway trees, is crucial for computer-aided diagnosis, radiotherapy, and surgical planning. However, significant challenges exist in algorithm design when faced with diverse sizes, complex topologies, and (often) inco…

Cited by 0SourcePDFScholar
2025

Is Your Autonomous Vehicle Safe? Understanding the Threat of Electromagnetic Signal Injection Attacks on Traffic Scene Perception

AAAI 2025technical

Autonomous vehicles rely on camera-based perception systems to comprehend their driving environment and make crucial decisions, thereby ensuring vehicles to steer safely. However, a significant threat known as Electromagnetic Signal Injection Attacks (ESIA) can distort the images captured by these c…

Cited by 0SourcePDFScholar
2025

Multi-Label Few-Shot Image Classification via Pairwise Feature Augmentation and Flexible Prompt Learning

AAAI 2025technical

Multi-label few-shot image classification is a crucial and challenging task due to limited annotated data and elusive category specificity. However, research on this topic is still in the rudimentary stage and few methods are available. Existing methods either leverage data augmentation to alleviate…

Cited by 0SourcePDFScholar
2025

SEP-MLDC: A Simple and Effective Paradigm for Multi-Label Document Classification

NAACL 2025findings

Multi-label document classification (MLDC) aims to allocate more than one label to each document and attracts increasing attention in many practical applications. However, previous studies have failed to pay sufficient attention to the lack of semantic information on labels and the long-tail problem…

Cited by 0SourcePDFScholar
2024

Audio Prompt Tuning for Universal Sound Separation

ICASSP 2024accepted

Universal sound separation (USS) is a task to separate arbitrary sounds from an audio mixture. Existing USS systems are capable of separating arbitrary sources, given a few examples of the target sources as queries. However, separating arbitrary sounds with a single system is challenging, and the ro…

Cited by 0SourceScholar
2024

Bandwidth-Efficient Inference for Nerual Image Compression

ICASSP 2024accepted

With neural networks growing deeper and feature maps growing larger, limited communication bandwidth with external memory (or DRAM) and power constraints become a bottle-neck in implementing network inference on mobile and edge devices. In this paper, we propose an end-to-end differentiable bandwidt…

Cited by 0SourceScholar
2024

Consistent and Relevant: Rethink the Query Embedding in General Sound Separation

ICASSP 2024accepted

The query-based audio separation usually employs specific queries to extract target sources from a mixture of audio signals. Currently, most query-based separation models need additional networks to obtain query embedding. In this way, separation model is optimized to be adapted to the distribution…

Cited by 0SourceScholar
2024

Idempotence and Perceptual Image Compression

ICLR 2024spotlight

Idempotence is the stability of image codec to re-compression. At the first glance, it is unrelated to perceptual image compression. However, we find that theoretically: 1) Conditional generative model-based perceptual codec satisfies idempotence; 2) Unconditional generative model with idempotence c…

2024

Identifiability Analysis of Linear ODE Systems with Hidden Confounders

NeurIPS 2024poster

The identifiability analysis of linear Ordinary Differential Equation (ODE) systems is a necessary prerequisite for making reliable causal inferences about these systems. While identifiability has been well studied in scenarios where the system is fully observable, the conditions for identifiability…

Cited by 0SourcePDFScholar
2024

UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task Learner

NeurIPS 2024poster

Large Language models (LLMs) have demonstrated supreme capabilities in textual understanding and generation, but cannot be directly applied to cross-modal tasks without fine-tuning. This paper proposes a cross-modal in-context learning approach, empowering the frozen LLMs to achieve multiple audio t…

2023

Bit Allocation using Optimization

ICML 2023poster

In this paper, we consider the problem of bit allocation in Neural Video Compression (NVC). First, we reveal a fundamental relationship between bit allocation in NVC and Semi-Amortized Variational Inference (SAVI). Specifically, we show that SAVI with GoP (Group-of-Picture)-level likelihood is equiv…

2023

DASA: Difficulty-Aware Semantic Augmentation for Speaker Verification

ICASSP 2023accepted

Data augmentation is vital to the generalization ability and robustness of deep neural networks (DNNs) models. Existing augmentation methods for speaker verification manipulate the raw signal, which are time-consuming and the augmented samples lack diversity. In this paper, we present a novel diffic…

Cited by 0SourceScholar
2023

Generator Identification for Linear SDEs with Additive and Multiplicative Noise

NeurIPS 2023poster

In this paper, we present conditions for identifying the generator of a linear stochastic differential equation (SDE) from the distribution of its solution process with a given fixed initial state. These identifiability conditions are crucial in causal inference using linear SDEs as they enable the…

Cited by 5SourcePDFScholar
2022

Optimization of Spring Constant of a Pneumatic Artificial Muscle-Spring Driven Antagonistic Structure

RA-L 2022

Pneumatic artificial muscles (PAMs) have been widely applied to robotic systems, especially assistive devices, which could benefit from PAMs’ intrinsic viscoelasticity. However, nonlinearity and hysteresis make it challenging to achieve high-accuracy control. Using pre-tensioned springs to replace o

Cited by 7SourceScholar
2022

Practical Learned Lossless JPEG Recompression With Multi-Level Cross-Channel Entropy Model in the DCT Domain

CVPR 2022poster

JPEG is a popular image compression method widely used by individuals, data center, cloud storage and network filesystems. However, most recent progress on image compression mainly focuses on uncompressed images while ignoring trillions of already-existing JPEG images. To compress these JPEG images…

Cited by 7PDFScholar
2021

Designing Soft Pneumatic Actuators for Thumb Movements

RA-L 2021

Studies have developed various types of soft robotic gloves for hand rehabilitation in recent years. Most soft actuators achieved a sufficient thumb flexion assist while lacking opposition support, which requires the coordination of thumb flexion and abduction-adduction. The difficulties for thumb s

Cited by 26SourceScholar