← Search

Yuguang Yang

12 accepted papers

2026

AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentation

ICLR 2026poster

Referring Image Segmentation (RIS) aims to segment the object in an image uniquely referred to by a natural language expression. However, RIS training often contains hard-to-align and instance-specific visual signals; optimizing on such pixels injects misleading gradients and drives the model in the…

Cited by 0SourcecodeScholar
2026

SURGE: Surrogate Gradient Adaptation in Binary Neural Networks

ICML 2026poster

The training of Binary Neural Networks (BNNs) is fundamentally based on gradient approximation for non-differentiable binarization operations (e.g., sign function). However, prevailing methods including the Straight-Through Estimator (STE) and its improved variants, rely on hand-crafted designs that…

Cited by 0SourceScholar
2025

DFM: Differentiable Feature Matching for Anomaly Detection

CVPR 2025poster

Feature matching methods for unsupervised anomaly detection have demonstrated impressive performance. Existing methods primarily rely on self-supervised training and handcrafted matching schemes for task adaptation. However, they can only achieve an inferior feature representation for anomaly detect…

Cited by 0SourcePDFScholar
2025

Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection

ICLR 2025poster

Zero-shot medical detection can further improve detection performance without relying on annotated medical images even upon the fine-tuned model, showing great clinical value. Recent studies leverage grounded vision-language models (GLIP) to achieve this by using detailed disease descriptions as pro…

Cited by 0SourcePDFScholar
2025

WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection

ICCV 2025poster

Leveraging the complementary characteristics of visible (RGB) and infrared (IR) imagery offers significant potential for improving object detection. In this paper, we propose WaveMamba, a cross-modality fusion method that efficiently integrates the unique and complementary frequency features of RGB…

Cited by 0SourcePDFScholar
2024

CLIP in Mirror: Disentangling text from visual images through reflection

NeurIPS 2024poster

The CLIP network excels in various tasks, but struggles with text-visual images i.e., images that contain both text and visual objects; it risks confusing textual and visual representations. To address this issue, we propose MirrorCLIP, a zero-shot framework, which disentangles the image features of…

2024

GEmo-CLAP: Gender-Attribute-Enhanced Contrastive Language-Audio Pretraining for Accurate Speech Emotion Recognition

ICASSP 2024accepted

Contrastive cross-modality pretraining has recently exhibited impressive success in diverse fields, whereas there is limited research on their merits in speech emotion recognition (SER). In this paper, we propose GEmo-CLAP, a kind of gender-attribute-enhanced contrastive language-audio pretraining (…

Cited by 0SourceScholar
2024

Promptvc: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts

ICASSP 2024accepted

Stylistic voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or reference speech to control the conversion process, which leads to limitations…

Cited by 0SourceScholar
2023

Hybridformer: Improving Squeezeformer with Hybrid Attention and NSR Mechanism

ICASSP 2023accepted

SqueezeFormer has recently shown impressive performance in automatic speech recognition (ASR). However, its inference speed suffers from the quadratic complexity of softmax-attention (SA). In addition, limited by the large convolution kernel size, the local modeling ability of SqueezeFormer is insuf…

Cited by 0SourceScholar
2022

Improving Fairness in Speaker Verification via Group-Adapted Fusion Network

ICASSP 2022accepted

Modern speaker verification models use deep neural networks to encode utterance audio into discriminative embedding vectors. During the training process, these networks are typically optimized to differentiate arbitrary speakers. This learning process biases the learning of fine voice characteristic…

Cited by 0SourceScholar
2022

Self-Supervised Speaker Recognition Training using Human-Machine Dialogues

ICASSP 2022accepted

Speaker recognition, recognizing speaker identities based on voice alone, enables important downstream applications, such as personalization and authentication. Learning speaker representations, in the context of supervised learning, heavily depends on both clean and sufficient labeled data, which i…

Cited by 0SourceScholar