← Search

Qiangchang Wang

7 accepted papers

2026

DVLA-RL: Dual-Level Vision–Language Alignment with Reinforcement Learning Gating for Few-Shot Learning

ICLR 2026poster

Few-shot learning (FSL) aims to generalize to novel categories with only a few samples. Recent approaches incorporate large language models (LLMs) to enrich visual representations with semantic embeddings derived from class names. However, they overlook progressive and adaptive alignment between vis…

Cited by 0SourceScholar
2025

Disparity-Guided Cross-View Transformer For Stereo Image Super-Resolution

ICASSP 2025accepted

Although transformer-based methods excel in stereo image super-resolution, the full potential of the distinctive, complementary information inherent in stereo images has not been fully utilized. We propose a Disparity-Guided Cross-View Transformer (DCT) to extract features across dimensions and view…

Cited by 0SourceScholar
2025

LOFI: Harnessing Attention Dynamics for Facial Expression Recognition with Noisy Labels

ICASSP 2025accepted

Facial expression recognition (FER) faces unique challenges from expression ambiguity and noisy labels, degrading performance in real-world applications. While leveraging attention, existing methods frequently neglect attention dynamic mechanism of dispersion followed by focus and the spatially stru…

Cited by 0SourceScholar
2025

VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning

NeurIPS 2025poster

Few-shot learning (FSL) aims to recognize novel concepts from only a few labeled support samples. Recent studies enhance support features by incorporating additional semantic information (e.g., class descriptions) or designing complex semantic fusion modules. However, these methods still suffer fro…

Cited by 0SourcecodeScholar
2021

TransFER: Learning Relation-Aware Facial Expression Representations With Transformers

ICCV 2021poster

Facial expression recognition (FER) has received increasing interest in computer vision. We propose the TransFER model which can learn rich relation-aware local representations. It mainly consists of three components: Multi-Attention Dropping (MAD), ViT-FER, and Multi-head Self-Attention Dropping (M…

Cited by 278PDFcodeScholar