← Search

Bingzhi Chen

20 accepted papers

2026

BayesVQA: Energy-Guided Bayesian Debiasing for Language-Bias-Robust Visual Question Answering

AAAI 2026technical

Numerous studies have demonstrated that Visual Question Answering (VQA) models are vulnerable to language priors and dataset biases, often leading to spurious correlations between questions and answers. As a result, these models excessively rely on linguistic cues, neglecting essential visual inform

Cited by 1SourcePDFScholar
2026

Coverage ≠ Exposure: Auditable Control of Same-Support Tail Failures under Multimodal Missingness

ICML 2026poster

Real-world multimodal systems inevitably face partial observability due to sensor dropout and degradation. Standard robustness methods can improve average performance, but they often remain unreliable in rare, adverse long-tail conditions. Under a locked same-support contract, we uncover a same-supp…

Cited by 0SourceScholar
2026

Cross Modal Fine-grained Alignment via Granularity-aware and Region-uncertain Modeling

AAAI 2026technical

Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained alignment requires precise correspondence between localized visual

Cited by 0SourcePDFScholar
2026

ReMoE: Region-Mixture Experts for Adversarially-Robust Vision Transformers

CVPR 2026

Vision Transformers (ViTs) achieve state-of-the-art performance on a wide range of vision tasks, yet they remain highly vulnerable to adversarial perturbations due to the lack of explicit region-level semantic modeling. Adversarial perturbations are typically local and spatially structured, whereas

Cited by 0SourcecodeScholar
2026

Selection-as-Nonlinearity: Bridging Attention and Activation via a Joint Game-Decision Lens for Interpretable, Discriminative Visual Representations

CVPR 2026

Self-attention with separate pre- and post-projections can be a universal approximator (on compact domains) under mild conditions. Yet we observe a striking gap: an attention-only Transformer (w/o FFN layers) exhibits a marked accuracy drop relative to its standard interleaved attention--FFN baselin

Cited by 0SourcecodeScholar
2026

Toward Principled Flexible Scaling for Self-Gated Neural Activation

ICLR 2026poster

Neural networks necessitate nonlinearities to achieve universal approximability. Traditional activation functions introduce nonlinearities through rigid feature rectifications. Recent self-gated variants improve traditional methods in fitting flexibility by incorporating learnable content-aware fact…

Cited by 0SourceScholar
2025

A2GP-SF: Enhancing Few-shot Class Incremental Learning via Attribute Generative Prompting and Adaptive Sharpness Flattening

ICASSP 2025accepted

Few-shot Class Incremental Learning (FSCIL) aims to incrementally learn new classes with limited examples while retaining knowledge of previously learned classes. Recent advancements in prompt tuning for large pre-trained models have shown promise in FSCIL. However, current FSCIL methods still suffe…

Cited by 0SourceScholar
2025

Advancing Few-Shot Class-Incremental Learning with Virtual Prototype Guidance Prompting

ICASSP 2025accepted

Few-Shot Class-Incremental Learning (FSCIL) aims to incrementally learn new class knowledge from limited samples while preserving previously knowledge from encountered classes. However, existing FSCIL methods encounter two primary challenges: (1) inadequate adaptation, where overfitting to new class…

Cited by 0SourceScholar
2025

Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language Biases

IJCAI 2025

Existing Medical Visual Question Answering (Med-VQA) models often suffer from language biases, where spurious correlations between question types and answer categories are inadvertently established. To address these issues, we propose a novel Cause-Effect Driven Optimization framework called CEDO, t

2025

Enhancing Incomplete Multimodal Learning via Modal Complementary Recovering

ICASSP 2025accepted

Multimodal learning presents significant challenges arising from the unpredictable absence of modalities during both training and testing phases. Existing recovery methods struggle to leverage the available data, which can introduce additional noise during the recovery process and degrade performanc…

Cited by 0SourceScholar
2025

Language‑Bias‑Resilient Visual Question Answering via Adaptive Multi‑Margin Collaborative Debiasing

NeurIPS 2025poster

Language bias in Visual Question Answering (VQA) arises when models exploit spurious statistical correlations between question templates and answers, particularly in out-of-distribution scenarios, thereby neglecting essential visual cues and compromising genuine multimodal reasoning. Despite numerou…

Cited by 0SourceScholar
2025

Learning Hierarchical Attribute Prompt for Vision-Language Models

ICASSP 2025accepted

Prompt learning is a common strategy for adapting Visual Language Models (VLMs) to downstream tasks by fine-tuning prompts for task-specific performance. However, existing methods face two key challenges: overfitting to base classes, which limits generalization to novel classes, and the dependence o…

Cited by 0SourceScholar
2025

OralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware Mining

CVPR 2025poster

In clinical practice, panoramic dental radiography is a widely employed imaging technique that can provide a detailed and comprehensive view of dental structures and surrounding tissues for identifying various oral anomalies. However, due to the complexity of oral anomalies and the scarcity of avail…

2025

Towards Differential Optimization: Rehearsal-Free Class-Incremental Learning with Slow Learners and Fast Adapters

ICASSP 2025accepted

Class-incremental learning (CIL) enables models to learn new tasks without forgetting previously acquired knowledge. However, existing CIL approaches often struggle with inadequate adaptation to task-specific feature spaces and catastrophic forgetting of previously-acquired knowledge, compromising t…

Cited by 0SourceScholar
2025

Towards Robust Visual Question Answering via Prompt-Driven Geometric Harmonization

AAAI 2025technical

Visual Question Answering (VQA) has garnered significant attention as a crucial link between vision and language, aimed at generating accurate responses to visual queries. However, current VQA models still struggle with the challenges of minority class collapse and spurious semantic correlations pos…

Cited by 0SourcePDFScholar
2024

CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive Learning

AAAI 2024technical

Dental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intrica…

2024

Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image Classification

ICASSP 2024accepted

The feature dispersion, arising from the inherent constraints of data scarcity, has emerged as a prominent challenge in the domain of few-shot learning. In this paper, we propose a novel Self-adaptive Distribution Regularization (SADR) approach, which can adaptively bridge the semantic gaps across d…

Cited by 0SourceScholar
2024

Enhancing Cross-Modal Retrieval via Visual-Textual Prompt Hashing

IJCAI 2024poster

Cross-modal hashing has garnered considerable research interest due to its rapid retrieval and low storage costs. However, the majority of existing methods suffer from the limitations of context loss and information redundancy, particularly in simulated textual environments enriched with manually an…

Cited by 3SourcePDFScholar
2024

Medical Vision-Language Representation Learning with Cross-Modal Multi-Teacher Contrastive Distillation

ICASSP 2024accepted

Medical vision-language representation learning has garnered considerable attention owing to its applicability to extracting generic representations from the image and text modality. However, it still remains challenging to acquire a more comprehensive understanding of intra- and inter-modal semanti…

Cited by 0SourceScholar
2021

Hierarchical Network Based on the Fusion of Static and Dynamic Features for Speech Emotion Recognition

ICASSP 2021accepted

Many studies on automatic speech emotion recognition (SER) have been devoted to extracting meaningful emotional features for generating emotion-relevant representations. However, they generally ignore the complementary learning of static and dynamic features, leading to limited performances. In this…

Cited by 0SourceScholar