← Search

Yishu Liu

11 accepted papers

2026

Coverage ≠ Exposure: Auditable Control of Same-Support Tail Failures under Multimodal Missingness

ICML 2026poster

Real-world multimodal systems inevitably face partial observability due to sensor dropout and degradation. Standard robustness methods can improve average performance, but they often remain unreliable in rare, adverse long-tail conditions. Under a locked same-support contract, we uncover a same-supp…

Cited by 0SourceScholar
2026

Cross Modal Fine-grained Alignment via Granularity-aware and Region-uncertain Modeling

AAAI 2026technical

Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained alignment requires precise correspondence between localized visual

Cited by 0SourcePDFScholar
2026

ReMoE: Region-Mixture Experts for Adversarially-Robust Vision Transformers

CVPR 2026

Vision Transformers (ViTs) achieve state-of-the-art performance on a wide range of vision tasks, yet they remain highly vulnerable to adversarial perturbations due to the lack of explicit region-level semantic modeling. Adversarial perturbations are typically local and spatially structured, whereas

Cited by 0SourcecodeScholar
2025

Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language Biases

IJCAI 2025

Existing Medical Visual Question Answering (Med-VQA) models often suffer from language biases, where spurious correlations between question types and answer categories are inadvertently established. To address these issues, we propose a novel Cause-Effect Driven Optimization framework called CEDO, t

2025

Language‑Bias‑Resilient Visual Question Answering via Adaptive Multi‑Margin Collaborative Debiasing

NeurIPS 2025poster

Language bias in Visual Question Answering (VQA) arises when models exploit spurious statistical correlations between question templates and answers, particularly in out-of-distribution scenarios, thereby neglecting essential visual cues and compromising genuine multimodal reasoning. Despite numerou…

Cited by 0SourceScholar
2025

OralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware Mining

CVPR 2025poster

In clinical practice, panoramic dental radiography is a widely employed imaging technique that can provide a detailed and comprehensive view of dental structures and surrounding tissues for identifying various oral anomalies. However, due to the complexity of oral anomalies and the scarcity of avail…

2025

Towards Robust Visual Question Answering via Prompt-Driven Geometric Harmonization

AAAI 2025technical

Visual Question Answering (VQA) has garnered significant attention as a crucial link between vision and language, aimed at generating accurate responses to visual queries. However, current VQA models still struggle with the challenges of minority class collapse and spurious semantic correlations pos…

Cited by 0SourcePDFScholar
2024

CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive Learning

AAAI 2024technical

Dental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intrica…

2024

Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image Classification

ICASSP 2024accepted

The feature dispersion, arising from the inherent constraints of data scarcity, has emerged as a prominent challenge in the domain of few-shot learning. In this paper, we propose a novel Self-adaptive Distribution Regularization (SADR) approach, which can adaptively bridge the semantic gaps across d…

Cited by 0SourceScholar
2024

Enhancing Cross-Modal Retrieval via Visual-Textual Prompt Hashing

IJCAI 2024poster

Cross-modal hashing has garnered considerable research interest due to its rapid retrieval and low storage costs. However, the majority of existing methods suffer from the limitations of context loss and information redundancy, particularly in simulated textual environments enriched with manually an…

Cited by 3SourcePDFScholar
2024

Medical Vision-Language Representation Learning with Cross-Modal Multi-Teacher Contrastive Distillation

ICASSP 2024accepted

Medical vision-language representation learning has garnered considerable attention owing to its applicability to extracting generic representations from the image and text modality. However, it still remains challenging to acquire a more comprehensive understanding of intra- and inter-modal semanti…

Cited by 0SourceScholar