← Search

Shuyi Ouyang

8 accepted papers

2026

Language-guided Frequency Modulation for Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual reasoning across diverse tasks. These tasks place different demands on visual representations: some prioritize high-level global context, while others emphasize fine-grained local details. However, most existing

Cited by 0SourceScholar
2026

ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps

CVPR 2026

Multimodal large language models (MLLMs) have demonstrated significant progress in semantic scene understanding and text-image alignment, with reasoning variants enhancing performance on more complex tasks involving mathematics and logic. However, their proficiency in tasks requiring both fine-grain

Cited by 0SourcecodeScholar
2026

Taming the Phantom: Token-Asymmetric Filtering for Hallucination Mitigation in Large Vision-Language Models

AAAI 2026technical

Hallucination in Large Vision-Language Models (LVLMs) remains a critical challenge, undermining their reliability in real-world applications. Existing studies have investigated the causes of hallucination at the modality level and proposed effective strategies. However, interaction patterns beyond

Cited by 0SourcePDFScholar
2025

M2OST: Many-to-one Regression for Predicting Spatial Transcriptomics from Digital Pathology Images

AAAI 2025technical

The advancement of Spatial Transcriptomics (ST) has facilitated the spatially-aware profiling of gene expressions based on histopathology images. Although ST data offers valuable insights into the micro-environment of tumors, its acquisition cost remains expensive. Therefore, directly predicting the…

2025

Region-aware Anchoring Mechanism for Efficient Referring Visual Grounding

ICCV 2025poster

Referring Visual Grounding (RVG) tasks revolve around utilizing vision-language interactions to incorporate object information from language expressions, thereby enabling targeted object detection or segmentation within images. Transformer-based methods have enabled effective interaction through att…

Cited by 0SourcePDFScholar
2024

IRLSG: Invariant Representation Learning for Single-Domain Generalization in Medical Image Segmentation

ICASSP 2024accepted

Single-domain generalization (SDG) can efficiently enhance model generalization while avoiding high annotation costs and privacy concerns. However, existing SDG methods are mainly based on data manipulation and meta-learning, which are not efficient enough due to the limited generalization performan…

Cited by 0SourceScholar
2023

MCKD: Mutually Collaborative Knowledge Distillation For Federated Domain Adaptation And Generalization

ICASSP 2023accepted

Conventional unsupervised domain adaptation (UDA) and domain generalization (DG) methods rely on the assumption that all source domains can be directly accessed and combined for model training. However, this centralized training strategy may violate privacy policies in many real-world applications.…

Cited by 0SourceScholar
2023

SLViT: Scale-Wise Language-Guided Vision Transformer for Referring Image Segmentation

IJCAI 2023poster

Referring image segmentation aims to segment an object out of an image via a specific language expression. The main concept is establishing global visual-linguistic relationships to locate the object and identify boundaries using details of the image. Recently, various Transformer-based techniques h…