← Search

Wan-Cyuan Fan

10 accepted papers

2026

To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models

ICLR 2026poster

Large Vision Language Models (LVLMs) have recently emerged as powerful architectures capable of understanding and reasoning over both visual and textual information. These models typically rely on two key components: a Vision Transformer (ViT) and a Large Language Model (LLM). ViT encodes visual con…

Cited by 0SourceScholar
2025

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

EMNLP 2025

Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particularly when models attend to irrelevant regions of the chart. In this work, we prese

Cited by 0SourcePDFScholar
2025

Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities

ACL 2025long

Vision-language Models (VLMs) have emerged as general-purpose tools for addressing a variety of complex computer vision problems. Such models have been shown to be highly capable, but, at the same time, lacking some basic visual understanding skills. In this paper, we set out to understand the limit…

Cited by 0SourcePDFScholar
2023

Frido: Feature Pyramid Diffusion for Complex Scene Image Synthesis

AAAI 2023technical

Diffusion models (DMs) have shown great potential for high-quality image synthesis. However, when it comes to producing images with complex scenes, how to properly describe both image global structures and object details remains a challenging task. In this paper, we present Frido, a Feature Pyramid…

2023

IoU-Aware Multi-Expert Cascade Network Via Dynamic Ensemble for Long-Tailed Object Detection

ICASSP 2023accepted

Object detection over a long-tailed large-scale dataset is practical, challenging, and comprehensively under-explored. Recently proposed methods mainly focus on eliminating the imbalanced classification problem. However, only a few attempts have been made to consider the quality of the predicted bou…

Cited by 0SourceScholar
2022

Cross-Modal Mutual Learning for Audio-Visual Speech Recognition and Manipulation

AAAI 2022technical

As a key characteristic in audio-visual speech recognition (AVSR), relating linguistic information observed across visual and audio data has been a challenge, benefiting not only audio/visual speech recognition (ASR/VSR) but also for manipulating data within/across modalities. In this paper, we pres…

Cited by 16SourcePDFScholar
2022

Paraphrasing Is All You Need for Novel Object Captioning

NeurIPS 2022accept

Novel object captioning (NOC) aims to describe images containing objects without observing their ground truth captions during training. Due to the absence of caption annotation, captioning models cannot be directly optimized via sequence-to-sequence training or CIDEr optimization. As a result, we pr…

Cited by 5SourcePDFScholar
2022

Scene Graph Expansion for Semantics-Guided Image Outpainting

CVPR 2022poster

In this paper, we address the task of semantics-guided image outpainting, which is to complete an image by generating semantically practical content. Different from most existing image outpainting works, we approach the above task by understanding and completing image semantics at the scene graph le…

Cited by 19PDFScholar
2021

LayoutTransformer: Scene Layout Generation With Conceptual and Spatial Diversity

CVPR 2021poster

When translating text inputs into layouts or images, existing works typically require explicit descriptions of each object in a scene, including their spatial information or the associated relationships. To better exploit the text input, so that implicit objects or relationships can be properly infe…

Cited by 42PDFcodeScholar