← Search

Pinyuan Feng

3 accepted papers

2026

Towards Interpretable Visual Decoding with Attention to Brain Representations

ICLR 2026poster

Recent work has demonstrated that complex visual stimuli can be decoded from human brain activity using deep generative models, helping brain science researchers interpret how the brain represents real-world scenes. However, most current approaches leverage mapping brain signals into intermediate im…

Cited by 0SourceScholar
2026

Vision Language Models Cannot Reason About Physical Transformation

ICML 2026poster

Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they genuinely understand physical transformations remains unclear. We introduce ***ConservationBench*** evaluating ***conse…

Cited by 0SourceScholar
2025

TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration

EMNLP 2025

Multimodal in-context learning (ICL) has emerged as a key mechanism for harnessing the capabilities of large vision–language models (LVLMs). However, its effectiveness remains highly sensitive to the quality of input ICL sequences, particularly for tasks involving complex reasoning or open-ended gen

Cited by 0SourcePDFScholar