← Search

Yoshiyuki Kobayashi

2 accepted papers

2026

Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs

ICML 2026poster

Recent Multimodal Large Language Models (MLLMs) have demonstrated significant progress in perceiving and reasoning over multimodal inquiries, ushering in a new research era for foundation models. However, vision-language misalignment in MLLMs has emerged as a critical challenge, where the textual re…

Cited by 0SourcecodeScholar
2019

Hierarchical Disentanglement of Discriminative Latent Features for Zero-Shot Learning

CVPR 2019poster

Most studies in zero-shot learning model the relationship, in the form of a classifier or mapping, between features from images of seen classes and their attributes. Therefore, the degree of a model's generalization ability for recognizing unseen images is highly constrained by that of image feature…

Cited by 72PDFScholar