← Search

Naoto Usuyama

6 accepted papers

2026

Learning Sparse Visual Representations via Spatial-Semantic Factorization

ICML 2026poster

Self-supervised learning (SSL) faces a fundamental conflict between semantic understanding and image reconstruction. High-level semantic SSL (e.g., DINO) relies on global tokens that are forced to be location-invariant for augmentation alignment, a process that inherently discards the spatial coordi…

Cited by 0SourceScholar
2025

Boltzmann Attention Sampling for Image Analysis with Small Objects

CVPR 2025poster

Detecting and segmenting small objects, such as lung nodules and tumor lesions, remains a critical challenge in image analysis. These objects often occupy less than 0.1% of an image, making traditional transformer architectures inefficient and prone to performance degradation due to redundant attent…

Cited by 0SourcePDFScholar
2023

Exploring the Boundaries of GPT-4 in Radiology

EMNLP 2023long main

The recent success of general-domain large language models (LLMs) has significantly changed the natural language processing paradigm towards a unified foundation model across domains and applications. In this paper, we focus on assessing the performance of GPT-4, the most capable LLM so far, on the…

Cited by 0SourceScholar
2023

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

NeurIPS 2023spotlight

Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversational AI has seen rapid progress by leveraging billions of image-text pairs from the public web, but such general-domain vi…

Cited by 828SourcePDFScholar
2022

Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing

ECCV 2022poster

"Multi-modal data abounds in biomedicine, such as radiology images and reports. Interpreting this data at scale is essential for improving clinical care and accelerating clinical research. Biomedical text with its complex semantics poses additional challenges in vision-language modelling compared to…

2021

Modular Self-Supervision for Document-Level Relation Extraction

EMNLP 2021main

Extracting relations across large text spans has been relatively underexplored in NLP, but it is particularly important for high-value domains such as biomedicine, where obtaining high recall of the latest findings is crucial for practical applications. Compared to conventional information extractio…

Cited by 10SourcePDFScholar