← Search

Kazuya Tateishi

2 accepted papers

2026

VIRTUE: Visual-Interactive Text-Image Universal Embedder

ICLR 2026poster

Multimodal representation learning models have demonstrated successful operation across complex tasks, and the integration of vision-language models (VLMs) has further enabled embedding models with instruction-following capabilities. However, existing embedding models lack visual-interactive capabil…

Cited by 0SourcecodeScholar
2023

An Attention-Based Approach to Hierarchical Multi-Label Music Instrument Classification

ICASSP 2023accepted

Although music is typically multi-label, many works have studied hierarchical music tagging with simplified settings such as single-label data. Moreover, there lacks a framework to describe various joint training methods under the multi-label setting. In order to discuss the above topics, we introdu…

Cited by 0SourceScholar