← Search

Yuta Nakashima

27 accepted papers

2026

EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories

CVPR 2026

The widespread adoption of text-to-image (T2I) generation has raised concerns about privacy, bias, and copyright violations. Concept erasure techniques offer a promising solution by selectively removing undesired concepts from pre-trained models without requiring full retraining. However, these meth

Cited by 0SourcecodeScholar
2026

Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMs

ICLR 2026poster

Temporally localizing user-queried events through natural language is a crucial capability for video models. Recent methods predominantly adapt video LLMs to generate event boundary timestamps for temporal localization tasks, which struggle to leverage LLMs' pre-trained semantic understanding capabi…

Cited by 0SourcecodeScholar
2026

QuMAB: Query-based Multi-annotator Behavior Pattern Learning

AAAI 2026technical

Multi-annotator learning traditionally aggregates diverse annotations to approximate a single “ground truth”, treating disagreements as noise. However, this paradigm faces fundamental challenges: subjective tasks often lack absolute ground truth, and sparse annotation coverage makes aggregation stat

Cited by 0SourcePDFScholar
2026

SimLabel: Similarity-Weighted Semi-supervision for Multi-annotator Learning with Missing Labels

AAAI 2026technical

Multi-annotator learning (MAL) aims to model annotator-specific labeling patterns. However, existing methods face a critical challenge: they simply skip updating annotator-specific model parameters when encountering missing labels—a common scenario in real-world crowdsourced datasets where each anno

Cited by 0SourcePDFScholar
2025

Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation

ICCV 2025poster

Gender bias in vision-language foundation models (VLMs) raises concerns about their safe deployment and is typically evaluated using benchmarks with gender annotations on real-world images. However, as these benchmarks often contain spurious correlations between gender and non-gender features, such…

Cited by 0SourcePDFScholar
2025

Processing and acquisition traces in visual encoders: What does CLIP know about your camera?

ICCV 2025poster

Prior work has analyzed the robustness of visual encoders to image transformations and corruptions, particularly in cases where such alterations are not seen during training. When this occurs, they introduce a form of distribution shift at test time, often leading to performance degradation. The pri…

2025

Putting People in LLMs’ Shoes: Generating Better Answers via Question Rewriter

AAAI 2025technical

Large Language Models (LLMs) have demonstrated significant capabilities, particularly in the domain of question answering (QA). However, their effectiveness in QA is often undermined by the vagueness of user questions. To address this issue, we introduce single-round instance-level prompt optimizat…

2025

ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training

COLING 2025main

Recent approaches for visually-rich document understanding (VrDU) uses manually annotated semantic groups, where a semantic group encompasses all semantically relevant but not obviously grouped words. As OCR tools are unable to automatically identify such grouping, we argue that current VrDU approac…

2025

SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP

ICLR 2025poster

Large-scale vision-language models, such as CLIP, are known to contain societal bias regarding protected attributes (e.g., gender, age). This paper aims to address the problems of societal bias in CLIP. Although previous studies have proposed to debias societal bias through adversarial learning or t…

Cited by 2SourcePDFScholar
2025

Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown

ICCV 2025poster

The real value of knowledge lies not just in its accumulation, but in its potential to be harnessed effectively to conquer the unknown. Although recent multimodal large language models (MLLMs) exhibit impressing multimodal capabilities, they often fail in rarely encountered domain-specific tasks due…

2024

DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models

NeurIPS 2024poster

Large language models (LLMs) have recently showcased remarkable capabilities, spanning a wide range of tasks and applications, including those in the medical domain. Models like GPT-4 excel in medical question answering but may face challenges in the lack of interpretability when handling complex ta…

2024

From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment

EMNLP 2024main

Large language models (LLMs) have enhanced the capacity of vision-language models to caption visual text. This generative approach to image caption enrichment further makes textual captions more descriptive, improving alignment with the visual context. However, while many studies focus on the benefi…

Cited by 3SourcePDFScholar
2024

Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes

EMNLP 2024main

We tackle societal bias in image-text datasets by removing spurious correlations between protected groups and image attributes. Traditional methods only target labeled attributes, ignoring biases from unlabeled ones. Using text-guided inpainting models, our approach ensures protected group independe…

Cited by 2SourcePDFScholar
2024

Would Deep Generative Models Amplify Bias in Future Models?

CVPR 2024poster

We investigate the impact of deep generative models on potential social biases in upcoming computer vision models. As the internet witnesses an increasing influx of AI-generated images concerns arise regarding inherent biases that may accompany them potentially leading to the dissemination of harmfu…

Cited by 11SourcePDFScholar
2023

Learning Bottleneck Concepts in Image Classification

CVPR 2023poster

Interpreting and explaining the behavior of deep neural networks is critical for many tasks. Explainable AI provides a way to address this challenge, mostly by providing per-pixel relevance to the decision. Yet, interpreting such explanations may require expert knowledge. Some recent attempts toward…

2023

Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation

CVPR 2023poster

Human evaluation is critical for validating the performance of text-to-image generative models, as this highly cognitive process requires deep comprehension of text and images. However, our survey of 37 recent papers reveals that many works rely solely on automatic measures (e.g., FID) or perform po…

2023

Uncurated Image-Text Datasets: Shedding Light on Demographic Bias

CVPR 2023highlight

The increasing tendency to collect large and uncurated datasets to train vision-and-language models has raised concerns about fair representations. It is known that even small but manually annotated datasets, such as MSCOCO, are affected by societal bias. This problem, far from being solved, may be…

2022

AxIoU: An Axiomatically Justified Measure for Video Moment Retrieval

CVPR 2022poster

Evaluation measures have a crucial impact on the direction of research. Therefore, it is of utmost importance to develop appropriate and reliable evaluation measures for new applications where conventional measures are not well suited. Video Moment Retrieval (VMR) is one such application, and the cu…

Cited by 2PDFScholar
2022

Deep Gesture Generation for Social Robots Using Type-Specific Libraries

IROS 2022poster

Body language such as conversational gesture is a powerful way to ease communication. Conversational gestures do not only make a speech more lively but also contain semantic meaning that helps to stress important information in the discussion. In the field of robotics, giving conversational agents (…

Cited by 6SourceScholar
2022

Optimal Correction Cost for Object Detection Evaluation

CVPR 2022poster

Mean Average Precision (mAP) is the primary evaluation measure for object detection. Although object detection has a broad range of applications, mAP evaluates detectors in terms of the performance of ranked instance retrieval. Such the assumption for the evaluation task does not suit some downstrea…

Cited by 19PDFcodeScholar
2021

Explain Me the Painting: Multi-Topic Knowledgeable Art Description Generation

ICCV 2021poster

Have you ever looked at a painting and wondered what is the story behind it? This work presents a framework to bring art closer to people by generating comprehensive descriptions of fine-art paintings. Generating informative descriptions for artworks, however, is extremely challenging, as it require…

Cited by 54PDFcodeScholar
2021

SCOUTER: Slot Attention-Based Classifier for Explainable Image Recognition

ICCV 2021poster

Explainable artificial intelligence has been gaining attention in the past few years. However, most existing methods are based on gradients or intermediate features, which are not directly involved in the decision-making process of the classifier. In this paper, we propose a slot attention-based cla…

Cited by 58PDFcodeScholar
2021

WRIME: A New Dataset for Emotional Intensity Estimation with Subjective and Objective Annotations

NAACL 2021long

We annotate 17,000 SNS posts with both the writer’s subjective emotional intensity and the reader’s objective one to construct a Japanese emotion analysis dataset. In this study, we explore the difference between the emotional intensity of the writer and that of the readers with this dataset. We fou…

2020

Knowledge-Based Video Question Answering with Unsupervised Scene Descriptions

ECCV 2020poster

To understand movies, humans constantly reason over the dialogues and actions shown in specific scenes and relate them to the overall storyline already seen. Inspired by this behaviour, we design ROLL, a model for knowledge-based video story question answering that leverages three crucial aspects of…