← Search

Mayu Otani

12 accepted papers

2026

Evaluating Cross-Modal Reasoning Ability and Problem Charactaristics with Multimodal Item Response Theory

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their ability for cross‑modal integration. However, current benchmarks are filled with shortcut questions, which can be solved usi…

Cited by 0SourcecodeScholar
2026

Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMs

ICLR 2026poster

Temporally localizing user-queried events through natural language is a crucial capability for video models. Recent methods predominantly adapt video LLMs to generate event boundary timestamps for temporal localization tasks, which struggle to leverage LLMs' pre-trained semantic understanding capabi…

Cited by 0SourcecodeScholar
2024

LayoutFlow: Flow Matching for Layout Generation

ECCV 2024poster

"Finding a suitable layout represents a crucial task for diverse applications in graphic design. Motivated by simpler and smoother sampling trajectories, we explore the use of Flow Matching as an alternative to current diffusion-based layout generation models. Specifically, we propose LayoutFlow, an…

2024

Robust Nearest Neighbors for Source-Free Domain Adaptation under Class Distribution Shift

ECCV 2024poster

"The goal of source-free domain adaptation (SFDA) is retraining a model fit on data from a source domain (drawings) to classify data from a target domain (photos) employing only the target samples. In addition to the domain shift, in a realistic scenario, the number of samples per class on source an…

Cited by 1SourcePDFScholar
2024

Would Deep Generative Models Amplify Bias in Future Models?

CVPR 2024poster

We investigate the impact of deep generative models on potential social biases in upcoming computer vision models. As the internet witnesses an increasing influx of AI-generated images concerns arise regarding inherent biases that may accompany them potentially leading to the dissemination of harmfu…

Cited by 11SourcePDFScholar
2023

LayoutDM: Discrete Diffusion Model for Controllable Layout Generation

CVPR 2023poster

Controllable layout generation aims at synthesizing plausible arrangement of element bounding boxes with optional constraints, such as type or position of a specific element. In this work, we try to solve a broad range of layout generation tasks in a single model that is based on discrete state-spac…

2023

Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation

CVPR 2023poster

Human evaluation is critical for validating the performance of text-to-image generative models, as this highly cognitive process requires deep comprehension of text and images. However, our survey of 37 recent papers reveals that many works rely solely on automatic measures (e.g., FID) or perform po…

2023

Towards Flexible Multi-Modal Document Models

CVPR 2023highlight

Creative workflows for generating graphical documents involve complex inter-related tasks, such as aligning elements, choosing appropriate fonts, or employing aesthetically harmonious colors. In this work, we attempt at building a holistic model that can jointly solve many different design tasks. Ou…

2022

AxIoU: An Axiomatically Justified Measure for Video Moment Retrieval

CVPR 2022poster

Evaluation measures have a crucial impact on the direction of research. Therefore, it is of utmost importance to develop appropriate and reliable evaluation measures for new applications where conventional measures are not well suited. Video Moment Retrieval (VMR) is one such application, and the cu…

Cited by 2PDFScholar
2022

Optimal Correction Cost for Object Detection Evaluation

CVPR 2022poster

Mean Average Precision (mAP) is the primary evaluation measure for object detection. Although object detection has a broad range of applications, mAP evaluates detectors in terms of the performance of ranked instance retrieval. Such the assumption for the evaluation task does not suit some downstrea…

Cited by 19PDFcodeScholar