← Search

Naoto Inoue

8 accepted papers

2026

Evaluating Cross-Modal Reasoning Ability and Problem Charactaristics with Multimodal Item Response Theory

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their ability for cross‑modal integration. However, current benchmarks are filled with shortcut questions, which can be solved usi…

Cited by 0SourcecodeScholar
2025

LayerD: Decomposing Raster Graphic Designs into Layers

ICCV 2025poster

Designers craft and edit graphic designs in a layer representation, but layer-based editing becomes impossible once composited into a raster image. In this work, we propose LayerD, a method to decompose raster graphic designs into layers for re-editable creative workflow. LayerD addresses the decomp…

2025

Type-R: Automatically Retouching Typos for Text-to-Image Generation

CVPR 2025highlight

While recent text-to-image models can generate photorealistic images from text prompts that reflect detailed instructions, they still face significant challenges in accurately rendering words in the image.In this paper, we propose to retouch erroneous text renderings in the post-processing pipeline.…

Cited by 1SourcePDFScholar
2024

Retrieval-Augmented Layout Transformer for Content-Aware Layout Generation

CVPR 2024poster

Content-aware graphic layout generation aims to automatically arrange visual elements along with a given content such as an e-commerce product image. In this paper we argue that the current layout generation approaches suffer from the limited training data for the high-dimensional layout structure.…

2023

LayoutDM: Discrete Diffusion Model for Controllable Layout Generation

CVPR 2023poster

Controllable layout generation aims at synthesizing plausible arrangement of element bounding boxes with optional constraints, such as type or position of a specific element. In this work, we try to solve a broad range of layout generation tasks in a single model that is based on discrete state-spac…

2023

Towards Flexible Multi-Modal Document Models

CVPR 2023highlight

Creative workflows for generating graphical documents involve complex inter-related tasks, such as aligning elements, choosing appropriate fonts, or employing aesthetically harmonious colors. In this work, we attempt at building a holistic model that can jointly solve many different design tasks. Ou…

2018

Cross-Domain Weakly-Supervised Object Detection Through Progressive Domain Adaptation

CVPR 2018poster

Can we detect common objects in a variety of image domains without instance-level annotations? In this paper, we present a framework for a novel task, cross-domain weakly supervised object detection, which addresses this question. For this paper, we have access to images with instance-level annotati…

2017

Object detection refinement using Markov random field based pruning and learning based rescoring

ICASSP 2017accepted

Contextual information such as the co-occurrence of objects and the location of objects has played an important role in object detection. We present candidate pruning and object rescoring methods that leverage contextual information and that can improve the state-of-the-art CNN-based object detectio…

Cited by 0SourceScholar