← Search

Akihiro Sugimoto

10 accepted papers

2026

SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation

AAAI 2026technical

State-of-the-art text-to-image models produce visually impressive results but often struggle with precise alignment to text prompts, leading to missing critical elements or unintended blending of distinct concepts. We propose a novel approach that learns a high-success-rate distribution conditioned

Cited by 0SourcePDFScholar
2024

EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension

CVPR 2024poster

Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently necessitating the requirement of sustaining up-to-date object knowledge for open-world comprehension. Instead of relying on large…

Cited by 25SourcePDFScholar
2023

A-Cap: Anticipation Captioning With Commonsense Knowledge

CVPR 2023poster

Humans possess the capacity to reason about the future based on a sparse collection of visual cues acquired over time. In order to emulate this ability, we introduce a novel task called Anticipation Captioning, which generates a caption for an unseen oracle image using a sparsely temporally-ordered…

Cited by 5SourcePDFScholar
2022

NOC-REK: Novel Object Captioning With Retrieved Vocabulary From External Knowledge

CVPR 2022poster

Novel object captioning aims at describing objects absent from training data, with the key ingredient being the provision of object vocabulary to the model. Although existing methods heavily rely on an object detection model, we view the detection step as vocabulary retrieval from an external knowle…

Cited by 21PDFScholar
2021

Agent-Environment Network for Temporal Action Proposal Generation

ICASSP 2021accepted

Temporal action proposal generation is an essential and challenging task that aims at localizing temporal intervals containing human actions in untrimmed videos. Most of existing approaches are unable to follow the human cognitive process of understanding the video context due to lack of attention m…

Cited by 0SourceScholar
2020

Minimal Rolling Shutter Absolute Pose with Unknown Focal Length and Radial Distortion

ECCV 2020poster

The internal geometry of most modern consumer cameras is not adequately described by the perspective projection. Almost all cameras exhibit some radial lens distortion and are equipped with electronic rolling shutter that induces distortions when the camera moves during the image capture. When focal…

2020

TetraTSDF: 3D Human Reconstruction From a Single Image With a Tetrahedral Outer Shell

CVPR 2020poster

Recovering the 3D shape of a person from its 2D appearance is ill-posed due to ambiguities. Nevertheless, with the help of convolutional neural networks (CNN) and prior knowledge on the 3D human body, it is possible to overcome such ambiguities to recover detailed 3D shapes of human bodies from sing…

Cited by 48PDFcodeScholar