← Search

Atsushi Hashimoto

10 accepted papers

2025

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning

ICCV 2025poster

An image captioning model flexibly switching its language pattern, e.g., descriptiveness and length, should be useful since it can be applied to diverse applications. However, despite the dramatic improvement in generative vision-language models, fine-grained control over the properties of generated…

2024

PolarDB: Formula-Driven Dataset for Pre-Training Trajectory Encoders

ICASSP 2024accepted

Formula-driven supervised learning (FDSL) is a growing research topic for finding simple mathematical formulas that generate synthetic data and labels for pre-training neural networks. The main advantage of FDSL is that there is no risk of generating data with ethical implications such as gender bia…

Cited by 0SourceScholar
2024

Vision-Language Interpreter for Robot Task Planning

ICRA 2024poster

Large language models (LLMs) are accelerating the development of language-guided robot planners. Meanwhile, symbolic planners offer the advantage of interpretability. This paper proposes a new task that bridges these two trends, namely, multimodal planning problem specification. The aim is to genera…

Cited by 62SourcecodeScholar
2024

Visuo-Tactile Zero-Shot Object Recognition with Vision-Language Model

IROS 2024poster

Tactile perception is vital, especially when distinguishing visually similar objects. We propose an approach to incorporate tactile data into a Vision-Language Model (VLM) for visuo-tactile zero-shot object recognition. Our approach leverages the zero-shot capability of VLMs to infer tactile propert…

Cited by 1SourceScholar
2023

Invertible Conditional GAN Revisited: Photo-to-Manga Face Translation with Modern Architectures (Student Abstract)

AAAI 2023technical

Recent style translation methods have extended their transferability from texture to geometry. However, performing translation while preserving image content when there is a significant style difference is still an open problem. To overcome this problem, we propose Invertible Conditional Fast GAN (I…

Cited by 1SourcePDFScholar
2023

Learning Food Picking without Food: Fracture Anticipation by Breaking Reusable Fragile Objects

ICRA 2023poster

Food picking is trivial for humans but not for robots, as foods are fragile. Presetting foods' physical properties does not help robots much due to the objects' inter- and intra-category diversity. A recent study proved that learning-based fracture anticipation with tactile sensors could overcome th…

Cited by 3SourceScholar
2022

Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe Flows

COLING 2022main

We present a new multimodal dataset called Visual Recipe Flow, which enables us to learn a cooking action result for each object in a recipe text. The dataset consists of object state changes and the workflow of the recipe text. The state change is represented as an image pair, while the workflow is…

Cited by 11SourcePDFScholar
2020

Partially-Shared Variational Auto-encoders for Unsupervised Domain Adaptation with Target Shift

ECCV 2020poster

This paper discusses unsupervised domain adaptation (UDA) with target shift, i.e., UDA with the non-identical label distributions of the source and target domains. In practice, this is an important problem; as we do not know labels in target domain datasets, we do not know whether or not its distrib…

2018

Photometric Stereo in Participating Media Considering Shape-Dependent Forward Scatter

CVPR 2018poster

Images captured in participating media such as murky water, fog, or smoke are degraded by scattered light. Thus, the use of traditional three-dimensional (3D) reconstruction techniques in such environments is difficult. In this paper, we propose a photometric stereo method for participating media. T…

Cited by 25SourcePDFScholar