← Search

Yoshitaka Ushiku

33 accepted papers

2026

Robust and Resilient Soft Robotic Object Insertion with Compliance-Enabled Contact Formation and Failure Recovery

ICRA 2026poster

We address robust and resilient object insertion using a passively compliant soft wrist that permits large deformations and safely absorbs contacts, without high-frequency control or force sensing. To improve robustness, we structure the task as compliance-enabled contact formations: a sequence of c…

2026

SCU-Hand with Integrated Single-Sheet Valve: A Funnel-Shaped Robotic Hand for Milligram-Scale Powder Handling

ICRA 2026poster

Laboratory Automation (LA) has the potential to accelerate solid-state materials discovery by enabling continuous robotic operation without human intervention. While robotic systems have been developed for tasks such as powder grinding and X-ray diffraction (XRD) analysis, fully automating powder ha…

2025

AgroBench: Vision-Language Model Benchmark in Agriculture

ICCV 2025poster

Precise automated understanding of agricultural tasks such as disease identification is essential for the sustainable crop production. Recent advances in vision-language models (VLMs) are expected to further expand the range of agricultural tasks by facilitating human-model interaction through easy,…

2025

CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning

ICCV 2025poster

An image captioning model flexibly switching its language pattern, e.g., descriptiveness and length, should be useful since it can be applied to diverse applications. However, despite the dramatic improvement in generative vision-language models, fine-grained control over the properties of generated…

2025

Rethinking the role of frames for SE(3)-invariant crystal structure modeling

ICLR 2025poster

Crystal structure modeling with graph neural networks is essential for various applications in materials informatics, and capturing SE(3)-invariant geometric features is a fundamental requirement for these networks. A straightforward approach is to model with orientation-standardized structures thro…

Cited by 1SourcePDFScholar
2025

SCU-Hand: Soft Conical Universal Robotic Hand for Scooping Granular Media from Containers of Various Sizes

ICRA 2025

Automating small-scale experiments in materials science presents challenges due to the heterogeneous nature of experimental setups. This study introduces the SCU-Hand (Soft Conical Universal Robot Hand), a novel end-effector designed to automate the task of scooping powdered samples from various con

Cited by 4SourceScholar
2025

Where is the answer? An empirical study of positional bias for parametric knowledge extraction in language model

NAACL 2025long

Language model (LM) stores diverse factual knowledge in their parameters, which is learned during self-supervised training on unlabeled documents and is made extractable by instruction-tuning. For knowledge-intensive tasks, it is essential to memorize information in a way that makes it extractable f…

Cited by 0SourcePDFScholar
2024

COM Kitchens: An Unedited Overhead-view Procedural Videos Dataset a Vision-Language Benchmark

ECCV 2024poster

"Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resources, making it challenging to query instructional contents from raw video observa…

2024

Crystalformer: Infinitely Connected Attention for Periodic Structure Encoding

ICLR 2024poster

Predicting physical properties of materials from their crystal structures is a fundamental problem in materials science. In peripheral areas such as the prediction of molecular properties, fully connected attention networks have been shown to be successful. However, unlike these finite atom arrangem…

Cited by 11SourcePDFScholar
2024

PolarDB: Formula-Driven Dataset for Pre-Training Trajectory Encoders

ICASSP 2024accepted

Formula-driven supervised learning (FDSL) is a growing research topic for finding simple mathematical formulas that generate synthetic data and labels for pre-training neural networks. The main advantage of FDSL is that there is no risk of generating data with ethical implications such as gender bia…

Cited by 0SourceScholar
2024

Vision-Language Interpreter for Robot Task Planning

ICRA 2024poster

Large language models (LLMs) are accelerating the development of language-guided robot planners. Meanwhile, symbolic planners offer the advantage of interpretability. This paper proposes a new task that bridges these two trends, namely, multimodal planning problem specification. The aim is to genera…

Cited by 62SourcecodeScholar
2023

Robotic Powder Grinding with Audio-Visual Feedback for Laboratory Automation in Materials Science

IROS 2023poster

This study focuses on the powder grinding process, which is a necessary step for material synthesis in materials science experiments. In material science, powder grinding is a time-consuming process that is typically executed by hand, as commercial grinding machines are unsuitable for samples of sma…

Cited by 1SourceScholar
2022

Robotic Powder Grinding with a Soft Jig for Laboratory Automation in Material Science

IROS 2022poster

Grinding materials into a fine powder is a time-consuming task in material science that is generally performed by hand, as current automated grinding machines might not be suitable for preparing small-sized samples. This study presents a robotic powder grinding system for laboratory automation in ma…

Cited by 11SourceScholar
2022

Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe Flows

COLING 2022main

We present a new multimodal dataset called Visual Recipe Flow, which enables us to learn a cooking action result for each object in a recipe text. The dataset consists of object state changes and the workflow of the recipe text. The state change is represented as an image pair, while the workflow is…

Cited by 11SourcePDFScholar
2019

Generating Easy-to-Understand Referring Expressions for Target Identifications

ICCV 2019poster

This paper addresses the generation of referring expressions that not only refer to objects correctly but also let humans find them quickly. As a target becomes relatively less salient, identifying referred objects itself becomes more difficult. However, the existing studies regarded all sentences t…

Cited by 30PDFcodeScholar
2019

Strong-Weak Distribution Alignment for Adaptive Object Detection

CVPR 2019poster

We propose an approach for unsupervised adaptation of object detectors from label-rich to label-poor domains which can significantly reduce annotation costs associated with detection. Recently, approaches that align distributions of source and target images using an adversarial loss have been proven…

Cited by 853PDFcodeScholar
2018

Customized Image Narrative Generation via Interactive Visual Question Generation and Answering

CVPR 2018poster

Image description task has been invariably examined in a static manner with qualitative presumptions held to be universally applicable, regardless of the scope or target of the description. In practice, however, different viewers may pay attention to different aspects of the image, and yield differe…

Cited by 10SourcePDFScholar
2018

Learning from Between-class Examples for Deep Sound Recognition

ICLR 2018poster

Deep learning methods have achieved high performance in sound recognition tasks. Deciding how to feed the training data is important for further performance improvement. We propose a novel learning method for deep sound recognition: Between-Class learning (BC learning). Our strategy is to learn a di…

2018

Maximum Classifier Discrepancy for Unsupervised Domain Adaptation

CVPR 2018poster

In this work, we present a method for unsupervised domain adaptation. Many adversarial learning methods train domain classifier networks to distinguish the features as either a source or target and train a feature generator network to mimic the discriminator. Two problems exist with these methods.…

2018

Visual Question Generation for Class Acquisition of Unknown Objects

ECCV 2018poster

Traditional image recognition methods only consider objects belonging to already learned classes. However, since training a recognition model with every object class in the world is unfeasible, a way of getting information on unknown objects (i.e., objects whose class has not been learned) is necess…

2017

MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes

IROS 2017poster

This work addresses the semantic segmentation of images of street scenes for autonomous vehicles based on a new RGB-Thermal dataset, which is also introduced in this paper. An increasing interest in self-driving vehicles has brought the adaptation of semantic segmentation to self-driving systems. Ho…

Cited by 617SourceScholar
2017

Spatio-Temporal Person Retrieval via Natural Language Queries

ICCV 2017poster

In this paper, we address the problem of spatio-temporal person retrieval from videos using a natural language query, in which we output a tube (i.e., a sequence of bounding boxes) which encloses the person described by the query. For this problem, we introduce a novel dataset consisting of videos c…

Cited by 71PDFcodeScholar
2015

Common Subspace for Model and Similarity: Phrase Learning for Caption Generation From Images

ICCV 2015poster

Generating captions to describe images is a fundamental problem that combines computer vision and natural language processing. Recent works focus on descriptive phrases, such as "a white dog" to explain the visual composites of an input image. The phrases can not only express objects, attributes, ev…

Cited by 81PDFcodeScholar