← Search

Taichi Nishimura

8 accepted papers

2026

CASTELLA: LONG AUDIO DATASET WITH CAPTIONS AND TEMPORAL BOUNDARIES

ICASSP 2026poster

We introduce CASTELLA, a human-annotated audio benchmark for the task of audio moment retrieval (AMR). Although AMR has various useful potential applications, there is still no established benchmark with real-world data. The initial study of AMR trained the models solely on synthetic datasets. Moreo…

Cited by 0SourcePDFScholar
2026

Developing Vision-Language-Action Model from Egocentric Videos

ICRA 2026poster

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models (VLAs), egocentric videos offer a scalable alternative. How…

2025

DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

ICASSP 2025accepted

Current audio-visual representation learning can capture rough object categories (e.g., "animals" and "instruments"), but it lacks the ability to recognize fine-grained details, such as specific categories like "dogs" and "flutes" within animals and instruments. To address this issue, we introduce D…

Cited by 0SourceScholar
2025

Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision

CVPR 2025highlight

Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation trajectories requires a large and diverse collection of detailed manipulation demonstration…

Cited by 0SourcePDFScholar
2024

Automatic Construction of a Large-Scale Corpus for Geoparsing Using Wikipedia Hyperlinks

COLING 2024main

Geoparsing is the task of estimating the latitude and longitude (coordinates) of location expressions in texts. Geoparsing must deal with the ambiguity of the expressions that indicate multiple locations with the same notation. For evaluating geoparsing systems, several corpora have been proposed in…

Cited by 0SourcePDFScholar
2024

Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection

EMNLP 2024system demonstrations

We propose Lighthouse, a user-friendly library for reproducible video moment retrieval and highlight detection (MR-HD). Although researchers proposed various MR-HD approaches, the research community holds two main issues. The first is a lack of comprehensive and reproducible experiments across vario…

2022

Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe Flows

COLING 2022main

We present a new multimodal dataset called Visual Recipe Flow, which enables us to learn a cooking action result for each object in a recipe text. The dataset consists of object state changes and the workflow of the recipe text. The state change is represented as an image pair, while the workflow is…

Cited by 11SourcePDFScholar