← Search

Kohei Uehara

4 accepted papers

2025

Memory-Maze: Scenario Driven Visual Language Navigation Benchmark for Guiding Blind People

RA-L 2025

Visual Language Navigation (VLN) powered robots have the potential to guide blind people by understanding route instructions provided by sighted passersby. This capability allows robots to operate in environments often unknown a prior. Existing VLN models are insufficient for the scenario of navigat

Cited by 3SourceScholar
2024

Content-Specific Humorous Image Captioning Using Incongruity Resolution Chain-of-Thought

NAACL 2024findings

Although automated image captioning methods have benefited considerably from the development of large language models (LLMs), generating humorous captions is still a challenging task. Humorous captions generated by humans are unique to the image and reflect the content of the image. However, caption…

2018

Visual Question Generation for Class Acquisition of Unknown Objects

ECCV 2018poster

Traditional image recognition methods only consider objects belonging to already learned classes. However, since training a recognition model with every object class in the world is unfeasible, a way of getting information on unknown objects (i.e., objects whose class has not been learned) is necess…