← Search

Daiki Kimura

10 accepted papers

2025

Text-Guided Few-Shot Semantic Segmentation with Training-Free Multimodal Feature Matching

ICASSP 2025accepted

This paper addresses few-shot semantic segmentation (FSS) guided by text, where we classify unseen novel classes using image and text references as in-context examples, without the need for training. We enhance the quality and stability of the segmentation masks generated by FSS by combining the cap…

Cited by 0SourceScholar
2024

A Surprisingly Simple Approach to Generalized Few-Shot Semantic Segmentation

NeurIPS 2024poster

The goal of *generalized* few-shot semantic segmentation (GFSS) is to recognize *novel-class* objects through training with a few annotated examples and the *base-class* model that learned the knowledge about the base classes. Unlike the classic few-shot semantic segmentation, GFSS aims to classify…

2024

SAR2NDVI: Pre-Training for SAR-to-NDVI Image Translation

ICASSP 2024accepted

Geospatial machine learning is of growing importance in various global remote-sensing applications, particularly in the realm of vegetation monitoring. However, acquiring accurate ground truth data for geospatial tasks remains a significant challenge, often entailing considerable time and effort. Fo…

Cited by 0SourceScholar
2023

Learning Neuro-Symbolic World Models with Conversational Proprioception

ACL 2023short

The recent emergence of Neuro-Symbolic Agent (NeSA) approaches to natural language-based interactions calls for the investigation of model-based approaches. In contrast to model-free approaches, which existing NeSAs take, learning an explicit world model has an interesting potential especially in th…

Cited by 1SourcePDFScholar
2023

Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement Learning

ACL 2023long

Text-based reinforcement learning agents have predominantly been neural network-based models with embeddings-based representation, learning uninterpretable policies that often do not generalize well to unseen games. On the other hand, neuro-symbolic methods, specifically those that leverage an inter…

2022

DiffG-RL: Leveraging Difference between Environment State and Common Sense

EMNLP 2022finding

Taking into account background knowledge as the context has always been an important part of solving tasks that involve natural language. One representative example of such tasks is text-based games, where players need to make decisions based on both description text previously shown in the game, an…

Cited by 0SourcePDFScholar
2022

X-FACTOR: A Cross-metric Evaluation of Factual Correctness in Abstractive Summarization

EMNLP 2022main

Abstractive summarization models often produce factually inconsistent summaries that are not supported by the original article. Recently, a number of fact-consistent evaluation techniques have been proposed to address this issue; however, a detailed analysis of how these metrics agree with one anoth…

Cited by 12SourcePDFScholar
2021

Neuro-Symbolic Approaches for Text-Based Policy Learning

EMNLP 2021main

Text-Based Games (TBGs) have emerged as important testbeds for reinforcement learning (RL) in the natural language domain. Previous methods using LSTM-based action policies are uninterpretable and often overfit the training games showing poor performance to unseen test games. We present SymboLic Act…

2021

Neuro-Symbolic Reinforcement Learning with First-Order Logic

EMNLP 2021main

Deep reinforcement learning (RL) methods often require many trials before convergence, and no direct interpretability of trained policies is provided. In order to achieve fast convergence and interpretability for the policy in RL, we propose a novel RL method for text-based games with a recent neuro…

Cited by 48SourcePDFScholar
2018

MaestROB: A Robotics Framework for Integrated Orchestration of Low-Level Control and High-Level Reasoning

ICRA 2018poster

This paper describes a framework called MaestROBe It is designed to make the robots perform complex tasks with high precision by simple high-level instructions given by natural language or demonstration. To realize this, it handles a hierarchical structure by using the knowledge stored in the forms…

Cited by 25SourceScholar