← Search

Chiori Hori

16 accepted papers

2026

SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs

ICLR 2026poster

We present SpinBench, a cognitively grounded diagnostic benchmark for evaluating spatial reasoning in vision language models (VLMs). SpinBench is designed around the core challenge of spatial reasoning: perspective taking, the ability to reason about how scenes and object relations change under vie…

Cited by 0SourcecodeScholar
2025

Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration Videos

ICASSP 2025accepted

Understanding human actions could allow robots to perform a large spectrum of complex manipulation tasks and make collaboration with humans easier. Recently, multimodal scene understanding using audio-visual Transformers has been used to generate robot action sequences from videos of human demonstra…

Cited by 0SourceScholar
2024

Generation or Replication: Auscultating Audio Latent Diffusion Models

ICASSP 2024accepted

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt at understanding the inner workings of audio latent diffusi…

Cited by 0SourceScholar
2024

Interactive Planning Using Large Language Models for Partially Observable Robotic Tasks

ICRA 2024poster

Designing robotic agents to perform open vocabulary tasks has been the long-standing goal in robotics and AI. Recently, Large Language Models (LLMs) have achieved impressive results in creating robotic agents for performing open vocabulary tasks. However, planning for these tasks in the presence of…

Cited by 31SourceScholar
2024

NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization

ICASSP 2024accepted

Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields (NFs) which map from sound source direction to HRTF have gained attention. Existing NF-based methods focused on estimat…

Cited by 0SourceScholar
2024

WI-FI based Indoor Monitoring Enhanced by Multimodal Fusion

ICASSP 2024accepted

Indoor monitoring systems are in high demand to protect vulnerable people, especially when they are alone at home, in nursing homes, hospitals, etc. Although surveillance systems in public spaces use cameras and microphones to find incidents, indoor monitoring in personal spaces needs to protect pri…

Cited by 0SourceScholar
2022

(2.5+1)D Spatio-Temporal Scene Graphs for Video Question Answering

AAAI 2022technical

Spatio-temporal scene-graph approaches to video-based reasoning tasks, such as video question-answering (QA), typically construct such graphs for every video frame. These approaches often ignore the fact that videos are essentially sequences of 2D ``views'' of events happening in a 3D space, and tha…

Cited by 25SourcePDFScholar
2022

Audio-Visual Scene-Aware Dialog and Reasoning Using Audio-Visual Transformers with Joint Student-Teacher Learning

ICASSP 2022accepted

In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technology Challenges (DSTC7, DSTC8). In these challenges, the best-performing systems…

Cited by 0SourceScholar
2021

Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled Transformers

AAAI 2021technical

Given an input video, its associated audio, and a brief caption, the audio-visual scene aware dialog (AVSD) task requires an agent to indulge in a question-answer dialog with a human about the audio-visual content. This task thus poses a challenging multi-modal representation learning and reasoning…

Cited by 50SourcePDFScholar
2020

Multi-Layer Content Interaction Through Quaternion Product for Visual Question Answering

ICASSP 2020accepted

Multi-modality fusion technologies have greatly improved the performance of neural network-based Video Description/Caption, Visual Question Answering (VQA) and Audio Visual Scene-aware Dialog (AVSD) over the recent years. Most previous approaches only explore the last layers of multiple layer featur…

Cited by 0SourceScholar
2019

Audio Visual Scene-Aware Dialog

CVPR 2019poster

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To answer successfully, agents must ground concepts from the question in the video whi…

Cited by 226PDFcodeScholar
2019

End-to-end Audio Visual Scene-aware Dialog Using Multimodal Attention-based Video Features

ICASSP 2019accepted

In order for machines interacting with the real world to have conversations with users about the objects and events around them, they need to understand dynamic audiovisual scenes. The recent revolution of neural network models allows us to combine various modules into a single end-to-end differenti…

Cited by 0SourceScholar
2017

Attention-Based Multimodal Fusion for Video Description

ICCV 2017poster

Current methods for video description are based on encoder-decoder sentence generation using recurrent neural networks (RNNs). Recent work has demonstrated the advantages of integrating temporal attention mechanisms into these models, in which the decoder network predicts each word in the descriptio…

Cited by 469PDFScholar
2016

Minimum word error training of long short-term memory recurrent neural network language models for speech recognition

ICASSP 2016accepted

This paper describes minimum word error (MWE) training of recurrent neural network language models (RNNLMs) for speech recognition. RNNLMs are usually trained to minimize a cross entropy of estimated word probabilities against the correct word sequence, which corresponds to maximum likelihood criter…

Cited by 0SourceScholar
2015

Speaker adaptive training for deep neural networks embedding linear transformation networks

ICASSP 2015accepted

Recently, a novel speaker adaptation method was proposed that applied the Speaker Adaptive Training (SAT) concept to a speech recognizer consisting of a Deep Neural Network (DNN) and a Hidden Markov Model (HMM), and its utility was demonstrated. This method implements the SAT scheme by allocating on…

Cited by 15SourceScholar