← Search

Takahito Kawanishi

8 accepted papers

2020

Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms

ICASSP 2020accepted

We propose a trilingual semantic embedding model that associates visual objects in images with segments of speech signals corresponding to spoken words in an unsupervised manner. Unlike the existing models, our model incorporates three different languages, namely, English, Hindi, and Japanese. To bu…

Cited by 0SourceScholar
2019

Learning Search Path for Region-level Image Matching

ICASSP 2019accepted

Finding a region of an image which matches to a query from a large number of candidates is a fundamental problem in image processing. The exhaustive nature of the sliding window approach has encouraged works that can reduce the run time by skipping unnecessary windows or pixels that do not play a su…

Cited by 0SourceScholar
2019

Seeing through Sounds: Predicting Visual Semantic Segmentation Results from Multichannel Audio Signals

ICASSP 2019accepted

Sounds provide us with vast amounts of information about surrounding objects and can even remind us visual images of them. Is it possible to implement this noteworthy human ability on machines? In this paper, we study a new task that consists of predicting image recognition results in the form of se…

Cited by 0SourceScholar
2019

Subspace Structure-Aware Spectral Clustering for Robust Subspace Clustering

ICCV 2019poster

Subspace clustering is the problem of partitioning data drawn from a union of multiple subspaces. The most popular subspace clustering framework in recent years is the graph clustering-based approach, which performs subspace clustering in two steps: graph construction and graph clustering. Although…

Cited by 7PDFScholar
2017

Deep salience map guided arbitrary direction scene text recognition

ICASSP 2017accepted

Irregular scene text such as curved, rotated or perspective texts commonly appear in natural scene images due to different camera view points, special design purposes etc. In this work, we propose a text salience map guided model to recognize these arbitrary direction scene texts. We train a deep Fu…

Cited by 0SourceScholar
2017

Edited film alignment via selective Hough transform and accurate template matching

ICASSP 2017accepted

Edited film alignment is the post-production process of finding small parts of unedited footage that temporally and spatially match an edited film. The huge amount of data to be processed makes significant downsampling of the videos essential in real-life applications. Simultaneously, professional u…

Cited by 0SourceScholar
2016

Scene text recognition with high performance CNN classifier and efficient word inference

ICASSP 2016accepted

The recognition of text in natural scene images is a practical yet challenging task due to the large variations in backgrounds, textures, fonts, and illumination conditions. In this paper, we propose a highly accurate character recognition model by utilizing the representational power of a specially…

Cited by 0SourceScholar