← Search

Kunio Kashino

23 accepted papers

2025

Hyperbolic PHATE: Visualizing Continuous Hierarchy of Latent Differentiation Structures

ICASSP 2025accepted

This paper proposes a method for embedding diffusion potentials into a hyperbolic space in order to visualize the differentiation structure consisting of diffusion and branching inherent in high-dimensional data. In recent years, the rapid development of single-cell sequencing in the field of biolog…

Cited by 0SourceScholar
2024

Warped Diffusion for Latent Differentiation Inference

AISTATS 2024poster

This paper proposes a Bayesian nonparametric diffusion model with a black-box warping function represented by a Gaussian process to infer potential diffusion structures latent in observed data, such as differentiation mechanisms of living cells and phylogenetic evolution processes of media informati…

2023

Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input

ICASSP 2023accepted

Masked Autoencoders is a simple yet powerful self-supervised learning method. However, it learns representations indirectly by reconstructing masked input patches. Several methods learn representations directly by predicting representations of masked patches; however, we think using all patches to e…

Cited by 0SourceScholar
2021

Reflectance-Oriented Probabilistic Equalization for Image Enhancement

ICASSP 2021accepted

Despite recent advances in image enhancement, it remains difficult for existing approaches to adaptively improve the brightness and contrast for both low-light and normal-light images. To solve this problem, we propose a novel 2D histogram equalization approach. It assumes intensity occurrence and c…

Cited by 0SourceScholar
2020

Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms

ICASSP 2020accepted

We propose a trilingual semantic embedding model that associates visual objects in images with segments of speech signals corresponding to spoken words in an unsupervised manner. Unlike the existing models, our model incorporates three different languages, namely, English, Hindi, and Japanese. To bu…

Cited by 0SourceScholar
2019

Learning Search Path for Region-level Image Matching

ICASSP 2019accepted

Finding a region of an image which matches to a query from a large number of candidates is a fundamental problem in image processing. The exhaustive nature of the sliding window approach has encouraged works that can reduce the run time by skipping unnecessary windows or pixels that do not play a su…

Cited by 0SourceScholar
2019

Prewarping Siamese Network: Learning Local Representations for Online Signature Verification

ICASSP 2019accepted

We propose a neural network-based framework for learning local representations of multivariate time series, and demonstrate its effectiveness for online signature verification. In contrast to related works that optimize a global distance objective, we incorporate a Siamese network into dynamic time…

Cited by 0SourceScholar
2019

Seeing through Sounds: Predicting Visual Semantic Segmentation Results from Multichannel Audio Signals

ICASSP 2019accepted

Sounds provide us with vast amounts of information about surrounding objects and can even remind us visual images of them. Is it possible to implement this noteworthy human ability on machines? In this paper, we study a new task that consists of predicting image recognition results in the form of se…

Cited by 0SourceScholar
2019

Subspace Structure-Aware Spectral Clustering for Robust Subspace Clustering

ICCV 2019poster

Subspace clustering is the problem of partitioning data drawn from a union of multiple subspaces. The most popular subspace clustering framework in recent years is the graph clustering-based approach, which performs subspace clustering in two steps: graph construction and graph clustering. Although…

Cited by 7PDFScholar
2018

Generating Sound Words from Audio Signals of Acoustic Events with Sequence-to-Sequence Model

ICASSP 2018accepted

Representing various sounds in language, such as sound words, or onomatopoeias, is not only useful as an auxiliary means for automatic speech recognition, but also essential in emerging fields such as natural human-machine communication, searching audio archives for acoustic events, and abnormality…

Cited by 0SourceScholar
2018

Generative Adversarial Image Synthesis With Decision Tree Latent Controller

CVPR 2018poster

This paper proposes the decision tree latent controller generative adversarial network (DTLC-GAN), an extension of a GAN that can learn hierarchically interpretable representations without relying on detailed supervision. To impose a hierarchical inclusion structure on latent variables, we incorpora…

Cited by 25SourcePDFScholar
2017

Deep salience map guided arbitrary direction scene text recognition

ICASSP 2017accepted

Irregular scene text such as curved, rotated or perspective texts commonly appear in natural scene images due to different camera view points, special design purposes etc. In this work, we propose a text salience map guided model to recognize these arbitrary direction scene texts. We train a deep Fu…

Cited by 0SourceScholar
2017

Edited film alignment via selective Hough transform and accurate template matching

ICASSP 2017accepted

Edited film alignment is the post-production process of finding small parts of unedited footage that temporally and spatially match an edited film. The huge amount of data to be processed makes significant downsampling of the videos essential in real-life applications. Simultaneously, professional u…

Cited by 0SourceScholar
2017

Fast algorithm for statistical phrase/accent command estimation based on generative model incorporating spectral features

ICASSP 2017accepted

An important challenge in speech processing involves extracting non-linguistic information from a fundamental frequency (F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> ) contour of speech. We propose a fast algorithm for estimating the model…

Cited by 0SourceScholar
2017

Generative Attribute Controller With Conditional Filtered Generative Adversarial Networks

CVPR 2017poster

We present a generative attribute controller (GAC), a novel functionality for generating or editing an image while intuitively controlling large variations of an attribute. This controller is based on a novel generative model called the conditional filtered generative adversarial network (CFGAN), wh…

Cited by 114PDFScholar
2017

Generative adversarial network-based postfilter for statistical parametric speech synthesis

ICASSP 2017accepted

We propose a postfilter based on a generative adversarial network (GAN) to compensate for the differences between natural speech and speech synthesized by statistical parametric speech synthesis. In particular, we focus on the differences caused by over-smoothing, which makes the sounds muffled. Ove…

Cited by 0SourceScholar
2016

Scene text recognition with high performance CNN classifier and efficient word inference

ICASSP 2016accepted

The recognition of text in natural scene images is a practical yet challenging task due to the large variations in backgrounds, textures, fonts, and illumination conditions. In this paper, we propose a highly accurate character recognition model by utilizing the representational power of a specially…

Cited by 0SourceScholar
2015

A fast audio search method based on skipping irrelevant signals by similarity upper-bound calculation

ICASSP 2015accepted

In this paper, we describe an approach to accelerate fingerprint techniques by skipping the search for irrelevant sections of the signal and demonstrate its application to the divide and locate (DAL) audio fingerprint method. The search result for the applied method, DAL3, is the same as that of DAL…

Cited by 0SourceScholar