← Search

Nakamasa Inoue

38 accepted papers

2026

ANIMALCLAP: TAXONOMY-AWARE LANGUAGE-AUDIO PRETRAINING FOR SPECIES RECOGNITION AND TRAIT INFERENCE

ICASSP 2026poster

Animal vocalizations provide crucial insights for wildlife assessment, particularly in complex environments such as forests, aiding species identification and ecological monitoring. Recent advances in deep learning have enabled automatic species classification from their vocalizations. However, clas…

Cited by 0SourcePDFScholar
2026

BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment

CVPR 2026

Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology.While recent biological models, such as BioCLIP, have demonstrated strong alignment between images and textual taxonomic information for species identification, the integr

Cited by 0SourcecodeScholar
2026

DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning

AAAI 2026technical

Large vision-language models (LVLMs) have shown impressive performance across a broad range of multimodal tasks. However, robust image caption evaluation using LVLMs remains challenging, particularly under domain-shift scenarios. To address this issue, we introduce the Distribution-Aware Score Decod

Cited by 0SourcePDFScholar
2026

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models

ICML 2026poster

While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, which requires integration of information from multiple viewpoints, remains challenging. Cognitive studies suggest that humans address such tasks through tw…

Cited by 0SourceScholar
2026

PowerCLIP: Powerset Alignment for Contrastive Pre-Training

CVPR 2026

Contrastive pre-training frameworks such as CLIP have demonstrated impressive zero-shot performance across a range of vision-language tasks. Recent studies have shown that aligning individual text tokens with specific image patches or regions enhances fine-grained compositional understanding. Howeve

Cited by 0SourcecodeScholar
2025

AgroBench: Vision-Language Model Benchmark in Agriculture

ICCV 2025poster

Precise automated understanding of agricultural tasks such as disease identification is essential for the sustainable crop production. Recent advances in vision-language models (VLMs) are expected to further expand the range of agricultural tasks by facilitating human-model interaction through easy,…

2025

AnimalClue: Recognizing Animals by their Traces

ICCV 2025poster

Wildlife observation plays an important role in biodiversity conservation, necessitating robust methodologies for monitoring wildlife populations and interspecies interactions. Recent advances in computer vision have significantly contributed to automating fundamental wildlife observation tasks, suc…

2025

Binary Stochastic Flip Optimization for Training Binary Neural Networks

ICASSP 2025accepted

For deploying deep neural networks on edge devices with limited resources, binary neural networks (BNNs) have attracted significant attention, due to their computational and memory efficiency. However, once a neural network is binarized, finetuning it on edge devices becomes challenging because most…

Cited by 0SourceScholar
2025

CityNav: A Large-Scale Dataset for Real-World Aerial Navigation

ICCV 2025poster

Vision-and-language navigation (VLN) aims to develop agents capable of navigating in realistic environments. While recent cross-modal training approaches have significantly improved navigation performance in both indoor and outdoor scenarios, aerial navigation over real-world cities remains underexp…

Cited by 0SourcePDFScholar
2025

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

ICCV 2025poster

The advancement of 3D language fields has enabled intuitive interactions with 3D scenes via natural language. However, existing approaches are typically limited to small-scale environments, lacking the scalability and compositional reasoning capabilities necessary for large, complex urban settings.…

Cited by 0SourcePDFScholar
2025

HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

ICLR 2025poster

Recently, Text-to-speech (TTS) models based on large language models (LLMs) that translate natural language text into sequences of discrete audio tokens have gained great research attention, with advances in neural audio codec (NAC) mod- els using residual vector quantization (RVQ). However, long-fo…

2025

Multi-Point Positional Insertion Tuning for Small Object Detection

ICASSP 2025accepted

Small object detection aims to localize and classify small objects within images. With recent advances in large-scale vision-language pretraining, finetuning pretrained object detection models has emerged as a promising approach. However, finetuning large models is computationally and memory expensi…

Cited by 0SourceScholar
2025

Rectified Lagrangian for Out-of-Distribution Detection in Modern Hopfield Networks

AAAI 2025technical

Modern Hopfield networks (MHNs) have recently gained significant attention in the field of artificial intelligence because they can store and retrieve a large set of patterns with an exponentially large memory capacity. A MHN is generally a dynamical system defined with Lagrangians of memory and fea…

Cited by 0SourcePDFScholar
2025

Referring Expression Comprehension for Small Objects

ICCV 2025poster

Referring expression comprehension (REC) aims to localize the target object described by a natural language expression.Recent advances in vision-language learning have led to significant performance improvements in REC tasks.However, localizing extremely small objects remains a considerable challeng…

2024

Cubic Knowledge Distillation for Speech Emotion Recognition

ICASSP 2024accepted

Speech Emotion Recognition (SER) can play an important role in human-computer interaction. In this paper, we propose a logit knowledge distillation method for SER, called Cubic KD, that distill the knowledge of fine-tuned self-supervised models to allow better performance of small models. By creatin…

Cited by 0SourceScholar
2024

Efficient Target Propagation by Deriving Analytical Solution

AAAI 2024technical

Exploring biologically plausible algorithms as alternatives to error backpropagation (BP) is a challenging research topic in artificial intelligence. It also provides insights into the brain's learning methods. Recently, when combined with well-designed feedback loss functions such as Local Differen…

Cited by 0SourcePDFScholar
2024

Formula-Supervised Visual-Geometric Pre-training

ECCV 2024poster

"Throughout the history of computer vision, while research has explored the integration of images (visual) and point clouds (geometric), many advancements in image and 3D object recognition have tended to process these modalities separately. We aim to bridge this divide by integrating images and poi…

2024

PolarDB: Formula-Driven Dataset for Pre-Training Trajectory Encoders

ICASSP 2024accepted

Formula-driven supervised learning (FDSL) is a growing research topic for finding simple mathematical formulas that generate synthetic data and labels for pre-training neural networks. The main advantage of FDSL is that there is no risk of generating data with ethical implications such as gender bia…

Cited by 0SourceScholar
2024

Pseudo-Outlier Synthesis Using Q-Gaussian Distributions for Out-of-Distribution Detection

ICASSP 2024accepted

Out-of-distribution (OOD) detection, which aims to determine whether an input is outside the training data distribution or not, is an indispensable task in many computer vision applications. In many of the previous studies on OOD detection for image classification, the class-conditional distribution…

Cited by 0SourceScholar
2024

Rethinking Image Super Resolution from Training Data Perspectives

ECCV 2024poster

"In this work, we investigate the understudied effect of the training data used for image super-resolution (SR). Most commonly, novel SR methods are developed and benchmarked on common training datasets such as DIV2K and DF2K. However, we investigate and rethink the training data from the perspectiv…

2023

CityRefer: Geography-aware 3D Visual Grounding Dataset on City-scale Point Cloud Data

NeurIPS 2023poster

City-scale 3D point cloud is a promising way to express detailed and complicated outdoor structures. It encompasses both the appearance and geometry features of segmented city components, including cars, streets, and buildings that can be utilized for attractive applications such as user-interactive…

2023

Fixed-Weight Difference Target Propagation

AAAI 2023technical

Target Propagation (TP) is a biologically more plausible algorithm than the error backpropagation (BP) to train deep networks, and improving practicality of TP is an open issue. TP methods require the feedforward and feedback networks to form layer-wise autoencoders for propagating the target value…

2023

Learning with Partial Forgetting in Modern Hopfield Networks

AISTATS 2023poster

It has been known by neuroscience studies that partial and transient forgetting of memory often plays an important role in the brain to improve performance for certain intellectual activities. In machine learning, associative memory models such as classical and modern Hopfield networks have been pro…

2023

Parameter Efficient Transfer Learning for Various Speech Processing Tasks

ICASSP 2023accepted

Fine-tuning of self-supervised models is a powerful transfer learning method in a variety of fields, including speech processing, since it can utilize generic feature representations obtained from large amounts of unlabeled data. Fine-tuning, however, requires a new parameter set for each downstream…

Cited by 0SourceScholar
2023

Pre-training Vision Transformers with Very Limited Synthesized Images

ICCV 2023poster

Formula-driven supervised learning (FDSL) is a pre-training method that relies on synthetic images generated from mathematical formulae such as fractals. Prior work on FDSL has shown that pre-training vision transformers on such synthetic datasets can yield competitive accuracy on a wide range of d…

Cited by 10PDFcodeScholar
2023

SegRCDB: Semantic Segmentation via Formula-Driven Supervised Learning

ICCV 2023poster

Pre-training is a strong strategy for enhancing visual models to efficiently train them with a limited number of labeled images. In semantic segmentation, creating annotation masks requires an intensive amount of labor and time, and therefore, a large-scale pre-training dataset with semantic labels…

Cited by 13PDFcodeScholar
2023

Visual Atoms: Pre-Training Vision Transformers With Sinusoidal Waves

CVPR 2023poster

Formula-driven supervised learning (FDSL) has been shown to be an effective method for pre-training vision transformers, where ExFractalDB-21k was shown to exceed the pre-training effect of ImageNet-21k. These studies also indicate that contours mattered more than textures when pre-training vision t…

Cited by 27SourcePDFScholar
2022

Can Vision Transformers Learn without Natural Images?

AAAI 2022technical

Is it possible to complete Vision Transformer (ViT) pre-training without natural images and human-annotated labels? This question has become increasingly relevant in recent months because while current ViT pre-training tends to rely heavily on a large number of natural images and human-annotated lab…

Cited by 38SourcePDFScholar
2022

PoF: Post-Training of Feature Extractor for Improving Generalization

ICML 2022spotlight

It has been intensively investigated that the local shape, especially flatness, of the loss landscape near a minimum plays an important role for generalization of deep models. We developed a training algorithm called PoF: Post-Training of Feature Extractor that updates the feature extractor part of…

2022

Replacing Labeled Real-Image Datasets With Auto-Generated Contours

CVPR 2022poster

In the present work, we show that the performance of formula-driven supervised learning (FDSL) can match or even exceed that of ImageNet-21k without the use of real images, human-, and self-supervision during the pre-training of Vision Transformers (ViTs). For example, ViT-Base pre-trained on ImageN…

Cited by 44PDFScholar
2019

Sequence-level Knowledge Distillation for Model Compression of Attention-based Sequence-to-sequence Speech Recognition

ICASSP 2019accepted

We investigate the feasibility of sequence-level knowledge distillation of Sequence-to-Sequence (Seq2Seq) models for Large Vocabulary Continuous Speech Recognition (LVCSR). We first use a pre-trained larger teacher model to generate multiple hypotheses per utterance with beam search. With the same i…

Cited by 0SourceScholar