← Search

Naoki Makishima

7 accepted papers

2026

Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models

AAAI 2026technical

Contrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image and text encoders. This paper aims to robustly fine-tune these vision-language models on in-distribution (ID) data witho

Cited by 0SourcePDFScholar
2025

Multimodal Fine-Grained Apparent Personality Trait Recognition: Joint Modeling of Big Five and Questionnaire Item-level Scores

AAAI 2025technical

This paper presents a novel method for automatically recognizing people's apparent personality traits as perceived by others. In previous studies, apparent personality trait recognition from multimodal human behavior is often modeled to directly estimate personality trait scores, i.e., the ``Big Fiv…

Cited by 0SourcePDFScholar
2023

Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness Tradeoff

ICCV 2023poster

This paper addresses the tradeoff between standard accuracy on clean examples and robustness against adversarial examples in deep neural networks (DNNs). Although adversarial training (AT) improves robustness, it degrades the standard accuracy, thus yielding the tradeoff. To mitigate this tradeoff…

Cited by 7PDFScholar
2022

Customer Satisfaction Estimation Using Unsupervised Representation Learning with Multi-Format Prediction Loss

ICASSP 2022accepted

We propose a new Customer Satisfaction Estimation (CSE) method that utilizes unsupervised representation learning. Though conventional methods have improved both the heuristic features and the estimation models, their performance is still insufficient as only small amounts of labeled training data c…

Cited by 0SourceScholar
2021

Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss

ICASSP 2021accepted

We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual speech separation is a technique to estimate the individual speech signals from a mi…

Cited by 0SourceScholar
2021

Hierarchical Transformer-Based Large-Context End-To-End ASR with Large-Context Knowledge Distillation

ICASSP 2021accepted

We present a novel large-context end-to-end automatic speech recognition (E2E-ASR) model and its effective training method based on knowledge distillation. Common E2E-ASR models have mainly focused on utterance-level processing in which each utterance is independently transcribed. On the other hand,…

Cited by 0SourceScholar
2021

MAPGN: Masked Pointer-Generator Network for Sequence-to-Sequence Pre-Training

ICASSP 2021accepted

This paper presents a self-supervised learning method for pointer-generator networks to improve spoken-text normalization. Spoken-text normalization that converts spoken-style text into style normalized text is becoming an important technology for improving subsequent processing such as machine tran…

Cited by 0SourceScholar