← Search

Takahiro Ogawa

52 accepted papers

2026

ASMIL: Attention-Stabilized Multiple Instance Learning for Whole-Slide Imaging

ICLR 2026poster

Attention-based multiple instance learning (MIL) has emerged as a powerful framework for whole slide image (WSI) diagnosis, leveraging attention to aggregate instance-level features into bag-level predictions. Despite this success, we find that such methods exhibit a new failure mode: unstable atte…

Cited by 0SourcecodeScholar
2026

Adversarial Perturbation Shield: Preventing Concept Bleed-through in Continual Learning of Personalized Generative Models

AAAI 2026technical

Personalized text-to-image diffusion models have gained increasing attention because they can generate images that contain unique concepts based on limited training data. However, in continual learning scenarios, these models suffer from concept bleed-through, where newly introduced concepts frequen

Cited by 0SourcePDFScholar
2026

Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV

AAAI 2026technical

Wide-angle videos in few-shot action recognition (FSAR) effectively express actions within specific scenarios. However, without a global understanding of both subjects and background, recognizing actions in such samples remains challenging because of the background distractions. Receptance Weighted

Cited by 0SourcePDFScholar
2026

Personalized Longitudinal Medical Report Generation via Temporally-Aware Federated Adaptation

CVPR 2026

Automatic medical report generation from multimodal longitudinal imaging is crucial for clinical diagnosis but remains challenging due to privacy constraints and evolving disease dynamics. While federated learning (FL) enables decentralized model training without data sharing, its extension to longi

Cited by 0SourcecodeScholar
2025

Continual Self-supervised Learning Considering Medical Domain Knowledge in Chest CT Images

ICASSP 2025accepted

We propose a novel continual self-supervised learning method (CSSL) considering medical domain knowledge in chest CT images. Our approach addresses the challenge of sequential learning by effectively capturing the relationship between previously learned knowledge and new information at different sta…

Cited by 0SourceScholar
2025

Generative Dataset Distillation Based on Self-knowledge Distillation

ICASSP 2025accepted

Dataset distillation is an effective technique for reducing the cost and complexity of model training while maintaining performance by compressing large datasets into smaller, more efficient versions. In this paper, we present a novel generative dataset distillation method that can improve the accur…

Cited by 0SourceScholar
2025

Gradient-Oriented Clustered Federated Learning With Efficient Knowledge Sharing in Non-IID Settings

ICASSP 2025accepted

We present a novel Clustered Federated Learning (CFL) approach that efficiently shares knowledge among clusters to address non-Independent and Identically Distributed (non-IID) settings. Although conventional CFL has demonstrated strong performance in non-IID settings, a fundamental challenge, the l…

Cited by 0SourceScholar
2025

Hyperboloid GPLVM for Discovering Continuous Hierarchies via Nonparametric Estimation

AISTATS 2025poster

Dimensionality reduction (DR) offers interpretable representations of complex high-dimensional data, and recent DR methods have leveraged hyperbolic geometry to obtain faithful low-dimensional embeddings of high-dimensional hierarchical relationships. However, existing methods are dependent on neigh…

Cited by 0SourceScholar
2025

Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-Sequence

AAAI 2025technical

In few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their application. Recent Mamba demonstrates efficiency in modeling long sequences, but dire…

2025

Robust Adversarial Defense Based on Non-Transferability of Attack Across Foundation Models

ICASSP 2025accepted

This paper presents a novel adversarial defense method that exploits non-transferability of attack across foundation models. Existing adversarial training methods have insufficient robustness to adversarial examples under powerful adversarial attacks in a white-box setting. We clarify that there is…

Cited by 0SourceScholar
2025

Triplet Synthesis for Enhancing Composed Image Retrieval via Counterfactual Image Generation

ICASSP 2025accepted

Composed Image Retrieval (CIR) provides an effective way to manage and access large-scale visual data. Construction of the CIR model utilizes triplets that consist of a reference image, modification text describing desired changes, and a target image that reflects these changes. For effectively trai…

Cited by 0SourceScholar
2024

Caption Unification for Multi-View Lifelogging Images Based on In-Context Learning with Heterogeneous Semantic Contents

ICASSP 2024accepted

This paper presents a new task of caption unification and a novel caption unification method for multi-view lifelogging images based on in-context learning with heterogeneous semantic contents. Most of the existing image captioning models target a single image and do not consider the common semantic…

Cited by 0SourceScholar
2024

Confidence-Aware Spatial-Temporal Attention Graph Convolutional Network for Skeleton-Based Expert-Novice Level Classification

ICASSP 2024accepted

A skeleton-based expert-novice level classification method via Confidence-aware Spatial-Temporal Attention Graph Convolutional Network (ConfSTA-GCN) is presented in this paper. The main contribution of this paper is the realization of the accurate expert-novice level classification by introducing a…

Cited by 0SourceScholar
2024

Enhancing Noisy Label Learning Via Unsupervised Contrastive Loss with Label Correction Based on Prior Knowledge

ICASSP 2024accepted

To alleviate the negative impacts of noisy labels, most of the noisy label learning (NLL) methods dynamically divide the training data into two types, "clean samples" and "noisy samples", in the training process. However, the conventional selection of clean samples heavily depends on the features le…

Cited by 0SourceScholar
2024

Multi-Object Editing in Personalized Text-To-Image Diffusion Model Via Segmentation Guidance

ICASSP 2024accepted

This paper presents a personalized text-to-image diffusion model for multiple object editing that can improve visual fidelity of the target image and editing ability with a segmentation-based restriction and continual learning. Multiple personalization tasks face the problem of destabilization, espe…

Cited by 0SourceScholar
2024

Privacy Preserving Gaze Estimation Via Federated Learning Adapted To Egocentric Video

ICASSP 2024accepted

This paper presents privacy preserving gaze estimation via federated learning adapted to egocentric videos. Gaze estimation with egocentric video stands out among many applications of wearable cameras and is closely related to many high-tech applications envisioned in the future. However, traditiona…

Cited by 0SourceScholar
2024

Prompt-Based Personalized Federated Learning for Medical Visual Question Answering

ICASSP 2024accepted

We present a novel prompt-based personalized federated learning (pFL) method to address data heterogeneity and privacy concerns in traditional medical visual question answering (VQA) methods. Specifically, we regard medical datasets from different organs as clients and use pFL to train personalized…

Cited by 0SourceScholar
2023

Binauralization Robust To Camera Rotation Using 360° Videos

ICASSP 2023accepted

We propose a novel binauralization method that is robust to camera rotation. Since binaural audio can bring a 3D sensation to the listener, it can enhance the immersive experience of the video. Researchers have been explored binaural audio generation from monaural audio to deepen the experience of a…

Cited by 0SourceScholar
2023

Defense Against Black-Box Adversarial Attacks Via Heterogeneous Fusion Features

ICASSP 2023accepted

This paper presents an effective approach for the adversarial defense task named a heterogeneous feature fusion network (HFFN). Inspired by the fact that humans can utilize multimodal information to help themselves perceive objects, we introduce the caption features into the classic convolutional ne…

Cited by 0SourceScholar
2023

Estimation of Visual Contents from Human Brain Signals via VQA Based on Brain-Specific Attention

ICASSP 2023accepted

This paper presents a method for estimation of visual cognitive contents from human brain signals via a newly derived visual question answering (VQA) model. The proposed method can estimate a wide range of cognitive contents from functional magnetic resonance imaging data when subjects viewed images…

Cited by 0SourceScholar
2023

Improving Dropout in Graph Convolutional Networks for Recommendation via Contrastive Loss

ICASSP 2023accepted

We propose a novel graph convolutional network (GCN)-based recommendation model that incorporates a contrastive loss. Although GCN-based recommendation models achieve high recommendation performance, existing models suffer from over-fitting since they explicitly encode the interactions as a graph. T…

Cited by 0SourceScholar
2023

Learning Graph Laplacian from Intrinsic Patterns via Gaussian Process

ICASSP 2023accepted

In this paper, we present a novel scheme to learn a graph topology named Laplacian constrained Gaussian process (LCGP). Previous graph learning methods directly use raw observed signals, which often degrades estimation accuracy due to the noise effects and the sample imbalances. LCGP tackles this pr…

Cited by 0SourceScholar
2022

Distributed Label Dequantized Gaussian Process Latent Variable Model for Multi-View Data Integration

ICASSP 2022accepted

In this paper, we present a novel method for multi-view data analysis, distributed label dequantized Gaussian process latent variable model (DLDGP). DLDGP can integrate multi-view data and class information into a common latent space. In the previous multiview methods, the dimension of label feature…

Cited by 0SourceScholar
2022

Divergence-Guided Feature Alignment for Cross-Domain Object Detection

ICASSP 2022accepted

Domain shift causes performance drop in cross-domain object detection. To alleviate the domain shift, a prevailing approach is global feature alignment with adversarial learning. However, such simple feature alignment has defects of unawareness of fore-ground/background regions and well-aligned/poor…

Cited by 0SourceScholar
2022

Generative Adversarial Network Including Referring Image Segmentation For Text-Guided Image Manipulation

ICASSP 2022accepted

This paper proposes a novel generative adversarial network to improve the performance of image manipulation using natural language descriptions that contain desired attributes. Text-guided image manipulation aims to semantically manipulate an image aligned with the text description while preserving…

Cited by 0SourceScholar
2022

Human Emotion Recognition Using Multi-Modal Biological Signals Based On Time Lag-Considered Correlation Maximization

ICASSP 2022accepted

A human emotion recognition using multi-modal biological signals based on time lag-considered correlation maximization is presented in this paper. Various multi-modal emotion recognition methods for visual stimuli have been studied and they focus on gaze and brain activity data. The visual stimuli c…

Cited by 0SourceScholar
2022

Self-Knowledge Distillation based Self-Supervised Learning for Covid-19 Detection from Chest X-Ray Images

ICASSP 2022accepted

The global outbreak of the Coronavirus 2019 (COVID-19) has overloaded worldwide healthcare systems. Computer-aided diagnosis for COVID-19 fast detection and patient triage is becoming critical. This paper proposes a novel self-knowledge distillation based self-supervised learning method for COVID-19…

Cited by 0SourceScholar
2022

Union-Set Multi-source Model Adaptation for Semantic Segmentation

ECCV 2022poster

"This paper solves a generalized version of the problem of multi-source model adaptation for semantic segmentation. Model adaptation is proposed as a new domain adaptation problem which requires access to a pre-trained model instead of data for the source domain. A general multi-source setting of mo…

2022

Variational Bayesian Graph Convolutional Network for Robust Collaborative Filtering

ICASSP 2022accepted

This paper presents a variational Bayesian graph convolutional network for robust collaborative filtering (VBGCF). Conventional graph convolutional network (GCN)-based recommendation models fully trust the observed interaction graph. However, the data used in real-world applications (e.g., video str…

Cited by 0SourceScholar
2021

Classification of Expert-Novice Level Using Eye Tracking And Motion Data via Conditional Multimodal Variational Autoencoder

ICASSP 2021accepted

Sensor data from wearable devices have been utilized to analyze differences between experts and novices. Previous studies attempted to classify the expert-novice level from sensor data based on supervised learning methods. However, these approaches need to collect enough training data covering vario…

Cited by 0SourceScholar
2021

Cross-Domain Semi-Supervised Deep Metric Learning for Image Sentiment Analysis

ICASSP 2021accepted

This paper presents a novel method on image sentiment analysis called cross-domain semi-supervised deep metric learning (CDSS-DML). The proposed method has two contributions. Firstly, since previous researches on image sentiment analysis suffer from the limit of a small amount of well-labeled data,…

Cited by 0SourceScholar
2021

Estimation of Visual Features of Viewed Image From Individual and Shared Brain Information Based on FMRI Data Using Probabilistic Generative Model

ICASSP 2021accepted

This paper presents a method for estimation of visual features based on brain responses measured when subjects view images. The pro-posed method estimates visual features of viewed images by using both individual and shared brain information from functional magnetic resonance imaging (fMRI) data whe…

Cited by 0SourceScholar
2021

Feature Integration via Semi-Supervised Ordinally Multi-Modal Gaussian Process Latent Variable Model

ICASSP 2021accepted

This paper presents a method of feature integration via semi-supervised ordinally multi-modal Gaussian process latent variable model (Semi-OMGP). The proposed method transforms multimodal features into common latent variables suitable for users' interest level estimation. For dealing with the multi-…

Cited by 0SourceScholar
2021

Human-Centered Favorite Music Classification Using EEG-Based Individual Music Preference Via Deep Time-Series CCA

ICASSP 2021accepted

A method to classify a user’s like or dislike musical pieces based on the extraction of his or her music preference is proposed in this paper. New scheme of Canonical Correlation Analysis (CCA), called Deep Time-series CCA (DTCCA), which can consider the correlation between two sets of input feature…

Cited by 0SourceScholar
2021

Multi-Modal Label Dequantized Gaussian Process Latent Variable Model for Ordinal Label Estimation

ICASSP 2021accepted

This paper presents multi-modal label dequantized Gaussian process latent variable model (mLDGP) for ordinal label estimation. mLDGP is constructed based on a probabilistic generative model via Gaussian process and realizes accurate calculation of common latent space from multi-view features includi…

Cited by 0SourceScholar
2021

Semantic-Aware Unpaired Image-to-Image Translation for Urban Scene Images

ICASSP 2021accepted

Unpaired image-to-image (I2I) translation methods have been developed for several years. Present methods do not take into consideration semantic information of the original image, which may perform well on simple datasets of uncomplicated scenes, however, fail in complex datasets of scenes involving…

Cited by 0SourceScholar
2020

Multi-View Bayesian Generative Model for Multi-Subject FMRI Data on Brain Decoding of Viewed Image Categories

ICASSP 2020accepted

Brain decoding studies have demonstrated that viewed image categories can be estimated from human functional magnetic resonance imaging (fMRI) activity. However, there are still limitations with the estimation performance because of the characteristics of fMRI data and the employment of only one mod…

Cited by 0SourceScholar
2020

Unsupervised Domain Adaptation for Semantic Segmentation with Symmetric Adaptation Consistency

ICASSP 2020accepted

Unsupervised domain adaptation, which leverages label information from other domains to solve tasks on a domain without any labels, can alleviate the problem of the scarcity of labels and expensive labeling costs faced by supervised semantic segmentation. In this paper, we utilize adversarial learni…

Cited by 0SourceScholar
2019

Estimating Viewed Image Categories from Human Brain Activity via Semi-supervised Fuzzy Discriminative Canonical Correlation Analysis

ICASSP 2019accepted

This paper presents a method to estimate viewed image categories from human brain activity via newly derived semi-supervised fuzzy discriminative canonical correlation analysis (Semi-FDCCA). The proposed method can estimate image categories from functional magnetic resonance imaging (fMRI) activity…

Cited by 0SourceScholar
2019

Multi-feature Fusion Based on Supervised Multi-view Multi-label Canonical Correlation Projection

ICASSP 2019accepted

This paper presents multi-feature fusion based on supervised multi-view multi-label canonical correlation projection (sM2CP). The proposed method applies sM2CP-based feature fusion to multiple features obtained from various convolutional neural networks (CNNs) whose characteristics are different. Si…

Cited by 0SourceScholar
2018

Sfemcca: Supervised Fractional-Order Embedding Multiview Canonical Correlation Analysis for Video Preference Estimation

ICASSP 2018accepted

In this paper, we present supervised fractional-order embedding multiview canonical correlation analysis (SFEMCCA). SFEMCCA is a CCA method realizing the following three points: (1) learning noisy data with small number of samples and large number of dimensions, (2) multiview learning that can integ…

Cited by 0SourceScholar
2017

Emotion estimation via tensor-based supervised decision-level fusion from multiple Brodmann areas

ICASSP 2017accepted

This paper presents a novel method that estimates human emotion based on tensor-based supervised decision-level fusion (TS-DLF) from multiple Brodmann areas (BAs). From multiple brain data corresponding to these BAs captured by functional magnetic resonance imaging (fMRI), our method performs genera…

Cited by 0SourceScholar
2017

Exemplar-based image completion via new quality measure based on phaseless texture features

ICASSP 2017accepted

This paper presents an exemplar-based image completion via a new quality measure based on phaseless texture features. The proposed method derives a new quality measure obtained by monitoring errors caused in power spectra, i.e., errors of phaseless texture features, converged through phase retrieval…

Cited by 0SourceScholar
2017

Personalized video preference estimation based on early fusion using multiple users' viewing behavior

ICASSP 2017accepted

This paper presents a novel method for personalized video preference estimation based on early fusion using multiple users' viewing behavior. The proposed method adopts supervised Multi-View Canonical Correlation Analysis (sMVCCA) to estimate correlation between different types of features. Specific…

Cited by 0SourceScholar
2016

Novel favorite music classification using EEG-based optimal audio features selected via KDLPCCA

ICASSP 2016accepted

This paper presents a novel method of favorite music classification using EEG-based optimal audio features. To select audio features related to user's music preference, our method utilizes a relationship between EEG features obtained from the user's EEG signals during listening to music and their co…

Cited by 0SourceScholar
2015

Missing intensity restoration via adaptive selection of perceptually optimized subspaces

ICASSP 2015accepted

A missing intensity restoration method via adaptive selection of perceptually optimized subspaces is presented in this paper. In order to realize adaptive and perceptually optimized restoration, the proposed method generates several subspaces of known textures optimized in terms of the structural si…

Cited by 0SourceScholar
2015

Novel image classification based on integration of EEG and visual features via MSLPCCA

ICASSP 2015accepted

This paper presents a novel image classification method based on integration of EEG and visual features. In the proposed method, we obtain classification results by separately using EEG and visual features. Furthermore, we merge the above classification results based on a kernelized version of Super…

Cited by 0SourceScholar