← Search

Shujian YU

34 accepted papers

2026

ADAPTING HFMCA TO GRAPH DATA: SELF-SUPERVISED LEARNING FOR GENERALIZABLE FMRI REPRESENTATIONS

ICASSP 2026poster

Functional magnetic resonance imaging (fMRI) analysis faces significant challenges due to limited dataset sizes and domain variability between studies. Traditional self-supervised learning methods inspired by computer vision often rely on positive and negative sample pairs, which can be problematic…

Cited by 0SourcePDFScholar
2026

Continual Learning for fMRI-Based Brain Disorder Diagnosis via Functional Connectivity Matrices Generative Replay

CVPR 2026

Functional magnetic resonance imaging (fMRI) is widely used for studying and diagnosing brain disorders, with functional connectivity (FC) matrices providing powerful representations of large-scale neural interactions. However, existing diagnostic models are trained either on a single site or under

Cited by 0SourcecodeScholar
2026

Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence

ICLR 2026poster

Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairwise samples across modalities while overlooking distributional differences.…

Cited by 0SourceScholar
2026

HFMCA: ORTHONORMAL FEATURE LEARNING FOR EEG-BASED BRAIN DECODING

ICASSP 2026poster

Electroencephalography (EEG) analysis is critical for brain-computer interfaces and neuroscience, but the intrinsic noise and high dimensionality of EEG signals hinder effective feature learning. We propose a self-supervised framework based on the Hierarchical Functional Maximal Correlation Algorith…

Cited by 0SourcePDFScholar
2026

Towards Uniformity and Alignment for Multimodal Representation Learning

ICML 2026poster

Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-based objectives introduce inherent conflicts that yield distribution gaps across modalities. In this work, we identify …

Cited by 0SourceScholar
2025

Aggregation of Dependent Expert Distributions in Multimodal Variational Autoencoders

ICML 2025poster

Multimodal learning with variational autoencoders (VAEs) requires estimating joint distributions to evaluate the evidence lower bound (ELBO). Current methods, the product and mixture of experts, aggregate single-modality distributions assuming independence for simplicity, which is an overoptimistic…

Cited by 0SourcePDFScholar
2025

Start Smart: Leveraging Gradients For Enhancing Mask-based XAI Methods

ICLR 2025poster

Mask-based explanation methods offer a powerful framework for interpreting deep learning model predictions across diverse data modalities, such as images and time series, in which the central idea is to identify an instance-dependent mask that minimizes the performance drop from the resulting masked…

Cited by 0SourcePDFScholar
2024

BAN: Detecting Backdoors Activated by Adversarial Neuron Noise

NeurIPS 2024poster

Backdoor attacks on deep learning represent a recent threat that has gained significant attention in the research community. Backdoor defenses are mainly based on backdoor inversion, which has been shown to be generic, model-agnostic, and applicable to practical threat scenarios. State-of-the-art b…

2024

Cauchy-Schwarz Divergence Information Bottleneck for Regression

ICLR 2024poster

The information bottleneck (IB) approach is popular to improve the generalization, robustness and explainability of deep neural networks. Essentially, it aims to find a minimum sufficient representation $\mathbf{t}$ by striking a trade-off between a compression term $I(\mathbf{x};\mathbf{t})$ and a…

2024

DIB-X: Formulating Explainability Principles for a Self-Explainable Model Through Information Theoretic Learning

ICASSP 2024accepted

The recent development of self-explainable deep learning approaches has focused on integrating well-defined explainability principles into learning process, with the goal of achieving these principles through optimization. In this work, we propose DIB-X, a self-explainable deep learning approach for…

Cited by 0SourceScholar
2024

Domain Adaptation with Cauchy-Schwarz Divergence

UAI 2024poster

Domain adaptation aims to use training data from one or multiple source domains to learn a hypothesis that can be generalized to a different, but related, target domain. As such, having a reliable measure for evaluating the discrepancy of both marginal and conditional distributions is crucial. We in…

2024

Jacobian Regularizer-based Neural Granger Causality

ICML 2024poster

With the advancement of neural networks, diverse methods for neural Granger causality have emerged, which demonstrate proficiency in handling complex data, and nonlinear relationships. However, the existing framework of neural Granger causality has several limitations. It requires the construction o…

2024

Rethinking Information-theoretic Generalization: Loss Entropy Induced PAC Bounds

ICLR 2024poster

Information-theoretic generalization analysis has achieved astonishing success in characterizing the generalization capabilities of noisy and iterative learning algorithms. However, current advancements are mostly restricted to average-case scenarios and necessitate the stringent bounded loss assump…

Cited by 2SourcePDFScholar
2023

Causal Recurrent Variational Autoencoder for Medical Time Series Generation

AAAI 2023technical

We propose causal recurrent variational autoencoder (CR-VAE), a novel generative model that is able to learn a Granger causal graph from a multivariate time series x and incorporates the underlying causal mechanism into its data generation process. Distinct to the classical recurrent VAEs, our CR-VA…

2023

Robust and Fast Measure of Information via Low-Rank Representation

AAAI 2023technical

The matrix-based Rényi's entropy allows us to directly quantify information measures from given data, without explicit estimation of the underlying probability distribution. This intriguing property makes it widely applied in statistical inference and machine learning tasks. However, this informatio…

2023

Towards a More Stable and General Subgraph Information Bottleneck

ICASSP 2023accepted

Graph Neural Networks (GNNs) have been widely applied to graph-structured data. However, the lack of interpretability impedes its practical deployment especially in high-risk areas such as medical diagnosis. Recently, the Information Bottleneck (IB) principle has been extended to GNNs to identify a…

Cited by 0SourceScholar
2022

Deep Deterministic Independent Component Analysis for Hyperspectral Unmixing

ICASSP 2022accepted

We develop a new neural network based independent component analysis (ICA) method by directly minimizing the dependence amongst all extracted components. Using the matrix-based Rényi’s α-order entropy functional, our network can be directly optimized by stochastic gradient descent (SGD), without any…

Cited by 0SourceScholar
2022

Principle of relevant information for graph sparsification

UAI 2022poster

Graph sparsification aims to reduce the number of edges of a graph while maintaining its structural properties. In this paper, we propose the first general and effective information-theoretic formulation of graph sparsification, by taking inspiration from the Principle of Relevant Information (PRI).…

2021

Deep Deterministic Information Bottleneck with Matrix-Based Entropy Functional

ICASSP 2021accepted

We introduce the matrix-based Rényi’s α-order entropy functional to parameterize Tishby et al. information bottleneck (IB) principle [1] with a neural network. We term our methodology Deep Deterministic Information Bottleneck (DIB), as it avoids variational inference and distribution assumption. We…

Cited by 0SourceScholar
2021

Information-Theoretic Methods in Deep Neural Networks: Recent Advances and Emerging Opportunities

IJCAI 2021poster

We present a review on the recent advances and emerging opportunities around the theme of analyzing deep neural networks (DNNs) with information-theoretic methods. We first discuss popular information-theoretic quantities and their estimators. We then introduce recent developments on information-the…

Cited by 20SourcePDFScholar
2021

Measuring Dependence with Matrix-based Entropy Functional

AAAI 2021technical

Measuring the dependence of data plays a central role in statistics and machine learning. In this work, we summarize and generalize the main idea of existing information-theoretic dependence measures into a higher-level perspective by the Shearer's inequality. Based on our generalization, we then pr…

2020

Composite Dynamic Texture Synthesis Using Hierarchical Linear Dynamical System

ICASSP 2020accepted

We demonstrate that a systematic inclusion of prior structural constraints on the states of a linear dynamical system significantly improves its ability to model complex multidimensional sequences. This constrained LDS, typically termed as the hierarchical linear dynamical system (HLDS), is a Kalman…

Cited by 0SourceScholar
2020

Measuring the Discrepancy between Conditional Distributions: Methods, Properties and Applications

IJCAI 2020poster

We propose a simple yet powerful test statistic to quantify the discrepancy between two conditional distributions. The new statistic avoids the explicit estimation of the underlying distributions in high-dimensional space and it operates on the cone of symmetric positive semidefinite (SPS) matrix usi…

2017

Autoencoders trained with relevant information: Blending Shannon and Wiener's perspectives

ICASSP 2017accepted

It is almost seventy years after the publication of Claude Shannon's “A Mathematical Theory of Communication” [1] and Norbert Wiener's “Extrapolation, Interpolation and Smoothing of Stationary Time Series” [2]. The pioneering works of Shannon and Wiener lay the foundation of communication, data stor…

Cited by 0SourceScholar
2017

Robust linear discriminant analysis with a Laplacian assumption on projection distribution

ICASSP 2017accepted

Linear discriminant analysis (LDA) is typically carried out using Fisher's method, which relies heavily on the estimation of sample mean vectors and covariance matrices. However, Fisher LDA is vulnerable to outliers as it happens to other multivariate statistical methods. In this paper, we analyzed…

Cited by 0SourceScholar