← Search

Shuigeng Zhou

47 accepted papers

2026

Conditional Independent Component Analysis For Estimating Causal Structure with Latent Variables

ICLR 2026poster

Identifying latent variables and their induced causal structure is fundamental in various scientific fields. Existing approaches often rely on restrictive structural assumptions (e.g., purity) and may become invalid when these assumptions are violated. We introduce Conditional Independent Component…

Cited by 0SourceScholar
2026

FedOpenMatch: Towards Semi-Supervised Federated Learning in Open-Set Environments

ICLR 2026poster

Semi-supervised federated learning (SSFL) has emerged as an effective approach to leverage unlabeled data distributed across multiple data owners for improving model generalization. Existing SSFL methods typically assume that labeled and unlabeled data share the same label space. However, in realist…

Cited by 0SourcecodeScholar
2026

GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation

AAAI 2026technical

Existing state-of-the-art image tokenization methods leverage diverse semantic features from pre-trained vision models for additional supervision, to expand the distribution of latent representations and thereby improve the quality of image reconstruction and generation. These methods employ a local

Cited by 0SourcePDFScholar
2026

GoR: A Unified and Extensible Generative Framework for Ordinal Regression

ICLR 2026poster

Ordinal Regression (OR), which predicts the target values with inherent order, underpins a wide spectrum of applications from computer vision to recommendation systems. The intrinsic ordinal structure and non-stationary inter-class boundaries make OR fundamentally more challenging than conventional…

Cited by 0SourceScholar
2026

ImmerIris: A Large-Scale Dataset and Benchmark for Off-Axis and Unconstrained Iris Recognition in Immersive Applications

CVPR 2026

Recently, iris recognition is regaining prominence in immersive applications such as extended reality as a means of seamless user identification. This application scenario introduces unique challenges compared to traditional iris recognition under controlled setups, as the ocular images are primaril

Cited by 0SourceScholar
2026

Invariant Feature Learning for Counterfactual Watch-time Prediction in Video Recommendation

AAAI 2026technical

Video recommendation systems heavily rely on user watch time feedback, making accurate watch time prediction a crucial task. However, this task inherently suffers from bias, as recommendation models tend to favor long-duration videos to maximize watch time. This issue, known as duration bias in the

Cited by 0SourcePDFScholar
2026

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource

ICLR 2026oral

Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable performance without increasing per-token compute. However, can MoEs surpass dense architectures under strictly equal resource constraints — that is, when the total parameter count, training compute, an…

Cited by 0SourceScholar
2026

Powerful and Theoretically Guaranteed Independence Testing on Heterogeneous Federated Clients

ICML 2026poster

In this paper, we present a novel federated independence testing method that addresses both theoretical and practical challenges arising from client heterogeneity. We begin by revisiting existing federated independence testing methods and showing why they fail to provide valid guarantees or maintain…

Cited by 0SourceScholar
2026

Streaming Covariate Balancing via Discrepancy-Based Feature Coresets

ICML 2026poster

Real-time estimation of average treatment effects (ATE) in streaming observational data poses two key challenges: strict memory constraints that preclude storing the full data history, and distributional shifts in both treatment assignment and outcome-generating process. Existing methods either requ…

Cited by 0SourceScholar
2026

Towards Policy-Adaptive Image Guardrail: Benchmark and Method

CVPR 2026

Accurate rejection of sensitive or harmful visual content, i.e., harmful image guardrail, is critical in many application scenarios. This task must continuously adapt to the evolving safety policies and content across various domains and over time. However, traditional classifiers, confined to fixed

Cited by 0SourceScholar
2025

A New Model for Prototype-based Continual Learning in Hyperspherical Space

ICASSP 2025accepted

The continuous emergence of new objects in the visual world poses a serious challenge to deep object recognition methods, which sparks the increasing study on continual or incremental learning. However, learning new tasks faces the tough catastrophic forgetting problem, i.e., dramatic performance de…

Cited by 0SourceScholar
2025

Data Synthesis with Diverse Styles for Face Recognition via 3DMM-Guided Diffusion

CVPR 2025poster

Identity-preserving face synthesis aims to generate synthetic face images of virtual subjects that can substitute real-world data for training face recognition models. While prior arts strive to create images with consistent identities and diverse styles, they face a trade-off between them. Identify…

2025

Effective Cloud Removal for Remote Sensing Images by an Improved Mean-Reverting Denoising Model with Elucidated Design Space

CVPR 2025poster

Cloud removal (CR) remains a challenging task in remote sensing image processing. Although diffusion models (DM) exhibit strong generative capabilities, their direct applications to CR are suboptimal, as they generate cloudless images from random noise, ignoring inherent information in cloudy inputs…

2025

Efficient Constraint-based Window Causal Graph Discovery in Time Series with Multiple Time Lags

IJCAI 2025

We address the identification of direct causes in time series with multiple time lags, and propose a constraint-based window causal graph discovery method. A key advantage of our method is that the number of required conditional independence (CI) tests scales quadratically with the number of sub-ser

Cited by 0SourcePDFScholar
2025

Exploring Inter-Variate and Long-Term Dependencies to Boost Multivariate Time Series Forecasting

ICASSP 2025accepted

Multivariate Time Series Forecasting (MTSF) is a critical task in various domains, and Large Language Models (LLMs) for MTSF have recently received considerable attention. Despite significant progress in large-scale time series models, particularly in fine-tuning pre-trained LLMs for MTSF, there are…

Cited by 0SourceScholar
2025

Identifying Causal Mechanism Shifts Under Additive Models with Arbitrary Noise

IJCAI 2025

In many real-world scenarios, the goal is to identify variables whose causal mechanisms change across related datasets. For example, detecting abnormal root nodes in manufacturing, and identifying key genes that influence cancer by analyzing differences in gene regulatory mechanisms between healthy

Cited by 0SourcePDFScholar
2025

Imagination-Limited Q-Learning for Offline Reinforcement Learning

IJCAI 2025

Offline reinforcement learning seeks to derive improved policies entirely from historical data but often struggles with over-optimistic value estimates for out-of-distribution (OOD) actions. This issue is typically mitigated via policy constraint or conservative value regularization methods. However

2025

Multi-matrix Factorization Attention

ACL 2025finding

We propose novel attention architectures, Multi-matrix Factorization Attention (MFA) and MFA-Key-Reuse (MFA-KR). Existing variants for standard Multi-Head Attention (MHA), including SOTA methods like MLA, fail to maintain as strong performance under stringent Key-Value cache (KV cache) constraints.…

Cited by 0SourcePDFScholar
2025

Predictable Scale (Part II) --- Farseer: A Refined Scaling Law in LLMs

NeurIPS 2025spotlight

Training Large Language Models (LLMs) is prohibitively expensive, creating a critical scaling gap where insights from small-scale experiments often fail to transfer to resource-intensive production systems, thereby hindering efficient innovation. To bridge this, we introduce Farseer, a novel and ref…

Cited by 0SourcecodeScholar
2025

SlerpFace: Face Template Protection via Spherical Linear Interpolation

AAAI 2025technical

Contemporary face recognition systems use feature templates extracted from face images to identify persons. To enhance privacy, face template protection techniques are widely employed to conceal sensitive identity and appearance information stored in the template. This paper identifies an emerging p…

Cited by 6SourcePDFScholar
2025

UIFace: Unleashing Inherent Model Capabilities to Enhance Intra-Class Diversity in Synthetic Face Recognition

ICLR 2025poster

Face recognition (FR) stands as one of the most crucial applications in computer vision. The accuracy of FR models has significantly improved in recent years due to the availability of large-scale human face datasets. However, directly using these datasets can inevitably lead to privacy and legal pr…

2024

Efficiently Learning Significant Fourier Feature Pairs for Statistical Independence Testing

NeurIPS 2024poster

We propose a novel method to efficiently learn significant Fourier feature pairs for maximizing the power of Hilbert-Schmidt Independence Criterion~(HSIC) based independence tests. We first reinterpret HSIC in the frequency domain, which reveals its limited discriminative power due to the inability…

Cited by 0SourcePDFScholar
2024

Learning Adaptive Kernels for Statistical Independence Tests

AISTATS 2024poster

We propose a novel framework for kernel-based statistical independence tests that enable adaptatively learning parameterized kernels to maximize test power. Our framework can effectively address the pitfall inherent in the existing signal-to-noise ratio criterion by modeling the change of the null d…

2024

Multivariate Time Series Forecasting with Causal-Temporal Attention Network

ICASSP 2024accepted

The task of multivariate time series (MTS) forecasting has attracted much attention in recent years. However, most existing methods overlook the causal relationship among different variables, which may lead to inaccurate forecasting results. In this paper, we incorporate causality into the forecasti…

Cited by 0SourceScholar
2024

Privacy-Preserving Face Recognition Using Trainable Feature Subtraction

CVPR 2024poster

The widespread adoption of face recognition has led to increasing privacy concerns as unauthorized access to face images can expose sensitive personal information. This paper explores face image protection against viewing and recovery attacks. Inspired by image compression we propose creating a visu…

2024

Tail Classes Matter: Long-Tailed Object Detection Revisited

ICASSP 2024accepted

Real-world data ubiquitously exhibit long-tailed distribution, which sparks the increasing interest in long-tailed object detection (LTOD). However, existing methods neglect that a lack of diverse data in tail classes will cause underrepresented tail class features, making their efforts for balancin…

Cited by 0SourceScholar
2024

Weakly Supervised Few-Shot Object Detection with DETR

AAAI 2024technical

In recent years, Few-shot Object Detection (FSOD) has become an increasingly important research topic in computer vision. However, existing FSOD methods require strong annotations including category labels and bounding boxes, and their performance is heavily dependent on the quality of box annotatio…

Cited by 3SourcePDFScholar
2023

Differentially Private Nonlinear Causal Discovery from Numerical Data

AAAI 2023technical

Recently, several methods such as private ANM, EM-PC and Priv-PC have been proposed to perform differentially private causal discovery in various scenarios including bivariate, multivariate Gaussian and categorical cases. However, there is little effort on how to conduct private nonlinear causal dis…

2023

DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery – a Focus on Affinity Prediction Problems with Noise Annotations

AAAI 2023technical

AI-aided drug discovery (AIDD) is gaining popularity due to its potential to make the search for new pharmaceuticals faster, less expensive, and more effective. Despite its extensive use in numerous fields (e.g., ADMET prediction, virtual screening), little research has been conducted on the out-of-…

Cited by 122SourcePDFScholar
2023

Multi-Level Wavelet Mapping Correlation for Statistical Dependence Measurement: Methodology and Performance

AAAI 2023technical

We propose a new criterion for measuring dependence between two real variables, namely, Multi-level Wavelet Mapping Correlation (MWMC). MWMC can capture the nonlinear dependencies between variables by measuring their correlation under different levels of wavelet mappings. We show that the empirical…

2023

Privacy-Preserving Face Recognition Using Random Frequency Components

ICCV 2023poster

The ubiquitous use of face recognition has sparked increasing privacy concerns, as unauthorized access to sensitive face images could compromise the information of individuals. This paper presents an in-depth study of the privacy protection of face images' visual information and against recovery. Dr…

Cited by 21PDFcodeScholar
2022

C3-STISR: Scene Text Image Super-resolution with Triple Clues

IJCAI 2022poster

Scene text image super-resolution (STISR) has been regarded as an important pre-processing task for text recognition from low-resolution scene text images. Most recent approaches use the recognizer's feedback as clues to guide super-resolution. However, directly using recognition clue has two proble…

2022

EPiDA: An Easy Plug-in Data Augmentation Framework for High Performance Text Classification

NAACL 2022long

Recent works have empirically shown the effectiveness of data augmentation (DA) in NLP tasks, especially for those suffering from data scarcity. Intuitively, given the size of generated data, their diversity and quality are crucial to the performance of targeted tasks. However, to the best of our kn…

2022

Residual Similarity Based Conditional Independence Test and Its Application in Causal Discovery

AAAI 2022technical

Recently, many regression based conditional independence (CI) test methods have been proposed to solve the problem of causal discovery. These methods provide alternatives to test CI by first removing the information of the controlling set from the two target variables, and then testing the independe…

2022

Towards Video Text Visual Question Answering: Benchmark and Baseline

NeurIPS 2022accept

There are already some text-based visual question answering (TextVQA) benchmarks for developing machine's ability to answer questions based on texts in images in recent years. However, models developed on these benchmarks cannot work effectively in many real-life scenarios (e.g. traffic monitoring,…

2021

Accurate Few-Shot Object Detection With Support-Query Mutual Guidance and Hybrid Loss

CVPR 2021poster

Most object detection methods require huge amounts of annotated data and can detect only the categories that appear in the training set. However, in reality acquiring massive annotated training data is both expensive and time-consuming. In this paper, we propose a novel two-stage detector for accura…

Cited by 75PDFScholar
2021

DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled Samples

NeurIPS 2021poster

The scarcity of labeled data is a critical obstacle to deep learning. Semi-supervised learning (SSL) provides a promising way to leverage unlabeled data by pseudo labels. However, when the size of labeled data is very small (say a few labeled samples per class), SSL performs poorly and unstably, pos…

Cited by 36SourcePDFScholar
2021

GIF Thumbnails: Attract More Clicks to Your Videos

AAAI 2021technical

With the rapid increase of mobile devices and online media, more and more people prefer posting/viewing videos online. Generally, these videos are presented on video streaming sites with image thumbnails and text titles. While facing huge amounts of videos, a viewer clicks through a certain video wi…

2021

Testing Independence Between Linear Combinations for Causal Discovery

AAAI 2021technical

Recently, regression based conditional independence (CI) tests have been employed to solve the problem of causal discovery. These methods provide an alternative way to test for CI by transforming CI to independence between residuals. Generally, it is nontrivial to check for independence when these r…

Cited by 20SourcePDFScholar
2021

Weakly-supervised Text Classification Based on Keyword Graph

EMNLP 2021main

Weakly-supervised text classification has received much attention in recent years for it can alleviate the heavy burden of annotating massive data. Among them, keyword-driven methods are the mainstream where user-provided keywords are exploited to generate pseudo-labels for unlabeled texts. However,…

2018

AON: Towards Arbitrarily-Oriented Text Recognition

CVPR 2018poster

Recognizing text from natural images is a hot research topic in computer vision due to its various applications. Despite the enduring research of several decades on optical character recognition (OCR), recognizing texts from natural images is still a challenging task. This is because scene texts are…

Cited by 360SourcePDFScholar
2017

Focusing Attention: Towards Accurate Text Recognition in Natural Images

ICCV 2017poster

Scene text recognition has been a hot research topic in computer vision due to its various applications. The state of the art is the attention-based encoder-decoder framework that learns the mapping between input images and output sequences in a purely data-driven way. However, we observe that exist…

Cited by 626PDFScholar