← Search

Si Wu

56 accepted papers

2026

Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection

CVPR 2026

Modern industrial quality control heavily relies on automated anomaly detection. While few-shot anomaly detection addresses the challenge of limited labeled data, real-world inspection faces a vast diversity of anomaly types, sizes, and shapes. We identify the primary cause for the anomaly detection

Cited by 0SourceScholar
2026

From Representation to Action: A Unified Laplacian Framework for Spatial Representation and Path Planning

ICML 2026poster

Navigation in complex environments relies on internal spatial representations that guide action. While the brain employs a diverse repertoire of spatial tuning cells—including grid, place, and head-direction cells—a normative theory linking these static neural codes to the dynamic process of navigat…

Cited by 0SourceScholar
2026

Predicting Context-Aware Transcriptional Responses to Unseen Genetic Perturbation Subject to Interactome Distance Constraints

IJCAI 2026

Predicting responses to genetic perturbation is pivotal for elucidating gene regulatory machinery. However, existing methods often rely on statistical perspectives to model differential expression, overlooking the constraints of the underlying molecular interactome, which renders predictions suscept

Cited by 0Scholar
2026

Refinement Contrastive Learning of Cell–Gene Associations for Unsupervised Cell Type Identification

AAAI 2026technical

Unsupervised cell type identification is crucial for uncovering and characterizing heterogeneous populations in single cell omics studies. Although a range of clustering methods have been developed, most focus exclusively on intrinsic cellular structure and ignore the pivotal role of cell-gene assoc

Cited by 0SourcePDFScholar
2026

Structure Abstraction and Generalization in a Hippocampus-Entorhinal Inspired World Model

ICML 2026poster

Humans abstract experiences into structured representations to facilitate pattern inference and knowledge transfer. While the hippocampal-entorhinal (HPC-MEC) circuit is known to represent both spatial and conceptual spaces, the mechanisms for concurrently extracting abstract structures from continu…

Cited by 0SourceScholar
2026

Syntactic Structure-Guided Visual Grounding with Subject-Centric Feature Enhancement and Verification

IJCAI 2026

Visual grounding aims to localize target objects based on natural language descriptions, and the core challenge lies in the cross-modal gap, which is partly caused by the significant differences in semantic structure between language and vision. Existing methods typically rely on holistic sentence-l

Cited by 0Scholar
2025

3DHumanEdit: Multi-modal Body Part-aware Conditioning Information Integration for 3D Human Manipulation

AAAI 2025technical

The rapid advancement of 3D Generative Adversarial Networks (GANs) has significantly enhanced the diversity and quality of generated 3D images. Despite these breakthroughs, the manipulation capabilities of 3D GANs remain unexplored, presenting substantial challenges for practical applications where…

Cited by 0SourcePDFScholar
2025

Discrete Prior-Based Temporal-Coherent Content Prediction for Blind Face Video Restoration

AAAI 2025technical

Blind face video restoration aims to restore high-fidelity details from videos subjected to complex and unknown degradations. This task poses a significant challenge of managing temporal heterogeneity while at the same time maintaining stable face attributes. In this paper, we introduce a Discrete P…

2025

Dynamic Content Prediction with Motion-aware Priors for Blind Face Video Restoration

CVPR 2025poster

Blind Face Video Restoration (BFVR) focuses on reconstructing high-quality facial image sequences from degraded video inputs. The main challenge is address unknown degradations, while maintaining temporal consistency across frames. Current blind face restoration methods are primarily designed for im…

Cited by 0SourcePDFScholar
2025

Facilitating Semi-Supervised Pedestrian Detection with Structurally Controllable Instance Synthesis

ICASSP 2025accepted

The performance of pedestrian detectors typically relies on sufficient labeled data, and semi-supervised learning is a promising way to address the deficiency in manual annotations by utilizing sufficient unlabeled images. In this work, we design a Structure-Controllable Pedestrian Instance Generati…

Cited by 0SourceScholar
2025

Prompt-augmented Feature with Cross-domain Contrastive Learning for Efficient Multi-domain Sentiment Analysis

ICASSP 2025accepted

Pre-trained language models (PrLMs) demonstrate impressive performance on the sentiment analysis task. However, the large number of trainable parameters brings about heavy computational costs, which become more serious in multi-domain scenarios. In this paper, we propose to extract multi-layer featu…

Cited by 0SourceScholar
2025

RetouchGPT: LLM-based Interactive High-Fidelity Face Retouching via Imperfection Prompting

AAAI 2025technical

Face retouching aims to remove facial imperfections from image and videos while at the same time preserving face attributes. The existing methods are designed to perform non-interactive end-to-end retouching, while the ability to interact with users is highly demanded in downstream applications. In…

Cited by 0SourcePDFScholar
2025

Self-Correcting Robot Manipulation via Gaussian-Splatted Foresight

AAAI 2025technical

Language-conditioned robotic manipulation in unstructured environments presents significant challenges for intelligent robotic systems. However, due to partial observation or imprecise action prediction, failure may be unavoidable for learned policies. Moreover, operational failures can lead to the…

Cited by 0SourcePDFScholar
2025

SpotDiff: Spatial Gene Expression Imputation Diffusion with Single-Cell RNA Sequencing Data Integration

AAAI 2025technical

The advent of Spatial Transcriptomics (ST) has revolutionized understanding of tissue architecture by creating high-resolution maps of gene expression patterns. However, the low capture rate of ST leads to significant sparsity. The aim of imputation is to recover biological signals by imputing the d…

Cited by 0SourcePDFScholar
2025

Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding

CVPR 2025poster

The goal of visual grounding is to establish connections between target objects and textual descriptions. Large Language Models (LLMs) have demonstrated strong comprehension abilities across a variety of visual tasks. To establish precise associations between the text and the corresponding visual re…

Cited by 0SourcePDFScholar
2025

Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Space

ACL 2025long

Imageability (potential of text to evoke a mental image) and concreteness (perceptibility of text) are two psycholinguistic properties that link visual and semantic spaces. It is little surprise that computational methods that estimate them do so using parallel visual and semantic spaces, such as co…

2025

Unfolding the Black Box of Recurrent Neural Networks for Path Integration

NeurIPS 2025poster

Path integration is essential for spatial navigation. Experimental studies have identified neural correlates for path integration, but exactly how the neural system accomplishes this computation remains unresolved. Here, we adopt recurrent neural networks (RNNs) trained to perform a path integration…

Cited by 0SourceScholar
2024

A differentiable brain simulator bridging brain simulation and brain-inspired computing

ICLR 2024poster

Brain simulation builds dynamical models to mimic the structure and functions of the brain, while brain-inspired computing (BIC) develops intelligent systems by learning from the structure and functions of the brain. The two fields are intertwined and should share a common programming framework to f…

Cited by 4SourcePDFScholar
2024

AttriHuman-3D: Editable 3D Human Avatar Generation with Attribute Decomposition and Indexing

CVPR 2024poster

Editable 3D-aware generation which supports user-interacted editing has witnessed rapid development recently. However existing editable 3D GANs either fail to achieve high-accuracy local editing or suffer from huge computational costs. We propose AttriHuman-3D an editable 3D human generation model w…

Cited by 9SourcePDFScholar
2024

Learning Degradation-unaware Representation with Prior-based Latent Transformations for Blind Face Restoration

CVPR 2024poster

Blind face restoration focuses on restoring high-fidelity details from images subjected to complex and unknown degradations while preserving identity information. In this paper we present a Prior-based Latent Transformation approach (PLTrans) which is specifically designed to learn a degradation-una…

Cited by 4SourcePDFScholar
2024

Relational Matching for Weakly Semi-Supervised Oriented Object Detection

CVPR 2024poster

Oriented object detection has witnessed significant progress in recent years. However the impressive performance of oriented object detectors is at the huge cost of labor-intensive annotations and deteriorates once the annotated data becomes limited. Semi-supervised learning in which sufficient unan…

Cited by 4SourcePDFScholar
2024

RetouchFormer: Semi-supervised High-Quality Face Retouching Transformer with Prior-Based Selective Self-Attention

AAAI 2024technical

Face retouching is to beautify a face image, while preserving the image content as much as possible. It is a promising yet challenging task to remove face imperfections and fill with normal skin. Generic image enhancement methods are hampered by the lack of imperfection localization, which often res…

Cited by 1SourcePDFScholar
2024

SCTrans: Multi-scale scRNA-seq Sub-vector Completion Transformer for Gene-selective Cell Type Annotation

IJCAI 2024poster

Cell type annotation is pivotal to single-cell RNA sequencing data (scRNA-seq)-based biological and medical analysis, e.g., identifying biomarkers, exploring cellular heterogeneity, and understanding disease mechanisms. The previous annotation methods typically learn a nonlinear mapping to infer cel…

Cited by 0SourcePDFScholar
2024

Text-conditional Attribute Alignment across Latent Spaces for 3D Controllable Face Image Synthesis

CVPR 2024poster

With the advent of generative models and vision language pretraining significant improvement has been made in text-driven face manipulation. The text embedding can be used as target supervision for expression control.However it is non-trivial to associate with its 3D attributesi.e. pose and illumina…

Cited by 0SourcePDFScholar
2024

The motion planning neural circuit in goal-directed navigation as Lie group operator search

NeurIPS 2024poster

The information processing in the brain and embodied agents form a sensory-action loop to interact with the world. An important step in the loop is motion planning which selects motor actions based on the current world state and task need. In goal-directed navigation, the brain chooses and generates…

Cited by 0SourcePDFScholar
2024

To Learn or Not to Learn, That is the Question — A Feature-Task Dual Learning Model of Perceptual Learning

NeurIPS 2024poster

Perceptual learning refers to the practices through which participants learn to improve their performance in perceiving sensory stimuli. Two seemingly conflicting phenomena of specificity and transfer have been widely observed in perceptual learning. Here, we propose a dual-learning model to recon…

Cited by 0SourcePDFScholar
2024

VRetouchEr: Learning Cross-frame Feature Interdependence with Imperfection Flow for Face Retouching in Videos

CVPR 2024poster

Face Video Retouching is a complex task that often requires labor-intensive manual editing. Conventional image retouching methods perform less satisfactorily in terms of generalization performance and stability when applied to videos without exploiting the correlation among frames. To address this i…

Cited by 1SourcePDFScholar
2023

A Recurrent Neural Circuit Mechanism of Temporal-scaling Equivariant Representation

NeurIPS 2023poster

Time perception is critical in our daily life. An important feature of time perception is temporal scaling (TS): the ability to generate temporal sequences (e.g., motor actions) at different speeds. However, it is largely unknown about the math principle underlying temporal scaling in recurrent circ…

Cited by 1SourcePDFScholar
2023

Blemish-Aware and Progressive Face Retouching With Limited Paired Data

CVPR 2023poster

Face retouching aims to remove facial blemishes, while at the same time maintaining the textual details of a given input image. The main challenge lies in distinguishing blemishes from the facial characteristics, such as moles. Training an image-to-image translation network with pixel-wise supervisi…

Cited by 5SourcePDFScholar
2023

Exploring Intra-Class Variation Factors With Learnable Cluster Prompts for Semi-Supervised Image Synthesis

CVPR 2023poster

Semi-supervised class-conditional image synthesis is typically performed by inferring and injecting class labels into a conditional Generative Adversarial Network (GAN). The supervision in the form of class identity may be inadequate to model classes with diverse visual appearances. In this paper, w…

Cited by 3SourcePDFScholar
2023

Learning and processing the ordinal information of temporal sequences in recurrent neural circuits

NeurIPS 2023poster

Temporal sequence processing is fundamental in brain cognitive functions. Experimental data has indicated that the representations of ordinal information and contents of temporal sequences are disentangled in the brain, but the neural mechanism underlying this disentanglement remains largely unclea…

Cited by 0SourcePDFScholar
2023

Slow and Weak Attractor Computation Embedded in Fast and Strong E-I Balanced Neural Dynamics

NeurIPS 2023spotlight

Attractor networks require neuronal connections to be highly structured in order to maintain attractor states that represent information, while excitation and inhibition balanced networks (E-INNs) require neuronal connections to be random and sparse to generate irregular neuronal firings. Despite be…

Cited by 3SourcePDFScholar
2023

Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image Manipulation

CVPR 2023poster

Great progress has been made in StyleGAN-based image editing. To associate with preset attributes, most existing approaches focus on supervised learning for semantically meaningful latent space traversal directions, and each manipulation step is typically determined for an individual attribute. To a…

Cited by 3SourcePDFScholar
2022

Adaptation Accelerating Sampling-based Bayesian Inference in Attractor Neural Networks

NeurIPS 2022accept

The brain performs probabilistic Bayesian inference to interpret the external world. The sampling-based view assumes that the brain represents the stimulus posterior distribution via samples of stochastic neuronal responses. Although the idea of sampling-based inference is appealing, it faces a crit…

Cited by 8SourcePDFScholar
2022

Oscillatory Tracking of Continuous Attractor Neural Networks Account for Phase Precession and Procession of Hippocampal Place Cells

NeurIPS 2022accept

Hippocampal place cells of freely moving rodents display an intriguing temporal organization in their responses known as `theta phase precession', in which individual neurons fire at progressively earlier phases in successive theta cycles as the animal traverses the place fields. Recent experimental…

Cited by 6SourcePDFScholar
2022

SphericGAN: Semi-Supervised Hyper-Spherical Generative Adversarial Networks for Fine-Grained Image Synthesis

CVPR 2022poster

Generative Adversarial Network (GAN)-based models have greatly facilitated image synthesis. However, the model performance may be degraded when applied to fine-grained data, due to limited training samples and subtle distinction among categories. Different from generic GANs, we address the issue fro…

Cited by 18PDFScholar
2022

Translation-equivariant Representation in Recurrent Networks with a Continuous Manifold of Attractors

NeurIPS 2022accept

Equivariant representation is necessary for the brain and artificial perceptual systems to faithfully represent the stimulus under some (Lie) group transformations. However, it remains unknown how recurrent neural circuits in the brain represent the stimulus equivariantly, nor the neural representat…

Cited by 7SourcePDFScholar
2021

High Fidelity GAN Inversion via Prior Multi-Subspace Feature Composition

AAAI 2021technical

Generative Adversarial Networks (GANs) have shown impressive gains in image synthesis. GAN inversion was recently studied to understand and utilize the knowledge it learns, where a real image is inverted back to a latent code and can thus be reconstructed by the generator. Although increasing the nu…

Cited by 0SourcePDFScholar
2021

Mask-Embedded Discriminator With Region-Based Semantic Regularization for Semi-Supervised Class-Conditional Image Synthesis

CVPR 2021poster

Semi-supervised generative learning (SSGL) makes use of unlabeled data to achieve a trade-off between the data collection/annotation effort and generation performance, when adequate labeled data are not available. Learning precise class semantics is crucial for class-conditional image synthesis with…

Cited by 6PDFScholar
2021

Noisy Adaptation Generates Lévy Flights in Attractor Neural Networks

NeurIPS 2021poster

Lévy flights describe a special class of random walks whose step sizes satisfy a power-law tailed distribution. As being an efficient searching strategy in unknown environments, Lévy flights are widely observed in animal foraging behaviors. Recent studies further showed that human cognitive function…

Cited by 16SourcePDFScholar
2021

Scalable Font Reconstruction with Dual Latent Manifolds

EMNLP 2021main

We propose a deep generative model that performs typography analysis and font reconstruction by learning disentangled manifolds of both font style and character shape. Our approach enables us to massively scale up the number of character types we can effectively model compared to previous methods. S…

2021

Semi-Supervised Single-Stage Controllable GANs for Conditional Fine-Grained Image Generation

ICCV 2021poster

Previous state-of-the-art deep generative models improve fine-grained image generation quality by designing hierarchical model structures and synthesizing images across multiple stages. The learning process is typically performed without any supervision in object categories. To address this issue, w…

Cited by 10PDFScholar
2020

An Attention-driven Two-stage Clustering Method for Unsupervised Person Re-Identification

ECCV 2020poster

The progressive clustering method and its variants, which iteratively generate pseudo labels for unlabeled data and perform feature learning, have shown great process in unsupervised person re-identification (re-id). However, they have an intrinsic problem of modeling the in-camera variability of im…

Cited by 65SourcePDFScholar
2020

Model Adaptation: Unsupervised Domain Adaptation Without Source Data

CVPR 2020poster

In this paper, we investigate a challenging unsupervised domain adaptation setting --- unsupervised model adaptation. We aim to explore how to rely only on unlabeled target data to improve performance of an existing source prediction model on the target domain, since labeled source data may not be a…

Cited by 0PDFScholar
2020

Regularizing Discriminative Capability of CGANs for Semi-Supervised Generative Learning

CVPR 2020poster

Semi-supervised generative learning aims to learn the underlying class-conditional distribution of partially labeled data. Generative Adversarial Networks (GANs) have led to promising progress in this task. However, it still needs to further explore the issue of imbalance between real labeled data a…

Cited by 29PDFScholar
2019

A Normative Theory for Causal Inference and Bayes Factor Computation in Neural Circuits

NeurIPS 2019poster

This study provides a normative theory for how Bayesian causal inference can be implemented in neural circuits. In both cognitive processes such as causal reasoning and perceptual inference such as cue integration, the nervous systems need to choose different models representing the underlying causa…

2019

Enhancing TripleGAN for Semi-Supervised Conditional Instance Synthesis and Classification

CVPR 2019poster

Learning class-conditional data distributions is crucial for Generative Adversarial Networks (GAN) in semi-supervised learning. To improve both instance synthesis and classification in this setting, we propose an enhanced TripleGAN (EnhancedTGAN) model in this work. We follow the adversarial trainin…

Cited by 41PDFScholar
2019

Mutual Learning of Complementary Networks via Residual Correction for Improving Semi-Supervised Classification

CVPR 2019oral

Deep mutual learning jointly trains multiple essential networks having similar properties to improve semi-supervised classification. However, the commonly used consistency regularization between the outputs of the networks may not fully leverage the difference between them. In this paper, we explore…

Cited by 44PDFScholar
2019

Push-pull Feedback Implements Hierarchical Information Retrieval Efficiently

NeurIPS 2019poster

Experimental data has revealed that in addition to feedforward connections, there exist abundant feedback connections in a neural pathway. Although the importance of feedback in neural information processing has been widely recognized in the field, the detailed mechanism of how it works remains larg…

2019

Semi-Supervised Pedestrian Instance Synthesis and Detection With Mutual Reinforcement

ICCV 2019poster

We propose a GAN-based scene-specific instance synthesis and classification model for semi-supervised pedestrian detection. Instead of collecting unreliable detections from unlabeled data, we adopt a class-conditional GAN for synthesizing pedestrian instances to alleviate the problem of insufficient…

Cited by 12PDFScholar
2016

“Congruent” and “Opposite” Neurons: Sisters for Multisensory Integration and Segregation

NeurIPS 2016poster

Experiments reveal that in the dorsal medial superior temporal (MSTd) and the ventral intraparietal (VIP) areas, where visual and vestibular cues are integrated to infer heading direction, there are two types of neurons with roughly the same number. One is “congruent” cells, whose preferred heading…

Cited by 10SourcePDFScholar