← Search

Wei Huang

126 accepted papers

2026

Absorbing Quantization Error by Deformable Noise Scheduler for Diffusion Models

ICML 2026poster

Diffusion models deliver state-of-the-art image quality but are expensive to deploy. Post-training quantization (PTQ) can shrink models and speed up inference, yet residual quantization errors distort the diffusion distribution (the timestep-wise marginal over $\vx_t$), degrading sample quality. We …

Cited by 0SourceScholar
2026

CoEvoer: Collaborative Evolution Transformer for Upper-Body Expressive Human Pose and Shape Estimation

AAAI 2026technical

Expressive Human Pose and Shape Estimation (EHPS) plays a crucial role in various AR/VR applications and has witnessed significant progress in recent years. However, current state-of-the-art methods still struggle with accurate parameter estimation for facial and hand regions and exhibit limited gen

Cited by 0SourcePDFScholar
2026

DualFete: Revisiting Teacher-Student Interactions from a Feedback Perspective for Semi-supervised Medical Image Segmentation

AAAI 2026technical

The teacher-student paradigm has emerged as a canonical framework in semi-supervised learning. When applied to medical image segmentation, the paradigm faces challenges due to inherent image ambiguities, making it particularly vulnerable to erroneous supervision. Crucially, the student

Cited by 0SourcePDFScholar
2026

Energy Waveify and Redistribution for Test-Time Adaptation: A Control System Perspective

CVPR 2026

This work tackles a key challenge in test-time energy adaptation: prohibitive time overhead arising from recent state-of-the-art test-time adaptation (TTA) methods, which are built on energy models relying on iterative Monte Carlo or Langevin dynamics sampling with multiple stochastic updates per te

Cited by 0SourcecodeScholar
2026

GradPruner: Gradient-guided Layer Pruning Enabling Efficient Fine-Tuning and Inference for LLMs

ICLR 2026poster

Fine-tuning Large Language Models (LLMs) with downstream data is often considered time-consuming and expensive. Structured pruning methods are primarily employed to improve the inference efficiency of pre-trained models. Meanwhile, they often require additional time and memory for training, knowledg…

Cited by 0SourceScholar
2026

Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models

CVPR 2026

Vision-language models (VLM) excel at general understanding yet remain weak at dynamic spatial reasoning (DSR), i.e., reasoning about the evolvement of object geometry and relationship in 3D space over time, largely due to the scarcity of scalable 4D-aware training resources. To bridge this gap acro

Cited by 0SourcecodeScholar
2026

LongLive: Real-time Interactive Long Video Generation

ICLR 2026poster

We present LongLive, a frame-level autoregressive (AR) framework for real-time and interactive long video generation. Long video generation presents challenges in both efficiency and quality. Diffusion and Diffusion-Forcing models can produce high-quality videos but suffer from low efficiency due to…

Cited by 188SourcecodeScholar
2026

MIMOMamba: From Scalar Duality to Matrix-Valued Attention

ICML 2026poster

The state space duality (SSD) framework, central to modern state-space models (SSMs) such as Mamba, has established an efficient attention-like mechanism by leveraging the commutative property of linear recurrences. However, existing formulations are limited to single-input single-output (SISO) syst…

Cited by 0SourceScholar
2026

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

ICLR 2026poster

Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study the design choices across model architecture and data curati…

Cited by 0SourcecodeScholar
2026

On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD

AAAI 2026technical

One crucial factor behind the success of deep learning lies in the implicit bias induced by noise inherent in gradient-based training algorithms. Motivated by empirical observations that training with noisy labels improves model generalization, we delve into the underlying mechanisms behind stochast

Cited by 0SourcePDFScholar
2026

Prior Refinement Is Better: Diffusion-Driven Graph Harmonization for Federated Graph Learning

AAAI 2026technical

Federated Graph Learning (FGL) has emerged as a compelling paradigm for collaboratively training a global model while preserving the privacy of multi-source graphs. Nonetheless, FGL faces a critical challenge of data heterogeneity, where semantic and structural discrepancies across clients significa

Cited by 0SourcePDFScholar
2026

Provable Sample Efficiency of Curriculum Post-Training for Transformer Reasoning

ICML 2026poster

Recent curriculum techniques in the post-training stage of LLMs have been empirically observed to outperform non-curriculum approaches in improving reasoning performance, yet a principled understanding of their effectiveness and limitations remains incomplete. To bridge this gap, we develop an abstr…

Cited by 0SourceScholar
2026

QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMs

ICLR 2026poster

We propose QeRL, a Quantization-enhanced Reinforcement Learning framework for large language models (LLMs). While RL is essential for LLMs' reasoning capabilities, it is resource-intensive, requiring substantial GPU memory and long rollout duration. QeRL addresses these issues by combining NVFP4 qua…

Cited by 0SourcecodeScholar
2026

SELF-HARMONY: LEARNING TO HARMONIZE SELF-SUPERVISION AND SELF-PLAY IN TEST-TIME REINFORCEMENT LEARNING

ICLR 2026poster

Test-time reinforcement learning (TTRL) offers a label-free paradigm for adapting models using only synthetic signals at inference, but its success hinges on constructing reliable learning signals. Standard approaches such as majority voting often collapse to spurious yet popular answers. We introdu…

Cited by 0SourceScholar
2026

Spectral Gradient Descent Mitigates Anisotropy-Driven Misalignment: A Case Study in Phase Retrieval

ICML 2026poster

Spectral gradient methods, such as the Muon optimizer, modify gradient updates by preserving directional information while discarding scale, and have shown strong empirical performance in deep learning. We investigate the mechanisms underlying these gains through a dynamical analysis of a nonlinear …

Cited by 0SourceScholar
2026

Towards Generative Graph Matching for Graph Edit Distance Computation

ICML 2026poster

Graph Edit Distance (GED), which aims to find an edit path with minimum number of edit operations to transform one graph into another, is a fundamental NP-hard problem and a widely used graph similarity measure. Recent matching-based hybrid approaches have demonstrated better scalability than A* sea…

Cited by 0SourceScholar
2026

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression

ICML 2026poster

Extended reasoning in large language models (LLMs) requires long and accurate decoding and creates severe KV cache memory bottlenecks. Leading KV cache compression methods estimate KV importance using attention scores from recent post-RoPE queries. However, queries rotate with position during RoPE, …

Cited by 0SourceScholar
2026

Two Modalities Are Better Than One: Efficient Adversarial Purification via Multimodal Diffusion Models

ICML 2026poster

Adversarial purification uses generative models to restore clean data distributions from unseen attacks without retraining classifiers. However, unimodal diffusion-based approaches struggle to preserve semantic consistency, while recent multimodal variants rely on computationally expensive adversari…

Cited by 0SourceScholar
2026

Whole-Body Path Following of an Active-Joint Active-Wheel Snake Robot Based on Head-Body Stepwise Control

RA-L 2026

This paper presents a two-dimensional planar path-following controller for active-joint active-wheel snake-like robots. Compared to passive-wheel snake robots, active-wheel snake robots can navigate narrower spaces and generate stronger driving forces, enabling adaptation to various terrains. Howeve

Cited by 0SourceScholar
2025

A Fully Probabilistic Perspective on Large Language Model Unlearning: Evaluation and Optimization

EMNLP 2025

Large Language Model Unlearning (LLMU) is a promising way to remove private or sensitive information from large language models. However, the comprehensive evaluation of LLMU remains underexplored. The dominant deterministic evaluation can yield overly optimistic assessments of unlearning efficacy.

Cited by 0SourcePDFScholar
2025

Breaking Data Silos in Parkinson’s Disease Diagnosis: An Adaptive Federated Learning Approach for Privacy-Preserving Facial Expression Analysis

AAAI 2025technical

The early diagnosis of Parkinson’s disease (PD) is crucial for potential patients to receive timely treatment and prevent disease progression. Recent studies have shown that PD is closely linked to impairments in facial muscle control, resulting in characteristic “masked face” symptoms. This discove…

Cited by 1SourcePDFScholar
2025

Breaking the Stage Barrier: A Novel Single-Stage Approach to Long Context Extension for Large Language Models

COLING 2025main

Recently, Large language models (LLMs) have revolutionized Natural Language Processing (NLP). Pretrained LLMs, due to limited training context size, struggle with handling long token sequences, limiting their performance on various downstream tasks. Current solutions toward long context modeling oft…

Cited by 3SourcePDFScholar
2025

Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images?

ICML 2025poster

Despite the remarkable success of diffusion models (DMs) in data generation, they exhibit specific failure cases with unsatisfactory outputs. We focus on one such limitation: the ability of DMs to learn hidden rules between image features. Specifically, for image data with dependent features ($\math…

Cited by 1SourcePDFScholar
2025

DPF-CM: A Data Processing Framework with Privacy-Preserving Vector Databases for Chinese Medical LLMs Training and Deployment

EMNLP 2025

Current open-source training pipelines for Chinese medical language models predominantly emphasize optimizing training methodologies to enhance the performance of large language models (LLMs), yet lack comprehensive exploration into training data processing. To address this gap, we propose DPF-CM, a

Cited by 0SourcePDFScholar
2025

DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression

EMNLP 2025

Large language models (LLMs) excel in general tasks but struggle with domain-specific ones, requiring fine-tuning with specific data. With many open-source LLMs available, selecting the best model for fine-tuning downstream tasks is challenging, primarily focusing on how to quickly identify the opti

Cited by 0SourcePDFScholar
2025

Fast Quiet-STaR: Thinking Without Thought Tokens

EMNLP 2025

Large Language Models (LLMs) have achieved impressive performance across a range of natural language processing tasks. However, recent advances demonstrate that further gains—particularly in complex reasoning tasks—require more than merely scaling up model sizes or training data. One promising direc

2025

From Layers to States: A State Space Model Perspective to Deep Neural Network Layer Dynamics

ICLR 2025poster

The depth of neural networks is a critical factor for their capability, with deeper models often demonstrating superior performance. Motivated by this, significant efforts have been made to enhance layer aggregation - reusing information from previous layers to better extract features at the current…

Cited by 0SourcePDFScholar
2025

GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs

ICML 2025poster

Large language model (LLM) unlearning has demonstrated its essential role in removing privacy and copyright-related responses, crucial for their legal and safe applications. However, the pursuit of complete unlearning often comes with substantial costs due to its compromises in their general functio…

Cited by 0SourcePDFScholar
2025

GaRe: Relightable 3D Gaussian Splatting for Outdoor Scenes from Unconstrained Photo Collections

ICCV 2025poster

We propose a 3D Gaussian splatting-based framework for outdoor relighting that leverages intrinsic image decomposition to precisely integrate sunlight, sky radiance, and indirect lighting from unconstrained photo collections. Unlike prior methods that compress the per-image global illumination into…

Cited by 0SourcePDFScholar
2025

GapMatch: Bridging Instance and Model Perturbations for Enhanced Semi-Supervised Medical Image Segmentation

AAAI 2025technical

Medical image segmentation provides detailed understanding and aids in diagnosis, treatment planning, and monitoring of diseases. Due to the high cost of acquiring labeled data in the field of medical image analysis, semi-supervised segmentation methods have garnered increasing attention. Benefiting…

Cited by 0SourcePDFScholar
2025

Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel

NeurIPS 2025poster

Gradient-based optimization methods have shown remarkable empirical success, yet their theoretical generalization properties remain only partially understood. In this paper, we establish a generalization bound for gradient flow that aligns with the classical Rademacher complexity bounds for kernel m…

Cited by 0SourceScholar
2025

How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?

NeurIPS 2025poster

The capacity of deep learning models is often large enough to both learn the underlying statistical signal and overfit to noise in the training set. This noise memorization can be harmful especially for data with a low signal-to-noise ratio (SNR), leading to poor generalization. Inspired by prior ob…

Cited by 0SourceScholar
2025

Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied Learning

NeurIPS 2025poster

Offline reinforcement learning (offline RL) is increasingly approached as a sequence modeling task, with methods leveraging advanced architectures like Transformers to capture trajectory dependencies. Despite significant progress, the mechanisms underlying their effectiveness and limitations remain…

Cited by 0SourcecodeScholar
2025

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

NeurIPS 2025poster

Recent text-to-image systems face limitations in handling multimodal inputs and complex reasoning tasks. We introduce MindOmni, a unified multimodal large language model that addresses these challenges by incorporating reasoning generation through reinforcement learning. MindOmni leverages a three-p…

Cited by 0SourcecodeScholar
2025

Mitigating Catastrophic Forgetting in Large Language Models with Forgetting-aware Pruning

EMNLP 2025

Recent advancements in large language models (LLMs) have shown impressive capabilities in various downstream tasks but typically face Catastrophic Forgetting (CF) during fine-tuning. In this paper, we propose the Forgetting-Aware Pruning Metric (FAPM), a novel pruning-based approach to balance CF an

2025

Mixture Compressor for Mixture-of-Experts LLMs Gains More

ICLR 2025poster

Mixture-of-Experts large language models (MoE-LLMs) marks a significant step forward of language models, however, they encounter two critical challenges in practice: 1) expert parameters lead to considerable memory consumption and loading latency; and 2) the current activated experts are redundant,…

2025

Multinoulli Extension: A Lossless Yet Effective Probabilistic Framework for Subset Selection over Partition Constraints

ICML 2025poster

Identifying the most representative subset for a close-to-submodular objective while satisfying the predefined partition constraint is a fundamental task with numerous applications in machine learning. However, the existing distorted local-search methods are often hindered by their prohibitive que…

Cited by 0SourcePDFScholar
2025

NLPrompt: Noise-Label Prompt Learning for Vision-Language Models

CVPR 2025highlight

The emergence of vision-language foundation models, such as CLIP, has revolutionized image-text representation, enabling a broad range of applications via prompt learning. Despite its promise, real-world datasets often contain noisy labels that can degrade prompt learning performance. In this paper,…

2025

On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent

ICLR 2025spotlight

The Adam optimizer is widely used for transformer optimization in practice, which makes understanding the underlying optimization mechanisms an important problem. However, due to the Adam's complexity, theoretical analysis of how it optimizes transformers remains a challenging task. Fortunately, Si…

Cited by 3SourcePDFScholar
2025

On the Role of Label Noise in the Feature Learning Process

ICML 2025poster

Deep learning with noisy labels presents significant challenges. In this work, we theoretically characterize the role of label noise from a feature learning perspective. Specifically, we consider a signal-noise data distribution, where each sample comprises a label-dependent signal and label-indepen…

2025

Provable In-Context Vector Arithmetic via Retrieving Task Concepts

ICML 2025poster

In-context learning (ICL) has garnered significant attention for its ability to grasp functions/tasks from demonstrations. Recent studies suggest the presence of a latent **task/function vector** in LLMs during ICL. Merullo et al. (2024) showed that LLMs leverage this vector alongside the residual s…

Cited by 0SourcePDFScholar
2025

Quantifying the Optimization and Generalization Advantages of Graph Neural Networks Over Multilayer Perceptrons

AISTATS 2025poster

Graph neural networks (GNNs) have demonstrated remarkable capabilities in learning from graph-structured data, often outperforming traditional Multilayer Perceptrons (MLPs) in numerous graph-based tasks. Although existing works have demonstrated the benefits of graph convolution through Laplacian sm…

Cited by 0SourceScholar
2025

Rethinking Out-of-Distribution Detection and Generalization with Collective Behavior Dynamics

NeurIPS 2025poster

Out-of-distribution (OOD) problems commonly occur when models process data with a distribution significantly deviates from the in-distribution (InD) training data. In this paper, we hypothesize that a $\textit{field}$ or $\textit{potential}$ more essential than features exists, and features are not…

Cited by 0SourceScholar
2025

Scaling Diffusion Transformers Efficiently via $\mu$P

NeurIPS 2025poster

Diffusion Transformers have emerged as the foundation for vision generative models, but their scalability is limited by the high cost of hyperparameter (HP) tuning at large scales. Recently, Maximal Update Parametrization ($\mu$P) was proposed for vanilla Transformers, which enables stable HP transf…

Cited by 0SourceScholar
2025

SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models

ICML 2025poster

Post-training quantization (PTQ) is an effective technique for compressing large language models (LLMs). However, while uniform-precision quantization is computationally efficient, it often compromises model performance. To address this, we propose SliM-LLM, a salience-driven mixed-precision quantiz…

2025

TIME-FS: Joint Learning of Tensorial Incomplete Multi-View Unsupervised Feature Selection and Missing-View Imputation

AAAI 2025technical

Multi-view unsupervised feature selection (MUFS) has received considerable attention in recent years. Existing MUFS methods for processing unlabeled incomplete multi-view data, where some samples are missing in certain views, first impute the missing values and then perform feature selection on the…

2025

Temporal Scaling Law for Large Language Models

EMNLP 2025

Recently, Large Language Models (LLMs) have been widely adopted in a wide range of tasks, leading to increasing attention towards the research on how scaling LLMs affects their performance. Existing works, termed Scaling Laws, have discovered that the final test loss of LLMs scales as power-laws wit

2025

Test-Time Graph Neural Dataset Search With Generative Projection

ICML 2025poster

In this work, we address the test-time adaptation challenge in graph neural networks (GNNs), focusing on overcoming the limitations in flexibility and generalization inherent in existing data-centric approaches. To this end, we propose a novel research problem, test-time graph neural dataset search,…

Cited by 0SourcePDFScholar
2025

Towards Unsupervised Training of Matching-based Graph Edit Distance Solver via Preference-aware GAN

NeurIPS 2025poster

Graph Edit Distance (GED) is a fundamental graph similarity metric widely used in various applications. However, computing GED is an NP-hard problem. Recent state-of-the-art hybrid GED solver has shown promising performance by formulating GED as a bipartite graph matching problem, then leveraging a…

Cited by 0SourceScholar
2025

Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression

NeurIPS 2025poster

State-space models (SSMs), particularly Mamba, emerge as an efficient Transformer alternative with linear complexity for long-sequence modeling. Recent empirical works demonstrate Mamba's in-context learning (ICL) capabilities competitive with Transformers, a critical capacity for large foundation m…

Cited by 0SourceScholar
2025

Understanding the Forgetting of (Replay-based) Continual Learning via Feature Learning: Angle Matters

ICML 2025poster

Continual learning (CL) is crucial for advancing human-level intelligence, but its theoretical understanding, especially regarding factors influencing forgetting, is still relatively limited. This work aims to build a unified theoretical framework for understanding CL using feature learning theory.…

Cited by 0SourcePDFScholar
2025

UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets

EMNLP 2025

Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image synthesis. However, progress in unified VLLMs remains constrained by the lack of

2025

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection

CVPR 2025poster

The advancement of Large Vision Language Models (LVLMs) has significantly improved multimodal understanding, yet challenges remain in video reasoning tasks due to the scarcity of high-quality, large-scale datasets. Existing video question-answering (VideoQA) datasets often rely on costly manual anno…

2024

A Fast, Performant, Secure Distributed Training Framework For LLM

ICASSP 2024accepted

The distributed (federated) LLM is an important method for co-training the domain-specific LLM using siloed data. However, maliciously stealing model parameters and data from the server or client side has become an urgent problem to be solved. In this paper, we propose a secure distributed LLM based…

Cited by 0SourceScholar
2024

A Variational Framework for Estimating Continuous Treatment Effects with Measurement Error

ICLR 2024poster

Estimating treatment effects has numerous real-world applications in various fields, such as epidemiology and political science. While much attention has been devoted to addressing the challenge using fully observational data, there has been comparatively limited exploration of this issue in cases w…

Cited by 2SourcePDFScholar
2024

BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

ICML 2024poster

Pretrained large language models (LLMs) exhibit exceptional general language processing capabilities but come with significant demands on memory and computational resources. As a powerful compression technology, binarization can extremely reduce model weights to a mere 1 bit, lowering the expensive…

2024

DaFoEs: Mixing Datasets Towards the Generalization of Vision-State Deep-Learning Force Estimation in Minimally Invasive Robotic Surgery

RA-L 2024

Precisely determining the contact force during safe interaction in Minimally Invasive Robotic Surgery (MIRS) is still an open research challenge. Inspired by post-operative qualitative analysis from surgical videos, the use of cross-modality data driven deep neural network models has been one of the

Cited by 7SourcecodeScholar
2024

Design and Implementation of A Robotized Hand-held Dissector for Endoscopic Pulmonary Endarterectomy

ICRA 2024poster

Severe chronic pulmonary endarterectomy needs a dissector to delicately remove proliferative intima located in the depth of the pulmonary artery. This work proposed a novel endoscopic robotized steerable dissector for this surgery, enabling easier access to curved deep artery branches. The handheld…

Cited by 0SourceScholar
2024

DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction

CVPR 2024poster

Audio-visual saliency prediction can draw support from diverse modality complements but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies denoising diffusion models have shown more promising in unifying task fra…

Cited by 6SourcePDFScholar
2024

Diffusion Models Demand Contrastive Guidance for Adversarial Purification to Advance

ICML 2024poster

In adversarial defense, adversarial purification can be viewed as a special generation task with the purpose to remove adversarial attacks and diffusion models excel in adversarial purification for their strong generative power. With different predetermined generation requirements, various types of…

Cited by 6SourcePDFScholar
2024

Early Diagnosing Parkinson's Disease Via a Deep Learning Model Based on Augmented Facial Expression Data

ICASSP 2024accepted

It is crucial to promptly diagnose potential Parkinson's disease (PD) patients in order to facilitate early treatment and prevent disease progression. In recent years, there has been growing interest in using facial expressions for in-vitro PD diagnosis due to the distinct "masked face" characterist…

Cited by 0SourceScholar
2024

Earthfarsser: Versatile Spatio-Temporal Dynamical Systems Modeling in One Model

AAAI 2024technical

Efficiently modeling spatio-temporal (ST) physical processes and observations presents a challenging problem for the deep learning community. Many recent studies have concentrated on meticulously reconciling various advantages, leading to designed models that are neither simple nor practical. To add…

2024

Enhancing Dual-Target Cross-Domain Recommendation with Federated Privacy-Preserving Learning

IJCAI 2024poster

Recently, dual-target cross-domain recommendation (DTCDR) has been proposed to alleviate the data sparsity problem by sharing the common knowledge across domains simultaneously. However, existing methods often assume that personal data containing abundant identifiable information can be directly acc…

Cited by 2SourcePDFScholar
2024

Federated Learning from Vision-Language Foundation Models: Theoretical Analysis and Method

NeurIPS 2024poster

Integrating pretrained vision-language foundation models like CLIP into federated learning has attracted significant attention for enhancing generalization across diverse tasks. Typically, federated learning of vision-language models employs prompt learning to reduce communication and computational…

2024

Global and Local Prompts Cooperation via Optimal Transport for Federated Learning

CVPR 2024poster

Prompt learning in pretrained visual-language models has shown remarkable flexibility across various downstream tasks. Leveraging its inherent lightweight nature recent research attempted to integrate the powerful pretrained models into federated learning frameworks to simultaneously reduce communic…

2024

Identifiability Analysis of Linear ODE Systems with Hidden Confounders

NeurIPS 2024poster

The identifiability analysis of linear Ordinary Differential Equation (ODE) systems is a necessary prerequisite for making reliable causal inferences about these systems. While identifiability has been well studied in scenarios where the system is fully observable, the conditions for identifiability…

Cited by 0SourcePDFScholar
2024

Learning Large-Factor EM Image Super-Resolution with Generative Priors

CVPR 2024poster

As the mainstream technique for capturing images of biological specimens at nanometer resolution electron microscopy (EM) is extremely time-consuming for scanning wide field-of-view (FOV) specimens. In this paper we investigate a challenging task of large-factor EM image super-resolution (EMSR) whic…

2024

Learning Multiscale Consistency for Self-Supervised Electron Microscopy Instance Segmentation

ICASSP 2024accepted

Electron microscopy (EM) images are notoriously challenging to segment due to their complex structures and lack of effective annotations. Fortunately, large-scale self-supervised pretraining offers a promising solution by allowing us to acquire prior knowledge of cell and subcellular tissue structur…

Cited by 0SourceScholar
2024

On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and Capability

NeurIPS 2024poster

Autoregressively trained transformers have brought a profound revolution to the world, especially with their in-context learning (ICL) ability to address downstream tasks. Recently, several studies suggest that transformers learn a mesa-optimizer during autoregressive (AR) pretraining to implement…

2024

On the Comparison between Multi-modal and Single-modal Contrastive Learning

NeurIPS 2024poster

Multi-modal contrastive learning with language supervision has presented a paradigm shift in modern machine learning. By pre-training on a web-scale dataset, multi-modal contrastive learning can learn high-quality representations that exhibit impressive robustness and transferability. Despite its em…

Cited by 6SourcePDFScholar
2024

Provable and Efficient Dataset Distillation for Kernel Ridge Regression

NeurIPS 2024poster

Deep learning models are now trained on increasingly larger datasets, making it crucial to reduce computational costs and improve data quality. Dataset distillation aims to distill a large dataset into a small synthesized dataset such that models trained on it can achieve similar performance to thos…

2024

Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples

ICML 2024poster

Neural Network-based active learning (NAL) is a cost-effective data selection technique that utilizes neural networks to select and train on a small subset of samples. While existing work successfully develops various effective or theory-justified NAL algorithms, the understanding of the two commonl…

Cited by 3SourcePDFScholar
2024

Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning

NeurIPS 2024poster

Transformer-based large language models (LLMs) have displayed remarkable creative prowess and emergence capabilities. Existing empirical studies have revealed a strong connection between these LLMs' impressive emergence abilities and their in-context learning (ICL) capacity, allowing them to solve n…

Cited by 0SourcePDFScholar
2024

SLTrain: a sparse plus low rank approach for parameter and memory efficient pretraining

NeurIPS 2024poster

Large language models (LLMs) have shown impressive capabilities across various tasks. However, training LLMs from scratch requires significant computational power and extensive memory capacity. Recent studies have explored low-rank structures on weights for efficient fine-tuning in terms of paramete…

2024

Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal Search

AAAI 2024technical

Deep cross-modal hashing technology provides an effective and efficient cross-modal unified representation learning solution for cross-modal search. However, the existing methods neglect the implicit fine-grained multimodal knowledge relations between these modalities such as when the image contains…

Cited by 7SourcePDFScholar
2024

Understanding Convergence and Generalization in Federated Learning through Feature Learning Theory

ICLR 2024poster

Federated Learning (FL) has attracted significant attention as an efficient privacy-preserving approach to distributed learning across multiple clients. Despite extensive empirical research and practical applications, a systematic way to theoretically understand the convergence and generalization pr…

Cited by 16SourcePDFScholar
2024

Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization

NeurIPS 2024poster

Transformers have demonstrated great power in the recent development of large foundational models. In particular, the Vision Transformer (ViT) has brought revolutionary changes to the field of vision, achieving significant accomplishments on the experimental side. However, their theoretical capabili…

Cited by 5SourcePDFScholar
2023

A Soma Segmentation Benchmark in Full Adult Fly Brain

CVPR 2023poster

Neuron reconstruction in a full adult fly brain from high-resolution electron microscopy (EM) data is regarded as a cornerstone for neuroscientists to explore how neurons inspire intelligence. As the central part of neurons, somas in the full brain indicate the origin of neurogenesis and neural func…

2023

Analyzing Generalization of Neural Networks through Loss Path Kernels

NeurIPS 2023poster

Deep neural networks have been increasingly used in real-world applications, making it critical to ensure their ability to adapt to new, unseen data. In this paper, we study the generalization capability of neural networks trained with (stochastic) gradient flow. We establish a new connection betwee…

Cited by 1SourcePDFScholar
2023

Automatic Segmentation of Nasopharyngeal Carcinoma in CT Images Using Dual Attention and Edge Detection

ICASSP 2023accepted

Nasopharyngeal carcinoma (NPC) is a malignant tumor with a high incidence. Accurate segmentation of the tumor region in Computed Tomography (CT) images of NPC is the key to treatment. However, the features of uneven grayscale values and hazy boundaries of NPC regions make accurate NPC segmentation p…

Cited by 0SourceScholar
2023

CASP-Net: Rethinking Video Saliency Prediction From an Audio-Visual Consistency Perceptual Perspective

CVPR 2023poster

Incorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP methods are capable of exploiting semantic correlation between vision and audio modalitie…

2023

Fed-CO$_{2}$: Cooperation of Online and Offline Models for Severe Data Heterogeneity in Federated Learning

NeurIPS 2023poster

Federated Learning (FL) has emerged as a promising distributed learning paradigm that enables multiple clients to learn a global model collaboratively without sharing their private data. However, the effectiveness of FL is highly dependent on the quality of the data that is being used for training.…

2023

Generator Identification for Linear SDEs with Additive and Multiplicative Noise

NeurIPS 2023poster

In this paper, we present conditions for identifying the generator of a linear stochastic differential equation (SDE) from the distribution of its solution process with a given fixed initial state. These identifiability conditions are crucial in causal inference using linear SDEs as they enable the…

Cited by 5SourcePDFScholar
2023

HiViT: A Simpler and More Efficient Design of Hierarchical Vision Transformer

ICLR 2023top-25%

There has been a debate on the choice of plain vs. hierarchical vision transformers, where researchers often believe that the former (e.g., ViT) has a simpler design but the latter (e.g., Swin) enjoys higher recognition accuracy. Recently, the emerge of masked image modeling (MIM), a self-supervised…

2023

Joint Microstrip Selection and Beamforming Design for MmWave Systems with Dynamic Metasurface Antennas

ICASSP 2023accepted

Dynamic metasurface antennas (DMAs) provide a new paradigm to realize large-scale antenna arrays for future wireless systems. In this paper, we study the downlink millimeter wave (mmWave) DMA systems with limited number of radio frequency (RF) chains. By using the specific DMA structure, an equivale…

Cited by 0SourceScholar
2023

Learning Cross-Representation Affinity Consistency for Sparsely Supervised Biomedical Instance Segmentation

ICCV 2023poster

Sparse instance-level supervision has recently been explored to address insufficient annotation in biomedical instance segmentation, which is easier to annotate crowded instances and better preserves instance completeness for 3D volumetric datasets compared to common semi-supervision.In this paper,…

Cited by 8PDFcodeScholar
2023

Learning Steerable Function for Efficient Image Resampling

CVPR 2023poster

Image resampling is a basic technique that is widely employed in daily applications. Existing deep neural networks (DNNs) have made impressive progress in resampling performance. Yet these methods are still not the perfect substitute for interpolation, due to the issues of efficiency and continuous…

Cited by 11SourcePDFScholar
2023

SAFARI: Versatile and Efficient Evaluations for Robustness of Interpretability

ICCV 2023poster

Interpretability of Deep Learning (DL) is a barrier to trustworthy AI. Despite great efforts made by the Explainable AI (XAI) community, explanations lack robustness--indistinguishable input perturbations may lead to different XAI results. Thus, it is vital to assess how robust DL interpretability i…

Cited by 32PDFcodeScholar
2023

Self-Supervised Neuron Segmentation with Multi-Agent Reinforcement Learning

IJCAI 2023poster

The performance of existing supervised neuron segmentation methods is highly dependent on the number of accurate annotations, especially when applied to large scale electron microscopy (EM) data. By extracting semantic information from unlabeled data, self-supervised methods can improve the performa…

2023

Style Projected Clustering for Domain Generalized Semantic Segmentation

CVPR 2023poster

Existing semantic segmentation methods improve generalization capability, by regularizing various images to a canonical feature space. While this process contributes to generalization, it weakens the representation inevitably. In contrast to existing methods, we instead utilize the difference betwee…

Cited by 40SourcePDFScholar
2023

Understanding and Improving Feature Learning for Out-of-Distribution Generalization

NeurIPS 2023poster

A common explanation for the failure of out-of-distribution (OOD) generalization is that the model trained with empirical risk minimization (ERM) learns spurious features instead of invariant features. However, several recent studies challenged this explanation and found that deep networks may have…

Cited by 47SourcePDFScholar
2022

Auto-scaling Vision Transformers without Training

ICLR 2022poster

This work targets automated designing and scaling of Vision Transformers (ViTs). The motivation comes from two pain spots: 1) the lack of efficient and principled methods for designing and scaling ViTs; 2) the tremendous computational cost of training ViT that is much heavier than its convolution co…

2022

Biological Instance Segmentation with a Superpixel-Guided Graph

IJCAI 2022poster

Recent advanced proposal-free instance segmentation methods have made significant progress in biological images. However, existing methods are vulnerable to local imaging artifacts and similar object appearances, resulting in over-merge and over-segmentation. To reduce these two kinds of errors, we…

2022

Deep Active Learning by Leveraging Training Dynamics

NeurIPS 2022accept

Active learning theories and methods have been extensively studied in classical statistical learning settings. However, deep active learning, i.e., active learning with deep learning models, is usually based on empirical criteria without solid theoretical justification, thus suffering from heavy dou…

Cited by 39SourcePDFScholar
2022

Deep Architecture Connectivity Matters for Its Convergence: A Fine-Grained Analysis

NeurIPS 2022accept

Advanced deep neural networks (DNNs), designed by either human or AutoML algorithms, are growing increasingly complex. Diverse operations are connected by complicated connectivity patterns, e.g., various types of skip connections. Those topological compositions are empirically effective and observed…

Cited by 11SourcePDFScholar
2022

Enhancing Adversarial Training With Second-Order Statistics of Weights

CVPR 2022poster

Adversarial training has been shown to be one of the most effective approaches to improve the robustness of deep neural networks. It is formalized as a min-max optimization over model weights and adversarial perturbations, where the weights can be optimized through gradient descent methods like SGD.…

Cited by 72PDFcodeScholar
2022

Exploring Label Hierarchy in a Generative Way for Hierarchical Text Classification

COLING 2022main

Hierarchical Text Classification (HTC), which aims to predict text labels organized in hierarchical space, is a significant task lacking in investigation in natural language processing. Existing methods usually encode the entire hierarchical structure and fail to construct a robust label-dependent m…

2022

Interpreting Operation Selection in Differentiable Architecture Search: A Perspective from Influence-Directed Explanations

NeurIPS 2022accept

The Differentiable ARchiTecture Search (DARTS) has dominated the neural architecture search community due to its search efficiency and simplicity. DARTS leverages continuous relaxation to convert the intractable operation selection problem into a continuous magnitude optimization problem which can b…

Cited by 3SourcePDFScholar
2022

Learning to Model Pixel-Embedded Affinity for Homogeneous Instance Segmentation

AAAI 2022technical

Homogeneous instance segmentation aims to identify each instance in an image where all interested instances belong to the same category, such as plant leaves and microscopic cells. Recently, proposal-free methods, which straightforwardly generate instance-aware information to group pixels into diffe…

2022

MissDAG: Causal Discovery in the Presence of Missing Data with Continuous Additive Noise Models

NeurIPS 2022accept

State-of-the-art causal discovery methods usually assume that the observational data is complete. However, the missing data problem is pervasive in many practical scenarios such as clinical trials, economics, and biology. One straightforward way to address the missing data problem is first to impute…

2022

Stacked Multi-Scale Attention Network for Image Colorization

ICASSP 2022accepted

Deep convolutional networks (CNNs) show their potential in image colorization for producing plausible results. Recently, the attention mechanism further boosts the performances of CNNs by constructing channel and spatial interactions. However, existing attention methods are performed in a single-sca…

Cited by 0SourceScholar
2022

Towards Deepening Graph Neural Networks: A GNTK-based Optimization Perspective

ICLR 2022poster

Graph convolutional networks (GCNs) and their variants have achieved great success in dealing with graph-structured data. Nevertheless, it is well known that deep GCNs suffer from the over-smoothing problem, where node representations tend to be indistinguishable as more layers are stacked up. The t…

Cited by 33SourcePDFScholar
2022

Weighted Mutual Learning with Diversity-Driven Model Compression

NeurIPS 2022accept

Online distillation attracts attention from the community as it simplifies the traditional two-stage knowledge distillation process into a single stage. Online distillation collaboratively trains a group of peer models, which are treated as students, and all students gain extra knowledge from each o…

Cited by 10SourcePDFScholar
2021

BayLIME: Bayesian local interpretable model-agnostic explanations

UAI 2021poster

Given the pressing need for assuring algorithmic transparency, Explainable AI (XAI) has emerged as one of the key areas of AI research. In this paper, we develop a novel Bayesian extension to the LIME framework, one of the most widely used approaches in XAI – which we call BayLIME. Compared to LIME,…

2021

Conformer: Local Features Coupling Global Representations for Visual Recognition

ICCV 2021poster

Within Convolutional Neural Network (CNN), the convolution operations are good at extracting local features but experience difficulty to capture global representations. Within visual transformer, the cascaded self-attention modules can capture long-distance feature dependencies but unfortunately det…

Cited by 890PDFcodeScholar
2021

DeFLOCNet: Deep Image Editing via Flexible Low-Level Controls

CVPR 2021poster

User-intended visual content fills the hole regions of an input image in the image editing scenario. The coarse lowlevel inputs, which typically consist of sparse sketch lines and color dots, convey user intentions for content creation (i.e., free-form editing). While existing methods combine an inp…

Cited by 42PDFcodeScholar
2021

Item Response Ranking for Cognitive Diagnosis

IJCAI 2021poster

Cognitive diagnosis, a fundamental task in education area, aims at providing an approach to reveal the proficiency level of students on knowledge concepts. Actually, monotonicity is one of the basic conditions in cognitive diagnosis theory, which assumes that student's proficiency is monotonic with…

Cited by 34SourcePDFScholar
2021

On the Equivalence between Neural Network and Support Vector Machine

NeurIPS 2021poster

Recent research shows that the dynamics of an infinitely wide neural network (NN) trained by gradient descent can be characterized by Neural Tangent Kernel (NTK) \citep{jacot2018neural}. Under the squared loss, the infinite-width NN trained by gradient descent with an infinitely small learning rate…

2021

On the Neural Tangent Kernel of Deep Networks with Orthogonal Initialization

IJCAI 2021poster

The prevailing thinking is that orthogonal weights are crucial to enforcing dynamical isometry and speeding up training. The increase in learning speed that results from orthogonal initialization in linear networks has been well-proven. However, while the same is believed to also hold for nonlinear…

Cited by 41SourcePDFScholar
2021

PD-GAN: Probabilistic Diverse GAN for Image Inpainting

CVPR 2021poster

We propose PD-GAN, a probabilistic diverse GAN forimage inpainting. Given an input image with arbitrary holeregions, PD-GAN produces multiple inpainting results withdiverse and visually realistic content. Our PD-GAN is builtupon a vanilla GAN which generates images based on random noise. During imag…

Cited by 287PDFcodeScholar
2020

KALM: Key Area Localization Mechanism for Abnormality Detection in Musculoskeletal Radiographs

ICASSP 2020accepted

Recently abnormality detection in musculoskeletal radio-graphs has attracted many attentions. For abnormality detection, it is crucial to locate the most important area in the musculoskeletal radiographs. To achieve this goal, we propose a key area localization mechanism (KALM) for abnormality detec…

Cited by 0SourceScholar
2020

Practical Verification of Neural Network Enabled State Estimation System for Robotics

IROS 2020poster

We study for the first time the verification problem on learning-enabled state estimation systems for robotics, which use Bayes filter for localisation, and use deep neural network to process sensory input into observations for the Bayes filter. Specifically, we are interested in a robustness proper…

Cited by 7SourceScholar
2020

Rethinking Image Inpainting via a Mutual Encoder-Decoder with Feature Equalizations

ECCV 2020poster

Deep encoder-decoder based CNNs have advanced image inpainting methods for hole filling. While existing methods recover structures and textures step-by-step in the hole regions, they typically use two encoder-decoders for separate recovery. The CNN features of each encoder are learned to capture eit…