← Search

Shudong Huang

21 accepted papers

2026

Intra-Modal Neighbors Never Lie: Rectifying Inter-Modal Noisy Correspondence via Graph-Based Intra-Modal Reasoning

ICML 2026poster

Large-scale web-harvested datasets have fueled the progress of cross-modal retrieval but inevitably suffer from \textit{noisy correspondence}, which severely degrades model generalization. Existing methods primarily address this by filtering out noise or seeking a substitute label, yet they predomin…

Cited by 0SourceScholar
2026

On the Power of Statistics in Class-Incremental Learning with Pretrained Models

ICML 2026poster

Recent class-incremental learning (CIL) methods built on large pre-trained vision models have shown that strong performance can be retained even under strict data access constraints. This raises a fundamental question: which properties of pre-trained representations make such recovery possible in th…

Cited by 0SourceScholar
2026

PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence Learning

AAAI 2026technical

Cross-modal retrieval aims to align different modalities via semantic similarity. However, existing methods often assume that image-text pairs are perfectly aligned, overlooking Noisy Correspondences in real data. These misaligned pairs misguide similarity learning and degrade retrieval performance.

Cited by 0SourcePDFScholar
2026

S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing

AAAI 2026technical

Dynamic Vision Sensor (DVS) asynchronously records sparse events triggered by changes in pixel intensity, offering high temporal resolution and low latency. Existing frame-based methods process event data densely, violating its inherent sparsity and introducing computational redundancy. While asynch

Cited by 0SourcePDFScholar
2026

Towards Safe and Optimal Online Bidding: A Modular Look-ahead Lyapunov Framework

ICLR 2026poster

This paper studies online bidding subject to simultaneous budget and return-on-investment (ROI) constraints, which encodes the goal of balancing high volume and profitability. We formulate the problem as a general constrained online learning problem that can be applied to diverse bidding settings (e…

Cited by 0SourceScholar
2025

Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching

ICCV 2025poster

Enabling Visual Semantic Models to effectively handle multi-view description matching has been a longstanding challenge. Existing methods typically learn a set of embeddings to find the optimal match for each view's text and compute similarity. However, the visual and text embeddings learned through…

2025

Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment

AAAI 2025technical

Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain textual information from multiple different views, which makes it…

2025

DONIS: Importance Sampling for Training Physics-Informed DeepONet

IJCAI 2025

Deep Operator Network (DeepONet) effectively learns complex operator mappings, especially for systems governed by differential equations. Physics-informed DeepONet (PI-DeepONet) extends these capabilities by integrating physical constraints, enabling robust performance with limited or no labeled dat

2025

Learning Dynamic Similarity by Bidirectional Hierarchical Sliding Semantic Probe for Efficient Text Video Retrieval

AAAI 2025technical

Text-video retrieval is a foundation task in multi-modal research which aims to align texts and videos in the embedding space. The key challenge is to learn the similarity between videos and texts. A conventional approach involves directly aligning video-text pairs using cosine similarity. However,…

Cited by 0SourcePDFScholar
2025

Multi-view Granular-ball Contrastive Clustering

AAAI 2025technical

Previous multi-view contrastive learning methods typically operate at two scales: instance-level and cluster-level. The former generally constructs positive and negative pairs based on the correspondence between samples and view instances. These methods aim to bring positive pairs closer and push…

2024

An Effective Augmented Lagrangian Method for Fine-Grained Multi-View Optimization

AAAI 2024technical

The significance of multi-view learning in effectively mitigating the intricate intricacies entrenched within heterogeneous data has garnered substantial attention in recent years. Notwithstanding the favorable achievements showcased by recent strides in this area, a confluence of noteworthy challen…

Cited by 3SourcePDFScholar
2024

Multi-View Clustering by Inter-cluster Connectivity Guided Reward

ICML 2024poster

Multi-view clustering has been widely explored for its effectiveness in harmonizing heterogeneity along with consistency in different views of data. Despite the significant progress made by recent works, the performance of most existing methods is heavily reliant on strong priori information regardi…

Cited by 1SourcePDFScholar
2024

Robust Contrastive Multi-view Kernel Clustering

IJCAI 2024poster

Multi-view kernel clustering (MKC) aims to fully reveal the consistency and complementarity of multiple views in a potential Hilbert space, thereby enhancing clustering performance. The clustering results of most MKC methods are highly sensitive to the quality of the constructed kernels, as traditio…

2024

With a Little Help from Language: Semantic Enhanced Visual Prototype Framework for Few-Shot Learning

IJCAI 2024poster

Few-shot learning (FSL) aims to recognize new categories given limited training samples. The core challenge is to avoid overfitting to the minimal data while ensuring good generalization to novel classes. One mainstream method employs prototypes from visual feature extractors as classifier weight an…

Cited by 0SourcePDFScholar
2023

PRIOR: Personalized Prior for Reactivating the Information Overlooked in Federated Learning.

NeurIPS 2023poster

Classical federated learning (FL) enables training machine learning models without sharing data for privacy preservation, but heterogeneous data characteristic degrades the performance of the localized model. Personalized FL (PFL) addresses this by synthesizing personalized models from a global mode…

2023

Self-Supervised Graph Attention Networks for Deep Weighted Multi-View Clustering

AAAI 2023technical

As one of the most important research topics in the unsupervised learning field, Multi-View Clustering (MVC) has been widely studied in the past decade and numerous MVC methods have been developed. Among these methods, the recently emerged Graph Neural Networks (GNN) shine a light on modeling both t…

Cited by 42SourcePDFScholar
2022

Multi-View Clustering on Topological Manifold

AAAI 2022technical

Multi-view clustering has received a lot of attentions in data mining recently. Though plenty of works have been investigated on this topic, it is still a severe challenge due to the complex nature of the multiple heterogeneous features. Particularly, existing multi-view clustering algorithms fail t…

Cited by 27SourcePDFScholar
2022

Multi-view Subspace Clustering on Topological Manifold

NeurIPS 2022accept

Multi-view subspace clustering aims to exploit a common affinity representation by means of self-expression. Plenty of works have been presented to boost the clustering performance, yet seldom considering the topological structure in data, which is crucial for clustering data on manifold. Orthogonal…

Cited by 31SourcePDFScholar