← Search

Heng Huang

169 accepted papers

2026

Beyond Buffer Limits: Energy-Based Data Reassembly for Continual Learning

ICML 2026poster

Continual learning (CL) aims to acquire new knowledge from a non-stationary data stream while retaining performance on previously learned tasks. Memory-based replay methods mitigate catastrophic forgetting by storing and revisiting past samples, but their effectiveness is fundamentally constrained b…

Cited by 0SourceScholar
2026

Catalog-Native LLM: Speaking Item-ID dialect with Less Entanglement for Recommendation

ICLR 2026poster

While collaborative filtering delivers predictive accuracy and efficiency, and Large Language Models (LLMs) enable expressive and generalizable reasoning, modern recommendation systems must bring these strengths together. Growing user expectations, such as natural-language queries and transparent ex…

Cited by 0SourceScholar
2026

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

ICML 2026poster

As policy catches up with the capabilities of generative AI, watermarking is central to content provenance efforts. Inference-time watermarks for autoregressive models are unfit for continuous modalities due to discretization inconsistencies. Existing methods overcome this by finetuning the modality…

Cited by 0SourceScholar
2026

Learning to Reason via Mixture-of-Thought for Logical Reasoning

ICLR 2026poster

Human beings naturally utilize multiple reasoning modalities to learn and solve logical problems, i.e., different representational formats such as natural language, code, and symbolic logic. In contrast, most existing LLM-based approaches operate with a single reasoning modality during training, typ…

Cited by 0SourcecodeScholar
2026

Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following

CVPR 2026

Large multimodal models (LMMs) are increasingly adopted as judges in multimodal evaluation systems due to their strong instruction following and consistency with human preferences. However, their ability to follow diverse, fine-grained evaluation criteria remains underexplored. We develop Multi-Crit

Cited by 0SourcecodeScholar
2026

New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis Framework

ICLR 2026poster

Fine-tuning Large Language Models (LLMs) typically involves either full fine-tuning, which updates all model parameters, or Parameter-Efficient Fine-Tuning (PEFT), which adjusts a small subset of parameters. However, both approaches have inherent limitations: full fine-tuning is computationally expe…

Cited by 0SourceScholar
2026

Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing

ICML 2026poster

Parallel thinking has emerged as a promising paradigm for reasoning, yet it imposes significant computational burdens. Existing efficiency methods primarily rely on local, per-trajectory signals and lack principled mechanisms to exploit global dynamics across parallel branches. We introduce 2D probi…

Cited by 0SourceScholar
2026

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

ICLR 2026poster

Parallel thinking has emerged as a novel approach for enhancing the reasoning capabilities of large language models (LLMs) by exploring multiple reasoning paths concurrently. However, activating such capabilities through training remains challenging. Existing methods mainly rely on supervised fine-t…

Cited by 0SourcecodeScholar
2026

RNA-FM: Flow-Matching Generative Model for Genome-wide RNA-Seq Prediction

ICML 2026poster

Histopathology whole-slide images (WSIs) are routinely acquired in clinical practice and contain rich tissue morphology but lack direct molecular architecture and functional programs defining pathological states, whereas RNA sequencing (RNA-seq) provides genome-wide transcriptional profiles at subst…

Cited by 0SourceScholar
2026

Riemannian Zeroth-Order Gradient Estimation with Structure-Preserving Metrics for Geodesically Incomplete Manifolds

ICLR 2026poster

In this paper, we study Riemannian zeroth-order optimization in settings where the underlying Riemannian metric $g$ is geodesically incomplete, and the goal is to approximate stationary points with respect to this incomplete metric. To address this challenge, we construct structure-preserving metric…

Cited by 0SourceScholar
2026

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation

CVPR 2026

The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional generation quality metrics to encompass aesthetic appeal. However, existing benchmarks remain largely focused on technical fidelity, leaving a sig

Cited by 0SourceScholar
2026

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

IJCAI 2026

The rapid advancement of AIGC video generation calls for evaluation frameworks that move beyond technical fidelity and incorporate human-centered aesthetic assessment. Existing benchmarks often overlook fine-grained perceptual qualities such as visual aesthetics, artistic style, and human preference

Cited by 0Scholar
2025

A Watermark for Order-Agnostic Language Models

ICLR 2025poster

Statistical watermarking techniques are well-established for sequentially decoded language models (LMs). However, these techniques cannot be directly applied to order-agnostic LMs, as the tokens in order-agnostic LMs are not generated sequentially. In this work, we introduce PATTERN-MARK, a pattern-…

Cited by 2SourcePDFScholar
2025

ARGUS: Hallucination and Omission Evaluation in Video-LLMs

ICCV 2025poster

Video large language models have not yet been widely deployed, largely due to their tendency to hallucinate. Typical benchmarks for Video-LLMs rely simply on multiple choice questions. Unfortunately, VideoLLMs hallucinate far more aggressively on freeform text generation tasks like video captioning…

2025

Asymmetric Conflict and Synergy in Post-training for LLM-based Multilingual Machine Translation

ACL 2025finding

The emergence of Large Language Models (LLMs) has advanced the multilingual machine translation (MMT), yet the Curse of Multilinguality (CoM) remains a major challenge. Existing work in LLM-based MMT typically mitigates this issue via scaling up training and computation budget, which raises a critic…

Cited by 0SourcePDFScholar
2025

Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models

CVPR 2025poster

Recent advances in diffusion generative models have yielded remarkable progress. While the quality of generated content continues to improve, these models have grown considerably in size and complexity. This increasing computational burden poses significant challenges, particularly in resource-const…

2025

Federated Continuous Category Discovery and Learning

ICCV 2025poster

Federated Learning (FL) studies often assume a static data distribution, whereas real-world scenarios involve dynamic changes. To address this gap, we study Federated Continuous Category Discovery and Learning (FC^2DL), an essential yet underexplored problem that enables FL models to evolve continuo…

Cited by 0SourcePDFScholar
2025

From Lists to Emojis: How Format Bias Affects Model Alignment

ACL 2025long

In this paper, we study format biases in reinforcement learning from human feedback (RLHF). We observe that many widely-used preference models—including human evaluators, GPT-4, and top-ranking models on the RewardBench benchmark—exhibit strong biases towards specific format patterns, such as lists,…

2025

GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning

ICCV 2025poster

Recent advances have shown that video generation models can enhance robot learning by deriving effective robot actions through inverse dynamics. However, these methods heavily depend on the quality of generated data and struggle with fine-grained manipulation due to the lack of environment feedback.…

2025

KLMN: Knowledge distillation based lightweight multi-clue image forgery detection and localization

ICASSP 2025accepted

Current image forensics methods often utilize image features from various frequency domains. However, the effective use of these features frequently depends on complex network architectures and a large number of parameters. In this paper, we introduce a lightweight Multi-Clue image forgery detection…

Cited by 0SourceScholar
2025

LLaVA-Critic: Learning to Evaluate Multimodal Models

CVPR 2025poster

We introduce LLaVA-Critic, the first open-source large multimodal model (LMM) designed as a generalist evaluator to assess performance across a wide range of multimodal tasks. LLaVA-Critic is trained using a high-quality critic instruction-following dataset that incorporates diverse evaluation crite…

Cited by 53SourcePDFScholar
2025

Not All Prompts Are Made Equal: Prompt-based Pruning of Text-to-Image Diffusion Models

ICLR 2025poster

Text-to-image (T2I) diffusion models have demonstrated impressive image generation capabilities. Still, their computational intensity prohibits resource-constrained organizations from deploying T2I models after fine-tuning them on their internal *target* data. While pruning techniques offer a potent…

2025

OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities

ICLR 2025poster

We introduce \textbf{OmnixR}, an evaluation suite designed to benchmark state-of-the-art Omni-modality Language Models (OLMs), such as GPT-4o and Gemini. Evaluating OLMs, which integrate multiple modalities such as text, vision, and audio, presents unique challenges. Particularly, the user message…

Cited by 5SourcePDFScholar
2025

Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations

ICLR 2025spotlight

In this paper, we explore the two-point zeroth-order gradient estimator and identify the distribution of random perturbations that minimizes the estimator's asymptotic variance as the perturbation stepsize tends to zero. We formulate it as a constrained functional optimization problem over the space…

Cited by 1SourcePDFScholar
2025

Robust Distortion-Free Watermark for Autoregressive Audio Generation Models

NeurIPS 2025poster

The rapid advancement of next-token-prediction models has led to widespread adoption across modalities, enabling the creation of realistic synthetic media. In the audio domain, while autoregressive speech models have propelled conversational interactions forward, the potential for misuse, such as im…

Cited by 0SourceScholar
2025

Robust Reinforcement Learning in Finance: Modeling Market Impact with Elliptic Uncertainty Sets

NeurIPS 2025poster

In financial applications, reinforcement learning (RL) agents are commonly trained on historical data, where their actions do not influence prices. However, during deployment, these agents trade in live markets where their own transactions can shift asset prices, a phenomenon known as market impact.…

Cited by 0SourceScholar
2025

SleeperMark: Towards Robust Watermark against Fine-Tuning Text-to-image Diffusion Models

CVPR 2025poster

Recent advances in large-scale text-to-image (T2I) diffusion models have enabled a variety of downstream applications. As T2I models require extensive resources for training, they constitute highly valued intellectual property (IP) for their legitimate owners, yet making them incentive targets for u…

2025

Towards Optimal Multi-draft Speculative Decoding

ICLR 2025poster

Large Language Models (LLMs) have become an indispensable part of natural language processing tasks. However, autoregressive sampling has become an efficiency bottleneck. Multi-Draft Speculative Decoding (MDSD) is a recent approach where, when generating each token, a small draft model generates mul…

Cited by 2SourcePDFScholar
2025

VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations

NeurIPS 2025poster

Video aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and multimodal fusion challenges hinder direct application of image-bas…

Cited by 0SourcecodeScholar
2025

Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language Models

EMNLP 2025

The protection of cyber Intellectual Property (IP) such as web content is an increasingly critical concern. The rise of large language models (LLMs) with online retrieval capabilities enables convenient access to information but often undermines the rights of original content creators. As users incr

Cited by 0SourcePDFScholar
2024

A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution

EMNLP 2024main

Authorship attribution aims to identify the origin or author of a document. Traditional approaches have heavily relied on manual features and fail to capture long-range correlations, limiting their effectiveness. Recent advancements leverage text embeddings from pre-trained language models, which re…

Cited by 0SourcePDFScholar
2024

A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models

ICML 2024poster

Watermarking techniques offer a promising way to identify machine-generated content via embedding covert information into the contents generated from language models. A challenge in the domain lies in preserving the distribution of original generated content after watermarking. Our research extends…

2024

APDDv2: Aesthetics of Paintings and Drawings Dataset with Artist Labeled Scores and Comments

NeurIPS 2024poster

Datasets play a pivotal role in training visual models, facilitating the development of abstract understandings of visual features through diverse image samples and multidimensional attributes. However, in the realm of aesthetic evaluation of artistic images, datasets remain relatively scarce. Exist…

2024

AlpaGasus: Training a Better Alpaca with Fewer Data

ICLR 2024poster

Large language models~(LLMs) strengthen instruction-following capability through instruction-finetuning (IFT) on supervised instruction/response data. However, widely used IFT datasets (e.g., Alpaca's 52k data) surprisingly contain many low-quality instances with incorrect or irrelevant responses, w…

2024

Auto-Train-Once: Controller Network Guided Automatic Network Pruning from Scratch

CVPR 2024poster

Current techniques for deep neural network (DNN) pruning often involve intricate multi-step processes that require domain-specific expertise making their widespread adoption challenging. To address the limitation the Only-Train-Once (OTO) and OTOv2 are proposed to eliminate the need for additional f…

2024

BilevelPruning: Unified Dynamic and Static Channel Pruning for Convolutional Neural Networks

CVPR 2024poster

Most existing dynamic or runtime channel pruning methods have to store all weights to achieve efficient inference which brings extra storage costs. Static pruning methods can reduce storage costs directly but their performance is limited by using a fixed sub-network to approximate the original model…

Cited by 5SourcePDFScholar
2024

Compressing Image-to-Image Translation GANs Using Local Density Structures on Their Learned Manifold

AAAI 2024technical

Generative Adversarial Networks (GANs) have shown remarkable success in modeling complex data distributions for image-to-image translation. Still, their high computational demands prohibit their deployment in practical scenarios like edge devices. Existing GAN compression methods mainly rely on know…

Cited by 10SourcePDFScholar
2024

Event3DGS: Event-Based 3D Gaussian Splatting for High-Speed Robot Egomotion

CoRL 2024poster

By combining differentiable rendering with explicit point-based scene representations, 3D Gaussian Splatting (3DGS) has demonstrated breakthrough 3D reconstruction capabilities. However, to date 3DGS has had limited impact on robotics, where high-speed egomotion is pervasive: Egomotion introduc…

Cited by 11SourceScholar
2024

Inevitable Trade-off between Watermark Strength and Speculative Sampling Efficiency for Language Models

NeurIPS 2024poster

Large language models are probabilistic models, and the process of generating content is essentially sampling from the output distribution of the language model. Existing watermarking techniques inject watermarks into the generated content without altering the output quality. On the other hand, exis…

Cited by 1SourcePDFScholar
2024

InstructZero: Efficient Instruction Optimization for Black-Box Large Language Models

ICML 2024poster

Large language models (LLMs) are instruction followers but the performance varies under different instructions. It is challenging to create the best instruction, especially for black-box LLMs on which backpropagation is forbidden. Instead of directly optimizing the discrete instruction, we optimize…

2024

Jointly Training and Pruning CNNs via Learnable Agent Guidance and Alignment

CVPR 2024poster

Structural model pruning is a prominent approach used for reducing the computational cost of Convolutional Neural Networks (CNNs) before their deployment on resource-constrained devices. Yet the majority of proposed ideas require a pretrained model before pruning which is costly to secure. In this p…

Cited by 5SourcePDFScholar
2024

Learning Sampling Policy to Achieve Fewer Queries for Zeroth-Order Optimization

AISTATS 2024poster

Zeroth-order (ZO) methods, which use the finite difference of two function evaluations (also called ZO gradient) to approximate first-order gradient, have attracted much attention recently in machine learning because of their broad applications. The accuracy of the ZO gradient highly depends on how…

Cited by 0SourcePDFScholar
2024

Mixture of Efficient Diffusion Experts Through Automatic Interval and Sub-Network Selection

ECCV 2024poster

"Diffusion probabilistic models can generate high-quality samples. Yet, their sampling process requires numerous denoising steps, making it slow and computationally intensive. We propose to reduce the sampling cost by pruning a pretrained diffusion model into a mixture of efficient experts. First, w…

2024

ODIN: Disentangled Reward Mitigates Hacking in RLHF

ICML 2024poster

In this work, we study the issue of reward hacking on the response length, a challenge emerging in Reinforcement Learning from Human Feedback (RLHF) on LLMs. A well-formatted, verbose but less helpful response from the LLMs can often deceive LLMs or even human evaluators and achieve high scores. The…

Cited by 57SourcePDFScholar
2024

Paintings and Drawings Aesthetics Assessment with Rich Attributes for Various Artistic Categories

IJCAI 2024poster

Image aesthetic evaluation is a highly prominent research domain in the field of computer vision. In recent years, there has been a proliferation of datasets and corresponding evaluation methodologies for assessing the aesthetic quality of photographic works, leading to the establishment of a relati…

2024

Prompting Language-Informed Distribution for Compositional Zero-Shot Learning

ECCV 2024poster

"Compositional zero-shot learning (CZSL) task aims to recognize unseen compositional visual concepts, , sliced tomatoes, where the model is learned only from the seen compositions, , sliced potatoes and red tomatoes. Thanks to the prompt tuning on large pre-trained visual language models such as CLI…

2024

Retrieval Across Any Domains via Large-scale Pre-trained Model

ICML 2024poster

In order to enhance the generalization ability towards unseen domains, universal cross-domain image retrieval methods require a training dataset encompassing diverse domains, which is costly to assemble. Given this constraint, we introduce a novel problem of data-free adaptive cross-domain retrieval…

Cited by 0SourcePDFScholar
2024

Revisiting Adaptive Cellular Recognition Under Domain Shifts: A Contextual Correspondence View

ECCV 2024oral

"Cellular nuclei recognition serves as a fundamental and essential step in the workflow of digital pathology. However, with disparate source organs and staining procedures among histology image clusters, the scanned tiles inherently conform to a non-uniform data distribution, which induces deteriora…

2024

Seeing Unseen: Discover Novel Biomedical Concepts via Geometry-Constrained Probabilistic Modeling

CVPR 2024poster

Machine learning holds tremendous promise for transforming the fundamental practice of scientific discovery by virtue of its data-driven nature. With the ever-increasing stream of research data collection it would be appealing to autonomously explore patterns and insights from observational data for…

Cited by 6SourcePDFScholar
2024

Towards Green AI in Fine-tuning Large Language Models via Adaptive Backpropagation

ICLR 2024poster

Fine-tuning is essential to adapting pre-trained large language models to downstream applications. With the increasing popularity of LLM-enabled applications, fine-tuning has been performed intensively worldwide, incurring a tremendous amount of computing costs that correspond to big carbon footprin…

2024

Unbiased Watermark for Large Language Models

ICLR 2024spotlight

The recent advancements in large language models (LLMs) have sparked a growing apprehension regarding the potential misuse. One approach to mitigating this risk is to incorporate watermarking techniques into LLMs, allowing for the tracking and attribution of model outputs. This study examines a cruc…

Cited by 129SourcePDFScholar
2024

Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection

ACL 2024findings

Data selection in instruction tuning emerges as a pivotal process for acquiring high-quality data and training instruction-following large language models (LLMs), but it is still a new and unexplored research area for vision-language models (VLMs). Existing data selection approaches on LLMs either r…

2024

ZeroMark: Towards Dataset Ownership Verification without Disclosing Watermark

NeurIPS 2024poster

High-quality public datasets significantly prompt the prosperity of deep neural networks (DNNs). Currently, dataset ownership verification (DOV), which consists of dataset watermarking and ownership verification, is the only feasible solution to protect their copyright by preventing unauthorized use…

2023

Adversarial Weight Perturbation Improves Generalization in Graph Neural Networks

AAAI 2023technical

A lot of theoretical and empirical evidence shows that the flatter local minima tend to improve generalization. Adversarial Weight Perturbation (AWP) is an emerging technique to efficiently and effectively find such minima. In AMP we minimize the loss w.r.t. a bounded worst-case perturbation of the…

2023

Communication-Efficient Federated Bilevel Optimization with Global and Local Lower Level Problems

NeurIPS 2023poster

Bilevel Optimization has witnessed notable progress recently with new emerging efficient algorithms. However, its application in the Federated Learning setting remains relatively underexplored, and the impact of Federated Learning's inherent challenges on the convergence of bilevel algorithms remain…

Cited by 11SourcePDFScholar
2023

Cooperation or Competition: Avoiding Player Domination for Multi-Target Robustness via Adaptive Budgets

CVPR 2023poster

Despite incredible advances, deep learning has been shown to be susceptible to adversarial attacks. Numerous approaches were proposed to train robust networks both empirically and certifiably. However, most of them defend against only a single type of attack, while recent work steps forward at defen…

Cited by 2SourcePDFScholar
2023

Domain Watermark: Effective and Harmless Dataset Copyright Protection is Closed at Hand

NeurIPS 2023poster

The prosperity of deep neural networks (DNNs) is largely benefited from open-source datasets, based on which users can evaluate and improve their methods. In this paper, we revisit backdoor-based dataset ownership verification (DOV), which is currently the only feasible approach to protect the copyr…

2023

EffConv: Efficient Learning of Kernel Sizes for Convolution Layers of CNNs

AAAI 2023technical

Determining kernel sizes of a CNN model is a crucial and non-trivial design choice and significantly impacts its performance. The majority of kernel size design methods rely on complex heuristic tricks or leverage neural architecture search that requires extreme computational resources. Thus, learni…

2023

Faster Fair Machine via Transferring Fairness Constraints to Virtual Samples

AAAI 2023technical

Fair classification is an emerging and important research topic in machine learning community. Existing methods usually formulate the fairness metrics as additional inequality constraints, and then embed them into the original objective. This makes fair classification problems unable to be effective…

Cited by 0SourcePDFScholar
2023

Federated Conditional Stochastic Optimization

NeurIPS 2023poster

Conditional stochastic optimization has found applications in a wide range of machine learning tasks, such as invariant learning, AUPRC maximization, and meta-learning. As the demand for training models with large-scale distributed data grows in these applications, there is an increasing need for co…

Cited by 12SourcePDFScholar
2023

Learning to Jointly Share and Prune Weights for Grounding Based Vision and Language Models

ICLR 2023poster

Transformers have seen growing interest in processing different modalities, including language and image data. As a result, we can process vision and language data using transformers that are architecturally similar. Leveraging this feature of transformers, we propose weight sharing across two tran…

Cited by 10SourcePDFScholar
2023

Learning with Diversity: Self-Expanded Equalization for Better Generalized Deep Metric Learning

ICCV 2023poster

Exploring good generalization ability is essential in deep metric learning (DML). Most existing DML methods focus on improving the model robustness against category shift to keep the performance on unseen categories. However, in addition to category shift, domain shift also widely exists in real-wor…

Cited by 8PDFScholar
2023

PTP: Boosting Stability and Performance of Prompt Tuning with Perturbation-Based Regularizer

EMNLP 2023long main

Recent studies show that prompt tuning can better leverage the power of large language models than fine-tuning on downstream natural language understanding tasks. However, the existing prompt tuning methods have training instability issues, as the variance of scores under different random seeds is q…

Cited by 0SourceScholar
2023

Resolving the Tug-of-War: A Separation of Communication and Learning in Federated Learning

NeurIPS 2023poster

Federated learning (FL) is a promising privacy-preserving machine learning paradigm over distributed data. In this paradigm, each client trains the parameter of a model locally and the server aggregates the parameter from clients periodically. Therefore, we perform the learning and communication ove…

Cited by 2SourcePDFScholar
2023

Solving a Class of Non-Convex Minimax Optimization in Federated Learning

NeurIPS 2023poster

The minimax problems arise throughout machine learning applications, ranging from adversarial training and policy evaluation in reinforcement learning to AUROC maximization. To address the large-scale distributed data challenges across multiple clients with communication-efficient distributed traini…

2023

Structural Alignment for Network Pruning through Partial Regularization

ICCV 2023poster

In this paper, we propose a novel channel pruning method to reduce the computational and storage costs of Convolutional Neural Networks (CNNs). Many existing one-shot pruning methods directly remove redundant structures, which brings a huge gap between the model before and after network pruning. Thi…

Cited by 18PDFScholar
2023

Taxonomy Adaptive Cross-Domain Adaptation in Medical Imaging via Optimization Trajectory Distillation

ICCV 2023poster

The success of automated medical image analysis depends on large-scale and expert-annotated training sets. Unsupervised domain adaptation (UDA) has been raised as a promising approach to alleviate the burden of labeled data collection. However, they generally operate under the closed-set adaptation…

Cited by 15PDFcodeScholar
2022

Closing the Generalization Gap of Cross-Silo Federated Medical Image Segmentation

CVPR 2022poster

Cross-silo federated learning (FL) has attracted much attention in medical imaging analysis with deep learning in recent years as it can resolve the critical issues of insufficient data, data privacy, and training efficiency. However, there can be a generalization gap between the model trained from…

Cited by 84PDFcodeScholar
2022

Doubly Sparse Asynchronous Learning for Stochastic Composite Optimization

IJCAI 2022poster

Parallel optimization has become popular for large-scale learning in the past decades. However, existing methods suffer from huge computational costs, memory usage, and communication burden in high-dimensional scenarios. To address the challenges, we propose a new accelerated doubly sparse asynchron…

Cited by 0SourcePDFScholar
2022

Interpretations Steered Network Pruning via Amortized Inferred Saliency Maps

ECCV 2022poster

"Convolutional Neural Networks (CNNs) compression is crucial to deploying these models in edge devices with limited resources. Existing channel pruning algorithms for CNNs have achieved plenty of success on complex models. They approach the pruning problem from various perspectives and use different…

2022

Learning Universal Adversarial Perturbation by Adversarial Example

AAAI 2022technical

Deep learning models have shown to be susceptible to universal adversarial perturbation (UAP), which has aroused wide concerns in the community. Compared with the conventional adversarial attacks that generate adversarial samples at the instance level, UAP can fool the target model for different ins…

2022

MetricFormer: A Unified Perspective of Correlation Exploring in Similarity Learning

NeurIPS 2022accept

Similarity learning can be significantly advanced by informative relationships among different samples and features. The current methods try to excavate the multiple correlations in different aspects, but cannot integrate them into a unified framework. In this paper, we provide to consider the multi…

Cited by 9SourcePDFScholar
2022

Noise Is Also Useful: Negative Correlation-Steered Latent Contrastive Learning

CVPR 2022poster

How to effectively handle label noise has been one of the most practical but challenging tasks in Deep Neural Networks (DNNs). Recent popular methods for training DNNs with noisy labels mainly focus on directly filtering out samples with low confidence or repeatedly mining valuable information from…

Cited by 27PDFScholar
2022

On the Convergence of Local Stochastic Compositional Gradient Descent with Momentum

ICML 2022spotlight

Federated Learning has been actively studied due to its efficiency in numerous real-world applications in the past few years. However, the federated stochastic compositional optimization problem is still underexplored, even though it has widespread applications in machine learning. In this paper, we…

Cited by 22SourcePDFScholar
2022

Recover Fair Deep Classification Models via Altering Pre-trained Structure

ECCV 2022poster

"There have been growing interest in algorithmic fairness for biased data. Although various pre-, in-, and post-processing methods are designed to address this problem, new learning paradigms designed for fair deep models are still necessary. Modern computer vision tasks usually involve large generi…

Cited by 11SourcePDFScholar
2021

A Faster Decentralized Algorithm for Nonconvex Minimax Problems

NeurIPS 2021poster

In this paper, we study the nonconvex-strongly-concave minimax optimization problem on decentralized setting. The minimax problems are attracting increasing attentions because of their popular practical applications such as policy evaluation and adversarial training. As training data become larger,…

Cited by 63SourcePDFScholar
2021

Communication-Efficient Frank-Wolfe Algorithm for Nonconvex Decentralized Distributed Learning

AAAI 2021technical

Recently decentralized optimization attracts much attention in machine learning because it is more communication-efficient than the centralized fashion. Quantization is a promising method to reduce the communication cost via cutting down the budget of each single communication using the gradient com…

Cited by 24SourcePDFScholar
2021

Large Batch Optimization for Deep Learning Using New Complete Layer-Wise Adaptive Rate Scaling

AAAI 2021technical

Training deep neural networks using a large batch size has shown promising results and benefits many real-world applications. Warmup is one of nontrivial techniques to stabilize the convergence of large batch training. However, warmup is an empirical method and it is still unknown whether there is a…

Cited by 22SourcePDFScholar
2021

Learning Better Visual Data Similarities via New Grouplet Non-Euclidean Embedding

ICCV 2021poster

In many computer vision problems, it is desired to learn the effective visual data similarity such that the prediction accuracy can be enhanced. Deep Metric Learning (DML) methods have been actively studied to measure the data similarity. Pair-based and proxy-based losses are the two major paradigms…

Cited by 16PDFcodeScholar
2021

On the Convergence of Communication-Efficient Local SGD for Federated Learning

AAAI 2021technical

Federated Learning (FL) has attracted increasing attention in recent years. A leading training algorithm in FL is local SGD, which updates the model parameter on each worker and averages model parameters across different workers only once in a while. Although it has fewer communication rounds than t…

Cited by 66SourcePDFScholar
2021

Secure Bilevel Asynchronous Vertical Federated Learning with Backward Updating

AAAI 2021technical

Vertical federated learning (VFL) attracts increasing attention due to the emerging demands of multi-party collaborative modeling and concerns of privacy leakage. In the real VFL applications, usually only one or partial parties hold labels, which makes it challenging for all parties to collaborativ…

Cited by 90SourcePDFScholar
2020

Binarized Neural Network for Single Image Super Resolution

ECCV 2020poster

Lighter model and faster inference are the focus of current single image super-resolution (SISR) research. However, existing methods are still hard to be applied in real-world applications due to the requirement of its heavy computation. Model quantization is an effective way to significantly reduce…

Cited by 90SourcePDFScholar
2020

Can Stochastic Zeroth-Order Frank-Wolfe Method Converge Faster for Non-Convex Problems?

ICML 2020poster

Frank-Wolfe algorithm is an efficient method for optimizing non-convex constrained problems. However, most of existing methods focus on the first-order case. In real-world applications, the gradient is not always available. To address the problem of lacking gradient in many applications, we propose…

Cited by 18SourcePDFScholar
2020

Discrete Model Compression With Resource Constraint for Deep Neural Networks

CVPR 2020poster

In this paper, we target to address the problem of compression and acceleration of Convolutional Neural Networks (CNNs). Specifically, we propose a novel structural pruning method to obtain a compact CNN with strong discriminative power. To find such networks, we propose an efficient discrete optimi…

Cited by 100PDFScholar
2020

Sinkhorn Regression

IJCAI 2020poster

This paper introduces a novel Robust Regression (RR) model, named Sinkhorn regression, which imposes Sinkhorn distances on both loss function and regularization. Traditional RR methods target at searching for an element-wise loss function (e.g., Lp-norm) to characterize the errors such that ou…

Cited by 0SourcePDFScholar
2020

Unsupervised Instance Segmentation in Microscopy Images via Panoptic Domain Adaptation and Task Re-Weighting

CVPR 2020poster

Unsupervised domain adaptation (UDA) for nuclei instance segmentation is important for digital pathology, as it alleviates the burden of labor-intensive annotation and domain shift across datasets. In this work, we propose a Cycle Consistency Panoptic Domain Adaptive Mask R-CNN (CyC-PDAM) architectu…

Cited by 98PDFcodeScholar
2019

Balanced Self-Paced Learning for Generative Adversarial Clustering Network

CVPR 2019oral

Clustering is an important problem in various machine learning applications, but still a challenging task when dealing with complex real data. The existing clustering algorithms utilize either shallow models with insufficient capacity for capturing the non-linear nature of data, or deep models with…

Cited by 126PDFScholar
2019

Faster Stochastic Alternating Direction Method of Multipliers for Nonconvex Optimization

ICML 2019oral

In this paper, we propose a faster stochastic alternating direction method of multipliers (ADMM) for nonconvex optimization by using a new stochastic path-integrated differential estimator (SPIDER), called as SPIDER-ADMM. Moreover, we prove that the SPIDER-ADMM achieves a record-breaking incremental…

Cited by 50SourcePDFScholar
2019

Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering

CVPR 2019poster

In this paper, we propose a novel end-to-end trainable Video Question Answering (VideoQA) framework with three major components: 1) a new heterogeneous memory which can effectively learn global context information from appearance and motion features; 2) a redesigned question memory which helps under…

Cited by 342PDFcodeScholar
2018

Direct Shape Regression Networks for End-to-End Face Alignment

CVPR 2018poster

Face alignment has been extensively studied in computer vision community due to its fundamental role in facial analysis, but it remains an unsolved problem. The major challenges lie in the highly nonlinear relationship between face images and associated facial shapes, which is coupled by underlying…

2018

Faster Derivative-Free Stochastic Algorithm for Shared Memory Machines

ICML 2018oral

Asynchronous parallel stochastic gradient optimization has been playing a pivotal role to solve large-scale machine learning problems in big data applications. Zeroth-order (derivative-free) methods estimate the gradient only by two function evaluations, thus have been applied to solve the problems…

Cited by 28SourcePDFScholar
2018

Unsupervised Deep Generative Adversarial Hashing Network

CVPR 2018poster

Unsupervised deep hash functions have not shown satisfactory improvements against the shallow alternatives, and usually, require supervised pretraining to avoid getting stuck in bad local minima. In this paper, we propose a deep unsupervised hashing function, called HashGAN, which outperforms unsupe…

Cited by 144SourcePDFScholar
2017

Deep Clustering via Joint Convolutional Autoencoder Embedding and Relative Entropy Minimization

ICCV 2017poster

In this paper, we propose a new clustering model, called DEeP Embedded RegularIzed ClusTering (DEPICT), which efficiently maps data into a discriminative embedding subspace and precisely predicts cluster assignments. DEPICT generally consists of a multinomial logistic regression function stacked on…

Cited by 585PDFcodeScholar
2017

Learning A Structured Optimal Bipartite Graph for Co-Clustering

NeurIPS 2017poster

Co-clustering methods have been widely applied to document clustering and gene expression analysis. These methods make use of the duality between features and samples such that the co-occurring structure of sample and feature clusters can be extracted. In graph based co-clustering methods, a biparti…

Cited by 176SourcePDFScholar
2017

Locally-Transferred Fisher Vectors for Texture Classification

ICCV 2017poster

Texture classification has been extensively studied in computer vision. Recent research shows that the combination of Fisher vector (FV) encoding and convolutional neural network (CNN) provides significant improvement in texture classification over the previous feature representation methods. Howeve…

Cited by 75PDFScholar
2017

Regularized Modal Regression with Applications in Cognitive Impairment Prediction

NeurIPS 2017poster

Linear regression models have been successfully used to function estimation and model selection in high-dimensional data analysis. However, most existing methods are built on least squares with the mean square error (MSE) criterion, which are sensitive to outliers and their performance may be degrad…

Cited by 41SourcePDFScholar
2015

Fusing Subcategory Probabilities for Texture Classification

CVPR 2015poster

Texture, as a fundamental characteristic of objects, has attracted much attention in computer vision research. Performance of texture classification is however still lacking for some challenging cases, largely due to the high intra-class variation and low inter-class distinction. To tackle these iss…

Cited by 22SourcePDFScholar