← Search

Yuheng Bu

21 accepted papers

2026

In-Context Watermarks for Large Language Models

ICLR 2026poster

The growing use of large language models (LLMs) for sensitive applications has highlighted the need for effective watermarking techniques to ensure the provenance and accountability of AI-generated text. However, most existing watermarking methods require access to the decoding process, limiting the…

Cited by 0SourcecodeScholar
2026

TrustEnergy: A Unified Framework for Accurate and Reliable User-level Energy Usage Prediction

AAAI 2026technical

Energy usage prediction is important for various real-world applications, including grid management, infrastructure planning, and disaster response. Although a plethora of deep learning approaches have been proposed to perform this task, most of them either overlook the essential spatial correlation

Cited by 0SourcePDFScholar
2025

Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective

ICML 2025poster

Despite substantial progress in promoting fairness in high-stake applications using machine learning models, existing methods often modify the training process, such as through regularizers or other interventions, but lack formal guarantees that fairness achieved during training will generalize to u…

Cited by 0SourcePDFScholar
2025

Image Watermarks are Removable using Controllable Regeneration from Clean Noise

ICLR 2025poster

Image watermark techniques provide an effective way to assert ownership, deter misuse, and trace content sources, which has become increasingly essential in the era of large generative models. A critical attribute of watermark techniques is their robustness against various manipulations. In this pap…

2025

Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive Approach

NeurIPS 2025poster

Watermarking has emerged as a crucial method to distinguish AI-generated text from human-created text. Current watermarking approaches often lack formal optimality guarantees or address the scheme and detector design separately. In this paper, we introduce a novel, unified theoretical framework for…

Cited by 0SourcecodeScholar
2024

Are Uncertainty Quantification Capabilities of Evidential Deep Learning a Mirage?

NeurIPS 2024poster

This paper questions the effectiveness of a modern predictive uncertainty quantification approach, called *evidential deep learning* (EDL), in which a single neural network model is trained to learn a meta distribution over the predictive distribution by minimizing a specific objective function. Des…

2024

Operator SVD with Neural Networks via Nested Low-Rank Approximation

ICML 2024poster

Computing eigenvalue decomposition (EVD) of a given linear operator, or finding its leading eigenvalues and eigenfunctions, is a fundamental task in many machine learning and scientific simulation problems. For high-dimensional eigenvalue problems, training neural networks to parameterize the eigenf…

2023

How Does Pseudo-Labeling Affect the Generalization Error of the Semi-Supervised Gibbs Algorithm?

AISTATS 2023poster

We provide an exact characterization of the expected generalization error (gen-error) for semi-supervised learning (SSL) with pseudo-labeling via the Gibbs algorithm. The gen-error is expressed in terms of the symmetrized KL information between the output hypothesis, the pseudo-labeled dataset, and…

Cited by 6SourcePDFScholar
2023

On Balancing Bias and Variance in Unsupervised Multi-Source-Free Domain Adaptation

ICML 2023poster

Due to privacy, storage, and other constraints, there is a growing need for unsupervised domain adaptation techniques in machine learning that do not require access to the data used to train a collection of source models. Existing methods for multi-source-free domain adaptation (MSFDA) typically tra…

2023

Post-hoc Uncertainty Learning Using a Dirichlet Meta-Model

AAAI 2023technical

It is known that neural networks have the problem of being over-confident when directly using the output label distribution to generate uncertainty measures. Existing methods mainly resolve this issue by retraining the entire model to impose the uncertainty quantification capability so that the lear…

2022

A Maximal Correlation Approach to Imposing Fairness in Machine Learning

ICASSP 2022accepted

As machine learning algorithms grow in popularity and diversify to many industries, ethical and legal concerns regarding their fairness have become increasingly relevant. We explore the problem of algorithmic fairness, taking an information-theoretic view. The maximal correlation framework is introd…

Cited by 0SourceScholar
2022

Characterizing and Understanding the Generalization Error of Transfer Learning with Gibbs Algorithm

AISTATS 2022poster

We provide an information-theoretic analysis of the generalization ability of Gibbs-based transfer learning algorithms by focusing on two popular empirical risk minimization (ERM) approaches for transfer learning, $\alpha$-weighted-ERM and two-stage-ERM. Our key result is an exact characterization o…

Cited by 17SourcePDFScholar
2022

Selective Regression under Fairness Criteria

ICML 2022spotlight

Selective regression allows abstention from prediction if the confidence to make an accurate prediction is not sufficient. In general, by allowing a reject option, one expects the performance of a regression model to increase at the cost of reducing coverage (i.e., by predicting on fewer samples). H…

2021

An Exact Characterization of the Generalization Error for the Gibbs Algorithm

NeurIPS 2021poster

Various approaches have been developed to upper bound the generalization error of a supervised learning algorithm. However, existing bounds are often loose and lack of guarantees. As a result, they may fail to characterize the exact generalization ability of a learning algorithm. Our main contributi…

Cited by 54SourcePDFScholar
2021

Fair Selective Classification Via Sufficiency

ICML 2021oral

Selective classification is a powerful tool for decision-making in scenarios where mistakes are costly but abstentions are allowed. In general, by allowing a classifier to abstain, one can improve the performance of a model at the cost of reducing coverage and classifying fewer samples. However, rec…

2016

Universal outlying sequence detection for continuous observations

ICASSP 2016accepted

The following detection problem is studied, in which there are M sequences of samples out of which one outlier sequence needs to be detected. Each typical sequence contains n independent and identically distributed (i.i.d.) continuous observations from a known distribution π, and the outlier sequenc…

Cited by 0SourceScholar