← Search

Dylan Sam

11 accepted papers

2025

Analyzing Similarity Metrics for Data Selection for Language Model Pretraining

NeurIPS 2025poster

Measuring similarity between training examples is critical for curating high-quality and diverse pretraining datasets for language models. However, similarity is typically computed with a generic off-the-shelf embedding model that has been trained for tasks such as retrieval. Whether these embeddi…

Cited by 0SourceScholar
2025

Predicting the Performance of Black-box Language Models with Follow-up Queries

NeurIPS 2025poster

Reliably predicting the behavior of language models---such as whether their outputs are correct or have been adversarially manipulated---is a fundamentally challenging task. This is often made even more difficult as frontier language models are offered only through closed-source APIs, providing only…

Cited by 0SourceScholar
2025

Safety Pretraining: Toward the Next Generation of Safe AI

NeurIPS 2025poster

As large language models (LLMs) are increasingly deployed in high-stakes settings, the risk of generating harmful or toxic content remains a central challenge. Post-hoc alignment methods are brittle: once unsafe patterns are learned during pretraining, they are hard to remove. In this work, we prese…

Cited by 0SourceScholar
2024

Auditing Fairness under Unobserved Confounding

AISTATS 2024poster

A fundamental problem in decision-making systems is the presence of inequity along demographic lines. However, inequity can be difficult to quantify, particularly if our notion of equity relies on hard-to-measure notions like risk (e.g., equal access to treatment for those who would die without it).…

2024

Computing Low-Entropy Couplings for Large-Support Distributions

UAI 2024poster

Minimum-entropy coupling (MEC)—the process of finding a joint distribution with minimum entropy for given marginals—has applications in areas such as causality and steganography. However, existing algorithms are either computationally intractable for large-support distributions or limited to specifi…

2024

Understanding prompt engineering may not require rethinking generalization

ICLR 2024poster

Zero-shot learning in prompted vision-language models, the practice of crafting prompts to build classifiers without an explicit training process, has achieved impressive performance in many settings. This success presents a seemingly surprising observation: these methods suffer relatively little fr…

Cited by 11SourcePDFScholar
2023

Learning with Explanation Constraints

NeurIPS 2023poster

As larger deep learning models are hard to interpret, there has been a recent focus on generating explanations of these black-box models. In contrast, we may have apriori explanations of how models should behave. In this paper, we formalize this notion as learning from explanation constraints and p…

Cited by 7SourcePDFScholar
2023

Losses over Labels: Weakly Supervised Learning via Direct Loss Construction

AAAI 2023technical

Owing to the prohibitive costs of generating large amounts of labeled data, programmatic weak supervision is a growing paradigm within machine learning. In this setting, users design heuristics that provide noisy labels for subsets of the data. These weak labels are combined (typically via a graphic…

2021

Adversarial Multi Class Learning under Weak Supervision with Performance Guarantees

ICML 2021spotlight

We develop a rigorous approach for using a set of arbitrarily correlated weak supervision sources in order to solve a multiclass classification task when only a very small set of labeled data is available. Our learning algorithm provably converges to a model that has minimum empirical risk with resp…

Cited by 40SourcePDFScholar
2021

Semi-Supervised Aggregation of Dependent Weak Supervision Sources With Performance Guarantees

AISTATS 2021poster

We develop a novel method that provides theoretical guarantees for learning from weak labelers without the (mostly unrealistic) assumption that the errors of the weak labelers are independent or come from a particular family of distributions. We show a rigorous technique for efficiently selecting sm…

Cited by 34SourcePDFScholar