← Search

Carey Priebe

7 accepted papers

2025

Statistical inference on black-box generative models in the data kernel perspective space

ACL 2025finding

Generative models are capable of producing human-expert level content across a variety of topics and domains. As the impact of generative models grows, it is necessary to develop statistical methods to understand collections of available models. These methods are particularly important in settings w…

Cited by 0SourcePDFScholar
2024

Tracking the perspectives of interacting language models

EMNLP 2024main

Large language models (LLMs) are capable of producing high quality information at unprecedented rates. As these models continue to entrench themselves in society, the content they produce will become increasingly pervasive in databases that are, in turn, incorporated into the pre-training data, fine…

Cited by 4SourcePDFScholar
2023

The Value of Out-of-Distribution Data

ICML 2023poster

Generalization error always improves with more in-distribution data. However, it is an open question what happens as we add out-of-distribution (OOD) data. Intuitively, if the OOD data is quite different, it seems more data would harm generalization error, though if the OOD data are sufficiently sim…

2022

Bilingual Lexicon Induction for Low-Resource Languages using Graph Matching via Optimal Transport

EMNLP 2022main

Bilingual lexicons form a critical component of various natural language processing applications, including unsupervised and semisupervised machine translation and crosslingual information retrieval. In this work, we improve bilingual lexicon induction performance across 40 language pairs with a gra…

2021

An Analysis of Euclidean vs. Graph-Based Framing for Bilingual Lexicon Induction from Word Embedding Spaces

EMNLP 2021finding

Much recent work in bilingual lexicon induction (BLI) views word embeddings as vectors in Euclidean space. As such, BLI is typically solved by finding a linear transformation that maps embeddings to a common space. Alternatively, word embeddings may be understood as nodes in a weighted graph. This f…

2018

Out-of-sample extension of graph adjacency spectral embedding

ICML 2018oral

Many popular dimensionality reduction procedures have out-of-sample extensions, which allow a practitioner to apply a learned embedding to observations not seen in the initial training sample. In this work, we consider the problem of obtaining an out-of-sample extension for the adjacency spectral em…

Cited by 22SourcePDFScholar