← Search

Skyler Seto

17 accepted papers

2026

Closing the Gap Between Text and Speech Understanding in LLMs

ICLR 2026poster

Large Language Models (LLMs) can be adapted to extend their text capabilities to speech inputs. However, these speech-adapted LLMs consistently underperform their text-based counterparts—and even cascaded pipelines—on language understanding tasks. We term this shortfall the text–speech understanding…

Cited by 0SourcecodeScholar
2026

Normalized Rewards for Preference Optimization

ICML 2026poster

Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their implicit reward model and decrease the likelihood of preferred responses. This results in a decrease in the total likel…

Cited by 0SourceScholar
2026

Optimal Splitting of Language Models from Mixtures to Specialized Domains

ICML 2026poster

Language models achieve impressive performance on a variety of knowledge, language, and reasoning tasks due to the scale and diversity of pretraining data available. The standard training recipe is a two-stage paradigm: pretraining first on the full corpus of data followed by specialization on a muc…

Cited by 0SourceScholar
2025

Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks

EMNLP 2025

Large language models (LLMs) are ubiquitous in modern day natural language processing. However, previous work has shown degraded LLM performance for under-represented English dialects. We analyze the effects of typifying “standard” American English language questions as non-”standard” dialectal vari

Cited by 0SourcePDFScholar
2025

Assessing the Role of Data Quality in Training Bilingual Language Models

EMNLP 2025

Bilingual and multilingual language models offer a promising path toward scaling NLP systems across diverse languages and users. However, their performance often varies wildly between languages as prior works show that adding more languages can degrade performance for some languages (such as English

2025

Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks

EMNLP 2025

We introduce a set of training-free ABX-style discrimination tasks to evaluate how multilingual language models represent language identity (form) and semantic content (meaning). Inspired from speech processing, these zero-shot tasks measure whether minimal differences in representation can be relia

Cited by 0SourcePDFScholar
2025

Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting

ICML 2025poster

Vision foundation models pre-trained on massive data encode rich representations of real-world concepts, which can be adapted to downstream tasks by fine-tuning. However, fine-tuning foundation models on one task often leads to the issue of *concept forgetting* on other tasks. Recent methods of robu…

Cited by 0SourcePDFScholar
2025

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging

ICML 2025poster

Machine learning models are routinely trained on a mixture of different data domains. Different domain weights yield very different downstream performances. We propose the Soup-of-Experts, a novel architecture that can instantiate a model at test time for any domain weights with minimal computation…

Cited by 0SourcePDFScholar
2025

Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models

ACL 2025long

Aligned representations across languages is a desired property in multilingual large language models (mLLMs), as alignment can improve performance in cross-lingual tasks. Typically alignment requires fine-tuning a model, which is computationally expensive, and sizable language data, which often may…

Cited by 0SourcePDFScholar
2025

Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling

ICLR 2025poster

Specialist language models (LMs) focus on a specific task or domain on which they often outperform generalist LMs of the same size. However, the specialist data needed to pretrain these models is only available in limited amount for most tasks. In this work, we build specialist models from large gen…

Cited by 3SourcePDFScholar
2025

Training Bilingual LMs with Data Constraints in the Targeted Language

ACL 2025finding

Large language models are trained on massive scrapes of the web, as required by current scaling laws. Most progress is made for English, given its abundance of high-quality pretraining data. For most other languages, however, such high quality pretraining data is unavailable. In this work, we study…

2024

Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP

NeurIPS 2024poster

Large pretrained vision-language models like CLIP have shown promising generalization capability, but may struggle in specialized domains (e.g., satellite imagery) or fine-grained classification (e.g., car models) where the visual concepts are unseen or under-represented during pretraining. Prompt l…

Cited by 0SourcePDFScholar
2024

Learning Spatially-Aware Language and Audio Embeddings

NeurIPS 2024poster

Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like "the lion roar came from right behind me!". For a machine to have the same degree of comprehension, the machine must know what a lion is (…

Cited by 1SourcePDFScholar
2024

On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

EMNLP 2024finding

Reinforcement Learning from Human Feedback (RLHF) is an effective approach for aligning language models to human preferences. Central to RLHF is learning a reward function for scoring human preferences. Two main approaches for learning a reward model are 1) training an EXplicit Reward Model (EXRM) a…

2024

Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling

ACL 2024long

Large language models are trained on massive scrapes of the web, which are often unstructured, noisy, and poorly phrased. Current scaling laws show that learning from such data requires an abundance of both compute and data, which grows with the size of the model being trained. This is infeasible bo…

Cited by 57SourcePDFScholar
2023

On the Role of LIP Articulation in Visual Speech Perception

ICASSP 2023accepted

Generating realistic lip motion from audio to simulate speech production is critical for driving natural character animation. Previous research has shown that traditional metrics used to optimize and assess models for generating lip motion from speech are not a good indicator of subjective opinion o…

Cited by 0SourceScholar