← Search

Suchi Saria

18 accepted papers

2026

Guiding Mixture-of-Experts with Temporal Multimodal Interactions

ICLR 2026poster

Mixture-of-Experts (MoE) architectures have become pivotal for large-scale multimodal models. However, their routing mechanisms typically overlook the informative, time-varying interaction dynamics between modalities. This limitation hinders expert specialization, as the model cannot explicitly leve…

Cited by 0SourceScholar
2026

Toward Calibrated Mixture-of-Experts Under Distribution Shift

ICML 2026poster

Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important toward understanding and trusting reported probabilities. Recent work shows that enforcing calibration at the level of individual predictors can substantially improve ensemble performa…

Cited by 0SourceScholar
2025

WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales

ICML 2025poster

Responsibly deploying artificial intelligence (AI) / machine learning (ML) systems in high-stakes settings arguably requires not only proof of system reliability, but also continual, post-deployment monitoring to quickly detect and address any unsafe behavior. Methods for nonparametric sequential te…

2024

Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)

ICML 2024poster

As artificial intelligence (AI) / machine learning (ML) gain widespread adoption, practitioners are increasingly seeking means to quantify and control the risk these systems incur. This challenge is especially salient when such systems have autonomy to collect their own data, such as in black-box op…

2024

FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion

NeurIPS 2024poster

As machine learning models in critical fields increasingly grapple with multimodal data, they face the dual challenges of handling a wide array of modalities, often incomplete due to missing elements, and the temporal irregularity and sparsity of collected samples. Successfully leveraging this compl…

Cited by 22SourcePDFScholar
2023

Data Augmentations for Improved (Large) Language Model Generalization

NeurIPS 2023poster

The reliance of text classifiers on spurious correlations can lead to poor generalization at deployment, raising concerns about their use in safety-critical domains such as healthcare. In this work, we propose to use counterfactual data augmentation, guided by knowledge of the causal structure of th…

Cited by 9SourcePDFScholar
2023

JAWS-X: Addressing Efficiency Bottlenecks of Conformal Prediction Under Standard and Feedback Covariate Shift

ICML 2023oral

We study the efficient estimation of predictive confidence intervals for black-box predictors when the common data exchangeability (e.g., i.i.d.) assumption is violated due to potentially feedback-induced shifts in the input data distribution. That is, we focus on standard and feedback covariate shi…

Cited by 7SourcePDFScholar
2021

Partial Identifiability in Discrete Data with Measurement Error

UAI 2021poster

When data contains measurement errors, it is necessary to make modeling assumptions relating the error-prone measurements to the unobserved true values. Work on measurement error has largely focused on models that fully identify the parameter of interest. As a result, many practically useful models…

Cited by 9SourcePDFScholar
2019

Active Learning for Decision-Making from Imbalanced Observational Data

ICML 2019oral

Machine learning can help personalized decision support by learning models to predict individual treatment effects (ITE). This work studies the reliability of prediction-based decision-making in a task of deciding which action $a$ to take for a target unit after observing its covariates $\tilde{x}$…

Cited by 38SourcePDFScholar
2019

Learning Models from Data with Measurement Error: Tackling Underreporting

ICML 2019oral

Measurement error in observational datasets can lead to systematic bias in inferences based on these datasets. As studies based on observational data are increasingly used to inform decisions with real-world impact, it is critical that we develop a robust set of techniques for analyzing and adjustin…

Cited by 18SourcePDFScholar
2019

Preventing Failures Due to Dataset Shift: Learning Predictive Models That Transport

AISTATS 2019poster

Classical supervised learning produces unreliable models when training and target distributions differ, with most existing solutions requiring samples from the target domain. We propose a proactive approach which learns a relationship in the training domain that will generalize to the target domain…

2015

A Framework for Individualizing Predictions of Disease Trajectories by Exploiting Multi-Resolution Structure

NeurIPS 2015poster

For many complex diseases, there is a wide variety of ways in which an individual can manifest the disease. The challenge of personalized medicine is to develop tools that can accurately predict the trajectory of an individual's disease, which can in turn enable clinicians to optimize treatments. We…

Cited by 116SourcePDFScholar