← Search

Yash Patel

10 accepted papers

2026

Continuum Transformers Perform In-Context Learning by Operator Gradient Descent

ICLR 2026poster

Transformers robustly exhibit the ability to perform in-context learning, whereby their predictive accuracy on a task can increase not by parameter updates but merely with the placement of training samples in their context windows. Recent works have shown that transformers achieve this by implementi…

Cited by 0SourcecodeScholar
2025

Conformal Prediction for Ensembles: Improving Efficiency via Score-Based Aggregation

NeurIPS 2025poster

Distribution-free uncertainty estimation for ensemble methods is increasingly desirable due to the widening deployment of multi-modal black-box predictive models. Conformal prediction is one approach that avoids such distributional assumptions. Methods for conformal aggregation have in turn been pro…

Cited by 0SourcecodeScholar
2025

Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch

ACL 2025long

To meet the demands of content moderation, online platforms have resorted to automated systems. Newer forms of real-time engagement (e.g., users commenting on live streams) on platforms like Twitch exert additional pressures on the latency expected of such moderation systems. Despite their prevalenc…

2024

Variational Inference with Coverage Guarantees in Simulation-Based Inference

ICML 2024poster

Amortized variational inference is an often employed framework in simulation-based inference that produces a posterior approximation that can be rapidly computed given any new observation. Unfortunately, there are few guarantees about the quality of these approximate posteriors. We propose Conformal…

2023

Contrastive Classification and Representation Learning with Probabilistic Interpretation

AAAI 2023technical

Cross entropy loss has served as the main objective function for classification-based tasks. Widely deployed for learning neural network classifiers, it shows both effectiveness and a probabilistic interpretation. Recently, after the success of self supervised contrastive representation learning me…

Cited by 7SourcePDFScholar
2023

Filtering, Distillation, and Hard Negatives for Vision-Language Pre-Training

CVPR 2023poster

Vision-language models trained with contrastive learning on large-scale noisy data are becoming increasingly popular for zero-shot recognition problems. In this paper we improve the following three aspects of the contrastive pre-training pipeline: dataset noise, model initialization and the training…

2017

Self-Supervised Learning of Visual Features Through Embedding Images Into Text Topic Spaces

CVPR 2017poster

End-to-end training from scratch of current deep architectures for new computer vision problems would require Imagenet-scale datasets, and this is not always possible. In this paper we present a method that is able to take advantage of freely available multi-modal content to train computer vision al…

Cited by 143PDFScholar