← Search

Manjot Bilkhu

2 accepted papers

2026

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights

ICML 2026poster

Large-scale web-crawled datasets contain noise, bias, and irrelevant information, necessitating data selection techniques. Existing methods depend on hand-crafted heuristics, downstream datasets, or require expensive influence-based computations---all of which limit scalability and introduce unwante…

Cited by 0SourceScholar
2021

Probabilistic Attention for Interactive Segmentation

NeurIPS 2021spotlight

We provide a probabilistic interpretation of attention and show that the standard dot-product attention in transformers is a special case of Maximum A Posteriori (MAP) inference. The proposed approach suggests the use of Expectation Maximization algorithms for on-line adaptation of key and value mod…