2025
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
EMNLP 2025
Recent advancements in large language models (LLMs) have shifted focus toward scaling inference-time compute—improving performance without retraining the model. A common approach is to sample multiple outputs in parallel, and select one of these as the final output. While existing work has focused o