← Search

Kou Misaki

4 accepted papers

2025

TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models

ICLR 2025spotlight

Causal language models have demonstrated remarkable capabilities, but their size poses significant challenges for deployment in resource-constrained environments. Knowledge distillation, a widely-used technique for transferring knowledge from a large teacher model to a small student model, presents…

Cited by 0SourcePDFScholar
2025

Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search

NeurIPS 2025spotlight

Recent advances demonstrate that increasing inference-time computation can significantly boost the reasoning capabilities of large language models (LLMs). Although repeated sampling (i.e., generating multiple candidate outputs) is a highly effective strategy, it does not leverage external feedback s…

Cited by 0SourceScholar