LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations
In this paper, we present a benchmark to pressure-test today’s frontier models’ multimodal decision-making capabilities in the very long-context regime (up to one million tokens) and investigate whether these models can learn from large numbers of expert demonstrations in their context. We evaluate…