Divid: Disentangled Spatial-Temporal Modeling within LLMs for Temporally Grounded Video Understanding
Recent advances in Video LLMs have improved video understanding performance, but temporally grounded understanding in long-form videos remains challenging. Most models encode video frames into a flat sequence of visual tokens, which are then processed together with textual input by the LLM. While ef…