← Search

Jialuo Li

6 accepted papers

2026

Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding

CVPR 2026

The application of Large Multimodal Models (LMMs) to long-form video understanding is constrained by limited context lengths and the computationally prohibitive cost of processing dense video tokens. Consequently, recent research has focused on query-aware frame selection, methods that often incur s

Cited by 0SourceScholar
2026

DuoGen: Towards Autonomous Interleaved Multimodal Generation

CVPR 2026

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of existing interleaved generation models under general instructions remains limited

Cited by 0SourceScholar
2026

SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning

CVPR 2026

As humans, we are natural any-horizon reasoners, i.e., we can decide whether to iteratively skim long videos or watch short ones in full when necessary for a given task. With this in mind, one would expect video reasoning models to reason flexibly across different durations. However, SOTA models are

Cited by 0SourcecodeScholar
2026

STEAM-LIVO: Spatio-Temporally Adaptive Manifold Lidar-Inertial-Visual Odometry for Sensor Degradation in Unstructured Natural Aquatic-Terrestrial Scenes

ICRA 2026poster

Sensor degradation in unstructured natural environments---manifesting as LiDAR point cloud sparsity or visual feature dropout---and out-of-sequence measurement challenges critically undermine localization robustness in autonomous systems. To address these limitations, we present STEAM-LIVO, a Spatio…

Cited by 0Scholar
2025

Science-T2I: Addressing Scientific Illusions in Image Synthesis

CVPR 2025poster

We present a novel approach to integrating scientific knowledge into generative models, enhancing their realism and consistency in image synthesis. First, we introduce Science-T2I, an expert-annotated adversarial dataset comprising adversarial 20k image pairs with 9k prompts, covering wide distinct…

Cited by 1SourcePDFScholar