← Search

Sibo Song

7 accepted papers

2026

Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos

CVPR 2026

The transition from image to video understanding requires vision-language models (VLMs) to shift from recognizing static patterns to reasoning over temporal dynamics such as motion trajectories, speed changes, and state transitions. Yet current post-training methods fall short due to two critical li

Cited by 0SourcecodeScholar
2026

Revisiting Multimodal Positional Encoding in Vision–Language Models

ICLR 2026poster

Multimodal position encoding is essential for vision-language models, yet there has been little systematic investigation into multimodal position encoding. We conduct a comprehensive analysis of multimodal Rotary Positional Embedding (RoPE) by examining its two core components: position design and f…

Cited by 0SourcecodeScholar
2024

OmniParser: A Unified Framework for Text Spotting Key Information Extraction and Table Recognition

CVPR 2024poster

Recently visually-situated text parsing (VsTP) has experienced notable advancements driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs) capable of processing document-based questions. Various methods have been proposed to…

2023

Modeling Entities As Semantic Points for Visual Information Extraction in the Wild

CVPR 2023poster

Recently, Visual Information Extraction (VIE) has been becoming increasingly important in both academia and industry, due to the wide range of real-world applications. Previously, numerous works have been proposed to tackle this problem. However, the benchmarks used to assess these methods are relat…

2022

Vision-Language Pre-Training for Boosting Scene Text Detectors

CVPR 2022poster

Recently, vision-language joint representation learning has proven to be highly effective in various scenarios. In this paper, we specifically adapt vision-language joint learning for scene text detection, a task that intrinsically involves cross-modal interaction between the two modalities: vision…

Cited by 37PDFcodeScholar
2016

Egocentric activity recognition with multimodal fisher vector

ICASSP 2016accepted

With the increasing availability of wearable devices, research on egocentric activity recognition has received much attention recently. In this paper, we build a Multimodal Egocentric Activity dataset which includes egocentric videos and sensor data of 20 fine-grained and diverse activity categories…

Cited by 0SourceScholar