← Search

Guorong Li

18 accepted papers

2026

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding

ICML 2026poster

Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream, often resulting in brittle numerics and inconsistent boundaries. To address this, we propose Foresee-to-Ground (F2G), a framework that enforces a ve…

Cited by 0SourceScholar
2026

HeroGS: Hierarchical Guidance for Robust 3D Gaussian Splatting under Sparse Views

CVPR 2026

3D Gaussian Splatting (3DGS) has recently emerged as a promising approach in novel view synthesis, combining photorealistic rendering with real-time efficiency. However, its success heavily relies on dense camera coverage; under sparse-view conditions, insufficient supervision leads to irregular Gau

Cited by 0SourceScholar
2025

Less Is More: Token Context-Aware Learning for Object Tracking

AAAI 2025technical

Recently, several studies have shown that utilizing contextual information to perceive target states is crucial for object tracking. They typically capture context by incorporating multiple video frames. However, these naive frame-context methods fail to consider the importance of each patch within…

2025

MambaLCT: Boosting Tracking via Long-term Context State Space Model

AAAI 2025technical

Effectively constructing context information with long-term dependencies from video sequences is crucial for object tracking. However, the context length constructed by existing work is limited, only considering object information from adjacent frames or video clips, leading to insufficient utilizat…

2024

Semantic-aware SAM for Point-Prompted Instance Segmentation

CVPR 2024highlight

Single-point annotation in visual tasks with the goal of minimizing labeling costs is becoming increasingly prominent in research. Recently visual foundation models such as Segment Anything (SAM) have gained widespread usage due to their robust zero-shot capabilities and exceptional annotation perfo…

2024

Weakly Supervised Video Individual Counting

CVPR 2024poster

Video Individual Counting (VIC) aims to predict the number of unique individuals in a single video. Existing methods learn representations based on trajectory labels for individuals which are annotation-expensive. To provide a more realistic reflection of the underlying practical challenge we introd…

2023

Exploiting Completeness and Uncertainty of Pseudo Labels for Weakly Supervised Video Anomaly Detection

CVPR 2023poster

Weakly supervised video anomaly detection aims to identify abnormal events in videos using only video-level labels. Recently, two-stage self-training methods have achieved significant improvements by self-generating pseudo labels and self-refining anomaly scores with these labels. As the pseudo labe…

Cited by 92SourcePDFScholar
2023

Spatial Self-Distillation for Object Detection with Inaccurate Bounding Boxes

ICCV 2023poster

Object detection via inaccurate bounding box supervision has boosted a broad interest due to the expensive high-quality annotation data or the occasional inevitability of low annotation quality (e.g. tiny objects). The previous works usually utilize multiple instance learning (MIL), which highly dep…

Cited by 18PDFcodeScholar
2022

Hierarchical Modular Network for Video Captioning

CVPR 2022poster

Video captioning aims to generate natural language descriptions according to the content, where representation learning plays a crucial role. Existing methods are mainly developed within the supervised learning framework via word-by-word comparison of the generated caption against the ground-truth t…

Cited by 113PDFcodeScholar
2022

Object Localization Under Single Coarse Point Supervision

CVPR 2022poster

Point-based object localization (POL), which pursues high-performance object sensing under low-cost data annotation, has attracted increased attention. However, the point annotation mode inevitably introduces semantic variance for the inconsistency of annotated points. Existing POL methods heavily r…

Cited by 33PDFcodeScholar
2021

Exploiting Sample Correlation for Crowd Counting With Multi-Expert Network

ICCV 2021poster

Crowd counting is a difficult task because of the diversity of scenes. Most of the existing crowd counting methods adopt complex structures with massive backbones to enhance the generalization ability. Unfortunately, the performance of existing methods on large-scale data sets is not satisfactory. I…

Cited by 38PDFScholar
2021

Learning To Filter: Siamese Relation Network for Robust Tracking

CVPR 2021poster

Despite the great success of Siamese-based trackers, their performance under complicated scenarios is still not satisfying, especially when there are distractors. To this end, we propose a novel Siamese relation network, which introduces two efficient modules, i.e. Relation Detector (RD) and Refinem…

Cited by 146PDFcodeScholar
2020

Reverse Perspective Network for Perspective-Aware Object Counting

CVPR 2020poster

One of the critical challenges of object counting is the dramatic scale variations, which is introduced by arbitrary perspectives. We propose a reverse perspective network to solve the scale variations of input images, instead of generating perspective maps to smooth final outputs. The reverse persp…

Cited by 168PDFScholar
2020

Siamese Box Adaptive Network for Visual Tracking

CVPR 2020poster

Most of the existing trackers usually rely on either a multi-scale searching scheme or pre-defined anchor boxes to accurately estimate the scale and aspect ratio of a target. Unfortunately, they typically call for tedious and heuristic configurations. To address this issue, we propose a simple yet e…

Cited by 1031PDFcodeScholar
2020

Weakly-Supervised Crowd Counting Learns from Sorting rather than Locations

ECCV 2020poster

In crowd counting datasets, the location labels are costly, yet, they are not taken into the evaluation metrics. Besides, existing multi-task approaches employ high-level tasks to improve counting accuracy. This research tendency increases the demand for more annotations. In this paper, we propose a…

Cited by 108SourcePDFScholar
2018

The Unmanned Aerial Vehicle Benchmark: Object Detection and Tracking

ECCV 2018poster

With the advantage of high mobility, Unmanned Aerial Vehicles (UAVs) are used to fuel numerous important applications in computer vision, delivering more efficiency and convenience than surveillance cameras with fixed camera angle, scale and view. However, very limited UAV datasets are proposed, and…

Cited by 974SourcePDFScholar