← Search

Kun Xia

7 accepted papers

2025

Diversifying Query: Region-Guided Transformer for Temporal Sentence Grounding

AAAI 2025technical

Temporal sentence grounding is a challenging task that aims to localize the moment spans relevant to a language description. Although recent DETR-based models have achieved notable progress by leveraging multiple learnable moment queries, they suffer from overlapped and redundant proposals, leading…

2025

Moment Quantization for Video Temporal Grounding

ICCV 2025poster

Video temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishing relevant and irrelevant moments. Previous methods focused on learning continuous features exhibit weak differentiation…

2025

SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World Models

NeurIPS 2025poster

World models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions due to disrupted spatial structure, inefficient decoding, and…

Cited by 0SourceScholar
2024

Analysis-by-Synthesis Transformer for Single-View 3D Reconstruction

ECCV 2024poster

"Deep learning approaches have made significant success in single-view 3D reconstruction, but they often rely on expensive 3D annotations for training. Recent efforts tackle this challenge by adopting an analysis-by-synthesis paradigm to learn 3D reconstruction with only 2D annotations. However, exi…

2024

Stepwise Multi-grained Boundary Detector for Point-supervised Temporal Action Localization

ECCV 2024poster

"Point-supervised temporal action localization pursues high-accuracy action detection under low-cost data annotation. Despite recent advances, a significant challenge remains: sparse labeling of individual frames leads to semantic ambiguity in determining action boundaries due to the lack of continu…

Cited by 0SourcePDFScholar
2023

Learning from Noisy Pseudo Labels for Semi-Supervised Temporal Action Localization

ICCV 2023poster

Semi-Supervised Temporal Action Localization (SS-TAL) aims to improve the generalization ability of action detectors with large-scale unlabeled videos. Albeit the recent advancement, one of the major challenges still remains: noisy pseudo labels hinder efficient learning on abundant unlabeled videos…

Cited by 10PDFcodeScholar
2022

Learning To Refactor Action and Co-Occurrence Features for Temporal Action Localization

CVPR 2022poster

The main challenge of Temporal Action Localization is to retrieve subtle human actions from various co-occurring ingredients, e.g., context and background, in an untrimmed video. While prior approaches have achieved substantial progress through devising advanced action detectors, they still suffer f…

Cited by 57PDFScholar