← Search

Yu Xiong

9 accepted papers

2025

Enhancing Teacher Classroom Behavior Descriptions: A Spatio-Temporal Graph-Based Method for Video Captioning

ICASSP 2025accepted

Teacher behavior description is an objective record of teachers’ behaviors in the teaching process, aiming to provide evidence for reflection on teaching behavior. Although video captioning technology can automatically generate behavior description, existing research on teacher behavior description…

Cited by 0SourceScholar
2024

MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

NeurIPS 2024poster

Sora's high-motion intensity and long consistent videos have significantly impacted the field of video generation, attracting unprecedented attention. However, existing publicly available datasets are inadequate for generating Sora-like videos, as they mainly contain short videos with low motion int…

Cited by 42SourcePDFScholar
2024

Robust Beamforming for Downlink Multi-Cell Systems: A Bilevel Optimization Perspective

AAAI 2024technical

Utilization of inter-base station cooperation for information processing has shown great potential in enhancing the overall quality of communication services (QoS) in wireless communication networks. Nevertheless, such cooperations require the knowledge of channel state information (CSI) at base sta…

Cited by 3SourcePDFScholar
2020

A Local-to-Global Approach to Multi-Modal Movie Scene Segmentation

CVPR 2020poster

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semantic understanding of movies. This is very challenging - compared to the videos st…

Cited by 155PDFcodeScholar
2020

MovieNet: A Holistic Dataset for Movie Understanding

ECCV 2020poster

Recent years have seen remarkable advances in visual understanding. However, how to understand a story-based long video with artistic styles, e.g. movie, remains challenging. In this paper, we introduce MovieNet -- a holistic dataset for movie understanding. MovieNet contains 1,100 movies with a lar…

2019

A Graph-Based Framework to Bridge Movies and Synopses

ICCV 2019oral

Inspired by the remarkable advances in video analytics, research teams are stepping towards a greater ambition - movie understanding. However, compared to those activity videos in conventional datasets, movies are significantly different. Generally, movies are much longer and consist of much richer…

Cited by 78PDFcodeScholar
2019

Hybrid Task Cascade for Instance Segmentation

CVPR 2019poster

Cascade is a classic yet powerful architecture that has boosted performance on various tasks. However, how to introduce cascade to instance segmentation remains an open question. A simple combination of Cascade R-CNN and Mask R-CNN only brings limited gain. In exploring a more effective approach, we…

Cited by 1727PDFcodeScholar
2018

Find and Focus: Retrieve and Localize Video Events with Natural Language Queries

ECCV 2018poster

The thriving of video sharing services brings new challenges to video retrieval, e.g. the rapid growth in video duration and content diversity. Meeting such challenges calls for new techniques that can effectively retrieve videos with natural language queries. Existing methods along this line, which…

Cited by 90SourcePDFScholar