← Search

Zeqian Li

5 accepted papers

2025

Learning Streaming Video Representation via Multitask Training

ICCV 2025poster

Understanding continuous video streams plays a fundamental role in real-time applications, including embodied AI and autonomous driving. Unlike offline video processing, streaming video understanding requires the ability to process video streams frame by frame, preserve historical information, and m…

Cited by 0SourcePDFScholar
2025

Universal Video Temporal Grounding with Generative Multi-modal Large Language Models

NeurIPS 2025poster

This paper presents a computational model for universal video temporal grounding, which accurately localizes temporal moments in videos based on natural language queries (e.g., questions or descriptions). Unlike existing methods that are often limited to specific video domains or durations, we prop…

Cited by 0SourcecodeScholar
2024

Multi-Sentence Grounding for Long-term Instructional Video

ECCV 2024poster

"In this paper, we aim to establish an automatic, scalable pipeline for denoising the large-scale instructional dataset and construct a high-quality video-text dataset with multiple descriptive steps supervision, named HowToStep. We make the following contributions: (i) improving the quality of sent…

Cited by 2SourcePDFScholar
2021

Compositional Embedding Models for Speaker Identification and Diarization with Simultaneous Speech From 2+ Speakers

ICASSP 2021accepted

We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding models contain a function f that separates speech from differe…

Cited by 0SourceScholar
2020

Toward Better Speaker Embeddings: Automated Collection of Speech Samples From Unknown Distinct Speakers

ICASSP 2020accepted

The accuracy of speaker verification and diarization models depends on the quality of the speaker embeddings used to separate audio samples from different speakers. With the goal of training better embedding models, we devise an automatic pipeline for large-scale collection of speech samples from un…

Cited by 0SourceScholar