← Search

Jingran Zhang

2 accepted papers

2022

Semi-Supervised Video Paragraph Grounding With Contrastive Encoder

CVPR 2022poster

Video events grounding aims at retrieving the most relevant moments from an untrimmed video in terms of a given natural language query. Most previous works focus on Video Sentence Grounding (VSG), which localizes the moment with a sentence query. Recently, researchers extended this task to Video Par…

Cited by 35PDFScholar
2021

Enhancing Audio-Visual Association with Self-Supervised Curriculum Learning

AAAI 2021technical

The recent success of audio-visual representations learning can be largely attributed to their pervasive concurrency property, which can be used as a self-supervision signal and extract correlation information. While most recent works focus on capturing the shared associations between the audio and…

Cited by 26SourcePDFScholar