← Search

Dian Li

11 accepted papers

2025

Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning

NeurIPS 2025poster

The rapid spread of multimodal misinformation on social media has raised growing concerns, while research on video misinformation detection remains limited due to the lack of large-scale, diverse datasets. Existing methods often overfit to rigid templates and lack deep reasoning over deceptive conte…

Cited by 0SourcecodeScholar
2025

Innovative Thinking, Infinite Humor: Humor Research of Large Language Models through Structured Thought Leaps

ICLR 2025poster

Humor is previously regarded as a gift exclusive to humans for the following reasons. Humor is a culturally nuanced aspect of human language, presenting challenges for its understanding and generation. Humor generation necessitates a multi-hop reasoning process, with each hop founded on proper ratio…

Cited by 1SourcePDFScholar
2024

Dynamic Label Smoothing Strategy for Biosignal Classification

ICASSP 2024accepted

Biological signals classification is essential for human machine interaction. Although previous research has achieved high classification performance, compensating for domain shift due to the intra and inter individual variations remains a challenge. In this paper, we propose a novel dynamic label s…

Cited by 0SourceScholar
2024

Humtrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond

ICASSP 2024accepted

This paper introduces the HumTrans dataset, which is publicly available and primarily designed for humming melody transcription. The dataset can also serve as a foundation for downstream tasks such as humming melody based music generation. It consists of 500 musical compositions of different genres…

Cited by 0SourceScholar
2024

Unified Pretraining Target Based Video-Music Retrieval with Music Rhythm and Video Optical Flow Information

ICASSP 2024accepted

Background music (BGM) can enhance the video’s emotion. However, selecting an appropriate BGM often requires domain knowledge. This has led to the development of video-music retrieval techniques. Most existing approaches utilize pretrained video/music feature extractors trained with different target…

Cited by 0SourceScholar
2023

Masked Image Modeling with Denoising Contrast

ICLR 2023poster

Since the development of self-supervised visual representation learning from contrastive learning to masked image modeling (MIM), there is no significant difference in essence, that is, how to design proper pretext tasks for vision dictionary look-up. MIM recently dominates this line of research wit…

2023

RILS: Masked Visual Reconstruction in Language Semantic Space

CVPR 2023poster

Both masked image modeling (MIM) and natural language supervision have facilitated the progress of transferable visual pre-training. In this work, we seek the synergy between two paradigms and study the emerging properties when MIM meets natural language supervision. To this end, we present a novel…

2022

Bridging Video-Text Retrieval With Multiple Choice Questions

CVPR 2022oral

Pre-training a model to learn transferable video-text representation for retrieval has attracted a lot of attention in recent years. Previous dominant works mainly adopt two separate encoders for efficient retrieval, but ignore local associations between videos and texts. Another line of research us…

Cited by 179PDFcodeScholar
2022

CA-SSL: Class-Agnostic Semi-Supervised Learning for Detection and Segmentation

ECCV 2022poster

"To improve instance-level detection/segmentation performance, existing self-supervised and semi-supervised methods extract either very task-unrelated or very task-specific training signals from unlabeled data. We argue that these two approaches, at the two extreme ends of the task-specificity spect…

2022

Tencent-MVSE: A Large-Scale Benchmark Dataset for Multi-Modal Video Similarity Evaluation

CVPR 2022poster

Multi-modal video similarity evaluation is important for video recommendation systems such as video de-duplication, relevance matching, ranking, and diversity control. However, there still lacks a benchmark dataset that can support supervised training and accurate evaluation. In this paper, we propo…

Cited by 8PDFcodeScholar
2021

Enhancing Self-Supervised Video Representation Learning via Multi-Level Feature Optimization

ICCV 2021poster

The crux of self-supervised video representation learning is to build general features from unlabeled videos. However, most recent works have mainly focused on high-level semantics and neglected lower-level representations and their temporal relationship which are crucial for general video understan…

Cited by 34PDFcodeScholar