← Search

Songlin Fan

5 accepted papers

2026

MDiff4STR: Mask Diffusion Model for Scene Text Recognition

AAAI 2026technical

Mask Diffusion Models (MDMs) have recently emerged as a promising alternative to auto-regressive models (ARMs) for vision-language tasks, owing to their flexible balance of efficiency and accuracy. In this paper, for the first time, we introduce MDMs into the Scene Text Recognition (STR) task. We sh

Cited by 0SourcePDFScholar
2025

A Learning-based Multi-Frame Visual Feature Framework for Real-Time Driver Fatigue Detection

NAACL 2025system demonstrations

Driver fatigue is a significant factor contributing to road accidents, highlighting the need for reliable and accurate detection methods. In this study, we introduce a novel learning-based multi-frame visual feature framework (LMVFF) designed for precise fatigue detection. Our methodology comprises…

Cited by 0SourcePDFScholar
2025

Stochasticity-aware No-Reference Point Cloud Quality Assessment

IJCAI 2025

The evolution of point cloud processing algorithms necessitates an accurate assessment for their quality. Previous works consistently regard point cloud quality assessment (PCQA) as a MOS regression problem and devise a deterministic mapping, ignoring the stochasticity in generating MOS from subject

Cited by 0SourcePDFScholar
2025

VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment

AAAI 2025technical

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics for video editing are still notably absent. To address this,…