← Search

Chunlin Zhong

3 accepted papers

2026

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward

AAAI 2026technical

Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation. However, existing methods often suffer from motion-detail imbalance, as models tend to overemphasize one aspect while neglecting

Cited by 0SourcePDFScholar
2025

Rethinking Detecting Salient and Camouflaged Objects in Unconstrained Scenes

ICCV 2025poster

While the human visual system employs distinct mechanisms to perceive salient and camouflaged objects, existing models struggle to disentangle these tasks. Specifically, salient object detection (SOD) models frequently misclassify camouflaged objects as salient, while camouflaged object detection (C…

2024

CoLA: Conditional Dropout and Language-driven Robust Dual-modal Salient Object Detection

ECCV 2024poster

"The depth/thermal information is beneficial for detecting salient object with conventional RGB images. However, in dual-modal salient object detection (SOD) model, the robustness against noisy inputs and modality missing is crucial but rarely studied. To tackle this problem, we introduce Conditiona…