CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning
Image captioning remains a fundamental task for vision-language understanding, yet ground-truth supervision still relies predominantly on human-annotated references.Because human annotations reflect subjective preferences and expertise, ground-truth captions are often incomplete or even incorrect, w