Subject-Agnostic Cross-View Referring Expression Comprehension via Reward-Driven Consistency Learning
Cross-view Referring Expression Comprehension (REC) requires a model to accurately localize objects in a spatial representation of one perspective based on descriptions generated from a different perspective. Essentially, it demands maintaining referential consistency under perspective shifts. This