2026
EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior Understanding
CVPR 2026
Vision-language models (VLMs) have achieved remarkable success across numerous domains, yet they lag significantly in animal behavior understanding due to severe data scarcity. Annotated animal behavior videos are prohibitively expensive and time-consuming to collect, requiring domain expertise and