← Search

Yinuo Jing

2 accepted papers

2026

EthoCLIP: Ontology-Enhanced Video-Language Pretraining for Animal Behavior Understanding

CVPR 2026

Vision-language models (VLMs) have achieved remarkable success across numerous domains, yet they lag significantly in animal behavior understanding due to severe data scarcity. Annotated animal behavior videos are prohibitively expensive and time-consuming to collect, requiring domain expertise and

Cited by 0SourcecodeScholar
2024

Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video Understanding

NeurIPS 2024poster

With the emergence of large pre-trained multimodal video models, multiple benchmarks have been proposed to evaluate model capabilities. However, most of the benchmarks are human-centric, with evaluation data and tasks centered around human applications. Animals are an integral part of the natural wo…