← Search

Risa Shinoda

5 accepted papers

2026

ANIMALCLAP: TAXONOMY-AWARE LANGUAGE-AUDIO PRETRAINING FOR SPECIES RECOGNITION AND TRAIT INFERENCE

ICASSP 2026poster

Animal vocalizations provide crucial insights for wildlife assessment, particularly in complex environments such as forests, aiding species identification and ecological monitoring. Recent advances in deep learning have enabled automatic species classification from their vocalizations. However, clas…

Cited by 0SourcePDFScholar
2026

BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment

CVPR 2026

Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology.While recent biological models, such as BioCLIP, have demonstrated strong alignment between images and textual taxonomic information for species identification, the integr

Cited by 0SourcecodeScholar
2025

AgroBench: Vision-Language Model Benchmark in Agriculture

ICCV 2025poster

Precise automated understanding of agricultural tasks such as disease identification is essential for the sustainable crop production. Recent advances in vision-language models (VLMs) are expected to further expand the range of agricultural tasks by facilitating human-model interaction through easy,…

2025

AnimalClue: Recognizing Animals by their Traces

ICCV 2025poster

Wildlife observation plays an important role in biodiversity conservation, necessitating robust methodologies for monitoring wildlife populations and interspecies interactions. Recent advances in computer vision have significantly contributed to automating fundamental wildlife observation tasks, suc…

2023

SegRCDB: Semantic Segmentation via Formula-Driven Supervised Learning

ICCV 2023poster

Pre-training is a strong strategy for enhancing visual models to efficiently train them with a limited number of labeled images. In semantic segmentation, creating annotation masks requires an intensive amount of labor and time, and therefore, a large-scale pre-training dataset with semantic labels…

Cited by 13PDFcodeScholar