← Search

Akshay Rao

1 accepted papers

2026

Learning complete and explainable visual representations from itemized text supervision

CVPR 2026

Training vision models with language supervision enables general and transferable representations. However, many visual domains, especially non-object-centric domains such as medical imaging and remote sensing, contain itemized text annotations: multiple text items describing distinct and semantical

Cited by 0SourcecodeScholar