← Search

Rintaro Yanagi

3 accepted papers

2026

CLIP-like Model as a Foundational Density Ratio Estimator

CVPR 2026

Density ratio estimation is a core concept in statistical machine learning because it provides a unified mechanism for tasks such as importance weighting, divergence estimation, and likelihood-free inference, but its potential in vision and language models has not been fully explored. Modern vision-

Cited by 0SourcecodeScholar
2026

PowerCLIP: Powerset Alignment for Contrastive Pre-Training

CVPR 2026

Contrastive pre-training frameworks such as CLIP have demonstrated impressive zero-shot performance across a range of vision-language tasks. Recent studies have shown that aligning individual text tokens with specific image patches or regions enhances fine-grained compositional understanding. Howeve

Cited by 0SourcecodeScholar
2025

Approximate Domain Unlearning for Vision-Language Models

NeurIPS 2025spotlight

Pre-trained Vision-Language Models (VLMs) exhibit strong generalization capabilities, enabling them to recognize a wide range of objects across diverse domains without additional training. However, they often retain irrelevant information beyond the requirements of specific target downstream tasks,…

Cited by 0SourcecodeScholar