← Search

Bardia Safaei

4 accepted papers

2025

Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning

CVPR 2025highlight

Visual instruction tuning (VIT) for large vision-language models (LVLMs) requires training on expansive datasets of image-instruction pairs, which can be costly. Recent efforts in VIT data selection aim to select a small subset of high-quality image-instruction pairs, reducing VIT runtime while main…

2025

Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models

CVPR 2025highlight

Zero-Shot Anomaly Detection (ZSAD) is an emerging AD paradigm. Unlike the traditional unsupervised AD setting that requires a large number of normal samples to train a model, ZSAD is more practical for handling data-restricted real-world scenarios. Recently, Multimodal Large Language Models (MLLMs)…