← Search

Abduljalil Radman

3 accepted papers

2025

CAMEL-Bench: A Comprehensive Arabic LMM Benchmark

NAACL 2025findings

Recent years have witnessed a significant interest in developing large multi-modal models (LMMs) capable of performing various visual reasoning and understanding tasks. This has led to the introduction of multiple LMM benchmarks to evaluate LMMs on different tasks. However, most existing LMM evaluat…

2025

Learning to Describe Implicit Changes: Noise-robust Pre-training for Image Difference Captioning

EMNLP 2025

Image Difference Captioning (IDC) methods have advanced in highlighting subtle differences between similar images, but their performance is often constrained by limited training data. Using Large Multimodal Models (LMMs) to describe changes in image pairs mitigates data limits but adds noise. These

Cited by 0SourcePDFScholar
2025

TSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual Segmentation

CVPR 2025poster

Referring audio-visual segmentation (Ref-AVS) aims to segment objects within audio-visual scenes using multimodal cues embedded in text expressions. While the Segment Anything Model (SAM) has revolutionized visual segmentation, its applicability to Ref-AVS, where multimodal cues act as novel prompts…

Cited by 0SourcePDFScholar