← Search

Reza Farahbakhsh

5 accepted papers

2026

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

IJCAI 2026

Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet incorrect statements, raising significant safety concerns. Existing medical hallucination benchmarks primarily focus on 2D imaging with one-shot diagnos

Cited by 0Scholar
2025

The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs

ACL 2025long

We present a novel class of jailbreak adversarial attacks on LLMs, termed Task-in-Prompt (TIP) attacks. Our approach embeds sequence-to-sequence tasks (e.g., cipher decoding, riddles, code execution) into the model’s prompt to indirectly generate prohibited inputs. To systematically assess the effec…

Cited by 0SourcePDFScholar
2025

Towards Cross-Lingual Audio Abuse Detection in Low-Resource Settings with Few-Shot Learning

COLING 2025main

Online abusive content detection, particularly in low-resource settings and within the audio modality, remains underexplored. We investigate the potential of pre-trained audio representations for detecting abusive language in low-resource languages, in this case, in Indian languages using Few Shot L…

2024

Improving Cross-lingual Transfer with Contrastive Negative Learning and Self-training

COLING 2024main

Recent studies improve the cross-lingual transfer learning by better aligning the internal representations within the multilingual model or exploring the information of the target language using self-training. However, the alignment-based methods exhibit intrinsic limitations such as non-transferabl…

Cited by 1SourcePDFScholar
2023

No offence, Bert - I insult only humans! Multilingual sentence-level attack on toxicity detection networks

EMNLP 2023short findings

We introduce a simple yet efficient sentence-level attack on black-box toxicity detector models. By adding several positive words or sentences to the end of a hateful message, we are able to change the prediction of a neural network and pass the toxicity detection system check. This approach is show…

Cited by 0SourceScholar