← Search

Chris Thomas

11 accepted papers

2026

Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization Dynamics

CVPR 2026

Fine-tuning has become the default way to adapt powerful foundation models, but this also enables low-cost repurposing for harmful objectives. Existing immunization methods try to optimize local geometry or simulate short attacker horizons, and penalize observed loss drops. However, in practice, dow

Cited by 0SourceScholar
2026

LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have achieved remarkable performance across vision-language tasks. Recent advancements allow these models to process multiple images as inputs. However, the vulnerabilities of multi-image MLLMs remain unexplored. Existing adversarial attacks focus on single-i

Cited by 0SourcePDFScholar
2025

Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models

EMNLP 2025

Large Vision-Language Models (LVLMs) have achieved strong performance on vision-language tasks, particularly Visual Question Answering (VQA). While prior work has explored unimodal biases in VQA, the problem of selection bias in Multiple-Choice Question Answering (MCQA), where models may favor speci

2025

Flexible-length Text Infilling for Discrete Diffusion Models

EMNLP 2025

Discrete diffusion models are a new class of text generators that offer advantages such as bidirectional context use, parallelizable generation, and flexible prompting compared to autoregressive models. However, a critical limitation of discrete diffusion models is their inability to perform flexibl

Cited by 0SourcePDFScholar
2025

Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

ACL 2025long

Cross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities. Traditional methods learn a single-vector embedding to represent semantics of each sample, but struggle to capture nuanced and diverse relationships that can exist…

Cited by 0SourcePDFScholar
2025

SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models

EMNLP 2025

This work introduces SteerVLM, a lightweight steering module designed to guide Vision-Language Models (VLMs) towards outputs that better adhere to desired instructions. Our approach learns from the latent embeddings of paired prompts encoding target and converse behaviors to dynamically adjust activ

Cited by 0SourcePDFScholar
2025

Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models

EMNLP 2025

Large Vision-Language Models (LVLMs) have demonstrated impressive performance on vision-language reasoning tasks. However, their potential for zero-shot fine-grained image classification, a challenging task requiring precise differentiation between visually similar categories, remains underexplored.

2024

JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images

NeurIPS 2024poster

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying on background language biases. Thus, strong performance on th…

2024

M3D: MultiModal MultiDocument Fine-Grained Inconsistency Detection

EMNLP 2024main

Fact-checking claims is a highly laborious task that involves understanding how each factual assertion within the claim relates to a set of trusted source materials. Existing approaches make sample-level predictions but fail to identify the specific aspects of the claim that are troublesome and the…

Cited by 0SourcePDFScholar
2024

MetaSumPerceiver: Multimodal Multi-Document Evidence Summarization for Fact-Checking

ACL 2024long

Fact-checking real-world claims often requires reviewing multiple multimodal documents in order to assess the claim’s truthfulness, a highly laborious and time-consuming task. In this paper, we present a summarization model crafted to generate claim-specific summaries useful for fact-checking from m…