← Search

Anushka Sivakumar

4 accepted papers

2025

Flexible-length Text Infilling for Discrete Diffusion Models

EMNLP 2025

Discrete diffusion models are a new class of text generators that offer advantages such as bidirectional context use, parallelizable generation, and flexible prompting compared to autoregressive models. However, a critical limitation of discrete diffusion models is their inability to perform flexibl

Cited by 0SourcePDFScholar
2025

Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

ACL 2025long

Cross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities. Traditional methods learn a single-vector embedding to represent semantics of each sample, but struggle to capture nuanced and diverse relationships that can exist…

Cited by 0SourcePDFScholar
2025

SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models

EMNLP 2025

This work introduces SteerVLM, a lightweight steering module designed to guide Vision-Language Models (VLMs) towards outputs that better adhere to desired instructions. Our approach learns from the latent embeddings of paired prompts encoding target and converse behaviors to dynamically adjust activ

Cited by 0SourcePDFScholar
2024

JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images

NeurIPS 2024poster

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying on background language biases. Thus, strong performance on th…