← Search

Andrew Zhang

6 accepted papers

2025

Flexible-length Text Infilling for Discrete Diffusion Models

EMNLP 2025

Discrete diffusion models are a new class of text generators that offer advantages such as bidirectional context use, parallelizable generation, and flexible prompting compared to autoregressive models. However, a critical limitation of discrete diffusion models is their inability to perform flexibl

Cited by 0SourcePDFScholar
2025

Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval

ACL 2025long

Cross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities. Traditional methods learn a single-vector embedding to represent semantics of each sample, but struggle to capture nuanced and diverse relationships that can exist…

Cited by 0SourcePDFScholar
2025

SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models

EMNLP 2025

This work introduces SteerVLM, a lightweight steering module designed to guide Vision-Language Models (VLMs) towards outputs that better adhere to desired instructions. Our approach learns from the latent embeddings of paired prompts encoding target and converse behaviors to dynamically adjust activ

Cited by 0SourcePDFScholar
2025

Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models

EMNLP 2025

Large Vision-Language Models (LVLMs) have demonstrated impressive performance on vision-language reasoning tasks. However, their potential for zero-shot fine-grained image classification, a challenging task requiring precise differentiation between visually similar categories, remains underexplored.

2024

Multistain Pretraining for Slide Representation Learning in Pathology

ECCV 2024poster

"Developing self-supervised learning (SSL) models that can learn universal and transferable representations of H&E gigapixel whole-slide images (WSIs) is becoming increasingly valuable in computational pathology. These models hold the potential to advance critical tasks such as few-shot classificati…

2023

Visual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology Images

CVPR 2023poster

Contrastive visual language pretraining has emerged as a powerful method for either training new language-aware image encoders or augmenting existing pretrained models with zero-shot visual recognition capabilities. However, existing works typically train on large datasets of image-text pairs and ha…