← Search

Ajay Patel

11 accepted papers

2025

Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution

COLING 2025main

Recent state-of-the-art authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications. We propose a novel approach for interpreting learned embeddings by identifying representative points in the la…

Cited by 0SourcePDFScholar
2025

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

CVPR 2025award

Today's most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational…

2025

Quantifying Misattribution Unfairness in Authorship Attribution

ACL 2025short

Authorship misattribution can have profound consequences in real life. In forensic settings simply being considered as one of the potential authors of an evidential piece of text or communication can result in undesirable scrutiny. This raises a fairness question: Is every author in the candidate po…

2025

Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation

ACL 2025long

Reasoning about images with rich text, such as charts and documents, is a critical application of vision-language models (VLMs). However, VLMs often struggle in these domains due to the scarcity of diverse text-rich vision-language data. To address this challenge, we present CoSyn, a framework that…

Cited by 0SourcePDFScholar
2025

StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples

NAACL 2025long

Style representations aim to embed texts with similar writing styles closely and texts with different styles far apart, regardless of content. However, the contrastive triplets often used for training these representations may vary in both style and content, leading to potential content leakage in t…

Cited by 2SourcePDFScholar
2025

mStyleDistance: Multilingual Style Embeddings and their Evaluation

ACL 2025finding

Style embeddings are useful for stylistic analysis and style transfer, yet they only exist for English. We introduce Multilingual StyleDistance (mStyleDistance), a method that can generate style embeddings in new languages using synthetic data and a contrastive loss. We create style embeddings in ni…

Cited by 0SourcePDFScholar
2024

DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows

ACL 2024long

Large language models (LLMs) have become a dominant and important tool for NLP researchers in a wide range of tasks. Today, many researchers use LLMs in synthetic data generation, task evaluation, fine-tuning, distillation, and other model-in-the-loop research workflows. However, challenges arise wh…

2024

ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer

AAAI 2024technical

Textual style transfer is the task of transforming stylistic properties of text while preserving meaning. Target "styles" can be defined in numerous ways, ranging from single attributes (e.g. formality) to authorship (e.g. Shakespeare). Previous unsupervised style-transfer approaches generally rely…

2024

TinyStyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings

EMNLP 2024finding

The goal of text style transfer is to transform the style of texts while preserving their original meaning, often with only a few examples of the target style. Existing style transfer methods generally rely on the few-shot capabilities of large language models or on complex controllable text generat…

2023

Bidirectional Language Models Are Also Few-shot Learners

ICLR 2023poster

Large language models such as GPT-3 (Brown et al., 2020) can perform arbitrary tasks without undergoing fine-tuning after being prompted with only a few labeled examples. An arbitrary task can be reformulated as a natural language prompt, and a language model can be asked to generate the completion,…

Cited by 66SourcePDFScholar
2023

Learning Interpretable Style Embeddings via Prompting LLMs

EMNLP 2023long findings

Style representation learning builds content-independent representations of author style in text. To date, no large dataset of texts with stylometric annotations on a wide range of style dimensions has been compiled, perhaps because the linguistic expertise to perform such annotation would be prohib…

Cited by 0SourceScholar