← Search

Lyle Ungar

23 accepted papers

2026

Correctness-Optimized Residual Activation Lens (CORAL): Transferrable and Calibration-Aware Inference-Time Steering

ICML 2026poster

Large language models (LLMs) exhibit persistent miscalibration, especially after instruction tuning and preference alignment. Modified training objectives can improve calibration, but retraining is expensive. Inference-time steering offers a lightweight alternative, yet most existing methods optimiz…

Cited by 0SourceScholar
2026

LaTeX2Layout: High-Fidelity, Scalable Document Layout Annotation Pipeline for Layout Detection

AAAI 2026technical

General-purpose Vision-Language Models (VLMs) are increasingly integral to modern AI systems for document understanding, yet their ability to perform fine-grained layout analysis remains severely underdeveloped. Overcoming this limitation requires large-scale, high-fidelity training datasets. Howeve

Cited by 0SourcePDFScholar
2026

ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training

ICLR 2026poster

Token-level attention tuning -- a class of training-free methods including Post-hoc Attention Steering (PASTA) and Attention Calibration (ACT) -- has emerged as a promising approach for improving frozen LLMs via interpretable interventions. However, these methods rely on auxiliary heuristics to iden…

Cited by 0SourceScholar
2025

ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations

EMNLP 2025

This work demonstrates that diffusion models can achieve font-controllable multilingual text rendering using just raw images without font label annotations. Visual text rendering remains a significant challenge. While recent methods condition diffusion on glyphs, it is impossible to retrieve exact f

2025

Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?

NeurIPS 2025spotlight

Object binding, the brain’s ability to bind the many features that collectively represent an object into a coherent whole, is central to human cognition. It groups low-level perceptual features into high‑level object representations, stores those objects efficiently and compositionally in memory, an…

Cited by 0SourceScholar
2025

Language-based Valence and Arousal Expressions between the United States and China: a Cross-Cultural Examination

NAACL 2025findings

While affective expressions on social media have been extensively studied, most research has focused on the Western context. This paper explores cultural differences in affective expressions by comparing valence and arousal on Twitter/X (geolocated to the US) and Sina Weibo (in Mainland China). Usin…

2025

Social Norms in Cinema: A Cross-Cultural Analysis of Shame, Pride and Prejudice

NAACL 2025long

Shame and pride are social emotions expressed across cultures to motivate and regulate people’s thoughts, feelings, and behaviors. In this paper, we introduce the first cross-cultural dataset of over 10k shame/pride-related expressions with underlying social expectations from ~5.4K Bollywood and Hol…

Cited by 1SourcePDFScholar
2025

Systematic Evaluation of Auto-Encoding and Large Language Model Representations for Capturing Author States and Traits

ACL 2025finding

Large Language Models (LLMs) are increasingly used in human-centered applications, yet their ability to model diverse psychological constructs is not well understood. In this study, we systematically evaluate a range of Transformer-LMs to predict psychological variables across five major dimensions:…

Cited by 0SourcePDFScholar
2025

The Illusion of Empathy: How AI Chatbots Shape Conversation Perception

AAAI 2025technical

As AI chatbots increasingly incorporate empathy, understanding user-centered perceptions of chatbot empathy and its impact on conversation quality remains essential yet under-explored. This study examines how chatbot identity and perceived empathy influence users' overall conversation experience. An…

2024

Building Knowledge-Guided Lexica to Model Cultural Variation

NAACL 2024long

Cultural variation exists between nations (e.g., the United States vs. China), but also within regions (e.g., California vs. Texas, Los Angeles vs. San Francisco). Measuring this regional cultural variation can illuminate how and why people think and behave differently. Historically, it has been dif…

2024

Modeling Human Subjectivity in LLMs Using Explicit and Implicit Human Factors in Personas

EMNLP 2024finding

Large language models (LLMs) are increasingly being used in human-centered social scientific tasks, such as data annotation, synthetic data creation, and engaging in dialog. However, these tasks are highly subjective and dependent on human factors, such as one’s environment, attitudes, beliefs, and…

Cited by 4SourcePDFScholar
2023

An Integrative Survey on Mental Health Conversational Agents to Bridge Computer Science and Medical Perspectives

EMNLP 2023long main

Mental health conversational agents (a.k.a. chatbots) are widely studied for their potential to offer accessible support to those experiencing mental health challenges. Previous surveys on the topic primarily consider papers published in either computer science or medicine, leading to a divide in un…

Cited by 0SourcecodeScholar
2023

Interactive Concept Learning for Uncovering Latent Themes in Large Text Collections

ACL 2023findings

Experts across diverse disciplines are often interested in making sense of large text collections. Traditionally, this challenge is approached either by noisy unsupervised techniques such as topic models, or by following a manual theme discovery process. In this paper, we expand the definition of a…

Cited by 17SourcePDFScholar
2022

A Holistic Framework for Analyzing the COVID-19 Vaccine Debate

NAACL 2022long

The Covid-19 pandemic has led to infodemic of low quality information leading to poor health decisions. Combating the outcomes of this infodemic is not only a question of identifying false claims, but also reasoning about the decisions individuals make. In this work we propose a holistic analysis fr…

2022

Inducing Generalizable and Interpretable Lexica

EMNLP 2022finding

Lexica – words and associated scores – are widely used as simple, interpretable, generalizable language features to predict sentiment, emotions, mental health, and personality. They also provide insight into the psychological features behind those moods and traits. Such lexica, historically created…

2022

Measuring the Language of Self-Disclosure across Corpora

ACL 2022findings

Being able to reliably estimate self-disclosure – a key component of friendship and intimacy – from language is important for many psychology studies. We build single-task models on five self-disclosure corpora, but find that these models generalize poorly; the within-domain accuracy of predicted me…

2021

Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica

EMNLP 2021main

People convey their intention and attitude through linguistic styles of the text that they write. In this study, we investigate lexicon usages across styles throughout two lenses: human perception and machine word importance, since words differ in the strength of the stylistic cues that they provide…

2021

WikiTalkEdit: A Dataset for modeling Editors’ behaviors on Wikipedia

NAACL 2021long

This study introduces and analyzes WikiTalkEdit, a dataset of conversations and edit histories from Wikipedia, for research in online cooperation and conversation modeling. The dataset comprises dialog triplets from the Wikipedia Talk pages, and editing actions on the corresponding articles being di…