← Search

Yu Hou

5 accepted papers

2026

Tokenize Once, Recommend Anywhere: Unified Item Tokenization for Multi-domain LLM-based Recommendation

AAAI 2026technical

Large language model (LLM)-based recommender systems have achieved high-quality performance by bridging the discrepancy between the item space and the language space through item tokenization. However, existing item tokenization methods typically require training separate models for each item domain

Cited by 0SourcePDFScholar
2025

GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration

ACL 2025long

Language models are often miscalibrated, leading to confidently incorrect answers. We introduce GRACE, a benchmark for language model calibration that incorporates comparison with human calibration. GRACE consists of question-answer pairs, in which each question contains a series of clues that gradu…

Cited by 0SourcePDFScholar
2025

Language Models Predict Empathy Gaps Between Social In-groups and Out-groups

NAACL 2025long

Studies of human psychology have demonstrated that people are more motivated to extend empathy to in-group members than out-group members (Cikara et al., 2011). In this study, we investigate how this aspect of intergroup relations in humans is replicated by LLMs in an emotion intensity prediction ta…

2025

Natural Language Inference Improves Compositionality in Vision-Language Models

ICLR 2025poster

Compositional reasoning in Vision-Language Models (VLMs) remains challenging as these models often struggle to relate objects, attributes, and spatial relationships. Recent methods aim to address these limitations by relying on the semantics of the textual description, using Large Language Models (L…

Cited by 3SourcePDFScholar
2023

ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos

EMNLP 2023long main

Multimodal counterfactual reasoning is a vital yet challenging ability for AI systems. It involves predicting the outcomes of hypothetical circumstances based on vision and language inputs, which enables AI models to learn from failures and explore hypothetical scenarios. Despite its importance, the…

Cited by 0SourcecodeScholar