← Search

Yunqi Hong

3 accepted papers

2026

When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models

ICML 2026poster

Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling is the Bradley-Terry (BT) loss, which learns from pairwise data consisting of a pair of chosen and rejected responses. In this work, we analyze the per-sampl…

Cited by 0SourceScholar
2025

QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models

EMNLP 2025

Recently, Multimodal Large Language Models (MLLMs) encounter two key issues in multi-image contexts: (1) a lack of fine-grained perception across disparate images, and (2) a diminished capability to effectively reason over and synthesize information from multiple visual inputs. However, while variou

Cited by 0SourcePDFScholar
2025

Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs

NeurIPS 2025poster

Despite Multimodal Large Language Models (MLLMs) showing promising results on general zero-shot image classification tasks, fine-grained image classification remains challenging. It demands precise attention to subtle visual details to distinguish between visually similar subcategories—details that…

Cited by 0SourceScholar