← Search

Yixiong Fang

3 accepted papers

2026

GameDevBench: Evaluating Agentic Capabilities Through Game Development

ICML 2026poster

While coding agents have advanced rapidly, progress on multimodal agents has lagged behind, largely due to a gap between the unimodal nature of code and other multimodal computer applications. Game development bridges the modality gap, mirroring software development's complexity in terms of large co…

Cited by 0SourceScholar
2025

Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding

NeurIPS 2025poster

Large vision-language models (LVLMs) excel at multimodal tasks but are prone to misinterpreting visual inputs, often resulting in hallucinations and unreliable outputs. We present Dropout Decoding, a novel inference-time approach that quantifies the uncertainty of visual tokens and selectively masks…

Cited by 0SourceScholar
2025

LastingBench: Defend Benchmarks Against Knowledge Leakage

EMNLP 2025

The increasing size and complexity of large language models (LLMs) raise concerns about their ability to “cheat” on standard Question Answering (QA) benchmarks by memorizing task-specific data. This undermines the validity of benchmark evaluations, as they no longer reflect genuine model capabilitie

Cited by 0SourcePDFScholar