← Search

Kaoyan Lu

2 accepted papers

2026

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

AAAI 2026technical

Human-defined creativity is highly abstract, posing a challenge for multimodal large language models (MLLMs) to comprehend and assess creativity that aligns with human judgments. The absence of an existing benchmark further exacerbates this dilemma. To this end, we propose CreBench, which consists o

Cited by 0SourcePDFScholar
2026

ERGeoBench: A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models

ICML 2026poster

Multimodal large language models (MLLMs) have shown strong potential for building embodied agents, yet embodied geo-localization remains underexplored due to the lack of fine-grained evaluation. We introduce ERGeoBench, a large-scale benchmark for vision-driven embodied geo-localization. ERGeoBench …

Cited by 0SourceScholar