← Search

Yifan Chang

7 accepted papers

2026

A High Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation

ICLR 2026poster

Recent advancements in Large Multimodal Models (LMMs) have significantly improved multimodal understanding and generation. However, these models still struggle to generate tightly interleaved image-text outputs, primarily due to the limited scale, quality and instructional richness of current traini…

Cited by 0SourceScholar
2026

Closing the Expression Gap in LLM Instructions via Socratic Questioning

ICML 2026poster

A fundamental bottleneck in human-AI collaboration is the "intention expression gap", the difficulty for humans to effectively convey complex, high-dimensional thoughts to AI. This challenge often traps users in inefficient trial-and-error loops and is exacerbated by the diverse expertise levels of …

Cited by 0SourceScholar
2026

ProSoftArena: Benchmarking Hierarchical Capabilities of Multi-modal Agents in Professional Software Environments

CVPR 2026

Multi-modal agents are making rapid progress on general computer-use tasks. However, existing benchmarks remain largely confined to web browsers and rudimentary applications, failing to capture the professional software workflows that dominate real-world scientific and industrial practices. To bridg

Cited by 0SourcecodeScholar
2026

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization

ICLR 2026poster

Vector quantization (VQ) is a key component in discrete tokenizers for image generation, but its training is often unstable due to straight-through estimation bias, one-step-behind updates, and sparse codebook gradients, which lead to suboptimal reconstruction performance and low codebook usage. In…

Cited by 0SourceScholar
2025

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

ICLR 2025poster

Autonomous agents that execute human tasks by controlling computers can enhance human productivity and application accessibility. However, progress in this field will be driven by realistic and reproducible benchmarks. We present AndroidWorld, a fully functional Android environment that provides rew…

2025

Rethinking Lanes and Points in Complex Scenarios for Monocular 3D Lane Detection

CVPR 2025poster

Monocular 3D lane detection is a fundamental task in autonomous driving. Although sparse-point methods lower computational load and maintain high accuracy in complex lane geometries, current methods fail to fully leverage the geometric structure of lanes in both lane geometry representations and mod…

Cited by 0SourcePDFScholar