← Search

Yijia Zhong

4 accepted papers

2026

Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs

ICLR 2026poster

Multimodal Large Language Models (MLLMs) combine the linguistic strengths of LLMs with the ability to process multimodal data, enabling them to address a broader range of tasks. This progression highlights a shift from language-only reasoning to integrated vision–language reasoning in children's dev…

Cited by 0SourcecodeScholar
2026

Evaluating Generative Models via One-Dimensional Code Distributions

CVPR 2026

Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality. We instead evaluate models in the spac

Cited by 0SourcecodeScholar
2026

GROW: Watermark Generation with Progressive Guidance for Diffusion Models

CVPR 2026

Digital watermarking is a cornerstone for copyright protection. With the rapid advancement of generative models like diffusion models, in-generation and training-free watermarking techniques have garnered more attention for their endogeneity and convenience. These methods typically embed a watermark

Cited by 0SourceScholar
2026

Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity

CVPR 2026

Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authentic to real-world photography remains challenging. This is partly due to biases in existing evaluation paradigms: human ratings and preference-trained metrics often

Cited by 0SourcecodeScholar