← Search

Zhenyang Cai

8 accepted papers

2026

MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation

ICLR 2026poster

Recent advances in video generation have opened new avenues for macroscopic simulation of complex dynamic systems, but their application to microscopic phenomena remains largely unexplored. Microscale simulation holds great promise for biomedical applications such as drug discovery, organ-on-chip sy…

Cited by 0SourcecodeScholar
2025

Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization

EMNLP 2025

Materials characterization is fundamental to acquiring materials information, revealing the processing-microstructure-property relationships that guide material design and optimization. While multimodal large language models (MLLMs) have recently shown promise in generative and predictive tasks with

2025

Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging

ACL 2025long

Medical imaging provides essential visual insights for diagnosis, and multimodal large language models (MLLMs) are increasingly utilized for its analysis due to their strong generalization capabilities; however, the underlying factors driving this generalization remain unclear. Current research sugg…

2025

LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture

EMNLP 2025

Expanding the long-context capabilities of Multi-modal Large Language Models (MLLMs) is critical for advancing video understanding and high-resolution image analysis. Achieving this requires systematic improvements in model architecture, data construction, and training strategies, particularly to ad

2025

Towards Medical Complex Reasoning with LLMs through Medical Verifiable Problems

ACL 2025finding

The breakthrough of OpenAI o1 highlights the potential of enhancing reasoning to improve LLM. Yet, most research in reasoning has focused on mathematical tasks, leaving domains like medicine underexplored. The medical domain, though distinct from mathematics, also demands robust reasoning to provide…

2025

UCL-Bench: A Chinese User-Centric Legal Benchmark for Large Language Models

NAACL 2025findings

Existing legal benchmarks focusing on knowledge and logic effectively evaluate LLMs on various tasks in legal domain. However, few have explored the practical application of LLMs by actual users. To further assess whether LLMs meet the specific needs of legal practitioners in real-world scenarios, w…

2024

Alignment at Pre-training! Towards Native Alignment for Arabic LLMs

NeurIPS 2024poster

The alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as `\textit{post alignment}'. We argue that alignment…

2024

Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

EMNLP 2024main

The rapid development of multimodal large language models (MLLMs), such as GPT-4V, has led to significant advancements. However, these models still face challenges in medical multimodal capabilities due to limitations in the quantity and quality of medical vision-text data, stemming from data privac…