← Search

Wenjie Zhou

8 accepted papers

2026

Benchmarking Real-Time Question Answering via Executable Code Workflows

IJCAI 2026

Retrieving real-time information is a fundamental capability for search-integrated agents in real-world applications. However, existing benchmarks are predominantly static and therefore fail to capture the temporal dynamics of information and the continuously evolving nature of real-world knowledge.

Cited by 0Scholar
2026

Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training

ICML 2026poster

Model merging has emerged as a lightweight paradigm for enhancing Large Language Models (LLMs), yet its underlying mechanisms remain poorly understood. In this work, we analyze late-stage pre-training trajectories and uncover a \textbf{Rank-1 Subspace} phenomenon: while raw optimization steps oscill…

Cited by 0SourceScholar
2026

The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-training

ICML 2026poster

Large language model pre-training typically exhibits a two-phase trajectory: a fast initial loss drop followed by a prolonged slow improvement. We identify an underlying spectral phenomenon, Stability of Singular Distribution (SoSD), where the trace-normalized singular value spectrum stabilizes earl…

Cited by 0SourceScholar
2025

BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training

EMNLP 2025

Recent studies (CITATION) highlight a fundamental dichotomy in deep learning optimization: Although parameter updates along the top eigendirections of the loss Hessian (Dom-space) capture most of the update magnitude, they often contribute minimally to loss reduction. In contrast, updates in the ort

2024

GOVERN: Gradient Orientation Vote Ensemble for Multi-Teacher Reinforced Distillation

EMNLP 2024industry

Pre-trained language models have become an integral component of question-answering systems, achieving remarkable performance. However, for practical deployment, it is crucial to perform knowledge distillation to maintain high performance while operating under computational constraints. In this pape…

Cited by 1SourcePDFScholar
2024

Revisiting the Self-Consistency Challenges in Multi-Choice Question Formats for Large Language Model Evaluation

COLING 2024main

Multi-choice questions (MCQ) are a common method for assessing the world knowledge of large language models (LLMs), demonstrated by benchmarks such as MMLU and C-Eval. However, recent findings indicate that even top-tier LLMs, such as ChatGPT and GPT4, might display inconsistencies when faced with s…

Cited by 8SourcePDFScholar