← Search

Jiaze Xu

3 accepted papers

2026

A Unified Framework for Knowledge Transfer in Bidirectional Model Scaling

CVPR 2026

Transferring pre-trained knowledge from a source model to a target model of a different architectural size is a key challenge for flexible and efficient model scaling. However, current parameter-space methods treat Small-to-Large (S2L) and Large-to-Small (L2S) scaling as separate, incompatible probl

Cited by 0SourceScholar
2026

Unlocking Pre-trained Weights: Parameter Inheritance for Zero-Shot Initialization

CVPR 2026

Appropriate parameter initialization is crucial for reducing the training cost of deep neural networks. Graph HyperNetworks (GHN) have emerged as a promising approach for initializing diverse architectures, with recent methods such as Task-Aware Learngene (TAL) further attempting to leverage pre-tra

Cited by 0SourcecodeScholar
2025

Learngene Tells You How to Customize: Task-Aware Parameter Initialization at Flexible Scales

ICML 2025poster

Appropriate parameter initialization strategies are essential for reducing the high computational costs of training large pretrained models in various task scenarios. Graph HyperNetwork (GHN), a parameter initialization method, has recently demonstrated strong performance in initializing models. How…

Cited by 0SourcePDFScholar