← Search

Ruiming Chen

3 accepted papers

2026

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

AAAI 2026technical

CLIP (Contrastive Language-Image Pre-training) has attracted widespread attention for its multimodal generalizable knowledge, which is significant for downstream tasks. However, the computational overhead of a large number of parameters and large-scale pre-training poses challenges of pre-training a

Cited by 0SourcePDFScholar
2024

Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale Models

NeurIPS 2024poster

Vision Transformers (ViTs) are widely used in a variety of applications, while they usually have a fixed architecture that may not match the varying computational resources of different deployment environments. Thus, it is necessary to adapt ViT architectures to devices with diverse computational ov…

Cited by 3SourcePDFScholar
2024

Transformer as Linear Expansion of Learngene

AAAI 2024technical

We propose expanding the shared Transformer module to produce and initialize Transformers of varying depths, enabling adaptation to diverse resource constraints. Drawing an analogy to genetic expansibility, we term such module as learngene. To identify the expansion mechanism, we delve into the rela…