← Search

Shengzhuang Chen

5 accepted papers

2025

Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts

ACL 2025long

We present Sparse Interpolated Mixture-of-Experts (SIMoE) instruction-tuning, an end-to-end algorithm designed to fine-tune a dense pre-trained Large Language Model (LLM) into a MoE-style model that possesses capabilities in multiple specialized domains. During instruction-tuning, SIMoE automaticall…

Cited by 0SourcePDFScholar
2025

CLDyB: Towards Dynamic Benchmarking for Continual Learning with Pre-trained Models

ICLR 2025poster

The emergence of the foundation model era has sparked immense research interest in utilizing pre-trained representations for continual learning~(CL), yielding a series of strong CL methods with outstanding performance on standard evaluation benchmarks. Nonetheless, there are growing concerns regardi…

2024

Learning Where to Edit Vision Transformers

NeurIPS 2024poster

Model editing aims to data-efficiently correct predictive errors of large pre-trained models while ensuring generalization to neighboring failures and locality to minimize unintended effects on unrelated examples. While significant progress has been made in editing Transformer-based large language m…

2024

Unleashing the Power of Meta-tuning for Few-shot Generalization Through Sparse Interpolated Experts

ICML 2024poster

Recent successes suggest that parameter-efficient fine-tuning of foundation models is becoming the state-of-the-art method for transfer learning in vision, gradually replacing the rich literature of alternatives such as meta-learning. In trying to harness the best of both worlds, meta-tuning introdu…

2023

Secure Out-of-Distribution Task Generalization with Energy-Based Models

NeurIPS 2023poster

The success of meta-learning on out-of-distribution (OOD) tasks in the wild has proved to be hit-and-miss. To safeguard the generalization capability of the meta-learned prior knowledge to OOD tasks, in particularly safety-critical applications, necessitates detection of an OOD task followed by adap…

Cited by 6SourcePDFScholar