← Search

Sanghyeok Chu

5 accepted papers

2026

Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling

CVPR 2026

Sparse Upcycling provides an efficient way to initialize a Mixture-of-Experts (MoE) model from pretrained dense weights instead of training from scratch. However, since all experts start from identical weights and the router is randomly initialized, the model suffers from expert symmetry and limited

Cited by 0SourceScholar
2021

Learning Debiased and Disentangled Representations for Semantic Segmentation

NeurIPS 2021poster

Deep neural networks are susceptible to learn biased models with entangled feature representations, which may lead to subpar performances on various downstream tasks. This is particularly true for under-represented classes, where a lack of diversity in the data exacerbates the tendency. This limitat…

Cited by 24SourcePDFScholar