← Search

Juntong Wu

4 accepted papers

2026

SERE: Similarity-based Expert Re-routing for Efficient Batch Decoding in MoE Models

ICLR 2026poster

Mixture-of-Experts (MoE) architectures employ sparse activation to deliver faster training and inference with higher accuracy than dense LLMs. However, in production serving, MoE models require batch inference to optimize hardware efficiency, which may cause excessive expert activation and thus slow…

Cited by 0SourceScholar
2026

Task-Aware Mechanism: Hybrid MoE Vision Tower Towards Holistic Video Understanding

ICML 2026poster

Does \emph{Comprehending the main idea of a 2-hour movie} and \emph{Counting the birds appearing in a 15-second clip} really warrant the same video processing pipeline? We present Task-Aware Mechanism (TAM), a hybrid-gated Mixture-of-Experts (MoE) vision tower that adapts frame count and resolution …

Cited by 0SourceScholar
2026

WaveFormer: Frequency-Time Decoupled Vision Modeling with Wave Equation

AAAI 2026technical

Vision modeling has advanced rapidly with Transformers, whose attention mechanisms capture visual dependencies but lack a principled account of how semantic information propagates spatially. We revisit this problem from a wave-based perspective: feature maps are treated as spatial signals whose evol

Cited by 0SourcePDFScholar
2025

Rethinking Text-based Protein Understanding: Retrieval or LLM?

EMNLP 2025

In recent years, protein-text models have gained significant attention for their potential in protein generation and understanding. Current approaches focus on integrating protein-related knowledge into large language models through continued pretraining and multi-modal alignment, enabling simultane