2025
SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMs
NeurIPS 2025poster
Recent multimodal large language models (MLLMs) marry modality-specific vision or audio encoders with a shared text decoder. While the encoder is compute- intensive but memory-light, the decoder is the opposite, yet state-of-the-art serving stacks still time-multiplex these complementary kernels, id…