CollectiveKV: Decoupling and Sharing Collaborative Information in Sequential Recommendation
Sequential recommendation models are widely used in applications, yet they face stringent latency requirements. Mainstream models leverage the Transformer attention mechanism to improve performance, but its computational complexity grows with the sequence length, leading to a latency challenge for…