INTERMITTENT SEMI-WORKING MASK: A NEW MASKING PARADIGM FOR LLMS
Multi-turn dialogues and context-intensive tasks challenge Large Language Models (LLMs) to integrate long histories without sacrificing generation quality. Although prefix LLMs can better exploit historical context via bidirectional attention on prefix tokens, they are rarely used in practice becaus…