Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization
This work focuses on the task of query-based meeting summarization in which the summary of a context (meeting transcript) is generated in response to a specific query. When using Large Language Models (LLMs) for this task, a new call to the LLM inference endpoint/API is required for each new query e…