← Search

Songshuo Lu

3 accepted papers

2026

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models

AAAI 2026technical

Large-scale alignment pipelines typically pair a policy model with a separately trained reward model whose parameters remain frozen during reinforcement learning (RL). This separation creates a complex, resource-intensive pipeline and leads to a performance ceiling. We propose a novel framework, Uni

Cited by 0SourcePDFScholar
2025

TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text

EMNLP 2025

Current Retrieval-Augmented Generation (RAG) systems concatenate and process numerous retrieved document chunks for prefill which requires a large volume of computation, therefore leading to significant latency in time-to-first-token (TTFT). To reduce the computation overhead as well as TTFT, we int

2024

A Full-duplex Speech Dialogue Scheme Based On Large Language Model

NeurIPS 2024poster

We present a generative dialogue system capable of operating in a full-duplex manner, allowing for seamless interaction. It is based on a large language model (LLM) carefully aligned to be aware of a perception module, a motor function module, and the concept of a simple finite state machine (called…

Cited by 15SourcePDFScholar