← Search

Qinya Li

4 accepted papers

2026

CONFCLIP: CONFIDENCE-WEIGHTED AND CLIPPED REWARD FOR REINFORCEMENT LEARNING IN LLMS

ICASSP 2026oral

Reinforcement learning (RL) has become a standard paradigm for refining large language models (LLMs) beyond pre-training and instruction tuning. A prominent line of work is RL with verifiable rewards (RLVR), which leverages automatically verifiable outcomes (e.g., correctness or executability) to ge…

Cited by 0SourcePDFScholar
2025

AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference

AAAI 2025technical

Long-context large language models (LLMs) inference is increasingly critical, motivating a number of studies devoted to alleviating the substantial storage and computational costs in such scenarios. Layer-wise skipping methods are promising optimizations but rarely explored in long-context inference…

2025

Adaptive Routing of Text-to-Image Generation Requests Between Large Cloud Model and Light-Weight Edge Model

ICCV 2025poster

Large text-to-image models demonstrate impressive generation capabilities; however, their substantial size necessitates expensive cloud servers for deployment. Conversely, light-weight models can be deployed on edge devices at lower cost but often with inferior generation quality for complex user pr…

Cited by 0SourcePDFScholar
2025

Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation

AAAI 2025technical

In an era overwhelmed by vast amounts of data, the effective curation of web-crawl datasets is essential for optimizing model performance. This paper tackles the challenges associated with the unstructured and heterogeneous nature of such datasets. Traditional heuristic curation methods often inadeq…

Cited by 0SourcePDFScholar