← Search

Lei Gao

12 accepted papers

2026

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing

ICML 2026poster

Modern LLM serving systems must sustain high throughput while meeting strict latency SLOs across two distinct inference phases: compute-intensive prefill and memory-bound decode phases. Existing approaches either (1) aggregate both phases on shared GPUs, leading to interference between prefill and d…

Cited by 0SourceScholar
2026

One4Many-StablePacker: An Efficient Deep Reinforcement Learning Framework for the 3D Bin Packing Problem

IJCAI 2026

The three-dimensional bin packing problem (3D-BPP) is widely applied in logistics and warehousing. However, existing learning-based approaches often neglect stability constraints and struggle to generalize across diverse bin dimensions. To address this, we propose a novel deep reinforcement learning

Cited by 0Scholar
2026

Segment-Aligned Policy Optimization for Multi-Modal Reasoning

ICML 2026poster

Existing reinforcement learning approaches for Large Language Models typically perform policy optimization at the granularity of individual tokens or entire response sequences. However, such formulations often misalign with the natural step-wise structure of reasoning processes, leading to suboptima…

Cited by 0SourceScholar
2025

Gradient-Guided Credit Assignment and Joint Optimization for Dependency-Aware Spatial Crowdsourcing

AAAI 2025technical

Dependency-aware spatial crowdsourcing (DASC) addresses the unique challenges posed by subtask dependencies in spatial task assignment. This paper investigates the task assignment problem in DASC and proposes a two-stage Recommend and Match Optimization (RMO) framework, leveraging multi-agent reinfo…

2025

KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation

ACL 2025finding

Inference for Large Language Models (LLMs) is computationally demanding. To reduce the cost of auto-regressive decoding, Key-Value (KV) cache is used to store intermediate activations, which significantly lowers the computational overhead for token generation. However, the memory required for the KV…

2025

MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines

EMNLP 2025

Large Language Models (LLMs) are currently pre-trained and fine-tuned on large cloud servers. The next frontier is LLM personalization, where a foundation model can be fine-tuned with user/task-specific data. Given the sensitive nature of such private data, it is desirable to fine-tune these models

2024

Ethos: Rectifying Language Models in Orthogonal Parameter Space

NAACL 2024findings

Language models (LMs) have greatly propelled the research on natural language processing. However, LMs also raise concerns regarding the generation of biased or toxic content and the potential disclosure of private information from the training dataset. In this work, we present a new efficient appro…

2023

Dialogue Medical Information Extraction with Medical-Item Graph and Dialogue-Status Enriched Representation

EMNLP 2023long findings

The multi-turn doctor-patient dialogue includes rich medical knowledge, like the symptoms of the patient, the diagnosis and medication suggested by the doctor. If mined and represented properly, such medical knowledge can benefit a large range of clinical applications, including diagnosis assistance…

Cited by 0SourceScholar
2022

MAG+: An Extended Multimodal Adaptation Gate for Multimodal Sentiment Analysis

ICASSP 2022accepted

Human multimodal sentiment analysis is a challenging task that devotes to extract and integrate information from multiple resources, such as language, acoustic and visual information. Recently, multimodal adaptation gate (MAG), an attachment to transformer-based pre-trained language representation m…

Cited by 0SourceScholar
2021

Question-Driven Span Labeling Model for Aspect–Opinion Pair Extraction

AAAI 2021technical

Aspect term extraction and opinion word extraction are two fundamental subtasks of aspect-based sentiment analysis. The internal relationship between aspect terms and opinion words is typically ignored, and information for the decision-making of buyers and sellers is insufficient. In this paper, we…

Cited by 80SourcePDFScholar
2016

Information fusion based on kernel entropy component analysis in discriminative canonical correlation space with application to audio emotion recognition

ICASSP 2016accepted

As an information fusion tool, Kernel Entropy Component Analysis (KECA) is realized by using descriptor of information entropy and optimized by entropy estimation. However, as an unsuper-vised method, it merely puts the information or features from different channels together without considering the…

Cited by 0SourceScholar