← Search

Xuri Ge

8 accepted papers

2026

DRIVE: Distributional and Retrieval-Augmented Bidding with Value Evaluation

ICML 2026poster

Auto-bidding is a core component of real-time advertising systems, where decisions must optimize long-term performance under budget and cost constraints, while online exploration is prohibitively risky. Offline reinforcement learning and, more recently, Transformer-based sequence modeling have shown…

Cited by 1SourceScholar
2026

From Traits to Roles: Consensus-Guided Composition of Orthogonal Experts for Cooperative MARL

IJCAI 2026

Parameter sharing is a central design choice in cooperative multi-agent reinforcement learning, yet it fundamentally conflicts with the need for role specialization in heterogeneous cooperative environments. Existing role-based methods typically learn monolithic role representations, which often suf

Cited by 0Scholar
2026

LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation

ICASSP 2026poster

In an era where micro-videos dominate platforms like TikTok and YouTube, AI-generated content is nearing cinematic quality. The next frontier is using large language models (LLMs) to autonomously create viral micro-videos, a largely untapped potential that could shape the future of AI-driven content…

Cited by 0SourcePDFScholar
2026

Vision-language Incremental Learning with Dual Class-individual Memory

AAAI 2026technical

The emergence of multimodal technologies has propelled Vision-Language Incremental Learning (VLIL) into a research spotlight. Current VLIL approaches predominantly inherit unimodal paradigms, failing to address fundamental distinctions between visual and linguistic modalities. Crucially, the semanti

Cited by 0SourcePDFScholar
2025

Double-Filter: Efficient Fine-tuning of Pre-trained Vision-Language Models via Patch&Layer Filtering

ICML 2025poster

In this paper, we present a novel approach, termed Double-Filter,to “slim down” the fine-tuning process of vision-language pre-trained (VLP) models via filtering redundancies in feature inputs and architectural components. We enhance the fine-tuning process using two approaches. First, we develop a…

Cited by 0SourcePDFScholar
2023

Continuous Interaction with A Smart Speaker via Low-Dimensional Embeddings of Dynamic Hand Pose

ICASSP 2023accepted

This paper presents a new continuous interaction strategy with visual feedback of hand pose and mid-air gesture recognition and control for a smart music speaker, which utilizes only 2 video frames to recognize gestures. Frame-based hand pose features from MediaPipe Hands, containing 21 landmarks, a…

Cited by 0SourceScholar
2023

Multi-Local Attention for Speech-Based Depression Detection

ICASSP 2023accepted

This article shows that an attention mechanism, the Multi-Local Attention, can improve a depression detection approach based on Long Short-Term Memory Networks. Besides leading to higher performance metrics (e.g., Accuracy and F1 Score), Multi-Local Attention improves two other aspects of the approa…

Cited by 0SourceScholar
2019

Variational Structured Semantic Inference for Diverse Image Captioning

NeurIPS 2019poster

Despite the exciting progress in image captioning, generating diverse captions for a given image remains as an open problem. Existing methods typically apply generative models such as Variational Auto-Encoder to diversify the captions, which however neglect two key factors of diverse expression, i.e…