← Search

Xi Liu

21 accepted papers

2026

Cross-Domain Policy Optimization via Bellman Consistency and Hybrid Critics

ICLR 2026poster

Cross-domain reinforcement learning (CDRL) is meant to improve the data efficiency of RL by leveraging the data samples collected from a source domain to facilitate the learning in a similar target domain. Despite its potential, cross-domain transfer in RL is known to have two fundamental and intert…

Cited by 0SourceScholar
2026

HAD: Hallucination-Aware Diffusion Priors for 3D Reconstruction

CVPR 2026

Diffusion priors have recently demonstrated strong capability in enhancing the quality of sparse-view 3D reconstruction by augmenting training views at novel viewpoints, but they inevitably introduce hallucinated content-- artifacts inconsistent with the input views -- into the final 3D model. To ad

Cited by 0SourceScholar
2025

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations

ICCV 2025poster

Face-to-face communication, as a common human activity, motivates the research on interactive head generation. A virtual agent can generate motion responses with both listening and speaking capabilities based on the audio or motion signals of the other user and itself. However, previous clip-wise ge…

2025

Bézier Splatting for Fast and Differentiable Vector Graphics Rendering

NeurIPS 2025poster

Differentiable vector graphics (VGs) are widely used in image vectorization and vector synthesis, while existing representations are costly to optimize and struggle to achieve high-quality rendering results for high-resolution images. This work introduces a new differentiable VG representation, dubb…

Cited by 0SourceScholar
2025

The Efficiency vs. Accuracy Trade-off: Optimizing RAG-Enhanced LLM Recommender Systems Using Multi-Head Early Exit

ACL 2025long

The deployment of Large Language Models (LLMs) in recommender systems for Click-Through Rate (CTR) prediction requires a careful balance between computational efficiency and predictive accuracy. This paper introduces OptiRAG-Rec, a comprehensive framework that integrates Retrieval-Augmented Generati…

Cited by 0SourcePDFScholar
2024

3DGS-Enhancer: Enhancing Unbounded 3D Gaussian Splatting with View-consistent 2D Diffusion Priors

NeurIPS 2024spotlight

Novel-view synthesis aims to generate novel views of a scene from multiple input images or videos, and recent advancements like 3D Gaussian splatting (3DGS) have achieved notable success in producing photorealistic renderings with efficient pipelines. However, generating high-quality novel views und…

2024

CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation

CVPR 2024poster

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion. The applications of listener agent generation in virtual interaction have promoted many works achieving diverse and fine-grained…

Cited by 4SourcePDFScholar
2024

Dual-Path Minimum-Phase and All-Pass Decomposition Network for Single Channel Speech Dereverberation

ICASSP 2024accepted

With the development of deep neural networks (DNN), many DNN-based speech dereverberation approaches have been proposed to achieve significant improvement over the traditional methods. However, most deep learning-based dereverberation methods solely focus on suppressing time-frequency domain reverbe…

Cited by 0SourceScholar
2024

Fearless Steps Apollo: Team Communications Based Community Resource Development for Science, Technology, Education, and Historical Preservation

ICASSP 2024accepted

The Fearless Steps Apollo (FS-APOLLO) resource is a collection of 150,000 hours of audio, associated meta-data, and supplemental speech technology infrastructure intended to benefit the (i) speech processing technology, (ii) communication science, team-based psychology, and (iii) education/STEM, his…

Cited by 0SourceScholar
2024

Wukong: Towards a Scaling Law for Large-Scale Recommendation

ICML 2024poster

Scaling laws play an instrumental role in the sustainable improvement in model quality. Unfortunately, recommendation models to date do not exhibit such laws similar to those observed in the domain of large language models, due to the inefficiencies of their upscaling mechanisms. This limitation pos…

Cited by 21SourcePDFScholar
2023

AccFlow: Backward Accumulation for Long-Range Optical Flow

ICCV 2023poster

Recent deep learning-based optical flow estimators have exhibited impressive performance in generating local flows between consecutive frames. However, the estimation of long-range flows between distant frames, particularly under complex object deformation and large motion occlusion, remains a chall…

Cited by 24PDFcodeScholar
2023

Coordinate Ascent for Off-Policy RL with Global Convergence Guarantees

AISTATS 2023poster

We revisit the domain of off-policy policy optimization in RL from the perspective of coordinate ascent. One commonly-used approach is to leverage the off-policy policy gradient to optimize a surrogate objective – the total discounted in expectation return of the target policy with respect to the st…

Cited by 0SourcePDFScholar
2023

Q-Pensieve: Boosting Sample Efficiency of Multi-Objective RL Through Memory Sharing of Q-Snapshots

ICLR 2023poster

Many real-world continuous control problems are in the dilemma of weighing the pros and cons, multi-objective reinforcement learning (MORL) serves as a generic framework of learning control policies for different preferences over objectives. However, the existing MORL methods either rely on multiple…

2022

Learning from Few Samples: Transformation-Invariant SVMs with Composition and Locality at Multiple Scales

NeurIPS 2022accept

Motivated by the problem of learning with small sample sizes, this paper shows how to incorporate into support-vector machines (SVMs) those properties that have made convolutional neural networks (CNNs) successful. Particularly important is the ability to incorporate domain knowledge of invariances,…

2021

Escaping from zero gradient: Revisiting action-constrained reinforcement learning via Frank-Wolfe policy optimization

UAI 2021poster

Action-constrained reinforcement learning (RL) is a widely-used approach in various real-world applications, such as scheduling in networked systems with resource constraints and control of a robot with kinematic constraints. While the existing projection-based approaches ensure zero constraint viol…

2021

Reinforced Few-Shot Acquisition Function Learning for Bayesian Optimization

NeurIPS 2021poster

Bayesian optimization (BO) conventionally relies on handcrafted acquisition functions (AFs) to sequentially determine the sample points. However, it has been widely observed in practice that the best-performing AF in terms of regret can vary significantly under different types of black-box functions…

Cited by 23SourcePDFScholar
2021

Reward-Biased Maximum Likelihood Estimation for Linear Stochastic Bandits

AAAI 2021technical

Modifying the reward-biased maximum likelihood method originally proposed in the adaptive control literature, we propose novel learning algorithms to handle the explore-exploit trade-off in linear bandits problems as well as generalized linear bandits problems. We develop novel index policies that w…

Cited by 15SourcePDFScholar
2020

Exploration Through Reward Biasing: Reward-Biased Maximum Likelihood Estimation for Stochastic Multi-Armed Bandits

ICML 2020poster

Inspired by the Reward-Biased Maximum Likelihood Estimate method of adaptive control, we propose RBMLE – a novel family of learning algorithms for stochastic multi-armed bandits (SMABs). For a broad range of SMABs including both the parametric Exponential Family as well as the non-parametric sub-Gau…

Cited by 16SourcePDFScholar
2019

Stay With Me: Lifetime Maximization Through Heteroscedastic Linear Bandits With Reneging

ICML 2019oral

Sequential decision making for lifetime maximization is a critical problem in many real-world applications, such as medical treatment and portfolio selection. In these applications, a “reneging” phenomenon, where participants may disengage from future interactions after observing an unsatisfiable ou…

Cited by 5SourcePDFScholar