← Search

Di Liu

27 accepted papers

2026

Deep Ensemble Clustering for Visual Representation Learning

ICML 2026poster

Recent advances in visual representation learning have seen the rise of clustering-based vision backbones, which adopt clustering as a core paradigm for feature extraction. However, existing clustering-based backbones typically rely on a single clustering algorithm, whose inherent inductive bias lim…

Cited by 0SourceScholar
2026

Enabling Crab Driving for Rear Steering-Limited Vehicles via Coordinated Direct Yaw Moment Control and Steering Allocation

RA-L 2026

Crab driving has significant application potential in complex driving conditions such as maneuvering in confined spaces, high-speed lane changes, and emergency obstacle avoidance. However, its practical application is impeded by small rear-wheel steering angles in mass-production vehicles. To overco

Cited by 0SourceScholar
2026

WorldGen: From Text to Traversable and Interactive 3D Worlds

CVPR 2026

We introduce WorldGen, a method for generating large, fully formed, navigable 3D worlds from a single text prompt. Existing approaches to 3D scene generation often trade off scene diversity, completeness, and correctness in different ways. We push this envelope by producing large scenes explicitly d

Cited by 0SourceScholar
2025

Enhancing Continual Learning for Medical Imaging: Efficient Knowledge Transfer and Multi-Disease Prediction

ICASSP 2025accepted

Deep learning models for medical disease detection require extensive labeled data, which is often scarce and expensive. Transfer learning can help by leveraging knowledge from large source domains, but directly fine-tuning these models can lead to catastrophic forgetting, making it impossible to reu…

Cited by 0SourceScholar
2025

Improved Training Technique for Latent Consistency Models

ICLR 2025poster

Consistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demonstrated impressive performance, achieving results on par with diffusion models in the pixel space. However, the success…

2025

LUCAS: Layered Universal Codec Avatars

CVPR 2025poster

Photorealistic 3D head avatar reconstruction faces critical challenges in modeling dynamic face-hair interactions and achieving cross-identity generalization, particularly during expressions and head movements. We present LUCAS, a novel Universal Prior Model (UPM) for codec avatar modeling that dise…

Cited by 0SourcePDFScholar
2025

Learning Causally Disentangled Representations for Fair Personality Detection

IJCAI 2025

Personality detection aims to identify the personality traits implied in social posts. Existing methods mainly focus on learning the mapping between user-generated posts and personality trait labels but inevitably suffer from potential harm caused by individual bias, as these posts are written by au

Cited by 0SourcePDFScholar
2025

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

NeurIPS 2025poster

Transformer-based Large Language Models (LLMs) have become increasingly important. However, scaling LLMs to longer contexts incurs slow inference speed and high GPU memory consumption for caching key-value (KV) vectors. This paper presents RetrievalAttention, a training-free approach to both acceler…

Cited by 0SourcecodeScholar
2025

Show and Segment: Universal Medical Image Segmentation via In-Context Learning

CVPR 2025poster

Medical image segmentation remains challenging due to the vast diversity of anatomical structures, imaging modalities, and segmentation tasks. While deep learning has made significant advances, current approaches struggle to generalize as they require task-specific training or fine-tuning on unseen…

Cited by 0SourcePDFScholar
2025

T2Bs: Text-to-Character Blendshapes via Video Generation

ICCV 2025poster

We present T2Bs, a framework for generating high-quality, animatable character head morphable models from text by combining static text-to-3D generation with video diffusion. Text-to-3D models produce detailed static geometry but lack motion synthesis, while video diffusion models generate motion wi…

Cited by 0SourcePDFScholar
2025

Temporal Action Localization with Cross Layer Task Decoupling and Refinement

AAAI 2025technical

Temporal action localization (TAL) involves dual tasks to classify and localize actions within untrimmed videos. However, the two tasks often have conflicting requirements for features. Existing methods typically employ separate heads for classification and localization tasks but share the same inpu…

2025

The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models Via Visual Information Steering

ICML 2025poster

Large Vision-Language Models (LVLMs) can reason effectively over both textual and visual inputs, but they tend to hallucinate syntactically coherent yet visually ungrounded contents. In this paper, we investigate the internal dynamics of hallucination by examining the tokens logits rankings througho…

2025

VISIAR: Empower MLLM for Visual Story Ideation

ACL 2025finding

Ideation, the process of forming ideas from concepts, is a big part of the content creation process. However, the noble goal of helping visual content creators by suggesting meaningful sequences of visual assets from a limited collection is challenging. It requires a nuanced understanding of visual…

2024

AC-EVAL: Evaluating Ancient Chinese Language Understanding in Large Language Models

EMNLP 2024finding

Given the importance of ancient Chinese in capturing the essence of rich historical and cultural heritage, the rapid advancements in Large Language Models (LLMs) necessitate benchmarks that can effectively evaluate their understanding of ancient contexts. To meet this need, we present AC-EVAL, an in…

2024

Instantaneous Perception of Moving Objects in 3D

CVPR 2024poster

The perception of 3D motion of surrounding traffic participants is crucial for driving safety. While existing works primarily focus on general large motions we contend that the instantaneous detection and quantification of subtle motions is equally important as they indicate the nuances in driving b…

Cited by 1SourcePDFScholar
2024

Layout-Agnostic Scene Text Image Synthesis with Diffusion Models

CVPR 2024poster

While diffusion models have significantly advanced the quality of image generation their capability to accurately and coherently render text within these images remains a substantial challenge. Conventional diffusion-based methods for scene text generation are typically limited by their reliance on…

Cited by 5SourcePDFScholar
2023

DeFormer: Integrating Transformers with Deformable Models for 3D Shape Abstraction from a Single Image

ICCV 2023poster

Explicit 3D shape abstraction from a single 2D image is a long-standing problem in computer vision and graphics. By leveraging a set of primitives to represent the target shape, recent methods have achieved promising results. However, these methods either use a relatively larger number of primitives…

Cited by 8PDFScholar
2023

Improving Generalization With Domain Convex Game

CVPR 2023poster

Domain generalization (DG) tends to alleviate the poor generalization capability of deep neural networks by learning model with multiple source domains. A classical solution to DG is domain augmentation, the common belief of which is that diversifying source domains will be conducive to the out-of-d…

2023

LEPARD: Learning Explicit Part Discovery for 3D Articulated Shape Reconstruction

NeurIPS 2023poster

Reconstructing the 3D articulated shape of an animal from a single in-the-wild image is a challenging task. We propose LEPARD, a learning-based framework that discovers semantically meaningful 3D parts and reconstructs 3D shapes in a part-based manner. This is advantageous as 3D parts are robust to…

Cited by 13SourcePDFScholar
2023

Multi-Layer Seasonal Perception Network for Time Series Forecasting

ICASSP 2023accepted

Seasonal time series contain rich long-term dependencies. How to make good use of the seasonal information to predict the future is still a challenging problem. In this paper, we propose a neural network model called Multilayer Seasonal Perception Network (MSPNet) to predict seasonal time series. Fi…

Cited by 0SourceScholar
2023

Pairwise GUI Dataset Construction Between Android Phones and Tablets

NeurIPS 2023poster

In the current landscape of pervasive smartphones and tablets, apps frequently exist across both platforms. Although apps share most graphic user interfaces (GUIs) and functionalities across phones and tablets, developers often rebuild from scratch for tablet versions, escalating costs and squanderi…

2022

A Centimeter-Scale Electrohydrodynamic Multi-Modal Robot Capable of Rolling, Hopping, and Taking Off

RA-L 2022

Insects and animals in nature generally have various modes of locomotion to adapt to complex environments, such as crawling, running, flying, and jumping. Achieving multi-locomotion in a centimeter-scale robot requires complex structures and mechanisms that are normally difficult to design and fabri

Cited by 7SourceScholar
2022

Causality Inspired Representation Learning for Domain Generalization

CVPR 2022oral

Domain generalization (DG) is essentially an out-of-distribution problem, aiming to generalize the knowledge learned from multiple source domains to an unseen target domain. The mainstream is to leverage statistical models to model the dependence between data and labels, intending to learn represent…

Cited by 213PDFcodeScholar
2022

Contrastive Graph Structure Learning via Information Bottleneck for Recommendation

NeurIPS 2022accept

Graph convolution networks (GCNs) for recommendations have emerged as an important research topic due to their ability to exploit higher-order neighbors. Despite their success, most of them suffer from the popularity bias brought by a small number of active users and popular items. Also, a real-worl…

Cited by 72SourcePDFScholar
2021

Low Voltage Control of Micro-Ionic Thrusters Using the Electrostatic Induced Potential of the Collector

RA-L 2021

This work presents a low voltage (5 V) control method for micro-ionic thrusters with high operating voltage (>1 kV). When the collector of a micro-ionic thruster is connected to the negative electrode through a 5V switch circuit, and a high DC voltage (1-10 kV) is applied to the emitter, we can get

Cited by 7SourceScholar