← Search

Dong Liu

79 accepted papers

2026

CADC: Content Adaptive Diffusion-Based Generative Image Compression

CVPR 2026

Diffusion-based generative image compression has demonstrated remarkable potential for achieving realistic reconstruction at ultra-low bitrates. The key to unlocking this potential lies in making the entire compression process content-adaptive, ensuring that the encoder's representation and the deco

Cited by 0SourceScholar
2026

Communication-Efficient Decentralized Optimization via Double-Communication Symmetric ADMM

ICLR 2026poster

This paper focuses on decentralized composite optimization over networks without a central coordinator. We propose a novel decentralized Symmetric ADMM algorithm that incorporates multiple communication rounds within each iteration, derived from a new constraint formulation that enables information…

Cited by 0SourceScholar
2026

Content-Adaptive Hierarchical Hyperprior for Neural Video Coding

CVPR 2026

While neural video codecs (NVCs) have recently demonstrated superior performance over traditional codecs through end-to-end learning, existing approaches primarily focus on architectural enhancements and coding module design, with limited exploration into optimizing hierarchical structures--specific

Cited by 0SourceScholar
2026

EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling

ICLR 2026poster

Instruction-guided image editing has achieved remarkable progress, yet current models still face challenges with complex instructions and often require multiple samples to produce a desired result. Reinforcement Learning (RL) offers a promising solution, but its adoption in image editing has been se…

Cited by 0SourcecodeScholar
2026

MM-ACT: Learn from Multimodal Parallel Generation to Act

CVPR 2026

A generalist robotic policy needs both semantic understanding for task planning and the ability to interact with the environment through predictive capabilities. To tackle this, we present MM-ACT, a unified Vision-Language-Action (VLA) model that integrates text, image, and action in shared token sp

Cited by 0SourcecodeScholar
2026

OmniGen2: Towards Instruction-Aligned Multimodal Generation

CVPR 2026

Multimodal generative models can process instructions in various modalities and demonstrate outstanding performance across a wide range of image generation tasks. However, their robustness in complex real-world scenarios remains limited due to insufficient generalized instruction alignment. We intro

Cited by 0SourcecodeScholar
2026

Perceptual Neural Video Compression with Color Separation and Rank Chain

CVPR 2026

Neural video compression (NVC) has achieved significant progress in recent years. The state-of-the-art (SOTA) NVC schemes, exemplified by the Deep Conditional Video Coding series, have focused on pursuing higher fidelity (e.g., PSNR), but lack sufficient exploitation of deep networks' advantages for

Cited by 0SourcecodeScholar
2026

Real-Time Neural Video Compression with Unified Intra and Inter Coding

CVPR 2026

Neural video compression (NVC) technologies have advanced rapidly in recent years, yielding state-of-the-art schemes such as DCVC-RT that offer superior compression efficiency to H.266/VVC and real-time encoding/decoding capabilities. Nonetheless, existing NVC schemes have several limitations, inclu

Cited by 0SourcecodeScholar
2026

Sample-Efficient Learning with Online Expert Correction for Autonomous Catheter Steering in Endovascular Bifurcation Navigation

ICRA 2026poster

Robot-assisted endovascular intervention offers a safe and effective solution for remote catheter manipulation, reducing radiation exposure while enabling precise navigation. Reinforcement learning (RL) has recently emerged as a promising approach for autonomous catheter steering; however, conventio…

2026

Transform-Free Feature Coding via Entropy-Constrained Vector Quantization

AAAI 2026technical

Feature coding has recently emerged as a key technique for efficient transmission of intermediate representations in distributed AI systems. Existing approaches largely follow a transform-based pipeline inherited from image and video coding, where the transform module is used to remove spatial struc

Cited by 0SourcePDFScholar
2025

A Plug-and-Play Diffusion-Styled Conversion Model for Domain Discrepancies in Medical Image Segmentation

ICASSP 2025accepted

Accurate segmentation is a crucial step in medical image analysis. However, models trained on one dataset often suffer from performance degradation when directly applied to a different domain with a different data distribution due to domain discrepancies. To address this issue, we introduce a novel…

Cited by 0SourceScholar
2025

Conditional Latent Coding with Learnable Synthesized Reference for Deep Image Compression

AAAI 2025technical

In this paper, we study how to synthesize a dynamic reference from an external dictionary to perform conditional coding of the input image in the latent domain and how to learn the conditional latent synthesis and coding modules in an end-to-end manner. Our approach begins by constructing a universa…

2025

Crafting Customisable Characters with LLMs: A Persona-Driven Role-Playing Agent Framework

EMNLP 2025

Large Language Models (LLMs) demonstrate remarkable ability to comprehend instructions and generate human-like text, enabling sophisticated agent simulation beyond basic behavior replication. However, the potential for creating freely customisable characters remains underexplored. We introduce the C

2025

Emphasising Structured Information: Integrating Abstract Meaning Representation into LLMs for Enhanced Open-Domain Dialogue Evaluation

EMNLP 2025

Automatic open-domain dialogue evaluation has attracted increasing attention, yet remains challenging due to the complexity of assessing response appropriateness. Traditional evaluation metrics, typically trained with true positive and randomly selected negative responses, tend to assign higher scor

2025

Exploiting Diffusion Prior for Real-World Image Dehazing with Unpaired Training

AAAI 2025technical

Unpaired training has been verified as one of the most effective paradigms for real scene dehazing by learning from unpaired real-world hazy and clear images. Although numerous studies have been proposed, current methods demonstrate limited generalization for various real scenes due to limited featu…

2025

Feature Coding in the Era of Large Models: Dataset, Test Conditions, and Benchmark

ICCV 2025poster

Large models have achieved remarkable performance across various tasks, yet they incur significant computational costs and privacy concerns during both training and inference. Distributed deployment has emerged as a potential solution, but it necessitates the exchange of intermediate information bet…

2025

Few-Shot Domain Adaptation for Learned Image Compression

AAAI 2025technical

Learned image compression (LIC) has achieved state-of-the-art rate-distortion performance, deemed promising for next-generation image compression techniques. However, pre-trained LIC models usually suffer from significant performance degradation when applied to out-of-training-domain images, implyin…

Cited by 0SourcePDFScholar
2025

Learned Image Compression with Hierarchical Progressive Context Modeling

ICCV 2025poster

Context modeling is essential in learned image compression for accurately estimating the distribution of latents. While recent advanced methods have expanded context modeling capacity, they still struggle to efficiently exploit long-range dependency and diverse context information across different c…

2025

MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation

ACL 2025long

Automatic evaluation of retrieval augmented generation (RAG) systems relies on fine-grained dimensions like faithfulness and relevance, as judged by expert human annotators. Meta-evaluation benchmarks support the development of automatic evaluators that correlate well with human judgement. However,…

2025

MK-Pose: Category-Level Object Pose Estimation via Multimodal-Based Keypoint Learning

IROS 2025

Category-level object pose estimation, which predicts the pose of objects within a known category without prior knowledge of individual instances, is essential in applications like warehouse automation and manufacturing. Existing methods relying on RGB images or point cloud data often struggle with

Cited by 5SourcecodeScholar
2025

MaskTwins: Dual-form Complementary Masking for Domain-Adaptive Image Segmentation

ICML 2025poster

Recent works have correlated Masked Image Modeling (MIM) with consistency regularization in Unsupervised Domain Adaptation (UDA). However, they merely treat masking as a special form of deformation on the input images and neglect the theoretical analysis, which leads to a superficial understanding o…

2025

SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling

CVPR 2025poster

Open-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two c…

2025

When Schrodinger Bridge Meets Real-World Image Dehazing with Unpaired Training

ICCV 2025poster

Recent advancements in unpaired dehazing, particularly those using GANs, show promising performance in processing real-world hazy images. However, these methods tend to face limitations due to the generator's limited transport mapping capability, which hinders the full exploitation of their effectiv…

Cited by 0SourcePDFScholar
2024

Arbitrary-Scale Video Super-resolution Guided by Dynamic Context

AAAI 2024technical

We propose a Dynamic Context-Guided Upsampling (DCGU) module for video super-resolution (VSR) that leverages temporal context guidance to achieve efficient and effective arbitrary-scale VSR. While most VSR research focuses on backbone design, the importance of the upsampling part is often overlooke…

Cited by 2SourcePDFScholar
2024

Deep Fusion of Shifted MLP and CNN for Medical Image Segmentation

ICASSP 2024accepted

Medical image segmentation is an important task in modern analysis of medical images. Current methods tend to extract either local features with convolutions or global features with Transformers. However, few of them are able to effectively fuse global and local features to facilitate segmentation.…

Cited by 0SourceScholar
2024

Generalizable Fourier Augmentation for Unsupervised Video Object Segmentation

AAAI 2024technical

The performance of existing unsupervised video object segmentation methods typically suffers from severe performance degradation on test videos when tested in out-of-distribution scenarios. The primary reason is that the test data in real- world may not follow the independent and identically distrib…

Cited by 6SourcePDFScholar
2024

HiLo: Detailed and Robust 3D Clothed Human Reconstruction with High-and Low-Frequency Information of Parametric Models

CVPR 2024poster

Reconstructing 3D clothed human involves creating a detailed geometry of individuals in clothing with applications ranging from virtual try-on movies to games. To enable practical and widespread applications recent advances propose to generate a clothed human from an RGB image. However they struggle…

2024

Language-Conditioned Robotic Manipulation with Fast and Slow Thinking

ICRA 2024poster

The language-conditioned robotic manipulation aims to transfer natural language instructions into executable actions, from simple "pick-and-place" to tasks requiring intent recognition and visual reasoning. Inspired by the dual-process theory in cognitive science—which suggests two parallel systems…

Cited by 17SourceScholar
2024

Mask-Based Modeling for Neural Radiance Fields

ICLR 2024spotlight

Most Neural Radiance Fields (NeRFs) exhibit limited generalization capabilities,which restrict their applicability in representing multiple scenes using a single model. To address this problem, existing generalizable NeRF methods simply condition the model on image features. These methods still stru…

2024

Object-Centric Instruction Augmentation for Robotic Manipulation

ICRA 2024poster

Humans interpret scenes by recognizing both the identities and positions of objects in their observations. For a robot to perform tasks such as "pick and place", understanding both what the objects are and where they are located is crucial. While the former has been extensively discussed in the lite…

Cited by 14SourceScholar
2024

Offline and Online Optical Flow Enhancement for Deep Video Compression

AAAI 2024technical

Video compression relies heavily on exploiting the temporal redundancy between video frames, which is usually achieved by estimating and using the motion information. The motion information is represented as optical flows in most of the existing deep video compression networks. Indeed, these network…

Cited by 20SourcePDFScholar
2023

Co-Salient Object Detection With Uncertainty-Aware Group Exchange-Masking

CVPR 2023poster

The traditional definition of co-salient object detection (CoSOD) task is to segment the common salient objects in a group of relevant images. Existing CoSOD models by default adopt the group consensus assumption. This brings about model robustness defect under the condition of irrelevant images in…

Cited by 24SourcePDFScholar
2023

On the Effectiveness of Spectral Discriminators for Perceptual Quality Improvement

ICCV 2023poster

Several recent studies advocate the use of spectral discriminators, which evaluate the Fourier spectra of images for generative modeling. However, the effectiveness of the spectral discriminators is not well interpreted yet. We tackle this issue by examining the spectral discriminators in the contex…

Cited by 14PDFcodeScholar
2023

PyramidFlow: High-Resolution Defect Contrastive Localization Using Pyramid Normalizing Flow

CVPR 2023poster

During industrial processing, unforeseen defects may arise in products due to uncontrollable factors. Although unsupervised methods have been successful in defect localization, the usual use of pre-trained models results in low-resolution outputs, which damages visual performance. To address this is…

Cited by 102SourcePDFScholar
2023

Unsupervised Video Object Segmentation with Online Adversarial Self-Tuning

ICCV 2023poster

The existing unsupervised video object segmentation methods depend heavily on the segmentation model trained offline on a labeled training video set, and cannot well generalize to the test videos from a different domain with possible distribution shifts. We propose to perform online fine-tuning on t…

Cited by 11PDFScholar
2022

Extrinsic Calibration of a 2D Laser Rangefinder and a Depth-camera Using an Orthogonal Trihedron

IROS 2022poster

2D laser range-finders and depth-cameras are usually equipped on service robots. But there are rarely calibration methods of them. This paper proposes an extrinsic calibration method of a 2D laser range-finder and a depth-camera using an orthogonal trihedron. The trihedron with orthogonal assumption…

Cited by 3SourceScholar
2022

Learning Pruning-Friendly Networks via Frank-Wolfe: One-Shot, Any-Sparsity, And No Retraining

ICLR 2022spotlight

We present a novel framework to train a large deep neural network (DNN) for only $\textit{once}$, which can then be pruned to $\textit{any sparsity ratio}$ to preserve competitive accuracy $\textit{without any re-training}$. Conventional methods often require (iterative) pruning followed by re-train…

Cited by 41SourcePDFScholar
2022

Recurrent Dynamic Embedding for Video Object Segmentation

CVPR 2022poster

Space-time memory (STM) based video object segmentation (VOS) networks usually keep increasing memory bank every several frames, which shows excellent performance. However, 1) the hardware cannot withstand the ever-increasing memory requirements as the video length increases. 2) Storing lots of info…

Cited by 95PDFcodeScholar
2021

Deep Transport Network for Unsupervised Video Object Segmentation

ICCV 2021poster

The popular unsupervised video object segmentation methods fuse the RGB frame and optical flow via a two-stream network. However, they cannot handle the distracting noises in each input modality, which may vastly deteriorate the model performance. We propose to establish the correspondence between t…

Cited by 67PDFScholar
2021

Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE

CVPR 2021poster

Given an incomplete image without additional constraint, image inpainting natively allows for multiple solutions as long as they appear plausible. Recently, multiple-solution inpainting methods have been proposed and shown the potential of generating diverse results. However, these methods have diff…

Cited by 274PDFcodeScholar
2021

Graph-Based 3D Multi-Person Pose Estimation Using Multi-View Images

ICCV 2021poster

This paper studies the task of estimating the 3D human poses of multiple persons from multiple calibrated camera views. Following the top-down paradigm, we decompose the task into two stages, i.e. person localization and pose estimation. Both stages are processed in coarse-to-fine manners. And we pr…

Cited by 66PDFcodeScholar
2021

Modeling of Planar Hydraulically Amplified Self-Healing Electrostatic Actuators

RA-L 2021

With the advantages of high actuation strain and specific power and ability of self-healing after dielectric breakdown, the planar hydraulically amplified self-healing electrostatic (pHASEL) actuators are promising for extensive emerging applications of soft robots. However, the relationship between

Cited by 0SourceScholar
2021

Motion-Focused Contrastive Learning of Video Representations

ICCV 2021poster

Motion, as the most distinct phenomenon in a video to involve the changes over time, has been unique and critical to the development of video representation learning. In this paper, we ask the question: how important is the motion particularly for self-supervised video representation learning. To th…

Cited by 47PDFcodeScholar
2021

PSD: Principled Synthetic-to-Real Dehazing Guided by Physical Priors

CVPR 2021poster

Deep learning-based methods have achieved remarkable performance for image dehazing. However, previous studies are mostly focused on training models with synthetic hazy images, which incurs performance drop when the models are used for real-world hazy images. We propose a Principled Synthetic-to-rea…

Cited by 334PDFcodeScholar
2021

Structured Multi-Level Interaction Network for Video Moment Localization via Language Query

CVPR 2021poster

We address the problem of localizing a specific moment described by a natural language query. Existing works interact the query with either video frame or moment proposal, and neglect the inherent structure of moment construction for both cross-modal understanding and video content comprehension, wh…

Cited by 102PDFScholar
2020

Hidden Markov Models for Sepsis Detection in Preterm Infants

ICASSP 2020accepted

We explore the use of traditional and contemporary hidden Markov models (HMMs) for sequential physiological data analysis and sepsis prediction in preterm infants. We investigate the use of classical Gaussian mixture model based HMM, and a recently proposed neural network based HMM. To improve the n…

Cited by 0SourceScholar
2020

Learning Trailer Moments in Full-Length Movies with Co-Contrastive Attention

ECCV 2020poster

A movie's key moments stand out of the screenplay to grab an audience's attention and make movie browsing efficient. But a lack of annotations makes the existing approaches not applicable to movie key moment detection. To get rid of human annotations, we leverage the officially-released trailers as…

Cited by 65SourcePDFScholar
2020

Photon-Efficient 3D Imaging with A Non-Local Neural Network

ECCV 2020poster

Photon-efficient imaging has enabled a number of applications relying on single-photon sensors that can capture a 3D image with as few as one photon per pixel. In practice, however, measurements of low photon counts are often mixed with heavy background noise, which poses a great challenge for exist…

2020

Transferring and Regularizing Prediction for Semantic Segmentation

CVPR 2020poster

Semantic segmentation often requires a large set of images with pixel-level annotations. In the view of extremely expensive expert labeling, recent research has shown that the models trained on photo-realistic synthetic data (e.g., computer games) with computer-generated annotations can be adapted t…

Cited by 47PDFScholar
2019

Customizable Architecture Search for Semantic Segmentation

CVPR 2019poster

In this paper, we propose a Customizable Architecture Search (CAS) approach to automatically generate a network architecture for semantic image segmentation. The generated network consists of a sequence of stacked computation cells. A computation cell is represented as a directed acyclic graph, in w…

Cited by 179PDFScholar
2019

DADA: Deep Adversarial Data Augmentation for Extremely Low Data Regime Classification

ICASSP 2019accepted

Deep learning has revolutionized the performance of classification, but meanwhile demands sufficient labeled data for training. Given insufficient data, while many techniques have been developed to help combat overfitting, the challenge remains if one tries to train deep networks, especially in the…

Cited by 0SourceScholar
2019

Deep High-Resolution Representation Learning for Human Pose Estimation

CVPR 2019poster

In this paper, we are interested in the human pose estimation problem with a focus on learning reliable high-resolution representations. Most existing methods recover high-resolution representations from low-resolution representations produced by a high-to-low resolution network. Instead, our propos…

Cited by 6183PDFcodeScholar
2019

Entropy-regularized Optimal Transport Generative Models

ICASSP 2019accepted

We investigate the use of entropy-regularized optimal transport (EOT) cost in developing generative models to learn implicit distributions. Two generative models are proposed. One uses EOT cost directly in an one-shot optimization problem and the other uses EOT cost iteratively in an adversarial gam…

Cited by 0SourceScholar
2018

Fully Convolutional Adaptation Networks for Semantic Segmentation

CVPR 2018poster

The recent advances in deep neural networks have convincingly demonstrated high capability in learning vision models on large datasets. Nevertheless, collecting expert labeled datasets especially with pixel-level annotations is an extremely expensive process. An appealing alternative is to render sy…

Cited by 429SourcePDFScholar
2017

Human Pose Estimation Using Global and Local Normalization

ICCV 2017poster

In this paper, we address the problem of estimating the positions of human joints, i.e., articulated pose estimation. Recent state-of-the-art solutions model two key issues, joint detection and spatial configuration refinement, together using convolutional neural networks. Our work mainly focuses on…

Cited by 82PDFScholar
2016

Comparative Deep Learning of Hybrid Representations for Image Recommendations

CVPR 2016poster

In many image-related tasks, learning expressive and discriminative representations of images is essential, and deep learning has been studied for automating the learning of such representations. Some user-centric tasks, such as image recommendations, call for effective representations of not only i…

Cited by 162PDFScholar