← Search

shan liu

44 accepted papers

2026

Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning

ICML 2026poster

Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drift (thus slow recovery), and leaving unanswered the principled question of how …

Cited by 0SourceScholar
2025

Point Cloud Semantic Segmentation with Sparse and Inhomogeneous Annotations

AAAI 2025technical

Utilizing uniformly distributed sparse annotations, weakly supervised learning alleviates the heavy reliance on fine-grained annotations in point cloud semantic segmentation tasks. However, few works discuss the inhomogeneity of sparse annotations, albeit it is common in real-world scenarios. Theref…

2024

Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment

CVPR 2024poster

No-reference point cloud quality assessment (NR-PCQA) aims to automatically evaluate the perceptual quality of distorted point clouds without available reference which have achieved tremendous improvements due to the utilization of deep neural networks. However learning-based NR-PCQA methods suffer…

Cited by 17SourcePDFScholar
2024

Distribution Guidance Network for Weakly Supervised Point Cloud Semantic Segmentation

NeurIPS 2024poster

Despite alleviating the dependence on dense annotations inherent to fully supervised methods, weakly supervised point cloud semantic segmentation suffers from inadequate supervision signals. In response to this challenge, we introduce a novel perspective that imparts auxiliary constraints by regulat…

Cited by 2SourcePDFScholar
2024

Layer-Wise Representation Fusion for Compositional Generalization

AAAI 2024technical

Existing neural models are demonstrated to struggle with compositional generalization (CG), i.e., the ability to systematically generalize to unseen compositions of seen components. A key reason for failure on CG is that the syntactic and semantic representations of sequences in both the uppermost l…

2024

Learned Lossless Image Compression based on Bit Plane Slicing

CVPR 2024poster

Autoregressive Initial Bits (ArIB) a framework that combines subimage autoregression and latent variable models has shown its advantages in lossless image compression. However in current methods the image splitting makes the information of latent variables being uniformly distributed in each subimag…

2024

Less Is More: Label Recommendation for Weakly Supervised Point Cloud Semantic Segmentation

AAAI 2024technical

Weak supervision has proven to be an effective strategy for reducing the burden of annotating semantic segmentation tasks in 3D space. However, unconstrained or heuristic weakly supervised annotation forms may lead to suboptimal label efficiency. To address this issue, we propose a novel label recom…

Cited by 20SourcePDFScholar
2024

SJTU-TMQA: A Quality Assessment Database for Static Mesh with Texture Map

ICASSP 2024accepted

In recent years, static meshes with texture maps have become one of the most prevalent digital representations of 3D shapes in various applications, such as animation, gaming, medical imaging, and cultural heritage applications. However, little research has been done on the quality assessment of tex…

Cited by 0SourceScholar
2024

ScanPCGC: Learning-Based Lossless Point Cloud Geometry Compression using Sequential Slice Representation

ICASSP 2024accepted

The efficient storage and transportation requirements of point clouds promote the development of point cloud compression algorithms. In this paper, we develop a novel point cloud geometry compression using sequential slice representation. Unlike the limited contexts in conventional codecs and other…

Cited by 0SourceScholar
2023

Learning to Compose Representations of Different Encoder Layers towards Improving Compositional Generalization

EMNLP 2023long findings

Recent studies have shown that sequence-to-sequence (seq2seq) models struggle with compositional generalization (CG), i.e., the ability to systematically generalize to unseen compositions of seen components. There is mounting evidence that one of the reasons hindering CG is the representation of the…

Cited by 0SourcecodeScholar
2023

PanelNet: Understanding 360 Indoor Environment via Panel Representation

CVPR 2023poster

Indoor 360 panoramas have two essential properties. (1) The panoramas are continuous and seamless in the horizontal direction. (2) Gravity plays an important role in indoor environment design. By leveraging these properties, we present PanelNet, a framework that understands indoor environments using…

Cited by 18SourcePDFScholar
2023

S3I-PointHop: SO(3)-Invariant PointHop for 3D Point Cloud Classification

ICASSP 2023accepted

Many point cloud classification methods are developed under the assumption that all point clouds in the dataset are well aligned with the canonical axes so that the 3D Cartesian point coordinates can be employed to learn features. When input point clouds are not aligned, the classification performan…

Cited by 0SourceScholar
2023

Surface-Sampling Based Objective Quality Assessment Metrics for Meshes

ICASSP 2023accepted

In this paper, we prove that it is feasible to perform mesh quality assessment by sampling it into point cloud. We propose a general and efficient surface-sampling based framework that can deal with various types and levels of distortions with less complexity. In this method, the original and distor…

Cited by 0SourceScholar
2022

Bi-Directional Normalization and Color Attention-Guided Generative Adversarial Network for Image Enhancement

ICASSP 2022accepted

Most existing image enhancement methods require paired images, and rarely consider the aesthetic quality. This paper proposes a bi-directional normalization and color attention-guided generative adversarial network (BNCAGAN) for unsupervised image enhancement. An auxiliary attention classifier (AAC)…

Cited by 0SourceScholar
2022

Coarse-To-Fine Deep Video Coding With Hyperprior-Guided Mode Prediction

CVPR 2022poster

The previous deep video compression approaches only use the single scale motion compensation strategy and rarely adopt the mode prediction technique from the traditional standards like H.264/H.265 for both motion and residual compression. In this work, we first propose a coarse-to-fine (C2F) deep vi…

Cited by 107PDFScholar
2022

ERDN: Equivalent Receptive Field Deformable Network for Video Deblurring

ECCV 2022poster

"Video deblurring aims to restore sharp frames from blurry video sequences. Existing methods usually adopt optical flow to compensate misalignment between reference frame and each neighboring frame. However, inaccurate flow estimation caused by large displacements will lead to artifacts in the warpe…

2022

Neural Texture Extraction and Distribution for Controllable Person Image Synthesis

CVPR 2022oral

We deal with the controllable person image synthesis task which aims to re-render a human from a reference image with explicit control over body pose and appearance. Observing that person images are highly structured, we propose to generate desired images by extracting and distributing semantic enti…

Cited by 94PDFcodeScholar
2022

OctAttention: Octree-Based Large-Scale Contexts Model for Point Cloud Compression

AAAI 2022technical

In point cloud compression, sufficient contexts are significant for modeling the point cloud distribution. However, the contexts gathered by the previous voxel-based methods decrease when handling sparse point clouds. To address this problem, we propose a multiple-contexts deep learning framework ca…

2022

Rate Control for Learned Video Compression

ICASSP 2022accepted

Rate control is a critical part for video compression, especially in bandwidth-limited tasks such as live and broadcast. The newly-rising learned video compression has shown advantageous rate-distortion (RD) performance in previous research, but lack of rate control heavily limits its usage in real…

Cited by 0SourceScholar
2022

Towards Joint Frame-Level and MOS Quality Predictions with Low-Complexity Objective Models

ICASSP 2022accepted

The evaluation of the quality of gaming content, with low-complexity and low-delay approaches is a major challenge raised by the emerging gaming video streaming and cloud-gaming services. Considering two existing and a newly created gaming databases this paper confirms that some low-complexity metri…

Cited by 0SourceScholar
2021

High Quality Disparity Remapping With Two-Stage Warping

ICCV 2021poster

A high quality disparity remapping method that preserves 2D shapes and 3D structures, and adjusts disparities of important objects in stereo image pairs is proposed. It is formulated as a constrained optimization problem, whose solution is challenging, since we need to meet multiple requirements of…

Cited by 2PDFScholar
2021

Learning Model-Blind Temporal Denoisers without Ground Truths

ICASSP 2021accepted

Denoisers trained with synthetic noises often fail to cope with the diversity of real noises, giving way to methods that can adapt to unknown noise without noise modeling or ground truth. Previous image-based method leads to noise overfitting if directly applied to temporal denoising, and has inadeq…

Cited by 0SourceScholar
2021

Nested Error Map Generation Network for No-Reference Image Quality Assessment

ICASSP 2021accepted

We propose a multi-task learning neural network for No-Reference image quality assessment (NR-IQA). The pro-posed architecture consists of a backbone feature extractor, a nested multi-task generative module and a quality regression module. We adopt a coarse-to-fine strategy to predict objective erro…

Cited by 0SourceScholar
2021

PIRenderer: Controllable Portrait Image Generation via Semantic Neural Rendering

ICCV 2021poster

Generating portrait images by controlling the motions of existing faces is an important task of great consequence to social media industries. For easy use and intuitive control, semantically meaningful and fully disentangled parameters should be used as modifications. However, many existing techniqu…

Cited by 252PDFcodeScholar
2021

SSD-GAN: Measuring the Realness in the Spatial and Spectral Domains

AAAI 2021technical

This paper observes that there is an issue of high frequencies missing in the discriminator of standard GAN, and we reveal it stems from downsampling layers employed in the network architecture. This issue makes the generator lack the incentive from the discriminator to learn high-frequency content…

2021

Structure-Transformed Texture-Enhanced Network for Person Image Synthesis

ICCV 2021poster

Pose-guided virtual try-on task aims to modify the fashion item based on pose transfer task. These two tasks that belong to person image synthesis have strong correlations and similarities. However, existing methods treat them as two individual tasks and do not explore correlations between them. Mor…

Cited by 8PDFScholar
2020

C3DVQA: Full-Reference Video Quality Assessment with 3D Convolutional Neural Network

ICASSP 2020accepted

Traditional video quality assessment (VQA) methods evaluate localized picture quality and video score is predicted by temporally aggregating frame scores. However, video quality exhibits different characteristics from static image quality due to the existence of temporal masking effects. In this pap…

Cited by 0SourceScholar
2020

ROIMIX: Proposal-Fusion Among Multiple Images for Underwater Object Detection

ICASSP 2020accepted

Generic object detection algorithms have proven their excellent performance in recent years. However, object detection on underwater datasets is still less explored. In contrast to generic datasets, underwater images usually have color shift and low contrast; sediment would cause blurring in underwa…

Cited by 0SourceScholar
2019

AttPool: Towards Hierarchical Feature Representation in Graph Convolutional Networks via Attention Mechanism

ICCV 2019poster

Graph convolutional networks (GCNs) are potentially short of the ability to learn hierarchical representation for graph embedding, which holds them back in the graph classification task. Here, we propose AttPool, which is a novel graph pooling module based on attention mechanism, to remedy the probl…

Cited by 87PDFcodeScholar
2019

BLP - Boundary Likelihood Pinpointing Networks for Accurate Temporal Action Localization

ICASSP 2019accepted

Despite tremendous progress achieved in temporal action detection, state-of-the-art methods still suffer from the sharp performance deterioration when localizing the starting and ending temporal action boundaries. Although most methods apply boundary regression paradigm to tackle this problem, we ar…

Cited by 0SourceScholar
2019

Boundary Information Matters More: Accurate Temporal Action Detection with Temporal Boundary Network

ICASSP 2019accepted

Temporal action detection in untrimmed videos is an important yet challenging task. How to locate complex actions accurately is still an open question due to the ambiguous boundaries between action instances and the background. Recently a newly proposed work exploits Structured Segment Networks (SSN…

Cited by 0SourceScholar
2019

Graph Convolutional Label Noise Cleaner: Train a Plug-And-Play Action Classifier for Anomaly Detection

CVPR 2019poster

Video anomaly detection under weak labels is formulated as a typical multiple-instance learning problem in previous works. In this paper, we provide a new perspective, i.e., a supervised learning task under noisy labels. In such a viewpoint, as long as cleaning away label noise, we can directly appl…

Cited by 590PDFcodeScholar
2019

Multi-mapping Image-to-Image Translation via Learning Disentanglement

NeurIPS 2019poster

Recent advances of image-to-image translation focus on learning the one-to-many mapping from two aspects: multi-modal translation and multi-domain translation. However, the existing methods only consider one of the two perspectives, which makes them unable to solve each other's problem. To address t…

2019

StructureFlow: Image Inpainting via Structure-Aware Appearance Flow

ICCV 2019poster

Image inpainting techniques have shown significant improvements by using deep neural networks recently. However, most of them may either fail to reconstruct reasonable structures or restore fine-grained textures. In order to solve this problem, in this paper, we propose a two-stage model which split…

Cited by 455PDFcodeScholar