← Search

Li Song

36 accepted papers

2026

Content-Aware Mamba for Learned Image Compression

ICLR 2026poster

Recent Learned image compression (LIC) leverages Mamba-style state-space models (SSMs) for global receptive fields with linear complexity. However, the standard Mamba adopts content-agnostic, predefined raster (or multi-directional) scans under strict causality. This rigidity hinders its ability to…

Cited by 0SourcecodeScholar
2026

D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint Videos

AAAI 2026technical

Free-Viewpoint Video (FVV) enables immersive 3D experiences, but efficient compression of dynamic 3D representation remains a major challenge. Existing dynamic 3D Gaussian Splatting methods couple reconstruction with optimization-dependent compression and customized motion formats, limiting generali

Cited by 0SourcePDFScholar
2026

OmniZip: Learning a Unified and Lightweight Lossless Compressor for Multi-Modal Data

CVPR 2026

Lossless compression is essential for efficient data storage and transmission. Although learning-based lossless compressors achieve strong results, most of them are designed for a single modality, leading to redundant compressor deployments in multi-modal settings. Designing a unified multi-modal co

Cited by 0SourcecodeScholar
2026

SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors

ICLR 2026poster

Reconstructing 3D scenes from sparse images remains a challenging task due to the difficulty of recovering accurate geometry and texture without optimization. Recent approaches leverage generalizable models to generate 3D scenes using 3D Gaussian Splatting (3DGS) primitive. However, they often fail…

Cited by 0SourcecodeScholar
2026

TaCo: A Benchmark for Lossless and Lossy Codecs of Heterogeneous Tactile Data

ICLR 2026poster

Tactile sensing is crucial for embodied intelligence, providing fine-grained perception and control in complex environments. However, efficient tactile data compression, which is essential for real-time robotic applications under strict bandwidth constraints, remains underexplored. The inherent hete…

Cited by 0SourceScholar
2026

WEBEXPERT: DOMAIN-AWARE WEB AGENTS WITH CRITIC-GUIDED EXPERT EXPERIENCE FOR HIGH-PRECISION SEARCH

ICASSP 2026poster

Specialized web tasks in finance, biomedicine, and pharmaceuticals remain challenging due to missing domain priors: queries drift, evidence is noisy, and reasoning is brittle. We present WebExpert, a domain-aware web agent that we implement end-to-end, featuring : (i) sentence-level experience retri…

Cited by 0SourcePDFScholar
2025

4DGC: Rate-Aware 4D Gaussian Compression for Efficient Streamable Free-Viewpoint Video

CVPR 2025poster

3D Gaussian Splatting (3DGS) has substantial potential for enabling photorealistic Free-Viewpoint Video (FVV) experiences. However, the vast number of Gaussians and their associated attributes poses significant challenges for storage and transmission. Existing methods typically handle dynamic 3DGS r…

Cited by 0SourcePDFScholar
2025

Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image Compression

AAAI 2025technical

Neural image compression often faces a challenging trade-off among rate, distortion and perception. While most existing methods typically focus on either achieving high pixel-level fidelity or optimizing for perceptual metrics, we propose a novel approach that simultaneously addresses both aspects f…

Cited by 1SourcePDFScholar
2025

H3D-DGS: Exploring Heterogeneous 3D Motion Representation for Deformable 3D Gaussian Splatting

NeurIPS 2025poster

Dynamic scene reconstruction poses a persistent challenge in 3D vision. Deformable 3D Gaussian Splatting has emerged as an effective method for this task, offering real-time rendering and high visual fidelity. This approach decomposes a dynamic scene into a static representation in a canonical space…

Cited by 0SourceScholar
2025

L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text Compression

AAAI 2025technical

Learning-based probabilistic models can be combined with an entropy coder for data compression. However, due to the high complexity of learning-based models, their practical application as text compressors has been largely overlooked. To address this issue, our work focuses on a low-complexity desig…

2025

Linear Attention Modeling for Learned Image Compression

CVPR 2025poster

Recent years, learned image compression has made tremendous progress to achieve impressive coding efficiency. Its coding gain mainly comes from non-linear neural network-based transform and learnable entropy modeling. However, most studies focus on a strong backbone, and few studies consider a low c…

2025

OmniRe: Omni Urban Scene Reconstruction

ICLR 2025spotlight

We introduce OmniRe, a comprehensive system for efficiently creating high-fidelity digital twins of dynamic real-world scenes from on-device logs. Recent methods using neural fields or Gaussian Splatting primarily focus on vehicles, hindering a holistic framework for all dynamic foregrounds demanded…

2025

VRVVC: Variable-Rate NeRF-Based Volumetric Video Compression

AAAI 2025technical

Neural Radiance Field (NeRF)-based volumetric video has revolutionized visual media by delivering photorealistic Free-Viewpoint Video (FVV) experiences that provide audiences with unprecedented immersion and interactivity. However, the substantial data volumes pose significant challenges for storage…

Cited by 0SourcePDFScholar
2024

Approaches and Challenges for Resolving Different Representations of Fictional Characters for Chinese Novels

COLING 2024main

Due to the huge scale of literary works, automatic text analysis technologies are urgently needed for literary studies such as Digital Humanities. However, the domain-generality of existing NLP technologies limits their effectiveness on in-depth literary studies. It is valuable to explore how to ada…

2024

Depth-Guided Robust and Fast Point Cloud Fusion NeRF for Sparse Input Views

AAAI 2024technical

Novel-view synthesis with sparse input views is important for real-world applications like AR/VR and autonomous driving. Recent methods have integrated depth information into NeRFs for sparse input synthesis, leveraging depth prior for geometric and spatial understanding. However, most existing work…

Cited by 6SourcePDFScholar
2024

Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition

ECCV 2024poster

"Parameter-efficient transfer learning (PETL) is a promising task, aiming to adapt the large-scale pre-trained model to downstream tasks with a relatively modest cost. However, current PETL methods struggle in compressing computational complexity and bear a heavy inference burden due to the complete…

Cited by 1SourcePDFScholar
2024

Hdrtvformer: Efficient Sdrtv-to-Hdrtv via Affine Transformation and Spatial-Aware Transformer

ICASSP 2024accepted

Recent works on reconstructing HDR videos in display format (HDRTV) suffer from high computational and memory requirements because they learn the SDRTV-to-HDRTV mapping directly in 4K resolution. This paper proposes an efficient SDRTV-to-HDRTV model (HDRTVFormer) that decomposes the HDRTV restoratio…

Cited by 0SourceScholar
2024

Neural Rate Control for Learned Video Compression

ICLR 2024poster

The learning-based video compression method has made significant progress in recent years, exhibiting promising compression performance compared with traditional video codecs. However, prior works have primarily focused on advanced compression architectures while neglecting the rate control techniqu…

Cited by 6SourcePDFScholar
2023

Boosting Video Object Segmentation via Space-Time Correspondence Learning

CVPR 2023poster

Current top-leading solutions for video object segmentation (VOS) typically follow a matching-based regime: for each query frame, the segmentation mask is inferred according to its correspondence to previously processed and the first annotated frames. They simply exploit the supervisory signals from…

2023

Divide and Conquer: a Two-Step Method for High Quality Face De-identification with Model Explainability

ICCV 2023poster

Face de-identification involves concealing the true identity of a face while retaining other facial characteristics. Current target-generic methods typically disentangle identity features in the latent space, using adversarial training to balance privacy and utility. However, this pattern often lead…

Cited by 22PDFcodeScholar
2022

A Codec Information Assisted Framework for Efficient Compressed Video Super-Resolution

ECCV 2022poster

"Online processing of compressed videos to increase their resolutions attracts increasing and broad attention. Video Super-Resolution (VSR) using recurrent neural network architecture is a promising solution due to its efficient modeling of long-range temporal dependencies. However, state-of-the-art…

Cited by 9SourcePDFScholar
2022

L-Tracing: Fast Light Visibility Estimation on Neural Surfaces by Sphere Tracing

ECCV 2022poster

"We introduce a highly efficient light visibility estimation method, called L-Tracing, for reflectance factorization on neural implicit surfaces. Light visibility is indispensable for modeling shadows and specular of high quality on object’s surface. For neural implicit representations, former metho…

Cited by 11SourcePDFScholar
2022

Multi-Mode Motion Control of Reconfigurable Vortex-Shaped Microrobot Swarms for Targeted Tumor Therapy

RA-L 2022

Micro-nano robots with low invasiveness and high drug utilization are considered a promising approach for tumor therapy. However, individual micro-nano robots are limited in terms of motility, drug-carrying capacity, and environmental adaptability. This study proposes a clustering control strategy f

Cited by 13SourceScholar
2022

PTSEFormer: Progressive Temporal-Spatial Enhanced TransFormer towards Video Object Detection

ECCV 2022poster

"Recent years have witnessed a trend of applying context frames to boost the performance of object detection as video object detection. Existing methods usually aggregate features at one stroke to enhance the feature. These methods, however, usually lack spatial information from neighboring frames a…

2021

A Portable Remote Optoelectronic Tweezer System for Microobjects Manipulation

IROS 2021poster

Non-contact manipulation technology has extensive application in the manipulation and fabrication of micro/nanomaterials. However, the manipulation devices are often precise and complex, operated only by professionals and subject to site constraints. We propose a simple optoelectronic tweezer platfo…

Cited by 1SourceScholar
2021

Dual Attention Guided Gaze Target Detection in the Wild

CVPR 2021poster

Gaze target detection aims to infer where each person in a scene is looking. Existing works focus on 2D gaze and 2D saliency, but fail to exploit 3D contexts. In this work, we propose a three-stage method to simulate the human gaze inference behavior in 3D space. In the first stage, we introduce a c…

Cited by 91PDFcodeScholar
2021

Personalized and Invertible Face De-Identification by Disentangled Identity Information Manipulation

ICCV 2021poster

The popularization of intelligent devices including smartphones and surveillance cameras results in more serious privacy issues. De-identification is regarded as an effective tool for visual privacy protection with the process of concealing or replacing identity information. Most of the existing de-…

Cited by 80PDFScholar
2021

Precise Control of Magnetized Macrophage Cell Robot for Targeted Drug Delivery

IROS 2021poster

Micro-nano-robots are considered to be a promising platform for drug delivery in biological organisms, but there are still urgent technical problems in biocompatibility and degradability of 3D-printed-based micro-robots that need to be solved. Therefore, in this paper, we design a magnetized bio-hyb…

Cited by 2SourceScholar
2021

Region-Aware Adaptive Instance Normalization for Image Harmonization

CVPR 2021poster

Image composition plays a common but important role in photo editing. To acquire photo-realistic composite images, one must adjust the appearance and visual style of the foreground to be compatible with the background. Existing deep learning methods for harmonizing composite images directly learn an…

Cited by 144PDFcodeScholar
2020

Magnetized Cell-robot Propelled by Magnetic Field for Cancer Killing

IROS 2020poster

In this paper, we present a magnetized cell-robot using macrophages as templates, which can be controlled under a strong gradient magnetic field, to approach and kill cancer cells in both vitro and vivo environment. Firstly, we establish a magnetic control system using only four coils which can gene…

Cited by 3SourceScholar
2020

Toward Fine-grained Facial Expression Manipulation

ECCV 2020poster

Facial expression manipulation aims at editing facial expression with a given condition. Previous methods edit an input image under the guidance of a discrete emotion label or absolute condition (e.g., facial action units) to possess the desired expression. However, these methods either suffer from…

2018

Learning an Inverse Tone Mapping Network with a Generative Adversarial Regularizer

ICASSP 2018accepted

Transferring a low-dynamic-range (LDR) image to a high-dynamic-range (HDR) image, which is the so-called inverse tone mapping (iTM), is an important imaging technique to improve visual effects of imaging devices. In this paper, we propose a novel deep learning-based iTM method, which learns an inver…

Cited by 0SourceScholar