← Search

Miaohui Wang

21 accepted papers

2026

Firing Bits Where It Matters: Spiking-Guided Just Recognizable Distortion Modeling for Machine-Centric Video Coding

AAAI 2026technical

Just recognizable distortion (JRD) has emerged as a promising paradigm for machine-centric video coding. However, existing JRD-guided coding methods are limited by coarse annotation granularity and high computational cost, which hinder their deployment. In this paper, we first investigate the impact

Cited by 0SourcePDFScholar
2026

FreeMem: Enhancing Consistency in Long Video Generation via Tuning-Free Memory

AAAI 2026technical

Text-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation resu

Cited by 0SourcePDFScholar
2026

MCI-Net: A Robust Multi-Domain Context Integration Network for Point Cloud Registration

AAAI 2026technical

Robust and discriminative feature learning is critical for high-quality point cloud registration. However, existing deep learning–based methods typically rely on Euclidean neighborhood-based strategies for feature extraction, which struggle to effectively capture the implicit semantics and structura

Cited by 0SourcePDFScholar
2026

Perceive More with Less: LiDAR Point Cloud Compression at Just Recognizable Distortion for 3D Scene Understanding

AAAI 2026technical

Existing LiDAR point cloud (LPC) data coding methods primarily focus on balancing compression efficiency and reconstruction quality according to the human vision system (HVS). However, these methods rarely consider the requirements of downstream scene understanding tasks from the perspective of the

Cited by 0SourcePDFScholar
2026

Point Cloud Quality Assessment via Multi-View Structure-Aware Feature Fusion

AAAI 2026technical

Point cloud quality assessment (PCQA) is essential for reliable 3D visual applications. While point-based methods face challenges in characterizing distortions due to point cloud disorder, projection-based approaches offer better efficiency but suffer from geometric distortion insensitivity and text

Cited by 0SourcePDFScholar
2026

The Last Byte: Learning Just Enough for Machine-Oriented Image Compression

AAAI 2026technical

Just recognizable distortion (JRD) has been introduced for image compression for machines, aiming to quantify the maximum coding distortion that can be tolerated by a specific perception model, thereby defining the upper bound of machine vision redundancy (MVR). However, existing JRD-based redundanc

Cited by 0SourcePDFScholar
2025

DDJND: Dual Domain Just Noticeable Difference in Multi-Source Content Images with Structural Discrepancy

AAAI 2025technical

Most existing just noticeable difference (JND) methods primarily integrate specific masking effects in a single domain. However, these single-domain JND methods struggle with the structural discrepancies in multi-source content images, limiting their effectiveness in visual redundancy estimation. To…

Cited by 0SourcePDFScholar
2025

mmFAS: Multimodal Face Anti-Spoofing Using Multi-Level Alignment and Switch-Attention Fusion

AAAI 2025technical

The increasing number of presentation attacks on reliable face matching has raised concerns and garnered attention towards face anti-spoofing (FAS). However, existing methods for FAS modeling commonly fuse multiple visual modalities (e.g., RGB, Depth, and Infrared) in a straightforward manner, disre…

Cited by 0SourcePDFScholar
2024

MetaJND: A Meta-Learning Approach for Just Noticeable Difference Estimation

IJCAI 2024poster

The modeling of just noticeable difference (JND) in supervised learning for visual signals has made significant progress. However, existing JND models often suffer from limited generalization due to the need for large-scale training data and their constraints to certain image types. Moreover, these…

Cited by 1SourcePDFScholar
2024

Visual Redundancy Removal for Composite Images: A Benchmark Dataset and a Multi-Visual-Effects Driven Incremental Method

AAAI 2024technical

Composite images (CIs) typically combine various elements from different scenes, views, and styles, which are a very important information carrier in the era of mixed media such as virtual reality, mixed reality, metaverse, etc. However, the complexity of CI content presents a significant challenge…

Cited by 1SourcePDFScholar
2024

Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene Completion

NeurIPS 2024poster

Semantic scene completion is a difficult task that involves completing the geometry and semantics of a scene from point clouds in a large-scale environment. Many current methods use 3D/2D convolutions or attention mechanisms, but these have limitations in directly constructing geometry and accuratel…

Cited by 1SourcePDFScholar
2024

msLPCC: A Multimodal-Driven Scalable Framework for Deep LiDAR Point Cloud Compression

AAAI 2024technical

LiDAR sensors are widely used in autonomous driving, and the growing storage and transmission demands have made LiDAR point cloud compression (LPCC) a hot research topic. To address the challenges posed by the large-scale and uneven-distribution (spatial and categorical) of LiDAR point data, this pa…

Cited by 4SourcePDFScholar
2023

3D Surface Super-resolution from Enhanced 2D Normal Images: A Multimodal-driven Variational AutoEncoder Approach

IJCAI 2023poster

3D surface super-resolution is an important technical tool in virtual reality, and it is also a research hotspot in computer vision. Due to the unstructured and irregular nature of 3D object data, it is usually difficult to obtain high-quality surface details and geometry textures via a low-cost har…

Cited by 1SourcePDFScholar
2023

CVSformer: Cross-View Synthesis Transformer for Semantic Scene Completion

ICCV 2023poster

Semantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the obj…

Cited by 9PDFcodeScholar
2023

Just Noticeable Visual Redundancy Forecasting: A Deep Multimodal-Driven Approach

AAAI 2023technical

Just noticeable difference (JND) refers to the maximum visual change that human eyes cannot perceive, and it has a wide range of applications in multimedia systems. However, most existing JND approaches only focus on a single modality, and rarely consider the complementary effects of multimodal info…

Cited by 5SourcePDFScholar
2022

Generative Status Estimation and Information Decoupling for Image Rain Removal

NeurIPS 2022accept

Image rain removal requires the accurate separation between the pixels of the rain streaks and object textures. But the confusing appearances of rains and objects lead to the misunderstanding of pixels, thus remaining the rain streaks or missing the object details in the result. In this paper, we pr…

Cited by 9SourcePDFScholar
2019

Surface Reconstruction From Normals: A Robust DGP-Based Discontinuity Preservation Approach

CVPR 2019poster

In 3D surface reconstruction from normals, discontinuity preservation is an important but challenging task. However, existing studies fail to address the discontinuous normal maps by enforcing the surface integrability in the continuous domain. This paper introduces a robust approach to preserve the…

Cited by 18PDFScholar