← Search

Chi-Man Pun

34 accepted papers

2026

GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors

ICML 2026poster

Reconstructing 3D scenes using 3D Gaussian Splatting (3DGS) from sparse views is an ill-posed problem due to insufficient information, often resulting in noticeable artifacts. While recent approaches have sought to leverage generative priors to complete information for under-constrained regions, the…

Cited by 0SourceScholar
2026

Glimpse: Geometry Learning of Multi-scale Structural Priors for 3D Pose Estimation

ICML 2026poster

Monocular 3D human pose estimation is fundamentally challenged by severe occlusion and inherent depth ambiguity. To address this, we propose Glimpse, a framework that learns robust 3D poses by explicitly modeling anatomical geometry from a single image. We recast the problem as geometry learning of …

Cited by 0SourceScholar
2026

IO-RAE: Information-Obfuscation Reversible Adversarial Example for Audio Privacy Protection

AAAI 2026technical

The rapid advancements in artificial intelligence have significantly accelerated the adoption of speech recognition technology, leading to its widespread integration across various applications. However, this surge in usage also highlights a critical issue: audio data is highly vulnerable to unautho

Cited by 0SourcePDFScholar
2026

Intrinsic Concept Extraction Based on Compositional Interpretability

CVPR 2026

Unsupervised Concept Extraction aims to extract concepts from a single image, yet existing methods suffer from the inability to extract composable intrinsic concepts. To address this, this paper introduces a new task called Compositional and Interpretable Intrinsic Concept Extraction (CI-ICE). The C

Cited by 0SourceScholar
2026

MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence

ICML 2026poster

Humans are born with vision-based 4D spatial-temporal intelligence, which enables us to perceive and reason about the evolution of 3D space over time from purely visual inputs. Despite its importance, this capability remains a significant bottleneck for current multimodal large language models (MLLM…

Cited by 0SourceScholar
2026

OTI: A Model-free and Visually Interpretable Measure of Image Attackability

AAAI 2026technical

Despite the tremendous success of neural networks, benign images can be corrupted by adversarial perturbations to deceive these models. Intriguingly, images differ in their attackability. Specifically, given an attack configuration, some images are easily corrupted, whereas others are more resistant

Cited by 0SourcePDFScholar
2026

PersonaLive! Expressive Portrait Image Animation for Live Streaming

CVPR 2026

Current diffusion-based portrait animation models predominantly focus on enhancing visual quality and expression realism, while overlooking generation latency and real-time performance, which restricts their application range in the live streaming scenario. We propose PersonaLive, a novel diffusion-

Cited by 0SourcecodeScholar
2026

ProMist-5K: A Comprehensive Dataset for Digital Emulation of Cinematic Pro-Mist Filter Effects

ICASSP 2026poster

Pro-Mist filters are widely used in cinematography for their ability to create soft halation, lower contrast, and produce a distinctive, atmospheric style. These effects are difficult to reproduce digitally due to the complex behavior of light diffusion. We present ProMist-5K, a dataset designed to…

Cited by 0SourcePDFScholar
2026

RevINN: An End-to-End Invertible Neural Network for Reversible Adversarial Examples Generation

CVPR 2026

Recent studies have shown that Reversible Adversarial Examples (RAE) can mislead unauthorized deep neural networks while remaining usable for authorized users, effectively preventing image data leakage. Existing RAE methods rely on reversibly embedding perturbation information into the original adve

Cited by 0SourcecodeScholar
2026

SNS-Grasp: Semantic-guided Noise Scaling for Grasp Generation

AAAI 2026technical

While diffusion models show promise for intent-based grasp generation, their isotropic noise schedules struggle with joint-specific sensitivity and task-aware variability. This limitation leads to grasps with suboptimal semantic alignment or physical feasibility. To address this challenge, we propos

Cited by 0SourcePDFScholar
2025

An Adaptive Framework for Multi-View Clustering Leveraging Conditional Entropy Optimization

ICASSP 2025accepted

Multi-view clustering (MVC) has emerged as a powerful technique for extracting valuable insights from data characterized by multiple perspectives or modalities. Despite significant advancements, existing MVC methods struggle with effectively quantifying the consistency and complementarity among view…

Cited by 0SourceScholar
2025

ForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and Localization

NeurIPS 2025poster

The field of Fake Image Detection and Localization (FIDL) is highly fragmented, encompassing four domains: deepfake detection (Deepfake), image manipulation detection and localization (IMDL), artificial intelligence-generated image detection (AIGC), and document image manipulation localization (Doc)…

Cited by 0SourcecodeScholar
2025

LensNet: An End-to-End Learning Framework for Empirical Point Spread Function Modeling and Lensless Imaging Reconstruction

IJCAI 2025

Lensless imaging stands out as a promising alternative to conventional lens-based systems, particularly in scenarios demanding ultracompact form factors and cost-effective architectures. However, such systems are fundamentally governed by the Point Spread Function (PSF), which dictates how a point s

2025

Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation Localization

AAAI 2025technical

The mesoscopic level serves as a bridge between the macroscopic and microscopic worlds, addressing gaps overlooked by both. Image manipulation localization (IML), a crucial technique to pursue truth from fake images, has long relied on low-level (microscopic-level) traces. However, in practice, most…

2025

Underwater Image Restoration via Polymorphic Large Kernel CNNs

ICASSP 2025accepted

Underwater Image Restoration (UIR) remains a challenging task in computer vision due to the complex degradation of images in underwater environments. While recent approaches have leveraged various deep learning techniques, including Transformers and complex, parameter-heavy models to achieve signifi…

Cited by 0SourceScholar
2024

DeformMLP: Dynamic Large-Scale Receptive Field MLP Networks for Human Motion Prediction

ICASSP 2024accepted

Predicting human motion requires addressing dependencies and errors for pose forecasting from sequences. The transformer’s self-attention aids this, but its complexity poses computational challenges. We present an efficient DeformMLP network without self-attention, using fully connected layers. Defo…

Cited by 0SourceScholar
2024

Devignet: High-Resolution Vignetting Removal via a Dual Aggregated Fusion Transformer with Adaptive Channel Expansion

AAAI 2024technical

Vignetting commonly occurs as a degradation in images resulting from factors such as lens design, improper lens hood usage, and limitations in camera sensors. This degradation affects image details, color accuracy, and presents challenges in computational photography. Existing vignetting removal alg…

2024

Generalized Uncertainty-Based Evidential Fusion with Hybrid Multi-Head Attention for Weak-Supervised Temporal Action Localization

ICASSP 2024accepted

Weakly supervised temporal action localization (WS-TAL) is a task of targeting at localizing complete action instances and categorizing them with video-level labels. Action-background ambiguity, primarily caused by background noise resulting from aggregation and intra-action variation, is a signific…

Cited by 0SourceScholar
2024

IMDL-BenCo: A Comprehensive Benchmark and Codebase for Image Manipulation Detection & Localization

NeurIPS 2024spotlight

A comprehensive benchmark is yet to be established in the Image Manipulation Detection \& Localization (IMDL) field. The absence of such a benchmark leads to insufficient and misleading model evaluations, severely undermining the development of this field. However, the scarcity of open-sourced basel…

2024

Local Optimization Networks for Multi-View Multi-Person Human Posture Estimation

ICASSP 2024accepted

With the growing applicability of multi-view multi-person 3D human pose estimation across diverse scenarios, the impact of external environmental factors and occlusion on accuracy has garnered substantial attention. In this research, we introduce a novel approach to multi-view multi-person 3D human…

Cited by 0SourceScholar
2024

PVALane: Prior-Guided 3D Lane Detection with View-Agnostic Feature Alignment

AAAI 2024technical

Monocular 3D lane detection is essential for a reliable autonomous driving system and has recently been rapidly developing. Existing popular methods mainly employ a predefined 3D anchor for lane detection based on front-viewed (FV) space, aiming to mitigate the effects of view transformations. Howev…

Cited by 6SourcePDFScholar
2023

A Large-Scale Film Style Dataset for Learning Multi-frequency Driven Film Enhancement

IJCAI 2023poster

Film, a classic image style, is culturally significant to the whole photographic industry since it marks the birth of photography. However, film photography is time-consuming and expensive, necessitating a more efficient method for collecting film-style photographs. Numerous datasets that have emerg…

2023

Boosting Face Recognition Performance with Synthetic Data and Limited Real Data

ICASSP 2023accepted

Face recognition is one of the most precise and straightforward methods to establish individual identity, and is important in our daily life. To solve the issues of privacy, bias, and collection difficulty caused by face recognition relying heavily on collecting a huge number of real face images fro…

Cited by 0SourceScholar
2023

CoordFill: Efficient High-Resolution Image Inpainting via Parameterized Coordinate Querying

AAAI 2023technical

Image inpainting aims to fill the missing hole of the input. It is hard to solve this task efficiently when facing high-resolution images due to two reasons: (1) Large reception field needs to be handled for high-resolution image inpainting. (2) The general encoder and decoder network synthesizes ma…

2023

High-Resolution Document Shadow Removal via A Large-Scale Real-World Dataset and A Frequency-Aware Shadow Erasing Net

ICCV 2023poster

Shadows often occur when we capture the document with casual equipment, which influences the visual quality and readability of the digital copies. Different from the algorithms for natural shadow removal, the algorithms in document shadow removal need to preserve the details of fonts and figures in…

Cited by 67PDFcodeScholar
2023

Locality Preserving Multiview Graph Hashing For Large Scale Remote Sensing Image Search

ICASSP 2023accepted

Hashing is very popular for remote sensing image search. This article proposes a multiview hashing with learnable parameters to retrieve the queried images for a large-scale remote sensing dataset. Existing methods always neglect that real-world remote sensing data lies on a low- dimensional manifol…

Cited by 0SourceScholar
2023

Shadocnet: Learning Spatial-Aware Tokens in Transformer for Document Shadow Removal

ICASSP 2023accepted

Shadow removal improves the visual quality and legibility of digital copies of documents. However, document shadow removal remains an unresolved subject. Traditional techniques rely on heuristics that vary from situation to situation. Given the quality and quantity of current public datasets, the ma…

Cited by 0SourceScholar
2022

Spatial-Separated Curve Rendering Network for Efficient and High-Resolution Image Harmonization

ECCV 2022poster

"Image harmonization aims to modify the color of the composited region according to the specific background. Previous works model this task as a pixel-wise image translation using UNet family structures. However, the model size and computational cost limit the ability of their models on edge devices…

2021

Split then Refine: Stacked Attention-guided ResUNets for Blind Single Image Visible Watermark Removal

AAAI 2021technical

Digital watermark is a commonly used technique to protect the copyright of medias. Simultaneously, to increase the robustness of watermark, attacking technique, such as watermark removal, also gets the attention from the community. Previous watermark removal methods require to gain the watermark loc…

2019

Audio Replay Spoof Attack Detection Using Segment-based Hybrid Feature and DenseNet-LSTM Network

ICASSP 2019accepted

At present, most automatic speaker verification (ASV) systems are vulnerable to replay spoof attacks. Therefore, this paper proposes a new approach for the detection of audio replay spoof attacks. Here, a segment-based hybrid feature extraction method is used, which includes the Mel-frequency cepstr…

Cited by 0SourceScholar