← Search

Qing Zhang

41 accepted papers

2026

A Game-Theoretic Framework for Measuring and Explaining Metric Compatibility in Fair Machine Learning

ICML 2026poster

Machine learning fairness research documents trade-offs but lacks quantitative frameworks to measure intrinsic metric compatibility without requiring causal graphs. We introduce a game-theoretic framework that decomposes metrics into interaction vectors, enabling compatibility measurement between me…

Cited by 0SourceScholar
2026

Beyond “Made with AI”: Visualizing Provenance Density to Mitigate the Transparency Penalty

IJCAI 2026

As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI-generated. Binary "Made with AI" labels r

Cited by 0Scholar
2026

DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts

AAAI 2026technical

Chart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test surface-level parsing, such as reading labels and legends, while overlooking deeper scientific reasoning. We propose Domai

Cited by 0SourcePDFScholar
2026

Learning Explicit Continuous Motion Representation for Dynamic Gaussian Splatting from Monocular Videos

CVPR 2026

We present an approach for high-quality dynamic Gaussian Splatting from monocular videos. To this end, we in this work go one step further beyond previous methods to explicitly model continuous position and orientation deformation of dynamic Gaussians, using an SE(3) B-spline motion bases with a com

Cited by 0SourcecodeScholar
2026

Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models

CVPR 2026

What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction, and interaction perception, which models how an agent's actio

Cited by 0SourceScholar
2026

REINPATH: A MULTIMODAL REINFORCEMENT LEARNING APPROACH FOR PATHOLOGY

ICASSP 2026poster

Interpretability is significant in computational pathology, leading to the development of multimodal information integration from histopathological image and corresponding text data.However, existing multimodal methods have limited interpretability due to the lack of high-quality dataset that suppor…

Cited by 0SourcePDFScholar
2026

RehearseVLA: Simulated Post-Training for VLAs with Physically-Consistent World Model

CVPR 2026

Vision-Language-Action (VLA) models trained via imitation learning suffer from significant performance degradation in data-scarce scenarios due to their reliance on large-scale demonstration datasets. Although reinforcement learning (RL)-based post-training has proven effective in addressing data sc

Cited by 0SourcecodeScholar
2026

SGS-Intrinsic: Semantic-Invariant Gaussian Splatting for Sparse-View Indoor Inverse Rendering

CVPR 2026

We present SGS-Intrinsic, an indoor inverse rendering framework that works well for sparse-view images. Unlike existing 3D Gaussian Splatting (3DGS) based methods that focus on object-centric reconstruction and fail to work under sparse view settings, our method allows to achieve high-quality geomet

Cited by 0SourcecodeScholar
2026

SynerDetect: Hierarchical Synergistic Learning for Generalizable AI-Generated Image Detection

AAAI 2026technical

The rapid advancement of generative models, which produce increasingly realistic synthetic images, urgently demands robust and generalizable detection methods. Consequently, research has largely pivoted to leveraging large-scale Vision Foundation Models (VFMs) for enhanced generalization. However, e

Cited by 0SourcePDFScholar
2026

You Only Erase Once: Erasing Anything without Bringing Unexpected Content

CVPR 2026

We present YOEO, an approach for object erasure. Unlike recent diffusion-based methods which struggle to erase target objects without generating unexpected content within the masked regions due to lack of sufficient paired training data and explicit constraint on content generation, our method allow

Cited by 0SourcecodeScholar
2025

CLIP-RestoreX: Restore Image Structure and Perception in Exposure Correction

AAAI 2025technical

Exposure correction aims to adjust the exposure of an under- and over-exposed image to enhance its overall visual quality. The core challenge of this task lies in that it requires to faithfully restore both the structure and perception information. In this work, we present a novel exposure correctio…

2025

DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse Rendering

ICCV 2025poster

Recent methods have shown that pre-trained diffusion models can be fine-tuned to enable generative inverse rendering by learning image-conditioned noise-to-intrinsic mapping. Despite their remarkable progress, they struggle to robustly produce high-quality results as the noise-to-intrinsic paradigm…

2025

Dynamic Gaussian Splatting from Defocused and Motion-blurred Monocular Videos

NeurIPS 2025poster

This paper presents a unified framework that allows high-quality dynamic Gaussian Splatting from both defocused and motion-blurred monocular videos. Due to the significant difference between the formation processes of defocus blur and motion blur, existing methods are tailored for either one of them…

Cited by 0SourcecodeScholar
2025

EntityErasure: Erasing Entity Cleanly via Amodal Entity Segmentation and Completion

CVPR 2025poster

This paper presents EntityErasure, a novel diffusion-based inpainting method that can effectively erase entities without inducing unwanted sundries. To this end, we propose to address this problem by dividing it into amodal entity segmentation and completion, such that the region to inpaint takes on…

2025

Evaluating Instructively Generated Statement by Large Language Models for Directional Event Causality Identification

ACL 2025finding

This paper aims to identify directional causal relations between events, including the existence and direction of causality. Previous studies mainly adopt prompt learning paradigm to predict a causal answer word based on a Pre-trained Language Model (PLM) for causality existence identification. Howe…

Cited by 0SourcePDFScholar
2025

Rethinking Camouflaged Object Detection via Foreground-Background Interactive Learning

ICASSP 2025accepted

Camouflaged object detection focuses on the challenge of segmenting objects that visually blend into their background. The effectiveness of camouflage strategies hinges on how well objects interact with their background to minimize their visibility. Based on this insight, we propose a novel Foregrou…

Cited by 0SourceScholar
2025

RoGSplat: Learning Robust Generalizable Human Gaussian Splatting from Sparse Multi-View Images

CVPR 2025poster

This paper presents RoGSplat, a novel approach for synthesizing high-fidelity novel views of unseen human from sparse multi-view images, while requiring no cumbersome per-subject optimization. Unlike previous methods that typically struggle with sparse views with few overlappings and are less effect…

2025

Stable-Hair: Real-World Hair Transfer via Diffusion Model

AAAI 2025technical

Current hair transfer methods struggle to handle diverse and intricate hairstyles, limiting their applicability in real-world scenarios. In this paper, we propose a novel diffusion-based hair transfer framework, named Stable-Hair, which robustly transfers a wide range of real-world hairstyles to use…

2025

Structural-Aware Disentangled Learning with CLIP for Hyperbolic Zero-Shot Sketch-Based Image Retrieval

ICASSP 2025accepted

The zero-shot sketch-based image retrieval task faces two key challenges: domain gap and knowledge transfer. Our innovation is recognizing that directly aligning cross-domain features weakens the discriminative ability of the model, as it overlooks the asymmetry between sketches and images. Addition…

Cited by 0SourceScholar
2025

Structure-Guided Diffusion Models for High-Fidelity Portrait Shadow Removal

ICCV 2025poster

We present a diffusion-based portrait shadow removal approach that can robustly produce high-fidelity results. Unlike previous methods, we cast shadow removal as diffusion-based inpainting. To this end, we first train a shadow-independent structure extraction network on a real-world portrait dataset…

2025

When CLIP Meets PHOC: A Dual-Branch Network for Historical Document Image Retrieval

ICASSP 2025accepted

In this paper, we leverage Contrastive Language-Image Pre-training (CLIP) for Historical Document Image Retrieval (HDIR). We are largely inspired by recent advances on CLIP and its exceptional generalization capabilities, but for the first time, we tailor it to benefit HDIR. We put forward a dual-br…

Cited by 0SourceScholar
2025

When Shadow Removal Meets Intrinsic Image Decomposition: A Joint Learning Framework Using Unpaired Data

AAAI 2025technical

We present a framework that achieves shadow removal by learning intrinsic image decomposition (IID) from unpaired shadow and shadow-free images. Although it is well-known that intrinsic images, \ie, illumination and reflectance, are highly beneficial to shadow removal, IID is rarely adopted by previ…

Cited by 0SourcePDFScholar
2024

HENet: Hyperbolic-Based Encoder-Decoder Network for Word Spotting in Historical Mongolian Documents

ICASSP 2024accepted

In the domain of historical Mongolian document image retrieval (HMDIR), word spotting poses a inherent challenge due to the frequent appearance of out-of-vocabulary (OOV) words. Existing methods have mainly focused on query-by-example (QBE), neglecting the query-by-string (QBS) approach. Meanwhile,…

Cited by 0SourceScholar
2024

Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh Recovery

NeurIPS 2024poster

Human Mesh Recovery (HMR) is the task of estimating a parameterized 3D human mesh from an image. There is a kind of methods first training a regression model for this problem, then further optimizing the pretrained regression model for any specific sample individually at test time. However, the pret…

2024

Interleaving One-Class and Weakly-Supervised Models with Adaptive Thresholding for Unsupervised Video Anomaly Detection

ECCV 2024poster

"Video Anomaly Detection (VAD) has been extensively studied under the settings of One-Class Classification (OCC) and Weakly-Supervised learning (WS), which however both require laborious human-annotated normal/abnormal labels. In this paper, we study Unsupervised VAD (UVAD) that does not depend on a…

2024

Multi-RoI Human Mesh Recovery with Camera Consistency and Contrastive Losses

ECCV 2024poster

"Besides a 3D mesh, Human Mesh Recovery (HMR) methods usually need to estimate a camera for computing 2D reprojection loss. Previous approaches may encounter the following problem: both the mesh and camera are not correct but the combination of them can yield a low reprojection loss. To alleviate th…

2023

Document Image Shadow Removal Guided by Color-Aware Background

CVPR 2023poster

Existing works on document image shadow removal mostly depend on learning and leveraging a constant background (the color of the paper) from the image. However, the constant background is less representative and frequently ignores other background colors, such as the printed colors, resulting in dis…

2023

PosterLayout: A New Benchmark and Approach for Content-Aware Visual-Textual Presentation Layout

CVPR 2023poster

Content-aware visual-textual presentation layout aims at arranging spatial space on the given canvas for pre-defined elements, including text, logo, and underlay, which is a key to automatic template-free creative graphic design. In practical applications, e.g., poster designs, the canvas is origina…

2023

Towards High-Quality Specular Highlight Removal by Leveraging Large-Scale Synthetic Data

ICCV 2023poster

This paper aims to remove specular highlights from a single object-level image. Although previous methods have made some progresses, their performance remains somewhat limited, particularly for real images with complex specular highlights. To this end, we propose a three-stage network to address the…

Cited by 14PDFcodeScholar
2022

Progressively Generating Better Initial Guesses Towards Next Stages for High-Quality Human Motion Prediction

CVPR 2022poster

This paper presents a high-quality human motion prediction method that accurately predicts future human poses given observed ones. Our method is based on the observation that a good initial guess of the future poses is very helpful in improving the forecasting accuracy. This motivates us to propose…

Cited by 140PDFcodeScholar
2021

A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame Prediction

ICCV 2021poster

In this paper, we propose HF2-VAD, a Hybrid framework that integrates Flow reconstruction and Frame prediction seamlessly to handle Video Anomaly Detection. Firstly, we design the network of ML-MemAE-SC (Multi-Level Memory modules in an Autoencoder with Skip Connections) to memorize normal patterns…

Cited by 278PDFcodeScholar
2021

A Multi-Task Network for Joint Specular Highlight Detection and Removal

CVPR 2021poster

Specular highlight detection and removal are fundamental and challenging tasks. Although recent methods achieve promising results on the two tasks by supervised training on synthetic training data, they are typically solely designed for highlight detection or removal, and their performance usually d…

Cited by 100PDFcodeScholar
2021

MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion Prediction

ICCV 2021poster

Human motion prediction is a challenging task due to the stochasticity and aperiodicity of future poses. Recently, graph convolutional network has been proven to be very effective to learn dynamic relations among pose joints, which is helpful for pose prediction. On the other hand, one can abstract…

Cited by 262PDFcodeScholar
2019

Deep Multi-Model Fusion for Single-Image Dehazing

ICCV 2019poster

This paper presents a deep multi-model fusion network to attentively integrate multiple models to separate layers and boost the performance in single-image dehazing. To do so, we first formulate the attentional feature integration module to maximize the integration of the convolutional neural networ…

Cited by 146PDFScholar
2019

Underexposed Photo Enhancement Using Deep Illumination Estimation

CVPR 2019oral

This paper presents a new neural network for enhancing underexposed photos. Instead of directly learning an image-to-image mapping as previous work, we introduce intermediate illumination in our network to associate the input with expected enhancement result, which augments the network's capability…

Cited by 1084PDFcodeScholar