← Search

Shiqi Wang

67 accepted papers

2026

CADC: Content Adaptive Diffusion-Based Generative Image Compression

CVPR 2026

Diffusion-based generative image compression has demonstrated remarkable potential for achieving realistic reconstruction at ultra-low bitrates. The key to unlocking this potential lies in making the entire compression process content-adaptive, ensuring that the encoder's representation and the deco

Cited by 0SourceScholar
2026

DeepRAHT: Learning Predictive RAHT for Point Cloud Attribute Compression

AAAI 2026technical

Regional Adaptive Hierarchical Transform (RAHT) is an effective point cloud attribute compression (PCAC) method. However, its application in deep learning lacks research. In this paper, we propose an end-to-end RAHT framework for lossy PCAC based on the sparse tensor, called DeepRAHT. The RAHT trans

Cited by 0SourcePDFScholar
2026

Dual Graph Regularized Deep Unfolding Network for Guided Depth Map Super-resolution

CVPR 2026

Depth map super-resolution with color guidance is a fundamental task in computer vision that aims to reconstruct high-resolution depth maps by leveraging structural correlations from corresponding guidance images. Recently, with the development of deep learning techniques, the performance of guided

Cited by 0SourceScholar
2026

MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation

ICML 2026poster

Generating realistic 3D Human-Object Interactions (HOI) is a fundamental task for applications ranging from embodied AI to virtual content creation, which requires harmonizing high-level semantic intent with strict low-level physical constraints. Existing methods excel at semantic alignment, however…

Cited by 0SourceScholar
2026

When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM Editing

AAAI 2026technical

Privacy leakage in Multimodal Large Language Models (MLLMs) has long been an intractable problem. Existing studies, though effectively obscure private information in MLLMs, often overlook the evaluation of authenticity and recovery quality of user privacy. To this end, this work uniquely focuses on

Cited by 0SourcePDFScholar
2025

AI-generated Image Quality Assessment in Visual Communication

AAAI 2025technical

Assessing the quality of artificial intelligence-generated images (AIGIs) plays a crucial role in their application in real-world scenarios. However, traditional image quality assessment (IQA) algorithms primarily focus on low-level visual perception, while existing IQA works on AIGIs overemphasize…

2025

An Information-Theoretic Regularizer for Lossy Neural Image Compression

ICCV 2025poster

Lossy image compression networks aim to minimize the latent entropy of images while adhering to specific distortion constraints. However, optimizing the neural network can be challenging due to its nature of learning quantized latent representations. In this paper, our key finding is that minimizing…

Cited by 0SourcePDFScholar
2025

Generating Commonsense Reasoning Questions with Controllable Complexity through Multi-step Structural Composition

COLING 2025main

This paper studies the task of generating commonsense reasoning questions (QG) with desired difficulty levels. Compared to traditional shallow questions that can be solved by simple term matching, ours are more challenging. Our answering process requires reasoning over multiple contextual and common…

Cited by 1SourcePDFScholar
2025

Joint Semantic and Rendering Enhancements in 3D Gaussian Modeling with Anisotropic Local Encoding

ICCV 2025poster

Recent works propose extending 3DGS with semantic feature vectors for simultaneous semantic segmentation and image rendering. However, these methods often treat the semantic and rendering branches separately, relying solely on 2D supervision while ignoring the 3D Gaussian geometry. Moreover, current…

2025

Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need

NeurIPS 2025poster

We have recently witnessed that ''Intelligence" and `''Compression" are the two sides of the same coin, where the language large model (LLM) with unprecedented intelligence is a general-purpose lossless compressor for various data modalities. This attribute is particularly appealing to the lossless…

Cited by 0SourcecodeScholar
2025

MR-FIQA: Face Image Quality Assessment with Multi-Reference Representations from Synthetic Data Generation

ICCV 2025poster

Recent advancements in Face Image Quality Assessment (FIQA) models trained on real large-scale face datasets are pivotal in guaranteeing precise face recognition in unrestricted scenarios. Regrettably, privacy concerns lead to the discontinuation of real datasets, underscoring the pressing need for…

2025

Making Old Film Great Again: Degradation-aware State Space Model for Old Film Restoration

CVPR 2025poster

Unlike modern native digital videos, the restoration of old films requires addressing specific degradations inherent to analog sources. However, existing specialized methods still fall short compared to general video restoration techniques. In this work, we propose a new baseline to re-examine the c…

2025

MicroEdit: Neuron-level Knowledge Disentanglement and Localization in Lifelong Model Editing

EMNLP 2025

Large language models (LLMs) require continual knowledge updates to keep pace with the evolving world. While various model editing methods have been proposed, most face critical challenges in the context of lifelong learning due to two fundamental limitations: (1) Edit Overshooting - parameter updat

2025

PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

ICML 2025poster

Recent research builds various patching agents that combine large language models (LLMs) with non-ML tools and achieve promising results on the state-of-the-art (SOTA) software patching benchmark, SWE-bench. Based on how to determine the patching workflows, existing patching agents can be categoriz…

2025

Planning-Aware Code Infilling via Horizon-Length Prediction

EMNLP 2025

Fill-in-the-Middle (FIM), or infilling, has become integral to code language models, enabling generation of missing code given both left and right contexts. However, the current FIM training paradigm which performs next-token prediction (NTP) over reordered sequence often leads to models struggling

Cited by 0SourcePDFScholar
2025

Q-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction Regression

CVPR 2025poster

In this work, we address the challenge of adaptive pediatric Left Ventricular Ejection Fraction (LVEF) assessment. While Test-time Training (TTT) approaches show promise for this task, they suffer from two significant limitations. Existing TTT works are primarily designed for classification tasks ra…

Cited by 0SourcePDFScholar
2025

Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval

NeurIPS 2025poster

The success of DeepSeek-R1 demonstrates the immense potential of using reinforcement learning (RL) to enhance LLMs' reasoning capabilities. This paper introduces Retrv-R1, the first R1-style MLLM specifically designed for multimodal universal retrieval, achieving higher performance by employing step…

Cited by 0SourceScholar
2025

Speed Master: Quick or Slow Play to Attack Speaker Recognition

AAAI 2025technical

Backdoor attacks pose a significant threat during the model's training phase. Attackers craft pre-defined triggers to break deep neural networks, ensuring the model accurately classifies clean samples during inference yet erroneously classifies samples added with these triggers. Recent studies have…

Cited by 0SourcePDFScholar
2025

Test-time Adaptation for Image Compression with Distribution Regularization

ICLR 2025poster

Current test- or compression-time adaptation image compression (TTA-IC) approaches, which leverage both latent and decoder refinements as a two-step adaptation scheme, have potentially enhanced the rate-distortion (R-D) performance of learned image compression models on cross-domain compression task…

Cited by 1SourcePDFScholar
2025

Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision

AAAI 2025technical

The image compression model has long struggled with adaptability and generalization, as the decoded bitstream typically serves only human or machine needs and fails to preserve information for unseen visual tasks. Therefore, this paper innovatively introduces supervision obtained from multimodal pre…

2025

daDPO: Distribution-Aware DPO for Distilling Conversational Abilities

ACL 2025finding

Large language models (LLMs) have demonstrated exceptional performance across various applications, but their conversational abilities decline sharply as model size decreases, presenting a barrier to their deployment in resource-constrained environments. Knowledge distillation (KD) with Direct Prefe…

2024

Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare

NeurIPS 2024spotlight

While recent advancements in large multimodal models (LMMs) have significantly improved their abilities in image quality assessment (IQA) relying on absolute quality rating, how to transfer reliable relative quality comparison outputs to continuous perceptual quality scores remains largely unexplore…

2024

CLIB-FIQA: Face Image Quality Assessment with Confidence Calibration

CVPR 2024poster

Face Image Quality Assessment (FIQA) is pivotal for guaranteeing the accuracy of face recognition in unconstrained environments. Recent progress in deep quality-fitting-based methods that train models to align with quality anchors has shown promise in FIQA. However these methods heavily depend on a…

2024

CodeFort: Robust Training for Code Generation Models

EMNLP 2024finding

Code generation models are not robust to small perturbations, which often lead to incorrect generations and significantly degrade the performance of these models. Although improving the robustness of code generation models is crucial to enhancing user experience in real-world applications, existing…

Cited by 1SourcePDFScholar
2024

DDR: Exploiting Deep Degradation Response as Flexible Image Descriptor

NeurIPS 2024poster

Image deep features extracted by pre-trained networks are known to contain rich and informative representations. In this paper, we present Deep Degradation Response (DDR), a method to quantify changes in image deep features under varying degradation conditions. Specifically, our approach facilitates…

2024

Image Coding for Analytics via Adversarially Augmented Adaptation

ICASSP 2024accepted

Image Coding for Machine (ICM) aims to compress an image so that the reconstructed one can meet the requirements of both human vision and machine vision. Existing methods apply the constraint from the downstream models to improve machine analytics performance while compromising the visual quality. T…

Cited by 0SourceScholar
2024

LeDex: Training LLMs to Better Self-Debug and Explain Code

NeurIPS 2024poster

In the domain of code generation, self-debugging is crucial. It allows LLMs to refine their generated code based on execution feedback. This is particularly important because generating correct solutions in one attempt proves challenging for complex tasks. Prior works on self-debugging mostly focus…

Cited by 4SourcePDFScholar
2024

Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence Regularization

AAAI 2024technical

In open-domain Question Answering (QA), dense text retrieval is crucial for finding relevant passages to generate answers. Typically, contrastive learning is used to train a retrieval model, which maps passages and queries to the same semantic space, making similar ones closer and dissimilar ones fu…

2024

Multimodal Clickbait Detection by De-confounding Biases Using Causal Representation Inference

EMNLP 2024main

This paper focuses on detecting clickbait posts on the Web. These posts often use eye-catching disinformation in mixed modalities to mislead users to click for profit. That affects the user experience and thus would be blocked by content provider. To escape detection, malicious creators use tricks t…

Cited by 0SourcePDFScholar
2024

Neuron Activation Coverage: Rethinking Out-of-distribution Detection and Generalization

ICLR 2024spotlight

The out-of-distribution (OOD) problem generally arises when neural networks encounter data that significantly deviates from the training data distribution, i.e., in-distribution (InD). In this paper, we study the OOD problem from a neuron activation view. We first formulate neuron activation states…

2024

PKU-DyMVHumans: A Multi-View Video Benchmark for High-Fidelity Dynamic Human Modeling

CVPR 2024poster

High-quality human reconstruction and photo-realistic rendering of a dynamic scene is a long-standing problem in computer vision and graphics. Despite considerable efforts invested in developing various capture systems and reconstruction algorithms recent advancements still struggle with loose or ov…

2024

ReTA: Recursively Thinking Ahead to Improve the Strategic Reasoning of Large Language Models

NAACL 2024long

Current logical reasoning evaluations of Large Language Models (LLMs) primarily focus on single-turn and static environments, such as arithmetic problems. The crucial problem of multi-turn, strategic reasoning is under-explored. In this work, we analyze the multi-turn strategic reasoning of LLMs thr…

2024

Reward Difference Optimization For Sample Reweighting In Offline RLHF

EMNLP 2024finding

With the wide deployment of Large Language Models (LLMs), aligning LLMs with human values becomes increasingly important. Although Reinforcement Learning with Human Feedback (RLHF) proves effective, it is complicated and highly resource-intensive. As such, offline RLHF has been introduced as an alte…

2024

ScreenAgent: A Vision Language Model-driven Computer Control Agent

IJCAI 2024poster

Large Language Models (LLM) can invoke a variety of tools and APIs to complete complex tasks. The computer, as the most powerful and universal tool, could potentially be controlled by a trained LLM agent. Powered by the computer, we can hopefully build a more generalized agent to assist humans in va…

2024

Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models

ACL 2024long

Large Language Models (LLMs) show promising results in language generation and instruction following but frequently “hallucinate”, making their outputs less reliable. Despite Uncertainty Quantification’s (UQ) potential solutions, implementing it accurately within LLMs is challenging. Our research in…

2024

Token Alignment via Character Matching for Subword Completion

ACL 2024findings

Generative models, widely utilized in various applications, can often struggle with prompts corresponding to partial tokens. This struggle stems from tokenization, where partial tokens fall out of distribution during inference, leading to incorrect or nonsensical outputs. This paper examines a techn…

Cited by 1SourcePDFScholar
2024

Towards Open-ended Visual Quality Comparison

ECCV 2024oral

"Comparative settings (pairwise choice, listwise ranking) have been adopted by a wide range of subjective studies for image quality assessment (IQA), as it inherently standardizes the evaluation criteria across different observers and offer more clear-cut responses. In this work, we extend the edge…

2023

Are Diffusion Models Vulnerable to Membership Inference Attacks?

ICML 2023poster

Diffusion-based generative models have shown great potential for image synthesis, but there is a lack of research on the security and privacy risks they may pose. In this paper, we investigate the vulnerability of diffusion models to Membership Inference Attacks (MIAs), a common privacy concern. Our…

2023

Generating Deep Questions with Commonsense Reasoning Ability from the Text by Disentangled Adversarial Inference

ACL 2023findings

This paper proposes a new task of commonsense question generation, which aims to yield deep-level and to-the-point questions from the text. Their answers need to reason over disjoint relevant contexts and external commonsense knowledge, such as encyclopedic facts and causality. The knowledge may not…

Cited by 9SourcePDFScholar
2023

Multi-lingual Evaluation of Code Generation Models

ICLR 2023top-25%

We present two new benchmarks, MBXP and Multilingual HumanEval, designed to evaluate code completion models in over 10 programming languages. These datasets are generated using a conversion framework that transpiles prompts and test cases from the original MBPP and HumanEval datasets into the corres…

2023

Optimization-Inspired Cross-Attention Transformer for Compressive Sensing

CVPR 2023poster

By integrating certain optimization solvers with deep neural networks, deep unfolding network (DUN) with good interpretability and high performance has attracted growing attention in compressive sensing (CS). However, existing DUNs often improve the visual quality at the price of a large number of p…

2023

ReCode: Robustness Evaluation of Code Generation Models

ACL 2023long

Code generation models have achieved impressive performance. However, they tend to be brittle as slight edits to a prompt could lead to very different generations; these robustness properties, critical for user experience when deployed in real-life applications, are not well understood. Most existin…

2023

ScanDMM: A Deep Markov Model of Scanpath Prediction for 360deg Images

CVPR 2023poster

Scanpath prediction for 360deg images aims to produce dynamic gaze behaviors based on the human visual perception mechanism. Most existing scanpath prediction methods for 360deg images do not give a complete treatment of the time-dependency when predicting human scanpath, resulting in inferior perfo…

2023

Troubleshooting Ethnic Quality Bias with Curriculum Domain Adaptation for Face Image Quality Assessment

ICCV 2023poster

Face Image Quality Assessment (FIQA) lays the foundation for ensuring the stability and accuracy of face recognition systems. However, existing FIQA methods mainly formulate quality relationships within the training set to yield quality scores, ignoring the generalization problem caused by ethnic qu…

Cited by 9PDFcodeScholar
2023

Two-Branch Multi-Scale Deep Neural Network for Generalized Document Recapture Attack Detection

ICASSP 2023accepted

The image recapture attack is an effective image manipulation method to erase certain forensic traces, and when targeting on personal document images, it poses a great threat to the security of e-commerce and other web applications. Considering the current learning-based methods suffer from serious…

Cited by 0SourceScholar
2022

A Branch and Bound Framework for Stronger Adversarial Attacks of ReLU Networks

ICML 2022spotlight

Strong adversarial attacks are important for evaluating the true robustness of deep neural networks. Most existing attacks search in the input space, e.g., using gradient descent, and may miss adversarial examples due to non-convexity. In this work, we systematically search adversarial examples in t…

2022

General Cutting Planes for Bound-Propagation-Based Neural Network Verification

NeurIPS 2022accept

Bound propagation methods, when combined with branch and bound, are among the most effective methods to formally verify properties of deep neural networks such as correctness, robustness, and safety. However, existing works cannot handle the general form of cutting plane constraints widely accepted…

2022

Generalizing to Evolving Domains with Latent Structure-Aware Sequential Autoencoder

ICML 2022spotlight

Domain generalization aims to improve the generalization capability of machine learning systems to out-of-distribution (OOD) data. Existing domain generalization techniques embark upon stationary and discrete environments to tackle the generalization issue caused by OOD data. However, many real-worl…

2022

Rethinking Attention-Model Explainability through Faithfulness Violation Test

ICML 2022spotlight

Attention mechanisms are dominating the explainability of deep models. They produce probability distributions over the input, which are widely deemed as feature-importance indicators. However, in this paper, we find one critical limitation in attention explanations: weakness in identifying the polar…

2021

Adaptive Verifiable Training Using Pairwise Class Similarity

AAAI 2021technical

Verifiable training has shown success in creating neural networks that are provably robust to a given amount of noise. However, despite only enforcing a single robustness criterion, its performance scales poorly with dataset complexity. On CIFAR10, a non-robust LeNet model has a 21.63% error rate, w…

Cited by 3SourcePDFScholar
2021

Beta-CROWN: Efficient Bound Propagation with Per-neuron Split Constraints for Neural Network Robustness Verification

NeurIPS 2021poster

Bound propagation based incomplete neural network verifiers such as CROWN are very efficient and can significantly accelerate branch-and-bound (BaB) based complete verification of neural networks. However, bound propagation cannot fully handle the neuron split constraints introduced by BaB commonly…

Cited by 306SourcePDFScholar
2021

Fast and Complete: Enabling Complete Neural Network Verification with Rapid and Massively Parallel Incomplete Verifiers

ICLR 2021poster

Formal verification of neural networks (NNs) is a challenging and important problem. Existing efficient complete solvers typically require the branch-and-bound (BaB) process, which splits the problem domain into sub-domains and solves each sub-domain using faster but weaker incomplete verifiers, suc…

2021

Teacher-Student Learning With Multi-Granularity Constraint Towards Compact Facial Feature Representation

ICASSP 2021accepted

In this paper, we propose a novel end-to-end feature compression scheme by leveraging the representation and learning capability of deep neural networks, towards intelligent front-end equipped analysis with promising accuracy and efficiency. In particular, the extracted features are compactly coded…

Cited by 0SourceScholar
2020

Domain Generalization for Medical Imaging Classification with Linear-Dependency Regularization

NeurIPS 2020poster

Recently, we have witnessed great progress in the field of medical imaging classification by adopting deep neural networks. However, the recent advanced models still require accessing sufficiently large and representative datasets for training, which is often unfeasible in clinically realistic envir…

2020

From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement

CVPR 2020poster

Under-exposure introduces a series of visual degradation, i.e. decreased visibility, intensive noise, and biased color, etc. To address these problems, we propose a novel semi-supervised learning approach for low-light image enhancement. A deep recursive band network (DRBN) is proposed to recover a…

Cited by 657PDFScholar
2020

HYDRA: Pruning Adversarially Robust Neural Networks

NeurIPS 2020poster

In safety-critical but computationally resource-constrained applications, deep learning faces two key challenges: lack of robustness against adversarial attacks and large neural network size (often millions of parameters). While the research community has extensively explored the use of robust train…

2020

Intra Frame Rate Control for Versatile Video Coding with Quadratic Rate-Distortion Modelling

ICASSP 2020accepted

With numerous coding tools adopted in the forthcoming Versatile Video Coding (VVC) standard, much less work has been dedicated to study the corresponding Rate-Distortion (R-D) characteristics. This paper proposes a new quadratic R-D model for Versatile Video Coding. In particular, based on the propo…

Cited by 0SourceScholar
2020

Just Noticeable Distortion Based Perceptually Lossless Intra Coding

ICASSP 2020accepted

Perceptual video coding plays a very important role in video codec optimization aiming at removing the perceptual redundancies in video content. In this paper, a just noticeable distortion (JND) guided perceptually lossless coding framework is proposed for Versatile Video Coding (VVC) intra coding.…

Cited by 0SourceScholar
2020

Self-Learning Video Rain Streak Removal: When Cyclic Consistency Meets Temporal Correspondence

CVPR 2020poster

In this paper, we address the problem of rain streaks removal in video by developing a self-learned rain streak removal method, which does not require any clean groundtruth images in the training process. The method is inspired by fact that the adjacent frames are highly correlated and can be regard…

Cited by 82PDFcodeScholar
2019

Learning to Explore Intrinsic Saliency for Stereoscopic Video

CVPR 2019poster

The human visual system excels at biasing the stereoscopic visual signals by the attention mechanisms. Traditional methods relying on the low-level features and depth relevant information for stereoscopic video saliency prediction have fundamental limitations. For example, it is cumbersome to model…

Cited by 6PDFScholar
2019

VERI-Wild: A Large Dataset and a New Method for Vehicle Re-Identification in the Wild

CVPR 2019poster

Vehicle Re-identification (ReID) is of great significance to the intelligent transportation and public security. However, many challenging issues of Vehicle ReID in real-world scenarios have not been fully investigated, e.g., the high viewpoint variations, extreme illumination conditions, complex ba…

Cited by 352PDFScholar
2018

Cluster-Based Point Cloud Coding with Normal Weighted Graph Fourier Transform

ICASSP 2018accepted

Point cloud has attracted more and more attention in 3D object representation, especially in free-view rendering. However, it is challenging to efficiently deploy the point cloud due to its huge data amount with multiple attributes including coordinates, normal and color. In order to represent point…

Cited by 0SourceScholar
2018

Domain Generalization With Adversarial Feature Learning

CVPR 2018poster

In this paper, we tackle the problem of domain generalization: how to learn a generalized feature representation for an “unseen” target domain by taking the advantage of multiple seen source-domain data. We present a novel framework based on adversarial autoencoders to learn a generalized latent fea…

Cited by 1574SourcePDFScholar
2018

Efficient Formal Safety Analysis of Neural Networks

NeurIPS 2018poster

Neural networks are increasingly deployed in real-world safety-critical domains such as autonomous driving, aircraft collision avoidance, and malware detection. However, these networks have been shown to often mispredict on inputs with minor adversarial or even accidental perturbations. Consequences…

2018

Image Quality Assessment Based Label Smoothing in Deep Neural Network Learning

ICASSP 2018accepted

For many computer vision problems, deep neural networks are trained and validated based on the assumption that the input images are pristine (i.e., artifact-free). However, digital images are subject to a wide range of distortions in real application scenarios, while the practical issues regarding i…

Cited by 0SourceScholar