← Search

jiawei zhang

116 accepted papers

2026

ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks

ICLR 2026poster

As vision-language models (VLMs) gain prominence, their multimodal interfaces also introduce new safety vulnerabilities, making the safety evaluation challenging and critical. Existing red-teaming efforts are either restricted to a narrow set of adversarial patterns or depend heavily on manual engin…

Cited by 0SourceScholar
2026

AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondence

CVPR 2026

Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations due to limited data diversity. Leveraging powerful 3D generative models and vision foundation models (VFMs), the proposed AffordGen framework overcome

Cited by 0SourceScholar
2026

Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth

ICLR 2026poster

Large Language Models (LLMs) exhibit strong but shallow alignment: they directly refuse harmful queries when a refusal is expected at the very start of an assistant turn, yet this protection collapses once a harmful continuation is underway (either through the adversarial attacks or via harmful assi…

Cited by 0SourceScholar
2026

Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework

ICLR 2026poster

Conventional preference learning methods often prioritize opinions held more widely when aggregating preferences from multiple evaluators. This may result in policies that are biased in favor of some types of opinions or groups and susceptible to strategic manipulation. To address this issue, we de…

Cited by 0SourceScholar
2026

Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning

CVPR 2026

Human video generation has advanced rapidly with the development of diffusion models, but the high computational cost and substantial memory consumption associated with training these models on high-resolution, multi-frame data pose significant challenges. In this paper, we propose Entropy-Guided Pr

Cited by 0SourcecodeScholar
2026

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

CVPR 2026

With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still require task-specific fine-tuning and large-scale audiovisual datasets, resulting in high computational costs that hinder accessibility for resource-c

Cited by 0SourcecodeScholar
2026

Mosaic: Unlocking Over 30$\times$ Context Length for Diffusion LLMs Inference via Global Memory Planning and Dynamic Peak Taming

ICML 2026poster

Diffusion-based large language models (dLLMs) have emerged as a promising alternative to autoregressive models, leveraging simultaneous denoising to enable global planning and iterative refinement. These properties make dLLMs particularly attractive for long-context generation. However, deploying dL…

Cited by 0SourceScholar
2026

Reducing Contextual Stochastic Bilevel Optimization via Structured Function Approximation

ICLR 2026poster

Contextual Stochastic Bilevel Optimization (CSBO) extends standard stochastic bilevel optimization (SBO) by incorporating context-dependent lower-level problems. CSBO problems are generally intractable since existing methods require solving a distinct lower-level problem for each sampled context, re…

Cited by 0SourceScholar
2026

Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction

ICML 2026poster

Surface reconstruction with differentiable rendering has achieved impressive performance in recent years, yet the pervasive photometric ambiguities have strictly bottlenecked existing approaches. This paper presents AmbiSuR, a framework that explores an intrinsic solution upon Gaussian Splatting for…

Cited by 0SourceScholar
2026

SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentation

CVPR 2026

In medical image segmentation tasks, the domain gap caused by the difference in data collection between training and testing data seriously hinders the deployment of pre-trained models in clinical practice. Continual Test-Time Adaptation (CTTA) aims to enable pre-trained models to adapt to continuou

Cited by 0SourcecodeScholar
2026

SparseSurf: Sparse-View 3D Gaussian Splatting for Surface Reconstruction

AAAI 2026technical

Recent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, such optimization is prone to overfitting, leading to suboptimal reconstruction quality. Existing approaches address this

Cited by 0SourcePDFScholar
2026

Stabilized Supralinear Networks Learn to Switch Coding Strategies Balancing Cost and Performance

ICML 2026poster

Lateral connections (LCs) are ubiquitous in the cortical circuits. While modern deep learning architectures have rich intralayer interactions (e.g., convolutional mixing, normalization, or attention) to support feature selectivity and contextual modulation, explicit excitatory and inhibitory (E-I) L…

Cited by 0SourceScholar
2026

Stage-wise Distortion–Perception Traversal in Zero-shot Inverse Problems with Diffusion Models

ICML 2026poster

The distortion–perception (D–P) tradeoff is a fundamental phenomenon of Bayesian inverse problems, which characterizes the inherent tension between distortion performance and perceptual quality. Enabling flexible traversal of the D-P tradeoff at inference time is crucial for practical applications. …

Cited by 0SourceScholar
2026

TextAtlas5M: A Large-Scale Dataset for Long Text Image Generation

ICML 2026poster

Text-conditioned image generation has made rapid progress, yet rendering images with long-form text remains challenging due to the limitations of existing datasets, which predominantly focus on short and simple text. We introduce TextAtlas5M, a large-scale dataset designed to evaluate long-text rend…

Cited by 0SourceScholar
2026

Virtual-Force Based Visual Servo for Multiple Peg-In-Hole Assembly with Tightly Coupled Multi-Manipulator

ICRA 2026poster

Multiple Peg-in-Hole (MPiH) assembly is one of the fundamental tasks in robotic assembly. In the MPiH tasks for large-size parts, it is challenging for a single manipulator to simultaneously align multiple distant pegs and holes, necessitating tightly coupled multi-manipulator systems. For such MPiH…

2026

Virtual-Force Based Visual Servo for Multiple Peg-in-Hole Assembly With Tightly Coupled Multi-Manipulator

RA-L 2026

Multiple Peg-in-Hole (MPiH) assembly is one of the fundamental tasks in robotic assembly. In the MPiH tasks for large-size parts, it is challenging for a single manipulator to simultaneously align multiple distant pegs and holes, necessitating tightly coupled multi-manipulator systems. For such MPiH

Cited by 6SourceScholar
2026

scChord: A Probabilistic Manifold Rectification Framework for RNA-to-Protein Translation

ICML 2026poster

Measuring single-cell protein abundance is essential for resolving biological mechanisms and disease progression with high resolution. However, due to the high costs and antibody throughput limitations of current proteomics, inferring protein levels from readily available RNA data has become a criti…

Cited by 0SourceScholar
2025

A Diffusion-Based Framework for Occluded Object Movement

AAAI 2025technical

Seamlessly moving objects within a scene is a common requirement for image editing, but it is still a challenge for existing editing methods. Especially for real-world images, the occlusion situation further increases the difficulty. The main difficulty is that the occluded portion needs to be compl…

Cited by 0SourcePDFScholar
2025

A Single-Loop First-Order Algorithm for Linearly Constrained Bilevel Optimization

NeurIPS 2025poster

We study bilevel optimization problems where the lower-level problems are strongly convex and have coupled linear constraints. To overcome the potential non-smoothness of the hyper-objective and the computational challenges associated with the Hessian matrix, we utilize penalty and augmented Lagrang…

Cited by 0SourcecodeScholar
2025

AdvAgent: Controllable Blackbox Red-teaming on Web Agents

ICML 2025poster

Foundation model-based agents are increasingly used to automate complex tasks, enhancing efficiency and productivity. However, their access to sensitive resources and autonomous decision-making also introduce significant security risks, where successful attacks could lead to severe consequences. To…

2025

Chemistry3D: Robotic Interaction Toolkit for Chemistry Experiments

ICRA 2025

The advent of simulation engines has revolutionized learning and operational efficiency for robots, offering cost-effective and swift pipelines. However, the lack of a universal simulation platform tailored for chemical scenarios impedes progress in robotic manipulation and visualization of reaction

Cited by 2SourcecodeScholar
2025

CoTD-PO: Chain-of-Thought Distillation with Preference Optimization

EMNLP 2025

Chain-of-Thought (CoT) distillation has emerged as a promising paradigm to enhance the reasoning ability of small language models by imitating the reasoning and outputs of larger teacher models. However, existing approaches suffer from a critical limitation: a distribution mismatch between teacher-g

2025

Contextual Optimization Under Model Misspecification: A Tractable and Generalizable Approach

ICML 2025poster

Contextual optimization problems are prevalent in decision-making applications where historical data and contextual features are used to learn predictive models that inform optimal actions. However, practical applications often suffer from model misspecification due to incomplete knowledge of the un…

Cited by 0SourcePDFScholar
2025

DeblurDiff: Real-Word Image Deblurring with Generative Diffusion Models

NeurIPS 2025poster

Diffusion models have achieved significant progress in image generation and the pre-trained Stable Diffusion (SD) models are helpful for image deblurring by providing clear image priors. However, directly using a blurry image or a pre-deblurred one as a conditional control for SD will either hinder…

Cited by 0SourceScholar
2025

DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution

ICCV 2025poster

Large-scale pre-trained diffusion models are becoming increasingly popular in solving the Real-World Image Super-Resolution (Real-ISR) problem because of their rich generative priors. The recent development of diffusion transformer (DiT) has witnessed overwhelming performance over the traditional UN…

Cited by 0SourcePDFScholar
2025

DiffRetouch: Using Diffusion to Retouch on the Shoulder of Experts

AAAI 2025technical

Image retouching aims to enhance the visual quality of photos. Considering the different aesthetic preferences of users, the target of retouching is subjective. However, current retouching methods mostly adopt deterministic models, which not only neglects the style diversity in the expert-retouched…

Cited by 0SourcePDFScholar
2025

EIA: ENVIRONMENTAL INJECTION ATTACK ON GENERALIST WEB AGENTS FOR PRIVACY LEAKAGE

ICLR 2025poster

Recently, generalist web agents have demonstrated remarkable potential in autonomously completing a wide range of tasks on real websites, significantly boosting human productivity. However, web tasks, such as booking flights, usually involve users' personally identifiable information (PII), which ma…

2025

Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

NeurIPS 2025poster

Balancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training separate preference models for helpfulness and safety, while framing safety as a constraint within a constrained Markov D…

Cited by 0SourcecodeScholar
2025

Eve3D: Elevating Vision Models for Enhanced 3D Surface Reconstruction via Gaussian Splatting

NeurIPS 2025poster

We present Eve3D, a novel framework for dense 3D reconstruction based on 3D Gaussian Splatting (3DGS). While most existing methods rely on imperfect priors derived from pre-trained vision models, Eve3D fully leverages these priors by jointly optimizing both them and the 3DGS backbone. This joint opt…

Cited by 0SourceScholar
2025

Event-guided HDR Reconstruction with Diffusion Priors

ICCV 2025poster

Events provide High Dynamic Range (HDR) intensity change which can guide Low Dynamic Range (LDR) image for HDR reconstruction. However, events only provide temporal intensity differences and it is still ill-posed in over-/under-exposed areas due to missing initial reference brightness and color info…

2025

FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video

CVPR 2025poster

Reconstructing high-fidelity, animatable 3D head avatars from effortlessly captured monocular videos is a pivotal yet formidable challenge. Although significant progress has been made in rendering performance and manipulation capabilities, notable challenges remain, including incomplete reconstructi…

2025

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles

AAAI 2025technical

Humans can perceive speakers’ characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech ( TTS ) scholars grounded their investigations on real-person faces, thereby restricting effec…

2025

GSV3D: Gaussian Splatting-based Geometric Distillation with Stable Video Diffusion for Single-Image 3D Object Generation

ICCV 2025poster

Image-based 3D generation has vast applications in robotics and gaming, where high-quality, diverse outputs and consistent 3D representations are crucial. However, existing methods have limitations: 3D diffusion models are limited by dataset scarcity and the absence of strong pre-trained priors, whi…

2025

GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction

NeurIPS 2025spotlight

Reconstructing accurate surfaces with radiance fields has achieved remarkable progress in recent years. However, prevailing approaches, primarily based on Gaussian Splatting, are increasingly constrained by representational bottlenecks. In this paper, we introduce GeoSVR, an explicit voxel-based fra…

Cited by 0SourcecodeScholar
2025

GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning

ICML 2025poster

The rapid advancement of large language model (LLM) agents has raised new concerns regarding their safety and security. In this paper, we propose GuardAgent, the first guardrail agent to protect target agents by dynamically checking whether their actions satisfy given safety guard requests. Specific…

2025

Improving Diffusion-based Inverse Algorithms under Few-Step Constraint via Linear Extrapolation

NeurIPS 2025poster

Diffusion-based inverse algorithms have shown remarkable performance across various inverse problems, yet their reliance on numerous denoising steps incurs high computational costs. While recent developments of fast diffusion ODE solvers offer effective acceleration for diffusion sampling without o…

Cited by 0SourcecodeScholar
2025

InsTaG: Learning Personalized 3D Talking Head from Few-Second Video

CVPR 2025poster

Despite exhibiting impressive performance in synthesizing lifelike personalized 3D talking heads, prevailing methods based on radiance fields suffer from high demands for training data and time for each new identity. This paper introduces InsTaG, a 3D talking head synthesis framework that allows a f…

2025

Local Path Optimization in The Latent Space Using Learned Distance Gradient

IROS 2025

Constrained motion planning is a common but challenging problem in robotic manipulation. In recent years, data-driven constrained motion planning algorithms have shown impressive planning speed and success rate. Among them, the latent motion method based on manifold approximation is the most efficie

Cited by 0SourceScholar
2025

MLEP: Multi-granularity Local Entropy Patterns for Generalized AI-generated Image Detection

NeurIPS 2025poster

Advances in image generation technologies have raised growing concerns about their potential misuse, particularly in producing misinformation and deepfakes. This creates an urgent demand for effective methods to detect AI-generated images (AIGIs). While progress has been made, achieving reliable per…

Cited by 0SourceScholar
2025

MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models

ICLR 2025poster

Multimodal foundation models (MMFMs) play a crucial role in various applications, including autonomous driving, healthcare, and virtual assistants. However, several studies have revealed vulnerabilities in these models, such as generating unsafe content by text-to-image models. Existing benchmarks o…

2025

Pixel2Feature Attack (P2FA): Rethinking the Perturbed Space to Enhance Adversarial Transferability

ICML 2025poster

Adversarial examples have been shown to deceive Deep Neural Networks (DNNs), raising widespread concerns about this security threat. More seriously, as different DNN models share critical features, feature-level attacks can generate transferable adversarial examples, thereby deceiving black-box mode…

Cited by 0SourcePDFScholar
2025

PolyGuard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset

NeurIPS 2025poster

As large language models (LLMs) become widespread across diverse applications, concerns about the security and safety of LLM interactions have intensified. Numerous guardrail models and benchmarks have been developed to ensure LLM content safety. However, existing guardrail benchmarks are often buil…

Cited by 0SourceScholar
2025

Rethinking Time Encoding via Learnable Transformation Functions

ICML 2025poster

Effectively modeling time information and incorporating it into applications or models involving chronologically occurring events is crucial. Real-world scenarios often involve diverse and complex time patterns, which pose significant challenges for time encoding methods. While previous methods focu…

2025

Retinex-Based Self-Conditioned Diffusion Model for Low-Light Image Enhancement

ICASSP 2025accepted

The conditional diffusion models have made significant progress in image synthesis, leveraging human annotations such as class labels or text descriptions to guide the generative process. However, different from image synthesis, low-light image enhancement(LLIE) lacks strictly calibrated conditional…

Cited by 0SourceScholar
2025

SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models

ICML 2025poster

Traditional autonomous driving systems often struggle to connect high-level reasoning with low-level control, leading to suboptimal and sometimes unsafe behaviors. Recent advances in multimodal large language models (MLLMs), which process both visual and textual data, offer an opportunity to unify p…

2025

Stochastic Smoothed Primal-Dual Algorithms for Nonconvex Optimization with Linear Inequality Constraints

ICML 2025spotlight

We propose smoothed primal-dual algorithms for solving stochastic nonconvex optimization problems with linear \emph{inequality} constraints. Our algorithms are single-loop and only require a single (or two) samples of stochastic gradients at each iteration. A defining feature of our algorithm is tha…

Cited by 0SourcePDFScholar
2025

TeRA: Rethinking Text-guided Realistic 3D Avatar Generation

ICCV 2025poster

Efficient 3D avatar creation is a significant demand in the metaverse, film/game, AR/VR, etc. In this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach…

Cited by 0SourcePDFScholar
2025

The Dual Nature of Plasticity Loss in Deep Continual Learning: Dissection and Mitigation

NeurIPS 2025poster

Loss of plasticity (LoP) is the primary cause of cognitive decline in normal aging brains next to cell loss. Recent works show that similar LoP also plagues neural networks during deep continual learning (DCL). While it has been shown that random perturbations of learned weights can alleviate LoP,…

Cited by 0SourceScholar
2025

UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning

ICML 2025poster

Large Language Model (LLM) agents equipped with external tools have become increasingly powerful for complex tasks such as web shopping, automated email replies, and financial trading. However, these advancements amplify the risks of adversarial attacks, especially when agents can access sensitive e…

2025

Unifying Text Semantics and Graph Structures for Temporal Text-attributed Graphs with Large Language Models

NeurIPS 2025poster

Temporal graph neural networks (TGNNs) have shown remarkable performance in temporal graph modeling. However, real-world temporal graphs often possess rich textual information, giving rise to temporal text-attributed graphs (TTAGs). Such combination of dynamic text semantics and evolving graph struc…

Cited by 0SourceScholar
2024

A Smoothed Bregman Proximal Gradient Algorithm for Decentralized Nonconvex Optimization

ICASSP 2024accepted

Decentralized computation has received considerable research interest lately, due to its wide applications in information processing systems. However, one key requirement to establish convergence for almost all decentralized algorithms, for convex and non-convex problems alike, is that the loss func…

Cited by 0SourceScholar
2024

A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and Feedback

ICML 2024poster

Inverse Reinforcement Learning (IRL) and Reinforcement Learning from Human Feedback (RLHF) are pivotal methodologies in reward learning, which involve inferring and shaping the underlying reward function of sequential decision-making problems based on observed human demonstrations and feedback. Most…

Cited by 1SourcePDFScholar
2024

An Efficient Transformer For Demosaicing Via Compressed Multi-Branch Attention Mechanism

ICASSP 2024accepted

Recent demosaicing approaches are not effective and efficient enough as they do not make full use of these two factors: (1) Capturing long-range spatial dependencies effiently. (2) Reducing the computational costs when utilizing channel attention. To take them into consideration, we propose an Effic…

Cited by 0SourceScholar
2024

ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous Vehicles

CVPR 2024poster

We present ChatScene a Large Language Model (LLM)-based agent that leverages the capabilities of LLMs to generate safety-critical scenarios for autonomous vehicles. Given unstructured language instructions the agent first generates textually described traffic scenarios using LLMs. These scenario des…

2024

CoR-GS: Sparse-View 3D Gaussian Splatting via Co-Regularization

ECCV 2024poster

"3D Gaussian Splatting (3DGS) creates a radiance field consisting of 3D Gaussians to represent a scene. With sparse training views, 3DGS easily suffers from overfitting, negatively impacting rendering. This paper introduces a new co-regularization perspective for improving sparse-view 3DGS. When tra…

2024

DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization

CVPR 2024poster

Radiance fields have demonstrated impressive performance in synthesizing novel views from sparse input views yet prevailing methods suffer from high training costs and slow inference speed. This paper introduces DNGaussian a depth-regularized framework based on 3D Gaussian radiance fields offering r…

2024

Diffusion-based Blind Text Image Super-Resolution

CVPR 2024poster

Recovering degraded low-resolution text images is challenging especially for Chinese text images with complex strokes and severe degradation in real-world scenarios. Ensuring both text fidelity and style realness is crucial for high-quality text image super-resolution. Recently diffusion models have…

2024

EPA: Neural Collapse Inspired Robust Out-of-distribution Detector

ICASSP 2024accepted

Out-of-distribution (OOD) detection plays a crucial role in ensuring the security of neural networks. Existing works have leveraged the fact that In-distribution (ID) samples form a subspace in the feature space, achieving state-of-the-art (SOTA) performance. However, the comprehensive characteristi…

Cited by 0SourceScholar
2024

Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs

ACL 2024findings

Large language models (LLMs), while exhibiting exceptional performance, suffer from hallucinations, especially on knowledge-intensive tasks. Existing works propose to augment LLMs with individual text units retrieved from external knowledge corpora to alleviate the issue. However, in many domains, t…

2024

Robust Synthetic-to-Real Transfer for Stereo Matching

CVPR 2024poster

With advancements in domain generalized stereo matching networks models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However few studies have investigated the robustness after fine-tuning them in real-world scenarios during which the domain generalization ability ca…

2024

Stochastic Smoothed Gradient Descent Ascent for Federated Minimax Optimization

AISTATS 2024poster

In recent years, federated minimax optimization has attracted growing interest due to its extensive applications in various machine learning tasks. While Smoothed Alternative Gradient Descent Ascent (Smoothed-AGDA) has proved successful in centralized nonconvex minimax optimization, how and whether…

Cited by 2SourcePDFScholar
2024

TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting

ECCV 2024poster

"Radiance fields have demonstrated impressive performance in synthesizing lifelike 3D talking heads. However, due to the difficulty in fitting steep appearance changes, the prevailing paradigm that presents facial motions by directly modifying point appearance may lead to distortions in dynamic regi…

2024

Uniformly Stable Algorithms for Adversarial Training and Beyond

ICML 2024poster

In adversarial machine learning, neural networks suffer from a significant issue known as robust overfitting, where the robust test accuracy decreases over epochs (Rice et al., 2020). Recent research conducted by Xing et al., 2021;Xiao et al., 2022 has focused on studying the uniform stability of ad…

2024

Unleashing the Denoising Capability of Diffusion Prior for Solving Inverse Problems

NeurIPS 2024poster

The recent emergence of diffusion models has significantly advanced the precision of learnable priors, presenting innovative avenues for addressing inverse problems. Previous works have endeavored to integrate diffusion priors into the maximum a posteriori estimation (MAP) framework and design optim…

2024

Unveiling the Magic: Investigating Attention Distillation in Retrieval-Augmented Generation

NAACL 2024short

Retrieval-augmented generation framework addresses the limitations of large language models by enabling real-time knowledge updates for more accurate answers. An efficient way in the training phase of retrieval-augmented models is attention distillation, which uses attention scores as supervision si…

Cited by 3SourcePDFScholar
2023

A New Method Combining Single-Point Push and Double-Point Complete Push for Partially Observable Scenarios

IROS 2023poster

Pushing objects into a target configuration is an important skill for robots. When there are obstacles in the scenario, the movement range of the objects will be limited, and objects are easily lost due to obscured by obstacles, which complicates the pushing task. This paper proposes a push path pla…

Cited by 0SourceScholar
2023

AMC-Net: An Effective Network for Automatic Modulation Classification

ICASSP 2023accepted

Automatic modulation classification (AMC) is a crucial stage in the spectrum management, signal monitoring, and control of wireless communication systems. The accurate classification of the modulation format plays a vital role in the subsequent decoding of the transmitted data. End-to-end deep learn…

Cited by 0SourceScholar
2023

DiffuSum: Generation Enhanced Extractive Summarization with Diffusion

ACL 2023findings

Extractive summarization aims to form a summary by directly extracting sentences from the source document. Existing works mostly formulate it as a sequence labeling problem by making individual sentence label predictions. This paper proposes DiffuSum, a novel paradigm for extractive summarization, b…

2023

Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis

ICCV 2023poster

This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art performance with small model size. Our idea is to explicitly exploit the unequal…

Cited by 84PDFcodeScholar
2023

HumanMAC: Masked Motion Completion for Human Motion Prediction

ICCV 2023poster

Human motion prediction is a classical problem in computer vision and computer graphics, which has a wide range of practical applications. Previous effects achieve great empirical performance based on an encoding-decoding style. The methods of this style work by first encoding previous motions to la…

Cited by 82PDFcodeScholar
2023

Linearly Constrained Bilevel Optimization: A Smoothed Implicit Gradient Approach

ICML 2023poster

This work develops analysis and algorithms for solving a class of bilevel optimization problems where the lower-level (LL) problems have linear constraints. Most of the existing approaches for constrained bilevel problems rely on value function-based approximate reformulations, which suffer from iss…

Cited by 22SourcePDFScholar
2023

Pruning Deep Neural Networks from a Sparsity Perspective

ICLR 2023poster

In recent years, deep network pruning has attracted significant attention in order to enable the rapid deployment of AI into small devices with computation and memory constraints. Pruning is often achieved by dropping redundant weights, neurons, or layers of a deep network while attempting to retain…

2023

Revisiting the Linear-Programming Framework for Offline RL with General Function Approximation

ICML 2023poster

Offline reinforcement learning (RL) aims to find an optimal policy for sequential decision-making using a pre-collected dataset, without further interaction with the environment. Recent theoretical progress has focused on developing sample-efficient offline RL algorithms with various relaxed assumpt…

Cited by 27SourcePDFScholar
2022

A Self-Supervised Mixed-Curvature Graph Neural Network

AAAI 2022technical

Graph representation learning received increasing attentions in recent years. Most of the existing methods ignore the complexity of the graph structures and restrict graphs in a single constant-curvature representation space, which is only suitable to particular kinds of graph structure indeed. Addi…

Cited by 44SourcePDFScholar
2022

Improving Certified Robustness via Statistical Learning with Logical Reasoning

NeurIPS 2022accept

Intensive algorithmic efforts have been made to enable the rapid improvements of certificated robustness for complex ML models recently. However, current robustness certification methods are only able to certify under a limited perturbation radius. Given that existing pure data-driven statistical ap…

2022

RestoreFormer: High-Quality Blind Face Restoration From Undegraded Key-Value Pairs

CVPR 2022poster

Blind face restoration is to recover a high-quality face image from unknown degradations. As face image contains abundant contextual information, we propose a method, RestoreFormer, which explores fully-spatial attentions to model contextual information and surpasses existing works that use local co…

Cited by 123PDFcodeScholar
2022

Revisiting Domain Generalized Stereo Matching Networks From a Feature Consistency Perspective

CVPR 2022poster

Despite recent stereo matching networks achieving impressive performance given sufficient training data, they suffer from domain shifts and generalize poorly to unseen domains. We argue that maintaining feature consistency between matching pixels is a vital factor for promoting the generalization ca…

Cited by 79PDFcodeScholar
2022

What is a Good Metric to Study Generalization of Minimax Learners?

NeurIPS 2022accept

Minimax optimization has served as the backbone of many machine learning problems. Although the convergence behavior of optimization algorithms has been extensively studied in minimax settings, their generalization guarantees, i.e., how the model trained on empirical data performs on the unseen test…

Cited by 16SourcePDFScholar
2021

Communication Efficient Primal-Dual Algorithm for Nonconvex Nonsmooth Distributed Optimization

AISTATS 2021poster

Decentralized optimization problems frequently appear in the large scale machine learning problems. However, few works work on the difficult nonconvex nonsmooth case. In this paper, we propose a decentralized primal-dual algorithm to solve this type of problem in a decentralized manner and the propo…

Cited by 19SourcePDFScholar
2021

Efficient Deep Image Denoising via Class Specific Convolution

AAAI 2021technical

Deep neural networks have been widely used in image denoising during the past few years. Even though they achieve great success on this problem, they are computationally inefficient which makes them inappropriate to be implemented in mobile devices. In this paper, we propose an efficient deep neural…

2021

Hyperbolic Variational Graph Neural Network for Modeling Dynamic Graphs

AAAI 2021technical

Learning representations for graphs plays a critical role in a wide spectrum of downstream applications. In this paper, we summarize the limitations of the prior works in three folds: representation space, modeling dynamics and modeling uncertainty. To bridge this gap, we propose to learn dynamic gr…

Cited by 81SourcePDFScholar
2021

Learning RAW-to-sRGB Mappings With Inaccurately Aligned Supervision

ICCV 2021poster

Learning RAW-to-sRGB mapping has drawn increasing attention in recent years, wherein an input raw image is trained to imitate the target sRGB image captured by another camera. However, the severe color inconsistency makes it very challenging to generate well-aligned training pairs of input raw and t…

Cited by 52PDFcodeScholar
2021

Learning To Know Where To See: A Visibility-Aware Approach for Occluded Person Re-Identification

ICCV 2021poster

Person re-identification (ReID) has gained an impressive progress in recent years. However, the occlusion is still a common and challenging problem for recent ReID methods. Several mainstream methods utilize extra cues (e.g., human pose information) to distinguish human parts from obstacles to allev…

Cited by 86PDFScholar
2021

Learning a Non-Blind Deblurring Network for Night Blurry Images

CVPR 2021poster

Deblurring night blurry images is difficult, because the common-used blur model based on the linear convolution operation does not hold in this situation due to the influence of saturated pixels. In this paper, we propose a non-blind deblurring network (NBDN) to restore night blurry images. To mitig…

Cited by 36PDFScholar
2021

Progressive-Scale Boundary Blackbox Attack via Projective Gradient Estimation

ICML 2021spotlight

Boundary based blackbox attack has been recognized as practical and effective, given that an attacker only needs to access the final model prediction. However, the query efficiency of it is in general high especially for high dimensional image data. In this paper, we show that such efficiency highly…

2021

When Expressivity Meets Trainability: Fewer than $n$ Neurons Can Work

NeurIPS 2021poster

Modern neural networks are often quite wide, causing large memory and computation costs. It is thus of great interest to train a narrower network. However, training narrow neural nets remains a challenging task. We ask two theoretical questions: Can narrow networks have as strong expressivity as wid…

Cited by 14SourcePDFScholar
2020

A Proximal Dual Consensus Method for Linearly Coupled Multi-Agent Non-Convex Optimization

ICASSP 2020accepted

Motivated by large-scale signal processing and machine learning applications, this paper considers the distributed multi-agent optimization problem for a linearly constrained non-convex problem. Each of the agents owns a local cost function and local variable, but are coupled with each other due to…

Cited by 0SourceScholar
2020

A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems

NeurIPS 2020poster

Nonconvex-concave min-max problem arises in many machine learning applications including minimizing a pointwise maximum of a set of nonconvex functions and robust adversarial training of neural networks. A popular approach to solve this problem is the gradient descent-ascent (GDA) algorithm which un…

Cited by 128SourcePDFScholar
2020

Attention-based Multi-level Feature Fusion for Named Entity Recognition

IJCAI 2020poster

Named entity recognition (NER) is a fundamental task in the natural language processing (NLP) area. Recently, representation learning methods (e.g., character embedding and word embedding) have achieved promising recognition results. However, existing models only consider partial features derived fr…

Cited by 0SourcePDFScholar
2020

BANANA: when Behavior ANAlysis meets social Network Alignment

IJCAI 2020poster

Recently, aligning users among different social networks has received significant attention. However, most of the existing studies do not consider users’ behavior information during the aligning procedure and thus still suffer from the poor learning performance. In fact, we observe that social netwo…

Cited by 0SourcePDFScholar
2020

Cross-Scale Internal Graph Neural Network for Image Super-Resolution

NeurIPS 2020poster

Non-local self-similarity in natural images has been well studied as an effective prior in image restoration. However, for single image super-resolution (SISR), most existing deep non-local methods (e.g., non-local neural networks) only exploit similar patches within the same scale of the low-resolu…

2020

EfficientFCN: Holistically-guided Decoding for Semantic Segmentation

ECCV 2020poster

Both performance and efficiency are important to semantic segmentation. State-of-the-art semantic segmentation algorithms are mostly based on dilated Fully Convolutional Networks (dilatedFCN), which adopt dilated convolutions in the backbone networks to extract high-resolution feature maps for achie…

Cited by 74SourcePDFScholar
2020

Learning Event-Driven Video Deblurring and Interpolation

ECCV 2020poster

Event-based sensors, which have a response if the change of pixel intensity exceeds a triggering threshold, can capture high-speed motion with microsecond accuracy. Assisted by an event camera, we can generate high frame-rate sharp videos from low frame-rate blurry ones captured by an intensity came…

Cited by 154SourcePDFScholar
2020

Learning a Reinforced Agent for Flexible Exposure Bracketing Selection

CVPR 2020poster

Automatically selecting exposure bracketing (images exposed differently) is important to obtain a high dynamic range image by using multi-exposure fusion. Unlike previous methods that have many restrictions such as requiring camera response function, sensor noise model, and a stream of preview image…

Cited by 24PDFcodeScholar
2020

OID: Outlier Identifying and Discarding in Blind Image Deblurring

ECCV 2020poster

Blind deblurring methods are sensitive to outliers, such as saturated pixels and non-Gaussian noise. Even a small amount of outliers can dramatically degrade the quality of the estimated blur kernel, because the outliers are not conforming to the linear formation of the blurring process. Prior arts…

Cited by 33SourcePDFScholar
2019

DAVANet: Stereo Deblurring With View Aggregation

CVPR 2019oral

Nowadays stereo cameras are more commonly adopted in emerging devices such as dual-lens smartphones and unmanned aerial vehicles. However, they also suffer from blurry images in dynamic scenes which leads to visual discomfort and hampers further image processing. Previous works have succeeded in mon…

Cited by 119PDFScholar
2019

Scalable Gaussian Process Using Inexact Admm for Big Data

ICASSP 2019accepted

Gaussian process (GP) for machine learning has been well studied over the past two decades and is now widely used in many sectors. However, the design of low-complexity GP models still remains a challenging research problem. In this paper, we propose a novel scalable GP regression model for processi…

Cited by 0SourceScholar
2019

Spatio-Temporal Filter Adaptive Network for Video Deblurring

ICCV 2019poster

Video deblurring is a challenging task due to the spatially variant blur caused by camera shake, object motions, and depth variations, etc. Existing methods usually estimate optical flow in the blurry video to align consecutive frames or approximate blur kernels. However, they tend to generate artif…

Cited by 247PDFScholar
2018

Deep Non-Blind Deconvolution via Generalized Low-Rank Approximation

NeurIPS 2018poster

In this paper, we present a deep convolutional neural network to capture the inherent properties of image degradation, which can handle different kernels and saturated pixels in a unified framework. The proposed neural network is motivated by the low-rank property of pseudo-inverse kernels. We firs…

Cited by 99SourcePDFScholar
2018

Dynamic Scene Deblurring Using Spatially Variant Recurrent Neural Networks

CVPR 2018poster

Due to the spatially variant blur caused by camera shake and object motions under different scene depths, deblurring images captured from dynamic scenes is challenging. Although recent works based on deep neural networks have shown great progress on this problem, their models are usually large and c…

Cited by 466SourcePDFScholar
2018

Gated Fusion Network for Single Image Dehazing

CVPR 2018poster

In this paper, we propose an efficient algorithm to directly restore a clear image from a hazy input. The proposed algorithm hinges on an end-to-end trainable neural network that consists of an encoder and a decoder. The encoder is exploited to capture the context of the derived input images, while…

Cited by 1032SourcePDFScholar
2018

Grassmann Pooling as Compact Homogeneous Bilinear Pooling for Fine-Grained Visual Classification

ECCV 2018poster

Designing discriminative and invariant features is the key to visual recognition. Recently, the bilinear pooled feature matrix of Convolutional Neural Network (CNN) has shown to achieve state-of-the-art performance on a range of fine-grained visual recognition tasks. The bilinear feature matrix coll…

Cited by 121SourcePDFScholar
2018

Learning Dual Convolutional Neural Networks for Low-Level Vision

CVPR 2018poster

In this paper, we propose a general dual convolutional neural network (DualCNN) for low-level vision problems, e.g., super-resolution, edge-preserving filtering, deraining and dehazing. These problems usually involve the estimation of two components of the target signals: structures and details. Mot…

Cited by 230SourcePDFScholar
2017

CREST: Convolutional Residual Learning for Visual Tracking

ICCV 2017poster

Discriminative correlation filters (DCFs) have \ryn been shown to perform superiorly in visual tracking. They \ryn only need a small set of training samples from the initial frame to generate an appearance model. However, existing DCFs learn the filters separately from feature extraction, and upda…

Cited by 652PDFScholar
2017

Learning Fully Convolutional Networks for Iterative Non-Blind Deconvolution

CVPR 2017poster

In this paper, we propose a fully convolutional network for iterative non-blind deconvolution. We decompose the non-blind deconvolution problem into image denoising and image deconvolution. We train a FCNN to remove noise in the gradient domain and use the learned gradients to guide the image deconv…

Cited by 215PDFScholar