← Search

Junping Zhang

32 accepted papers

2026

GRPO-based Cluster Decision Agent for Unknown-$\boldsymbol{K}$ Multi-view Clustering

ICML 2026poster

Existing contrastive multi-view clustering methods rely on a pre-defined cluster number, limiting their flexibility in real-world scenarios lacking prior knowledge. To address this, we propose GROK, a novel framework driven by a cluster decision agent for unknown-$K$ multi-view clustering. It pionee…

Cited by 0SourceScholar
2026

MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and Classification

AAAI 2026technical

Nucleus detection and classification (NDC) in histopathology analysis is a fundamental task that underpins a wide range of high-level pathology applications. However, existing methods heavily rely on labor-intensive nucleus-level annotations and struggle to fully exploit large-scale unlabeled data f

Cited by 0SourcePDFScholar
2026

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

ICLR 2026poster

We propose Low-Rank Sparse Attention (Lorsa), a sparse replacement model of Transformer attention layers to disentangle original Multi Head Self Attention (MHSA) into individually comprehensible components. Lorsa is designed to address the challenge of \textit{attention superposition} to understand…

Cited by 0SourcecodeScholar
2026

UniAPO: Unified Multimodal Automated Prompt Optimization

AAAI 2026technical

Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, demonstrating effectiveness primarily in text-only input scenarios. However, extending existing APO methods to multimodal t

Cited by 0SourcePDFScholar
2025

A2DO: Adaptive Anti-Degradation Odometry with Deep Multi-Sensor Fusion for Autonomous Navigation

ICRA 2025

Accurate localization is essential for the safe and effective navigation of autonomous vehicles, and Simultaneous Localization and Mapping (SLAM) is a cornerstone technology in this context. However, The performance of the SLAM system can deteriorate under challenging conditions such as low light, a

Cited by 1SourceScholar
2025

InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention

NeurIPS 2025poster

Diffusion models have demonstrated remarkable capabilities in generating high-quality images. Recent advancements in Layout-to-Image (L2I) generation have leveraged positional conditions and textual descriptions to facilitate precise and controllable image synthesis. Despite overall progress, curren…

Cited by 0SourcecodeScholar
2025

PROTOCOL: Partial Optimal Transport-enhanced Contrastive Learning for Imbalanced Multi-view Clustering

ICML 2025poster

While contrastive multi-view clustering has achieved remarkable success, it implicitly assumes balanced class distribution. However, real-world multi-view data primarily exhibits class imbalance distribution. Consequently, existing methods suffer performance degradation due to their inability to pe…

Cited by 0SourcePDFScholar
2024

Denoising Diffusion Path: Attribution Noise Reduction with An Auxiliary Diffusion Model

NeurIPS 2024poster

The explainability of deep neural networks (DNNs) is critical for trust and reliability in AI systems. Path-based attribution methods, such as integrated gradients (IG), aim to explain predictions by accumulating gradients along a path from a baseline to the target image. However, noise accumulated…

Cited by 1SourcePDFScholar
2024

Point Segment and Count: A Generalized Framework for Object Counting

CVPR 2024poster

Class-agnostic object counting aims to count all objects in an image with respect to example boxes or class names a.k.a few-shot and zero-shot counting. In this paper we propose a generalized framework for both few-shot and zero-shot object counting based on detection. Our framework combines the sup…

2024

Semantic Latent Decomposition with Normalizing Flows for Face Editing

ICASSP 2024accepted

Navigating in the latent space of StyleGAN has shown effectiveness for face editing. However, the resulting methods usually encounter challenges in complicated navigation due to the entanglement among different attributes in the latent space. To address this issue, this paper proposes a novel framew…

Cited by 0SourceScholar
2023

Adaptive Nonlinear Latent Transformation for Conditional Face Editing

ICCV 2023poster

Recent works for face editing usually manipulate the latent space of StyleGAN via the linear semantic directions. However, they usually suffer from the entanglement of facial attributes, need to tune the optimal editing strength, and are limited to binary attributes with strong supervision signals.…

Cited by 7PDFcodeScholar
2023

Cross-Head Supervision for Crowd Counting with Noisy Annotations

ICASSP 2023accepted

Noisy annotations such as missing annotations and location shifts often exist in crowd counting datasets due to multi-scale head sizes, high occlusion, etc. These noisy annotations severely affect the model training, especially for density map-based methods. To alleviate the negative impact of noisy…

Cited by 0SourceScholar
2023

DO-FAM: Disentangled Non-Linear Latent Navigation For Facial Attribute Manipulation

ICASSP 2023accepted

Facial attribute manipulation (FAM) aims to edit the semantic attributes of facial images according to the user’s requirements. Unfortunately, the majority of existing FAM methods struggle in meeting at least one of the two requirements: high reconstruction quality and high irrelevance preservation.…

Cited by 0SourceScholar
2023

From Node Interaction To Hop Interaction: New Effective and Scalable Graph Learning Paradigm

CVPR 2023poster

Existing Graph Neural Networks (GNNs) follow the message-passing mechanism that conducts information interaction among nodes iteratively. While considerable progress has been made, such node interaction paradigms still have the following limitation. First, the scalability limitation precludes the br…

2023

Gaitcotr: Improved Spatial-Temporal Representation for Gait Recognition with a Hybrid Convolution-Transformer Framework

ICASSP 2023accepted

This work presents a novel hybrid convolution-transformer framework for gait recognition, termed GaitCoTr. The developed framework captures the appearance and short-term temporal features by convolution and extracts the long-term temporal features by transformer architecture, achieving a comprehensi…

Cited by 0SourceScholar
2023

LICO: Explainable Models with Language-Image COnsistency

NeurIPS 2023poster

Interpreting the decisions of deep learning models has been actively studied since the explosion of deep neural networks. One of the most convincing interpretation approaches is salience-based visual interpretation, such as Grad-CAM, where the generation of attention maps depends merely on categoric…

2023

Learning to Distill Global Representation for Sparse-View CT

ICCV 2023poster

Sparse-view computed tomography (CT)---using a small number of projections for tomographic reconstruction---enables much lower radiation dose to patients and accelerated data acquisition. The reconstructed images, however, suffer from strong artifacts, greatly limiting their diagnostic value. Curren…

Cited by 12PDFcodeScholar
2023

Motion Matters: A Novel Motion Modeling for Cross-View Gait Feature Learning

ICASSP 2023accepted

As a unique biometric that can be perceived at a distance, gait has broad applications in person authentication, social security and so on. Existing gait recognition methods suffer from changes in viewpoint and clothing and barely consider extracting diverse motion features, a fundamental characteri…

Cited by 0SourceScholar
2023

Mutual Information Based Reweighting for Precipitation Nowcasting

ICASSP 2023accepted

Precipitation nowcasting uses previous rainfall observations to forecast future rainfall intensities in a local area. In rainfall data, the rain-less samples usually well exceed the heavy rainfall samples, and it causes the data imbalance problem in precipitation nowcasting tasks. In this paper, we…

Cited by 2SourceScholar
2023

Online Prototype Learning for Online Continual Learning

ICCV 2023poster

Online continual learning (CL) studies the problem of learning continuously from a single-pass data stream while adapting to new data and mitigating catastrophic forgetting. Recently, by storing a small subset of old data, replay-based methods have shown promising performance. Unlike previous method…

Cited by 63PDFcodeScholar
2022

VR-FAM: Variance-Reduced Encoder with Nonlinear Transformation for Facial Attribute Manipulation

ICASSP 2022accepted

Facial attribute manipulation (FAM) aims to infer desired facial images by modifying specific attributes while keeping others unchanged. Existing works suffer from the entanglement of facial attributes, leading to unexpected artifacts and the loss of facial identity information after editing. To all…

Cited by 0SourceScholar
2021

AgeFlow: Conditional Age Progression and Regression with Normalizing Flows

IJCAI 2021poster

Age progression and regression aim to synthesize photorealistic appearance of a given face image with aging and rejuvenation effects, respectively. Existing generative adversarial networks (GANs) based methods suffer from the following three major issues: 1) unstable training introducing strong ghos…

2021

Private Image Reconstruction from System Side Channels Using Generative Models

ICLR 2021poster

System side channels denote effects imposed on the underlying system and hardware when running a program, such as its accessed CPU cache lines. Side channel analysis (SCA) allows attackers to infer program secrets based on observed side channel signals. Given the ever-growing adoption of machine lea…

2021

Routinggan: Routing Age Progression and Regression with Disentangled Learning

ICASSP 2021accepted

Although impressive results have been achieved for age progression and regression, there remain two major issues in generative adversarial networks (GANs)-based methods: 1) conditional GANs (cGANs)-based methods can learn various effects between any two age groups in a single model, but are insuffic…

Cited by 0SourceScholar
2021

Selfgait: A Spatiotemporal Representation Learning Method for Self-Supervised Gait Recognition

ICASSP 2021accepted

Gait recognition plays a vital role in human identification since gait is a unique biometric feature that can be perceived at a distance. Although existing gait recognition methods can learn gait features from gait sequences in different ways, the performance of gait recognition suffers from insuffi…

Cited by 0SourceScholar
2021

When Age-Invariant Face Recognition Meets Face Age Synthesis: A Multi-Task Learning Framework

CVPR 2021poster

To minimize the effects of age variation in face recognition, previous work either extracts identity-related discriminative features by minimizing the correlation between identity- and age-related features, called age-invariant face recognition (AIFR), or removes age variation by transforming the fa…

Cited by 154PDFcodeScholar
2020

Look Globally, Age Locally: Face Aging With an Attention Mechanism

ICASSP 2020accepted

Face aging is of great importance for cross-age recognition and entertainment-related applications. Recently, conditional generative adversarial networks (cGANs) have achieved impressive results for face aging. Existing cGANs-based methods usually require a pixel-wise loss to keep the identity and b…

Cited by 0SourceScholar