← Search

Wenjun Zhang

51 accepted papers

2026

Behavior Cloning-Enhanced Deep Reinforcement Learning for Robot Navigation Via Dynamic Reward and Adaptive Replanning

ICRA 2026poster

Deep reinforcement learning (DRL) is a core technology for mobile robot navigation in diverse environments, yet existing Behavior Cloning (BC)-enhanced DRL methods suffer two critical challenges: fixed imitation constraints suppress autonomous exploration in late training stages despite stabilizing …

Cited by 0Scholar
2026

Content-Aware Mamba for Learned Image Compression

ICLR 2026poster

Recent Learned image compression (LIC) leverages Mamba-style state-space models (SSMs) for global receptive fields with linear complexity. However, the standard Mamba adopts content-agnostic, predefined raster (or multi-directional) scans under strict causality. This rigidity hinders its ability to…

Cited by 0SourcecodeScholar
2026

RATE-DISTORTION OPTIMIZED COMMUNICATION FOR COLLABORATIVE PERCEPTION

ICLR 2026poster

Collaborative perception emphasizes enhancing environmental understanding by enabling multiple agents to share visual information with limited bandwidth resources. While prior work has explored the empirical trade-off between task performance and communication volume, a significant gap remains in th…

Cited by 0SourceScholar
2026

SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors

ICLR 2026poster

Reconstructing 3D scenes from sparse images remains a challenging task due to the difficulty of recovering accurate geometry and texture without optimization. Recent approaches leverage generalizable models to generate 3D scenes using 3D Gaussian Splatting (3DGS) primitive. However, they often fail…

Cited by 0SourcecodeScholar
2025

4DGCPro: Efficient Hierarchical 4D Gaussian Compression for Progressive Volumetric Video Streaming

NeurIPS 2025poster

Achieving seamless viewing of high-fidelity volumetric video, comparable to 2D video experiences, remains an open challenge. Existing volumetric video compression methods either lack the flexibility to adjust quality and bitrate within a single model for efficient streaming across diverse networks a…

Cited by 0SourceScholar
2025

Beyond Point Annotation: A Weakly Supervised Network Guided by Multi-Level Labels Generated from Four-Point Annotation for Thyroid Nodule Segmentation in Ultrasound Image

ICASSP 2025accepted

Weakly supervised methods typically guided the pixel-wise training by comparing the predictions to single-level labels containing diverse segmentation-related information at once, but struggled to represent subtle feature differences between nodule and background regions and confused incorrect infor…

Cited by 0SourceScholar
2025

Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image Compression

AAAI 2025technical

Neural image compression often faces a challenging trade-off among rate, distortion and perception. While most existing methods typically focus on either achieving high pixel-level fidelity or optimizing for perceptual metrics, we propose a novel approach that simultaneously addresses both aspects f…

Cited by 1SourcePDFScholar
2025

H3D-DGS: Exploring Heterogeneous 3D Motion Representation for Deformable 3D Gaussian Splatting

NeurIPS 2025poster

Dynamic scene reconstruction poses a persistent challenge in 3D vision. Deformable 3D Gaussian Splatting has emerged as an effective method for this task, offering real-time rendering and high visual fidelity. This approach decomposes a dynamic scene into a static representation in a canonical space…

Cited by 0SourceScholar
2025

Instance Correlation Graph-based Naive Bayes

ICML 2025spotlight

Due to its simplicity, effectiveness and robustness, naive Bayes (NB) has continued to be one of the top 10 data mining algorithms. To improve its performance, a large number of improved algorithms have been proposed in the last few decades. However, in addition to Gaussian naive Bayes (GNB), there…

2025

InstantSticker: Realistic Decal Blending via Disentangled Object Reconstruction

AAAI 2025technical

We present InstantSticker, a disentangled reconstruction pipeline based on Image-Based Lighting (IBL), which focuses on highly realistic decal blending, simulates stickers attached to the reconstructed surface, and allows for instant editing and real-time rendering. To achieve stereoscopic impressio…

2025

Knowledge Distillation for Learned Image Compression

ICCV 2025poster

Recently, learned image compression (LIC) models have achieved remarkable rate-distortion (RD) performance, yet their high computational complexity severely limits practical deployment. To overcome this challenge, we propose a novel Stage-wise Modular Distillation framework, SMoDi, which efficiently…

Cited by 0SourcePDFScholar
2025

Label Distribution Propagation-based Label Completion for Crowdsourcing

ICML 2025poster

In real-world crowdsourcing scenarios, most workers often annotate a few instances only, which results in a significantly sparse crowdsourced label matrix and subsequently harms the performance of label integration algorithms. Recent work called worker similarity-based label completion (WSLC) has be…

2025

Neural Block Compression: Variable Bitrates Feature Blocks for Texture Representation

AAAI 2025technical

The imperative for compression of material textures emerges from the critical demand for high-quality rendering, which necessitates sophisticated textures that, in turn, require substantial storage and memory resources. Thus, low-bitrate compression is crucial, especially in modern games demanding h…

Cited by 0SourcePDFScholar
2025

OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?

ICLR 2025poster

We introduce OBI-Bench, a holistic benchmark crafted to systematically evaluate large multi-modal models (LMMs) on whole-process oracle bone inscriptions (OBI) processing tasks demanding expert-level domain knowledge and deliberate cognition. OBI-Bench includes 5,523 meticulously collected diverse-s…

Cited by 2SourcePDFScholar
2025

SinGS: Animatable Single-Image Human Gaussian Splats with Kinematic Priors

CVPR 2025poster

Despite significant advances in accurately estimating geometry in contemporary single-image 3D human reconstruction, creating a high-quality, efficient, and animatable 3D avatar remains an open challenge. Two key obstacles persist: incomplete observation and inconsistent 3D priors. To address these…

2025

TLLC: Transfer Learning-based Label Completion for Crowdsourcing

ICML 2025spotlight

Label completion serves as a preprocessing approach to handling the sparse crowdsourced label matrix problem, significantly boosting the effectiveness of the downstream label aggregation. In recent advances, worker modeling has been proved to be a powerful strategy to further improve the performance…

2025

UMFN: Unified Multi-Domain Face Normalization for Joint Cross-domain Prototype Learning and Heterogeneous Face Recognition

CVPR 2025poster

Face normalization aims to enhance the robustness and effectiveness of face recognition systems by mitigating intra-personal variations in expressions, poses, occlusions, illuminations, and domains. Existing methods face limitations in handling multiple variations and adapting to cross-domain scenar…

Cited by 0SourcePDFScholar
2024

GAIA: Rethinking Action Quality Assessment for AI-Generated Videos

NeurIPS 2024spotlight

Assessing action quality is both imperative and challenging due to its significant impact on the quality of AI-generated videos, further complicated by the inherently ambiguous nature of actions within AI-generated video (AIGV). Current action quality assessment (AQA) algorithms predominantly focus…

2024

IWBVT: Instance Weighting-based Bias-Variance Trade-off for Crowdsourcing

NeurIPS 2024poster

In recent years, a large number of algorithms for label integration and noise correction have been proposed to infer the unknown true labels of instances in crowdsourcing. They have made great advances in improving the label quality of crowdsourced datasets. However, due to the presence of intractab…

Cited by 1SourcePDFScholar
2024

Motion Control of Magnetic-Controlled Spiral Microrobots for In-vitro Plaque Removal

RA-L 2024

Non-contact magnetic-controlled microrobots exhibit great potential in vascular intervention medicine due to their efficiency and safety. This paper introduces a novel helical magnetic microrobot designed and manufactured with integrated consideration for structural functionality and controllability

Cited by 6SourceScholar
2023

AudioEar: Single-View Ear Reconstruction for Personalized Spatial Audio

AAAI 2023technical

Spatial audio, which focuses on immersive 3D sound rendering, is widely applied in the acoustic industry. One of the key problems of current spatial audio rendering methods is the lack of personalization based on different anatomies of individuals, which is essential to produce accurate sound source…

2023

Boosting Point Clouds Rendering via Radiance Mapping

AAAI 2023technical

Recent years we have witnessed rapid development in NeRF-based image rendering due to its high quality. However, point clouds rendering is somehow less explored. Compared to NeRF-based rendering which suffers from dense spatial sampling, point clouds rendering is naturally less computation intensive…

2023

Boosting Video Object Segmentation via Space-Time Correspondence Learning

CVPR 2023poster

Current top-leading solutions for video object segmentation (VOS) typically follow a matching-based regime: for each query frame, the segmentation mask is inferred according to its correspondence to previously processed and the first annotated frames. They simply exploit the supervisory signals from…

2023

Frequency-Modulated Point Cloud Rendering With Easy Editing

CVPR 2023highlight

We develop an effective point cloud rendering pipeline for novel view synthesis, which enables high fidelity local detail reconstruction, real-time rendering and user-friendly editing. In the heart of our pipeline is an adaptive frequency modulation module called Adaptive Frequency Net (AFNet), whic…

2023

Learning Shape Primitives via Implicit Convexity Regularization

ICCV 2023poster

Shape primitives decomposition has been an important and long-standing task in 3D shape analysis. Prior arts heavily rely on 3D point clouds or voxel data for shape primitives extraction, which are less practical in real-world scenarios. This paper proposes to learn shape primitives from multi-view…

Cited by 3PDFcodeScholar
2022

Remember Intentions: Retrospective-Memory-Based Trajectory Prediction

CVPR 2022poster

To realize trajectory prediction, most previous methods adopt the parameter-based approach, which encodes all the seen past-future instance pairs into model parameters. However, in this way, the model parameters come from all seen instances, which means a huge amount of irrelevant seen instances mig…

Cited by 151PDFcodeScholar
2022

Representation-Agnostic Shape Fields

ICLR 2022poster

3D shape analysis has been widely explored in the era of deep learning. Numerous models have been developed for various 3D data representation formats, e.g., MeshCNN for meshes, PointNet for point clouds and VoxNet for voxels. In this study, we present Representation-Agnostic Shape Fields (RASF), a…

2021

3D Human Action Representation Learning via Cross-View Consistency Pursuit

CVPR 2021poster

In this work, we propose a Cross-view Contrastive Learning framework for unsupervised 3D skeleton-based action representation (CrosSCLR), by leveraging multi-view complementary supervision signal. CrosSCLR consists of both single-view contrastive learning (SkeletonCLR) and cross-view consistent know…

Cited by 240PDFcodeScholar
2021

Learning Distilled Collaboration Graph for Multi-Agent Perception

NeurIPS 2021poster

To promote better performance-bandwidth trade-off for multi-agent perception, we propose a novel distilled collaboration graph (DiscoGraph) to model trainable, pose-aware, and adaptive collaboration among agents. Our key novelties lie in two aspects. First, we propose a teacher-student framework to…

2021

Progressive Stage-Wise Learning for Unsupervised Feature Representation Enhancement

CVPR 2021poster

Unsupervised learning methods have recently shown their competitiveness against supervised training. Typically, these methods use a single objective to train the entire network. But one distinct advantage of unsupervised over supervised learning is that the former possesses more variety and freedom…

Cited by 6PDFScholar
2021

Skeleton2Mesh: Kinematics Prior Injected Unsupervised Human Mesh Recovery

ICCV 2021poster

In this paper, we decouple unsupervised human mesh recovery into the well-studied problems of unsupervised 3D pose estimation, and human mesh recovery from estimated 3D skeletons, focusing on the latter task. The challenges of the latter task are two folds: (1) pose failure (i.e., pose mismatching -…

Cited by 29PDFcodeScholar
2021

Towards Alleviating the Modeling Ambiguity of Unsupervised Monocular 3D Human Pose Estimation

ICCV 2021poster

In this work, we study the ambiguity problem in the task of unsupervised 3D human pose estimation from 2D counterpart. On one hand, without explicit annotation, the scale of 3D pose is difficult to be accurately captured (scale ambiguity). On the other hand, one 2D pose might correspond to multiple…

Cited by 49PDFScholar
2020

Cross-Domain Detection via Graph-Induced Prototype Alignment

CVPR 2020oral

Applying the knowledge of an object detector trained on a specific domain directly onto a new domain is risky, as the gap between two domains can severely degrade model's performance. Furthermore, since different instances commonly embody distinct modal information in object detection scenario, the…

Cited by 302PDFcodeScholar
2020

Deep Kinematics Analysis for Monocular 3D Human Pose Estimation

CVPR 2020poster

For monocular 3D pose estimation conditioned on 2D detection, noisy/unreliable input is a key obstacle in this task. Simple structure constraints attempting to tackle this problem, e.g., symmetry loss and joint angle limit, could only provide marginal improvements and are commonly treated as auxilia…

Cited by 233PDFScholar
2020

HDMFH: Hypergraph Based Discrete Matrix Factorization Hashing for Multimodal Retrieval

ICASSP 2020accepted

In recent years, hashing based cross-modal retrieval methods have attracted considerable attention for the high retrieval efficiency and low storage cost. However, most of the existing methods neglect the high-order relationship among data samples. In addition, most of them can only deal with two mo…

Cited by 0SourceScholar
2020

Learning to Combine: Knowledge Aggregation for Multi-Source Domain Adaptation

ECCV 2020poster

Transferring knowledges learned from multiple source domains to target domain is a more practical and challenging task than conventional single-source domain adaptation. Furthermore, the increase of modalities brings more difficulty in aligning feature distributions among multiple domains. To mitiga…

2019

Variational Convolutional Neural Network Pruning

CVPR 2019poster

We propose a variational Bayesian scheme for pruning convolutional neural networks in channel level. This idea is motivated by the fact that deterministic value based pruning methods are inherently improper and unstable. In a nutshell, variational technique is introduced to estimate distribution of…

Cited by 455PDFScholar
2018

Learning an Inverse Tone Mapping Network with a Generative Adversarial Regularizer

ICASSP 2018accepted

Transferring a low-dynamic-range (LDR) image to a high-dynamic-range (HDR) image, which is the so-called inverse tone mapping (iTM), is an important imaging technique to improve visual effects of imaging devices. In this paper, we propose a novel deep learning-based iTM method, which learns an inver…

Cited by 0SourceScholar
2018

Multi-Scale Spatially-Asymmetric Recalibration for Image Classification

ECCV 2018poster

Convolution is spatially-symmetric, i.e., the visual features are independent of its position in the image, which limits its ability to use spatial information. This paper addresses this issue by a recalibration process, which refers to the surrounding region of each neuron, computes an importance v…

Cited by 17SourcePDFScholar
2018

Online Multi-Object Tracking with Dual Matching Attention Networks

ECCV 2018poster

In this paper, we propose an online Multi-Object Tracking (MOT) approach which integrates the merits of single object tracking and data association methods in a unified framework to handle noisy detections and frequent interactions between targets. Specifically, for applying single object tracking i…

Cited by 455SourcePDFScholar
2017

Performance Guaranteed Network Acceleration via High-Order Residual Quantization

ICCV 2017poster

Input binarization has shown to be an effective way for network acceleration. However, previous binarization scheme could be regarded as simple pixel-wise thresholding operations (i.e., order-one approximation) and suffers a big accuracy loss. In this paper, we propose a high-order binarization sche…

Cited by 137PDFScholar
2017

SORT: Second-Order Response Transform for Visual Recognition

ICCV 2017poster

In this paper, we reveal the importance and benefits of introducing second-order operations into deep neural networks. We propose a novel approach named Second-Order Response Transform (SORT), which appends element-wise product transform to the linear sum of a two-branch network module. A direct adv…

Cited by 67PDFcodeScholar