← Search

Quan Zhang

33 accepted papers

2026

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs

CVPR 2026

Although reinforcement learning (RL) has significantly advanced reasoning capabilities in large multimodal language models (MLLMs), its efficacy remains limited for lightweight models essential for edge deployments. To address this issue, we leverage causal analysis and experiment to reveal the unde

Cited by 0SourcecodeScholar
2026

Deploying Atmospheric and Oceanic AI Models on Chinese Hardware and Framework: Migration Strategies, Performance Optimization and Analysis

AAAI 2026technical

With the growing role of artificial intelligence in climate and weather research, efficient model training and inference are in high demand. Current models like FourCastNet and AI-GOMS depend heavily on GPUs, limiting hardware independence, especially for Chinese domestic hardware and frameworks. To

Cited by 0SourcePDFScholar
2026

Frequency-Aware Affinity for Weakly Supervised Semantic Segmentation

CVPR 2026

Weakly Supervised Semantic Segmentation (WSSS) typically utilizes Class Activation Maps (CAMs) to provide the pixel-wise localization. However, CAMs tend to activate only the most discriminative regions, leading to suboptimal WSSS performance. Although existing CAM refinement methods leverage pair-w

Cited by 0SourceScholar
2026

Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization

ICML 2026poster

Multi-Hop Fact Verification (MHFV) necessitates complex reasoning across disparate evidence, posing significant challenges for Large Language Models (LLMs) which often suffer from hallucinations and fractured logical chains. Existing methods, while improving transparency via Chain-of-Thought (CoT), …

Cited by 0SourceScholar
2026

Integrated Hierarchical Decision-Making in Inverse Kinematic Planning and Control

RSS 2026poster

This work presents a novel and efficient non-linear programming framework that tightly integrates hierarchical decision-making with inverse kinematic planning and control. Decision-making plays a central role in many aspects of robotics, from sparse inverse kinematic control with a minimal number of…

Cited by 0SourceScholar
2026

Leveraging Class Distributions in CLIP for Weakly Supervised Semantic Segmentation

CVPR 2026

Image-level Weakly Supervised Semantic Segmentation (WSSS) typically leverages Class Activation Maps (CAMs) for pixel-wise localization. However, existing CLIP-based methods often yield under-activated CAMs, primarily due to the inaccurate semantic relationships in the affinity-based refinement. In

Cited by 0SourceScholar
2026

Structure-Aware Riemannian Flow Matching for Registration and Fusion of Hyperspectral and Multispectral Images

ICML 2026poster

Precise alignment is a prerequisite for hyperspectral and multispectral image fusion, yet existing methods struggle with complex non-rigid deformations. Existing techniques either suffer from inter-task error accumulation by treating registration and fusion as disjoint processes or neglect the geome…

Cited by 0SourceScholar
2026

Towards Robust Text-Attributed Federated Graph Learning: Multimodal Threats and Defense

AAAI 2026technical

Text-Attributed Graphs (TAGs) are graphs where both nodes and edges are associated with text attributes. To leverage their semantic richness, recent efforts have integrated large language models (LLMs) with graph neural networks, leading to the development of GraphLLMs. However, many real-world data

Cited by 0SourcePDFScholar
2026

View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification

CVPR 2026

Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existing methods typically follow a view-invariant paradigm, aligning shared features across views to achieve robustness. However, view-invariant inherent

Cited by 0SourcecodeScholar
2025

Beyond Invisibility: Learning Robust Visible Watermarks for Stronger Copyright Protection

UAI 2025

As AI advances, copyrighted content faces growing risk of unauthorized use, whether through model training or direct misuse. Building upon invisible adversarial perturbation, recent works developed copyright protections against specific AI techniques such as unauthorized personalization through Drea

2025

Continuous renal calculi tracking for autonomous robotic ureteroscopic lithotripsy

IROS 2025

Renal calculi, while not inherently life-threatening, can induce excruciating pain during acute episodes. The predominant clinical treatment – ureteroscopic lithotripsy (URS) –currently faces challenges including restricted maneuverability, frequent manual adjustments during dynamic calculi movement

Cited by 0SourceScholar
2025

Elastic Representation: Mitigating Spurious Correlations for Group Robustness

AISTATS 2025poster

Deep learning models can suffer from severe performance degradation when relying on spurious correlations between input features and labels, making the models perform well on training data but have poor prediction accuracy for minority groups. This problem arises especially when training data are li…

Cited by 0SourceScholar
2025

FFR: Frequency Feature Rectification for Weakly Supervised Semantic Segmentation

CVPR 2025poster

Image-level Weakly Supervised Semantic Segmentation (WSSS) has garnered significant attention due to its low annotation costs. Current single-stage state-of-the-art WSSS methods mainly rely on V ision T ransformer (ViT) to extract features from input images, generating more complete segmentation r…

2025

Free Lunch of Image-mask Alignment for Anomaly Image Generation and Segmentation

IJCAI 2025

This paper aims at generating anomalous images and their segmentation labels to address the lack of real-world anomaly samples and privacy issues. Departing from conventional approaches that use masks solely to guide the generation of anomaly images, we propose a dual-branch training strategy for th

2025

GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and Localization

AAAI 2025technical

The extraordinary ability of generative models emerges as a new trend in image editing and generating realistic images, posing a serious threat to the trustworthiness of multimedia data and driving the research of image manipulation detection and location (IMDL). However, the lack of a large-scale d…

2025

IMDPrompter: Adapting SAM to Image Manipulation Detection by Cross-View Automated Prompt Learning

ICLR 2025poster

Using extensive training data from SA-1B, the Segment Anything Model (SAM) has demonstrated exceptional generalization and zero-shot capabilities, attracting widespread attention in areas such as medical image segmentation and remote sensing image segmentation. However, its performance in the field…

Cited by 0SourcePDFScholar
2025

Multi-Chamber Origami Actuator via Dual Fabric Layers for Dexterous Motions

RA-L 2025

Multi-chamber soft actuators have demonstrated significant potential in minimally invasive surgery, robotic manipulation, and wearable devices, due to their high flexibility. However, conventional multi-chamber soft actuators are usually fabricated from soft materials and complex molding techniques,

Cited by 1SourceScholar
2025

Rethinking Pseudo-Label Guided Learning for Weakly Supervised Temporal Action Localization from the Perspective of Noise Correction

AAAI 2025technical

Pseudo-label learning methods have been widely applied in weakly-supervised temporal action localization. Existing works directly utilize weakly-supervised base model to generate instance-level pseudo-labels for training the fully-supervised detection head. We argue that the noise in pseudo-labels w…

Cited by 1SourcePDFScholar
2025

SMTPD: A New Benchmark for Temporal Prediction of Social Media Popularity

CVPR 2025poster

Social media popularity prediction task aims to predict the popularity of posts on social media platforms, which has a positive driving effect on application scenarios such as content optimization, digital marketing and online advertising. Though many studies have made significant progress, few of t…

2025

Seeing Beyond Noise: Joint Graph Structure Evaluation and Denoising for Multimodal Recommendation

AAAI 2025technical

Multimodal Recommendation Systems (MRSs) boost traditional user-item interaction-based methods by incorporating multimodal information. However, existing methods ignore the inherent noise brought by (1) noisy semantic priors in multimodal content, and (2) noisy user interactions in history records,…

Cited by 0SourcePDFScholar
2025

Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models

CVPR 2025poster

Recent breakthroughs in Multimodal Large Language Models (MLLMs) have gained significant recognition within the deep learning community, where the fusion of the Video Foundation Models (VFMs) and Large Language Models(LLMs) has proven instrumental in constructing robust video understanding systems,…

Cited by 0SourcePDFScholar
2024

A Facile one-step injection novel composite sensor for robot tactile assistance

IROS 2024poster

Tactile information is the research hotspot of wearable flexible sensors due to its importance and complexity. With the innovation of wearable technology and robotics in healthcare, researchers are increasingly integrating wearable flexible sensors on the front end of robots to reproduce the hand ta…

Cited by 0SourceScholar
2024

Unsupervised Group Re-identification via Adaptive Clustering-Driven Progressive Learning

AAAI 2024technical

Group re-identification (G-ReID) aims to correctly associate groups with the same members captured by different cameras. However, supervised approaches for this task often suffer from the high cost of cross-camera sample labeling. Unsupervised methods based on clustering can avoid sample labeling, b…

Cited by 8SourcePDFScholar
2024

View-decoupled Transformer for Person Re-identification under Aerial-ground Camera Network

CVPR 2024poster

Existing person re-identification methods have achieved remarkable advances in appearance-based identity association across homogeneous cameras such as ground-ground matching. However as a more practical scenario aerial-ground person re-identification (AGPReID) among heterogeneous cameras has receiv…

2022

Designing a QAM Signal Detector for Massive Mimo Systems via PS-ADMM Approach

ICASSP 2022accepted

This paper presents an efficient quadrature amplitude modulation (QAM) signal detector for massive multiple-input multiple-output (MIMO) communication systems via the penalty-sharing alternating direction method of multipliers (PS-ADMM). The content of the paper is summarized as follows: first, we f…

Cited by 0SourceScholar
2022

Modeling 3D Layout for Group Re-Identification

CVPR 2022poster

Group re-identification (GReID) attempts to correctly associate groups with the same members under different cameras. The main challenge is how to resist the membership and layout variations. Existing works attempt to incorporate layout modeling on the basis of appearance features to achieve robust…

Cited by 24PDFcodeScholar
2022

Uncertainty Modeling with Second-Order Transformer for Group Re-identification

AAAI 2022technical

Group re-identification (G-ReID) focuses on associating the group images containing the same persons under different cameras. The key challenge of G-ReID is that all the cases of the intra-group member and layout variations are hard to exhaust. To this end, we propose a novel uncertainty modeling, w…

Cited by 22SourcePDFScholar
2021

A Low-Complexity Admm-Based Massive Mimo Detectors Via Deep Neural Networks

ICASSP 2021accepted

An alternate direction method of multipliers (ADMM)-based detectors can achieve good performance in both small and large-scale multiple-input multiple-output (MIMO) systems. However, due to the difficulty of choosing the optimal penalty parameters, their performance is limited. This paper presents a…

Cited by 0SourceScholar
2018

Nonparametric Bayesian Lomax delegate racing for survival analysis with competing risks

NeurIPS 2018poster

We propose Lomax delegate racing (LDR) to explicitly model the mechanism of survival under competing risks and to interpret how the covariates accelerate or decelerate the time to event. LDR explains non-monotonic covariate effects by racing a potentially infinite number of sub-risks, and consequent…

2017

Modelling and analysis of the passive planar rimless wheel mechanism in universal domain

IROS 2017poster

The planar rimless wheel (PRW) is a classic and simple passive dynamic mechanism to simulate biped walking, different simplified PRW models have respective descriptions and limited applications. This paper focus on constructing the general PRW model, and analyzing the intrinsic and mathematical rela…

Cited by 4SourceScholar