← Search

Zhen Xu

39 accepted papers

2026

CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement Learning

AAAI 2026technical

Compositional reasoning is a critical capability for multimodal models, enabling systematic understanding of complex scenes through structured combinations of objects, attributes, and relations. However, existing research on this ability primarily focuses on vision-language models (VLMs, e.g., CLIP

Cited by 0SourcePDFScholar
2026

Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection

CVPR 2026

Modern industrial quality control heavily relies on automated anomaly detection. While few-shot anomaly detection addresses the challenge of limited labeled data, real-world inspection faces a vast diversity of anomaly types, sizes, and shapes. We identify the primary cause for the anomaly detection

Cited by 0SourceScholar
2026

Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review

ICML 2026poster

The formal integration of Large Language Models (LLMs) and Multimodal LLMs (MLLMs) into scientific peer-review workflows introduces novel and significant risks. Their safety against adversarial manipulation remains critically underexplored, especially given the multimodal nature of scientific papers…

Cited by 0SourceScholar
2026

ST-SAM: Multimodal Scene Text Segmentation with Dense Visual and Sparse Textual Prompts via SAM

AAAI 2026technical

Scene text segmentation is a critical preprocessing step in various text-based applications. Specialist text segmentation methods, often relying on a detect-then-segment paradigm, tend to exhibit reduced robustness and can lead to cascading errors. The introduction of the Segment Anything Model (SAM

Cited by 0SourcePDFScholar
2026

Syntactic Structure-Guided Visual Grounding with Subject-Centric Feature Enhancement and Verification

IJCAI 2026

Visual grounding aims to localize target objects based on natural language descriptions, and the core challenge lies in the cross-modal gap, which is partly caused by the significant differences in semantic structure between language and vision. Existing methods typically rely on holistic sentence-l

Cited by 0Scholar
2025

4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos

NeurIPS 2025spotlight

We propose 4DGT, a 4D Gaussian-based Transformer model for dynamic scene reconstruction, trained entirely on real-world monocular posed videos. Using 4D Gaussian as an inductive bias, 4DGT unifies static and dynamic components, enabling the modeling of complex, time-varying environments with varying…

Cited by 0SourceScholar
2025

Adaptive Layered-Trust Robust Defense Mechanism for Personalized Federated Learning

ICASSP 2025accepted

Personalized Federated Learning (PFL) is confronted with escalating security threats, yet existing defense strategies primarily concentrate on traditional federated learning, lacking robust defense mechanisms tailored for PFL. To fortify the robustness of PFL against stealthy malicious attacks, we p…

Cited by 0SourceScholar
2025

Anchoring-Guidance Fine-Tuning (AnGFT): Elevating Professional Response Quality in Role-Playing Conversational Agents

EMNLP 2025

Large Language Models (LLMs) have demonstrated significant advancements in various fields, notably in Role-Playing Conversational Agents (RPCAs). However, when confronted with role-specific professional inquiries, LLMs-based RPCAs tend to underperform due to their excessive emphasis on the conversat

2025

Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants’ Question-Answering in Asynchronous Learning Environments

EMNLP 2025

Virtual Teaching Assistants (VTAs) can reduce the workload of teaching teams in Asynchronous Learning Environments (ALEs) where timely, personalized support is often limited. As VTA systems grow more capable, rigorous and pedagogically sound evaluation becomes essential. Existing assessments often r

2025

DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos

NeurIPS 2025poster

We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction has gained significant attention for its ability to rapidly cr…

Cited by 0SourceScholar
2025

Denoising Trajectory Biases for Zero-Shot AI-Generated Image Detection

NeurIPS 2025poster

The rapid advancement of generative models has led to the widespread emergence of highly realistic synthetic images, making the detection of AI-generated content increasingly critical. In particular, diffusion models have recently achieved unprecedented levels of visual fidelity, further raising con…

Cited by 0SourceScholar
2025

Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models

ICCV 2025poster

This paper addresses the challenge of high-fidelity view synthesis of humans with sparse-view videos as input. Previous methods solve the issue of insufficient observation by leveraging 4D diffusion models to generate videos at novel viewpoints. However, the generated videos from these models often…

2025

ERNet: Efficient Non-Rigid Registration Network for Point Sequences

ICCV 2025poster

Registering an object shape to a sequence of point clouds undergoing non-rigid deformation is a long-standing challenge. The key difficulties stem from two factors: (i) the presence of local minima due to the non-convexity of registration objectives, especially under noisy or partial inputs, which h…

Cited by 0SourcePDFScholar
2025

EnvGS: Modeling View-Dependent Appearance with Environment Gaussian

CVPR 2025poster

Reconstructing complex reflections in real-world scenes from 2D images is essential for achieving photorealistic novel view synthesis. Existing methods that utilize environment maps to model reflections from distant lighting often struggle with high-frequency reflection details and fail to account f…

2025

FreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene Reconstruction

CVPR 2025poster

This paper addresses the challenge of reconstructing dynamic 3D scenes with complex motions. Some recent works define 3D Gaussian primitives in the canonical space and use deformation fields to map canonical primitives to observation spaces, achieving real-time dynamic view synthesis. However, these…

Cited by 0SourcePDFScholar
2025

Hierarchy UGP: Hierarchy Unified Gaussian Primitive for Large-Scale Dynamic Scene Reconstruction

ICCV 2025poster

Recent advances in differentiable rendering have significantly improved dynamic street scene reconstruction. However, the complexity of large-scale scenarios and dynamic elements, such as vehicles and pedestrians, remains a substantial challenge. Existing methods often struggle to scale to large sce…

Cited by 0SourcePDFScholar
2025

MT-Fusion: Multi-Task Learning for Degradation-Aware Infrared and Visible Image Fusion

IROS 2025

The effective fusion of infrared and visible images could enhance environment perception during robot rescue mission by combining complementary information from both sensors. However, most existing fusion methods are developed for images captured under normal conditions, which limits their performan

Cited by 0SourceScholar
2025

MixHD: A Method for Detecting Hallucinations Based on the Internal State and Output Probability of Large Language Models

ICASSP 2025accepted

This paper presents a novel hallucination detection method based on the internal states and output probabilities of large language models (LLMs) to address the common issue of hallucinations in model-generated content. We designed a new detection framework that extracts internal features such as hid…

Cited by 0SourceScholar
2025

StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models

CVPR 2025poster

This paper aims to tackle the problem of photorealistic view synthesis from vehicle sensors data. Recent advancements in neural scene representation have achieved notable success in rendering high-quality autonomous driving scenes,but the performance significantly degrades as the viewpoint deviates…

Cited by 7SourcePDFScholar
2025

Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding

CVPR 2025poster

The goal of visual grounding is to establish connections between target objects and textual descriptions. Large Language Models (LLMs) have demonstrated strong comprehension abilities across a variety of visual tasks. To establish precise associations between the text and the corresponding visual re…

Cited by 0SourcePDFScholar
2024

4K4D: Real-Time 4D View Synthesis at 4K Resolution

CVPR 2024poster

This paper targets high-fidelity and real-time view synthesis of dynamic 3D scenes at 4K resolution. Recent methods on dynamic view synthesis have shown impressive rendering quality. However their speed is still limited when rendering high-resolution images. To overcome this problem we propose 4K4D…

2024

A Multi-Scale Bimodal Fusion Network for Robust and Accurate Online Handwriting Recognition

ICASSP 2024accepted

Online handwriting recognition based on sensor trajectory information faces several unresolved challenges: 1) sensor signals lack sufficient global spatial context; 2) different recognition tasks have inconsistent requirements for feature receptive fields. This is due to the inconsistent scales of t…

Cited by 0SourceScholar
2024

Mutual Information Based Noise Scale Optimization for Gradient Leakage Resistant Federated Learning

ICASSP 2024accepted

Federated learning decentralizes the learning process, yet it does not provide adequate privacy protection. Current countermeasures predominantly rely on Local Differential Privacy (LDP) techniques. While larger noise injection offers stronger privacy, it also leads to a degradation in model perform…

Cited by 0SourceScholar
2024

Relightable and Animatable Neural Avatar from Sparse-View Video

CVPR 2024highlight

This paper tackles the problem of creating relightable and animatable neural avatars from sparse-view (or monocular) videos of dynamic humans under unknown illumination. Previous neural human reconstruction methods produce animatable avatars from sparse views using deformed Signed Distance Fields (S…

2024

Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement

EMNLP 2024main

Large language models (LLMs) demonstrate exceptional instruct-following ability to complete various downstream tasks. Although this impressive ability makes LLMs flexible task solvers, their performance in solving tasks also heavily relies on instructions. In this paper, we reveal that LLMs are over…

Cited by 4SourcePDFScholar
2023

Blemish-Aware and Progressive Face Retouching With Limited Paired Data

CVPR 2023poster

Face retouching aims to remove facial blemishes, while at the same time maintaining the textual details of a given input image. The main challenge lies in distinguishing blemishes from the facial characteristics, such as moles. Training an image-to-image translation network with pixel-wise supervisi…

Cited by 5SourcePDFScholar
2023

Learning Neural Volumetric Representations of Dynamic Humans in Minutes

CVPR 2023poster

This paper addresses the challenge of efficiently reconstructing volumetric videos of dynamic humans from sparse multi-view videos. Some recent works represent a dynamic human as a canonical neural radiance field (NeRF) and a motion field, which are learned from input videos through differentiable r…

2023

Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image Manipulation

CVPR 2023poster

Great progress has been made in StyleGAN-based image editing. To associate with preset attributes, most existing approaches focus on supervised learning for semantically meaningful latent space traversal directions, and each manipulation step is typically determined for an individual attribute. To a…

Cited by 3SourcePDFScholar
2022

360-Attack: Distortion-Aware Perturbations From Perspective-Views

CVPR 2022poster

The application of deep neural networks (DNNs) on 360-degree images has achieved remarkable progress in the recent years. However, DNNs have been demonstrated to be vulnerable to well-crafted adversarial examples, which may trigger severe safety problems in the real-world applications based on 360-d…

Cited by 6PDFScholar
2022

Confidence Propagation Cluster: Unleash Full Potential of Object Detectors

CVPR 2022poster

It's been a long history that most object detection methods obtain objects by using the non-maximum suppression (NMS) and its improved versions like Soft-NMS to remove redundant bounding boxes. We challenge those NMS-based methods from three aspects: 1) The bounding box with highest confidence value…

Cited by 16PDFcodeScholar
2022

SP Attack: Single-Perspective Attack for Generating Adversarial Omnidirectional Images

ICASSP 2022accepted

The safety of Deep Neural Networks (DNNs) processing omnidirectional images (ODIs) is an under-researched topic. In this paper, we propose a novel sparse attack, named Single-Perspective (SP) Attack, towards fooling these models by perturbing only one perspective image (PI) rendered from the target…

Cited by 0SourceScholar
2021

Empowering Adaptive Early-Exit Inference with Latency Awareness

AAAI 2021technical

With the capability of trading accuracy for latency on-the-fly, the technique of adaptive early-exit inference has emerged as a promising line of research to accelerate the deep learning inference. However, studies in this line of research commonly use a group of thresholds to control the accuracy-l…

2021

MUFASA: Multimodal Fusion Architecture Search for Electronic Health Records

AAAI 2021technical

One important challenge of applying deep learning to electronic health records (EHR) is the complexity of their multimodal structure. EHR usually contains a mixture of structured (codes) and unstructured (free-text) data with sparse and irregular longitudinal features -- all of which doctors utilize…

Cited by 69SourcePDFScholar
2021

MatchVIE: Exploiting Match Relevancy between Entities for Visual Information Extraction

IJCAI 2021poster

Visual Information Extraction (VIE) task aims to extract key information from multifarious document images (e.g., invoices and purchase receipts). Most previous methods treat the VIE task simply as a sequence labeling problem or classification problem, which requires models to carefully identify eac…

Cited by 33SourcePDFScholar
2020

Flow Contrastive Estimation of Energy-Based Models

CVPR 2020oral

This paper studies a training method to jointly estimate an energy-based model and a flow-based model, in which the two models are iteratively updated based on a shared adversarial value function. This joint training method has the following traits. (1) The update of the energy-based model is based…

Cited by 133PDFScholar
2017

Illumination-robust face recognition with Block-based Local Contrast Patterns

ICASSP 2017accepted

This paper proposes a novel facial image representation Block-based Local Contrast Patterns (BLCP) for illumination-robust face recognition. This method is based on an effective texture descriptor local contrast patterns (LCP). We use the directed and undirected difference masks to calculate three t…

Cited by 0SourceScholar
2016

Using Social Dynamics to Make Individual Predictions: Variational Inference with a Stochastic Kinetic Model

NeurIPS 2016poster

Social dynamics is concerned primarily with interactions among individuals and the resulting group behaviors, modeling the temporal evolution of social systems via the interactions of individuals within these systems. In particular, the availability of large-scale data from social networks and senso…

Cited by 17SourcePDFScholar