← Search

Xudong Jiang

64 accepted papers

2026

ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos

CVPR 2026

Temporal forgery localization aims to temporally identify manipulated segments in videos. Most existing benchmarks focus on appearance-level forgeries, such as face swapping and object removal. However, recent advances in video generation have driven the emergence of activity-level forgeries that mo

Cited by 0SourcecodeScholar
2026

From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge

AAAI 2026technical

Large-scale Video Foundation Models (VFMs) have significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces critical security risks, as adversaries can exploit full knowl

Cited by 0SourcePDFScholar
2026

It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks

ICML 2026poster

Time series foundation models (TSFMs) are revolutionizing the forecasting landscape from specific dataset modeling to generalizable task evaluation. However, we contend that existing benchmarks exhibit common limitations in four dimensions: constrained data composition dominated by reused legacy sou…

Cited by 0SourceScholar
2026

Monocular Normal Estimation via Shading Sequence Estimation

ICLR 2026oral

Monocular normal estimation aims to estimate normal map from a single RGB image of an object under arbitrary lighting. Existing methods rely on deep models to directly predict normal maps. However, they often suffer from 3D misalignment: while the estimated normal maps may appear to have an overall…

Cited by 0SourcecodeScholar
2026

OneHOI: Unifying Human-Object Interaction Generation and Editing

CVPR 2026

Human-Object Interaction (HOI) modelling captures how humans act upon and relate to objects, typically expressed as <person, action, object> triplets. Existing approaches split into two disjoint families: HOI generation synthesises scenes from structured triplets and layout, but fails to integrate m

Cited by 0SourcecodeScholar
2026

SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision

AAAI 2026technical

Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, t

Cited by 0SourcePDFScholar
2026

SURE: Semi-Dense Uncertainty-REfined Feature Matching

ICRA 2026poster

Establishing reliable image correspondences is essential for many robotic vision problems. However, existing methods often struggle in challenging scenarios with large viewpoint changes or textureless regions, where incorrect correspondences may still receive high similarity scores. This is mainly b…

2026

Time Is All It Takes: Spike-Retiming Attacks on Event-Driven Spiking Neural Networks

ICLR 2026poster

Spiking neural networks (SNNs) compute with discrete spikes and exploit temporal structure, yet most adversarial attacks change intensities or event counts instead of timing. We study a timing-only adversary that retimes existing spikes while preserving spike counts and amplitudes in event-driven SN…

Cited by 0SourcecodeScholar
2026

Video-KTR: Reinforcing Video Reasoning via Key Token Attribution

ICLR 2026poster

Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models (MLLMs), yet existing video reasoning methods often rely on coarse sequence-level rewards or single-factor token selection. Such approaches neglect fine-grained links among visual input…

Cited by 5SourcecodeScholar
2026

When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models are vulnerable to adversarial attacks, yet universal and transferable attacks remain underexplored, as most existing patches overfit to a single model and fail in black-box settings. To address this gap, we present a systematic study of universal, transferable adv

Cited by 0SourcecodeScholar
2025

Advancing Expert Specialization for Better MoE

NeurIPS 2025oral

Mixture-of-Experts (MoE) models enable efficient scaling of large language models (LLMs) by activating only a subset of experts per input. However, we observe that the commonly used auxiliary load balancing loss often leads to expert overlap and overly uniform routing, which hinders expert speciali…

Cited by 0SourceScholar
2025

DARR: A Dual-Branch Arithmetic Regression Reasoning Framework for Solving Machine Number Reasoning

AAAI 2025technical

Abstract visual reasoning (AVR) is a critical ability of humans, and it has been widely studied, but arithmetic visual reasoning, a unique task in AVR to reason over number sense, is less studied in the literature. To facilitate this research, we construct a Machine Number Reasoning (MNR) dataset to…

2025

DBCR: Exploiting Both Intra-cluster and Extra-cluster Relations for Compositional Reasoning

ICASSP 2025accepted

Most existing models for abstract visual reasoning perform poorly in compositional visual reasoning (CVR), due to complex nature of compositional rules and difficulties in distinguishing tiny rule differences between outliers and normal images. To tackle the challenges, we propose a Dual-Branch Comp…

Cited by 0SourceScholar
2025

DSRF: A Dynamic and Scalable Reasoning Framework for Solving RPMs

NeurIPS 2025poster

Abstract Visual Reasoning (AVR) entails discerning latent patterns in visual data and inferring underlying rules. Existing solutions often lack scalability and adaptability, as deep architectures tend to overfit training data, and static neural networks fail to dynamically capture diverse rules. To…

Cited by 0SourcecodeScholar
2025

ERL-MPP: Evolutionary Reinforcement Learning with Multi-head Puzzle Perception for Solving Large-scale Jigsaw Puzzles of Eroded Gaps

AAAI 2025technical

Solving jigsaw puzzles has been extensively studied. While most existing models focus on solving either small-scale puzzles or puzzles with no gap between fragments, solving large-scale puzzles with gaps presents distinctive challenges in both image understanding and combinatorial optimization. To t…

Cited by 0SourcePDFScholar
2025

EvHDR-GS: Event-guided HDR Video Reconstruction with 3D Gaussian Splatting

AAAI 2025technical

High Dynamic Range (HDR) video reconstruction seeks to accurately restore the extensive dynamic range present in real-world scenes and is widely employed in downstream applications. Existing methods typically operate on one or a small number of consecutive frames, which often leads to inconsistent b…

Cited by 0SourcePDFScholar
2025

Exploiting Temporal State Space Sharing for Video Semantic Segmentation

CVPR 2025poster

Video semantic segmentation (VSS) plays a vital role in understanding the temporal evolution of scenes. Traditional methods often segment videos frame-by-frame or in a short temporal window, leading to limited temporal context, redundant computations, and heavy memory requirements. To this end, we i…

2025

FIRM: Flexible Interactive Reflection ReMoval

AAAI 2025technical

Removing reflection from a single image is challenging due to the absence of general reflection priors. Although existing methods incorporate extensive user guidance for satisfactory performance, they often lack the flexibility to adapt user guidance in different modalities, and dense user interacti…

2025

Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension

AAAI 2025technical

In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC extends the scope to a more practical setting by further encompassing no-target and…

Cited by 1SourcePDFScholar
2025

Jointly Optimizing Data Discretization and Naive Bayes Classifier via Multi-Objective Optimization

ICASSP 2025accepted

Data discretization plays a critical role in enhancing the performance of the naive Bayes classifier. Traditional data discretization methods often utilize a two-stage framework, where data discretization and classification are optimized separately, leading to sub-optimal performance. To tackle the…

Cited by 0SourceScholar
2025

Multi-Scale Finetuning for Encoder-based Time Series Foundation Models

NeurIPS 2025poster

Time series foundation models (TSFMs) demonstrate impressive zero-shot performance for time series forecasting. However, an important yet underexplored challenge is how to effectively finetune TSFMs on specific downstream tasks. While naive finetuning can yield performance gains, we argue that it fa…

Cited by 0SourcecodeScholar
2025

R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localization

CVPR 2025poster

Learning-based visual localization methods that use scene coordinate regression (SCR) offer the advantage of smaller map sizes. However, on datasets with complex illumination changes or image-level ambiguities, it remains a less robust alternative to feature matching methods. This work aims to close…

2025

Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference Systems

CVPR 2025highlight

By locally encoding raw data into intermediate features, collaborative inference enables end users to leverage powerful deep learning models without exposure of sensitive raw data to cloud servers. However, recent studies have revealed that these intermediate features may not sufficiently preserve p…

2024

Connecting Consistency Distillation to Score Distillation for Text-to-3D Generation

ECCV 2024poster

"Although recent advancements in text-to-3D generation have significantly improved generation quality, issues like limited level of detail and low fidelity still persist, which requires further improvement. To understand the essence of those issues, we thoroughly analyze current score distillation m…

2024

DocMSU: A Comprehensive Benchmark for Document-Level Multimodal Sarcasm Understanding

AAAI 2024technical

Multimodal Sarcasm Understanding (MSU) has a wide range of applications in the news field such as public opinion analysis and forgery detection. However, existing MSU benchmarks and approaches usually focus on sentence-level MSU. In document-level news, sarcasm clues are sparse or small and are of…

2024

GLACE: Global Local Accelerated Coordinate Encoding

CVPR 2024poster

Scene coordinate regression (SCR) methods are a family of visual localization methods that directly regress 2D-3D matches for camera pose estimation. They are effective in small-scale scenes but face significant challenges in large-scale scenes that are further amplified in the absence of ground tru…

2024

InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models

CVPR 2024poster

Large-scale text-to-image (T2I) diffusion models have showcased incredible capabilities in generating coherent images based on textual descriptions enabling vast applications in content generation. While recent advancements have introduced control over factors such as object localization posture and…

2024

Mitigating the Curse of Dimensionality for Certified Robustness via Dual Randomized Smoothing

ICLR 2024poster

Randomized Smoothing (RS) has been proven a promising method for endowing an arbitrary image classifier with certified robustness. However, the substantial uncertainty inherent in the high-dimensional isotropic Gaussian noise imposes the curse of dimensionality on RS. Specifically, the upper bound o…

2024

Pano-NeRF: Synthesizing High Dynamic Range Novel Views with Geometry from Sparse Low Dynamic Range Panoramic Images

AAAI 2024technical

Panoramic imaging research on geometry recovery and High Dynamic Range (HDR) reconstruction becomes a trend with the development of Extended Reality (XR). Neural Radiance Fields (NeRF) provide a promising scene representation for both tasks without requiring extensive prior data. How- ever, in the c…

2024

Regression Residual Reasoning with Pseudo-labeled Contrastive Learning for Uncovering Multiple Complex Compositional Relations

IJCAI 2024poster

Abstract Visual Reasoning (AVR) has been widely studied in literature. Our study reveals that AVR models tend to rely on appearance matching rather than a genuine understanding of underlying rules. We hence develop a challenging benchmark, Multiple Complex Compositional Reasoning (MC2R), composed of…

Cited by 4SourcePDFScholar
2024

Spin-UP: Spin Light for Natural Light Uncalibrated Photometric Stereo

CVPR 2024poster

Natural Light Uncalibrated Photometric Stereo (NaUPS) relieves the strict environment and light assumptions in classical Uncalibrated Photometric Stereo (UPS) methods. However due to the intrinsic ill-posedness and high-dimensional ambiguities addressing NaUPS is still an open question. Existing wor…

2024

Transferable Adversarial Attacks on SAM and Its Downstream Models

NeurIPS 2024poster

The utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical applications, their open accessibility also poses threats of adverse usage. This paper, for the first time, explores th…

2023

Class-Incremental Learning on Multivariate Time Series Via Shape-Aligned Temporal Distillation

ICASSP 2023accepted

Class-incremental learning (CIL) on multivariate time series (MTS) is an important yet understudied problem. Based on practical privacy-sensitive circumstances, we propose a novel distillation-based strategy using a single-headed classifier without saving historical samples. We propose to exploit So…

Cited by 0SourceScholar
2023

DANI-Net: Uncalibrated Photometric Stereo by Differentiable Shadow Handling, Anisotropic Reflectance Modeling, and Neural Inverse Rendering

CVPR 2023poster

Uncalibrated photometric stereo (UPS) is challenging due to the inherent ambiguity brought by the unknown light. Although the ambiguity is alleviated on non-Lambertian objects, the problem is still difficult to solve for more general objects with complex shapes introducing irregular shadows and gene…

2023

Dual-Stream Siamese Vision Transformer With Mutual Attention For Radar Gait Verification

ICASSP 2023accepted

The inconspicuousness of human gait characteristic in radar signal makes it hard to differentiate different identities. In this work, a Dual-stream Siamese Vision Transformer with Mutual Attention is proposed to verify whether a pair of radar gait sequences originate from the same person or not. The…

Cited by 0SourceScholar
2023

Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning

AAAI 2023technical

Raven’s Progressive Matrices (RPMs) have been widely used to evaluate the visual reasoning ability of humans. To tackle the challenges of visual perception and logic reasoning on RPMs, we propose a Hierarchical ConViT with Attention-based Relational Reasoner (HCV-ARR). Traditional solution methods o…

2023

MOSE: A New Dataset for Video Object Segmentation in Complex Scenes

ICCV 2023poster

Video object segmentation (VOS) aims at segmenting a particular object throughout the entire video clip sequence. The state-of-the-art VOS methods have achieved excellent performance (e.g., 90+% J&F) on existing datasets. However, since the target objects in these existing datasets are usually relat…

Cited by 148PDFcodeScholar
2023

MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

ICCV 2023poster

This paper strives for motion expressions guided video segmentation, which focuses on segmenting objects in video content based on a sentence describing the motion of the objects. Existing referring video object datasets typically focus on salient objects and use language expressions that contain ex…

Cited by 110PDFcodeScholar
2023

Solving Jigsaw Puzzle of Large Eroded Gaps Using Puzzlet Discriminant Network

ICASSP 2023accepted

Solving Jigsaw puzzles has recently become an emerging research topic. Traditionally, boundary similarities are utilized for puzzle reassembly. In this paper, we solve Jigsaw Puzzles of Large Eroded Gaps (JPLEG), where boundary similarities are weak and image semantics are the only feasible clues. I…

Cited by 0SourceScholar
2022

Attention-Based Dual-Stream Vision Transformer for Radar Gait Recognition

ICASSP 2022accepted

Radar gait recognition is robust to light variations and less infringement on privacy. Previous studies often utilize either spectrograms or cadence velocity diagrams. While the former shows the time-frequency patterns, the latter encodes the repetitive frequency patterns. In this work, a dual-strea…

Cited by 0SourceScholar
2021

Interaction via Bi-Directional Graph of Semantic Region Affinity for Scene Parsing

ICCV 2021poster

In this work, we devote to address the challenging problem of scene parsing. Previous methods, though capture context to exploit global clues, handle scene parsing as a pixel-independent task. However, it is well known that pixels in an image are highly correlated with each other, especially those f…

Cited by 19PDFScholar
2021

Single Image Reflection Removal With Absorption Effect

CVPR 2021poster

In this paper, we consider the absorption effect for the problem of single image reflection removal. We show that the absorption effect can be numerically approximated by the average of refractive amplitude coefficient map. We then reformulate the image formation model and propose a two-step solutio…

Cited by 54PDFcodeScholar
2021

Vision-Language Transformer and Query Generation for Referring Segmentation

ICCV 2021poster

In this work, we address the challenging task of referring segmentation. The query expression in referring segmentation typically indicates the target object by describing its relationship with others. Therefore, to find the target one among all instances in the image, the model must have a holistic…

Cited by 297PDFcodeScholar
2020

PhraseClick: Toward Achieving Flexible Interactive Segmentation by Phrase and Click

ECCV 2020poster

Existing interactive object segmentation methods mainly take spatial interactions such as bounding boxes or clicks as input. However, these interactions do not contain information about explicit attributes of the target-of-interest and thus cannot quickly specify what the selected object exactly is,…

Cited by 64SourcePDFScholar
2020

Temporal Distinct Representation Learning for Action Recognition

ECCV 2020poster

Motivated by the previous success of Two-Dimensional Convolutional Neural Network (2D CNN) on image recognition, researchers endeavor to leverage it to characterize videos. However, one limitation of applying 2D CNN to analyze videos is that different frames of a video share the same 2D CNN kernels,…

Cited by 38SourcePDFScholar
2020

What Does Plate Glass Reveal About Camera Calibration?

CVPR 2020poster

This paper aims to calibrate the orientation of glass and the field of view of the camera from a single reflection-contaminated image. We show how a reflective amplitude coefficient map can be used as a calibration cue. Different from existing methods, the proposed solution is free from image conten…

Cited by 19PDFScholar
2019

Boundary-Aware Feature Propagation for Scene Segmentation

ICCV 2019poster

In this work, we address the challenging issue of scene segmentation. To increase the feature similarity of the same object while keeping the feature discrimination of different objects, we explore to propagate information throughout the image under the control of objects' boundaries. To this end, w…

Cited by 286PDFcodeScholar
2019

SPLINE-Net: Sparse Photometric Stereo Through Lighting Interpolation and Normal Estimation Networks

ICCV 2019poster

This paper solves the Sparse Photometric stereo through Lighting Interpolation and Normal Estimation using a generative Network (SPLINE-Net). SPLINE-Net contains a lighting interpolation network to generate dense lighting observations given a sparse set of lights as inputs followed by a normal estim…

Cited by 90PDFScholar
2019

Semantic Correlation Promoted Shape-Variant Context for Segmentation

CVPR 2019oral

Context is essential for semantic segmentation. Due to the diverse shapes of objects and their complex layout in various scene images, the spatial scales and shapes of contexts for different objects have very large variation. It is thus ineffective or inefficient to aggregate various context informa…

Cited by 215PDFcodeScholar
2018

An Ensemble Learning Method Based on Random Subspace Sampling for Palmprint Identification

ICASSP 2018accepted

Palmprint recognition is an important and widely used biometric modality with high reliability, stability and user acceptability. In this paper we propose a simple and effective ensemble learning method for palmprint identification based on Random Subspace Sampling (RSS). To achieve it, we rely on 2…

Cited by 0SourceScholar
2018

Context Contrasted Feature and Gated Multi-Scale Aggregation for Scene Segmentation

CVPR 2018poster

Scene segmentation is a challenging task as it need label every pixel in the image. It is crucial to exploit discriminative context and aggregate multi-scale features to achieve better segmentation. In this paper, we first propose a novel context contrasted local feature that not only leverages the…

Cited by 450SourcePDFScholar
2018

Deformable Pose Traversal Convolution for 3D Action and Gesture Recognition

ECCV 2018poster

The representation of 3D pose plays a critical role for 3D body action and hand gesture recognition. Rather than directly representing the 3D pose using its joint locations, in this paper, we propose Deformable Pose Traversal Convolution which applies one-dimensional convolution to traverse the 3D p…

Cited by 87SourcePDFScholar
2017

Laplace gradient based Discriminative and Contrast Invertible descriptor

ICASSP 2017accepted

The performance of local descriptors such as SIFT drops under severe illumination changes. In this paper, we propose a Discriminative and Contrast Invertible (DCI) local feature descriptor. In order to increase the discriminative ability of the descriptor under illumination changes, a Laplace gradie…

Cited by 0SourceScholar