← Search

Changshuo Wang

31 accepted papers

2026

Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake Detection

CVPR 2026

We investigate state-of-the-art deepfake detectors that leverage ViT-based vision foundation models and discover that the [CLS] token suffers from the Pre-trained Information Bias (PIB), i.e., it tends to mainly focus on global semantics due to the knowledge dominated by pre-trained model parameters

Cited by 0SourceScholar
2026

Biologically-Inspired Evolutionary Domain Symbiosis for Few-shot and Zero-shot Point Cloud Semantic Segmentation

AAAI 2026technical

Few-shot and zero-shot point cloud semantic segmentation aim to accurately segment novel categories using limited or no labeled samples, respectively. However, existing methods face significant challenges including domain shifts between support and query sets and the inability to handle both few-sho

Cited by 0SourcePDFScholar
2026

Breaking Manifold Continuity: Vector Quantized Modeling for Real-Centric Deepfake Detection

ICML 2026poster

The increasingly realistic and diverse generative data has led some deepfake detection methods to shift towards learning robust real content, \textit{e.g.}, via reconstruction-based tasks. However, most existing approaches rely primarily on prevalent continuous modeling (\textit{e.g.}, GMMs, VAEs, D…

Cited by 0SourceScholar
2026

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning

CVPR 2026

Multi-modal Retrieval-Augmented Generation (MMRAG) has emerged as a powerful paradigm for enhancing Multimodal Large Language Models (MLLMs) in knowledge-intensive question answering by integrating external visual, textual, and structural knowledge. However, existing MMRAG frameworks suffer from cri

Cited by 0SourceScholar
2026

DiffStyle3D: Consistent 3D Gaussian Stylization via Attention Optimization

ICML 2026poster

3D style transfer enables the creation of visually expressive 3D content, enriching the visual appearance of 3D scenes and objects. However, existing VGG- and CLIP-based methods struggle to model multi-view consistency within the model itself, while diffusion-based approaches can capture such consis…

Cited by 0SourceScholar
2026

Efficient Multimodal Spatial Reasoning via Dynamic and Asymmetric Routing

ICLR 2026poster

Recently, visualization-of-thought (VoT) has unlocked new opportunities for complex spatial reasoning in multimodal large language models (MLLMs) by complementing verbal reasoning with visual thinking. However, the autoregressive accumulation of lengthy and redundant tokens substantially increases c…

Cited by 0SourceScholar
2026

FantasyStyle: Controllable Stylized Distillation for 3D Gaussian Splatting

AAAI 2026technical

The success of 3DGS in generative and editing applications has sparked growing interest in 3DGS-based style transfer. However, current methods still face two major challenges: (1) multi-view inconsistency often leads to style conflicts, resulting in appearance smoothing and distortion; and (2) heavy

Cited by 0SourcePDFScholar
2026

From Coarse to Fine: Deep Prototype Refinement Network for Few-Shot Point Cloud Semantic Segmentation

ICML 2026poster

Few-shot point cloud semantic segmentation (FS-PCSS) aims to achieve precise segmentation of novel categories using only limited labeled samples. Existing prototype-based methods typically rely on shallow feature fusion strategies, failing to adequately model the feature distribution shift between s…

Cited by 0SourceScholar
2026

Image-to-Point Cloud Feature Back-Projection for Multimodal Training of 3D Semantic Segmentation

CVPR 2026

The effective integration and utilization of multimodal data acquired from image cameras and LiDAR is of paramount importance for perception systems. This paper proposes **I**mage-to-**P**oint Cloud **F**eature Back-**P**rojection (**IPFP**), a novel method for training multimodal fusion networks th

Cited by 0SourceScholar
2026

MMPG: MoE-based Adaptive Multi-Perspective Graph Fusion for Protein Representation Learning

AAAI 2026technical

Graph Neural Networks (GNNs) have been widely adopted for Protein Representation Learning (PRL), as residue interaction networks can be naturally represented as graphs. Current GNN-based PRL methods typically rely on single-perspective graph construction strategies, which capture partial properties

Cited by 0SourcePDFScholar
2026

Rethinking Serialization in Linear 3D Vision: Decoupling Anisotropic Geometry from Isotropic Semantics

ICML 2026poster

Current linear State-Space Models for 3D point clouds typically rely on 1D serialization (e.g., Hilbert curves) for global modeling. Such rigid ordering disrupts spatial continuity in dense scenes, introducing what we term Serialization Bias. We propose AnIsoNet, a framework that decouples anisotrop…

Cited by 0SourceScholar
2026

Rethinking Video-Language Model from the Language Input Perspective

AAAI 2026technical

Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between videos and texts. Although previous VLM works have made significant progress, almost all of them implicitly assume that all the texts are predefine

Cited by 0SourcePDFScholar
2026

R²D-LPCC: Relevance-Ranking Guided Region-Adaptive Dynamic LiDAR Point Cloud Compression

AAAI 2026technical

Dynamic LiDAR point cloud compression (LPCC) is crucial for the efficient transmission and storage of large-scale three-dimensional data in applications such as autonomous driving. However, many existing methods, which primarily focus on compressing geometric or motion information, face a fundamenta

Cited by 0SourcePDFScholar
2026

Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models

ICML 2026poster

Pre-trained vision-language models (VLMs) such as CLIP have demonstrated strong zero-shot capabilities across diverse domains, yet remain highly vulnerable to adversarial perturbations that disrupt image-text alignment and compromise reliability. Existing defenses typically rely on adversarial fine-…

Cited by 0SourceScholar
2026

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning

ICML 2026poster

Multimodal reasoning often relies on long chains of intermediate textual and visual thoughts, where accumulating visual tokens and dense cross-modal attention incur substantial computation and memory overhead. To address this challenge, we propose Spectral-Progressive Thought Flow (*SpecFlow*), a *n…

Cited by 0SourceScholar
2026

SplitFlux: Learning to Decouple Content and Style from a Single Image

CVPR 2026

Disentangling image content and style is essential for customized image generation. Existing SDXL-based methods struggle to achieve high-quality results, while the recently proposed Flux model fails to achieve effective content-style separation due to its underexplored characteristics. To address th

Cited by 0SourcecodeScholar
2026

TVDRNet: Text-driven Viewpoint Optimization via Differentiable Rendering for 3D Reasoning Segmentation

ICML 2026poster

Three-dimensional (3D) reasoning segmentation aims to segment target objects based on text instructions and 3D spatial cues. Recent efforts in 3D reasoning leverage Multimodal Large Language Models (MLLMs) to bridge the gap between text and 3D data. However, since MLLMs are primarily trained on text…

Cited by 0SourceScholar
2026

TopAdapter: Topology-Aware Prompt Tuning for Efficient Point Cloud Understanding

ICML 2026poster

Point cloud data, with its inherent geometric and topological structures, plays a critical role in 3D vision tasks. However, existing parameter-efficient fine-tuning (PEFT) methods predominantly focus on input token prompting, overlooking the intrinsic geometric information. To address this limitati…

Cited by 0SourceScholar
2026

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

AAAI 2026technical

Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these VLMs are task-specific and assume that both video and language inputs are complete. However, real-world VLM applications might face challenges due

Cited by 0SourcePDFScholar
2026

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

AAAI 2026technical

Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness against adversarial attacks—particularly those exploiting both modalities—remain

Cited by 0SourcePDFScholar
2025

DyPolySeg: Taylor Series-Inspired Dynamic Polynomial Fitting Network for Few-shot Point Cloud Semantic Segmentation

ICML 2025poster

Few-shot point cloud semantic segmentation effectively addresses data scarcity by identifying unlabeled query samples through semantic prototypes generated from a small set of labeled support samples. However, pre-training-based methods suffer from domain shifts and increased training time. Addition…

Cited by 0SourcePDFScholar
2025

FlexUOD: The Answer to Real-world Unsupervised Image Outlier Detection

CVPR 2025poster

How many outliers are within an unlabeled and contaminated dataset? Despite a series of unsupervised outlier detection (UOD) approaches have been proposed, they cannot correctly answer this critical question, resulting in their performance instability across various real-world (varying contamination…

2025

Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation

NeurIPS 2025poster

Vision-Language Navigation in Continuous Environments (VLN-CE) poses a formidable challenge for autonomous agents, requiring seamless integration of natural language instructions and visual observations to navigate complex 3D indoor spaces. Existing approaches often falter in long-horizon tasks due…

Cited by 0SourceScholar
2025

Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer Network

AAAI 2025technical

Given some video-query pairs with untrimmed videos and sentence queries, temporal sentence grounding (TSG) aims to locate query-relevant segments in these videos. Although previous respectable TSG methods have achieved remarkable success, they train each video-query pair separately and ignore the re…

Cited by 4SourcePDFScholar
2025

Point Clouds Meets Physics: Dynamic Acoustic Field Fitting Network for Point Cloud Understanding

CVPR 2025poster

While existing pre-training-based methods have enhanced point cloud model performance, they have not fundamentally resolved the challenge of local structure representation in point clouds. The limited representational capacity of pure point cloud models continues to constrain the potential of cross-…

Cited by 1SourcePDFScholar
2025

Reasoning Beyond Points: A Visual Introspective Approach for Few-Shot 3D Segmentation

NeurIPS 2025poster

Point Cloud Few-Shot Semantic Segmentation (PC-FSS) aims to segment unknown categories in query samples using only a small number of annotated support samples. However, scene complexity and insufficient representation of local geometric structures pose significant challenges to PC-FSS. To address th…

Cited by 0SourcecodeScholar
2025

ReferSplat: Referring Segmentation in 3D Gaussian Splatting

ICML 2025oral

We introduce Referring 3D Gaussian Splatting Segmentation (R3DGS), a new task that aims to segment target objects in a 3D Gaussian scene based on natural language descriptions, which often contain spatial relationships or object attributes. This task requires the model to identify newly described o…

2025

Synonymous Variational Inference for Perceptual Image Compression

ICML 2025poster

Recent contributions of semantic information theory reveal the set-element relationship between semantic and syntactic information, represented as synonymous relationships. In this paper, we propose a synonymous variational inference (SVI) method based on this synonymity viewpoint to re-analyze the…

2025

Taylor Series-Inspired Local Structure Fitting Network for Few-shot Point Cloud Semantic Segmentation

AAAI 2025technical

Few-shot point cloud semantic segmentation aims to accurately segment "unseen" new categories in point cloud scenes using limited labeled data. However, pretraining-based methods not only introduce excessive time overhead but also overlook the local structure representation among irregular point clo…

2023

Compositional Zero-Shot Artistic Font Synthesis

IJCAI 2023poster

Recently, many researchers have made remarkable achievements in the field of artistic font synthesis, with impressive glyph style and effect style in the results. However, due to less exploration in style disentanglement, it is difficult for existing methods to envision a kind of unseen style (glyph…