← Search

Shuo Xu

10 accepted papers

2026

Introducing Decomposed Causality with Spatiotemporal Object-Centric Representation for Video Classification

AAAI 2026technical

Video classification requires event-level representations of objects and their interactions. Existing methods typically rely on data-driven approaches, which either learn such features from whole frames or object-centric visual regions. Therefore, the modeling of spatiotemporal interactions among ob

Cited by 0SourcePDFScholar
2026

TACOcc: Target-Adaptive Cross-Modal Fusion with Sequential Volume Rendering for 3D Semantic Occupancy Prediction

ICRA 2026poster

Multi-modal 3D semantic occupancy prediction remains challenged by two fundamental issues: (i) geometric--semantic misalignment introduced by fixed-neighborhood fusion under heterogeneous sensing distributions, and (ii) feature degradation with prediction inconsistency in dynamic scenes caused by sp…

Cited by 0Scholar
2025

CarbonGlobe: A Global-Scale, Multi-Decade Dataset and Benchmark for Carbon Forecasting in Forest Ecosystems

NeurIPS 2025poster

Forest ecosystems play a critical role in the Earth system as major carbon sinks that are essential for carbon neutralization and climate change mitigation. However, the Earth has undergone significant deforestation and forest degradation, and the remaining forested areas are also facing increasing…

Cited by 0SourcecodeScholar
2025

Clean-Label Graph Backdoor Attack in the Node Classification Task

AAAI 2025technical

Graph neural networks (GNNs) have achieved impressive results in various graph learning tasks. Backdoor attacks pose a significant threat to GNNs, with a focus on dirty-label attacks. However, these attacks often necessitate the inclusion of blatantly incorrect inputs into the training set, renderin…

Cited by 0SourcePDFScholar
2025

Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuning

CVPR 2025poster

Modeling label correlations has always played a pivotal role in multi-label image classification (MLC), attracting significant attention from researchers. However, recent studies have overemphasized co-occurrence relationships among labels, which can lead to overfitting risk on this overemphasis, re…

Cited by 0SourcePDFScholar
2025

Curriculum Conditioned Diffusion for Multimodal Recommendation

AAAI 2025technical

Multimodal recommendation (MMRec) aims to integrate multimodal information of items to address the inherent data sparsity issue in collaborative-based recommendation. Traditional MMRec methods typically capture the structure-level item representations from the observed user behaviors within the mult…

Cited by 1SourcePDFScholar
2025

FlowPrune: Accelerating Attention Flow Calculation by Pruning Flow Network

NeurIPS 2025poster

The Transformer architecture serves as the foundation of modern AI systems, powering recent advances in Large Language Models (LLMs) and Large Multimodal Models (LMMs). Central to these models, attention mechanisms capture contextual dependencies via token interactions. Beyond inference, attention h…

Cited by 0SourceScholar
2024

Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models

NeurIPS 2024poster

As Archimedes famously said, ``Give me a lever long enough and a fulcrum on which to place it, and I shall move the world'', in this study, we propose to use a tiny Language Model (LM), \eg, a Transformer with 67M parameters, to lever much larger Vision-Language Models (LVLMs) with 9B parameters. Sp…

Cited by 8SourcePDFScholar
2024

SimFair: Physics-Guided Fairness-Aware Learning with Simulation Models

AAAI 2024technical

Fairness-awareness has emerged as an essential building block for the responsible use of artificial intelligence in real applications. In many cases, inequity in performance is due to the change in distribution over different regions. While techniques have been developed to improve the transferabili…

Cited by 9SourcePDFScholar