← Search

Xiang Xiang

13 accepted papers

2026

Dismantling the Illusion of Vision-Language-Action Models Competence via Explicit Distributional Shifts

ICML 2026poster

Given that simulation can never exhaustively enumerate reality, generalization is the determining factor for whether Vision-Language-Action (VLA) models can translate benchmark success into real-world functionality. However, current evaluation protocols often incentivize mechanical memorization rath…

Cited by 0SourceScholar
2026

Random Amalgamation of Adapters for Flatter Loss Landscapes: Towards Class-Incremental Learning with Better Stability

AAAI 2026technical

Class-incremental learning (CIL) enables models to continuously learn from streaming data while mitigating catastrophic forgetting of prior knowledge. Our research reveals that the CIL performance of pre-trained models (PTMs) varies significantly across different datasets, a phenomenon underexplored

Cited by 0SourcePDFScholar
2025

Decoupled Entropy Minimization

NeurIPS 2025poster

Entropy Minimization (EM) is beneficial to reducing class overlap, bridging domain gap, and restricting uncertainty for various tasks in machine learning, yet its potential is limited. To study the internal mechanism of EM, we reformulate and decouple the classical EM into two parts with opposite ef…

Cited by 0SourceScholar
2025

Overcoming Shortcut Problem in VLM for Robust Out-of-Distribution Detection

CVPR 2025highlight

Vision-language models (VLMs), such as CLIP, have shown remarkable capabilities in downstream tasks. However, the coupling of semantic information between the foreground and the background in images leads to significant shortcut issues that adversely affect out-of-distribution (OOD) detection abilit…

2024

Aligning Logits Generatively for Principled Black-Box Knowledge Distillation

CVPR 2024poster

Black-Box Knowledge Distillation (B2KD) is a formulated problem for cloud-to-edge model compression with invisible data and models hosted on the server. B2KD faces challenges such as limited Internet exchange and edge-cloud disparity of data distributions. In this paper we formalize a two-step workf…

2024

Enhancing the General Agent Capabilities of Low-Paramter LLMs through Tuning and Multi-Branch Reasoning

NAACL 2024findings

Open-source pre-trained Large Language Models (LLMs) exhibit strong language understanding and generation capabilities, making them highly successful in a variety of tasks. However, when used as agents for dealing with complex problems in the real world, their performance is far inferior to large co…

2022

Coarse-to-Fine Incremental Few-Shot Learning

ECCV 2022poster

"Different from fine-tuning models pre-trained on a large-scale dataset of preset classes, class-incremental learning (CIL) aims to recognize novel classes over time without forgetting pre-trained classes. However, a given model will be challenged by test images with finer-grained classes, e.g., a b…

2022

Hierarchical Memory Learning for Fine-Grained Scene Graph Generation

ECCV 2022poster

"Regarding Scene Graph Generation (SGG), coarse and fine predicates mix in the dataset due to the crowd-sourced labeling, and the long-tail problem is also pronounced. Given this tricky situation, many existing SGG methods treat the predicates equally and learn the model under the supervision of mix…

Cited by 31SourcePDFScholar
2015

Hierarchical Sparse and Collaborative Low-Rank representation for emotion recognition

ICASSP 2015accepted

In this paper, we design a Collaborative-Hierarchical Sparse and Low-Rank (C-HiSLR) model that is natural for recognizing human emotion in visual data. Previous attempts require explicit expression components, which are often unavailable and difficult to recover. Instead, our model exploits the low-…

Cited by 0SourceScholar