← Search

Canran Xiao

16 accepted papers

2026

Affordance-First Decomposition for Continual Learning in Video-Language Understanding

CVPR 2026

Continual learning for video--language understanding is increasingly important as models face non-stationary data, domains, and query styles, yet prevailing solutions blur what should stay stable versus what should adapt, rely on static routing/capacity, or require replaying past videos. We aim to e

Cited by 0SourceScholar
2026

CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated Rewards

AAAI 2026technical

Large-scale Chinese spelling correction (CSC) remains critical for real-world text processing, yet existing LLMs and supervised methods lack robustness to novel errors and rely on costly annotations. We introduce CEC-Zero, a zerosupervision reinforcement learning framework that addresses this by ena

Cited by 0SourcePDFScholar
2026

CoMem: Compositional Concept-Graph Memory for Vision–Language Adaptation

ICLR 2026poster

Continual vision–language learning is crucial for multimodal tasks such as image–text retrieval, visual question answering, and grounded reasoning in dynamic environments, yet deployed systems must learn from non-stationary streams under strict privacy and memory budgets, where naïve finetuning forg…

Cited by 0SourceScholar
2026

From Points to Coalitions: Hierarchical Contrastive Shapley Values for Prioritizing Data Samples

AAAI 2026technical

How should we quantify the value of each training example when datasets are large, heterogeneous, and geometrically structured? Classical Data-Shapley answers in principle, but its O(n!) complexity and point-wise perspective are ill-suited to modern scales. We propose Hierarchical Contrastive Data V

Cited by 0SourcePDFScholar
2026

Influence-Disentangled Federated Training: Learning Models That Are Easy to Unlearn

ICML 2026poster

Federated learning increasingly faces deletion requests that require client-level unlearning without sacrificing model quality, yet a client’s influence is often deeply entangled after many rounds of aggregation. We aim to make unlearning fast, stable, and predictable by reducing the gap to leave-on…

Cited by 0SourceScholar
2026

Meta-UCF: Unified Task-Conditioned LoRA Generation for Continual Learning in Large Language Models

ICLR 2026poster

Large language models are increasingly deployed in settings where newtasks arrive continuously, yet existing parameter-efficient finetuning (PEFT) methods either bloat linearly with the task horizon or sacrifice deep adaptation, leaving catastrophic forgetting unresolved. We aim to achieve memory-co…

Cited by 0SourceScholar
2026

Path Matters: Unveiling Geometric Implicit Bias via Curvature-Aware Sparse View Optimization

ICLR 2026poster

3D Gaussian Splatting (3DGS) has recently emerged as a powerful approach for novel view synthesis by reconstructing scenes as sets of Gaussian ellipsoids. Despite its success in scenarios with dense input images, 3DGS faces critical challenges in sparse view settings, often resulting in geometric in…

Cited by 0SourceScholar
2026

Pi-CCA: Prompt-Invariant CCA Certificates for Replay-Free Continual Multimodal Learning

ICLR 2026poster

When deployed on non-stationary data streams, foundation vision-language models require continual updates without access to past data. However, naive fine-tuning undermines their zero-shot recognition capabilities and prompt robustness. We seek a replay-free principle that preserves pre-trained cros…

Cited by 0SourceScholar
2026

Reversible Primitive–Composition Alignment for Continual Vision–Language Learning

ICLR 2026poster

Vision-language (VL) models are increasingly deployed in non-stationary settings, yet under sequential adaptation they often preserve primitive recognition while losing compositional structure, especially with tight rehearsal budgets and no task IDs. We address this gap by asking how a continual VL…

Cited by 0SourceScholar
2026

Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented Generation

AAAI 2026technical

Large language models (LLMs) equipped with retrieval—the Retrieval-Augmented Generation (RAG) paradigm—should combine their parametric knowledge with external evidence, yet in practice they often hallucinate, over-trust noisy snippets, or ignore vital context. We introduce TCR (Transparent Conflict

Cited by 0SourcePDFScholar
2026

Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation

CVPR 2026

Multilingual text-to-image (T2I) models have advanced rapidly in terms of visual realism and semantic alignment, and are now widely utilised. Yet outputs vary across cultural contexts: because language carries cultural connotations, images synthesized from multilingual prompts should preserve cross-

Cited by 0SourceScholar
2025

Diffusion-Based Self-Supervised Imitation Learning from Imperfect Visual Servoing Demonstrations for Robotic Glass Installation

ICRA 2025

Heavy-duty glass installation is a high-risk, precision-critical task in modern construction, traditionally performed through labor-intensive and error-prone manual methods. This paper presents a novel robotic framework that leverages diffusion-based self-supervised imitation learning from imperfect

Cited by 14SourceScholar
2024

Confusion-Resistant Federated Learning via Diffusion-Based Data Harmonization on Non-IID Data

NeurIPS 2024poster

Federated learning has become a pivotal distributed learning paradigm, involving collaborative model updates across multiple nodes with private data. However, handling non-i.i.d. (not identically and independently distributed) data and ensuring model consistency across heterogeneous environments pre…

Cited by 3SourcePDFScholar