← Search

Xinyuan Song

10 accepted papers

2026

DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces

CVPR 2026

Articulated object pose estimation is a core task in embodied AI and computer vision. Existing methods typically regress poses in a continuous space, but often struggle with 1) navigating a large, complex search space and 2) failing to incorporate intrinsic kinematic constraints. In this paper, we i

Cited by 0SourceScholar
2026

Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning

ICASSP 2026poster

Recent studies suggest that the deeper layers of Large Language Models (LLMs) contribute little to representation learning and can often be removed without significant performance loss. However, such claims are typically drawn from narrow evaluations and may overlook important aspects of model behav…

Cited by 0SourcePDFScholar
2026

GAHMN: A Generative Approach for High-Dimensional Mediation Analysis

AAAI 2026technical

High-dimensional mediation analysis (HMA) seeks to uncover complex causal mechanisms involving numerous mediators and plays a crucial role in scientific and social sciences. In this work, we introduce the Generative Adversarial High-dimensional Mediation Network (GAHMN), a novel, scalable structured

Cited by 0SourcePDFScholar
2026

Newton-coupled Dual-Teacher Semi-supervised Learning Framework

ICML 2026poster

Most semi-supervised learning frameworks rely on a single teacher that transfers zero-order supervision through pseudo-labels, constraining the student to imitate categorical outputs without perceiving the loss geometry. This design often leads to unstable optimization and limited generalization und…

Cited by 0SourceScholar
2026

When Does Sparsity Mitigate the Curse of Depth in LLMs

ICML 2026poster

Recent work has demonstrated the curse of depth in large language models (LLMs), where later layers contribute less to learning and representation than earlier layers. Such under-utilization is linked to the accumulated growth of variance in Pre-Layer Normalization, which can push deep blocks toward…

Cited by 0SourceScholar
2025

BiAssemble: Learning Collaborative Affordance for Bimanual Geometric Assembly

ICML 2025poster

Shape assembly, the process of combining parts into a complete whole, is a crucial skill for robots with broad real-world applications. Among the various assembly tasks, geometric assembly—where broken parts are reassembled into their original form (e.g., reconstructing a shattered bowl)—is particul…

Cited by 0SourcePDFScholar
2025

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling

NeurIPS 2025poster

Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers,…

Cited by 0SourcecodeScholar
2025

SOLAR: Serendipity Optimized Language Model Aligned for Recommendation

EMNLP 2025

Recently, Large Language Models (LLMs) have shown strong potential in recommendation tasks due to their broad world knowledge and reasoning capabilities. However, applying them to serendipity-oriented recommendation remains challenging, mainly due to a domain gap of LLMs in modeling personalized use

2025

TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning

NeurIPS 2025poster

We introduce TRiCo, a novel triadic game-theoretic co-training framework that rethinks the structure of semi-supervised learning by incorporating a teacher, two students, and an adversarial generator into a unified training paradigm. Unlike existing co-training or teacher-student approaches, TRiCo f…

Cited by 0SourceScholar
2025

The Curse of Depth in Large Language Models

NeurIPS 2025poster

In this paper, we re-introduce the Curse of Depth, a concept that re-introduces, explains, and addresses the recent observation in modern Large Language Models (LLMs) where deeper layers are much less effective than expected. We first confirm the wide existence of this phenomenon across the most pop…

Cited by 0SourceScholar