← Search

Shu Yang

37 accepted papers

2026

Dissecting Representation Misalignment in Contrastive Learning via Influence Function

ICLR 2026poster

Contrastive learning, commonly applied in large-scale multimodal models, often relies on data from diverse and often unreliable sources, which can include misaligned or mislabeled text-image pairs. This frequently leads to robustness issues and hallucinations, ultimately causing performance degradat…

Cited by 0SourceScholar
2026

ISTER: LINEAR TRANSFORMER FOR EFFICIENT MULTIVARIATE TIME SERIES FORECASTING

ICASSP 2026poster

Transformer-based models have achieved remarkable success in multivariate time series forecasting (MTSF) by capturing long-range dependencies. However, their widespread adoption is hindered by the quadratic computational complexity of self-attention, which limits scalability on high-dimensional sequ…

Cited by 0SourcePDFScholar
2026

Meta-Router: Bridging Gold-standard and Preference-based Evaluations in LLM Routing

ICLR 2026poster

In language tasks requiring extensive human-model interaction, the inference cost of large language models (LLMs) can be substantial. To reduce expenses while preserving the quality of the responses, an LLM router selects among candidate models to balance between the expected response quality and t…

Cited by 0SourcecodeScholar
2026

Neuron-Aware Data Selection in Instruction Tuning for Large Language Models

ICLR 2026poster

Instruction Tuning (IT) has been proven to be an effective approach to unlock the powerful capabilities of large language models (LLMs). Recent studies indicate that excessive IT data can degrade LLMs performance, while carefully selecting a small subset of high-quality IT data can significantly en…

Cited by 0SourceScholar
2026

Privacy-Protected Causal Survival Analysis Under Distribution Shift

ICLR 2026poster

Causal inference across multiple data sources can improve the generalizability and reproducibility of scientific findings. However, for time-to-event outcomes, data integration methods remain underdeveloped, especially when populations are heterogeneous and privacy constraints prevent direct data po…

Cited by 0SourceScholar
2026

Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models

ICML 2026poster

Large Reasoning Models (LRMs) suffer from sycophantic behavior, where models tend to agree with users' incorrect beliefs and follow misinformation rather than maintain independent reasoning. This behavior undermines model reliability and poses societal risks. Mitigating LRM sycophancy requires monit…

Cited by 0SourceScholar
2026

When Thinking Backfires: Mechanistic Insights into Reason-induced Misalignment

ICLR 2026poster

With the growing accessibility and wide adoption of large language models, concerns about their safety and alignment with human values have become paramount. In this paper, we identify a concerning phenomenon: Reasoning-Induced Misalignment (RIM), in which misalignment emerges when reasoning capabil…

Cited by 0SourceScholar
2026

When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models

AAAI 2026technical

Large Language Models (LLMs) often exhibit sycophantic behavior, agreeing with user-stated opinions even when those contradict factual knowledge. While prior work has documented this tendency, the internal mechanisms that enable such behavior remain poorly understood. In this paper, we provide a mec

Cited by 0SourcePDFScholar
2026

nD-RoPE: A Generalized RoPE for n-Dimensional Position Embedding

ICML 2026poster

Rotary Position Embedding (RoPE) is widely adopted in Transformer models, yet its extension to high-dimensional domains lacks a unified theoretical formulation. Most existing approaches either apply rotations independently along each axis or mix frequencies empirically, which limits cross-dimensiona…

Cited by 0SourceScholar
2025

ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models

AAAI 2025technical

Large Language Models (LLMs) have revolutionized natural language processing tasks. However, their practical application is constrained by substantial memory and computational demands. Post-training quantization (PTQ) is considered an effective method to accelerate LLM inference. Despite its growing…

2025

Abstain Mask Retain Core: Time Series Prediction by Adaptive Masking Loss with Representation Consistency

NeurIPS 2025spotlight

Time series forecasting plays a pivotal role in critical domains such as energy management and financial markets. Although deep learning-based approaches (e.g., MLP, RNN, Transformer) have achieved remarkable progress, the prevailing "long-sequence information gain hypothesis" exhibits inherent limi…

Cited by 0SourcecodeScholar
2025

Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluation

EMNLP 2025

Suicide remains a major global mental health challenge, and early intervention hinges on recognizing signs of suicidal ideation. In private conversations, such ideation is often expressed in subtle or conflicted ways, making detection especially difficult. Existing data sets are mainly based on publ

Cited by 0SourcePDFScholar
2025

Doubly Protected Estimation for Survival Outcomes Utilizing External Controls for Randomized Clinical Trials

ICML 2025poster

Censored survival data are common in clinical trials, but small control groups can pose challenges, particularly in rare diseases or where balanced randomization is impractical. Recent approaches leverage external controls from historical studies or real-world data to strengthen treatment evaluation…

Cited by 2SourcePDFScholar
2025

Doubly Robust Fusion of Many Treatments for Policy Learning

ICML 2025poster

Individualized treatment rules/recommendations (ITRs) aim to improve patient outcomes by tailoring treatments to the characteristics of each individual. However, in high-dimensional treatment settings, existing methods face significant challenges due to data sparsity within treatment groups and high…

Cited by 0SourcePDFScholar
2025

EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification

NeurIPS 2025poster

Understanding the internal mechanisms of transformer-based language models remains challenging. Mechanistic interpretability based on circuit discovery aims to reverse engineer neural networks by analyzing their internal processes at the level of computational subgraphs. In this paper, we revisit ex…

Cited by 0SourceScholar
2025

Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacement

CVPR 2025poster

Spiking Neural Networks (SNNs) have garnered considerable attention as a potential alternative to Artificial Neural Networks (ANNs). Recent studies have highlighted SNNs' potential on large-scale datasets. For SNN training, two main approaches exist: direct training and ANN-to-SNN (ANN2SNN) conversi…

Cited by 0SourcePDFScholar
2025

Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep Deployment

ICML 2025poster

Spiking Neural Networks (SNNs) are emerging as a brain-inspired alternative to traditional Artificial Neural Networks (ANNs), prized for their potential energy efficiency on neuromorphic hardware. Despite this, SNNs often suffer from accuracy degradation compared to ANNs and face deployment challeng…

2025

Enhancing Statistical Validity and Power in Hybrid Controlled Trials: A Randomization Inference Approach with Conformal Selective Borrowing

ICML 2025poster

External controls from historical trials or observational data can augment randomized controlled trials when large-scale randomization is impractical or unethical, such as in drug evaluation for rare diseases. However, non-randomized external controls can introduce biases, and existing Bayesian and…

2025

Evaluating and Learning Optimal Dynamic Treatment Regimes under Truncation by Death

NeurIPS 2025poster

Truncation by death, a prevalent challenge in critical care, renders traditional dynamic treatment regime (DTR) evaluation inapplicable due to ill-defined potential outcomes. We introduce a principal stratification-based method, focusing on the always-survivor value function. We derive a semiparamet…

Cited by 0SourceScholar
2025

Fraud-R1 : A Multi-Round Benchmark for Assessing the Robustness of LLM Against Augmented Fraud and Phishing Inducements

ACL 2025finding

With the increasing integration of large language models (LLMs) into real-world applications such as finance, e-commerce, and recommendation systems, their susceptibility to misinformation and adversarial manipulation poses significant risks. Existing fraud detection benchmarks primarily focus on si…

2025

Rethinking Prompt-based Debiasing in Large Language Model

ACL 2025finding

Investigating bias in large language models (LLMs) is crucial for developing trustworthy AI. While prompt-based through prompt engineering is common, its effectiveness relies on the assumption that models inherently understand biases. Our study systematically analyzed this assumption using the BBQ a…

Cited by 0SourcePDFScholar
2025

Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks

CVPR 2025poster

Spiking Neural Networks (SNNs), inspired by the human brain, offer significant computational efficiency through discrete spike-based information transfer. Despite their potential to reduce inference energy consumption, a performance gap persists between SNNs and Artificial Neural Networks (ANNs), pr…

2025

Understanding How Value Neurons Shape the Generation of Specified Values in LLMs

EMNLP 2025

Rapid integration of large language models (LLMs) into societal applications has intensified concerns about their alignment with universal ethical principles, as their internal value representations remain opaque despite behavioral alignment advancements. Current approaches struggle to systematicall

Cited by 0SourcePDFScholar
2025

Understanding the Repeat Curse in Large Language Models from a Feature Perspective

ACL 2025finding

Large language models (LLMs) have made remarkable progress in various domains, yet they often suffer from repetitive text generation, a phenomenon we refer to as the ”Repeat Curse”. While previous studies have proposed decoding strategies to mitigate repetition, the underlying mechanism behind this…

2025

Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore

COLING 2025main

The efficacy of detectors for texts generated by large language models (LLMs) substantially depends on the availability of large-scale training data. However, white-box zero-shot detectors, which require no such data, are limited by the accessibility of the source model of the LLM-generated text. In…

2024

Causal Customer Churn Analysis with Low-rank Tensor Block Hazard Model

ICML 2024poster

This study introduces an innovative method for analyzing the impact of various interventions on customer churn, using the potential outcomes framework. We present a new causal model, the tensorized latent factor block hazard model, which incorporates tensor completion methods for a principled causal…

2024

DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer’s Disease Questions with Scientific Literature

EMNLP 2024finding

Recent advancements in large language models (LLMs) have achieved promising performances across various applications. Nonetheless, the ongoing challenge of integrating long-tail knowledge continues to impede the seamless adoption of LLMs in specialized domains. In this work, we introduce DALK, a.k.a…

2024

DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios

NeurIPS 2024poster

Detecting text generated by large language models (LLMs) is of great recent interest. With zero-shot methods like DetectGPT, detection capabilities have reached impressive levels. However, the reliability of existing detectors in real-world applications remains underexplored. In this study, we prese…

2024

Positivity-free Policy Learning with Observational Data

AISTATS 2024poster

Policy learning utilizing observational data is pivotal across various domains, with the objective of learning the optimal treatment assignment policy while adhering to specific constraints such as fairness, budget, and simplicity. This study introduces a novel positivity-free (stochastic) policy le…

2023

Enhancing Treatment Effect Estimation: A Model Robust Approach Integrating Randomized Experiments and External Controls using the Double Penalty Integration Estimator

UAI 2023poster

Randomized experiments (REs) are the cornerstone for treatment effect evaluation. However, due to practical considerations, REs may encounter difficulty recruiting sufficient patients. External controls (ECs) can supplement REs to boost estimation efficiency. Yet, there may be incomparability betwee…

Cited by 12SourcePDFScholar
2021

Learning Motion-Appearance Co-Attention for Zero-Shot Video Object Segmentation

ICCV 2021poster

How to make the appearance and motion information interact effectively to accommodate complex scenarios is a fundamental issue in flow-based zero-shot video object segmentation. In this paper, we propose an Attentive Multi-Modality Collaboration Network (AMC-Net) to utilize appearance and motion inf…

Cited by 78PDFcodeScholar
2020

A Programmably Compliant Origami Mechanism for Dynamically Dexterous Robots

RA-L 2020

We present an approach to overcoming challenges in dynamical dexterity for robots through programmably compliant origami mechanisms. Our work leverages a one-parameter family of flat sheet crease patterns that folds into origami bellows, whose axial compliance can be tuned to select desired stiffnes

Cited by 30SourceScholar