← Search

Yuan Zhao

17 accepted papers

2026

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance

ICLR 2026poster

Vision-Language-Action (VLA) models pre-trained on large, diverse datasets show remarkable potential for general-purpose robotic manipulation. However, a primary bottleneck remains in adapting these models to downstream tasks, especially when the robot's embodiment or the task itself differs from th…

Cited by 0SourcecodeScholar
2026

Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents

AAAI 2026technical

Graphical User Interface (GUI) task automation constitutes a critical frontier in artificial intelligence research. While effective GUI agents synergistically integrate planning and grounding capabilities, current methodologies exhibit two fundamental limitations: (1) insufficient exploitation of cr

Cited by 0SourcePDFScholar
2026

Complementary Prototype Mapping for Efficient Multimodal Anomaly Detection

CVPR 2026

Multimodal unsupervised anomaly detection has garnered increasing attention for robust defect localization.Recent approaches rely on establishing cross-modal matching relationships under normal conditions without explicit guidance.However, in practice, a single modality may have multiple distinct re

Cited by 0SourcecodeScholar
2026

Importance-Aware Data Selection for Efficient LLM Instruction Tuning

AAAI 2026technical

Instruction tuning plays a critical role in enhancing the performance and efficiency of Large Language Models (LLMs). Its success depends not only on the quality of the instruction data but also on the inherent capabilities of the LLM itself. Some studies suggest that even a small amount of high-qua

Cited by 0SourcePDFScholar
2026

UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression

CVPR 2026

Existing anomaly detection methods often treat the modality and class as independent factors. Although this paradigm has enriched the development of AD research branches and produced many specialized models, it has also led to fragmented solutions and excessive memory overhead. Moreover, reconstruct

Cited by 0SourcecodeScholar
2025

GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models

ACL 2025long

The rapid growth of large language models (LLMs) with traditional centralized fine-tuning emerges as a key technique for adapting these models to domain-specific challenges, yielding privacy risks for both model and data owners. One promising solution, called offsite-tuning (OT), is proposed to addr…

2025

ScaleOT: Privacy-utility-scalable Offsite-tuning with Dynamic LayerReplace and Selective Rank Compression

AAAI 2025technical

Offsite-tuning is a privacy-preserving method for tuning large language models (LLMs) by sharing a lossy compressed emulator from the LLM owners with data owners for downstream task tuning. This approach protects the privacy of both the model and data owners. However, current offsite tuning methods…

2024

Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models

EMNLP 2024finding

Parameter-Efficient Fine-Tuning (PEFT) methods have gained significant popularity for adapting pre-trained Large Language Models (LLMs) to downstream tasks, primarily due to their potential to significantly reduce memory and computational overheads. However, a common limitation in most PEFT approach…

2024

Multi-Signal Fusion of Social Diffusion Graph with Bi-Directional Semantic Consistency

ICASSP 2024accepted

Devising diffusion graph to learn user representations is a crucial step in studying information propagation prediction. However, previous works mainly focused on structural and temporal features. To better incorporate content features, we introduce the Backward Decomposition and Forward Preservatio…

Cited by 0SourceScholar
2024

eXponential FAmily Dynamical Systems (XFADS): Large-scale nonlinear Gaussian state-space modeling

NeurIPS 2024poster

State-space graphical models and the variational autoencoder framework provide a principled apparatus for learning dynamical systems from data. State-of-the-art probabilistic approaches are often able to scale to large problems at the cost of flexibility of the variational posterior or expressivity…

2023

Linear Time GPs for Inferring Latent Trajectories from Neural Spike Trains

ICML 2023poster

Latent Gaussian process (GP) models are widely used in neuroscience to uncover hidden state evolutions from sequential observations, mainly in neural activity recordings. While latent GP models provide a principled and powerful solution in theory, the intractable posterior in non-conjugate settings…

Cited by 8SourcePDFScholar
2023

Real-time variational method for learning neural trajectory and its dynamics

ICLR 2023top-25%

Latent variable models have become instrumental in computational neuroscience for reasoning about neural computation. This has fostered the development of powerful offline algorithms for extracting latent neural trajectories from neural recordings. However, despite the potential of real-time alter…

Cited by 9SourcePDFScholar
2023

Towards Lightweight, Model-Agnostic and Diversity-Aware Active Anomaly Detection

ICLR 2023poster

Active Anomaly Discovery (AAD) is flourishing in the anomaly detection research area, which aims to incorporate analysts’ feedback into unsupervised anomaly detectors. However, existing AAD approaches usually prioritize the samples with the highest anomaly scores for user labeling, which hinders the…

Cited by 1SourcePDFScholar