← Search

Yixian Shen

11 accepted papers

2026

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech

ICASSP 2026poster

Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker identity within a single utterance. This underexplored phenomenon undermines the coherence of synthetic speech, especially…

Cited by 0SourcePDFScholar
2026

Efficient Multimodal Spatial Reasoning via Dynamic and Asymmetric Routing

ICLR 2026poster

Recently, visualization-of-thought (VoT) has unlocked new opportunities for complex spatial reasoning in multimodal large language models (MLLMs) by complementing verbal reasoning with visual thinking. However, the autoregressive accumulation of lengthy and redundant tokens substantially increases c…

Cited by 0SourceScholar
2026

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning

ICML 2026poster

Multimodal reasoning often relies on long chains of intermediate textual and visual thoughts, where accumulating visual tokens and dense cross-modal attention incur substantial computation and memory overhead. To address this challenge, we propose Spectral-Progressive Thought Flow (*SpecFlow*), a *n…

Cited by 0SourceScholar
2026

TopAdapter: Topology-Aware Prompt Tuning for Efficient Point Cloud Understanding

ICML 2026poster

Point cloud data, with its inherent geometric and topological structures, plays a critical role in 3D vision tasks. However, existing parameter-efficient fine-tuning (PEFT) methods predominantly focus on input token prompting, overlooking the intrinsic geometric information. To address this limitati…

Cited by 0SourceScholar
2025

AdaDCP: Learning an Adapter with Discrete Cosine Prior for Clear-to-Adverse Domain Generalization

ICCV 2025poster

Vision Foundation Model (VFM) provides an inherent generalization ability to unseen domains for downstream tasks. However, fine-tuning VFM to parse various adverse scenes (e.g., fog, snow, night) is particularly challenging, as these samples are difficult to collect. Using easy-to-acquire clear scen…

Cited by 0SourcePDFScholar
2025

Degradation-Aware Dynamic Schrödinger Bridge for Unpaired Image Restoration

NeurIPS 2025poster

Image restoration is a fundamental task in computer vision and machine learning, which learns a mapping between the clear images and the degraded images under various conditions (e.g., blur, low-light, haze). Yet, most existing image restoration methods are highly restricted by the requirement of de…

Cited by 0SourceScholar
2025

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training

AAAI 2025technical

Large Language Models (LLMs) have shown remarkable performance across various tasks, but the escalating demands on computational resources pose significant challenges, particularly in the extensive utilization of full fine-tuning for downstream tasks. To address this, parameter-efficient fine-tuning…

2025

MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection

ACL 2025long

We present a new adaptation method MaCP, Minimal yet Mighty adaptive Cosine Projection, that achieves exceptional performance while requiring minimal parameters and memory for fine-tuning large foundation models.Its general idea is to exploit the superior energy compaction and decorrelation properti…

Cited by 0SourcePDFScholar
2025

NeuroAda: Activating Each Neuron’s Potential for Parameter-Efficient Fine-Tuning

EMNLP 2025

Existing parameter-efficient fine-tuning (PEFT) methods primarily fall into two categories: addition-based and selective in-situ adaptation. The former, such as LoRA, introduce additional modules to adapt the model to downstream tasks, offering strong memory efficiency. However, their representation

2025

Reasoning Beyond Points: A Visual Introspective Approach for Few-Shot 3D Segmentation

NeurIPS 2025poster

Point Cloud Few-Shot Semantic Segmentation (PC-FSS) aims to segment unknown categories in query samples using only a small number of annotated support samples. However, scene complexity and insufficient representation of local geometric structures pose significant challenges to PC-FSS. To address th…

Cited by 0SourcecodeScholar
2025

SSH: Sparse Spectrum Adaptation via Discrete Hartley Transformation

NAACL 2025long

Low-rank adaptation (LoRA) has been demonstrated effective in reducing the trainable parameter number when fine-tuning a large foundation model (LLM). However, it still encounters computational and memory challenges when scaling to larger models or addressing more complex task adaptation.In this wor…

Cited by 0SourcePDFScholar