← Search

Yuxin Huang

19 accepted papers

2026

Consensus-Aligned Neuron Efficient Fine-Tuning Large Language Models for Multi-Domain Machine Translation

AAAI 2026technical

Multi-domain machine translation (MDMT) aims to build a unified model capable of translating content across diverse domains. Despite the impressive machine translation capabilities demonstrated by large language models (LLMs), domain adaptation still remains a challenge for LLMs. Existing MDMT metho

Cited by 0SourcePDFScholar
2026

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

AAAI 2026technical

Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention to target regions. Instead, visual attention is always dispe

Cited by 0SourcePDFScholar
2026

Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress

CVPR 2026

Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models (VLMs) based methods primarily leverage their video understanding capabilities, while neglecting their complex reasoning

Cited by 0SourceScholar
2026

SageLM: A Multi-aspect and Explainable Large Language Model for Speech Judgement

AAAI 2026technical

Speech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. We propose SageLM, an end-to-end, multi-aspect, and explainable speech LLM for co

Cited by 0SourcePDFScholar
2025

3R: Enhancing Sentence Representation Learning via Redundant Representation Reduction

EMNLP 2025

Sentence representation learning (SRL) aims to learn sentence embeddings that conform to the semantic information of sentences. In recent years, fine-tuning methods based on pre-trained models and contrastive learning frameworks have significantly advanced the quality of sentence representations. Ho

2025

A Mixed-Language Multi-Document News Summarization Dataset and a Graphs-Based Extract-Generate Model

NAACL 2025long

Existing research on news summarization primarily focuses on single-language single-document (SLSD), single-language multi-document (SLMD) or cross-language single-document (CLSD). However, in real-world scenarios, news about an international event often involves multiple documents in different lang…

2025

Dynamic Syntactic Feature Filtering and Injecting Networks for Cross-lingual Dependency Parsing

AAAI 2025technical

Pre-trained language models enhanced parsers have achieved outstanding performance in rich-resource languages. Cross-lingual dependency parsing aims to learn useful knowledge from high-resource languages to alleviate data scarcity in low-resource languages. However, effectively reducing the syntacti…

2025

GRPO-Guided Modality Selection Enhanced LoRA-Tuned LLMs for Multimodal Emotion Recognition

EMNLP 2025

Multimodal emotion recognition in conversation (MERC) aims to identify speakers’ emotional states by utilizing text, audio, and visual modalities. Although recent large language model (LLM)-based methods have demonstrated strong performance, they typically adopt static fusion strategies that integra

2025

Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation

ACL 2025finding

The success of building textless speech-to-speech translation (S2ST) models has attracted much attention. However, S2ST still faces two main challenges: 1) extracting linguistic features for various speech signals, called cross-modal (CM), and 2) learning alignment of difference languages in long se…

2025

Memory-enhanced Large Language Model for Cross-lingual Dependency Parsing via Deep Hierarchical Syntax Understanding

EMNLP 2025

Large language models (LLMs) demonstrate remarkable text generation and syntax parsing capabilities in high-resource languages. However, their performance notably declines in low-resource languages due to memory forgetting stemming from semantic interference across languages. To address this issue,

2025

Multilingual Generative Retrieval via Cross-lingual Semantic Compression

EMNLP 2025

Generative Information Retrieval is an emerging retrieval paradigm that exhibits remarkable performance in monolingual scenarios. However, applying these methods to multilingual retrieval still encounters two primary challenges, cross-lingual identifier misalignment and identifier inflation. To addr

2025

SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models

AAAI 2025technical

With the rapid advancement of large language models (LLMs), discrete speech representations have become crucial for integrating speech into LLMs. Existing methods for speech representation discretization rely on a predefined codebook size and Euclidean distance-based quantization. However, 1) the si…

2025

Voice Conversion via Structural Entropy

ICASSP 2025accepted

Voice conversion (VC) aims to transform a person’s voice to resemble that of another person while maintaining the original linguistic content. Existing methods suffer from the blurring of speech representations and the leakage of prosody information. To address this issue, this study introduces SEVC…

Cited by 0SourceScholar
2024

DETS: End-to-End Single-Stage Text-to-Speech Via Hierarchical Diffusion Gan Models

ICASSP 2024accepted

End-to-end single-stage text-to-speech models have garnered significant attention in recent research, surpassing the performance of conventional two-stage pipeline systems. While prior single-stage models have made substantial advancements, there remains room for improvement in addressing intermitte…

Cited by 0SourceScholar
2024

Representation Alignment and Adversarial Networks for Cross-lingual Dependency Parsing

EMNLP 2024finding

With the strong representational capabilities of pre-trained language models, dependency parsing in resource-rich languages has seen significant advancements. However, the parsing accuracy drops sharply when the model is transferred to low-resource language due to distribution shifts. To alleviate t…

2023

Layer-wise Fusion with Modality Independence Modeling for Multi-modal Emotion Recognition

ACL 2023long

Multi-modal emotion recognition has gained increasing attention in recent years due to its widespread applications and the advances in multi-modal learning approaches. However, previous studies primarily focus on developing models that exploit the unification of multiple modalities. In this paper, w…

2023

Non-parallel Accent Transfer based on Fine-grained Controllable Accent Modelling

EMNLP 2023long findings

Existing accent transfer works rely on parallel data or speech recognition models. This paper focuses on the practical application of accent transfer and aims to implement accent transfer using non-parallel datasets. The study has encountered the challenge of speech representation disentanglement an…

Cited by 0SourceScholar
2022

PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit

NAACL 2022system demonstrations

PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code structure. This paper describes the design philosophy and core architecture of PaddleS…

2020

Multi-Branch Learning for Weakly-Labeled Sound Event Detection

ICASSP 2020accepted

There are two sub-tasks implied in the weakly-supervised SED: audio tagging and event boundary detection. Current methods which combine multi-task learning with SED requires annotations both for these two sub-tasks. Since there are only annotations for audio tagging available in weakly-supervised SE…

Cited by 0SourceScholar