← Search

Junan Zhang

5 accepted papers

2026

Multi-Metric Preference Alignment for Generative Speech Restoration

AAAI 2026technical

Recent generative models have significantly advanced speech restoration tasks, yet their training objectives often misalign with human perceptual preferences, resulting in suboptimal quality. While post-training alignment has proven effective in other generative domains like text and image generatio

Cited by 0SourcePDFScholar
2025

LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models

ICLR 2025spotlight

With the rapid development of AI-generated content, the future internet may be inundated with synthetic data, making the discrimination of authentic and credible multimodal data increasingly challenging. Synthetic data detection has thus garnered widespread attention, and the performance of large mu…

2025

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

NeurIPS 2025poster

We introduce ***Metis***, a foundation model for unified speech generation. Unlike previous task-specific or multi-task models, Metis follows a pre-training and fine-tuning paradigm. It is pre-trained on large-scale unlabeled speech data using masked generative modeling and then fine-tuned to adapt…

Cited by 0SourcecodeScholar
2025

TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling

NeurIPS 2025poster

Speech tokenizers serve as foundational components for speech language models, yet current designs exhibit several limitations, including: (1) dependence on multi-layer residual vector quantization structures or high frame rates, (2) reliance on auxiliary pre-trained models for semantic distillatio…

Cited by 0SourcecodeScholar