← Search

Fan Fan

6 accepted papers

2026

Multi-Metric Preference Alignment for Generative Speech Restoration

AAAI 2026technical

Recent generative models have significantly advanced speech restoration tasks, yet their training objectives often misalign with human perceptual preferences, resulting in suboptimal quality. While post-training alignment has proven effective in other generative domains like text and image generatio

Cited by 0SourcePDFScholar
2025

Multimodal Image Matching Based on Cross-Modality Completion Pre-training

IJCAI 2025

The differences in imaging devices cause multimodal images to have modal differences and geometric distortions, complicating the matching task. Deep learning-based matching methods struggle with multimodal images due to the lack of large annotated multimodal datasets. To address these challenges, we

Cited by 0SourcePDFScholar
2025

Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features

ICASSP 2025accepted

Melody preservation is crucial in singing voice conversion (SVC). However, in many scenarios, audio is often accompanied with background music (BGM), which can cause audio distortion and interfere with the extraction of melody and other key features, significantly degrading SVC performance. Previous…

Cited by 0SourceScholar
2024

Multi-View Midivae: Fusing Track- and Bar-View Representations for Long Multi-Track Symbolic Music Generation

ICASSP 2024accepted

Variational Autoencoders (VAEs) constitute a crucial component of neural symbolic music generation, among which some works have yielded outstanding results and attracted considerable attention. Nevertheless, previous VAEs still encounter issues with overly long feature sequences and generated result…

Cited by 0SourceScholar
2021

Unsupervised Stacked Capsule Autoencoder for Hyperspectral Image Classification

ICASSP 2021accepted

Since CapsNet [1] shattered all previous records of algorithms for image recognition, the capsule's conception has attracted bright attention. It interprets an object by the geometrical arrangement of parts. We think it can be transferred to hyperspectral images. In a hyperspectral data cube, each p…

Cited by 0SourceScholar