← Search

Feilong Bao

16 accepted papers

2025

A High-Precision Character Cartoon Style Transfer Method Based on VToonify and Diffusion Models

ICASSP 2025accepted

In recent years, the technology for converting character images into cartoons has gained widespread attention on short video platforms, with diffusion models particularly standing out in the field of image style transfer. However, existing methods still have shortcomings in detail handling and chara…

Cited by 0SourceScholar
2025

Distance-Adaptive Quaternion Knowledge Graph Embedding with Bidirectional Rotation

COLING 2025main

Quaternion contains one real part and three imaginary parts, which provided a more expressive hypercomplex space for learning knowledge graph. Existing quaternion embedding models measure the plausibility of a triplet either through semantic matching or distance scoring functions. However, it appear…

2025

Multilingual Parameter-Sharing Adapters: A Method for Optimizing Low-Resource Neural Machine Translation

ICASSP 2025accepted

Adapter-based Multilingual Neural Machine Translation (MNMT) has become a significant approach in low-resource language translation by mitigating data imbalances between high-resource and low-resource language pairs and reducing training costs. However, existing adapter-based methods lack generaliza…

Cited by 0SourceScholar
2025

OTMEA : Multi-modal Entity Alignment via Optimal Transport

ICASSP 2025accepted

Multi-modal Entity Alignment (MMEA) aims to identify the same entities exhibited in different knowledge graphs (KGs), where the entities are enriched by structure and visual information. Existing MMEA methods learn multi-modal joint entity embeddings by encompassing both modality interaction and mod…

Cited by 0SourceScholar
2025

Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience

EMNLP 2025

Large language models (LLMs) generate human-aligned content under certain safety constraints. However, the current known technique “jailbreak prompt” can circumvent safety-aligned measures and induce LLMs to output malicious content. Research on Jailbreaking can help identify vulnerabilities in LLMs

Cited by 0SourcePDFScholar
2025

Structural-Aware Disentangled Learning with CLIP for Hyperbolic Zero-Shot Sketch-Based Image Retrieval

ICASSP 2025accepted

The zero-shot sketch-based image retrieval task faces two key challenges: domain gap and knowledge transfer. Our innovation is recognizing that directly aligning cross-domain features weakens the discriminative ability of the model, as it overlooks the asymmetry between sketches and images. Addition…

Cited by 0SourceScholar
2025

Unifying Dual-Space Embedding for Entity Alignment via Contrastive Learning

COLING 2025main

Entity alignment (EA) aims to match identical entities across different knowledge graphs (KGs). Graph neural network-based entity alignment methods have achieved promising results in Euclidean space. However, KGs often contain complex local and hierarchical structures, which are hard to represent in…

2024

Lˆ2GC:Lorentzian Linear Graph Convolutional Networks for Node Classification

COLING 2024main

Linear Graph Convolutional Networks (GCNs) are used to classify the node in the graph data. However, we note that most existing linear GCN models perform neural network operations in Euclidean space, which do not explicitly capture the tree-like hierarchical structure exhibited in real-world dataset…

2022

Alignment-Learning Based Single-Step Decoding for Accurate and Fast Non-Autoregressive Speech Recognition

ICASSP 2022accepted

Non-autoregressive transformer (NAT) based speech recognition models have gained more and more attention since they perform faster inference speed compared with autoregressive counterparts, especially when the single-step decoding is applied. However, the single-step decoding process with length pre…

Cited by 0SourceScholar
2021

Joint Alignment Learning-Attention Based Model for Grapheme-to-Phoneme Conversion

ICASSP 2021accepted

Sequence-to-sequence attention-based models for grapheme-to-phoneme (G2P) conversion have gained significant interests. The attention-based encoder-decoder framework learns the mapping of input to output tokens by selectively focusing on relevant information, and has been shown well performance. How…

Cited by 0SourceScholar
2020

A Multi-Scaled Receptive Field Learning Approach for Medical Image Segmentation

ICASSP 2020accepted

Biomedical image segmentation has been widely studied, and lots of methods have been proposed. Among these methods, attention U-Net has achieved a promising performance. However, it has drawbacks of extracting the multi-scaled receptive field features at the high-level feature maps, resulting in the…

Cited by 0SourceScholar
2020

Incorporating Inner-word and Out-word Features for Mongolian Morphological Segmentation

COLING 2020main

Mongolian morphological segmentation is regarded as a crucial preprocessing step in many Mongolian related NLP applications and has received extensive attention. Recently, end-to-end segmentation approaches with long short-term memory networks (LSTM) have achieved excellent results. However, the inn…

Cited by 1SourcePDFScholar
2020

Masking and Inpainting: A Two-Stage Speech Enhancement Approach for Low SNR and Non-Stationary Noise

ICASSP 2020accepted

Currently, low signal-to-noise ratio (SNR) and non-stationary noise cause severe performance degradation for most of speech enhancement models. For better speech enhancement at the above scenarios, this paper proposes a two-stage approach that consists of binary masking and spectrogram inpainting. I…

Cited by 0SourceScholar
2020

Teacher-Student Training For Robust Tacotron-Based TTS

ICASSP 2020accepted

While neural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways, the exposure bias problem in the autoregressive models remains an issue to be resolved. The exposure bias problem arises from the mismatch between the training and inference process, that resul…

Cited by 0SourceScholar