← Search

Jia-Ching Wang

22 accepted papers

2025

A Key to Effective Multi-task Learning: Separate Query Selection for Task-Synergized Handling and Node Utilization

ICASSP 2025accepted

In the realm of computer vision, effectively handling multi-tasks simultaneously presents a challenge that necessitates innovative solutions. To better address multiple vision problems, we introduce SeTano, an integrated Graph Neural Network (GNN)-based framework. This framework comprises a Dynamic…

Cited by 0SourceScholar
2025

HistoFS: Non-IID Histopathologic Whole Slide Image Classification via Federated Style Transfer with RoI-Preserving

CVPR 2025poster

Federated learning for pathological whole slide image (WSI) classification allows multiple clients to train a global multiple instance learning (MIL) model without sharing their privacy-sensitive WSIs. To accommodate the non-independent and identically distributed (non-i.i.d.) feature shifts, cross-…

2025

Impact of Glyph Information on Latent Space Diffusion Models for Accurate Handwritten Text Generation

ICASSP 2025accepted

The generation of high-quality stylized handwritten text images is a challenging task in computer vision and artificial intelligence. While advanced approaches using Latent Diffusion Models (LDMs) for generating stylized handwritten text have shown effectiveness, they often struggle with maintaining…

Cited by 0SourceScholar
2025

Mixture of Ordered Scoring Experts for Cross-prompt Essay Trait Scoring

ACL 2025long

Automated Essay Scoring (AES) plays a crucial role in language assessment. In particular, cross-prompt essay trait scoring provides learners with valuable feedback to improve their writing skills. However, due to the scarcity of prompts, most existing methods overlook critical information, such as c…

2023

CNEG-VC: Contrastive Learning Using Hard Negative Example In Non-Parallel Voice Conversion

ICASSP 2023accepted

Contrastive learning has advantages for non-parallel voice conversion, but the previous conversion results could be better and more preserved. In previous techniques, negative samples were randomly selected in the features vector from different locations. A positive example could not be effectively…

Cited by 0SourceScholar
2023

Code-Switching Speech Synthesis Based on Self-Supervised Learning and Domain Adaptive Speaker Encoder

ICASSP 2023accepted

Recently, end-to-end speech synthesis models based on deep learning have made great progress in speech quality, and gradually replaced traditional speech synthesis methods into the mainstream. However, these methods are still challenging to synthesize highly natural speech. In order to solve the abo…

Cited by 0SourceScholar
2023

Dense Adversarial Transfer Learning Based On Class-Invariance

ICASSP 2023accepted

This work proposes the dense adversarial transfer learning based on class-invariance, which is a novel, unsupervised, conditional adversarial domain adaptation approach. The proposed framework concatenates feature maps from the last layer of each backbone’s block to improve transfer learning; these…

Cited by 0SourceScholar
2023

Discriminative Vector Learning with Application to Single Channel Speech Separation

ICASSP 2023accepted

In this paper, we introduce a discriminative vector learning method and apply it to single-channel speech separation. First, speech samples are transformed into discriminative vectors using two backbone networks. These vectors are easily separated by simple clustering algorithms. Among them, vectors…

Cited by 0SourceScholar
2022

Selective Mutual Learning: An Efficient Approach for Single Channel Speech Separation

ICASSP 2022accepted

Mutual learning, the related idea to knowledge distillation, is a group of untrained lightweight networks, which simultaneously learn and share knowledge to perform tasks together during training. In this paper, we propose a novel mutual learning approach, namely selective mutual learning. This is t…

Cited by 0SourceScholar
2019

Speaker Characterization Using TDNN-LSTM Based Speaker Embedding

ICASSP 2019accepted

In this paper we propose speaker characterization using time delay neural networks and long short-term memory neural networks (TDNN-LSTM) speaker embedding. Three types of front-end feature extraction are investigated to find good features for speaker embedding. Three kinds of data augmentation are…

Cited by 0SourceScholar
2018

Complex-Valued Gaussian Process Latent Variable Model for Phase-Incorporating Speech Enhancement

ICASSP 2018accepted

Traditional speech enhancement techniques modify the magnitude of a speech in time-frequency domain, and use the phase of a noisy speech to resynthesize a time domain speech. This work proposes a complex-valued Gaussian process latent variable model (CGPLVM) to enhance directly the complex-valued no…

Cited by 0SourceScholar
2018

Image Representation Using Supervised and Unsupervised Learning Methods on Complex Domain

ICASSP 2018accepted

Matrix factorization (MF) and its extensions have been intensively studied in computer vision and machine learning. In this paper, unsupervised and supervised learning methods based on MF technique on complex domain are introduced. Projective complex matrix factorization (PCMF) and discriminant proj…

Cited by 0SourceScholar
2018

Locality-Preserving Complex-Valued Gaussian Process Latent Variable Model for Robust Face Recognition

ICASSP 2018accepted

Learning a low-dimensional image representation yields effective and efficient face recognition. The use of such a representation helps to weaken the curse of dimensionality. However, the traditional facial representation method is not robust against partial occlusions or variations of expression. T…

Cited by 0SourceScholar
2017

Dynamic tracking attention model for action recognition

ICASSP 2017accepted

This paper proposes a dynamic tracking attention model (DTAM), which mainly comprises a motion attention mechanism, a convolutional neural network (CNN) and long short-term memory (LSTM), to recognize human action in a video sequence. In the motion attention mechanism, the local dynamic tracking is…

Cited by 0SourceScholar
2017

Exemplar-embed complex matrix factorization for facial expression recognition

ICASSP 2017accepted

This paper presents an image representation approach which is based on matrix factorization in the complex domain and called exemplar-embed complex matrix factorization (EE-CMF). The proposed EE-CMF approach can very effectively improve the performance of facial expression recognition. Moreover, Wir…

Cited by 0SourceScholar
2017

Fully complex deep neural network for phase-incorporating monaural source separation

ICASSP 2017accepted

Deep neural network (DNN) have become a popular means of separating a target source from a mixed signal. Most of DNN-based methods modify only the magnitude spectrum of the mixture. The phase spectrum is left unchanged, which is inherent in the short-time Fourier transform (STFT) coefficients of the…

Cited by 0SourceScholar
2017

Hierarchical joint-guided networks for semantic image segmentation

ICASSP 2017accepted

Semantic image segmentation is now an exciting area of research owing to its various useful applications in daily life. This paper introduces a hierarchical joint-guided network (HJGN) which is mainly composed of proposed hierarchical joint learning convolutional networks (HJLCNs) and proposed joint…

Cited by 0SourceScholar
2017

Kernel weighted Fisher sparse analysis on multiple maps for audio event recognition

ICASSP 2017accepted

This work presents a novel approach for audio event recognition. The approach develops a weighted kernel fisher sparse analysis method based on multiple maps. The proposed method consists of maps extraction and kernel weighted Fisher sparse analysis. Two maps are firstly extracted from each audio fi…

Cited by 0SourceScholar