← Search

Ning Jiang

26 accepted papers

2026

Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question Answering

AAAI 2026technical

Multi-hop question answering (MHQA) requires integrating knowledge scattered across multiple passages to derive the correct answer. Traditional retrieval-augmented generation (RAG) methods primarily focus on coarse-grained textual semantic similarity and ignore structural associations among disperse

Cited by 0SourcePDFScholar
2026

KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls

ICML 2026poster

Large Multimodal Models encode extensive factual knowledge in their pre-trained weights. However, its knowledge remains static and limited, unable to keep pace with real-world developments, which hinders continuous knowledge acquisition. Effective knowledge injection thus becomes critical, involving…

Cited by 0SourceScholar
2026

Massive Memorization with Hundreds of Trillions of Parameters for Sequential Transducer Generative Recommenders

ICLR 2026poster

Modern large-scale recommendation systems rely heavily on user interaction history sequences to enhance the model performance. The advent of large language models and sequential modeling techniques, particularly transformer architectures, has led to significant advancements (e.g., HSTU, SIM, and TW…

Cited by 0SourcecodeScholar
2026

When Large Multimodal Models Confront Evolving Knowledge: Challenges and Explorations

ICLR 2026poster

Large Multimodal Models (LMMs) store vast amounts of pretrained knowledge but struggle to remain aligned with real-world updates, making it difficult to avoid capability degradation when acquiring evolving knowledge. Furthermore, most current work focuses on exploring static textual knowledge inject…

Cited by 0SourceScholar
2025

Device-aware Optical Adversarial Attack for a Portable Projector-camera System

ICASSP 2025accepted

Deep-learning-based face recognition (FR) systems are susceptible to adversarial examples in both digital and physical domains. Physical attacks present a greater threat to deployed systems as adversaries can easily access the input channel, allowing them to provide malicious inputs to impersonate a…

Cited by 0SourceScholar
2025

Improving Multi-Position Training Performance on Reducing Limb Condition Effect in Wrist Myoelectric Control

RA-L 2025

Compared to the forearm, the wrist is suitable for the combination of myoelectric control with the popular wearable devices, enabling human machine interaction in an intuitive and effortless way. The change of limb condition is a common disturbance degrading the performance of wrist myoelectric cont

Cited by 2SourceScholar
2025

Learn from Balance: Rectifying Knowledge Transfer for Long-Tailed Scenarios

ICASSP 2025accepted

Knowledge Distillation (KD) transfers knowledge from a large pre-trained teacher network to a compact and efficient student network, making it suitable for deployment on resource-limited media terminals. However, traditional KD methods require balanced data to ensure robust training, which is often…

Cited by 0SourceScholar
2025

MS-UFAD: A Large-Scale Dataset for Real-world Unified Face Attack Detection with Text Descriptions

ICASSP 2025accepted

As deepfake and adversarial attacks evolve, facial recognition systems are encountering increasingly diverse threats. Most existing face liveness detection algorithms focus on single tasks, like spoofing or deepfake attack detection. The corresponding datasets have limited coverage of attack methods…

Cited by 0SourceScholar
2025

Realistic Real-Time Talking Head Synthesis with Grid Encoding and Progressive Conditioning

ICASSP 2025accepted

Dynamic NeRFs have recently been used for 3D talking portrait synthesis, but challenges remain in improving efficiency and effectiveness. We introduce R2-Talker, an efficient and effective framework for real-time talking head synthesis. Using multi-resolution hash grids, we losslessly encode facial…

Cited by 0SourceScholar
2024

Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation

COLING 2024main

Previous Sign Language Translation (SLT) methods achieve superior performance by relying on gloss annotations. However, labeling high-quality glosses is a labor-intensive task, which limits the further development of SLT. Although some approaches work towards gloss-free SLT through jointly training…

Cited by 14SourcePDFScholar
2024

Multi-Objective Progressive Clustering for Semi-Supervised Domain Adaptation in Speaker Verification

ICASSP 2024accepted

Utilizing the pseudo-labeling algorithm with large-scale unlabeled data becomes crucial for semi-supervised domain adaptation in speaker verification tasks. In this paper, we propose a novel pseudo-labeling method named Multi-objective Progressive Clustering (MoPC), specifically designed for semi-su…

Cited by 0SourceScholar
2024

SELM: Speech Enhancement using Discrete Tokens and Language Models

ICASSP 2024accepted

Language models (LMs) have recently shown superior performances in various speech generation tasks, demonstrating their powerful ability for semantic context modeling. Given the intrinsic similarity between speech generation and speech enhancement, harnessing semantic information is advantageous for…

Cited by 0SourceScholar
2024

Structured Optimal Brain Pruning for Large Language Models

EMNLP 2024main

The massive parameters and computational demands hinder the widespread application of Large Language Models (LLMs). Network pruning provides a practical solution to this problem. However, existing pruning works for LLMs mainly focus on unstructured pruning or necessitate post-pruning fine-tuning. Th…

Cited by 1SourcePDFScholar
2024

VL-FAS: Domain Generalization via Vision-Language Model For Face Anti-Spoofing

ICASSP 2024accepted

Recent approaches have demonstrated the effectiveness of Vision Transformer (ViT) with attention mechanisms for domain generalization of Face Anti-Spoofing (FAS). However, current attention algorithms highlight all the salient objects (e.g., background objects, hair, glasses), which results in the f…

Cited by 0SourceScholar
2024

Voxblink: A Large Scale Speaker Verification Dataset on Camera

ICASSP 2024accepted

In this paper, we introduce a large-scale and high-quality audiovisual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains 1.45M utterances from 38K speakers. Due to the inherent nature o…

Cited by 0SourceScholar
2023

MADI: Inter-Domain Matching and Intra-Domain Discrimination for Cross-Domain Speech Recognition

ICASSP 2023accepted

End-to-end automatic speech recognition (ASR) usually suffers from performance degradation when applied to a new domain due to domain shift. Unsupervised domain adaptation (UDA) aims to improve the performance on the unlabeled target domain by transferring knowledge from the source to the target dom…

Cited by 0SourceScholar
2022

MoSE: Modality Split and Ensemble for Multimodal Knowledge Graph Completion

EMNLP 2022main

Multimodal knowledge graph completion (MKGC) aims to predict missing entities in MKGs. Previous works usually share relation representation across modalities. This results in mutual interference between modalities during training, since for a pair of entities, the relation from one modality probably…

2022

Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension

ACL 2022long

Procedural Multimodal Documents (PMDs) organize textual instructions and corresponding images step by step. Comprehending PMDs and inducing their representations for the downstream reasoning tasks is designated as Procedural MultiModal Machine Comprehension (M3C). In this study, we approach Procedur…

2022

Overcoming Language Priors in Visual Question Answering via Distinguishing Superficially Similar Instances

COLING 2022main

Despite the great progress of Visual Question Answering (VQA), current VQA models heavily rely on the superficial correlation between the question type and its corresponding frequent answers (i.e., language priors) to make predictions, without really understanding the input. In this work, we define…

2022

PM2F2N: Patient Multi-view Multi-modal Feature Fusion Networks for Clinical Outcome Prediction

EMNLP 2022finding

Clinical outcome prediction is critical to the condition prediction of patients and management of hospital capacities. There are two kinds of medical data, including time series signals recorded by various devices and clinical notes in electronic health records (EHR), which are used for two common p…

2022

PPDL: Predicate Probability Distribution Based Loss for Unbiased Scene Graph Generation

CVPR 2022poster

Scene Graph Generation (SGG) has attracted more and more attention from visual researchers in recent years, since Scene Graph (SG) is valuable in many downstream tasks due to its rich structural-semantic details. However, the application value of SG on downstream tasks is severely limited by the pre…

Cited by 74PDFScholar
2021

Drawgan: Text to Image Synthesis with Drawing Generative Adversarial Networks

ICASSP 2021accepted

In this paper, we propose a novel drawing generative adversarial networks (DrawGAN) for text-to-image synthesis. The whole model divides the image synthesis into three stages by imitating the process of drawing. The first stage synthesizes the simple contour image based on the text description, the…

Cited by 0SourceScholar
2021

GMH: A General Multi-hop Reasoning Model for KG Completion

EMNLP 2021main

Knowledge graphs are essential for numerous downstream natural language processing applications, but are typically incomplete with many facts missing. This results in research efforts on multi-hop reasoning task, which can be formulated as a search process and current models typically perform short…

Cited by 17SourcePDFScholar
2021

TEMP: Taxonomy Expansion with Dynamic Margin Loss through Taxonomy-Paths

EMNLP 2021main

As an essential form of knowledge representation, taxonomies are widely used in various downstream natural language processing tasks. However, with the continuously rising of new concepts, many existing taxonomies are unable to maintain coverage by manual expansion. In this paper, we propose TEMP, a…