← Search

Xiaofeng Liu

38 accepted papers

2026

CoFact: Dynamic Coordination of Attention Heads for Improving Factual Consistency in LLMs

AAAI 2026technical

Large language models (LLMs) frequently generate fluent yet factually inaccurate content, a phenomenon known as hallucination. Recent inference-time approaches aim to improve truthfulness by steering model activations toward semantically meaningful directions. While effective to some extent, these m

Cited by 0SourcePDFScholar
2026

Complete Multi-Domain Decoupled Fusion Model for EEG-Based Person Identification

ICRA 2026poster

Electroencephalogram (EEG) signals have unique individual characteristics and have broad application prospects in identity authentication. At present, person identification (PI) based on EEG using the temporal-spatial-spectral feature extraction framework has achieved remarkable success. However, th…

Cited by 0codeScholar
2026

Explainable Depression Assessment from Face Videos by Weakly Supervised Learning

AAAI 2026technical

Existing video-based automatic depression assessment (ADA) approaches frequently achieve video-level depression assessment by aggregating features or predictions of individual frames or equal-length segments within the given video. While their performances have been largely enhanced by recent advanc

Cited by 0SourcePDFScholar
2026

GTFMN: Guided Texture and Feature Modulation Network for Low-Light Image Enhancement and Super-Resolution

ICASSP 2026oral

Low-light image super-resolution (LLSR) is a challenging task due to the coupled degradation of low resolution and poor illumination. To address this, we propose the Guided Texture and Feature Modulation Network (GTFMN), a novel framework that decouples the LLSR task into two sub-problems: illuminat…

Cited by 0SourcePDFScholar
2026

SI-IGCL: Subject Invariance-aware Inverse Graph Contrastive Learning for Psychiatric Disorder Identification

ICML 2026poster

Functional brain network analysis plays an important role in understanding and diagnosing psychiatric disorders. However, current methods struggle with subject variations, impairing the model’s generalization ability to the test set. To address this issue, we propose the Subject Invariance-aware Inv…

Cited by 0SourceScholar
2025

Domain Adaptive Diabetic Retinopathy Grading with Model Absence and Flowing Data

CVPR 2025poster

Domain shift (the difference between source and target domains) poses a significant challenge in clinical applications, e.g., Diabetic Retinopathy (DR) grading. Despite considering certain clinical requirements, like source data privacy, conventional transfer methods are predominantly model-centered…

2025

Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization

ACL 2025long

Video dubbing aims to translate original speech in visual media programs from the source language to the target language, relying on neural machine translation and text-to-speech technologies. Due to varying information densities across languages, target speech often mismatches the source speech dur…

2025

Hierarchical Multimodal Decoupling-Fusion Framework for offline Multiple Appropriate Facial Reaction Generation

ICASSP 2025accepted

Facial reactions convey crucial emotional information and coordinating interpersonal relationships in human dyadic interactions. While existing Multiple Appropriate Facial Reaction Generation (MAFRG) methods focus on generating multiple reasonable facial reactions, none of these approaches combines…

Cited by 0SourceScholar
2025

In-situ Value-aligned Human-Robot Interactions with Physical Constraints

IROS 2025

Equipped with Large Language Models (LLMs), human-centered robots are now capable of performing a wide range of tasks that were previously deemed challenging or unattainable. However, merely completing tasks is insufficient for cognitive robots, who should learn and apply human preferences to future

Cited by 0SourceScholar
2025

Progressive Compositionality in Text-to-Image Generative Models

ICLR 2025spotlight

Despite the impressive text-to-image (T2I) synthesis capabilities of diffusion models, they often struggle to understand compositional relationships between objects and attributes, especially in complex settings. Existing approaches through building compositional architectures or generating difficul…

2025

Rethinking Evaluation of Infrared Small Target Detection

NeurIPS 2025poster

As an essential vision task, infrared small target detection (IRSTD) has seen significant advancements through deep learning. However, critical limitations in current evaluation protocols impede further progress. First, existing methods rely on fragmented pixel- and target-level specific met…

Cited by 0SourceScholar
2025

Theoretical Insights into In-context Learning with Unlabeled Data

NeurIPS 2025poster

Recent research shows that in-context learning (ICL) can be effective even when demonstrations have missing or incorrect labels. To shed light on this capability, we examine a canonical setting where the demonstrations are drawn according to a binary Gaussian mixture model (GMM) and a certain fracti…

Cited by 0SourceScholar
2025

Towards Robust Deterministic and Probabilistic Modeling for Predictive Learning

IJCAI 2025

Predictive modeling of unannotated spatiotemporal data presents inherent challenges, primarily due to the highly entangled visual dynamics in real-world scenes. To tackle these complexities, we introduce a novel insight through Disentangling Deterministic and Probabilistic (DDP) modeling. We note a

Cited by 0SourcePDFScholar
2025

UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation

NeurIPS 2025poster

Multi-modal image segmentation faces real-world deployment challenges from incomplete/corrupted modalities degrading performance. While existing methods address training-inference modality gaps via specialized per-combination models, they introduce high deployment costs by requiring exhaustive mode…

Cited by 0SourcecodeScholar
2024

Label-Efficient Few-Shot Semantic Segmentation with Unsupervised Meta-Training

AAAI 2024technical

The goal of this paper is to alleviate the training cost for few-shot semantic segmentation (FSS) models. Despite that FSS in nature improves model generalization to new concepts using only a handful of test exemplars, it relies on strong supervision from a considerable amount of labeled training da…

2024

Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models

EMNLP 2024finding

Fine-tuning pre-trained Vision-Language Models (VLMs) has shown remarkable capabilities in medical image and textual depiction synergy. Nevertheless, many pre-training datasets are restricted by patient privacy concerns, potentially containing noise that can adversely affect downstream performance.…

Cited by 2SourcePDFScholar
2022

Cmri2spec: Cine MRI Sequence to Spectrogram Synthesis via A Pairwise Heterogeneous Translator

ICASSP 2022accepted

Multimodal representation learning using visual movements from cine magnetic resonance imaging (MRI) and their acoustics has shown great potential to learn shared representation and to predict one modality from another. Here, we propose a new synthesis framework to translate from cine MRI sequences…

Cited by 0SourceScholar
2021

Adversarial Unsupervised Domain Adaptation With Conditional and Label Shift: Infer, Align and Iterate

ICCV 2021poster

In this work, we propose an adversarial unsupervised domain adaptation (UDA) approach with the inherent conditional and label shifts, in which we aim to align the distributions w.r.t. both p(x|y) and p(y). Since the label is inaccessible in the target domain, the conventional adversarial UDA assumes…

Cited by 96PDFScholar
2021

BACO: A Background Knowledge- and Content-Based Framework for Citing Sentence Generation

ACL 2021long

In this paper, we focus on the problem of citing sentence generation, which entails generating a short text to capture the salient information in a cited paper and the connection between the citing and cited paper. We present BACO, a BAckground knowledge- and COntent-based framework for citing sente…

Cited by 41SourcePDFScholar
2021

Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models

AAAI 2021technical

AI Safety is a major concern in many deep learning applications such as autonomous driving. Given a trained deep learning model, an important natural problem is how to reliably verify the model's prediction. In this paper, we propose a novel framework --- deep verifier networks (DVN) to detect unrel…

Cited by 67SourcePDFScholar
2021

Domain Generalization under Conditional and Label Shifts via Variational Bayesian Inference

IJCAI 2021poster

In this work, we propose a domain generalization (DG) approach to learn on several labeled source domains and transfer knowledge to a target domain that is inaccessible in training. Considering the inherent conditional and label shifts, we would expect the alignment of p(x|y) and p(y). However, the…

Cited by 33SourcePDFScholar
2021

Embedding Semantic Hierarchy in Discrete Optimal Transport for Risk Minimization

ICASSP 2021accepted

The widely-used cross-entropy (CE) loss-based deep networks achieved significant progress w.r.t. the classification accuracy. However, the CE loss can essentially ignore the risk of misclassification which is usually measured by the distance between the prediction and label in a semantic hierarchica…

Cited by 0SourceScholar
2021

PNS: Population-Guided Novelty Search for Reinforcement Learning in Hard Exploration Environments

IROS 2021poster

Reinforcement Learning (RL) has made remarkable achievements, but it still suffers from inadequate exploration strategies, sparse reward signals, and deceptive reward functions. To alleviate these problems, a Population-guided Novelty Search (PNS) parallel learning method is proposed in this paper.…

Cited by 10SourceScholar
2021

Recursively Conditional Gaussian for Ordinal Unsupervised Domain Adaptation

ICCV 2021poster

The unsupervised domain adaptation (UDA) has been widely adopted to alleviate the data scalability issue, while the existing works usually focus on classifying independently discrete labels. However, in many tasks (e.g., medical diagnosis), the labels are discrete and successively distributed. The U…

Cited by 29PDFScholar
2021

Subtype-aware Unsupervised Domain Adaptation for Medical Diagnosis

AAAI 2021technical

Recent advances in unsupervised domain adaptation (UDA) show that transferable prototypical learning presents a powerful means for class conditional alignment, which encourages the closeness of cross-domain class centroids. However, the cross-domain inner-class compactness and the underlying fine-gr…

2020

AUTO3D: Novel view synthesis through unsupervisely learned variational viewpoint and global 3D representation

ECCV 2020poster

This paper targets on learning-based novel view synthesis from a single or limited 2D images without the pose supervision. In the viewer-centered coordinates, we construct an end-to-end trainable conditional variational framework to disentangle the unsupervisely learned relative-pose/rotation and im…

Cited by 26SourcePDFScholar
2020

High-Accuracy Classification of Attention Deficit Hyperactivity Disorder with L2, 1-Norm Linear Discriminant Analysis

ICASSP 2020accepted

Attention Deficit Hyperactivity Disorder (ADHD) is a high incidence of neurobehavioral disease in school-age children. Its neurobiological classification is meaningful for clinicians. The existing ADHD classification methods suffer from two problems, i.e., insufficient data and noise disturbance. He…

Cited by 0SourceScholar
2020

Rnn-Transducer with Stateless Prediction Network

ICASSP 2020accepted

The RNN-Transducer (RNNT) outperforms classic Automatic Speech Recognition (ASR) systems when a large amount of supervised training data is available. For low-resource languages, the RNNT models overfit, and can not directly take advantage of additional large text corpora as in classic ASR systems.W…

Cited by 0SourceScholar
2020

Severity-Aware Semantic Segmentation With Reinforced Wasserstein Training

CVPR 2020poster

Semantic segmentation is a class of methods to classify each pixel in an image into semantic classes, which is critical for autonomous vehicles and surgery systems. Cross-entropy (CE) loss-based deep neural networks (DNN) achieved great success w.r.t. the accuracy-based metrics, e.g., mean Intersect…

Cited by 35PDFScholar
2020

Sparse CSP Algorithm via Joint Spatio-Temporal Filtering

ICASSP 2020accepted

Common spatial pattern (CSP) is widely used in motor imagery classification tasks. Classical CSP depends only on spatial filters. To improve its performance, a novel and efficient spatio-temporal filtering strategy is proposed in this paper to extract discriminative features. Common temporal filters…

Cited by 0SourceScholar
2019

Feature-Level Frankenstein: Eliminating Variations for Discriminative Recognition

CVPR 2019poster

Recent successes of deep learning-based recognition rely on maintaining the content related to the main-task label. However, how to explicitly dispel the noisy signals for better generalization remains an open issue. We systematically summarize the detrimental factors as task-relevant/irrelevant sem…

Cited by 48PDFcodeScholar
2019

Permutation-Invariant Feature Restructuring for Correlation-Aware Image Set-Based Recognition

ICCV 2019poster

We consider the problem of comparing the similarity of image sets with variable-quantity, quality and un-ordered heterogeneous images. We use feature restructuring to exploit the correlations of both inner&inter-set images. Specifically, the residual self-attention can effectively restructure the fe…

Cited by 38PDFScholar
2018

Contextual-based Image Inpainting: Infer, Match, and Translate

ECCV 2018poster

We study the task of image inpainting, which is to fill in the missing region of an incomplete image with plausible contents. To this end, we propose a learning-based approach to generate visually coherent completion given a high-resolution image with missing components. In order to overcome the dif…

Cited by 346SourcePDFScholar
2018

Dependency-aware Attention Control for Unconstrained Face Recognition with Image Sets

ECCV 2018poster

This paper targets the problem of image set-based face verification and identification. Unlike traditional single media (an image or video) setting, we encounter a set of heterogeneous contents containing orderless images and videos. The importance of each image is usually considered either equal or…

Cited by 53SourcePDFScholar
2017

Large-scale audio event discovery in one million YouTube videos

ICASSP 2017accepted

Internet videos provide a virtually boundless source of audio with a conspicuous lack of localized annotations, presenting an ideal setting for unsupervised methods. With this motivation, we perform an unprecedented exploration into the large-scale discovery of recurring audio events in a diverse co…

Cited by 0SourceScholar