← Search

Xinrong Hu

13 accepted papers

2026

MACRec: A Multi-View Subspace Alignment Framework for Contrastive Sampling Calibration in Recommendation

AAAI 2026technical

Graph Contrastive Learning (GCL) has proven effective in mitigating data sparsity and enhancing representation learning for recommendation. Yet, most GCL frameworks indiscriminately treat all non-anchor nodes as negatives during contrastive sampling, often leading to the false negative problem where

Cited by 0SourcePDFScholar
2025

DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and Actions

ACL 2025long

In this paper, we propose contextualized and situated text-to-speech (CS-TTS), a novel TTS task to promote more accurate and customized speech generation using prompts with Dialogues, Narratives, and Actions (DNA). While prompt-based TTS methods facilitate controllable speech generation, existing TT…

2025

Importance-Awareness Masking Network for Robust Document Retrieval

ICASSP 2025accepted

In this paper, we introduce the IMPortance-awaReness maskIng NeTwork (IMPRINT), a novel approach to enhance the robustness of document retrieval systems against query variations, particularly those containing misspellings. Unlike previous models that treat all query components (words/features) equal…

Cited by 0SourceScholar
2024

Masking the Unknown: Leveraging Masked Samples for Enhanced Data Augmentation

UAI 2024poster

Data Augmentation (DA) has become a widely adopted strategy for addressing data scarcity in numerous NLP tasks, especially in scenarios with limited resources or imbalanced classes. However, many existing augmentation techniques rely on randomness or additional resources, presenting challenges in bo…

Cited by 0SourcePDFScholar
2024

Open-Vocabulary RGB-Thermal Semantic Segmentation

ECCV 2024poster

"RGB-Thermal (RGB-T) semantic segmentation is an important research branch of multi-modal image segmentation. The current RGB-T semantic segmentation methods generally have two unsolved and typical shortcomings. First, they do not have the open-vocabulary recognition ability, which significantly lim…

2024

PMDI: Combining Parametric-Model and Depth-Aware Implicit Function for Single-View Human Reconstruction

ICASSP 2024accepted

3D human reconstruction from a single image has achieved great progress with recent deep neural networks. However, conventional approaches still struggle with the issues of over-smoothing details and wrong limb poses. To this end, we propose PMDI, a method that combines parametric-model and depth-aw…

Cited by 0SourceScholar
2024

SGM: A Dataset for 3D Garment Reconstruction from Single Hand-Drawn Sketch

ICASSP 2024accepted

High-fidelity garment reconstruction is essential for various applications such as garment design and virtual try-on. While image-based reconstruction methods have made significant progress with deep generative models, generating 3D models from hand-drawn sketches to meet design intentions remains c…

Cited by 0SourceScholar
2022

Multi-Pose Virtual Try-On Via Self-Adaptive Feature Filtering

ICASSP 2022accepted

With the growing trend of virtual try-on, multi-pose tasks attract researchers due to their higher commercial value. Prior methods lack an effective geometric deformation to maintain the original image details resulting in many details loss in the head and garment. To address this problem, we propos…

Cited by 0SourceScholar
2022

Realistic Monocular-To-3d Virtual Try-On Via Multi-Scale Characteristics Capture

ICASSP 2022accepted

3D virtual try-on receives widespread attention from scholars due to its great practical and commercial values. In prior methods, the fundamental problems lie in the limitations on texture retention during garment deformation and the lack of feature context capture during depth estimation. To addres…

Cited by 0SourceScholar
2022

Seeing the wood for the trees: a contrastive regularization method for the low-resource Knowledge Base Question Answering

NAACL 2022findings

Given a context knowledge base (KB) and a corresponding question, the Knowledge Base Question Answering task aims to retrieve correct answer entities from this KB. Despite sophisticated retrieval algorithms, the impact of the low-resource (incomplete) KB is not fully exploited, where contributing co…

2021

A Triplet Appearance Parsing Network for Person Re-Identification

ICASSP 2021accepted

As one of the specific vision tasks, person re-identification has become a prevalent research topic in the field of multimedia and computer vision. However, existing feature extraction methods, originating from the quality of the bounding boxes which could cause the inhomogeneity and incoherence of…

Cited by 0SourceScholar
2021

DP-VTON: Toward Detail-Preserving Image-Based Virtual Try-on Network

ICASSP 2021accepted

Image-based virtual try-on systems with the goal of transferring a target clothing item onto the corresponding region of a person have received great attention recently. However, it is still a challenge for the existing methods to generate photo-realistic try-on images while preserving non-target de…

Cited by 0SourceScholar