← Search

Xiulong Liu

11 accepted papers

2026

GATCL: An Adaptive Contrastive Learning Framework Based on MHGAT for Spatial Domain Identification in Spatial Transcriptomics

AAAI 2026technical

Recent advances in spatial transcriptomics have enabled the simultaneous measurement of gene expression profiles and spatial location information, offering a more comprehensive and in-depth view for studying the tissue microenvironment. Spatial domain identification is a crucial step in analyzing sp

Cited by 0SourcePDFScholar
2026

SAMGTD: Spatial-Aware Masked Graph Transformer-Diffusion Model for Enhanced Cell Type Deconvolution in Spatial Transcriptomics

AAAI 2026technical

Recent advances in spatial transcriptomics have enabled the integration of gene expression profiles with precise spatial coordinates, which have facilitated the exploration of tumor occurrence and development mechanisms, as well as the development of more effective targeted and immunotherapy approac

Cited by 0SourcePDFScholar
2025

Hearing Anywhere in Any Environment

CVPR 2025poster

In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR) estimation, most existing methods are limited to the single environ…

Cited by 0SourcePDFScholar
2025

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

NeurIPS 2025oral

3D spatial reasoning in dynamic, audio-visual environments is a cornerstone of human cognition yet remains largely unexplored by existing Audio-Visual Large Language Models (AV-LLMs) and benchmarks, which predominantly focus on static or 2D scenes. We introduce SAVVY-Bench, the first benchmark for 3…

Cited by 0SourceScholar
2024

CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy Environments

AAAI 2024technical

Audio-visual navigation of an agent towards locating an audio goal is a challenging task especially when the audio is sporadic or the environment is noisy. In this paper, we present CAVEN, a Conversation-based Audio-Visual Embodied Navigation framework in which the agent may interact with a human/o…

Cited by 5SourcePDFScholar
2024

From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

ICML 2024poster

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interplay between audio and visual elements. Previous studies of audio-visual modalitie…

2024

MuseChat: A Conversational Music Recommendation System for Videos

CVPR 2024highlight

Music recommendation for videos attracts growing interest in multi-modal research. However existing systems focus primarily on content compatibility often ignoring the users' preferences. Their inability to interact with users for further refinements or to provide explanations leads to a less satisf…

2024

Tell What You Hear From What You See - Video to Audio Generation Through Text

NeurIPS 2024poster

The content of visual and audio scenes is multi-faceted such that a video stream can be paired with various audio streams and vice-versa. Thereby, in video-to-audio generation task, it is imperative to introduce steering approaches for controlling the generated audio. While Video-to-Audio generation…

2021

How Does it Sound?

NeurIPS 2021poster

One of the primary purposes of video is to capture people and their unique activities. It is often the case that the experience of watching the video can be enhanced by adding a musical soundtrack that is in-sync with the rhythmic features of these activities. How would this soundtrack sound? Such a…

Cited by 42SourcePDFScholar