← Search

Yanyu Xu

19 accepted papers

2026

Cancer Survival Prediction by Cyclic Generation and Multi-grained Alignment

AAAI 2026technical

Cancer survival analysis with multimodal data is crucial for precise treatments and patient benefits. However, the following challenges prohibit integrating histopathology and genomics: (i) multimodal data is not always complete, especially for the more costly genomics data; (ii) intricate interacti

Cited by 0SourcePDFScholar
2025

Aligning Contrastive Multiple Clusterings with User Interests

IJCAI 2025

Multiple clustering approaches aim to partition complex data in different ways. These methods often exhibit a one-to-many relationship in their results, and relying solely on the data context may be insufficient to capture the patterns relevant to the user. User’s expectation is key for the multiple

Cited by 0SourcePDFScholar
2025

Emergence-Inspired Multi-Granularity Causal Learning

AAAI 2025technical

Existing causal learning algorithms focus on micro-level causal discovery, confronting significant challenges in identifying the influence of macro systems, composed of micro-level variables, on other variables. This difficulty arises because the causal relationships in macro systems are often media…

Cited by 0SourcePDFScholar
2025

Semantic and Sequential Alignment for Referring Video Object Segmentation

CVPR 2025poster

Referring video object segmentation (RVOS) seeks to segment the objects within a video referred by linguistic expressions. Existing RVOS solutions follow a "fuse then select" paradigm: establishing semantic correlation between visual and linguistic feature, and performing frame-level query interacti…

2024

BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays

NeurIPS 2024poster

Medical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and facilitate adapting task-specific models to new setups using fe…

2024

MeshSegmenter: Zero-Shot Mesh Segmentation via Texture Synthesis

ECCV 2024poster

"We present MeshSegmenter, a simple yet effective framework designed for zero-shot 3D semantic segmentation. This model successfully extends the powerful capabilities of 2D segmentation models to 3D meshes, delivering accurate 3D segmentation across diverse meshes and segment descriptions. Specifica…

2023

Generative Gradient Inversion via Over-Parameterized Networks in Federated Learning

ICCV 2023poster

Federated learning has gained recognitions as a secure approach for safeguarding local private data in collaborative learning. But the advent of gradient inversion research has posed significant challenges to this premise by enabling a third-party to recover groundtruth images via gradients. While p…

Cited by 13PDFcodeScholar
2021

Accurate depth estimation from a hybrid event-RGB stereo setup

IROS 2021poster

Event-based visual perception is becoming increasingly popular owing to interesting sensor characteristics enabling the handling of difficult conditions such as highly dynamic motion or challenging illumination. The mostly complementary nature of event cameras however still means that best results a…

Cited by 11SourceScholar
2021

Amodal Segmentation Based on Visible Region Segmentation and Shape Prior

AAAI 2021technical

Almost all existing amodal segmentation methods make the inferences of occluded regions by using features corresponding to the whole image. This is against the human's amodal perception, where human uses the visible part and the shape prior knowledge of the target to infer the occluded region. To mi…

2021

Crowd Counting With Partial Annotations in an Image

ICCV 2021poster

To fully leverage the data captured from different scenes with different view angles while reducing the annotation cost, this paper studies a novel crowd counting setting, i.e. only using partial annotations in each image as training data. Inspired by the repetitive patterns in the annotated and una…

Cited by 57PDFcodeScholar
2021

Layout-Guided Novel View Synthesis From a Single Indoor Panorama

CVPR 2021poster

Existing view synthesis methods mainly focus on the perspective images and have shown promising results. However, due to the limited field-of-view of the pinhole camera, the performance quickly degrades when large camera movements are adopted. In this paper, we make the first attempt to generate nov…

Cited by 28PDFcodeScholar
2020

Geometric Structure Based and Regularized Depth Estimation From 360 Indoor Imagery

CVPR 2020poster

Motivated by the correlation between the depth and the geometric structure of a 360 indoor image, we propose a novel learning-based depth estimation framework that leverages the geometric structure of a scene to conduct depth estimation. Specifically, we represent the geometric structure of an indoo…

Cited by 86PDFScholar
2020

SIRI: Spatial Relation Induced Network For Spatial Description Resolution

NeurIPS 2020poster

Spatial Description Resolution, as a language-guided localization task, is proposed for target location in a panoramic street view, given corresponding language descriptions. Explicitly characterizing an object-level relationship while distilling spatial relationships are currently absent but crucia…

2019

PPGNet: Learning Point-Pair Graph for Line Segment Detection

CVPR 2019poster

In this paper, we present a novel framework to detect line segments in man-made environments. Specifically, we propose to describe junctions, line segments and relationships between them with a simple graph, which is more structured and informative than end-point representation used in existing line…

Cited by 110PDFcodeScholar
2018

Encoding Crowd Interaction With Deep Neural Network for Pedestrian Trajectory Prediction

CVPR 2018poster

Pedestrian trajectory prediction is a challenging task because of the complex nature of humans. In this paper, we tackle the problem within a deep learning framework by considering motion information of each pedestrian and its interaction with the crowd. Specifically, motivated by the residual learn…

Cited by 336SourcePDFScholar
2018

Gaze Prediction in Dynamic 360° Immersive Videos

CVPR 2018poster

This paper explores gaze prediction in dynamic $360^circ$ immersive videos, emph{i.e.}, based on the history scan path and VR contents, we predict where a viewer will look at an upcoming time. To tackle this problem, we first present the large-scale eye-tracking in dynamic VR scene dataset. Our data…