← Search

Xuying Zhang

7 accepted papers

2025

AR-1-to-3: Single Image to Consistent 3D Object via Next-View Prediction

ICCV 2025poster

Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views, especially when there is a significant camera pose difference, leading to poor-quality 3D geometries and textures. We…

2025

OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation

NeurIPS 2025poster

Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for multiple visual modalities remains unexplored. In this paper, we propose a novel multi-modal learning framework, termed…

Cited by 0SourcecodeScholar
2025

TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

ICCV 2025poster

We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQVAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate the multimodal unification and promising learning capabiliti…

2024

DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation

ICLR 2024poster

We present DFormer, a novel RGB-D pretraining framework to learn transferable representations for RGB-D segmentation tasks. DFormer has two new key innovations: 1) Unlike previous works that encode RGB-D information with RGB pretrained backbone, we pretrain the backbone using image-depth pairs from…

Cited by 58SourcePDFScholar
2024

TeMO: Towards Text-Driven 3D Stylization for Multi-Object Meshes

CVPR 2024poster

Recent progress in the text-driven 3D stylization of a single object has been considerably promoted by CLIP-based methods. However the stylization of multi-object 3D scenes is still impeded in that the image-text pairs used for pre-training CLIP mostly consist of an object. Meanwhile the local detai…

Cited by 9SourcePDFScholar
2022

DIFNet: Boosting Visual Information Flow for Image Captioning

CVPR 2022poster

Current Image captioning (IC) methods predict textual words sequentially based on the input visual information from the visual feature extractor and the partially generated sentence information. However, for most cases, the partially generated sentence may dominate the target word prediction due to…

Cited by 62PDFScholar
2021

RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual Words

CVPR 2021poster

Recent progress on visual question answering has explored the merits of grid features for vision language tasks. Meanwhile, transformer-based models have shown remarkable performance in various sequence prediction problems. However, the spatial information loss of grid features caused by flattening…

Cited by 286PDFcodeScholar