← Search

Xiaoning Zhang

5 accepted papers

2026

Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completion

CVPR 2026

Monocular Semantic Scene Completion (SSC) aims to reconstruct complete 3D semantic scenes from a single RGB image, offering a cost-effective solution for autonomous driving and robotics. However, the inherently imbalanced nature of voxel distributions--where over 93% of voxels are empty and foregrou

Cited by 0SourcecodeScholar
2025

VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible Regions

ICCV 2025poster

This paper introduces VisHall3D, a novel two-stage framework for monocular semantic scene completion that aims to address the issues of feature entanglement and geometric inconsistency prevalent in existing methods. VisHall3D decomposes the scene completion task into two stages: reconstructing the v…

2023

Layer-wise Fusion with Modality Independence Modeling for Multi-modal Emotion Recognition

ACL 2023long

Multi-modal emotion recognition has gained increasing attention in recent years due to its widespread applications and the advances in multi-modal learning approaches. However, previous studies primarily focus on developing models that exploit the unification of multiple modalities. In this paper, w…

2020

Adaptively Multi-Objective Adversarial Training for Dialogue Generation

IJCAI 2020poster

Naive neural dialogue generation models tend to produce repetitive and dull utterances. The promising adversarial models train the generator against a well-designed discriminator to push it to improve towards the expected direction. However, assessing dialogues requires consideration of many aspects…

Cited by 0SourcePDFScholar
2018

Progressive Attention Guided Recurrent Network for Salient Object Detection

CVPR 2018poster

Effective convolutional features play an important role in saliency estimation but how to learn powerful features for saliency is still a challenging task. FCN-based methods directly apply multi-level convolutional features without distinction, which leads to sub-optimal results due to the distracti…

Cited by 756SourcePDFScholar