← Search

Bo Miao

7 accepted papers

2026

Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation

ICLR 2026poster

Generating realistic 3D Human-Human Interaction (HHI) requires coherent modeling of the physical plausibility of the agents and their interaction semantics. Existing methods compress all motion information into a single latent representation, limiting their ability to capture fine-grained actions an…

Cited by 0SourcecodeScholar
2025

CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation

NeurIPS 2025poster

Outdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in this direction are constrained by the absence of publicly available, well-annotated datasets. We introduce SketchSem3D, th…

Cited by 0SourceScholar
2025

Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval

IJCAI 2025

Current text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieve

Cited by 0SourcePDFScholar
2024

External Knowledge Enhanced 3D Scene Generation from Sketch

ECCV 2024poster

"Generating realistic 3D scenes is challenging due to the complexity of room layouts and object geometries. We propose a sketch based knowledge enhanced diffusion architecture (SEK) for generating customized, diverse, and plausible 3D scenes. SEK conditions the denoising process with a hand-drawn sk…

Cited by 6SourcePDFScholar
2024

Referring Human Pose and Mask Estimation In the Wild

NeurIPS 2024poster

We introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to pr…

2023

Spectrum-guided Multi-granularity Referring Video Object Segmentation

ICCV 2023poster

Current referring video object segmentation (R-VOS) techniques extract conditional kernels from encoded (low-resolution) vision-language features to segment the decoded high-resolution features. We discovered that this causes significant feature drift, which the segmentation kernels struggle to perc…

Cited by 54PDFcodeScholar
2021

Object-to-Scene: Learning to Transfer Object Knowledge to Indoor Scene Recognition

IROS 2021poster

Accurate perception of the surrounding scene is helpful for robots to make reasonable judgments and behaviours. Therefore, developing effective scene representation and recognition methods are of significant importance in robotics. Currently, a large body of research focuses on developing novel auxi…

Cited by 34SourcecodeScholar