← Search

Xinliang Zhang

6 accepted papers

2026

Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context

CVPR 2026

Scene-consistent video generation aims to create videos that explore 3D scenes based on a camera trajectory. Previous methods rely on video generation models with external memory for consistency, or iterative 3D reconstruction and inpainting, which accumulate errors during inference due to incorrect

Cited by 0SourceScholar
2025

V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer

AAAI 2025technical

Concept Bottleneck Models (CBMs) offer inherent interpretability by initially translating images into human-comprehensible concepts, followed by a linear combination of these concepts for classification. However, the annotation of concepts for visual recognition tasks requires extensive expert knowl…

2024

Scribble Hides Class: Promoting Scribble-Based Weakly-Supervised Semantic Segmentation with Its Class Label

AAAI 2024technical

Scribble-based weakly-supervised semantic segmentation using sparse scribble supervision is gaining traction as it reduces annotation costs when compared to fully annotated alternatives. Existing methods primarily generate pseudo-labels by diffusing labeled pixels to unlabeled ones with local cues f…

2023

3D Implicit Transporter for Temporally Consistent Keypoint Discovery

ICCV 2023oral

Keypoint-based representation has proven advantageous in various visual and robotic tasks. However, the existing 2D and 3D methods for detecting keypoints mainly rely on geometric consistency to achieve spatial alignment, neglecting temporal consistency. To address this issue, the Transporter method…

Cited by 16PDFcodeScholar
2018

Compressed Sensing Mask Feature in Time-Frequency Domain for Civil Flight Radar Emitter Recognition

ICASSP 2018accepted

Specific emitter identification (SEI) is gaining popularity since it can distinguish different individuals in same type of radar emitter under complex electromagnetic environment. However, classification of signals is still a challenging task when the feature has low physical representation. In this…

Cited by 0SourceScholar