← Search

Wankou Yang

23 accepted papers

2026

DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation

CVPR 2026

Referring video object segmentation (RVOS) aims to segment objects within a video according to natural language expressions. Unlike earlier works focusing on static single-object scenarios, recent studies address more complex motion scenes. Previous methods typically adopt a query-based, logically m

Cited by 0SourceScholar
2026

EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens

AAAI 2026technical

Efficient Multimodal Large Language Models (MLLMs) compress vision tokens to reduce resource consumption, but the loss of visual information can degrade comprehension capabilities. Although some priors introduce Knowledge Distillation to enhance student models, they overlook the fundamental differen

Cited by 0SourcePDFScholar
2026

Towards Streaming Referring Video Segmentation via Large Language Model

CVPR 2026

Current referring video segmentation methods typically operate in an offline manner, where sparse frames are first selected for image-level referring segmentation, and the resulting masks are then propagated across the video. Although video sampling captures global context, its isolated processing s

Cited by 0SourcecodeScholar
2026

VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

ICML 2026poster

Reasoning Video Object Segmentation (RVOS) demands a sophisticated integration of temporal dynamics, spatial details, and linguistic reasoning to achieve precise pixel-level localization. Existing methods are limited to reasoning over fixed initial inputs and lack the capacity to actively acquire fu…

Cited by 0SourceScholar
2025

DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy

ICCV 2025poster

Referring Image Segmentation (RIS) is a challenging task that aims to segment objects in an image based on natural language expressions. While prior studies have predominantly concentrated on improving vision-language interactions and achieving fine-grained localization, a systematic analysis of the…

2025

Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving

ICRA 2025

The autoregressive world model exhibits robust generalization capabilities in vectorized scene understanding but encounters difficulties in deriving actions due to insufficient uncertainty modeling and self-delusion. In this paper, we explore the feasibility of deriving decisions from an autoregres-

Cited by 7SourcecodeScholar
2025

Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints

AAAI 2025technical

Multi-task visual grounding involves the simultaneous execution of localization and segmentation in images based on textual expressions. The majority of advanced methods predominantly focus on transformer-based multimodal fusion, aiming to extract robust multimodal representations. However, ambiguit…

2025

PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination

ICCV 2025poster

Recent advances in visual grounding have largely shifted away from traditional proposal-based two-stage frameworks due to their inefficiency and high computational complexity, favoring end-to-end direct reference paradigms. However, these methods rely exclusively on the referred target for supervisi…

2025

SpatialBot: Precise Spatial Understanding with Vision Language Models

ICRA 2025

Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding; however, they still struggle with spatial understanding, which is fundamental to embodied AI. In this paper, we propose SpatialBot, a model designed to enhance spatial understanding by utilizing both RGB an

Cited by 167SourcecodeScholar
2024

SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion

NeurIPS 2024poster

Visual grounding is a common vision task that involves grounding descriptive sentences to the corresponding regions of an image. Most existing methods use independent image-text encoding and apply complex hand-crafted modules or encoder-decoder architectures for modal interaction and query reasoning…

2023

Capturing the Motion of Every Joint: 3D Human Pose and Shape Estimation with Independent Tokens

ICLR 2023top-25%

In this paper we present a novel method to estimate 3D human pose and shape from monocular videos. This task requires directly recovering pixel-alignment 3D human pose and body shape from monocular images or videos, which is challenging due to its inherent ambiguity. To improve precision, existing m…

2023

GPA-3D: Geometry-aware Prototype Alignment for Unsupervised Domain Adaptive 3D Object Detection from Point Clouds

ICCV 2023poster

LiDAR-based 3D detection has made great progress in recent years. However, the performance of 3D detectors is considerably limited when deployed in unseen environments, owing to the severe domain gap problem. Existing domain adaptive 3D detection methods do not adequately consider the problem of the…

Cited by 15PDFcodeScholar
2023

LiftedCL: Lifting Contrastive Learning for Human-Centric Perception

ICLR 2023poster

Human-centric perception targets for understanding human body pose, shape and segmentation. Pre-training the model on large-scale datasets and fine-tuning it on specific tasks has become a well-established paradigm in human-centric perception. Recently, self-supervised learning methods have re-inves…

Cited by 9SourcePDFScholar
2023

PTC-Net: Point-Wise Transformer With Sparse Convolution Network for Place Recognition

RA-L 2023

In the point-cloud-based place recognition area, the existing hybrid architectures combining both convolutional networks and transformers have shown promising performance. They mainly apply the voxel-wise transformer after the sparse convolution (SPConv). However, they can induce information loss by

Cited by 22SourcecodeScholar
2022

SimCC: A Simple Coordinate Classification Perspective for Human Pose Estimation

ECCV 2022poster

"The 2D heatmap-based approaches have dominated Human Pose Estimation (HPE) for years due to high performance. However, the long-standing quantization error problem in the 2D heatmap-based methods leads to several well-known drawbacks: 1) The performance for the low-resolution inputs is limited; 2)…

2021

TokenPose: Learning Keypoint Tokens for Human Pose Estimation

ICCV 2021poster

Human pose estimation deeply relies on visual clues and anatomical constraints between parts to locate keypoints. Most existing CNN-based methods do well in visual representation, however, lacking in the ability to explicitly learn the constraint relationships between keypoints. In this paper, we pr…

Cited by 386PDFcodeScholar
2017

Semantic Regularisation for Recurrent Image Annotation

CVPR 2017poster

The "CNN-RNN" design pattern is increasingly widely applied in a variety of image annotation tasks including multi-label classification and captioning. Existing models use the weakly semantic CNN hidden layer or its transform as the image embedding that provides the interface between the CNN and RN…

Cited by 136PDFScholar