← Search

Wenbing Tao

26 accepted papers

2026

PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training

CVPR 2026

Open-Set Object Detection (OSOD) enables recognition of novel categories beyond fixed classes but faces challenges in aligning text representations with complex visual concepts and the scarcity of image-text pairs for rare categories. This results in suboptimal performance in specialized domains or

Cited by 0SourcecodeScholar
2026

VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range Consistency

ICML 2026spotlight

Despite recent progress in calibration-free monocular SLAM via 3D vision foundation models, scale drift remains severe on long sequences. Motion-agnostic partitioning breaks contextual coherence and causes zero-motion drift, while conventional geometric alignment is computationally expensive. To add…

Cited by 0SourceScholar
2025

Disentangling Instance and Scene Contexts for 3D Semantic Scene Completion

ICCV 2025poster

3D Semantic Scene Completion (SSC) has gained increasing attention due to its pivotal role in 3D perception. Recent advancements have primarily focused on refining voxel-level features to construct 3D scenes. However, treating voxels as the basic interaction units inherently limits the utilization o…

2025

FIELD: Fast Information-driven Autonomous Exploration using Larger Perception Distance

IROS 2025

Autonomous exploration is a critical challenge for various unmanned aerial vehicle (UAV) applications. Existing methods often suffer from low exploration rates due to limitations such as inefficient global coverage and inadequate sensor data utilization. In this paper, we introduce FIELD, a Fast Inf

Cited by 0SourceScholar
2025

High-Fidelity Lightweight Mesh Reconstruction from Point Clouds

CVPR 2025highlight

Recently, learning signed distance functions (SDFs) from point clouds has become popular for reconstruction. To ensure accuracy, most methods require using high-resolution Marching Cubes for surface extraction. However, this results in redundant mesh elements, making the mesh inconvenient to use. To…

Cited by 0SourcePDFScholar
2025

OVTR: End-to-End Open-Vocabulary Multiple Object Tracking with Transformer

ICLR 2025poster

Open-vocabulary multiple object tracking aims to generalize trackers to unseen categories during training, enabling their application across a variety of real-world scenarios. However, the existing open-vocabulary tracker is constrained by its framework structure, isolated frame-level perception, an…

2025

Perception-R1: Pioneering Perception Policy with Reinforcement Learning

NeurIPS 2025poster

Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, our initial experiments reveal that incorporating a thinking process through RL does not consistently lead to performance…

Cited by 0SourcecodeScholar
2025

Unhackable Temporal Reward for Scalable Video MLLMs

ICLR 2025poster

In the pursuit of superior video-processing MLLMs, we have encountered a perplexing paradox: the “anti-scaling law”, where more data and larger models lead to worse performance. This study unmasks the culprit: “temporal hacking”, a phenomenon where models shortcut by fixating on select frames, missi…

Cited by 0SourcePDFScholar
2024

A Comprehensive Survey and Taxonomy on Point Cloud Registration Based on Deep Learning

IJCAI 2024poster

Point cloud registration (PCR) involves determining a rigid transformation that aligns one point cloud to another. Despite the plethora of outstanding deep learning (DL)-based registration methods proposed, comprehensive and systematic studies on DL-based PCR techniques are still lacking. In this pa…

2024

Delving into the Trajectory Long-tail Distribution for Muti-object Tracking

CVPR 2024poster

Multiple Object Tracking (MOT) is a critical area within computer vision with a broad spectrum of practical implementations. Current research has primarily focused on the development of tracking algorithms and enhancement of post-processing techniques. Yet there has been a lack of thorough examinati…

2024

Few-shot NeRF by Adaptive Rendering Loss Regularization

ECCV 2024poster

"Novel view synthesis with sparse inputs poses great challenges to Neural Radiance Field (NeRF). Recent works demonstrate that the frequency regularization of Positional Encoding (PE) can achieve promising results for few-shot NeRF. In this work, we reveal that there exists an inconsistency between…

2024

IINet: Implicit Intra-inter Information Fusion for Real-Time Stereo Matching

AAAI 2024technical

Recently, there has been a growing interest in 3D CNN-based stereo matching methods due to their remarkable accuracy. However, the high complexity of 3D convolution makes it challenging to strike a balance between accuracy and speed. Notably, explicit 3D volumes contain considerable redundancy. In t…

Cited by 8SourcePDFScholar
2024

QTrack: Embracing Quality Clues for Robust 3D Multi-Object Tracking

IROS 2024poster

3D Multi-Object Tracking (MOT) has achieved tremendous achievement thanks to the rapid development of 3D object detection and 2D MOT. Recent advanced works generally employ a series of object attributes, e.g., position, size, velocity, and appearance, to provide the clues for the association in 3D M…

Cited by 1SourceScholar
2023

Generalizing Multiple Object Tracking to Unseen Domains by Introducing Natural Language Representation

AAAI 2023technical

Although existing multi-object tracking (MOT) algorithms have obtained competitive performance on various benchmarks, almost all of them train and validate models on the same domain. The domain generalization problem of MOT is hardly studied. To bridge this gap, we first draw the observation that th…

2022

DeTarNet: Decoupling Translation and Rotation by Siamese Network for Point Cloud Registration

AAAI 2022technical

Point cloud registration is a fundamental step for many tasks. In this paper, we propose a neural network named DetarNet to decouple the translation t and rotation R, so as to overcome the performance degradation due to their mutual interference in point cloud registration. First, a Siamese Network…

2022

Geo-Neus: Geometry-Consistent Neural Implicit Surfaces Learning for Multi-view Reconstruction

NeurIPS 2022accept

Recently, neural implicit surfaces learning by volume rendering has become popular for multi-view reconstruction. However, one key challenge remains: existing approaches lack explicit multi-view geometry constraints, hence usually fail to generate geometry-consistent surface reconstruction. To addre…

Cited by 294SourcePDFScholar
2022

One-Inlier is First: Towards Efficient Position Encoding for Point Cloud Registration

NeurIPS 2022accept

Transformer architecture has shown great potential for many visual tasks, including point cloud registration. As an order-aware module, position encoding plays an important role in Transformer architecture applied to point cloud registration task. In this paper, we propose OIF-PCR, a one-inlier base…

Cited by 35SourcePDFScholar
2022

SC2-PCR: A Second Order Spatial Compatibility for Efficient and Robust Point Cloud Registration

CVPR 2022poster

In this paper, we present a second order spatial compatibility (SC^2) measure based method for efficient and robust point cloud registration (PCR), called SC^2-PCR. Firstly, we propose a second order spatial compatibility (SC^2) measure to compute the similarity between correspondences. It considers…

Cited by 166PDFcodeScholar
2021

Cascade Network with Guided Loss and Hybrid Attention for Finding Good Correspondences

AAAI 2021technical

Finding good correspondences is a critical prerequisite in many feature based tasks. Given a putative correspondence set of an image pair, we propose a neural network which finds correct correspondences by a binary-class classifier and estimates relative pose through classified correspondences. Firs…

2021

DeepDT: Learning Geometry From Delaunay Triangulation for Surface Reconstruction

AAAI 2021technical

In this paper, a novel learning-based network, named DeepDT, is proposed to reconstruct the surface from Delaunay triangulation of point cloud. DeepDT learns to predict inside/outside labels of Delaunay tetrahedrons directly from a point cloud and corresponding Delaunay triangulation. The local geom…