← Search

Dongyan Guo

9 accepted papers

2026

Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object Tracking

CVPR 2026

Natural language provides valuable auxiliary information for enhancing visual object tracking. While existing vision-language tracking methods explicitly leverage linguistic descriptions to aid tracking, they suffer from two critical limitations: the inability to dynamically adapt descriptions to th

Cited by 0SourceScholar
2026

GeoMotion: Rethinking Motion Segmentation via Latent 4D Geometry

CVPR 2026

Motion segmentation in dynamic scenes is highly challenging, as conventional methods heavily rely on estimating camera poses and point correspondences from inherently noisy motion cues. Existing statistical inference or iterative optimization techniques that struggle to mitigate the cumulative error

Cited by 0SourcecodeScholar
2025

DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation

AAAI 2025technical

Monocular camera calibration is a key precondition for numerous 3D vision applications. Despite considerable advancements, existing methods often hinge on specific assumptions and struggle to generalize across varied real-world scenarios, and the performance is limited by insufficient training data.…

Cited by 4SourcePDFScholar
2024

PointAttN: You Only Need Attention for Point Cloud Completion

AAAI 2024technical

Point cloud completion referring to completing 3D shapes from partial 3D point clouds is a fundamental problem for 3D point cloud analysis tasks. Benefiting from the development of deep neural networks, researches on point cloud completion have made great progress in recent years. However, the expli…

2023

Multi-Stage Aggregation Transformer for Medical Image Segmentation

ICASSP 2023accepted

Capturing rich multi-scale features is essential for resolving complex variations in medical image segmentation. In this paper, we explore how to fully utilize the advantages of Convolutional neural networks (CNN) and Transformer, and propose a novel multi-stage aggregation architecture named MA-Tra…

Cited by 0SourceScholar
2021

Collaborative Visual Inertial SLAM for Multiple Smart Phones

ICRA 2021poster

The efficiency and accuracy of mapping are crucial in a large scene and long-term AR applications. Multi-agent cooperative SLAM is the precondition of multi-user AR interaction. The cooperation of multiple smart phones has the potential to improve efficiency and robustness of task completion and can…

Cited by 16SourceScholar
2021

Consistency-Aware Graph Network for Human Interaction Understanding

ICCV 2021poster

Compared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main cause is that recent approaches learn human interactive relations via shallow graphical models…

Cited by 12PDFcodeScholar
2020

SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking

CVPR 2020oral

By decomposing the visual tracking task into two subproblems as classification for pixel category and regression for object bounding box at this pixel, we propose a novel fully convolutional Siamese network to solve visual tracking end-to-end in a per-pixel manner. The proposed framework SiamCAR con…

Cited by 980PDFcodeScholar