← Search

Yandong Guo

44 accepted papers

2026

Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

ICML 2026poster

We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly assume that task-relevant objects are always visible, leading to brittle and reactive behaviors when targets fall outside the camera’s field of view. S…

Cited by 0SourceScholar
2025

3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment

ICRA 2025

The 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide model learning. This setting presents two primary challenges: category-level ambiguity and instance-level complexity. Catego

Cited by 2SourceScholar
2025

Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning

NeurIPS 2025poster

Generalized policy and execution efficiency constitute the two critical challenges in robotic manipulation. While recent foundation policies benefit from the common-sense reasoning capabilities of internet-scale pretrained vision-language models (VLMs), they often suffer from low execution frequency…

Cited by 0SourcecodeScholar
2025

LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

AAAI 2025technical

Recently, Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have shown promise in instruction following and image understanding. While these models are powerful, they have not yet been developed to comprehend the more challenging 3D geometric and physical scenes, especially w…

2025

MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation

NeurIPS 2025poster

Reasoning segmentation aims to segment target objects in complex scenes based on human intent and spatial reasoning. While recent multimodal large language models (MLLMs) have demonstrated impressive 2D image reasoning segmentation, adapting these capabilities to 3D scenes remains underexplored. In…

Cited by 0SourceScholar
2025

Surprise3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes

NeurIPS 2025poster

The integration of language and 3D perception is critical for embodied AI and robotic systems to perceive, understand, and interact with the physical world. Spatial reasoning, a key capability for understanding spatial relationships between objects, remains underexplored in current 3D vision-languag…

Cited by 0SourceScholar
2025

Towards Accurate Time Series Forecasting via Implicit Decoding

NeurIPS 2025poster

Recent booming time series models have demonstrated remarkable forecasting performance. However, these methods often place greater focus on more effectively modelling the historical series, largely neglecting the forecasting phase, which generates long-term forecasts by separately predicting multipl…

Cited by 0SourcecodeScholar
2024

BEVUDA: Multi-geometric Space Alignments for Domain Adaptive BEV 3D Object Detection

ICRA 2024poster

Vision-centric bird-eye-view (BEV) perception has shown promising potential in autonomous driving. Recent works mainly focus on improving efficiency or accuracy but neglect the challenges when facing environment changing, resulting in severe degradation of transfer performance. For BEV perception, w…

Cited by 5SourcecodeScholar
2024

Continual-MAE: Adaptive Distribution Masked Autoencoders for Continual Test-Time Adaptation

CVPR 2024poster

Continual Test-Time Adaptation (CTTA) is proposed to migrate a source pre-trained model to continually changing target distributions addressing real-world dynamism. Existing CTTA methods mainly rely on entropy minimization or teacher-student pseudo-labeling schemes for knowledge extraction in unlabe…

Cited by 11SourcePDFScholar
2024

Debiased Novel Category Discovering and Localization

AAAI 2024technical

In recent years, object detection in deep learning has experienced rapid development. However, most existing object detection models perform well only on closed-set datasets, ignoring a large number of potential objects whose categories are not defined in the training set. These objects are often id…

Cited by 6SourcePDFScholar
2024

NTO3D: Neural Target Object 3D Reconstruction with Segment Anything

CVPR 2024poster

Neural 3D reconstruction from multi-view images has recently attracted increasing attention from the community. Existing methods normally learn a neural field for the whole scene while it is still under-explored how to reconstruct a target object indicated by users. Considering the Segment Anything…

2024

RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

NeurIPS 2024poster

A fundamental objective in robot manipulation is to enable models to comprehend visual scenes and execute actions. Although existing Vision-Language-Action (VLA) models for robots can handle a range of basic tasks, they still face challenges in two areas: (1) insufficient reasoning ability to tackle…

Cited by 5SourcePDFScholar
2024

Tag2Text: Guiding Vision-Language Model via Image Tagging

ICLR 2024poster

This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a…

Cited by 84SourcePDFScholar
2024

ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation

ICLR 2024poster

Since real-world machine systems are running in non-stationary environments, Continual Test-Time Adaptation (CTTA) task is proposed to adapt the pre-trained model to continually changing target domains. Recently, existing methods mainly focus on model-based adaptation, which aims to leverage a self-…

2023

A Comprehensive Comparison of Projections in Omnidirectional Super-Resolution

ICASSP 2023accepted

Super-Resolution (SR) has gained increasing research attention over the past few years. With the development of Deep Neural Networks (DNNs), many super-resolution methods based on DNNs have been proposed. Although most of these methods are aimed at ordinary frames, there are few works on super-resol…

Cited by 0SourceScholar
2023

BEV-SAN: Accurate BEV 3D Object Detection via Slice Attention Networks

CVPR 2023poster

Bird's-Eye-View (BEV) 3D Object Detection is a crucial multi-view technique for autonomous driving systems. Recently, plenty of works are proposed, following a similar paradigm consisting of three essential components, i.e., camera feature extraction, BEV feature construction, and task heads. Among…

Cited by 28SourcePDFScholar
2023

Box-Level Active Detection

CVPR 2023highlight

Active learning selects informative samples for annotation within budget, which has proven efficient recently on object detection. However, the widely used active detection benchmarks conduct image-level evaluation, which is unrealistic in human workload estimation and biased towards crowded images.…

2023

CABM: Content-Aware Bit Mapping for Single Image Super-Resolution Network With Large Input

CVPR 2023poster

With the development of high-definition display devices, the practical scenario of Super-Resolution (SR) usually needs to super-resolve large input like 2K to higher resolution (4K/8K). To reduce the computational and memory cost, current methods first split the large input into local patches and th…

2023

CloSET: Modeling Clothed Humans on Continuous Surface With Explicit Template Decomposition

CVPR 2023poster

Creating animatable avatars from static scans requires the modeling of clothing deformations in different poses. Existing learning-based methods typically add pose-dependent deformations upon a minimally-clothed mesh template or a learned implicit template, which have limitations in capturing detail…

Cited by 28SourcePDFScholar
2023

ContrastMotion: Self-supervised Scene Motion Learning for Large-Scale LiDAR Point Clouds

IJCAI 2023poster

In this paper, we propose a novel self-supervised motion estimator for LiDAR-based autonomous driving via BEV representation. Different from usually adopted self-supervised strategies for data-level structure consistency, we predict scene motion via feature-level consistency between pillars in conse…

2023

Data-Driven Based Cascading Orientation and Translation Estimation for Inertial Navigation

IROS 2023poster

Recently, data-driven approaches have brought both opportunities and challenges for Inertial Navigation Systems. In this paper, we propose a novel data-driven method which is composed of cascading orientation and translation estimation with IMU-only measurements. For robust orientation estimation, w…

Cited by 4SourceScholar
2023

Learning Audio-Visual Source Localization via False Negative Aware Contrastive Learning

CVPR 2023poster

Self-supervised audio-visual source localization aims to locate sound-source objects in video frames without extra annotations. Recent methods often approach this goal with the help of contrastive learning, which assumes only the audio and visual contents from the same video are positive samples for…

2023

Mosaic Representation Learning for Self-supervised Visual Pre-training

ICLR 2023top-25%

Self-supervised learning has achieved significant success in learning visual representations without the need for manual annotation. To obtain generalizable representations, a meticulously designed data augmentation strategy is one of the most crucial parts. Recently, multi-crop strategies utilizing…

2023

PiMAE: Point Cloud and Image Interactive Masked Autoencoders for 3D Object Detection

CVPR 2023poster

Masked Autoencoders learn strong visual representations and achieve state-of-the-art results in several independent modalities, yet very few works have addressed their capabilities in multi-modality settings. In this work, we focus on point cloud and RGB image data, two modalities that are often pre…

2023

SELVO: A Semantic-Enhanced Lidar-Visual Odometry

IROS 2023poster

In the face of complex external environment, single sensor information can no longer meet the accuracy requirements of low-drift SLAM. In this paper, we focus on the fusion scheme of cameras and lidar, and explore the gain of semantic information to SLAM system. A Semantic-Enhanced Lidar-Visual Odom…

Cited by 2SourceScholar
2023

Ultra Real-Time Portrait Matting via Parallel Semantic Guidance

ICASSP 2023accepted

Most existing portrait matting models either require expensive auxiliary information or try to decompose the task into sub-tasks that are usually resource-hungry. These challenges limit its application on low-power computing devices. In this paper, we propose an ultra-light-weighted portrait matting…

Cited by 0SourceScholar
2022

Adaptive Patch Exiting for Scalable Single Image Super-Resolution

ECCV 2022poster

"Since the future of computing is heterogeneous, scalability is a crucial problem for single image super-resolution. Recent works try to train one network, which can be deployed on platforms with different capacities. However, they rely on the pixel-wise sparse convolution, which is not hardware-fri…

2022

CRIS: CLIP-Driven Referring Image Segmentation

CVPR 2022poster

Referring image segmentation aims to segment a referent via a natural linguistic expression. Due to the distinct data properties between text and image, it is challenging for a network to well align text and pixel-level features. Existing approaches use pretrained models to facilitate learning, yet…

Cited by 441PDFcodeScholar
2022

Efficient Meta-Tuning for Content-Aware Neural Video Delivery

ECCV 2022poster

"Recently, Deep Neural Networks (DNNs) are utilized to reduce the bandwidth and improve the quality of Internet video delivery. Existing methods train corresponding content-aware super-resolution (SR) model for each video chunk on the server, and stream low-resolution (LR) video chunks along with SR…

2022

On the Efficacy of Small Self-Supervised Contrastive Models without Distillation Signals

AAAI 2022technical

It is a consensus that small models perform quite poorly under the paradigm of self-supervised contrastive learning. Existing methods usually adopt a large off-the-shelf model to transfer knowledge to the small one via distillation. Despite their effectiveness, distillation-based methods may not be…

2022

Personalized Image Aesthetics Assessment With Rich Attributes

CVPR 2022poster

Personalized image aesthetics assessment (PIAA) is challenging due to its highly subjective nature. People's aesthetic tastes depend on diversified factors, including image characteristics and subject characters. The existing PIAA databases are limited in terms of annotation diversity, especially th…

Cited by 80PDFScholar
2022

Pose Refinement with Joint Optimization of Visual Points and Lines

IROS 2022poster

High-precision camera re-localization technology in a pre-established 3D environment map is the basis for many tasks, such as Augmented Reality, Robotics and Autonomous Driving. The point-based visual re-localization approaches are well-developed in recent decades, but are insufficient in some featu…

Cited by 22SourceScholar
2022

SDETR: Attention-Guided Salient Object Detection with Transformer

ICASSP 2022accepted

Most existing CNN-based salient object detection methods can identify fine-grained segmentation details like hair and animal fur, but often mispredict the salient object due to lack of global contextual information caused by locality convolution layers. The limited training data of the current SOD t…

Cited by 0SourceScholar
2022

Self-Distillation From the Last Mini-Batch for Consistency Regularization

CVPR 2022poster

Knowledge distillation (KD) shows a bright promise as a powerful regularization strategy to boost generalization ability by leveraging learned sample-level soft targets. Yet, employing a complex pre-trained teacher network or an ensemble of peer students in existing KD is both time-consuming and com…

Cited by 93PDFcodeScholar
2022

Single-Stage Is Enough: Multi-Person Absolute 3D Pose Estimation

CVPR 2022poster

The existing multi-person absolute 3D pose estimation methods are mainly based on two-stage paradigm, i.e., top-down or bottom-up, leading to redundant pipelines with high computation cost. We argue that it is more desirable to simplify such two-stage paradigm to a single-stage one to promote both e…

Cited by 56PDFScholar
2022

Structured Local Radiance Fields for Human Avatar Modeling

CVPR 2022poster

It is extremely challenging to create an animatable clothed human avatar from RGB videos, especially for loose clothes due to the difficulties in motion modeling. To address this problem, we introduce a novel representation on the basis of recent neural scene rendering techniques. The core of our re…

Cited by 218PDFcodeScholar
2021

Retrieval and Localization with Observation Constraints

ICRA 2021poster

Accurate visual re-localization is very critical to many artificial intelligence applications, such as augmented reality, virtual reality, robotics and autonomous driving. To accomplish this task, we propose an integrated visual re-localization method called RLOCS by combining image retrieval, seman…

Cited by 10SourceScholar
2021

Virtual Multi-Modality Self-Supervised Foreground Matting for Human-Object Interaction

ICCV 2021poster

Most existing human matting algorithms tried to separate pure human-only foreground from the background. In this paper, we propose a Virtual Multi-modality Foreground Matting (VMFM) method to learn human-object interactive foreground (human and objects interacted with him or her) from a raw RGB imag…

Cited by 7PDFcodeScholar
2017

Model-Based Iterative Restoration for Binary Document Image Compression With Dictionary Learning

CVPR 2017poster

The inherent noise in the observed (e.g., scanned) binary document image degrades the image quality and harms the compression ratio through breaking the pattern repentance and adding entropy to the document images. In this paper, we design a cost function in Bayesian framework with dictionary learni…

Cited by 10PDFScholar