← Search

Weiyu Guo

10 accepted papers

2026

Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection

AAAI 2026technical

Source-Free Object Detection (SFOD) aims to adapt a source-pretrained object detector to a target domain without access to source data. However, existing SFOD methods predominantly rely on internal knowledge from the source model, which limits their capacity to generalize across domains and often re

Cited by 0SourcePDFScholar
2026

Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

ICML 2026poster

We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly assume that task-relevant objects are always visible, leading to brittle and reactive behaviors when targets fall outside the camera’s field of view. S…

Cited by 0SourceScholar
2025

Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding

NeurIPS 2025poster

Understanding long video content is a complex endeavor that often relies on densely sampled frame captions or end-to-end feature selectors, yet these techniques commonly overlook the logical relationships between textual queries and visual elements. In practice, computational constraints necessitate…

Cited by 0SourceScholar
2025

Revisiting Noise Resilience Strategies in Gesture Recognition: Short-Term Enhancement in sEMG Analysis

ICML 2025poster

Gesture recognition based on surface electromyography (sEMG) has been gaining importance in many 3D Interactive Scenes. However, sEMG is easily influenced by various forms of noise in real-world environments, leading to challenges in providing long-term stable interactions through sEMG. Existing met…

Cited by 0SourcePDFScholar
2025

See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model

NeurIPS 2025poster

We introduce See&Trek, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMs) under vision-only constraints. While prior efforts have incorporated modalities like depth or point clouds to improve spatial reasoning, purely…

Cited by 0SourceScholar
2024

SpGesture: Source-Free Domain-adaptive sEMG-based Gesture Recognition with Jaccard Attentive Spiking Neural Network

NeurIPS 2024poster

Surface electromyography (sEMG) based gesture recognition offers a natural and intuitive interaction modality for wearable devices. Despite significant advancements in sEMG-based gesture recognition models, existing methods often suffer from high computational latency and increased energy consumptio…

2022

Context-Enhanced Stereo Transformer

ECCV 2022poster

"Stereo depth estimation is of great interest for computer vision research. However, existing methods struggles to generalize and predict reliably in hazardous regions, such as large uniform regions. To overcome these limitations, we propose Context Enhanced Path (CEP). CEP improves the generalizati…

2022

CrowdFormer: An Overlap Patching Vision Transformer for Top-Down Crowd Counting

IJCAI 2022poster

Crowd counting methods typically predict a density map as an intermediate representation of counting, and achieve good performance. However, due to the perspective phenomenon, there is a scale variation in real scenes, which causes the density map-based methods suffer from a severe scene generalizat…

2021

A Bi-Directional LSTM Network for Estimating Continuous Upper Limb Movement From Surface Electromyography

RA-L 2021

In human-machine interaction systems, continuous movement estimation methods occupy an important position because they are more natural and intuitive than pattern-recognition methods. Essentially, arm position is decided by the shoulder and elbow joint angles. However, the various deformations of mu

Cited by 74SourceScholar
2021

A Novel and Efficient Feature Extraction Method for Deep Learning Based Continuous Estimation

RA-L 2021

Simultaneous and proportional control (SPC) methods based on surface electromyogram (sEMG) can provide a more intuitive and natural interaction in rehabilitation assistive robots and prostheses. Recently, an increasing number of researchers have utilized deep learning methods to continuously estimat

Cited by 27SourceScholar