← Search

Li Guo

17 accepted papers

2026

SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition

ICLR 2026poster

Visual Place Recognition (VPR) requires robust retrieval of geotagged images despite large appearance, viewpoint, and environmental variation. Prior methods focus on descriptor fine-tuning or fixed sampling strategies yet neglect the dynamic interplay between spatial context and visual similarity d…

Cited by 0SourcecodeScholar
2026

Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI

CVPR 2026

Reproducible closed-loop evaluation remains a major bottleneck in Embodied AI such as visual navigation. A promising path forward is high-fidelity simulation that combines photorealistic sensor rendering with geometrically grounded interaction in complex, open-world urban environments. Although rece

Cited by 0SourcecodeScholar
2025

3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering

IROS 2025

With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by lever-aging the strengths of foundational models. The framework integrates key com

Cited by 12SourcecodeScholar
2025

Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion

ICRA 2025

Camera-based occupancy prediction is a main-stream approach for 3D perception in autonomous driving, aiming to infer complete 3D scene geometry and semantics from 2D images. Almost existing methods focus on improving performance through structural modifications, such as lightweight backbones and com

Cited by 1SourceScholar
2025

Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition

AAAI 2025technical

Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual…

2025

Hyperbolic-PDE GNN: Spectral Graph Neural Networks in the Perspective of A System of Hyperbolic Partial Differential Equations

ICML 2025poster

Graph neural networks (GNNs) leverage message passing mechanisms to learn the topological features of graph data. Traditional GNNs learns node features in a spatial domain unrelated to the topology, which can hardly ensure topological features. In this paper, we formulates message passing as a syste…

2025

PQNAS: Mixed-precision Quantization-aware Neural Architecture Search with Pseudo Quantizer

ICASSP 2025accepted

Quantization-aware neural architecture search is an efficient way to automatically search for the best quantized model that can meet the limited resource constraints on edge devices. Existing methods utilize the straight-through estimator for training the quantized supernet, but lead to oscillation…

Cited by 0SourceScholar
2025

SCS: Spatially Consistent Self-Supervised approach for One-Shot Anatomical Landmark Detection

ICASSP 2025accepted

Landmark detection is essential in medical image analysis, serving as the foundation for many downstream tasks. In recent years, supervised anatomical landmark detection models have achieved remarkable success, but typically require large amounts of labeled data for training, which is challenging to…

Cited by 0SourceScholar
2024

Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic Segmentation

AAAI 2024technical

Recently, CLIP has found practical utility in the domain of pixel-level zero-shot segmentation tasks. The present landscape features two-stage methodologies beset by issues such as intricate pipelines and elevated computational costs. While current one-stage approaches alleviate these concerns and…

2024

The Prevalence of Neural Collapse in Neural Multivariate Regression

NeurIPS 2024poster

Recently it has been observed that neural networks exhibit Neural Collapse (NC) during the final stage of training for the classification problem. We empirically show that multivariate regression, as employed in imitation learning and other applications, exhibits Neural Regression Collapse (NRC), a…

Cited by 4SourcePDFScholar
2023

Accurate MRI Reconstruction via Multi-Domain Recurrent Networks

IJCAI 2023poster

In recent years, deep convolutional neural networks (CNNs) have become dominant in MRI reconstruction from undersampled k-space. However, most existing CNNs methods reconstruct the undersampled images either in the spatial domain or in the frequency domain, and neglecting the correlation between the…

Cited by 7SourcePDFScholar
2023

Divide, Conquer, and Combine: Mixture of Semantic-Independent Experts for Zero-Shot Dialogue State Tracking

ACL 2023long

Zero-shot transfer learning for Dialogue State Tracking (DST) helps to handle a variety of task-oriented dialogue domains without the cost of collecting in-domain data. Existing works mainly study common data- or model-level augmentation methods to enhance the generalization but fail to effectively…

Cited by 19SourcePDFScholar
2023

Meta Omnium: A Benchmark for General-Purpose Learning-To-Learn

CVPR 2023poster

Meta-learning and other approaches to few-shot learning are widely studied for image recognition, and are increasingly applied to other vision tasks such as pose estimation and dense prediction. This naturally raises the question of whether there is any few-shot meta-learning algorithm capable of ge…

2021

A Periodic Frame Learning Approach for Accurate Landmark Localization in M-Mode Echocardiography

ICASSP 2021accepted

Anatomical landmark localization has been a key challenge for medical image analysis. Existing researches mostly adopt CNN as the main architecture for landmark localization while they are not applicable to process image modalities with periodic structure. In this paper, we propose a novel two-stage…

Cited by 0SourceScholar
2020

A Relation-Specific Attention Network for Joint Entity and Relation Extraction

IJCAI 2020poster

Joint extraction of entities and relations is an important task in natural language processing (NLP), which aims to capture all relational triplets from plain texts. This is a big challenge due to some of the triplets extracted from one sentence may have overlapping entities. Most existing methods p…

2020

Document-level Relation Extraction with Dual-tier Heterogeneous Graph

COLING 2020main

Document-level relation extraction (RE) poses new challenges over its sentence-level counterpart since it requires an adequate comprehension of the whole document and the multi-hop reasoning ability across multiple sentences to reach the final result. In this paper, we propose a novel graph-based mo…

Cited by 75SourcePDFScholar
2018

A Parallel Fusion Approach to Piano Music Transcription Based on Convolutional Neural Network

ICASSP 2018accepted

In this paper, a supervised approach based on Convolutional Neural Networks (CNN) for polyphonic piano transcription is presented. The system consists of pitch detection model, onset/offset detection model, and note search model. The pitch detection model is a single-channel CNN predicting the proba…

Cited by 0SourceScholar