← Search

Gui-song Xia

56 accepted papers

2026

Event-Guided Super-Resolving Blurry Image via Asymmetric Integral Driven Consistency

AAAI 2026technical

Super-Resolution from a Blurry low-resolution image (SRB) constitutes a severely ill-posed inverse problem. Current learning-based SRB approaches primarily rely on synthetic, well-labeled paired datasets to regularize solution spaces, yet they exhibit limited generalizability in practical applicatio

Cited by 0SourcePDFScholar
2026

I2D-LocX: An Efficient, Precise and Robust Method for Camera Localization in LiDAR Maps

ICRA 2026poster

Camera localization within LiDAR maps has gained significant attention due to its potential for accurate positioning with low-cost and lightweight sensors compared to LiDAR-based systems. However, existing methods often prioritize localization accuracy, sometimes compromising efficiency, which can l…

Cited by 0SourceScholar
2026

SGE-GLoc: Semantic Gaussian Ellipsoid Scene Graphs for Efficient LiDAR Global Localization

RA-L 2026

Global localization, encompassing robust place recognition and precise transformation estimation, is crucial for mobile robot navigation when the global navigation satellite system (GNSS) is unavailable. While LiDAR-based approaches are favored for their accuracy in 3D perception and resilience to i

Cited by 0SourceScholar
2026

Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image

ICLR 2026poster

Generating a street-level 3D scene from a single satellite image is a crucial yet challenging task. Current methods present a stark trade-off: geometry-colorization models achieve high geometric fidelity but are typically building-focused and lack semantic diversity. In contrast, proxy-based models…

Cited by 0SourcecodeScholar
2025

AdaDCP: Learning an Adapter with Discrete Cosine Prior for Clear-to-Adverse Domain Generalization

ICCV 2025poster

Vision Foundation Model (VFM) provides an inherent generalization ability to unseen domains for downstream tasks. However, fine-tuning VFM to parse various adverse scenes (e.g., fog, snow, night) is particularly challenging, as these samples are difficult to collect. Using easy-to-acquire clear scen…

Cited by 0SourcePDFScholar
2025

Bypass Back-propagation: Optimization-based Structural Pruning for Large Language Models via Policy Gradient

ACL 2025long

Recent Large-Language Models (LLMs) pruning methods typically operate at the post-training phase without the expensive weight finetuning, however, their pruning criteria often rely on **heuristically hand-crafted metrics**, potentially leading to suboptimal performance. We instead propose a novel **…

Cited by 0SourcePDFScholar
2025

Exploring Scene Affinity for Semi-Supervised LiDAR Semantic Segmentation

CVPR 2025poster

This paper explores scene affinity (AIScene), namely intra-scene consistency and inter-scene correlation, for semi-supervised LiDAR semantic segmentation in driving scenes. Adopting teacher-student training, AIScene employs a teacher network to generate pseudo-labeled scenes from unlabeled data, whi…

2025

Fuse Before Transfer: Knowledge Fusion for Heterogeneous Distillation

ICCV 2025poster

Most knowledge distillation (KD) methods focus on teacher-student pairs with similar architectures, such as both being CNN models. The potential and flexibility of KD can be greatly improved by expanding it to Cross-Architecture KD (CAKD), where the knowledge of homogeneous and heterogeneous teacher…

2025

Holistic Large-Scale Scene Reconstruction via Mixed Gaussian Splatting

NeurIPS 2025poster

Recent advances in 3D Gaussian Splatting have shown remarkable potential for novel view synthesis. However, most existing large-scale scene reconstruction methods rely on the divide-and-conquer paradigm, which often leads to the loss of global scene information and requires complex parameter tuning…

Cited by 0SourcecodeScholar
2025

I2D-LocX: An Efficient, Precise and Robust Method for Camera Localization in LiDAR Maps

RA-L 2025

Camera localization within LiDAR maps has gained significant attention due to its potential for accurate positioning with low-cost and lightweight sensors compared to LiDAR-based systems. However, existing methods often prioritize localization accuracy, sometimes compromising efficiency, which can l

Cited by 2SourceScholar
2025

Learning Fine-grained Domain Generalization via Hyperbolic State Space Hallucination

AAAI 2025technical

Fine-grained domain generalization (FGDG) aims to learn a fine-grained representation that can be well generalized to unseen target domains when only trained on the source domain data. Compared with generic domain generalization, FGDG is particularly challenging in that the fine-grained category can…

2025

Robust Visual Odometry Using Rigidly-Bundled Arbitrarily-Arranged Multi-Cameras

RA-L 2025

Making multi-camera visual SLAM systems easier to set up and more robust to the environment is attractive for vision robots. Existing monocular and binocular vision SLAM systems have narrow sensing Field-of-View (FoV), resulting in degenerated accuracy and limited robustness in textureless environme

Cited by 0SourcecodeScholar
2025

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation

CVPR 2025poster

Recently, text-to-image generation models have achieved remarkable advancements, particularly with diffusion models facilitating high-quality image synthesis from textual descriptions. However, these models often struggle with achieving precise control over pixel-level layouts, object appearances, a…

2025

VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis

AAAI 2025technical

This paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis. VHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest instruction dataset comprising both factual and deceptive questions (HnstD). Un…

2024

3D Building Reconstruction from Monocular Remote Sensing Images with Multi-level Supervisions

CVPR 2024poster

3D building reconstruction from monocular remote sensing images is an important and challenging research problem that has received increasing attention in recent years owing to its low cost of data acquisition and availability for large-scale applications. However existing methods rely on expensive…

2024

Anchor-based Robust Finetuning of Vision-Language Models

CVPR 2024poster

We aim at finetuning a vision-language model without hurting its out-of-distribution (OOD) generalization. We address two types of OOD generalization i.e. i) domain shift such as natural to sketch images and ii) zero-shot capability to recognize the category that was not contained in the finetune da…

Cited by 9SourcePDFScholar
2024

Aux-NAS: Exploiting Auxiliary Labels with Negligibly Extra Inference Cost

ICLR 2024poster

We aim at exploiting additional auxiliary labels from an independent (auxiliary) task to boost the primary task performance which we focus on, while preserving a single task inference cost of the primary task. While most existing auxiliary learning methods are optimization-based relying on loss weig…

2024

DMTG: One-Shot Differentiable Multi-Task Grouping

ICML 2024poster

We aim to address Multi-Task Learning (MTL) with a large number of tasks by Multi-Task Grouping (MTG). Given $N$ tasks, we propose to **simultaneously identify the best task groups from $2^N$ candidates and train the model weights simultaneously in one-shot**, with **the high-order task-affinity ful…

2024

I2D-Loc++: Camera Pose Tracking in LiDAR Maps With Multi-View Motion Flows

RA-L 2024

Camera localization in LiDAR maps has become increasingly popular due to its promising ability to handle complex scenarios, surpassing the limitations of visual-only localization methods. However, existing approaches mostly focus on addressing the cross-modal 2D–3D gaps while overlooking the relatio

Cited by 4SourceScholar
2024

NEAT: Distilling 3D Wireframes from Neural Attraction Fields

CVPR 2024poster

This paper studies the problem of structured 3D recon- struction using wireframes that consist of line segments and junctions focusing on the computation of structured boundary geometries of scenes. Instead of leveraging matching-based solutions from 2D wireframes (or line segments) for 3D wireframe…

2024

QuadricsNet: Learning Concise Representation for Geometric Primitives in Point Clouds

ICRA 2024poster

This paper presents a novel framework to learn a concise geometric primitive representation for 3D point clouds. Different from representing each type of primitive individually, we focus on the challenging problem of how to achieve a concise and uniform representation robustly. We employ quadrics to…

Cited by 5SourcecodeScholar
2024

Toward Robust Keypoint Detection and Tracking: A Fusion Approach With Event-Aligned Image Features

RA-L 2024

Robust keypoint detection and tracking are crucial for various robotic tasks. However, conventional cameras struggle under rapid motion and lighting changes, hindering local and edge feature extraction essential for keypoint detection and tracking. Event cameras offer advantages in such scenarios du

Cited by 12SourceScholar
2024

Unleashing Unlabeled Data: A Paradigm for Cross-View Geo-Localization

CVPR 2024poster

This paper investigates the effective utilization of unlabeled data for large-area cross-view geo-localization (CVGL) encompassing both unsupervised and semi-supervised settings. Common approaches to CVGL rely on ground-satellite image pairs and employ label-driven supervised training. However the c…

2023

ConDaFormer: Disassembled Transformer with Local Structure Enhancement for 3D Point Cloud Understanding

NeurIPS 2023poster

Transformers have been recently explored for 3D point cloud understanding with impressive progress achieved. A large number of points, over 0.1 million, make the global self-attention infeasible for point cloud data. Thus, most methods propose to apply the transformer in a local region, e.g., spheri…

2023

Dynamic Coarse-To-Fine Learning for Oriented Tiny Object Detection

CVPR 2023poster

Detecting arbitrarily oriented tiny objects poses intense challenges to existing detectors, especially for label assignment. Despite the exploration of adaptive label assignment in recent oriented object detectors, the extreme geometry shape and limited feature of oriented tiny objects still induce…

2023

Few-Shot Object Detection via Variational Feature Aggregation

AAAI 2023technical

As few-shot object detectors are often trained with abundant base samples and fine-tuned on few-shot novel examples, the learned models are usually biased to base classes and sensitive to the variance of novel examples. To address this issue, we propose a meta-learning framework with two novel featu…

2023

Generalizing Event-Based Motion Deblurring in Real-World Scenarios

ICCV 2023poster

Event-based motion deblurring has shown promising results by exploiting low-latency events. However, current approaches are limited in their practical usage, as they assume the same spatial resolution of inputs and specific blurriness distributions. This work addresses these limitations and aims to…

Cited by 29PDFcodeScholar
2023

HGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic Segmentation

CVPR 2023poster

Current semantic segmentation models have achieved great success under the independent and identically distributed (i.i.d.) condition. However, in real-world applications, test data might come from a different domain than training data. Therefore, it is important to improve model robustness against…

2023

Level-S$^2$fM: Structure From Motion on Neural Level Set of Implicit Surfaces

CVPR 2023poster

This paper presents a neural incremental Structure-from-Motion (SfM) approach, Level-S2fM, which estimates the camera poses and scene geometry from a set of uncalibrated images by learning coordinate MLPs for the implicit surfaces and the radiance fields from the established keypoint correspondences…

2023

OmniCity: Omnipotent City Understanding With Multi-Level and Multi-View Images

CVPR 2023poster

This paper presents OmniCity, a new dataset for omnipotent city understanding from multi-level and multi-view images. More precisely, OmniCity contains multi-view satellite images as well as street-level panorama and mono-view images, constituting over 100K pixel-wise annotated images that are well-…

2023

Sat2Density: Faithful Density Learning from Satellite-Ground Image Pairs

ICCV 2023poster

This paper aims to develop an accurate 3D geometry representation of satellite images using satellite-ground image pairs. Our focus is on the challenging problem of 3D-aware ground-views synthesis from a satellite image. We draw inspiration from the density field representation used in volumetric ne…

Cited by 17PDFcodeScholar
2022

Expanding Low-Density Latent Regions for Open-Set Object Detection

CVPR 2022poster

Modern object detectors have achieved impressive progress under the close-set setup. However, open-set object detection (OSOD) remains challenging since objects of unknown categories are often misclassified to existing known classes. In this work, we propose to identify unknown objects by separating…

Cited by 82PDFcodeScholar
2022

Learning Local-Global Contextual Adaptation for Multi-Person Pose Estimation

CVPR 2022poster

This paper studies the problem of multi-person pose estimation in a bottom-up fashion. With a new and strong observation that the localization issue of the center-offset formulation can be remedied in a local-window search scheme in an ideal situation, we propose a multi-person pose estimation appro…

Cited by 48PDFcodeScholar
2022

Partial Wasserstein Adversarial Network for Non-rigid Point Set Registration

ICLR 2022poster

Given two point sets, the problem of registration is to recover a transformation that matches one set to the other. This task is challenging due to the presence of large number of outliers, the unknown non-rigid deformations and the large sizes of point sets. To obtain strong robustness against outl…

Cited by 5SourcePDFScholar
2022

RFLA: Gaussian Receptive Field Based Label Assignment for Tiny Object Detection

ECCV 2022poster

"Detecting tiny objects is one of the main obstacles hindering the development of object detection. The performance of generic object detectors tends to drastically deteriorate on tiny object detection tasks. In this paper, we point out that either box prior in the anchor-based detector or point pri…

2022

Revisiting Document Image Dewarping by Grid Regularization

CVPR 2022poster

This paper addresses the problem of document image dewarping, which aims at eliminating the geometric distortion in document images for document digitization. Instead of designing a better neural network to approximate the optical flow fields between the inputs and outputs, we pursue the best readab…

Cited by 34PDFcodeScholar
2021

3D Building Reconstruction From Monocular Remote Sensing Images

ICCV 2021poster

3D building reconstruction from monocular remote sensing imagery is an important research problem and an economic solution to large-scale city modeling, compared with reconstruction from LiDAR data and multi-view imagery. However, several challenges such as the partial invisibility of building footp…

Cited by 36PDFcodeScholar
2021

Event-Based Synthetic Aperture Imaging With a Hybrid Network

CVPR 2021poster

Synthetic aperture imaging (SAI) is able to achieve the see through effect by blurring out the off-focus foreground occlusions and reconstructing the in-focus occluded targets from multi-view images. However, very dense occlusions and extreme lighting conditions may bring significant disturbances to…

Cited by 39PDFcodeScholar
2020

FGN: Fully Guided Network for Few-Shot Instance Segmentation

CVPR 2020poster

Few-shot instance segmentation (FSIS) conjoins the few-shot learning paradigm with general instance segmentation, which provides a possible way of tackling instance segmentation in the lack of abundant labeled data for training. This paper presents a Fully Guided Network (FGN) for few-shot instance…

Cited by 89PDFScholar
2020

Holistically-Attracted Wireframe Parsing

CVPR 2020poster

This paper presents a fast and parsimonious parsing method to accurately and robustly detect a vectorized wireframe in an input image with a single forward pass. The proposed method is end-to-end trainable, consisting of three components: (i) line segment and junction proposal generation, (ii) line…

Cited by 136PDFcodeScholar
2019

Learning Attraction Field Representation for Robust Line Segment Detection

CVPR 2019poster

This paper presents a region-partition based attraction field dual representation for line segment maps, and thus poses the problem of line segment detection (LSD) as the region coloring problem. The latter is then addressed by learning deep convolutional neural networks (ConvNets) for accur…

Cited by 158PDFcodeScholar
2019

Learning RoI Transformer for Oriented Object Detection in Aerial Images

CVPR 2019poster

Object detection in aerial images is an active yet challenging task in computer vision because of the bird's-eye view perspective, the highly complex backgrounds, and the variant appearances of objects. Especially when detecting densely packed objects in aerial images, methods relying on horizontal…

Cited by 1326PDFcodeScholar
2018

DOTA: A Large-Scale Dataset for Object Detection in Aerial Images

CVPR 2018poster

Object detection is an important and challenging problem in computer vision. Although the past decade has witnessed major advances in object detection in natural scenes, such successes have been slow to aerial imagery, not only because of the huge variation in the scale, orientation and shape of the…

2018

Rotation-Sensitive Regression for Oriented Scene Text Detection

CVPR 2018poster

Text in natural images is of arbitrary orientations, requiring detection in terms of oriented bounding boxes. Normally, a multi-oriented text detector often involves two key tasks: 1) text presence detection, which is a classification problem disregarding text orientation; 2) oriented bounding box r…