← Search

Weixin Li

11 accepted papers

2026

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

CVPR 2026

Embodied navigation is a fundamental capability for robotic agents operating. Real-world deployment requires open vocabulary generalization and low training overhead, motivating zero-shot methods rather than task-specific RL training. However, existing zero-shot methods that build explicit 3D scene

Cited by 0SourceScholar
2026

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation

CVPR 2026

Memory-efficient transfer learning (METL) approaches have recently achieved promising performance in adapting pre-trained models to downstream tasks. They avoid applying gradient backpropagation in large backbones, thus significantly reducing the number of trainable parameters and high memory consum

Cited by 0SourcecodeScholar
2026

RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models

AAAI 2026technical

Pre-trained Vision-Language Models (VLMs), e.g. CLIP, have become essential tools in multimodal transfer learning. However, fine-tuning VLMs in few-shot scenarios poses significant challenges in balancing task-specific adaptation and generalization in the obtained model. Meanwhile, current researc

Cited by 0SourcePDFScholar
2025

Generating Editable Head Avatars with 3D Gaussian GANs

ICASSP 2025accepted

Generating animatable and editable 3D head avatars is essential for various applications in computer vision and graphics. Traditional 3D-aware generative adversarial networks (GANs), often using implicit fields like Neural Radiance Fields (NeRF), achieve photo-realistic and view-consistent 3D head s…

Cited by 0SourceScholar
2024

Leveraging Predicate and Triplet Learning for Scene Graph Generation

CVPR 2024poster

Scene Graph Generation (SGG) aims to identify entities and predict the relationship triplets <subject predicate object> in visual scenes. Given the prevalence of large visual variations of subject-object pairs even in the same predicate it can be quite challenging to model and refine predicate repre…

2024

Read, Spell and Repeat: Scene Text Recognition with Vision-Language Circular Refinement

ICASSP 2024accepted

Scene Text Recognition (STR) has long been considered an important yet challenging task in the field of computer vision. Recent works have demonstrated that utilizing language information is effective for the visually difficult images, like ones with occultation or blurring. However, the use of lang…

Cited by 0SourceScholar
2023

Causal Intervention and Counterfactual Reasoning for Multi-modal Fake News Detection

ACL 2023long

Due to the rapid upgrade of social platforms, most of today’s fake news is published and spread in a multi-modal form. Most existing multi-modal fake news detection methods neglect the fact that some label-specific features learned from the training set cannot generalize well to the testing set, thu…

2023

Transcendental Idealism of Planner: Evaluating Perception from Planning Perspective for Autonomous Driving

ICML 2023poster

Evaluating the performance of perception modules in autonomous driving is one of the most critical tasks in developing the complex intelligent system. While module-level unit test metrics adopted from traditional computer vision tasks are feasible to some extent, it remains far less explored to meas…

2021

MIEHDR CNN: Main Image Enhancement based Ghost-Free High Dynamic Range Imaging using Dual-Lens Systems

AAAI 2021technical

We study the High Dynamic Range (HDR) imaging problem using two Low Dynamic Range (LDR) images that are shot from dual-lens systems in a single shot time with different exposures. In most of the related HDR imaging methods, the problem is usually solved by Multiple Images Merging, i.e. the final HDR…

Cited by 8SourcePDFScholar
2016

VLAD3: Encoding Dynamics of Deep Features for Action Recognition

CVPR 2016poster

Previous approaches to action recognition with deep features tend to process video frames only within a small temporal region, and do not model long-range dynamic information explicitly. However, such information is important for the accurate recognition of actions, especially for the discrimination…

Cited by 111PDFScholar