← Search

Qi Wei

21 accepted papers

2026

Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks

ICLR 2026poster

Group-based reinforcement learning (RL), such as GRPO, has advanced the capabilities of large language models on long-horizon agentic tasks. To enable more fine-grained policy updates, recent research has increasingly shifted toward stepwise group-based policy optimization, which treats each step in…

Cited by 0SourcecodeScholar
2026

InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models

ICLR 2026poster

General SVG modeling remains challenging due to fragmented datasets, limited transferability of methods across tasks, and the difficulty of handling structural complexity. In response, we leverage the strong transfer and generalization capabilities of multimodal large language models (MLLMs) to achi…

Cited by 0SourcecodeScholar
2026

InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models

ICLR 2026poster

Recent benchmarks and datasets have been proposed to improve spatial reasoning in vision-language models (VLMs), yet existing open resources remain limited in scale, visual diversity, and instruction expressiveness. In this work, we introduce InternSpatial, the largest open-source dataset for spatia…

Cited by 0SourceScholar
2025

A Natural Human-Robot Interaction System for Teleoperation Based on Noncontact Haptic Feedback

IROS 2025

In order to provide natural and immersive interactive experience for teleoperation in the context of human-robot collaboration and interaction, this work introduces a natural human-robot interaction system for teleoperation based on ultrasonic haptic feedback. Specifically, our system can accurately

Cited by 0SourceScholar
2025

Influence-Based Fair Selection for Sample-Discriminative Backdoor Attack

AAAI 2025technical

Backdoor attacks have posed a serious threat in machine learning models, wherein adversaries can poison training samples with maliciously crafted triggers to compromise the victim model. Advanced backdoor attack methods have focused on selectively poisoning more vulnerable training samples, achievin…

Cited by 0SourcePDFScholar
2025

Representation Surgery in Model Merging with Probabilistic Modeling

ICML 2025poster

Model merging aims to achieve multitask performance by merging multiple expert models without the need to access the raw training data. Recent research identified the \textit{representation bias} of model merging, characterized by a discrepancy in the representation distribution between the merged a…

Cited by 0SourcePDFScholar
2025

Test-Time Multimodal Backdoor Detection by Contrastive Prompting

ICML 2025poster

While multimodal contrastive learning methods (e.g., CLIP) can achieve impressive zero-shot classification performance, recent research has revealed that these methods are vulnerable to backdoor attacks. To defend against backdoor attacks on CLIP, existing defense methods focus on either the pre-tra…

Cited by 0SourcePDFScholar
2024

Candidate Pseudolabel Learning: Enhancing Vision-Language Models by Prompt Tuning with Unlabeled Data

ICML 2024oral

Fine-tuning vision-language models (VLMs) with abundant unlabeled data recently has attracted increasing attention. Existing methods that resort to the pseudolabeling strategy would suffer from heavily incorrect hard pseudolabels when VLMs exhibit low zero-shot performance in downstream tasks. To al…

2023

Design, Implementation, and Observer-Based Output Control of a Super-Coiled Polymer-Driven Two Degree-of-Freedom Robotic Eye

RA-L 2023

The prevalence of ineffective corrective surgeries for ocular motor disorders calls for a robotic eye platform in aiding ophthalmologists to better understand the biomechanisms of human eye movement. This letter presents the first hardware design and implementation of a 2-DOF robotic eye driven by s

Cited by 4SourceScholar
2022

OCTOANTS: A Heterogeneous Lightweight Intelligent Multi-Robot Collaboration System with Resource-constrained IoT Devices

IROS 2022poster

As the focus on highly intelligent robots continues, a problem that cannot be ignored has emerged: resource con-straints. Considering the game problem of resource limitation and the level of intelligence, we focus on lightweight intelligence. This work is a further refinement of our previous work, a…

Cited by 2SourceScholar
2022

Self-Filtering: A Noise-Aware Sample Selection for Label Noise with Confidence Penalization

ECCV 2022poster

"Sample selection is an effective strategy to mitigate the effect of label noise in robust learning. Typical strategies commonly apply the small-loss criterion to identify clean samples. However, those samples lying around the decision boundary with large losses usually entangle with noisy examples,…

2021

RaP-Net: A Region-wise and Point-wise Weighting Network to Extract Robust Features for Indoor Localization

IROS 2021poster

Feature extraction plays an important role in visual localization. Unreliable features on dynamic objects or repetitive regions will interfere with feature matching and challenge indoor localization greatly. To address the problem, we propose a novel network, RaP-Net, to simultaneously predict regio…

Cited by 7SourcecodeScholar
2020

DXSLAM: A Robust and Efficient Visual SLAM System with Deep Features

IROS 2020poster

A robust and efficient Simultaneous Localization and Mapping (SLAM) system is essential for robot autonomy. For visual SLAM algorithms, though the theoretical framework has been well established for most aspects, feature extraction and association is still empirically designed in most cases, and can…

Cited by 155SourcecodeScholar
2019

Concrete: A Per-layer Configurable Framework for Evaluating DNN with Approximate Operators

ICASSP 2019accepted

Approximate computing has drawn considerable attention to both academia and industry in the area of DNN hardware. Despite substantial efforts to design approximate circuits and building blocks, the resilience of DNN layers and structures remains an untapped field to explore. This paper presents an e…

Cited by 0SourceScholar
2018

DS-SLAM: A Semantic Visual SLAM towards Dynamic Environments

IROS 2018poster

Simultaneous Localization and Mapping (SLAM) is considered to be a fundamental capability for intelligent mobile robots. Over the past decades, many impressed SLAM systems have been developed and achieved good performance under certain circumstances. However, some problems are still not well solved,…

Cited by 1126SourceScholar
2017

An inner-loop free solution to inverse problems using deep neural networks

NeurIPS 2017poster

We propose a new method that uses deep learning techniques to accelerate the popular alternating direction method of multipliers (ADMM) solution for inverse problems. The ADMM updates consist of a proximity operator, a least squares regression that includes a big matrix inversion, and an explicit so…

Cited by 27SourcePDFScholar
2017

Change detection between multi-band images using a robust fusion-based approach

ICASSP 2017accepted

This paper proposes a robust fusion-based strategy to detect changes between two multi-band optical images with different spatial and spectral resolutions, e.g., a multispectral high spatial resolution image and a hyperspectral low spatial resolution image. The dissimilarity between sensor resolutio…

Cited by 4SourceScholar
2016

A precision-improved processing architecture of physical computing for energy-efficient SIFT feature extraction

ICASSP 2016accepted

A precision-improved processing architecture of physical computing for energy-efficient SIFT feature extraction algorithm has been proposed in this paper. With the novel physical computing technology of active resistor network (PC: ARN), the SIFT algorithm could be processed in analog signal domain…

Cited by 0SourceScholar