← Search

Qingyi Gu

20 accepted papers

2026

Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval

ICLR 2026poster

Segment Anything Model 2 (SAM2) shows excellent performance in video object segmentation tasks; however, the heavy computational burden hinders its application in real-time video processing. Although there have been efforts to improve the efficiency of SAM2, most of them focus on retraining a lightw…

Cited by 0SourceScholar
2026

K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge

ICLR 2026poster

The rapid development of visual generative models raises the need for more scalable and human-aligned evaluation methods. While the crowdsourced Arena platforms offer human preference assessments by collecting human votes, they are costly and time-consuming, inherently limiting their scalability. Le…

Cited by 0SourcecodeScholar
2026

MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic Segmentation

AAAI 2026technical

Semantic segmentation is a fundamental task in computer vision with wide-ranging applications, including autonomous driving and robotics. While RGB-based methods have achieved strong performance with CNNs and Transformers, their effectiveness degrades under fast motion, low-light, or high dynamic ra

Cited by 0SourcePDFScholar
2026

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization

ICML 2026poster

Large Language Models (LLMs) have demonstrated remarkable capabilities in understanding and generation tasks. However, their massive parameter scale leads to significant resource consumption and latency during inference. Post-training weight-only quantization offers a promising solution by reducing …

Cited by 0SourceScholar
2026

PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models

ICLR 2026poster

AutoRegressive Visual Generation (ARVG) models retain an architecture compatible with language models, while achieving performance comparable to diffusion-based models. Quantization is commonly employed in neural networks to reduce model size and computational latency. However, applying quantization…

Cited by 0SourcecodeScholar
2026

SAQ-SAM: Semantically-Aligned Quantization for Segment Anything Model

AAAI 2026technical

Segment Anything Model (SAM) exhibits remarkable zero-shot segmentation capability; however, its prohibitive computational costs make edge deployment challenging. Although post-training quantization (PTQ) offers a promising compression solution, existing methods yield unsatisfactory results when app

Cited by 0SourcePDFScholar
2025

K-Sort Arena: Efficient and Reliable Benchmarking for Generative Models via K-wise Human Preferences

CVPR 2025poster

The rapid advancement of visual generative models necessitates efficient and reliable evaluation methods. Arena platform, which gathers user votes on model comparisons, can rank models with human preferences. However, traditional Arena methods, while established, require an excessive number of compa…

Cited by 4SourcePDFScholar
2025

SLTNet: Efficient Event-based Semantic Segmentation with Spike-driven Lightweight Transformer-based Networks

IROS 2025

Event-based semantic segmentation has great potential in autonomous driving and robotics due to the advantages of event cameras, such as high dynamic range, low latency, and low power cost. Unfortunately, current artificial neural network (ANN)-based segmentation methods suffer from high computation

Cited by 1SourcecodeScholar
2024

A Novel Wide-Area Multiobject Detection System with High-Probability Region Searching

ICRA 2024poster

In recent years, wide-area visual surveillance systems have been widely applied in various industrial and transportation scenarios. These systems, however, face significant challenges when implementing multi-object detection due to conflicts arising from the need for high-resolution imaging, efficie…

Cited by 5SourceScholar
2023

RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision Transformers

ICCV 2023poster

Post-training quantization (PTQ), which only requires a tiny dataset for calibration without end-to-end retraining, is a light and practical model compression technique. Recently, several PTQ schemes for vision transformers (ViTs) have been presented; unfortunately, they typically suffer from non-tr…

Cited by 105PDFcodeScholar
2022

A Flexible Calibration Algorithm for High-speed Bionic Vision System based on Galvanometer

IROS 2022poster

Traditional gimbal-based bionic eye systems usually use a multi-degree-of-freedom mechanical platform to move the camera freely, which makes the structure complex and bulky. The galvanometer-based reflective bionic eye system uses a galvanometer to replace the traditional mechanical rotation structu…

Cited by 8SourceScholar
2022

Patch Similarity Aware Data-Free Quantization for Vision Transformers

ECCV 2022poster

"Vision transformers have recently gained great success on various computer vision tasks; nevertheless, their high model complexity makes it challenging to deploy on resource-constrained devices. Quantization is an effective approach to reduce model complexity, and data-free quantization, which can…

2020

Angle-based Search Space Shrinking for Neural Architecture Search

ECCV 2020poster

In this work, we present a simple and general search space shrinking method, called Angle-Based search space Shrinking (ABS), for Neural Architecture Search (NAS). Our approach progressively simplifies the original search space by dropping unpromising candidates, thus can reduce difficulties for exi…

Cited by 83SourcePDFScholar
2020

Natural Scene Facial Expression Recognition with Dimension Reduction Network

ICRA 2020poster

As an external manifestation of human emotions, expression recognition plays an important role in human-computer interaction. Although existing expression recognition methods performs perfectly on constrained frontal faces, there are still many challenges in expression recognition in natural scenes…

Cited by 1SourceScholar
2017

12,000-fps Multi-object detection using HOG descriptor and SVM classifier

IROS 2017poster

This paper describes a high-frame-rate (HFR) vision system that can detect multiple objects in an image of 512 × 512 pixels at 12,000 frames per seconds (fps). An optimized algorithm is proposed based on conventional Histograms of Oriented Gradient (HOG) descriptor and Support Vector Machine (SVM) c…

Cited by 13SourceScholar
2016

Control scheme of nongrasping manipulation based on virtual connecting constraint

ICRA 2016

The research field of nongrasping manipulation is a maturing area in robotic motion control. However, the common principles of motion planning for nongrasping manipulation systems have not yet been established. This paper proposes the concept of virtual connecting manipulation as a generalized motio

Cited by 1SourceScholar
2015

A scheme for manipulating a passive object using an active plate

ICRA 2015poster

We propose a novel scheme for manipulating a passive object using an active plate. The objective of this study is to control an object's orientation with respect to the gravitational force direction by using an active plate for realizing hitherto unrealized object motion. In this context, motions of…

Cited by 3SourceScholar
2015

Realization of flower stick rotation using robotic arm

IROS 2015poster

Flower stick juggling is a dexterous task done by skillful jugglers. We aim to realize dexterous tasks done by humans using robotic systems. This work focuses on flower stick juggling and proposes a feedback control strategy for a flower stick juggling task called “propeller” as one of the robotic d…

Cited by 18SourceScholar