← Search

Yufan Liu

21 accepted papers

2026

AutoQVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization

ICLR 2026poster

The advent of Vision-Language-Action (VLA) models represents a significant leap for embodied intelligence, yet their immense computational demands critically hinder deployment on resource-constrained robotic platforms. Intuitively, low-bit quantization is a prevalent and preferred technique for larg…

Cited by 0SourcecodeScholar
2026

Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream Tasks

AAAI 2026technical

In recent years, the development of burst imaging technology has improved the capture and processing capabilities of visual data, enabling a wide range of applications. However, the redundancy in burst images leads to the increased storage and transmission demands, as well as reduced efficiency of d

Cited by 0SourcePDFScholar
2026

GauSem-SLAM: Gaussian Semantic Submaps with Loop Closure for Globally Consistent SLAM

ICRA 2026poster

3DGS has shown outstanding performance in multi-view geometry, driving its adoption in visual SLAM. However, real-time semantic 3DGS mapping faces challenges. Current methods typically treat semantics as external priors, making it hard to integrate them into SLAM tracking or loop closure correction.…

Cited by 0Scholar
2025

AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts

ICCV 2025poster

Despite rapid advancements in text-to-image (T2I) models, their safety mechanisms are vulnerable to adversarial prompts, which maliciously generate unsafe images. Current red-teaming methods for proactively assessing such vulnerabilities usually require white-box access to T2I models, and rely on in…

Cited by 0SourcePDFScholar
2025

Each Complexity Deserves a Pruning Policy

NeurIPS 2025poster

The established redundancy in visual tokens within large vision–language models (LVLMs) allows for pruning to effectively reduce their substantial computational demands. Empirical evidence from previous works indicates that visual tokens in later decoder stages receive less attention than shallow la…

Cited by 0SourcecodeScholar
2025

Full Body Motion Control for Aerial Manipulation Based on Critical Dynamic Constraints

RA-L 2025

This paper is concerned with the motion control problem of an aerial manipulator that consists of a quadrotor base and a manipulator. We propose a novel control strategy that incorporates a simplified model considering the critical dynamics, and a linear-MPC-based trajectory converter to ensure prec

Cited by 3SourceScholar
2025

GaussR-SLAM: Gaussian Robust SLAM in Data Loss and Interference Environments

RA-L 2025

Recent advancements in 3DGS-based explicit mapping have significantly improved SLAM performance, achieving more realistic environment reconstruction and faster processing. However, issues such as data loss caused by unstable data transmission, textureless and repetitive-texture often occur in real-w

Cited by 0SourceScholar
2025

MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural Networks

NeurIPS 2025poster

Brain-inspired spiking neural networks (SNNs) provide energy-efficient computation through event-driven processing. However, the shared weights across multiple timesteps lead to serious temporal feature redundancy, limiting both efficiency and performance. This issue is further aggravated when proce…

Cited by 0SourcecodeScholar
2025

Visual-Instructed Degradation Diffusion for All-in-One Image Restoration

CVPR 2025poster

Image restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose Defusion, a novel all-in-one image restoration framework that…

2025

WuKong: Design, Modeling and Control of a Compact Flexible Hybrid Aerial-Aquatic Vehicle

RA-L 2025

The significant differences in the physical properties of air and water pose a substantial challenge for the development of hybrid aerial-aquatic vehicle (HAAV), which leading to increased prototype size, heavier thrusters, and reduced efficiency or under-actuation in one of the mediums. This letter

Cited by 9SourceScholar
2024

Learning to Route Among Specialized Experts for Zero-Shot Generalization

ICML 2024poster

Recently, there has been a widespread proliferation of "expert" language models that are specialized to a specific task or domain through parameter-efficient fine-tuning. How can we recycle large collections of expert language models to improve zero-shot generalization to unseen tasks? In this work,…

2023

AUNet: Learning Relations Between Action Units for Face Forgery Detection

CVPR 2023poster

Face forgery detection becomes increasingly crucial due to the serious security issues caused by face manipulation techniques. Recent studies in deepfake detection have yielded promising results when the training and testing face forgeries are from the same domain. However, the problem remains chall…

Cited by 56SourcePDFScholar
2023

Learning from the Raw Domain: Cross Modality Distillation for Compressed Video Action Recognition

ICASSP 2023accepted

Video action recognition is faced with the challenges of both huge computation burden and performance requirements. Using compressed domain data, which saves much decoding computation, is a possible solution. Unfortunately, existing compressed-domain-based (CD) methods fail to obtain high performanc…

Cited by 0SourceScholar
2023

Nasty-SFDA: Source Free Domain Adaptation from a Nasty Model

ICASSP 2023accepted

A challenging problem called Nasty Source Free Domain Adaptation (Nasty-SFDA) is proposed in this work, where only a nasty source model and unlabeled target samples are available for DA. Further, after DA, the target model is expected to be a nasty model. In order to deal with Nasty-SFDA, Nasty HypO…

Cited by 0SourceScholar
2022

Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video Classification

IJCAI 2022poster

Compared with image few-shot learning, most of the existing few-shot video classification methods perform worse on feature matching, because they fail to sufficiently exploit the temporal information and relation. Specifically, frames are usually evenly sampled, which may miss important frames. On t…

Cited by 16SourcePDFScholar
2021

DPFPS: Dynamic and Progressive Filter Pruning for Compressing Convolutional Neural Networks from Scratch

AAAI 2021technical

Filter pruning is a commonly used method for compressing Convolutional Neural Networks (ConvNets), due to its friendly hardware supporting and flexibility. However, existing methods mostly need a cumbersome procedure, which brings many extra hyper-parameters and training epochs. This is because only…

2020

Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model

ECCV 2020poster

Recently, video streams have occupied a large proportion of Internet traffic, most of which contain human faces. Hence, it is necessary to predict saliency on multiple-face videos, which can provide attention cues for many content based applications. However, most of multiple-face prediction works o…

2019

Knowledge Distillation via Instance Relationship Graph

CVPR 2019poster

The key challenge of knowledge distillation is to extract general, moderate and sufficient knowledge from a teacher network to guide a student network. In this paper, a novel Instance Relationship Graph (IRG) is proposed for knowledge distillation. It models three kinds of knowledge, including insta…

Cited by 371PDFScholar