← Search

Zeyu Liu

24 accepted papers

2026

InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization

AAAI 2026technical

The emergence of Multimodal Large Language Models (MLLMs) has propelled the development of autonomous agents that operate on Graphical User Interfaces (GUIs) using pure visual input. A fundamental challenge is robustly grounding natural language instructions. This requires a precise spatial alignmen

Cited by 0SourcePDFScholar
2026

Linearizing Vision Transformer with Test-Time Training

ICML 2026poster

While linear-complexity attention mechanisms offer a promising alternative to Softmax attention for overcoming the quadratic bottleneck, training such models from scratch remains prohibitively expensive. Inheriting weights from pretrained Transformers provides an appealing shortcut, yet the fundamen…

Cited by 0SourceScholar
2026

Local Covariate Selection for Average Causal Effect Estimation without Pretreatment and Causal Sufficiency Assumptions

ICML 2026spotlight

Causal effect estimation is a fundamental task in many scientific fields. Selecting appropriate covariates for adjustment is crucial for obtaining unbiased causal effects. However, most existing methods either rely on learning the global causal structure, assume the absence of latent variables, or i…

Cited by 0SourceScholar
2025

CODA: Repurposing Continuous VAEs for Discrete Tokenization

ICCV 2025poster

Discrete visual tokenizers transform images into a sequence of tokens, enabling token-based visual generation akin to language models. However, this process is inherently challenging, as it requires both compressing visual signals into a compact representation and discretizing them into a fixed set…

Cited by 0SourcePDFScholar
2025

Dynamic SpikFormer: Low-Latency & Energy-Efficient Spiking Neural Networks with Dynamic Time Steps for Vision Transformers

ICASSP 2025accepted

Spiking Neural Networks (SNNs) have emerged as a popular spatio-temporal computing paradigm for complex vision tasks. Recently proposed SNN training algorithms have significantly reduced the number of time steps (down to 1) for improved latency and energy efficiency, however, they target only convol…

Cited by 0SourceScholar
2025

EIC Framework for Hand Exoskeletons Based on a Multimodal Large Language Model

IROS 2025

Current hand exoskeleton interaction methods primarily focus on recognizing a limited range of hand motion intentions and rely on pre-programmed control to execute predefined commands. However, these approaches face significant limitations when confronted with unanticipated or non-predefined scenari

Cited by 1SourceScholar
2025

Joint Scheduling of Causal Prompts and Tasks for Multi-Task Learning

CVPR 2025poster

Multi-task prompt learning has emerged as a promising technique for fine-tuning pre-trained Vision-Language Models (VLMs) to various downstream tasks. However, existing methods ignore challenges caused by spurious correlations and dynamic task relationships, which may reduce the model performance. T…

Cited by 0SourcePDFScholar
2025

LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling

EMNLP 2025

Although transformer architectures have achieved state-of-the-art performance across diverse domains, their quadratic computational complexity with respect to sequence length remains a significant bottleneck, particularly for latency-sensitive long-context applications. While recent linear-complexit

Cited by 0SourcePDFScholar
2025

Local Identifying Causal Relations in the Presence of Latent Variables

ICML 2025spotlight

We tackle the problem of identifying whether a variable is the cause of a specified target using observational data. State-of-the-art causal learning algorithms that handle latent variables typically rely on identifying the global causal structure, often represented as a partial ancestral graph (PAG…

Cited by 0SourcePDFScholar
2024

AFLoRA: Adaptive Freezing of Low Rank Adaptation in Parameter Efficient Fine-Tuning of Large Models

ACL 2024short

We present a novel Parameter-Efficient Fine-Tuning (PEFT) method, dubbed as Adaptive Freezing of Low-Rank Adaptation (AFLoRA). Specifically, for each pre-trained frozen weight tensor, we add a parallel path of trainable low-rank matrices, namely a down-projection and an up-projection matrix, each of…

Cited by 14SourcePDFScholar
2024

Can we get the best of both Binary Neural Networks and Spiking Neural Networks for Efficient Computer Vision?

ICLR 2024poster

Binary Neural networks (BNN) have emerged as an attractive computing paradigm for a wide range of low-power vision tasks. However, state-of-the-art (SOTA) BNNs do not yield any sparsity, and induce a significant number of non-binary operations. On the other hand, activation sparsity can be provided…

2024

FOCWS: A High Sensitive Flexible Optical Curvature Sensor Inspired by Arthropod Sensory Systems

IROS 2024poster

Flexible sensors for joint angle measurement play a crucial role in various human-robot interaction applications. In previous studies, sensors with various sensing mechanisms have been developed. Among them, optical waveguide sensors exhibit high resistance to environmental factors (such as temperat…

Cited by 0SourceScholar
2024

Glyph-ByT5: A Customized Text Encoder for Accurate Visual Text Rendering

ECCV 2024poster

"Visual text rendering poses a fundamental challenge for contemporary text-to-image generation models, with the core problem lying in text encoder deficiencies. To achieve accurate text rendering, we identify two crucial requirements for text encoders: character awareness and alignment with glyphs.…

2024

LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory Units

ICLR 2024poster

Transformer models have demonstrated high accuracy in numerous applications but have high complexity and lack sequential processing capability making them ill-suited for many streaming applications at the edge where devices are heavily resource-constrained. Thus motivated, many researchers have prop…

2023

A Two-Dimensional Reticular Core Optical Waveguide Sensor for Tactile and Positioning Sensing

IROS 2023poster

Tactile sensors based on optical waveguides are highly sensitive to pressure, possess good chemical inertness and electromagnetic resistance, and are unaffected by temperature changes in the surrounding environment. Researchers have developed various waveguide structures with multi-level cores to si…

Cited by 3SourceScholar
2023

Dynamic Perceiver for Efficient Visual Recognition

ICCV 2023poster

Early exiting has become a promising approach to im- proving the inference efficiency of deep networks. By structuring models with multiple classifiers (exits), predictions for "easy" samples can be generated at earlier exits, negating the need for executing deeper layers. Current multi-exit network…

Cited by 36PDFcodeScholar
2023

FeatureBooster: Boosting Feature Descriptors With a Lightweight Neural Network

CVPR 2023poster

We introduce a lightweight network to improve descriptors of keypoints within the same image. The network takes the original descriptors and the geometric properties of keypoints as the input, and uses an MLP-based self-boosting stage and a Transformer-based cross-boosting stage to enhance the descr…

2023

In-Sensor & Neuromorphic Computing Are all You Need for Energy Efficient Computer Vision

ICASSP 2023accepted

Due to the high activation sparsity and use of accumulates (AC) instead of expensive multiply-and-accumulates (MAC), neuromorphic spiking neural networks (SNNs) have emerged as a promising low-power alternative to traditional DNNs for several computer vision (CV) applications. However, most existing…

Cited by 0SourceScholar
2023

Towards A Unified View of Sparse Feed-Forward Network in Pretraining Large Language Model

EMNLP 2023long main

Large and sparse feed-forward layers (S-FFN) such as Mixture-of-Experts (MoE) have proven effective in scaling up Transformers model size for pretraining large language models. By only activating part of the FFN parameters conditioning on input, S-FFN improves generalization performance while keepin…

Cited by 0SourceScholar
2022

Synthesizing Diverse and Physically Stable Grasps With Arbitrary Hand Structures Using Differentiable Force Closure Estimator

RA-L 2022

Existing grasp synthesis methods are either analytical or data-driven. The former one is oftentimes limited to specific application scope. The latter one depends heavily on demonstrations, thus suffers from generalization issues; <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="ht

Cited by 155SourceScholar
2021

Probing Across Time: What Does RoBERTa Know and When?

EMNLP 2021finding

Models of language trained on very large corpora have been demonstrated useful for natural language processing. As fixed artifacts, they have become the object of intense study, with many researchers “probing” the extent to which they acquire and readily demonstrate linguistic abstractions, factual…

2021

StructDepth: Leveraging the Structural Regularities for Self-Supervised Indoor Depth Estimation

ICCV 2021poster

Self-supervised monocular depth estimation has achieved impressive performance on outdoor datasets. Its performance however degrades notably in indoor environments because of the lack of textures. Without rich textures, the photometric consistency is too weak to train a good depth network. Inspired…

Cited by 78PDFcodeScholar