← Search

Liang Liu

46 accepted papers

2026

Improving Batch Normalization with Test-Time Adaptation for Robust Object Detection in Self-Driving

AAAI 2026technical

In open real-world autonomous driving scenarios, challenges such as sensor failure and extreme weather hinder the generalization of current autonomous driving perception models to these unseen domain, due to the domain shifts between the test and training data. As the parameter scale of autonomous d

Cited by 0SourcePDFScholar
2026

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

AAAI 2026technical

The recent DeepSeek-R1 has showcased the emergence of reasoning capabilities in large language models (LLMs) through reinforcement learning (RL) with rule-based rewards. Despite its success in language tasks, its application in multimodal domains, particularly in graphic user interface (GUI) agent t

Cited by 0SourcePDFScholar
2025

A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos

ACL 2025finding

Large Reasoning Models (LRMs) have significantly advanced beyond traditional Large Language Models (LLMs) with their exceptional logical reasoning capabilities, yet these improvements introduce heightened safety risks. When subjected to jailbreak attacks, their ability to generate more targeted and…

2025

AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

ACL 2025finding

AI agents have drawn increasing attention mostly on their ability to perceive environments, understand tasks, and autonomously achieve goals. To advance research on AI agents in mobile scenarios, we introduce the Android Multi-annotation EXpo (AMEX), a comprehensive, large-scale dataset designed for…

Cited by 0SourcePDFScholar
2025

FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User Data

EMNLP 2025

Mobile GUI agents have attracted tremendous research participation recently. Traditional approaches to mobile agent training rely on centralized data collection, leading to high cost and limited scalability. Distributed training utilizing federated learning offers an alternative by harnessing real-w

2025

Identical-Delay Based 2-D DOA and Frequency Joint Estimation With Sub-Nyquist Sampling for URA

ICASSP 2025accepted

As spectrum congestion intensifies in wireless communication, efficient spectrum utilization through advanced sensing techniques has become increasingly important. This paper proposes a joint carrier frequency and two-dimensional (2-D) Direction of Arrival (DOA) estimation algorithm with signal reco…

Cited by 0SourceScholar
2025

MemFreezing: A Novel Adversarial Attack on Temporal Graph Neural Networks under Limited Future Knowledge

ICML 2025poster

Temporal graph neural networks (TGNN) have achieved significant momentum in many real-world dynamic graph tasks. While most existing TGNN attack methods assume worst-case scenarios where attackers have complete knowledge of the input graph, the assumption may not always hold in real-world situations…

Cited by 0SourcePDFScholar
2025

T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving

CVPR 2025poster

Understanding the traffic scenes and then generating high-definition (HD) maps present significant challenges in autonomous driving. In this paper, we defined a novel \underline T raffic \underline T opology \underline S cene \underline G raph (\text T ^2\text SG ), a unified scene graph explicitl…

2025

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents

NeurIPS 2025poster

In this paper, we introduce UI-Genie, a self-improving framework addressing two key challenges in GUI agents: verification of trajectory outcome is challenging and high-quality training data are not scalable. These challenges are addressed by a reward model and a self-improving pipeline, respectivel…

Cited by 0SourcecodeScholar
2024

A CCM-Based Joint DOA-Frequency Estimation and Signal Recovery with Efficient Sub-Nyquist Sampling

ICASSP 2024accepted

This paper addresses key challenges caused by high sampling rates in wideband joint spectrum sensing applications. A joint Direction of Arrival (DOA) and frequency estimation algorithm is proposed by utilizing the Cross-Covariance Matrix (CCM) constructed from the outputs of an efficient undersampli…

Cited by 0SourceScholar
2024

A Robotic-centric Paradigm for 3D Human Tracking Under Complex Environments Using Multi-modal Adaptation

IROS 2024poster

The goal of this paper is to strike a feasible tracking paradigm that can make 3D human trackers applicable on robot platforms and enable more high-level tasks. Till now, two fundamental problems haven’t been adequately addressed. One is the computational cost lightweight enough for robotic deployme…

Cited by 0SourceScholar
2024

AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model

AAAI 2024technical

Anomaly inspection plays an important role in industrial manufacture. Existing anomaly inspection methods are limited in their performance due to insufficient anomaly data. Although anomaly generation methods have been proposed to augment the anomaly data, they either suffer from poor generation aut…

2024

DOA Estimation for Switch-Element Arrays Based on Sparse Representation

ICASSP 2024accepted

In the context of perceiving spatial information, researchers extensively investigate the use of large-scale arrays due to their numerous advantages such as high precision and resolution, as well as increased degrees of freedom. However, large-scale arrays may be impractical in certain applications…

Cited by 0SourceScholar
2024

LEVI: Generalizable Fine-tuning via Layer-wise Ensemble of Different Views

ICML 2024poster

Fine-tuning is becoming widely used for leveraging the power of pre-trained foundation models in new downstream tasks. While there are many successes of fine-tuning on various tasks, recent studies have observed challenges in the generalization of fine-tuned models to unseen distributions (i.e., out…

Cited by 1SourcePDFScholar
2024

Learning Unified Reference Representation for Unsupervised Multi-class Anomaly Detection

ECCV 2024poster

"In the field of multi-class anomaly detection, reconstruction-based methods derived from single-class anomaly detection face the well-known challenge of “learning shortcuts”, wherein the model fails to learn the patterns of normal samples as it should, opting instead for shortcuts such as identity…

2024

Multi-modal 3D Human Tracking for Robots in Complex Environment with Siamese Point-Video Transformer

ICRA 2024poster

Tracking a specific person in 3D scene is gaining momentum due to its numerous applications in robotics. Currently, most 3D trackers focus on driving scenarios with neglected jitter and uncomplicated surroundings, which results in their severe degeneration in complex environments, especially on jolt…

Cited by 3SourceScholar
2024

Rethinking Reverse Distillation for Multi-Modal Anomaly Detection

AAAI 2024technical

In recent years, there has been significant progress in employing color images for anomaly detection in industrial scenarios, but it is insufficient for identifying anomalies that are invisible in RGB images alone. As a supplement, introducing extra modalities such as depth and surface normal maps c…

Cited by 16SourcePDFScholar
2024

SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

CVPR 2024poster

Multimodal Visual Object Tracking (VOT) has recently gained significant attention due to its robustness. Early research focused on fully fine-tuning RGB-based trackers which was inefficient and lacked generalized representation due to the scarcity of multimodal data. Therefore recent studies have ut…

2024

Self-Supervised Likelihood Estimation with Energy Guidance for Anomaly Segmentation in Urban Scenes

AAAI 2024technical

Robust autonomous driving requires agents to accurately identify unexpected areas (anomalies) in urban scenes. To this end, some critical issues remain open: how to design advisable metric to measure anomalies, and how to properly generate training samples of anomaly data? Classical effort in anomal…

2024

Self-supervised Feature Adaptation for 3D Industrial Anomaly Detection

ECCV 2024poster

"Industrial anomaly detection is generally addressed as an unsupervised task that aims at locating defects with only normal training samples. Recently, numerous 2D anomaly detection methods have been proposed and have achieved promising results, however, using only the 2D RGB data as input is not su…

2024

The LuViRA Dataset: Synchronized Vision, Radio, and Audio Sensors for Indoor Localization

ICRA 2024poster

We present a synchronized multisensory dataset for accurate and robust indoor localization: the Lund University Vision, Radio, and Audio (LuViRA) Dataset. The dataset includes color images, corresponding depth maps, inertial measurement unit (IMU) readings, channel response between a 5G massive mult…

Cited by 2SourcecodeScholar
2024

User-Assisted Networked Sensing in OFDM Cellular Network with Erroneous Anchor Position Information

ICASSP 2024accepted

In the sixth-generation (6G) integrated sensing and communication (ISAC) cellular network, base stations (BSs) can collaborate with each other to reap not only the cooperative communication gain, but also the networked sensing gain. In contrast to cooperative communication where both line-of-sight (…

Cited by 0SourceScholar
2023

Calibrated Teacher for Sparsely Annotated Object Detection

AAAI 2023technical

Fully supervised object detection requires training images in which all instances are annotated. This is actually impractical due to the high labor and time costs and the unavoidable missing annotations. As a result, the incomplete annotation in each image could provide misleading supervision and ha…

2023

Joint Data Association, NLOS Mitigation, and Clutter Suppression for Networked Device-Free Sensing in 6G Cellular Network

ICASSP 2023accepted

Recently, there is a growing interest in achieving integrated sensing and communication (ISAC) in the sixth-generation (6G) cellular network. Inspired by the success of cooperative communication in cloud radio access network, this paper considers a networked device-free sensing architecture based on…

Cited by 0SourceScholar
2023

Learning From Noisy Labels With Decoupled Meta Label Purifier

CVPR 2023poster

Training deep neural networks (DNN) with noisy labels is challenging since DNN can easily memorize inaccurate labels, leading to poor generalization ability. Recently, the meta-learning based label correction strategy is widely adopted to tackle this problem via identifying and correcting potential…

2023

MixTeacher: Mining Promising Labels With Mixed Scale Teacher for Semi-Supervised Object Detection

CVPR 2023poster

Scale variation across object instances is one of the key challenges in object detection. Although modern detection models have achieved remarkable progress in dealing with the scale variation, it still brings trouble in the semi-supervised case. Most existing semi-supervised object detection method…

2023

Phasic Content Fusing Diffusion Model with Directional Distribution Consistency for Few-Shot Model Adaption

ICCV 2023poster

Training a generative model with limited number of samples is a challenging task. Current methods primarily rely on few-shot model adaption to train the network. However, in scenarios where data is extremely limited (less than 10), the generative network tends to overfit and suffers from content deg…

Cited by 14PDFcodeScholar
2023

Remembering Normality: Memory-guided Knowledge Distillation for Unsupervised Anomaly Detection

ICCV 2023poster

Knowledge distillation (KD) has been widely explored in unsupervised anomaly detection (AD). The student is assumed to constantly produce representations of typical patterns within trained data, named "normality", and the representation discrepancy between the teacher and student model is identified…

Cited by 48PDFScholar
2023

Rethinking Mobile Block for Efficient Attention-based Models

ICCV 2023poster

This paper focuses on developing modern, efficient, lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Inverted Residual Block (IRB) serves as the infrastructure for lightweight CNNs, but no counterpart has been recognized by attention-based studies. This…

Cited by 180PDFcodeScholar
2023

Understanding and Defending Patched-based Adversarial Attacks for Vision Transformer

ICML 2023poster

Vision Transformer (ViT) is an attention-based model architecture that has demonstrated superior performance on many computer vision tasks. However, its security properties, in particular, the robustness against adversarial attacks, are yet to be thoroughly studied. Recent works have shown that ViT…

Cited by 5SourcePDFScholar
2022

Defending Against Universal Attack Via Curvature-Aware Category Adversarial Training

ICASSP 2022accepted

Adversarial training can defend against universal adversarial perturbation (UAP) by injecting corresponding adversarial samples during training. However, adversarial samples used by existing methods, such as UAP, inevitably include excessive perturbations related to other categories due to its inher…

Cited by 0SourceScholar
2022

Efficiently and Globally Solving Joint Beamforming and Compression Problem in the Cooperative Cellular Network Via Lagrangian Duality

ICASSP 2022accepted

Consider the joint beamforming and quantization problem in the cooperative cellular network, where multiple relay-like base stations (BSs) connected to the central processor (CP) via rate-limited fronthaul links cooperatively serve the users. This problem can be formulated as the minimization of the…

Cited by 0SourceScholar
2022

ISDNet: Integrating Shallow and Deep Networks for Efficient Ultra-High Resolution Segmentation

CVPR 2022poster

The huge burden of computation and memory are two obstacles in ultra-high resolution image segmentation. To tackle these issues, most of the previous works follow the global-local refinement pipeline, which pays more attention to the memory consumption but neglects the inference speed. In comparison…

Cited by 60PDFcodeScholar
2022

Iterative Few-shot Semantic Segmentation from Image Label Text

IJCAI 2022poster

Few-shot semantic segmentation aims to learn to segment unseen class objects with the guidance of only a few support images. Most previous methods rely on the pixel-level label of support images. In this paper, we focus on a more challenging setting, in which only the image-level labels are availabl…

2021

An Efficient Algorithm For Device Detection And Channel Estimation In Asynchronous IOT Systems

ICASSP 2021accepted

A great amount of endeavour has recently been devoted to the joint device activity detection and channel estimation problem in massive machine-type communications. This paper targets at two practical issues along this line that have not been addressed before: asynchronous transmission from uncoordin…

Cited by 0SourceScholar
2021

HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation

AAAI 2021technical

Self-supervised learning shows great potential in monocular depth estimation, using image sequences as the only source of supervision. Although people try to use the high-resolution image for depth estimation, the accuracy of prediction has not been significantly improved. In this work…

2020

APB2FACE: Audio-Guided Face Reenactment with Auxiliary Pose and Blink Signals

ICASSP 2020accepted

Audio-guided face reenactment aims at generating photorealistic faces using audio information while maintaining the same facial movement as when speaking to a real person. However, existing methods can not generate vivid face images or only reenact low-resolution faces, which limits the application…

Cited by 0SourceScholar
2020

DTVNet: Dynamic Time-lapse Video Generation via Single Still Image

ECCV 2020poster

This paper presents a novel end-to-end dynamic time-lapse video generation framework, named DTVNet, to generate diversified time-lapse videos from a single landscape image, which are conditioned on normalized motion vectors. The proposed DTVNet consists of two submodules: mph{Optical Flow Encoder} (…

2020

Learning by Analogy: Reliable Supervision From Transformations for Unsupervised Optical Flow Estimation

CVPR 2020poster

Unsupervised learning of optical flow, which leverages the supervision from view synthesis, has emerged as a promising alternative to supervised methods. However, the objective of unsupervised learning is likely to be unreliable in challenging scenes. In this work, we present a framework to use more…

Cited by 213PDFcodeScholar
2020

Weighing Counts: Sequential Crowd Counting by Reinforcement Learning

ECCV 2020poster

We formulate counting as a sequential decision problem and present a novel crowd counting model solvable by deep reinforcement learning. In contrast to existing counting models that directly output count values, we divide one-step estimation into a sequence of much easier and more tractable sub-deci…

Cited by 97SourcePDFScholar
2019

From Open Set to Closed Set: Counting Objects by Spatial Divide-and-Conquer

ICCV 2019poster

Visual counting, a task that predicts the number of objects from an image/video, is an open-set problem by nature, i.e., the number of population can vary in [0,+[?]) in theory. However, the collected images and labeled count values are limited in reality, which means only a small closed set is obse…

Cited by 214PDFcodeScholar