← Search

Zihao Liu

17 accepted papers

2026

DiGraphHal-Bench: Evaluating Multimodal Large Language Models on Complex Directed Graphs

CVPR 2026

While prior research on Multimodal Large Language Model (MLLM) hallucinations has primarily examined cross-modal inconsistencies in natural images, hallucination over complex graph structures remains underexplored.Concurrently, there is a lack of robust evaluation for fine-grained reasoning integrat

Cited by 0SourcecodeScholar
2026

Language-guided Open-world Video Anomaly Detection under Weak Supervision

ICLR 2026poster

Video anomaly detection (VAD) aims to detect anomalies that deviate from what is expected. In open-world scenarios, the expected events may change as requirements change. For example, not wearing a mask may be considered abnormal during a flu outbreak but normal otherwise. However, existing methods…

Cited by 0SourcecodeScholar
2026

STAGE: STyle-Controllable Action GEneration for Personalized Autonomous Driving

RA-L 2026

Driving style refers to the behavioral preferences that drivers maintain during driving, shaped by their diverse experiences, habits, and needs, and is typically reflected in varying levels of aggressiveness. If humans choose to use autonomous driving systems, they would expect the driving style of

Cited by 0SourcecodeScholar
2026

STAGE: STyle-Controllable Action GEneration for Personalized Autonomous Driving

ICRA 2026poster

Driving style refers to the behavioral preferences that drivers maintain during driving, shaped by their diverse experiences, habits, and needs, and is typically reflected in varying levels of aggressiveness. If humans choose to use autonomous driving systems, they would expect the driving style of …

2026

Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training

AAAI 2026technical

Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated datasets. Although some zero-shot methods attempt to trajectory control in the l

Cited by 0SourcePDFScholar
2025

Frequency-Biased Synergistic Design for Image Compression and Compensation

CVPR 2025poster

Compression artifacts removal (CAR), an effective post-processing method to reduce compression distortion in edge-side codecs, demonstrates remarkable results by utilizing convolutional neural networks (CNNs) on high computational power cloud side. Traditional image compression reduces redundancy in…

Cited by 0SourcePDFScholar
2024

Asynchronous Event-Inertial Odometry using a Unified Gaussian Process Regression Framework

IROS 2024poster

Recent works have combined monocular event camera and inertial measurement unit to estimate the SE(3) trajectory. However, the asynchronicity of event cameras brings a great challenge to conventional fusion algorithms. In this paper, we present an asynchronous event-inertial odometry under a unified…

Cited by 2SourceScholar
2024

Enhancing Vectorized Map Perception with Historical Rasterized Maps

ECCV 2024poster

"In autonomous driving, there is growing interest in end-to-end online vectorized map perception in bird’s-eye-view (BEV) space, with an expectation that it could replace traditional high-cost offline high-definition (HD) maps. However, the accuracy and robustness of these methods can be easily comp…

2024

Leveraging Enhanced Queries of Point Sets for Vectorized Map Construction

ECCV 2024poster

"In autonomous driving, the high-definition (HD) map plays a crucial role in localization and planning. Recently, several methods have facilitated end-to-end online map construction in DETR-like frameworks. However, little attention has been paid to the potential capabilities of exploring the query…

2024

Self-reconfiguration Strategies for Space-distributed Spacecraft

IROS 2024poster

This paper proposes a distributed on-orbit spacecraft assembly algorithm, where future spacecraft can assemble modules with different functions on orbit to form a spacecraft structure with specific functions. This form of spacecraft organization has the advantages of reconfigurability, fast mission…

Cited by 0SourceScholar
2024

Tactile Active Inference Reinforcement Learning for Efficient Robotic Manipulation Skill Acquisition

IROS 2024poster

Robotic manipulation holds the potential to replace humans in the execution of tedious or dangerous tasks. However, control-based approaches are not suitable due to the difficulty of formally describing open-world manipulation in reality, and the inefficiency of existing learning methods. Therefore,…

Cited by 1SourceScholar
2024

Zero-Shot Structure-Preserving Diffusion Model for High Dynamic Range Tone Mapping

CVPR 2024highlight

Tone mapping techniques aiming to convert high dynamic range (HDR) images to high-quality low dynamic range (LDR) images for display play a more crucial role in real-world vision systems with the increasing application of HDR images. However obtaining paired HDR and high-quality LDR images is diffic…

2020

ShiftAddNet: A Hardware-Inspired Deep Network

NeurIPS 2020poster

Multiplication (e.g., convolution) is arguably a cornerstone of modern deep neural networks (DNNs). However, intensive multiplications cause expensive resource costs that challenge DNNs' deployment on resource-constrained edge devices, driving several attempts for multiplication-less deep networks.…

2019

Feature Distillation: DNN-Oriented JPEG Compression Against Adversarial Examples

CVPR 2019poster

Image compression-based approaches for defending against the adversarial-example attacks, which threaten the safety use of deep neural networks (DNN), have been investigated recently. However, prior works mainly rely on directly tuning parameters like compression rate, to blindly reduce image featur…

Cited by 337PDFScholar
2019

Machine Vision Guided 3D Medical Image Compression for Efficient Transmission and Accurate Segmentation in the Clouds

CVPR 2019poster

Cloud based medical image analysis has become popular recently due to the high computation complexities of various deep neural network (DNN) based frameworks and the increasingly large volume of medical images that need to be processed. It has been demonstrated that for medical images the transmissi…

Cited by 48PDFScholar
2019

Online Single Person Tracking for Unmanned Aerial Vehicles: Benchmark and New Baseline

ICASSP 2019accepted

Online tracking a specific person from a low-altitude unmanned aerial vehicle (UAV) is a very interesting and challenging problem to be solved. However, there exists no large-scale aerial video dataset regarding this online single person tracking (OSPT) task. To promote the study of the OSPT problem…

Cited by 0SourceScholar