← Search

Jiawei Shao

13 accepted papers

2026

GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents

AAAI 2026technical

Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot visual inputs and passive perception, lacking the ability to adaptively determine when, whether, and how to observe the int

Cited by 0SourcePDFScholar
2026

Streaming Video Crime Anticipation with Spatio-Temporal Causal Reasoning

CVPR 2026

Crime anticipation enables proactive public safety interventions, yet existing video security systems remain largely reactive, unable to detect precursors of crime. While current visual language models (VLM)-based video understanding methods show promise in high-level reasoning, they are not designe

Cited by 0SourceScholar
2025

Be Confident: Uncovering Overfitting in MLLM Multi-Task Tuning

ICML 2025poster

Fine-tuning Multimodal Large Language Models (MLLMs) in multi-task learning scenarios has emerged as an effective strategy for achieving cross-domain specialization. However, multi-task fine-tuning frequently induces performance degradation on open-response datasets. We posit that free-form answer g…

Cited by 0SourcePDFScholar
2025

CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression

AAAI 2025technical

Existing learning-based stereo image codec adopt sophisticated transformation with simple entropy models derived from single image codecs to encode latent representations. However, those entropy models struggle to effectively capture the spatial-disparity characteristics inherent in stereo images, w…

2025

CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs

NeurIPS 2025poster

Speculative decoding has become a widely adopted as an effective technique for lossless inference acceleration when deploying large language models (LLMs). While on-the-fly self-speculative methods offer seamless integration and broad utility, they often fall short of the speed gains achieved by met…

Cited by 0SourceScholar
2025

Graph Neural Network Enhanced Retrieval for Question Answering of Large Language Models

NAACL 2025long

Retrieval augmented generation has revolutionized large language model (LLM) outputs by providing factual supports. Nevertheless, it struggles to capture all the necessary knowledge for complex reasoning questions. Existing retrieval methods typically divide reference documents into passages, treati…

2025

KV-Edit: Training-Free Image Editing for Precise Background Preservation

ICCV 2025poster

Background consistency remains a significant challenge in image editing tasks. Despite extensive developments, existing works still face a trade-off between maintaining similarity to the original image and generating content that aligns with the target. Here, we propose KV-Edit, a training-free appr…

2025

What You Have is What You Track: Adaptive and Robust Multimodal Tracking

ICCV 2025poster

Multimodal data is known to be helpful for visual tracking by improving robustness to appearance variations. However, sensor synchronization challenges often compromise data availability, particularly in video settings where shortages can be temporal. Despite its importance, this area remains undere…

2024

Bidirectional Stereo Image Compression with Cross-Dimensional Entropy Model

ECCV 2024poster

"With the rapid advancement of stereo vision technologies, stereo image compression has emerged as a crucial field that continues to draw significant attention. Previous approaches have primarily employed a unidirectional paradigm, where the compression of one view is dependent on the other, resulti…

2024

Embodied Intelligence: Bionic Robot Controller Integrating Environment Perception, Autonomous Planning, and Motion Control

RA-L 2024

This letter proposes a bionic robot controller equipped with intelligent perception and autonomous planning modules to address the manufacturing industry's requirements for small-batch, customized, and autonomous task. Three crucial components: motion control module, vision perception module, and au

Cited by 18SourceScholar
2022

DReS-FL: Dropout-Resilient Secure Federated Learning for Non-IID Clients via Secret Data Sharing

NeurIPS 2022accept

Federated learning (FL) strives to enable collaborative training of machine learning models without centrally collecting clients' private data. Different from centralized training, the local datasets across clients in FL are non-independent and identically distributed (non-IID). In addition, the dat…

Cited by 47SourcePDFScholar
2021

Branchy-GNN: A Device-Edge Co-Inference Framework for Efficient Point Cloud Processing

ICASSP 2021accepted

The recent advancements of three-dimensional (3D) data acquisition devices have spurred a new breed of applications that rely on point cloud data processing. However, processing a large volume of point cloud data brings a significant workload on resource-constrained mobile devices, prohibiting from…

Cited by 0SourceScholar