← Search

Dan Zhao

14 accepted papers

2026

Imagine with Layout and Sketch: Enhancing Vision-Language Retrieval with Dual-Stream Multi-Modal Query Refinement

AAAI 2026technical

Vision-Language Retrieval (VLR) aims to retrieve relevant visual or textual information from multimodal data using language or image queries. However, traditional VLR methods often rely on data-driven shallow semantic alignment and fail to understand the deeper structural and fine-grained entity fea

Cited by 0SourcePDFScholar
2026

Suit the Remedy to the Retriever: Interpretable Query Optimization with Retriever Preference Alignment for Vision-Language Retrieval

AAAI 2026technical

Vision-language retrieval (VLR), which uses text or image queries to retrieve corresponding cross-modal content, plays a crucial role in multimedia and computer vision tasks. However, challenging concepts in queries often confuse retrievers, limiting their ability to align concepts with visual conte

Cited by 0SourcePDFScholar
2026

WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference

ICLR 2026poster

The ever-increasing computational demands of large language models (LLMs) make efficient inference a central challenge. While recent advances leverage specialized architectures or selective activation, they typically require (re)training or architectural modifications, limiting their broad applicabi…

Cited by 0SourcecodeScholar
2025

Layer by Layer: Uncovering Hidden Representations in Language Models

ICML 2025oral

From extracting features to generating text, the outputs of large language models (LLMs) typically rely on their final layers, following the conventional wisdom that earlier layers capture only low-level cues. However, our analysis shows that intermediate layers can encode even richer representation…

Cited by 5SourcePDFScholar
2025

SandboxSocial: A Sandbox for Social Media Using Multimodal AI Agents

IJCAI 2025

The online information ecosystem enables influence campaigns of unprecedented scale and impact. We urgently need empirically grounded approaches to counter the growing threat of malicious campaigns, now amplified by generative AI. But, developing defenses in real-world settings is impractical. Socia

2025

SpecCoT: Accelerating Chain-of-Thought Reasoning through Speculative Exploration

EMNLP 2025

Large Reasoning Models (LRMs) demonstrate strong performance on complex tasks through chain-of-thought (CoT) reasoning. However, they suffer from high inference latency due to lengthy reasoning chains. In this paper, we propose SpecCoT, a collaborative framework that combines large and small models

Cited by 0SourcePDFScholar
2025

VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks

ICLR 2025poster

Videos are often used to learn or extract the necessary information to complete tasks in ways different than what text or static imagery can provide. However, many existing agent benchmarks neglect long-context video understanding, instead focus- ing on text or static image inputs. To bridge this ga…

Cited by 3SourcePDFScholar
2025

WinSpot: GUI Grounding Benchmark with Multimodal Large Language Models

ACL 2025short

Graphical User Interface (GUI) automation relies on accurate GUI grounding. However, obtaining large-scale, high-quality labeled data remains a key challenge, particularly in desktop environments like Windows Operating System (OS). Existing datasets primarily focus on structured web-based elements,…

2025

Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

ICML 2025poster

Large language models (LLMs) show potential as computer agents, enhancing productivity and software accessibility in multi-modal tasks. However, measuring agent performance in sufficiently realistic and complex environments becomes increasingly challenging as: (i) most benchmarks are limited to sp…

2024

Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models

ACL 2024long

Finetuning large language models (LLMs) has been empirically effective on a variety of downstream tasks. Existing approaches to finetuning an LLM either focus on parameter-efficient finetuning, which only updates a small number of trainable parameters, or attempt to reduce the memory footprint durin…

2023

Interpreting Unsupervised Anomaly Detection in Security via Rule Extraction

NeurIPS 2023poster

Many security applications require unsupervised anomaly detection, as malicious data are extremely rare and often only unlabeled normal data are available for training (i.e., zero-positive). However, security operators are concerned about the high stakes of trusting black-box models due to their lac…

2023

Metis: Understanding and Enhancing In-Network Regular Expressions

NeurIPS 2023poster

Regular expressions (REs) offer one-shot solutions for many networking tasks, e.g., network intrusion detection. However, REs purely rely on expert knowledge and cannot utilize labeled data for better accuracy. Today, neural networks (NNs) have shown superior accuracy and flexibility, thanks to thei…

2022

Fldp: Flexible Strategy For Local Differential Privacy

ICASSP 2022accepted

Local differential privacy (LDP), a technique applying unbiased statistical estimations instead of real data, is often adopted in data collection. In particular, this technique is used in frequency oracles (FO) because it can protect each user’s privacy and prevent leakage of sensitive information.…

Cited by 0SourceScholar