← Search

Kaiwen Zhou

37 accepted papers

2026

$PhyWorldBench$: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

ICLR 2026oral

Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. This paper presents $PhyWorldBench$ , a comprehensive benchmark designed to evalua…

Cited by 16SourcecodeScholar
2026

A Unified Perspective on Adversarial Membership Manipulation in Vision Models

CVPR 2026

Membership inference attacks (MIAs) aim to determine whether a specific data point was part of a model's training set, serving as effective tools for evaluating privacy leakage of vision models. However, existing MIAs implicitly assume honest query inputs, and their adversarial robustness remains un

Cited by 0SourcecodeScholar
2026

A²Flow: Automating Agentic Workflow Generation via Self-Adaptive Abstraction Operators

AAAI 2026technical

Large language models (LLMs) have shown strong potential in automating the design of agentic workflows. However, existing methods still rely heavily on manually predefined operators, limiting generalization and scalability. To address this issue, we propose A²Flow, a fully automated framework for ag

Cited by 0SourcePDFScholar
2026

Boosting Cross-problem Generalization in Diffusion-Based Neural Combinatorial Solver via Inference Time Adaptation

AAAI 2026technical

Diffusion-based Neural Combinatorial Optimization (NCO) has demonstrated effectiveness in solving NP-complete (NPC) problems by learning discrete diffusion models for solution generation, eliminating hand-crafted domain knowledge. Despite their success, existing NCO methods face significant challeng

Cited by 0SourcePDFScholar
2026

HATS: Hardness-Aware Trajectory Synthesis for GUI Agents

CVPR 2026

Graphical user interface (GUI) agents powered by large vision-language models (VLMs) have shown remarkable potential in automating digital tasks, highlighting the need for high-quality trajectory data to support effective agent training. Yet existing trajectory synthesis pipelines often yield agents

Cited by 0SourcecodeScholar
2026

HiconAgent: History Context-aware Policy Optimization for GUI Agents

CVPR 2026

Graphical User Interface (GUI) agents require effective utilization of historical context to perform sequential navigation tasks. While incorporating past actions and observations can significantly improve decision-making, naively using full history leads to excessive computational overhead and pote

Cited by 0SourcecodeScholar
2026

Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations

ICLR 2026poster

The promotion of academic papers has become an important means of enhancing research visibility. where the appeal of dissemination largely determines its effectiveness. However, existing automated methods struggle limited storytelling, insufficient aesthetic quality, and constrained self-adjustment,…

Cited by 0SourcecodeScholar
2026

SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation

AAAI 2026technical

Vision-Language-Action (VLA) models have advanced in robotic manipulation, yet practical deployment remains hindered by two key limitations: **1) perceptual redundancy**, where irrelevant visual inputs are processed inefficiently, and **2) superficial instruction-vision alignment**, which hampers se

Cited by 0SourcePDFScholar
2025

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

ICCV 2025poster

The incorporation of high-resolution visual input equips multimodal large language models (MLLMs) with enhanced visual perception capabilities for real-world tasks. However, most existing high-resolution MLLMs rely on a cropping-based approach to process images, which leads to fragmented visual enco…

2025

GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents

NeurIPS 2025poster

Recent Graphical User Interface (GUI) agents replicate the R1-Zero paradigm, coupling online Reinforcement Learning (RL) with explicit chain-of-thought reasoning prior to object grounding and thereby achieving substantial performance gains. In this paper, we first conduct extensive analysis experime…

Cited by 0SourcecodeScholar
2025

GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent

ACL 2025long

GUI automation faces critical challenges in dynamic environments. MLLMs suffer from two key issues: misinterpreting UI components and outdated knowledge. Traditional fine-tuning methods are costly for app-specific knowledge updates. We propose GUI-explorer, a training-free GUI agent that incorporate…

2025

Less is More: Empowering GUI Agent with Context-Aware Simplification

ICCV 2025poster

The research focus of GUI agents is shifting from text-dependent to pure-vision-based approaches, which, though promising, prioritize comprehensive pre-training data collection while neglecting contextual modeling challenges. We probe the characteristics of element and history contextual modeling in…

2025

Multimodal Situational Safety

ICLR 2025poster

Multimodal Large Language Models (MLLMs) are rapidly evolving, demonstrating impressive capabilities as multimodal assistants that interact with both humans and their environments. However, this increased sophistication introduces significant safety concerns. In this paper, we present the first eval…

Cited by 48SourcePDFScholar
2025

SPA-BENCH: A COMPREHENSIVE BENCHMARK FOR SMARTPHONE AGENT EVALUATION

ICLR 2025spotlight

Smartphone agents are increasingly important for helping users control devices efficiently, with (Multimodal) Large Language Model (MLLM)-based approaches emerging as key contenders. Fairly comparing these agents is essential but challenging, requiring a varied task scope, the integration of agents…

2025

SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning

EMNLP 2025

Large Reasoning Models (LRMs) introduce a new generation paradigm of explicitly reasoning before answering, leading to remarkable improvements in complex tasks. However, they pose great safety risks against harmful queries and adversarial attacks. While recent mainstream safety efforts on LRMs, supe

2024

Enhancing Evolving Domain Generalization through Dynamic Latent Representations

AAAI 2024technical

Domain generalization is a critical challenge for machine learning systems. Prior domain generalization methods focus on extracting domain-invariant features across several stationary domains to enable generalization to new domains. However, in non-stationary tasks where new domains evolve in an und…

Cited by 5SourcePDFScholar
2024

Enhancing Neural Subset Selection: Integrating Background Information into Set Representations

ICLR 2024poster

Learning neural subset selection tasks, such as compound selection in AI-aided drug discovery, have become increasingly pivotal across diverse applications. The existing methodologies in the field primarily concentrate on constructing models that capture the relationship between utility function val…

Cited by 1SourcePDFScholar
2024

HORSE: Hierarchical Representation for Large-Scale Neural Subset Selection

NeurIPS 2024poster

Subset selection tasks, such as anomaly detection and compound selection in AI-assisted drug discovery, are crucial for a wide range of applications. Learning subset-valued functions with neural networks has achieved great success by incorporating permutation invariance symmetry into the architectur…

Cited by 0SourcePDFScholar
2024

Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA

ACL 2024long

Multipanel images, commonly seen as web screenshots, posters, etc., pervade our daily lives. These images, characterized by their composition of multiple subfigures in distinct layouts, effectively convey information to people. Toward building advanced multimodal AI applications, such as agents that…

Cited by 20SourcePDFScholar
2024

Navigation as Attackers Wish? Towards Building Robust Embodied Agents under Federated Learning

NAACL 2024long

Federated embodied agent learning protects the data privacy of individual visual environments by keeping data locally at each client (the individual environment) during training. However, since the local data is inaccessible to the server under federated learning, attackers may easily poison the tra…

Cited by 2SourcePDFScholar
2024

RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models

NeurIPS 2024poster

Natural images captured by mobile devices often suffer from multiple types of degradation, such as noise, blur, and low light. Traditional image restoration methods require manual selection of specific tasks, algorithms, and execution sequences, which is time-consuming and may yield suboptimal resul…

Cited by 6SourcePDFScholar
2024

ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models

ACL 2024findings

In our work, we explore the synergistic capabilities of pre-trained vision-and-language models (VLMs) and large language models (LLMs) on visual commonsense reasoning (VCR) problems. We find that VLMs and LLMs-based decision pipelines are good at different kinds of VCR problems. Pre-trained VLMs exh…

Cited by 9SourcePDFScholar
2023

Does Invariant Graph Learning via Environment Augmentation Learn Invariance?

NeurIPS 2023poster

Invariant graph representation learning aims to learn the invariance among data from different environments for out-of-distribution generalization on graphs. As the graph environment partitions are usually expensive to obtain, augmenting the environment information has become the de facto approach.…

Cited by 47SourcePDFScholar
2023

ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation

ICML 2023poster

The ability to accurately locate and navigate to a specific object is a crucial capability for embodied agents that operate in the real world and interact with objects to complete tasks. Such object navigation tasks usually require large-scale training in visual environments with labeled objects, wh…

Cited by 115SourcePDFScholar
2023

Pareto Invariant Risk Minimization: Towards Mitigating the Optimization Dilemma in Out-of-Distribution Generalization

ICLR 2023poster

Recently, there has been a growing surge of interest in enabling machine learning systems to generalize well to Out-of-Distribution (OOD) data. Most efforts are devoted to advancing optimization objectives that regularize models to capture the underlying invariance; however, there often are compromi…

2023

Understanding and Improving Feature Learning for Out-of-Distribution Generalization

NeurIPS 2023poster

A common explanation for the failure of out-of-distribution (OOD) generalization is that the model trained with empirical risk minimization (ERM) learns spurious features instead of invariant features. However, several recent studies challenged this explanation and found that deep networks may have…

Cited by 47SourcePDFScholar
2022

Fast and Reliable Evaluation of Adversarial Robustness with Minimum-Margin Attack

ICML 2022spotlight

The AutoAttack (AA) has been the most reliable method to evaluate adversarial robustness when considerable computational resources are available. However, the high computational cost (e.g., 100 times more than that of the project gradient descent attack) makes AA infeasible for practitioners with li…

2022

On the Finite-Time Complexity and Practical Computation of Approximate Stationarity Concepts of Lipschitz Functions

ICML 2022spotlight

We report a practical finite-time algorithmic scheme to compute approximately stationary points for nonconvex nonsmooth Lipschitz functions. In particular, we are interested in two kinds of approximate stationarity notions for nonconvex nonsmooth problems, i.e., Goldstein approximate stationarity (G…

Cited by 41SourcePDFScholar
2022

Practical Schemes for Finding Near-Stationary Points of Convex Finite-Sums

AISTATS 2022poster

In convex optimization, the problem of finding near-stationary points has not been adequately studied yet, unlike other optimality measures such as the function value. Even in the deterministic case, the optimal method (OGM-G, due to Kim and Fessler (2021)) has just been discovered recently. In this…

Cited by 14SourcePDFScholar
2020

Amortized Nesterov’s Momentum: A Robust Momentum and Its Application to Deep Learning

UAI 2020poster

This work proposes a novel momentum technique, the Amortized Nesterov’s Momentum, for stochastic convex optimization. The proposed method can be regarded as a smooth transition between Nesterov’s method and mirror descent. By tuning only a single parameter, users can trade Nesterov’s acceleration fo…

Cited by 8SourcePDFScholar
2020

Boosting First-Order Methods by Shifting Objective: New Schemes with Faster Worst-Case Rates

NeurIPS 2020poster

We propose a new methodology to design first-order methods for unconstrained strongly convex problems. Specifically, instead of tackling the original objective directly, we construct a shifted objective function that has the same minimizer as the original objective and encodes both the smoothness an…

Cited by 7SourcePDFScholar
2019

Direct Acceleration of SAGA using Sampled Negative Momentum

AISTATS 2019poster

Variance reduction is a simple and effective technique that accelerates convex (or non-convex) stochastic optimization. Among existing variance reduction methods, SVRG and SAGA adopt unbiased gradient estimators and are the most popular variance reduction methods in recent years. Although various ac…

Cited by 62SourcePDFScholar
2018

A Simple Stochastic Variance Reduced Algorithm with Fast Convergence Rates

ICML 2018oral

Recent years have witnessed exciting progress in the study of stochastic variance reduced gradient methods (e.g., SVRG, SAGA), their accelerated variants (e.g, Katyusha) and their extensions in many different settings (e.g., online, sparse, asynchronous, distributed). Among them, accelerated methods…

Cited by 103SourcePDFScholar
2018

Guaranteed Sufficient Decrease for Stochastic Variance Reduced Gradient Optimization

AISTATS 2018poster

In this paper, we propose a novel sufficient decrease technique for stochastic variance reduced gradient descent methods such as SVRG and SAGA. In order to make sufficient decrease for stochastic optimization, we design a new sufficient decrease criterion, which yields sufficient decrease versions o…

Cited by 0SourcePDFScholar