← Search

Zaijing Li

6 accepted papers

2026

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

ICML 2026oral

Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varying environments. While existing approaches attempt to construct behavior representations through action-centric latent v…

Cited by 0SourceScholar
2026

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

CVPR 2026

Hierarchical Vision-Language-Action (VLA) models have rapidly become a dominant paradigm for robotic manipulation. It typically comprising a Vision-Language backbone for perception and understanding, together with a generative policy for action generation. However, its performance is increasingly bo

Cited by 0SourcecodeScholar
2026

HiconAgent: History Context-aware Policy Optimization for GUI Agents

CVPR 2026

Graphical User Interface (GUI) agents require effective utilization of historical context to perform sequential navigation tasks. While incorporating past actions and observations can significantly improve decision-making, naively using full history leads to excessive computational overhead and pote

Cited by 0SourcecodeScholar
2025

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data

NeurIPS 2025spotlight

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial uncertainty and data scarcity, limiting the 3D spatial reasoning capab…

Cited by 0SourceScholar
2024

Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks

NeurIPS 2024poster

Building a general-purpose agent is a long-standing vision in the field of artificial intelligence. Existing agents have made remarkable progress in many domains, yet they still struggle to complete long-horizon tasks in an open world. We attribute this to the lack of necessary world knowledge and m…

2022

EmoCaps: Emotion Capsule based Model for Conversational Emotion Recognition

ACL 2022findings

Emotion recognition in conversation (ERC) aims to analyze the speaker’s state and identify their emotion in the conversation. Recent works in ERC focus on context modeling but ignore the representation of contextual emotional tendency. In order to extract multi-modal information and the emotional te…

Cited by 108SourcePDFScholar