← Search

Kai Zhou

15 accepted papers

2026

Action-and-object Aware Alignment for Partially Relevant Video Retrieval

AAAI 2026technical

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing relevant moments for a given text query. This task is extremely challenging, as untrimmed videos often include numerous actions and objects unrelated to the query. However, existing methods usually struggle with f

Cited by 0SourcePDFScholar
2026

HandX: Scaling Bimanual Motion and Interaction Generation

CVPR 2026

Synthesizing human motion has advanced rapidly, yet realistic hand motion and bimanual interaction remain underexplored. Whole-body models often miss the fine-grained cues that drive dexterous behavior, finger articulation, contact timing, and inter-hand coordination, and existing resources lack hig

Cited by 0SourcecodeScholar
2026

Instance-level Visual Active Tracking with Occlusion-Aware Planning

CVPR 2026

Visual Active Tracking (VAT) aims to control cameras to follow a target in 3D space, which is critical for applications like drone navigation and security surveillance. However, it faces two key bottlenecks in real-world deployment: confusion from visually similar distractors caused by insufficient

Cited by 0SourcecodeScholar
2026

RoSE: A Role Correlation Structure-Enhanced Model for Multi-Event Argument Extraction

AAAI 2026technical

Event co-occurrences have been proven effective for event argument extraction (EAE) in previous studies; however, few have considered intra- and inter-event role correlations. Since role varies among different event types, event structure heterogeneity and overlap pose significant challenges to EAE

Cited by 0SourcePDFScholar
2025

A Token-level Text Image Foundation Model for Document Understanding

ICCV 2025poster

In recent years, general visual foundation models (VFMs) have witnessed increasing adoption, particularly as image encoders for popular multi-modal large language models (MLLMs). However, without semantically fine-grained supervision, these models still encounter fundamental prediction errors in the…

2025

Efficient Dynamic Ensembling for Multiple LLM Experts

IJCAI 2025

LLMs have demonstrated impressive performance across various language tasks. However, the strengths of LLMs can vary due to different architectures, model sizes, areas of training data, etc. Therefore, ensemble reasoning for the strengths of different LLM experts is critical to achieving consistent

2025

GraphProt: Certified Black-Box Shielding Against Backdoored Graph Models

IJCAI 2025

Graph learning models have been empirically proven to be vulnerable to backdoor threats, wherein adversaries submit trigger-embedded inputs to manipulate the model predictions. Current graph backdoor defenses manifest several limitations: 1) dependence on model-related details, 2) necessitation of a

Cited by 0SourcePDFScholar
2025

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples

COLING 2025main

Aligning Large Language Models (LLMs) with human feedback is crucial for their development. Existing preference optimization methods such as DPO and KTO, while improved based on Reinforcement Learning from Human Feedback (RLHF), are inherently derived from PPO, requiring a reference model that adds…

Cited by 0SourcePDFScholar
2025

Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review

ACL 2025finding

The recent emergence of Multi-modal Large Language Models (MLLMs) has introduced a new dimension to the Text-rich Image Understanding (TIU) field, with models demonstrating impressive and inspiring performance. However, their rapid evolution and widespread adoption have made it increasingly challeng…

Cited by 0SourcePDFScholar
2024

Collective Certified Robustness against Graph Injection Attacks

ICML 2024poster

We investigate certified robustness for GNNs under graph injection attacks. Existing research only provides sample-wise certificates by verifying each node independently, leading to very limited certifying performance. In this paper, we present the first collective certificate, which certifies a set…

2024

Uncovering Strong Ties: A Study of Indirect Sybil Attack on Signed Social Network

ICASSP 2024accepted

The Fairness and Goodness Algorithm (FGA) is a widely used trust system in signed directed networks. However, attackers can manipulate trust scores on FGA by launching indirect Sybil attacks and exploiting strong ties. In this work, we propose a novel attack method vicinage-attack that formulates th…

Cited by 0SourceScholar