← Search

Zhiping Cai

19 accepted papers

2026

Propagating Unsafe Actions in LLM Controlled Multi-Robot Collaboration via Single Robot Compromise

IJCAI 2026

Large language models (LLMs) are increasingly used as general planners in embodied intelligence, enabling high level coordination and low level task planning for both single robot and multi-robot collaboration. This increasing reliance on embodied LLM planners also raises critical security concerns,

Cited by 0Scholar
2026

StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image References

CVPR 2026

Despite the advancements in diffusion-based image style transfer, existing methods are commonly limited by 1) semantic gap: the style reference could miss proper content semantics, causing uncontrollable stylization; 2) reliance on extra constraints (e.g., semantic masks) restricting applicability;

Cited by 0SourcecodeScholar
2026

TOSC: Task-Oriented Shape Completion for Open-World Dexterous Grasp Generation from Partial Point Clouds

AAAI 2026technical

Task-oriented dexterous grasping remains challenging in robotic manipulations of open-world objects under severe partial observation, where significant missing data invalidates generic shape completion. In this paper, to overcome this limitation, we study \emph{Task-Oriented Shape Completion}, a new

Cited by 0SourcePDFScholar
2025

ALLVB: All-in-One Long Video Understanding Benchmark

AAAI 2025technical

From image to video understanding, the capabilities of Multi-modal LLMs (MLLMs) are increasingly powerful. However, most existing video understanding benchmarks are relatively short, which makes them inadequate for effectively evaluating the long-sequence modeling capabilities of MLLMs. This highlig…

Cited by 0SourcePDFScholar
2025

Does One-shot Give the Best Shot? Mitigating Model Inconsistency in One-shot Federated Learning

ICML 2025poster

Turning the multi-round vanilla Federated Learning into one-shot FL (OFL) significantly reduces the communication burden and makes a big leap toward practical deployment. However, this work empirically and theoretically unravels that existing OFL falls into a garbage (inconsistent one-shot local mod…

2025

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly

ICCV 2025poster

Numerous synthesized videos from generative models, especially human-centric ones that simulate realistic human actions, pose significant threats to human information security and authenticity. While progress has been made in binary forgery video detection, the lack of fine-grained understanding of…

Cited by 0SourcePDFScholar
2024

Attribute-Missing Graph Clustering Network

AAAI 2024technical

Deep clustering with attribute-missing graphs, where only a subset of nodes possesses complete attributes while those of others are missing, is an important yet challenging topic in various practical applications. It has become a prevalent learning paradigm in existing studies to perform data imputa…

2024

DiffusionEdge: Diffusion Probabilistic Model for Crisp Edge Detection

AAAI 2024technical

Limited by the encoder-decoder architecture, learning-based edge detectors usually have difficulty predicting edge maps that satisfy both correctness and crispness. With the recent success of the diffusion probabilistic model (DPM), we found it is especially suitable for accurate and crisp edge dete…

2024

Text-Based Occluded Person Re-identification via Multi-Granularity Contrastive Consistency Learning

AAAI 2024technical

Text-based Person Re-identification (T-ReID), which aims at retrieving a specific pedestrian image from a collection of images via text-based information, has received significant attention. However, previous research has overlooked a challenging yet practical form of T-ReID: dealing with image gall…

2023

Efficient Personalized Federated Learning on Selective Model Training

ICASSP 2023accepted

Personalized Federated Learning (FL) handles the data heterogeneous problem by tailoring local models for each distributed data owner. Previous studies first train a highly-adaptable global model and then transfer it for personalization. However, the additional training aggravates burden of resource…

Cited by 0SourceScholar
2023

NEF: Neural Edge Fields for 3D Parametric Curve Reconstruction From Multi-View Images

CVPR 2023poster

We study the problem of reconstructing 3D feature curves of an object from a set of calibrated multi-view images. To do so, we learn a neural implicit field representing the density distribution of 3D edges which we refer to as Neural Edge Field (NEF). Inspired by NeRF, NEF is optimized with a view-…

2023

Tracking Targets in Hyper-Scale Cameras Using Movement Predication

ICASSP 2023accepted

Hyper-scale surveillance cameras have enabled seamless target (e.g., vehicles, individuals) tracking in urban scenarios, significantly improving everyday safety and emergency response capacity. However, practically tracking multiple targets among cameras would incur prohibitive huge computation cost…

Cited by 0SourceScholar
2022

Deep Anomaly Discovery From Unlabeled Videos via Normality Advantage and Self-Paced Refinement

CVPR 2022poster

While classic video anomaly detection (VAD) requires labeled normal videos for training, emerging unsupervised VAD (UVAD) aims to discover anomalies directly from fully unlabeled videos. However, existing UVAD methods still rely on shallow models to perform detection or initialization, and they are…

Cited by 51PDFcodeScholar
2022

Initializing Then Refining: A Simple Graph Attribute Imputation Network

IJCAI 2022poster

Representation learning on the attribute-missing graphs, whose connection information is complete while the attribute information of some nodes is missing, is an important yet challenging task. To impute the missing attributes, existing methods isolate the learning processes of attribute and structu…

Cited by 32SourcePDFScholar
2021

Deep Fusion Clustering Network

AAAI 2021technical

Deep clustering is a fundamental yet challenging task for data analysis. Recently we witness a strong tendency of combining autoencoder and graph neural networks to exploit structure information for clustering performance enhancement. However, we observe that existing literature 1) lacks a dynamic f…

2021

Dynamic Modeling Cross- and Self-Lattice Attention Network for Chinese NER

AAAI 2021technical

Word-character lattice models have been proved to be effective for Chinese named entity recognition (NER), in which word boundary information is fused into character sequences for enhancing character representations. However, prior approaches have only used simple methods such as feature concatenati…

2020

Modeling Dense Cross-Modal Interactions for Joint Entity-Relation Extraction

IJCAI 2020poster

Joint extraction of entities and their relations benefits from the close interaction between named entities and their relation information. Therefore, how to effectively model such cross-modal interactions is critical for the final performance. Previous works have used simple methods such as label-…

Cited by 0SourcePDFScholar