← Search

Junjie Cao

10 accepted papers

2026

STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification

AAAI 2026technical

Multi-modal object Re-Identification (ReID) aims to exploit complementary information from different modalities to retrieve specific objects. However, existing methods often rely on hard token filtering or simple fusion strategies, which can lead to the loss of discriminative cues and increased back

Cited by 0SourcePDFScholar
2025

DisPose: Disentangling Pose Guidance for Controllable Human Image Animation

ICLR 2025poster

Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleton pose), recent works have attempted to introduce additional dense conditions (e.g., depth map) to ensure motion alignme…

2025

PIAD: Pose and Illumination agnostic Anomaly Detection

CVPR 2025poster

We introduce the Pose and Illumination agnostic Anomaly Detection (PIAD) problem, a generalization of pose-agnostic anomaly detection (PAD). Being illumination agnostic is critical, as it relaxes the assumption that training data for an object has to be acquired in the same light configuration of th…

2024

A Distributed Pipeline for Collaborative Pursuit in the Target Guarding Problem

RA-L 2024

The target guarding problem (TGP) is a classical combat game where pursuers aim to capture evaders to protect a territory from intrusion. This paper proposes a distributed pipeline for multi-pursuer multi-evader TGP with the capability to accommodate varying numbers of evaders and criteria for succe

Cited by 8SourceScholar
2024

Can Large Language Models Grasp Legal Theories? Enhance Legal Reasoning with Insights from Multi-Agent Collaboration

EMNLP 2024finding

Large Language Models (LLMs) could struggle to fully understand legal theories and perform complex legal reasoning tasks. In this study, we introduce a challenging task (confusing charge prediction) to better evaluate LLMs’ understanding of legal theories and reasoning capabilities. We also propose…

2024

Hierarchical Search-Based Cooperative Motion Planning

IROS 2024poster

Cooperative path planning, a crucial aspect of multi-agent systems research, serves a variety of sectors, including military, agriculture, and industry. Many existing algorithms, however, come with certain limitations, such as simplified kinematic models and inadequate support for multiple group sce…

Cited by 0SourcecodeScholar
2023

Large Scale Pursuit-Evasion Under Collision Avoidance Using Deep Reinforcement Learning

IROS 2023poster

This paper examines a pursuit-evasion game (PEG) involving multiple pursuers and evaders. The decentralized pursuers aim to collaborate to capture the faster evaders while avoiding collisions. The policies of all agents are learning-based and are subjected to kinematic constraints that are specific…

Cited by 5SourceScholar
2023

Shunted Collision Avoidance for Multi-UAV Motion Planning with Posture Constraints

ICRA 2023poster

This paper investigates the problem of fixed-wing unmanned aerial vehicles (UAV s) motion planning with posture constraints and the problem of the more general symmetrical situations where UAVs have more than one optimal solution. In this paper, the posture constraints are formulated in the 3D Dubin…

Cited by 4SourcecodeScholar
2021

Entity Relation Extraction as Dependency Parsing in Visually Rich Documents

EMNLP 2021main

Previous works on key information extraction from visually rich documents (VRDs) mainly focus on labeling the text within each bounding box (i.e.,semantic entity), while the relations in-between are largely unexplored. In this paper, we adapt the popular dependency parsing model, the biaffine parser…

Cited by 34SourcePDFScholar
2021

Moving Forward in Formation: A Decentralized Hierarchical Learning Approach to Multi-Agent Moving Together

IROS 2021poster

Multi-agent path finding in formation has many potential real-world applications like mobile warehouse robotics. However, previous multi-agent path finding (MAPF) methods hardly take formation into consideration. Further-more, they are usually centralized planners and require the whole state of the…

Cited by 7SourceScholar