← Search

Kaiqi Chen

9 accepted papers

2026

DOUBT: Decoupled Object-level Understanding and Bridging via vMF-based Trustworthiness for Hallucination Detection in MLLMs

ICML 2026oral

Multimodal Large Language Models (MLLMs) frequently produce hallucinations (i.e., assertions that contradict the image or facts), undermining reliability in high-risk applications. Existing detection approaches typically feed images and texts jointly and estimate hallucination scores by measuring th…

Cited by 0SourceScholar
2025

DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference

ICLR 2025spotlight

Large language models (LLMs) are increasingly employed for complex tasks that process multiple generation calls in a tree structure with shared prefixes of tokens, including few-shot prompting, multi-step reasoning, speculative decoding, etc. However, existing inference systems for tree-based applic…

2025

Imitation Learning with Limited Actions via Diffusion Planners and Deep Koopman Controllers

ICRA 2025

Recent advances in diffusion-based robot policies have demonstrated significant potential in imitating multi-modal behaviors. However, these approaches typically require large quantities of demonstration data paired with corresponding robot action labels, creating a substantial data collection burde

Cited by 4SourcecodeScholar
2024

Don't Start From Scratch: Behavioral Refinement via Interpolant-based Policy Diffusion

RSS 2024poster

Imitation learning empowers artificial agents to mimic behavior by learning from demonstrations. Recently, diffusion models, which have the ability to model high-dimensional and multimodal distributions, have shown impressive performance on imitation learning tasks. These models learn to shape a pol…

2023

Latent Emission-Augmented Perspective-Taking (LEAPT) for Human-Robot Interaction

IROS 2023poster

Perspective-taking is the ability to perceive or understand a situation or concept from another individual's point of view, and is crucial in daily human interactions. Enabling robots to perform perspective-taking remains an unsolved problem; existing approaches that use deterministic or handcrafted…

Cited by 0SourceScholar
2022

MIRROR: Differentiable Deep Social Projection for Assistive Human-Robot Communication

RSS 2022poster

Communication is a hallmark of intelligence. In this work, we present MIRROR, an approach to (i) quickly learn human models from human demonstrations, and (ii) use the models for subsequent communication planning in assistive shared-control settings. MIRROR is inspired by social projection theory, w…

2022

Robust and Accurate Multi-Agent SLAM with Efficient Communication for Smart Mobiles

ICRA 2022poster

In a long-term large-scenario application, the multi-agent collaborative SLAM is expected to improve the robustness and efficiency of executing tasks for mobile agents. In this paper, a multi-agent collaborative visual-inertial SLAM system is proposed based on a centralized client-server (CS) archit…

Cited by 8SourceScholar
2021

Collaborative Visual Inertial SLAM for Multiple Smart Phones

ICRA 2021poster

The efficiency and accuracy of mapping are crucial in a large scene and long-term AR applications. Multi-agent cooperative SLAM is the precondition of multi-user AR interaction. The cooperation of multiple smart phones has the potential to improve efficiency and robustness of task completion and can…

Cited by 16SourceScholar
2021

Multi-Modal Mutual Information (MuMMI) Training for Robust Self-Supervised Deep Reinforcement Learning

ICRA 2021poster

This work focuses on learning useful and robust deep world models using multiple, possibly unreliable, sensors. We find that current methods do not sufficiently encourage a shared representation between modalities; this can cause poor performance on downstream tasks and over-reliance on specific sen…

Cited by 26SourceScholar