← Search

Hongkai Chen

9 accepted papers

2026

MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models

ICLR 2026poster

Unified Multimodal Large Language Models (U-MLLMs) have garnered considerable interest for their ability to seamlessly integrate generation and comprehension tasks. However, existing research lacks a unified evaluation standard, often relying on isolated benchmarks to assess these capabilities. More…

Cited by 0SourceScholar
2025

ARS-SLAM: Accurate Robust Spinning LiDAR SLAM for a Quadruped Robot in Large-Scale Scenario

ICRA 2025

It is challenging to employ a quadruped robot for real-time mapping and positioning in a large range of scenes. The significant vibration and instability of the quadruped robot during mobility, as well as the quantity of computation required to convey a wide variety of complex landscapes, result in

Cited by 0SourceScholar
2025

ContextAgent: Context-Aware Proactive LLM Agents with Open-world Sensory Perceptions

NeurIPS 2025poster

Recent advances in Large Language Models (LLMs) have propelled intelligent agents from reactive responses to proactive support. While promising, existing proactive agents either rely exclusively on observations from enclosed environments (e.g., desktop UIs) with direct LLM inference or employ rule-…

Cited by 0SourcecodeScholar
2025

DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning

EMNLP 2025

Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language. In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies

Cited by 0SourcePDFScholar
2025

Tri-Ergon: Fine-Grained Video-to-Audio Generation with Multi-Modal Conditions and LUFS Control

AAAI 2025technical

Video-to-audio (V2A) generation utilizes visual-only video features to produce realistic sounds that correspond to the scene. However, current V2A models often lack fine-grained control over the generated audio, especially in terms of loudness variation and the incorporation of multi-modal condition…

Cited by 2SourcePDFScholar
2022

ASpanFormer: Detector-Free Image Matching with Adaptive Span Transformer

ECCV 2022poster

"Generating robust and reliable correspondences across images is a fundamental task for a diversity of applications. To capture context at both global and local granularity, we propose ASpanFormer, a Transformer-based detector-free matcher that is built on hierarchical attention structure, adopting…

2021

Learning To Match Features With Seeded Graph Matching Network

ICCV 2021poster

Matching local features across images is a fundamental problem in computer vision. Targeting towards high accuracy and efficiency, we propose Seeded Graph Matching Network, a graph neural network with sparse structure to reduce redundant connectivity and learn compact representation. The network con…

Cited by 142PDFcodeScholar
2021

PointDSC: Robust Point Cloud Registration Using Deep Spatial Consistency

CVPR 2021poster

Removing outlier correspondences is one of the critical steps for successful feature-based point cloud registration. Despite the increasing popularity of introducing deep learning methods in this field, spatial consistency, which is essentially established by a Euclidean transformation between point…

Cited by 355PDFcodeScholar
2020

ASLFeat: Learning Local Features of Accurate Shape and Localization

CVPR 2020poster

This work focuses on mitigating two limitations in the joint learning of local feature detectors and descriptors. First, the ability to estimate the local shape (scale, orientation, etc.) of feature points is often neglected during dense feature extraction, while the shape-awareness is crucial to ac…

Cited by 379PDFcodeScholar