← Search

Ya-Li Li

14 accepted papers

2026

LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction Detection

ICLR 2026poster

Human-Object Interaction (HOI) detection with vision-language models (VLMs) has progressed rapidly, yet a trade-off persists between specialization and generalization. Two major challenges remain: (1) the sparsity of supervision, which hampers effective transfer of foundation models to HOI tasks, a…

Cited by 0SourceScholar
2026

Preserving Topological and Geometric Embeddings for Point Cloud Recovery

AAAI 2026technical

Recovering point clouds involves the sequential process of sampling and restoration, yet existing methods struggle to effectively leverage both topological and geometric attributes. To address this, we propose an end-to-end architecture named TopGeoFormer, which maintains these critical properties t

Cited by 0SourcePDFScholar
2026

TCoT: Trajectory Chain-of-Thoughts for Robotic Manipulation with Failure Recovery in Vision-Language-Action Model

AAAI 2026technical

Recent advances in vision-language-action (VLA) models have demonstrated impressive generalization for robotic manipulation. However, these models often operate by directly mapping visual and linguistic inputs to subsequent actions, lacking intermediate task planning, along with failure detection an

Cited by 0SourcePDFScholar
2026

TRM-VLA: Temporal-Aware Chain-of-Thought Reasoning and Memorization for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for general robotic manipulation. However, existing approaches typically omit intermediate reasoning steps and directly regress actions, limiting reasoning interpretability and performance in long-horizon or compositional tasks.

Cited by 0SourceScholar
2025

LIBA: Language Instructed Multi-granularity Bridge Assistant for 3D Visual Grounding

AAAI 2025technical

3D Vision Grounding (3D-VG) seeks to unravel referential language and identify targets in 3D physical world. Prevailing methods align with the 2D-VG's pipeline to pinpoint the referred object in a categorical multi-modal reasoning manner. However, the geometric complexities of 3D scenes and the nuan…

Cited by 0SourcePDFScholar
2024

OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation

ECCV 2024poster

"In the current state of 3D object detection research, the severe scarcity of annotated 3D data, substantial disparities across different data modalities, and the absence of a unified architecture, have impeded the progress towards the goal of universality. In this paper, we propose OV-Uni3DETR, a u…

2024

One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection

NeurIPS 2024poster

The current trend in computer vision is to utilize one universal model to address all various tasks. Achieving such a universal model inevitably requires incorporating multi-domain data for joint training to learn across multiple problem scenarios. In point cloud based 3D object detection, however,…

Cited by 2SourcePDFScholar
2024

Risk-Aware Self-Consistent Imitation Learning for Trajectory Planning in Autonomous Driving

ECCV 2024poster

"Planning for the ego vehicle is the ultimate goal of autono-mous driving. Although deep learning-based methods have been widely applied to predict future trajectories of other agents in traffic scenes, directly using them to plan for the ego vehicle is often unsatisfactory. This is due to misaligne…

Cited by 1SourcePDFScholar
2023

Uni3DETR: Unified 3D Detection Transformer

NeurIPS 2023poster

Existing point cloud based 3D detectors are designed for the particular scene, either indoor or outdoor ones. Because of the substantial differences in object distribution and point density within point clouds collected from various environments, coupled with the intricate nature of 3D metrics, ther…

2022

Delving into Probabilistic Uncertainty for Unsupervised Domain Adaptive Person Re-identification

AAAI 2022technical

Clustering-based unsupervised domain adaptive (UDA) person re-identification (ReID) reduces exhaustive annotations. However, owing to unsatisfactory feature embedding and imperfect clustering, pseudo labels for target domain data inherently contain an unknown proportion of wrong ones, which would mi…

2021

Combating Noise: Semi-supervised Learning by Region Uncertainty Quantification

NeurIPS 2021poster

Semi-supervised learning aims to leverage a large amount of unlabeled data for performance boosting. Existing works primarily focus on image classification. In this paper, we delve into semi-supervised learning for object detection, where labeled data are more labor-intensive to collect. Current met…

Cited by 33SourcePDFScholar
2021

Do Different Tracking Tasks Require Different Appearance Models?

NeurIPS 2021poster

Tracking objects of interest in a video is one of the most popular and widely applicable problems in computer vision. However, with the years, a Cambrian explosion of use cases and benchmarks has fragmented the problem in a multitude of different experimental setups. As a consequence, the literature…

2020

Video Super-Resolution With Temporal Group Attention

CVPR 2020poster

Video super-resolution, which aims at producing a high-resolution video from its corresponding low-resolution version, has recently drawn increasing attention. In this work, we propose a novel method that can effectively incorporate temporal information in a hierarchical way. The input sequence is d…

Cited by 220PDFcodeScholar