← Search

Yali Li

36 accepted papers

2026

HSI-GPT2: A Dual-Granularity Large Motion Reasoning Model with Diffusion Refinement for Human-Scene Interaction

CVPR 2026

Unified interpreting and synthesizing human behaviors within 3D environments is vital for advancing spatial intelligence and humanoid robotics. Despite recent advancements (e.g., HSI-GPT), two fundamental capabilities expected of a unified model--understanding and generation--still lag behind specia

Cited by 0SourceScholar
2026

VES-RFT: Rewarding Visual Evidence Sensitivity to Mitigate Hallucinations in Large Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs) often over-rely on linguistic priors even when images are provided, leading to object hallucinations. We revisit object-wise hallucination from the perspective of how visual evidence shapes the model's uncertainty. For each input, we measure decision uncertainty with an

Cited by 0SourceScholar
2025

Attention Augmented Structure-centric Bias Mitigation with Feature Disentanglement

ICASSP 2025accepted

Image classification models often rely on superficial visual features, such as textures or colors, leading to undesired bias. This can compromise the robustness and reliability of deep models, particularly their performance on out-of-distribution (o.o.d.) datasets. Existing approaches, focusing on d…

Cited by 0SourceScholar
2025

BookBot: A Robotic Manipulation Benchmark for Voice-Driven Book Recognition and Grasping in Cluttered Environments

IROS 2025

Books, as enduring repositories of cultural heritage as well as knowledge, play a fundamental role in human development. Although advances in embodied AI and robotics revolutionize automation in domains, e.g., manufacturing and logistics, robotic book manipulation remains an underexplored frontier.

Cited by 0SourcecodeScholar
2025

Dynamic Object Queries for Transformer-based Incremental Object Detection

ICASSP 2025accepted

Incremental object detection (IOD) aims to sequentially learn new classes, while maintaining the capability to locate and identify old ones. Prior methodologies mainly tackle catastrophic forgetting through knowledge distillation and exemplar replay, ignoring the conflict between limited model capac…

Cited by 0SourceScholar
2025

Find Details in Long Videos: Tower-of-Thoughts and Self-Retrieval Augmented Generation for Video Understanding

ICASSP 2025accepted

The Large Vision-Language Model (LVLM) has achieved impressive performance in the field of visual-language understanding. However, its ability to understand longer videos is still limited due to the length and information diversity of multi-modal videos. Moreover, accurately matching detailed conten…

Cited by 0SourceScholar
2025

HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene Interaction

CVPR 2025highlight

While flourishing developments have been witnessed in text-to-motion generation, synthesizing physically realistic, controllable, language-conditioned Human Scene Interactions (HSI) remains a relatively underexplored landscape. Current HSI methods naively rely on conditional Variational AutoEncoder…

Cited by 0SourcePDFScholar
2024

Exploring Pose-Aware Human-Object Interaction via Hybrid Learning

CVPR 2024poster

Human-Object Interaction (HOI) detection plays a crucial role in visual scene comprehension. In recent advancements two-stage detectors have taken a prominent position. However they are encumbered by two primary challenges. First the misalignment between feature representation and relation reasoning…

Cited by 4SourcePDFScholar
2024

G^3-LQ: Marrying Hyperbolic Alignment with Explicit Semantic-Geometric Modeling for 3D Visual Grounding

CVPR 2024poster

Grounding referred objects in 3D scenes is a burgeoning vision-language task pivotal for propelling Embodied AI as it endeavors to connect the 3D physical world with free-form descriptions. Compared to the 2D counterparts challenges posed by the variability of 3D visual grounding remain relatively u…

Cited by 8SourcePDFScholar
2024

Learning Generalizable Visual Representations via Self-Supervised Information Bottleneck

ICASSP 2024accepted

Numerous approaches have recently emerged in the realm of self-supervised visual representation learning. While these methods have demonstrated empirical success, a theoretical foundation that understands and unifies these diverse techniques remains to be established. In this work, we draw inspirati…

Cited by 0SourceScholar
2023

Detecting Everything in the Open World: Towards Universal Object Detection

CVPR 2023poster

In this paper, we formally address universal object detection, which aims to detect every scene and predict every category. The dependence on human annotations, the limited visual information, and the novel categories in the open world severely restrict the universality of traditional detectors. We…

2023

Identity-Seeking Self-Supervised Representation Learning for Generalizable Person Re-Identification

ICCV 2023oral

This paper aims to learn a domain-generalizable (DG) person re-identification (ReID) representation from large-scale videos without any annotation. Prior DG ReID methods employ limited labeled data for training due to the high cost of annotation, which restricts further advances. To overcome the bar…

Cited by 28PDFcodeScholar
2023

VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor Scenes

IROS 2023poster

Robotic grasping faces new challenges in human-robot-interaction scenarios. We consider the task that the robot grasps a target object designated by human's language directives. The robot not only needs to locate a target based on vision-and-language information, but also needs to predict the reason…

Cited by 23SourcecodeScholar
2022

CRPN: Distinguish Novel Categories Via Class-Relevant Region Proposal Network for Few-Shot Object Detection

ICASSP 2022accepted

Few-shot object detection (FSOD) has attracted more attention in computer vision, where only very few training examples are presented during model learning process. A commonly-overlooked issue in FSOD is that novel classes are usually classified as background clutters in the pre-training process. An…

Cited by 0SourceScholar
2022

GraphCSPN: Geometry-Aware Depth Completion via Dynamic GCNs

ECCV 2022poster

"Image guided depth completion aims to recover per-pixel dense depth maps from sparse depth measurements with the help of aligned color images, which has a wide range of applications from robotics to autonomous driving. However, the 3D nature of sparse-to-dense depth completion has not been fully ex…

2022

Hybrid Physical Metric For 6-DoF Grasp Pose Detection

ICRA 2022poster

6-DoF grasp pose detection of multi-grasp and multi-object is a challenge task in the field of intelligent robot. To imitate human reasoning ability for grasping objects, data driven methods are widely studied. With the introduction of large-scale datasets, we discover that a single physical metric…

Cited by 22SourcecodeScholar
2022

Noisy Boundaries: Lemon or Lemonade for Semi-Supervised Instance Segmentation?

CVPR 2022poster

Current instance segmentation methods rely heavily on pixel-level annotated images. The huge cost to obtain such fully-annotated images restricts the dataset scale and limits the performance. In this paper, we formally address semi-supervised instance segmentation, where unlabeled images are employe…

Cited by 40PDFcodeScholar
2022

OSKDet: Orientation-Sensitive Keypoint Localization for Rotated Object Detection

CVPR 2022poster

Rotated object detection is a challenging issue in computer vision field. Inadequate rotated representation and the confusion of parametric regression have been the bottleneck for high performance rotated detection. In this paper, we propose an orientation-sensitive keypoint based rotated detector O…

Cited by 24PDFScholar
2022

Progressive-Granularity Retrieval Via Hierarchical Feature Alignment for Person Re-Identification

ICASSP 2022accepted

Person re-identification (re-ID) aims to match pedestrian images from non-overlapping cameras. It is a challenging task because of the feature misalignment problem caused by occlusion. In this paper, inspired by the coarse-to-fine nature of human perception, we propose a novel Progressive-Granularit…

Cited by 0SourceScholar
2022

Reliability-Aware Prediction via Uncertainty Learning for Person Image Retrieval

ECCV 2022poster

"Current person image retrieval methods have achieved great improvements in accuracy metrics. However, they rarely describe the reliability of the prediction. In this paper, we propose an Uncertainty-Aware Learning (UAL) method to remedy this issue. UAL aims at providing reliability-aware prediction…

2021

A2-FPN: Attention Aggregation Based Feature Pyramid Network for Instance Segmentation

CVPR 2021poster

Learning pyramidal feature representations is crucial for recognizing object instances at different scales. Feature Pyramid Network (FPN) is the classic architecture to build a feature pyramid with high-level semantics throughout. However, intrinsic defects in feature extraction and fusion inhibit F…

Cited by 122PDFScholar
2021

Data-Uncertainty Guided Multi-Phase Learning for Semi-Supervised Object Detection

CVPR 2021poster

In this paper, we delve into semi-supervised object detection where unlabeled images are leveraged to break through the upper bound of fully-supervised object detection models. Previous semi-supervised methods based on pseudo labels are severely degenerated by noise and prone to overfit to noisy lab…

Cited by 101PDFScholar
2021

Disentangled Representation for Age-Invariant Face Recognition: A Mutual Information Minimization Perspective

ICCV 2021poster

General face recognition has seen remarkable progress in recent years. However, large age gap still remains a big challenge due to significant alterations in facial appearance and bone structure. Disentanglement plays a key role in partitioning face representations into identity-dependent and age-de…

Cited by 35PDFScholar
2021

Partial Off-Policy Learning: Balance Accuracy and Diversity for Human-Oriented Image Captioning

ICCV 2021poster

Human-oriented image captioning with both high diversity and accuracy is a challenging task in vision+language modeling. The reinforcement learning (RL) based frameworks promote the accuracy of image captioning, yet seriously hurt the diversity. In contrast, other methods based on variational auto-e…

Cited by 9PDFScholar
2020

CycAs: Self-supervised Cycle Association for Learning Re-identifiable Descriptions

ECCV 2020poster

This paper proposes a self-supervised learning method for the person re-identification (re-ID) problem, where existing unsupervised methods usually rely on pseudo labels, such as those from video tracklets or clustering. A potential drawback of using pseudo labels is that errors may accumulate and i…

Cited by 116SourcePDFScholar
2019

Perceive Where to Focus: Learning Visibility-Aware Part-Level Features for Partial Person Re-Identification

CVPR 2019poster

This paper considers a realistic problem in person re-identification (re-ID) task, i.e., partial re-ID. Under partial re-ID scenario, the images may contain a partial observation of a pedestrian. If we directly compare a partial pedestrian image with a holistic one, the extreme spatial misalignment…

Cited by 460PDFcodeScholar
2016

Bagging regularized common spatial pattern with hybrid motor imagery and myoelectric signal

ICASSP 2016accepted

Common Spatial Pattern(CSP) is a widely used algorithm in BCI application. However, it is sensitive to noise and artifact. In this paper, we propose a bagging regularized common spatial pattern (Bagging RCSP) approach for BCI with hybrid motor imagery and myoelectric signal. We divide the training s…

Cited by 0SourceScholar
2016

Weakly Supervised Object Localization With Progressive Domain Adaptation

CVPR 2016poster

We address the problem of weakly supervised object localization where only image-level annotations are available for training. Many existing approaches tackle this problem through object proposal mining. However, a substantial amount of noise in object proposals causes ambiguities for learning discr…

Cited by 257PDFScholar