← Search

Di Yang

12 accepted papers

2026

A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps

CVPR 2026

Few-shot object detection (FSOD) is challenging due to unstable optimization and limited generalization arising from the scarcity of training samples. To address these issues, we propose a hybrid ensemble decoder that enhances generalization during fine-tuning. Inspired by ensemble learning, the dec

Cited by 1SourcecodeScholar
2026

MoVie: Broaden Your Views with Human Motion for Action Detection

CVPR 2026

Human action detection in videos requires both semantic recognition and accurate modeling of motion. While recent video foundation models have advanced visual semantics, they still struggle to capture complex and compositional actions due to the limited representation ability of motion. Human skelet

Cited by 0SourceScholar
2026

PRISM: Learning a Shared Primitive Space for Transferable Skeleton Action Representation

CVPR 2026

Real-world human action understanding remains challenging due to long-tailed label distributions, compositional motion patterns, and viewpoint variations. Existing skeleton-based methods often lack a structured and transferable representation of motion, and task-specific models for generation, class

Cited by 0SourceScholar
2025

An Effective Levelling Paradigm for Unlabeled Scenarios

NeurIPS 2025poster

Advancements in direct-integration fine-tuning frameworks have underscored their potential to enhance the performance of labeled scenarios and tasks. To enhance the generalization of different categories in the same dataset, some methods have added visual loss to these frameworks for unlabeled scena…

Cited by 0SourceScholar
2025

Decoder-Only LLMs can be Masked Auto-Encoders

ACL 2025short

Modern NLP workflows (e.g., RAG systems) require different models for generation and embedding tasks, where bidirectional pre-trained encoders and decoder-only Large Language Models (LLMs) dominate respective tasks. Structural differences between models result in extra development costs and limit kn…

2024

Architecture-Agnostic Iterative Black-Box Certified Defense Against Adversarial Patches

ICASSP 2024accepted

The adversarial patch attack aims to fool image classifiers within a bounded, contiguous region of arbitrary changes. To address this problem in a trustworthy way, the certified patch defense methods are proposed. However, the state-of-the-art certified defenses inevitably needed to access the size…

Cited by 0SourceScholar
2024

CPsyCoun: A Report-based Multi-turn Dialogue Reconstruction and Evaluation Framework for Chinese Psychological Counseling

ACL 2024findings

Using large language models (LLMs) to assist psychological counseling is a significant but challenging task at present. Attempts have been made on improving empathetic conversations or acting as effective assistants in the treatment with LLMs. However, the existing datasets lack consulting knowledge…

2023

LAC - Latent Action Composition for Skeleton-based Action Segmentation

ICCV 2023poster

Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processing them by a temporal model to classify frame-wise actions. However, their perfo…

Cited by 14PDFScholar
2023

Prompting Neural Machine Translation with Translation Memories

AAAI 2023technical

Improving machine translation (MT) systems with translation memories (TMs) is of great interest to practitioners in the MT community. However, previous approaches require either a significant update of the model architecture and/or additional training efforts to make the models well-behaved when TMs…

Cited by 14SourcePDFScholar
2023

Self-Supervised Video Representation Learning via Latent Time Navigation

AAAI 2023technical

Self-supervised video representation learning aimed at maximizing similarity between different temporal segments of one video, in order to enforce feature persistence over time. This leads to loss of pertinent information related to temporal relationships, rendering actions such as `enter' and `leav…

Cited by 11SourcePDFScholar
2022

Latent Image Animator: Learning to Animate Images via Latent Space Navigation

ICLR 2022poster

Due to the remarkable progress of deep generative models, animating images has become increasingly efficient, whereas associated results have become increasingly realistic. Current animation-approaches commonly exploit structure representation extracted from driving videos. Such structure representa…

Cited by 178SourcePDFScholar
2019

On-line 3D active pose-graph SLAM based on key poses using graph topology and sub-maps

ICRA 2019poster

In this paper, we present an on-line active pose-graph simultaneous localization and mapping (SLAM) frame-work for robots in three-dimensional (3D) environments using graph topology and sub-maps. This framework aims to find the best trajectory for loop-closure by re-visiting old poses based on the T…

Cited by 18SourceScholar