← Search

Danyang Tu

6 accepted papers

2026

Rounded or Streamlined Head? Bridging Concept Bottleneck Models and Attribute-Described Object Parts

CVPR 2026

A faithful decision-making process requires models to ground human-understandable concepts both spatially (where they appear in the image) and causally (how they influence the prediction). Recent advances in Vision-Language Models (VLMs) enable concept-level alignment and have inspired Concept Bottl

Cited by 0SourceScholar
2023

MD-VQA: Multi-Dimensional Quality Assessment for UGC Live Videos

CVPR 2023poster

User-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers before being distributed to end-users. Because of the flourishing…

2022

End-to-End Human-Gaze-Target Detection With Transformers

CVPR 2022poster

In this paper, we propose an effective and efficient method for Human-Gaze-Target (HGT) detection, i.e., gaze following. Current approaches decouple the HGT detection task into separate branches of salient object detection and human gaze prediction, employing a two-stage framework where human head l…

Cited by 65PDFScholar
2022

Iwin: Human-Object Interaction Detection via Transformer with Irregular Windows

ECCV 2022poster

"This paper presents a new vision Transformer, named Iwin Transformer, which is specifically designed for human-object interaction (HOI) detection, a detailed scene understanding task involving a sequential process of human/object detection and interaction recognition. Iwin Transformer is a hierarch…

Cited by 28SourcePDFScholar
2022

Video-based Human-Object Interaction Detection from Tubelet Tokens

NeurIPS 2022accept

We present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatial-temporal representations, for video-based human-object interaction (V-HOI) detection. The tubelet tokens structurize videos by agglomerating and linking semantically-related…

Cited by 18SourcePDFScholar