← Search

Hong-Yuan Mark Liao

8 accepted papers

2025

YOLO-RD: Introducing Relevant and Compact Explicit Knowledge to YOLO by Retriever-Dictionary

ICLR 2025poster

Identifying and localizing objects within images is a fundamental challenge, and numerous efforts have been made to enhance model accuracy by experimenting with diverse architectures and refining training strategies. Nevertheless, a prevalent limitation in existing models is overemphasizing the curr…

2024

YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information

ECCV 2024poster

"Today’s deep learning methods focus on how to design the objective functions to make the prediction as close as possible to the target. Meanwhile, an appropriate neural network architecture has to be designed. Existing methods ignore a fact that when input data undergoes layer-by-layer feature tran…

2023

YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors

CVPR 2023poster

Real-time object detection is one of the most important research topics in computer vision. As new approaches regarding architecture optimization and training optimization are continually being developed, we have found two research topics that have spawned when dealing with these latest state-of-the…

2022

Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation From Monocular Video

CVPR 2022poster

Learning to capture human motion is essential to 3D human pose and shape estimation from monocular video. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the ability to capture non-local context relations of human mot…

Cited by 108PDFcodeScholar
2018

How Sampling Rate Affects Cross-Domain Transfer Learning for Video Description

ICASSP 2018accepted

Translating video to language is very challenging due to diversified video contents originated from multiple activities and complicated integration of spatio-temporal information. There are two urgent issues associated with the video-to-language translation problem. First, how to transfer knowledge…

Cited by 0SourceScholar
2017

Deep-net fusion to classify shots in concert videos

ICASSP 2017accepted

Varying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director to convey the emotion, ideas, and art. To classify such types of shots from images, we present a new framework that facilitates the intriguing task by addressing two key issues. W…

Cited by 0SourceScholar
2016

Precise player segmentation in team sports videos using contrast-aware co-segmentation

ICASSP 2016accepted

Player segmentation in team sports videos is challenging but crucial to video semantic understanding, such as player interaction identification and tactic analysis. We leverage the appearance similarity among players of the same team, and cast this task as a co-segmentation problem. In this way, the…

Cited by 0SourceScholar