← Search

Yutian Lin

11 accepted papers

2026

RiskProp: Collision-Anchored Self-Supervised Risk Propagation For Early Accident Anticipation

CVPR 2026

Accident anticipation aims to predict impending collisions from dashcam videos and trigger early alerts. Existing methods rely on binary supervision with manually annotated "anomaly onset" frames, which are subjective and inconsistent, leading to inaccurate risk estimation. In contrast, we propose R

Cited by 0SourcecodeScholar
2025

Accident Anticipation via Temporal Occurrence Prediction

NeurIPS 2025poster

Accident anticipation aims to predict potential collisions in an online manner, enabling timely alerts to enhance road safety. Existing methods typically predict frame-level risk scores as indicators of hazard. However, these approaches rely on ambiguous binary supervision—labeling all frames in acc…

Cited by 0SourcecodeScholar
2025

Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation

IJCAI 2025

Video-to-Audio (V2A) Generation achieves significant progress and plays a crucial role in film and video post-production. However, current methods overlook the cinematic language, a critical component of artistic expression in filmmaking. As a result, their performance deteriorates in scenarios wher

Cited by 0SourcePDFScholar
2024

DifTraj: Diffusion Inspired by Intrinsic Intention and Extrinsic Interaction for Multi-Modal Trajectory Prediction

IJCAI 2024poster

Recent years have witnessed the success of generative adversarial networks and diffusion models in multi-model trajectory prediction. However, prevailing algorithms only explicitly consider human interaction, but ignore the modeling of human intention, yielding that the generated results deviate lar…

Cited by 1SourcePDFScholar
2024

Improving Bird's Eye View Semantic Segmentation by Task Decomposition

CVPR 2024poster

Semantic segmentation in bird's eye view (BEV) plays a crucial role in autonomous driving. Previous methods usually follow an end-to-end pipeline directly predicting the BEV segmentation map from monocular RGB inputs. However the challenge arises when the RGB inputs and BEV targets from distinct per…

2024

Toward Real Ultra Image Segmentation: Leveraging Surrounding Context to Cultivate General Segmentation Model

NeurIPS 2024poster

Existing ultra image segmentation methods suffer from two major challenges, namely the scalability issue (i.e. they lack the stability and generality of standard segmentation models, as they are tailored to specific datasets), and the architectural issue (i.e. they are incompatible with real-world u…

Cited by 1SourcePDFScholar
2023

Revisit Weakly-Supervised Audio-Visual Video Parsing from the Language Perspective

NeurIPS 2023poster

We focus on the weakly-supervised audio-visual video parsing task (AVVP), which aims to identify and locate all the events in audio/visual modalities. Previous works only concentrate on video-level overall label denoising across modalities, but overlook the segment-level label noise, where adjacent…

2020

Unsupervised Person Re-Identification via Softened Similarity Learning

CVPR 2020poster

Person re-identification (re-ID) is an important topic in computer vision. This paper studies the unsupervised setting of re-ID, which does not require any labeled information and thus is freely deployed to new scenarios. There are very few studies under this setting, and one of the best approach ti…

Cited by 346PDFScholar
2018

Exploit the Unknown Gradually: One-Shot Video-Based Person Re-Identification by Stepwise Learning

CVPR 2018poster

We focus on the one-shot learning for video-based person re-Identification (re-ID). Unlabeled tracklets for the person re-ID tasks can be easily obtained by pre-processing, such as pedestrian detection and tracking. In this paper, we propose an approach to exploiting unlabeled tracklets by gradually…

Cited by 456SourcePDFScholar