← Search

Mingda Jia

5 accepted papers

2026

Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction

AAAI 2026technical

Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competitive performance. However, these pipelines rely on implicit modeling that uses fra

Cited by 0SourcePDFScholar
2025

AccidentX: A Large-Scale Multimodal BEV Dataset for Traffic Accident Analysis and Prevention

IROS 2025

With the rapid development and widespread application of autonomous driving technology, the accurate analysis and prevention of traffic accidents have become critical challenges. However, current traffic accident datasets are often constrained by limited scale and diversity, impeding progress in thi

Cited by 0SourceScholar
2025

ContextHOI: Spatial Context Learning for Human-Object Interaction Detection

AAAI 2025technical

Spatial contexts, such as the backgrounds and surroundings, are considered critical in Human-Object Interaction (HOI) recognition, especially when the instance-centric foreground is blurred or occluded. Recent advancements in HOI detectors are usually built upon detection transformer pipelines. Whil…

Cited by 1SourcePDFScholar
2025

Orchestrating the Symphony of Prompt Distribution Learning for Human-Object Interaction Detection

AAAI 2025technical

Human-object interaction (HOI) detectors with popular query-transformer architecture have achieved promising performance. However, accurately identifying uncommon visual patterns and distinguishing between ambiguous HOIs continue to be difficult for them. We observe that these difficulties may arise…

Cited by 1SourcePDFScholar