← Search

Yuan-Fang Wang

8 accepted papers

2025

SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models

ICLR 2025poster

Multimodal Large Language Models (MLLMs) are advancing the ability to reason about complex sports scenarios by integrating textual and visual information. To comprehensively evaluate their capabilities, we introduce SPORTU, a benchmark designed to assess MLLMs across multi-level sports reasoning tas…

2024

SportQA: A Benchmark for Sports Understanding in Large Language Models

NAACL 2024long

A deep understanding of sports, a field rich in strategic and dynamic content, is crucial for advancing Natural Language Processing (NLP). This holds particular significance in the context of evaluating and advancing Large Language Models (LLMs), given the existing gap in specialized benchmarks. To…

2019

MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment

CVPR 2019poster

This research strives for natural language moment retrieval in long, untrimmed video streams. The problem is not trivial especially when a video contains multiple moments of interests and the language describes complex temporal dependencies, which often happens in real scenarios. We identify two cr…

Cited by 372PDFScholar
2019

Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation

CVPR 2019oral

Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalizati…

Cited by 649PDFScholar
2019

VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research

ICCV 2019oral

We present a new large-scale multilingual video description dataset, VATEX, which contains over 41,250 videos and 825,000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation pairs. Compared to the widely-used MSR-VTT dataset, \vatex i…

Cited by 669PDFScholar
2018

Video Captioning via Hierarchical Reinforcement Learning

CVPR 2018poster

Video captioning is the task of automatically generating a textual description of the actions in a video. Although previous work (e.g. sequence-to-sequence model) has shown promising results in abstracting a coarse description of a short video, it is still very challenging to caption a video contain…

Cited by 320SourcePDFScholar
2017

Multimodal Transfer: A Hierarchical Deep Convolutional Neural Network for Fast Artistic Style Transfer

CVPR 2017poster

Transferring artistic styles onto everyday photographs has become an extremely popular task in both academia and industry. Recently, offline training has replaced online iterative optimization, enabling nearly real-time stylization. When those stylization networks are applied directly to high-reso…

Cited by 217PDFScholar