← Search

Zipeng Li

3 accepted papers

2024

TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding

CVPR 2024poster

Humans commonly work with multiple objects in daily life and can intuitively transfer manipulation skills to novel objects by understanding object functional regularities. However existing technical approaches for analyzing and synthesizing hand-object manipulation are mostly limited to handling a s…

2023

Background Disturbance Mitigation for Video Captioning Via Entity-Action Relocation

ICASSP 2023accepted

Video captioning aims to generate sentences to accurately describe the video content, in which video background plays the role of prompts. State-of-the-art methods tend to explore richer video representations adequately, fusing with language to improve caption quality, which has shown great success.…

Cited by 0SourceScholar
2023

Refined Semantic Enhancement towards Frequency Diffusion for Video Captioning

AAAI 2023technical

Video captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual representations in encode phase or improving the decoding ability. However, the long-tailed problem hinders these attempts at…