← Search

Duc Minh Vo

7 accepted papers

2024

A Compact Dynamic 3D Gaussian Representation for Real-Time Dynamic View Synthesis

ECCV 2024poster

"3D Gaussian Splatting (3DGS) has shown remarkable success in synthesizing novel views given multiple views of a static scene. Yet, 3DGS faces challenges when applied to dynamic scenes because 3D Gaussian parameters need to be updated per timestep, requiring a large amount of memory and at least a d…

Cited by 13SourcePDFScholar
2024

EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension

CVPR 2024poster

Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently necessitating the requirement of sustaining up-to-date object knowledge for open-world comprehension. Instead of relying on large…

Cited by 25SourcePDFScholar
2023

A-Cap: Anticipation Captioning With Commonsense Knowledge

CVPR 2023poster

Humans possess the capacity to reason about the future based on a sparse collection of visual cues acquired over time. In order to emulate this ability, we introduce a novel task called Anticipation Captioning, which generates a caption for an unseen oracle image using a sparsely temporally-ordered…

Cited by 5SourcePDFScholar
2023

Partition-And-Debias: Agnostic Biases Mitigation via a Mixture of Biases-Specific Experts

ICCV 2023poster

Bias mitigation in image classification has been widely researched, and existing methods have yielded notable results. However, most of these methods implicitly assume that a given image contains only one type of known or unknown bias, failing to consider the complexities of real-world biases. We in…

Cited by 3PDFcodeScholar
2022

NOC-REK: Novel Object Captioning With Retrieved Vocabulary From External Knowledge

CVPR 2022poster

Novel object captioning aims at describing objects absent from training data, with the key ingredient being the provision of object vocabulary to the model. Although existing methods heavily rely on an object detection model, we view the detection step as vocabulary retrieval from an external knowle…

Cited by 21PDFScholar