← Search

Tao You

2 accepted papers

2024

DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction

CVPR 2024poster

Audio-visual saliency prediction can draw support from diverse modality complements but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies denoising diffusion models have shown more promising in unifying task fra…

Cited by 6SourcePDFScholar
2024

LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic Surgery

ICRA 2024poster

Visual question answering (VQA) can be fundamentally crucial for promoting robotic-assisted surgical education. In practice, the needs of trainees are constantly evolving, such as learning more surgical types and adapting to new surgical instruments/techniques. Therefore, continually updating the VQ…

Cited by 17SourcecodeScholar