← Search

Guolong Wang

5 accepted papers

2026

EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision

AAAI 2026technical

Egocentric visual query localization is vital for embodied AI and VR/AR, yet remains challenging due to camera motion, viewpoint changes, and appearance variations. We present EAGLE, a novel framework that leverages episodic appearance- and geometry-aware memory to achieve unified 2D-3D visual query

Cited by 0SourcePDFScholar
2024

Boosting Text-to-Video Generative Model with MLLMs Feedback

NeurIPS 2024poster

Recent advancements in text-to-video generative models, such as Sora, have showcased impressive capabilities. These models have attracted significant interest for their potential applications. However, they often rely on extensive datasets of variable quality, which can result in generated videos th…

Cited by 5SourcePDFScholar
2024

Keep Knowledge in Perception: Zero-Shot Image Aesthetic Assessment

ICASSP 2024accepted

Image aesthetic assessment is an important issue in multimedia, but most existing studies employ supervised learning methods that rely on large-scale annotated data. However, aesthetic scoring annotations are difficult to obtain in large quantities. Therefore, this paper explores zero-shot image aes…

Cited by 0SourceScholar
2024

Multimodal Large Language Models Make Text-to-Image Generative Models Align Better

NeurIPS 2024poster

Recent studies have demonstrated the exceptional potentials of leveraging human preference datasets to refine text-to-image generative models, enhancing the alignment between generated images and textual prompts. Despite these advances, current human preference datasets are either prohibitively expe…

Cited by 3SourcePDFScholar
2023

Instance-Aware Hierarchical Structured Policy for Prompt Learning in Vision-Language Models

ICASSP 2023accepted

In recent years, learnable prompts have emerged as a major prompt learning paradigm, enhancing the performance of large-scale vision-language pre-trained models in few-shot image classification. However, enhancing methods are often time-consuming and inflexible because 1) class-specific prompts are…

Cited by 0SourceScholar