← Search

Xiaoying Xing

7 accepted papers

2026

View Invariant Learning for Vision-Language Navigation in Continuous Environments

RA-L 2026

Vision-Language Navigation in Continuous Environments (VLNCE), where an agent follows instructions and moves freely to reach a destination, is a key research problem in embodied AI. However, most existing approaches are sensitive to viewpoint changes, i.e. variations in camera height and viewing ang

Cited by 2SourcecodeScholar
2025

Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation

CVPR 2025highlight

Text-to-image (T2I) generation has made significant advances in recent years, but challenges still remain in the generation of perceptual artifacts, misalignment with complex prompts, and safety. The prevailing approach to address these issues involves collecting human feedback on generated images,…

Cited by 2SourcePDFScholar
2025

GPVK-VL: Geometry-Preserving Virtual Keyframes for Visual Localization under Large Viewpoint Changes

CVPR 2025poster

Visual localization, the task of determining the position and orientation of a camera, typically involves three core components: offline construction of a keyframe database, efficient online keyframes retrieval, and robust local feature matching. However, significant challenges arise when there are…

Cited by 0SourcePDFScholar
2025

SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing

ICCV 2025poster

Due to the challenges of manually collecting accurate editing data, existing datasets are typically constructed using various automated methods, leading to noisy supervision signals caused by the mismatch between editing instructions and original-edited image pairs. Recent efforts attempt to improve…

2024

Active Open-Vocabulary Recognition: Let Intelligent Moving Mitigate CLIP Limitations

CVPR 2024poster

Active recognition which allows intelligent agents to explore observations for better recognition performance serves as a prerequisite for various embodied AI tasks such as grasping navigation and room arrangements. Given the evolving environment and the multitude of object classes it is impractical…

Cited by 4SourcePDFScholar
2024

Learning to Ask Denotative and Connotative Questions for Knowledge-based VQA

EMNLP 2024finding

Large language models (LLMs) have attracted increasing attention due to its prominent performance on various tasks. Recent works seek to leverage LLMs on knowledge-based visual question answering (VQA) tasks which require common sense knowledge to answer the question about an image, since LLMs have…

Cited by 0SourcePDFScholar