SC-VLMaps: Depth-Free Visual–Language Mapping Via Scene Coordinate Regression
The ability to connect visual observations with human language is increasingly valuable for embodied agents in tasks such as navigation and semantic mapping. Existing visual–language map (VLMaps) approach enables this connection but typically depends on depth images to project semantic features into…