← Search

Lv Tang

10 accepted papers

2026

UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model

AAAI 2026technical

Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions—remains highly challenging. Recent research on enhancing language-guided navigation reasoning using pre-trained large language models (LLMs) has show

Cited by 0SourcePDFScholar
2025

Boosting Vision State Space Model with Fractal Scanning

AAAI 2025technical

Recently, foundational models have significantly advanced in different tasks, accompanied by Transformer as the general backbone. However, Transformer's quadratic complexity poses challenges for handling longer sequences and higher resolution images, which may limit foundational models further devel…

Cited by 0SourcePDFScholar
2022

Detecting Camouflaged Object in Frequency Domain

CVPR 2022poster

Camouflaged object detection (COD) aims to identify objects that are perfectly embedded in their environment, which has various downstream applications in fields such as medicine, art, and agriculture. However, it is an extremely challenging task to spot camouflaged objects with the perception abili…

Cited by 217PDFcodeScholar
2021

Fast: Feature Aggregation for Detecting Salient Object in Real-Time

ICASSP 2021accepted

This paper introduces a method named FAST for real-time salient object detection with an extremely efficient CNN architecture. Our proposed network starts from a single lightweight backbone and aggregates discriminative features through network-level and phase-level respectively. Based on the multi-…

Cited by 0SourceScholar