← Search

Qi Wu

144 accepted papers

2026

Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learning

CVPR 2026

We contend that embodied learning is fundamentally a lifecycle problem rather than a single-stage optimization. Systems that optimize only one link (data collection, simulation, learning, or deployment) rarely sustain improvement or generalize beyond narrow settings. We introduce Arcadia, a closed-l

Cited by 0SourceScholar
2026

EVOLVING ROLLOUTS: Harnessing Historical Experience for Web Agent Evolution in Reinforcement Learning

ICML 2026poster

Agentic reinforcement learning (RL) for web search is prohibitively expensive due to long context lengths and costly environment interactions, and this inefficiency is further exacerbated by GRPO-based optimization, which discards learning signals from entire rollout groups with zero reward variance…

Cited by 0SourceScholar
2026

Embodied Navigation Foundation Model

ICLR 2026poster

Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments. To achieve such intelligence, recent advanced works leverage Vision-Language Models (VLMs), which demonstrate strong generalizability and possess a wel…

Cited by 0SourcecodeScholar
2026

Fast-SmartWay: Panoramic-Free End-To-End Zero-Shot Vision-And-Language Navigation

ICRA 2026poster

Recent advances in Vision-and-Language Navigation in Continuous Environments (VLN-CE) have leveraged multimodal large language models (MLLMs) to achieve zero-shot navigation. However, existing methods often rely on panoramic observations and two-stage pipelines involving waypoint predictors, which i…

2026

FisherPoser: Human Motion Estimation from Sparse Observations with Hierarchical Region-Wise Fisher-Matrix Uncertainty Modeling

CVPR 2026

Full-body motion estimation from sparse VR observations is an inherently under-constrained problem, with only three 6-DoF trackers (HMD and controllers) available to infer a full skeletal pose. To address this ambiguity, we introduce a probabilistic framework that models joint orientations as distri

Cited by 0SourceScholar
2026

LightMover: Generative Light Movement with Color and Intensity Controls

CVPR 2026

We present LightMover, a framework for controllable light manipulation in single images that leverages video diffusion priors to produce physically plausible illumination changes without re-rendering the scene. We formulate light editing as a sequence-to-sequence prediction problem in visual token s

Cited by 0SourceScholar
2026

MEDCUTMIX: A DATA-CENTRIC APPROACH TO IMPROVE RADIOLOGY VISION-LANGUAGE PRE-TRAINING WITH DISEASE AWARENESS

ICASSP 2026oral

Vision-Language Pre-training (VLP) is drawing increasing interest for its ability to minimize manual annotation requirements while enhancing semantic understanding in downstream tasks. However, its reliance on image-text datasets poses challenges due to privacy concerns and the high cost of obtainin…

Cited by 0SourcePDFScholar
2026

Manipulation Intention Understanding for Zero-Shot Composed Image Retrieval

AAAI 2026technical

Zero-shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with varied visual manipulation intents across domains, scenes, objects, and attributes. A key challenge is that existing datasets contain limited intent-relevant annotations, making it hard for models to infer human intent from text

Cited by 0SourcePDFScholar
2026

OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs

AAAI 2026technical

Existing sparse attention methods primarily target inference-time acceleration by selecting critical tokens under predefined sparsity patterns. However, they often fail to bridge the training–inference gap and lack the capacity for fine-grained token selection across multiple dimensions—such as quer

Cited by 0SourcePDFScholar
2026

RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-wave Point Cloud Sequence

AAAI 2026technical

Millimeter-wave radar offers a privacy-preserving and environment-robust alternative to vision-based sensing, enabling human motion analysis in challenging conditions such as low light, occlusions, rain, or smoke. However, its sparse point clouds pose significant challenges for semantic understandin

Cited by 0SourcePDFScholar
2026

SimULi: Real-Time LiDAR and Camera Simulation with Unscented Transforms

ICLR 2026poster

Rigorous testing of autonomous robots, such as self-driving vehicles, is essential to ensure their safety in real-world deployments. This requires building high-fidelity simulators to test scenarios beyond those that can be safely or exhaustively collected in the real-world. Existing neural renderin…

Cited by 0SourceScholar
2026

Sparsity Forcing: Reinforcing Token Sparsity of MLLMs

ICLR 2026poster

Sparse attention mechanisms aim to reduce computational overhead with minimal accuracy loss by selectively processing salient tokens. Despite their effectiveness, most methods merely exploit a model’s inherent sparsity and thus plateau at moderate budgets (about 50\% token reduction), with little he…

Cited by 0SourceScholar
2025

360Recon: An Accurate Reconstruction Method based on Depth Fusion from 360 Images

IROS 2025

Accurate 3D reconstruction is crucial for AR and VR applications. Compared with traditional pinhole camera-based methods, 360° image-based reconstruction can achieve higher precision with fewer input images, making it especially effective in low-texture environments. However, the severe distortion r

Cited by 2SourcecodeScholar
2025

3DGUT: Enabling Distorted Cameras and Secondary Rays in Gaussian Splatting

CVPR 2025poster

3D Gaussian Splatting (3DGS) enables efficient reconstruction and high-fidelity real-time rendering of complex scenes on consumer hardware. However, due to its rasterization-based formulation, 3DGS is constrained to ideal pinhole cameras and lacks support for secondary lighting effects. Recent meth…

2025

COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language Navigation

ICCV 2025poster

Vision-and-Language Navigation (VLN) tasks have gained prominence within artificial intelligence research due to their potential application in fields like home assistants. Many contemporary VLN approaches, while based on transformer architectures, have increasingly incorporated additional component…

2025

Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational Data

AAAI 2025technical

Using offline observational data for policy evaluation and learning allows decision-makers to evaluate and learn a policy that connects characteristics and interventions. Most existing literature has focused on either discrete treatment spaces or assumed no difference in the distributions between th…

Cited by 0SourcePDFScholar
2025

EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty Modeling

CVPR 2025poster

Estimating full-body motion using the tracking signals of head and hands from VR devices holds great potential for various applications. However, the sparsity and unique distribution of observations present a significant challenge, resulting in an ill-posed problem with multiple feasible solutions (…

2025

General Scene Adaptation for Vision-and-Language Navigation

ICLR 2025poster

Vision-and-Language Navigation (VLN) tasks mainly evaluate agents based on one-time execution of individual instructions across multiple environments, aiming to develop agents capable of functioning in any environment in a zero-shot manner. However, real-world navigation robots often operate in pers…

2025

Ground-Level Viewpoint Vision-and-Language Navigation in Continuous Environments

ICRA 2025

Vision-and-Language Navigation (VLN) empowers agents to associate time-sequenced visual observations with corresponding instructions to make sequential decisions. However, dealing with visually diverse scenes or transitioning from simulated environments to real-world deployment is still challenging.

Cited by 7SourceScholar
2025

Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models

IROS 2025

Learning-Based methods have achieved strong performance for quadrupedal locomotion. However, several challenges prevent quadrupeds from learning helpful indoor skills that require interaction with environments and humans: lack of end-effectors for manipulation, limited semantic under-standing using

Cited by 15SourceScholar
2025

LiDARDustX: A LiDAR Dataset for Dusty Unstructured Road Environments

ICRA 2025

Autonomous driving datasets are essential for validating the progress of intelligent vehicle algorithms, which include localization, perception, and prediction. However, existing datasets are predominantly focused on structured urban environments, which limits the exploration of unstructured and spe

Cited by 1SourcecodeScholar
2025

MFL-Owner: Ownership Protection for Multi-modal Federated Learning via Orthogonal Transform Watermark

AAAI 2025technical

Multi-modal Federated Learning (MFL) is a distributed machine learning paradigm that enables multiple participants with multi-modal data to collaboratively train a global model for multi-modal tasks without sharing their local data. MFL typically deploys the trained global model as an Embedding-as-a…

2025

Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System

ACL 2025long

The rapid advancement of scientific progress requires innovative tools that can accelerate knowledge discovery. Although recent AI methods, particularly large language models (LLMs), have shown promise in tasks such as hypothesis generation and experimental design, they fall short of replicating the…

2025

MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge Distillation

ICRA 2025

In recent years, Embodied Artificial Intelligence (Embodied AI) has advanced rapidly, yet the increasing size of models conflicts with the limited computational capabilities of Embodied AI platforms. To address this challenge, we aim to achieve both high model performance and practical deployability

Cited by 5SourceScholar
2025

Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval

CVPR 2025poster

Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manipulation intent across domain, scene, object, and attribute. The key challenge for ZS-CIR tasks is to modify a reference image according to manipulation text to accurately retrieve a target im…

2025

Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs

ICLR 2025poster

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce **Motion-Agent**, an efficient conversational framework designed for general human motion generati…

2025

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments remains underexplored. We present NavBench, a benchmark to evaluate the embodied navigation capabilities of MLLMs under zero…

Cited by 0SourceScholar
2025

Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs

ICRA 2025

Vision-and-Language Navigation (VLN) tasks require an agent to follow textual instructions to navigate through 3D environments. Traditional approaches use supervised learning methods, relying heavily on domain-specific datasets to train VLN models. Recent methods try to utilize closedsource large la

Cited by 49SourceScholar
2025

REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion Constraints

NeurIPS 2025poster

Articulated objects, as prevalent entities in human life, their 3D representations play crucial roles across various applications. However, achieving both high-fidelity textured surface reconstruction and dynamic generation for articulated objects remains challenging for existing methods. In this pa…

Cited by 0SourceScholar
2025

Realistic Noise Synthesis with Diffusion Models

AAAI 2025technical

Deep denoising models require extensive real-world training data, which is challenging to acquire. Current noise synthesis techniques struggle to accurately model complex noise distributions. We propose a novel Realistic Noise Synthesis Diffusor (RNSD) method using diffusion models to address these…

2025

Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval

CVPR 2025highlight

Composed Image Retrieval (CIR) aims to retrieve target images that closely resemble a reference image while integrating user-specified textual modifications, thereby capturing user intent more accurately. Existing training-free zero-shot CIR (ZS-CIR) methods often employ a two-stage process: they fi…

2025

SADA-3D: Structure-Aware Unsupervised Domain Adaptation Segmentation of 3D Point Clouds

RA-L 2025

Domain adaptive LiDAR point cloud segmentation aims to develop an effective target segmentation model using labeled source data and unlabeled target data. Existing domain adaptation methods for segmentation primarily focus on global feature alignment, often neglecting critical structural cues, espec

Cited by 0SourceScholar
2025

SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts

ICCV 2025poster

The academic field of learning instruction-guided visual navigation can be generally categorized into high-level category-specific search and low-level language-guided navigation, depending on the granularity of language instruction, in which the former emphasizes the exploration process, while the…

2025

Secure and Efficient Watermarking for Latent Diffusion Models in Model Distribution Scenarios

IJCAI 2025

Latent diffusion models have exhibited considerable potential in generative tasks. Watermarking is considered to be an alternative to safeguard the copyright of generative models and prevent their misuse. However, in the context of model distribution scenarios, the accessibility of models to large s

2025

SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation

IROS 2025

Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE frameworks rely on a two-stage approach: a waypoint predictor to generate waypoints and a navigator to execute movement

Cited by 13SourceScholar
2025

Suite-IN: Aggregating Motion Features from Apple Suite for Robust Inertial Navigation

ICRA 2025

With the rapid development of wearable technology, devices like smartphones, smartwatches, and headphones equipped with IMUs have become essential for applications such as pedestrian positioning. However, traditional pedestrian dead reckoning (PDR) methods struggle with diverse motion patterns, whil

Cited by 3SourceScholar
2025

mmDEAR: mmWave Point Cloud Density Enhancement for Accurate Human Body Reconstruction

ICRA 2025

Millimeter-wave (mmWave) radar offers robust sensing capabilities in diverse environments, making it a highly promising solution for human body reconstruction due to its privacy-friendly and non-intrusive nature. However, the significant sparsity of mm Wave point clouds limits the estimation accurac

Cited by 4SourceScholar
2024

Audio-Aided Learning Framework for Image Classification with Limited Training Images

ICASSP 2024accepted

It is challenging to train a generalizable deep learning classifier with limited training images. Existing few-shot learning approaches try to improve classification performance largely by transferring prior knowledge from upstream large-sample tasks to the current small-sample task. Besides upstrea…

Cited by 0SourceScholar
2024

Augmented Commonsense Knowledge for Remote Object Grounding

AAAI 2024technical

The vision-and-language navigation (VLN) task necessitates an agent to perceive the surroundings, follow natural language instructions, and act in photo-realistic unseen environments. Most of the existing methods employ the entire image or object features to represent navigable viewpoints. However,…

2024

Context-I2W: Mapping Images to Context-Dependent Words for Accurate Zero-Shot Composed Image Retrieval

AAAI 2024technical

Different from the Composed Image Retrieval task that requires expensive labels for training task-specific models, Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manipulation intent that could be related to domain, scene, object, and attribute…

2024

Continual Self-supervised Learning: Towards Universal Multi-modal Medical Data Representation Learning

CVPR 2024highlight

Self-supervised learning (SSL) is an efficient pre-training method for medical image analysis. However current research is mostly confined to certain modalities consuming considerable time and resources without achieving universality across different modalities. A straightforward solution is combini…

2024

Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-training Framework

CVPR 2024poster

Medical vision language pre-training (VLP) has emerged as a frontier of research enabling zero-shot pathological recognition by comparing the query image with the textual descriptions for each disease. Due to the complex semantics of biomedical texts current methods struggle to align medical images…

2024

Dynamic Inertial Poser (DynaIP): Part-Based Motion Dynamics Learning for Enhanced Human Pose Estimation with Sparse Inertial Sensors

CVPR 2024poster

This paper introduces a novel human pose estimation approach using sparse inertial sensors addressing the shortcomings of previous methods reliant on synthetic data. It leverages a diverse array of real inertial motion capture data from different skeleton formats to improve motion diversity and mode…

2024

Dynamicity-aware Social Bot Detection with Dynamic Graph Transformers

IJCAI 2024poster

Detecting social bots has evolved into a pivotal yet intricate task, aimed at combating the dissemination of misinformation and preserving the authenticity of online interactions. While earlier graph-based approaches, which leverage topological structure of social networks, yielded notable outcomes,…

2024

Everyday Object Meets Vision-and-Language Navigation Agent via Backdoor

NeurIPS 2024poster

Vision-and-Language Navigation (VLN) requires an agent to dynamically explore environments following natural language. The VLN agent, closely integrated into daily lives, poses a substantial threat to the security of privacy and property upon the occurrence of malicious behavior. However, this serio…

Cited by 0SourcePDFScholar
2024

Explicit Interaction for Fusion-Based Place Recognition

IROS 2024poster

Fusion-based place recognition is an emerging technique jointly utilizing multi-modal perception data, to recognize previously visited places in GPS-denied scenarios for robots and autonomous vehicles. Recent fusion-based place recognition methods combine multi-modal features in implicit manners. Wh…

Cited by 2SourcecodeScholar
2024

G-NeRF: Geometry-enhanced Novel View Synthesis from Single-View Images

CVPR 2024poster

Novel view synthesis aims to generate new view images of a given view image collection. Recent attempts address this problem relying on 3D geometry priors (e.g. shapes sizes and positions) learned from multi-view images. However such methods encounter the following limitations: 1) they require a set…

2024

GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning

NeurIPS 2024poster

Large Language Models (LLMs) are increasingly used for various tasks with graph structures. Though LLMs can process graph information in a textual format, they overlook the rich vision modality, which is an intuitive way for humans to comprehend structural information and conduct general graph reaso…

2024

HumanPlus: Humanoid Shadowing and Imitation from Humans

CoRL 2024poster

One of the key arguments for building robots that have similar form factors to human beings is that we can leverage the massive human data for training.Yet, doing so has remained challenging in practice due to the complexities in humanoid perception and control, lingering physical gaps between human…

Cited by 109SourceScholar
2024

KPA-Tracker: Towards Robust and Real-Time Category-Level Articulated Object 6D Pose Tracking

AAAI 2024technical

Our life is populated with articulated objects. Current category-level articulation estimation works largely focus on predicting part-level 6D poses on static point cloud observations. In this paper, we tackle the problem of category-level online robust and real-time 6D pose tracking of articulated…

2024

LLM as Copilot for Coarse-grained Vision-and-Language Navigation

ECCV 2024poster

"Vision-and-Language Navigation (VLN) involves guiding an agent through indoor environments using human-provided textual instructions. Coarse-grained VLN, with short and high-level instructions, has gained popularity as it closely mirrors real-world scenarios. However, a significant challenge is the…

Cited by 9SourcePDFScholar
2024

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

RSS 2024poster

Vision-and-language navigation (VLN) stands as a key research problem of Embodied AI, aiming at enabling agents to navigate in unseen environments following linguistic instructions. In this field, generalization is a long-standing challenge, either to out-of-distribution scenes or from Sim to Real.…

Cited by 80SourcePDFScholar
2024

NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models

ECCV 2024poster

"Capitalizing on the remarkable advancements in Large Language Models (LLMs), there is a burgeoning initiative to harness LLMs for instruction following robotic navigation. Such a trend underscores the potential of LLMs to generalize navigational reasoning and diverse language understanding. However…

2024

NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

AAAI 2024technical

Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling. Such a trend underscored the potential of training LLMs with unlimited language data, advancing the development of a universal…

2024

PairAug: What Can Augmented Image-Text Pairs Do for Radiology?

CVPR 2024poster

Current vision-language pre-training (VLP) methodologies predominantly depend on paired image-text datasets a resource that is challenging to acquire in radiology due to privacy considerations and labelling complexities. Data augmentation provides a practical solution to overcome the issue of data s…

2024

Sparse Bayesian Deep Learning for Cross Domain Medical Image Reconstruction

AAAI 2024technical

Cross domain medical image reconstruction aims to address the issue that deep learning models trained solely on one source dataset might not generalize effectively to unseen target datasets from different hospitals. Some recent methods achieve satisfactory reconstruction performance, but often at th…

Cited by 4SourcePDFScholar
2024

Stepwise Multi-grained Boundary Detector for Point-supervised Temporal Action Localization

ECCV 2024poster

"Point-supervised temporal action localization pursues high-accuracy action detection under low-cost data annotation. Despite recent advances, a significant challenge remains: sparse labeling of individual frames leads to semantic ambiguity in determining action boundaries due to the lack of continu…

Cited by 0SourcePDFScholar
2024

The Causal Impact of Credit Lines on Spending Distributions

AAAI 2024technical

Consumer credit services offered by electronic commerce platforms provide customers with convenient loan access during shopping and have the potential to stimulate sales. To understand the causal impact of credit lines on spending, previous studies have employed causal estimators, (e.g., direct regr…

2024

Thermal-NeRF: Neural Radiance Fields from an Infrared Camera

IROS 2024poster

In recent years, Neural Radiance Fields (NeRFs) have demonstrated significant potential in encoding highly-detailed 3D geometry and environmental appearance, positioning themselves as a promising alternative to traditional explicit representation for 3D scene reconstruction. However, the predominant…

Cited by 13SourcecodeScholar
2024

Unveiling the Potential of Robustness in Selecting Conditional Average Treatment Effect Estimators

NeurIPS 2024poster

The growing demand for personalized decision-making has led to a surge of interest in estimating the Conditional Average Treatment Effect (CATE). Various types of CATE estimators have been developed with advancements in machine learning and causal inference. However, selecting the desirable CATE est…

2024

Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles

NeurIPS 2024poster

While advancements in NLP have significantly improved the performance of Large Language Models (LLMs) on tasks requiring vertical thinking, their lateral thinking capabilities remain under-explored and challenging to measure due to the complexity of assessing creative thought processes and the scarc…

2024

WebVLN: Vision-and-Language Navigation on Websites

AAAI 2024technical

Vision-and-Language Navigation (VLN) task aims to enable AI agents to accurately understand and follow natural language instructions to navigate through real-world environments, ultimately reaching specific target locations. We recognise a promising opportunity to extend VLN to a comparable navigati…

2024

Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts

IJCAI 2024poster

Current Vision-and-Language Navigation (VLN) tasks mainly employ textual instructions to guide agents. However, being inherently abstract, the same textual instruction can be associated with different visual signals, causing severe ambiguity and limiting the transfer of prior knowledge in the vision…

2024

mmBaT: A Multi-Task Framework for Mmwave-Based Human Body Reconstruction and Translation Prediction

ICASSP 2024accepted

Human body reconstruction with Millimeter Wave (mmWave) radar point clouds has gained significant interest due to its ability to work in adverse environments and its capacity to mitigate privacy concerns associated with traditional camera-based solutions. Despite pioneering efforts in this field, tw…

Cited by 0SourceScholar
2023

A Unified Perspective on Regularization and Perturbation in Differentiable Subset Selection

AISTATS 2023poster

Subset selection, i.e., finding a bunch of items from a collection to achieve specific goals, has wide applications in information retrieval, statistics, and machine learning. To implement an end-to-end learning framework, different relaxed differentiable operators of subset selection are proposed.…

2023

AerialVLN: Vision-and-Language Navigation for UAVs

ICCV 2023poster

Recently emerged Vision-and-Language Navigation(VLN) tasks have drawn significant attention in both computer vision and natural language processing communities. Existing VLN tasks are built for agents that navigate on the ground, either indoors or outdoors. However, many tasks require intelligent ag…

Cited by 47PDFcodeScholar
2023

DeLELSTM: Decomposition-based Linear Explainable LSTM to Capture Instantaneous and Long-term Effects in Time Series

IJCAI 2023poster

Time series forecasting is prevalent in various real-world applications. Despite the promising results of deep learning models in time series forecasting, especially the Recurrent Neural Networks (RNNs), the explanations of time series models, which are critical in high-stakes applications, have rec…

2023

Digging out Discrimination Information from Generated Samples for Robust Visual Question Answering

ACL 2023findings

Visual Question Answering (VQA) aims to answer a textual question based on a given image. Nevertheless, recent studies have shown that VQA models tend to capture the biases to answer the question, instead of using the reasoning ability, resulting in poor generalisation ability. To alleviate the issu…

Cited by 9SourcePDFScholar
2023

Learning To Dub Movies via Hierarchical Prosody Models

CVPR 2023poster

Given a piece of text, a video clip and a reference audio, the movie dubbing (also known as visual voice clone, V2C) task aims to generate speeches that match the speaker's emotion presented in the video using the desired speaker voice as reference. V2C is more challenging than conventional text-to-…

2023

LoRA: A Logical Reasoning Augmented Dataset for Visual Question Answering

NeurIPS 2023poster

The capacity to reason logically is a hallmark of human cognition. Humans excel at integrating multimodal information for locigal reasoning, as exemplified by the Visual Question Answering (VQA) task, which is a challenging multimodal task. VQA tasks and large vision-and-language models aim to tackl…

2023

March in Chat: Interactive Prompting for Remote Embodied Referring Expression

ICCV 2023poster

Many Vision-and-Language Navigation (VLN) tasks have been proposed in recent years, from room-based to object-based and indoor to outdoor. The REVERIE (Remote Embodied Referring Expression) is interesting since it only provides high-level instructions to the agent, which are closer to human commands…

Cited by 39PDFcodeScholar
2023

NeRF-LOAM: Neural Implicit Representation for Large-Scale Incremental LiDAR Odometry and Mapping

ICCV 2023poster

Simultaneously odometry and mapping using LiDAR data is an important task for mobile systems to achieve full autonomy in large-scale environments. However, most existing LiDAR-based methods prioritize tracking quality over reconstruction quality. Although the recently developed neural radiance field…

Cited by 75PDFcodeScholar
2023

Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval

ICCV 2023poster

In text-video retrieval, recent works have benefited from the powerful learning capabilities of pre-trained text-image foundation models (e.g., CLIP) by adapting them to the video domain. A critical problem for them is how to effectively capture the rich semantics inside the video using the image en…

Cited by 40PDFcodeScholar
2023

S3C: Semi-Supervised VQA Natural Language Explanation via Self-Critical Learning

CVPR 2023poster

VQA Natural Language Explanation (VQA-NLE) task aims to explain the decision-making process of VQA models in natural language. Unlike traditional attention or gradient analysis, free-text rationales can be easier to understand and gain users' trust. Existing methods mostly use post-hoc or self-ratio…

Cited by 10SourcePDFScholar
2023

Scaling Data Generation in Vision-and-Language Navigation

ICCV 2023oral

Recent research in language-guided visual navigation has demonstrated a significant demand for the diversity of traversable environments and the quantity of supervision for training generalizable agents. To tackle the common data scarcity issue in existing vision-and-language navigation datasets, we…

Cited by 80PDFcodeScholar
2023

Towards Balanced Representation Learning for Credit Policy Evaluation

AISTATS 2023poster

Credit policy evaluation presents profitable opportunities for E-commerce platforms through improved decision-making. The core of policy evaluation is estimating the causal effects of the policy on the target outcome. However, selection bias presents a key challenge in estimating causal effects from…

2022

A Simple and Robust Correlation Filtering Method for Text-Based Person Search

ECCV 2022poster

"Text-based person search aims to associate pedestrian images with natural language descriptions. In this task, extracting differentiated representations and aligning them among identities and descriptions is an essential yet challenging problem. Most of the previous methods depend on additional lan…

2022

Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language Navigation

CVPR 2022poster

Most existing works in vision-and-language navigation (VLN) focus on either discrete or continuous environments, training agents that cannot generalize across the two. Although learning to navigate in continuous spaces is closer to the real-world, training such an agent is significantly more difficu…

Cited by 86PDFcodeScholar
2022

Diagnosing Vision-and-Language Navigation: What Really Matters

NAACL 2022long

Vision-and-language navigation (VLN) is a multimodal task where an agent follows natural language instructions and navigates in visual environments. Multiple setups have been proposed, and researchers apply new model architectures or training techniques to boost navigation performance. However, ther…

2022

HOP: History-and-Order Aware Pre-Training for Vision-and-Language Navigation

CVPR 2022poster

Pre-training has been adopted in a few of recent works for Vision-and-Language Navigation (VLN). However, previous pre-training methods for VLN either lack the ability to predict future actions or ignore the trajectory contexts, which are essential for a greedy navigation process. In this work, to p…

Cited by 94PDFcodeScholar
2022

Learning the Dynamics of Visual Relational Reasoning via Reinforced Path Routing

AAAI 2022technical

Reasoning is a dynamic process. In cognitive theories, the dynamics of reasoning refers to reasoning states over time after successive state transitions. Modeling the cognitive dynamics is of utmost importance to simulate human reasoning capability. In this paper, we propose to learn the reasoning d…

Cited by 8SourcePDFScholar
2022

Maintaining Reasoning Consistency in Compositional Visual Question Answering

CVPR 2022poster

A compositional question refers to a question that contains multiple visual concepts (e.g., objects, attributes, and relationships) and requires compositional reasoning to answer. Existing VQA models can answer a compositional question well, but cannot work well in terms of reasoning consistency in…

Cited by 29PDFcodeScholar
2022

MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-Based Visual Question Answering

CVPR 2022poster

Knowledge-based visual question answering requires the ability of associating external knowledge for open-ended cross-modal scene understanding. One limitation of existing solutions is that they capture relevant knowledge from text-only knowledge bases, which merely contain facts expressed by first-…

Cited by 136PDFcodeScholar
2022

UniMiSS: Universal Medical Self-Supervised Learning via Breaking Dimensionality Barrier

ECCV 2022poster

"Self-supervised learning (SSL) opens up huge opportunities for medical image analysis that is well known for its lack of annotations. However, aggregating massive (unlabeled) 3D medical images like computerized tomography (CT) remains challenging due to its high imaging cost and privacy restriction…

2022

Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions

ACL 2022long

A long-term goal of AI research is to build intelligent agents that can communicate with humans in natural language, perceive the environment, and perform real-world tasks. Vision-and-Language Navigation (VLN) is a fundamental and interdisciplinary research topic towards this goal, and receives incr…

2021

Communicative Learning with Natural Gestures for Embodied Navigation Agents with Human-in-the-Scene

IROS 2021poster

Human-robot collaboration is an essential re-search topic in artificial intelligence (AI), enabling researchers to devise cognitive AI systems and affords an intuitive means for users to interact with the robot. Of note, communication plays a central role. To date, prior studies in embodied agent na…

Cited by 24SourceScholar
2021

Confidence-aware Non-repetitive Multimodal Transformers for TextCaps

AAAI 2021technical

When describing an image, reading text in the visual scene is crucial to understand the key information. Recent work explores the TextCaps task, i.e. image captioning with reading Optical Character Recognition (OCR) tokens, which requires models to read text and cover them in generated captions. Exi…

2021

Debiased Visual Question Answering from Feature and Sample Perspectives

NeurIPS 2021poster

Visual question answering (VQA) is designed to examine the visual-textual reasoning ability of an intelligent agent. However, recent observations show that many VQA models may only capture the biases between questions and answers in a dataset rather than showing real reasoning abilities. For example…

2021

Jo-SRC: A Contrastive Approach for Combating Noisy Labels

CVPR 2021poster

Due to the memorization effect in Deep Neural Networks (DNNs), training with noisy labels usually results in inferior model performance. Existing state-of-the-art methods primarily adopt a sample selection strategy, which selects small-loss samples for subsequent training. However, prior literature…

Cited by 190PDFScholar
2021

Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision

NeurIPS 2021poster

In Vision-and-Language Navigation (VLN) task, an agent is asked to navigate inside 3D indoor environments following given instructions. Cross-modal alignment is one of the most critical challenges in VLN because the predicted trajectory needs to match the given instruction accurately. In this paper,…

2021

Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation

CVPR 2021poster

Semantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of segmentation. However, existing works mainly concentrate on expanding the seed of…

Cited by 248PDFcodeScholar
2021

Proposal-free One-stage Referring Expression via Grid-Word Cross-Attention

IJCAI 2021poster

Referring Expression Comprehension (REC) has become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as visual question answering. However, it has not been widely used in many downstream tasks because it suffers 1) two-stage m…

Cited by 12SourcePDFScholar
2021

Room-and-Object Aware Knowledge Reasoning for Remote Embodied Referring Expression

CVPR 2021poster

The Remote Embodied Referring Expression (REVERIE) is a recently raised task that requires an agent to navigate to and localise a referred remote object according to a high-level language instruction. Different from related VLN tasks, the key to REVERIE is to conduct goal-oriented exploration instea…

Cited by 91PDFcodeScholar
2021

Simple is not Easy: A Simple Strong Baseline for TextVQA and TextCaps

AAAI 2021technical

Texts appearing in daily scenes that can be recognized by OCR (Optical Character Recognition) tools contain significant information, such as street name, product brand and prices. Two tasks -- text-based visual question answering and text-based image captioning, with a text extension from existing v…

2021

Smooth-RRT*: Asymptotically Optimal Motion Planning for Mobile Robots under Kinodynamic Constraints

ICRA 2021poster

Nowadays, various algorithms based on the Rapidly-exploring Random Tree (RRT) methods are utilized to solve motion planning problems. Based on the RRT*, we developed a novel reconnection method that enables the planner to directly generate a smooth curved trajectory. Meanwhile, kinodynamic constrain…

Cited by 11SourceScholar
2021

The Causal Learning of Retail Delinquency

AAAI 2021technical

This paper focuses on the expected difference in borrower's repayment when there is a change in the lender's credit decisions. Classical estimators overlook the confounding effects and hence the estimation error can be magnificent. As such, we propose another approach to construct the estimators suc…

Cited by 8SourcePDFScholar
2021

The Road To Know-Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language Navigation

ICCV 2021poster

Vision-and-Language Navigation (VLN) requires an agent to find a path to a remote location on the basis of natural-language instructions and a set of photo-realistic panoramas. Most existing methods take the words in the instructions and the discrete views of each panorama as the minimal unit of enc…

Cited by 87PDFcodeScholar
2021

Towards Accurate Text-Based Image Captioning With Content Diversity Exploration

CVPR 2021poster

Text-based image captioning (TextCap) which aims to read and reason images with texts is crucial for a machine to understand a detailed and complex scene environment, considering that texts are omnipresent in daily life. This task, however, is very challenging because an image often contains complex…

Cited by 84PDFcodeScholar
2021

VLN BERT: A Recurrent Vision-and-Language BERT for Navigation

CVPR 2021poster

Accuracy of many visiolinguistic tasks has benefited significantly from the application of vision-and-language (V&L) BERT. However, its application for the task of vision-and-language navigation (VLN) remains limited. One reason for this is the difficulty adapting the BERT architecture to the partia…

Cited by 319PDFcodeScholar
2020

Cops-Ref: A New Dataset and Task on Compositional Referring Expression Comprehension

CVPR 2020poster

Referring expression comprehension (REF) aims at identifying a particular object in a scene by a natural language expression. It requires joint reasoning over the textual and visual domains to solve the problem. Some popular referring expression datasets, however, fail to provide an ideal test bed f…

Cited by 75PDFScholar
2020

DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue

IJCAI 2020poster

Visual Dialogue task requires an agent to be engaged in a conversation with human about an image. The ability of generating detailed and non-repetitive responses is crucial for the agent to achieve human-like conversation. In this paper, we propose a novel generative decoding architecture to generat…

2020

Gold Seeker: Information Gain From Policy Distributions for Goal-Oriented Vision-and-Langauge Reasoning

CVPR 2020poster

As Computer Vision moves from passive analysis of pixels to active analysis of semantics, the breadth of information algorithms need to reason over has expanded significantly. One of the key challenges in this vein is the ability to identify the information required to make a decision, and select an…

Cited by 6PDFScholar
2020

Intelligent Home 3D: Automatic 3D-House Design From Linguistic Descriptions Only

CVPR 2020poster

Home design is a complex task that normally requires architects to finish with their professional skills and tools. It will be fascinating that if one can produce a house plan intuitively without knowing much knowledge about home design and experience of using complex designing tools, for example, v…

Cited by 47PDFcodeScholar
2020

Language and Visual Entity Relationship Graph for Agent Navigation

NeurIPS 2020poster

Vision-and-Language Navigation (VLN) requires an agent to navigate in a real-world environment following natural language instructions. From both the textual and visual perspectives, we find that the relationships among the scene, its objects, and directional cues are essential for the agent to inte…

2020

Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering

IJCAI 2020poster

Fact-based Visual Question Answering (FVQA) requires external knowledge beyond the visible content to answer questions about an image. This ability is challenging but indispensable to achieve general VQA. One limitation of existing FVQA solutions is that they jointly embed all kinds of information w…

2020

Object-and-Action Aware Model for Visual Language Navigation

ECCV 2020poster

Vision-and-Language Navigation (VLN) is unique in that it requires turning relatively general natural-language instructions into robot agent actions, on the basis of visible environments. This requires to extract value from two very different types of natural-language information. The first is objec…

Cited by 133SourcePDFScholar
2020

REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments

CVPR 2020oral

One of the long-term challenges of robotics is to enable robots to interact with humans in the visual world via natural language, as humans are visual animals that communicate through language. Overcoming this challenge requires the ability to perform a wide variety of complex tasks in response to m…

Cited by 374PDFcodeScholar
2020

Say As You Wish: Fine-Grained Control of Image Caption Generation With Abstract Scene Graphs

CVPR 2020oral

Humans are able to describe image contents with coarse to fine details as they wish. However, most image captioning models are intention-agnostic which cannot generate diverse descriptions according to different user intentions initiatively. In this work, we propose the Abstract Scene Graph (ASG) st…

Cited by 287PDFcodeScholar
2020

Semantic Equivalent Adversarial Data Augmentation for Visual Question Answering

ECCV 2020poster

Visual Question Answering (VQA) has achieved great success thanks to the fast development of deep neural networks (DNN). On the other hand, the data augmentation, as one of the major tricks for DNN, has been widely used in many computer vision tasks. However, there are few works studying the data au…

2019

Mind Your Neighbours: Image Annotation With Metadata Neighbourhood Graph Co-Attention Networks

CVPR 2019poster

As the visual reflections of our daily lives, images are frequently shared on the social network, which generates the abundant 'metadata' that records user interactions with images. Due to the diverse contents and complex styles, some images can be challenging to recognise when neglecting the contex…

Cited by 25PDFScholar
2019

Neighbourhood Watch: Referring Expression Comprehension via Language-Guided Graph Attention Networks

CVPR 2019poster

The task in referring expression comprehension is to localize the object instance in an image described by a referring expression phrased in natural language. As a language-to-vision matching task, the key to this problem is to learn a discriminative object feature that can adapt to the expression u…

Cited by 305PDFScholar
2019

What's to Know? Uncertainty as a Guide to Asking Goal-Oriented Questions

CVPR 2019poster

One of the core challenges in Visual Dialogue problems is asking the question that will provide the most useful information towards achieving the required objective. Encouraging an agent to ask the right questions is difficult because we don't know a-priori what information the agent will need to a…

Cited by 22PDFScholar
2018

Are You Talking to Me? Reasoned Visual Dialog Generation Through Adversarial Learning

CVPR 2018poster

The Visual Dialogue task requires an agent to engage in a conversation about an image with a human. It represents an extension of the Visual Question Answering task in that the agent needs to answer a question about an image, but it needs to do so in light of the previous dialogue that has taken pl…

Cited by 148SourcePDFScholar
2018

Goal-Oriented Visual Question Generation via Intermediate Rewards

ECCV 2018poster

Despite significant progress in a variety of vision-and-language problems, developing a method capable of asking intelligent, goal-oriented questions about images is proven to be an inscrutable challenge. Towards this end, we propose a Deep Reinforcement Learning framework based on three new interme…

Cited by 47SourcePDFScholar
2018

Learning Semantic Concepts and Order for Image and Sentence Matching

CVPR 2018poster

Image and sentence matching has made great progress recently, but it remains challenging due to the large visual semantic discrepancy. This mainly arises from that the representation of pixel-level image usually lacks of high-level semantic information as in its matched sentence. In this work, we pr…

Cited by 409SourcePDFScholar
2018

Parallel Attention: A Unified Framework for Visual Object Discovery Through Dialogs and Queries

CVPR 2018poster

Recognising objects according to a pre-defined fixed set of class labels has been well studied in the Computer Vision. There are a great many practical applications where the subjects that may be of interest are not known beforehand, or so easily delineated, however. In many of these cases natural l…

Cited by 158SourcePDFScholar
2018

Parsimonious Quantile Regression of Financial Asset Tail Dynamics via Sequential Learning

NeurIPS 2018poster

We propose a parsimonious quantile regression framework to learn the dynamic tail behaviors of financial asset returns. Our model captures well both the time-varying characteristic and the asymmetrical heavy-tail property of financial time series. It combines the merits of a popular sequential neura…

Cited by 31SourcePDFScholar
2018

Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments

CVPR 2018poster

A robot that can carry out a natural-language instruction has been a dream since before the Jetsons cartoon series imagined a life of leisure mediated by a fleet of attentive robot helpers. It is a dream that remains stubbornly distant. However, recent advances in vision and language methods have m…

2018

Visual Question Answering With Memory-Augmented Networks

CVPR 2018poster

In this paper, we exploit memory-augmented neural networks to predict accurate answers to visual questions, even when those answers rarely occur in the training set. The memory network incorporates both internal and external memory blocks and selectively pays attention to each training exemplar. We…

Cited by 134SourcePDFScholar
2017

The VQA-Machine: Learning How to Use Existing Vision Algorithms to Answer New Questions

CVPR 2017poster

One of the most intriguing features of the Visual Question Answering (VQA) challenge is the unpredictability of the questions. Extracting the information required to answer them demands a variety of image operations from detection and counting, to segmentation and reconstruction. To train a method t…

Cited by 105PDFcodeScholar
2016

Ask Me Anything: Free-Form Visual Question Answering Based on Knowledge From External Sources

CVPR 2016spotlight

We propose a method for visual question answering which combines an internal representation of the content of an image with information extracted from a general knowledge base to answer a broad range of image-based questions. This allows more complex questions to be answered using the predominant ne…

Cited by 475PDFScholar
2016

What Value Do Explicit High Level Concepts Have in Vision to Language Problems?

CVPR 2016poster

Much recent progress in Vision-to-Language (V2L) problems has been achieved through a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This approach does not explicitly represent high-level semantic concepts, but rather seeks to progress directly from image f…

Cited by 561PDFScholar