← Search

Long Chen

160 accepted papers

2026

AKCMamba-YOLO: Selective State Space Models For Real-Time Object Detection

CVPR 2026

The YOLO (You Only Look Once) series has been a cornerstone in real-time object detection, renowned for its efficient convolutional design and rapid inference. However, its reliance on convolutional operations inherently limits its ability to capture long-range dependencies and rich contextual infor

Cited by 0SourcecodeScholar
2026

AdaThinkDrive: Adaptive Thinking Via Reinforcement Learning for Autonomous Driving

ICRA 2026poster

While reasoning technology like Chain-of-Thought (CoT) has been widely adopted in Vision-Language-Action (VLA) models, it demonstrates promising capabilities in end-to-end autonomous driving. However, recent efforts to integrate CoT reasoning often fall short in simple scenarios, introducing unneces…

2026

AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture -of-Transformers for End-to-End Autonomous Driving

ICML 2026poster

Integrating vision-language models (VLMs) into end-to-end (E2E) autonomous driving (AD) systems has shown promise in improving scene understanding. However, existing integration strategies suffer from several limitations: they either struggle to resolve distribution misalignment between reasoning an…

Cited by 0SourcecodeScholar
2026

BTSP-CAM: A Brain-Inspired Geometric Memory for Class-Incremental Learning

ICML 2026poster

Gradient-based optimization in class-incremental learning (CIL) often faces the plasticity–stability dilemma, since continuous weight updates can distort decision boundaries learned from earlier tasks. We revisit this problem from the viewpoint of stochastic geometric memory allocation and propose B…

Cited by 0SourceScholar
2026

Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction

CVPR 2026

Integrating segmentation into Multimodal Large Language Models (MLLMs) presents a core trilemma: simultaneously preserving dialogue ability, achieving high segmentation performance, and ensuring fast inference. Prevailing paradigms are forced into a compromise. Embedding prediction methods introduce

Cited by 0SourcecodeScholar
2026

DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images

CVPR 2026

Autonomous driving needs fast, scalable 4D reconstruction and re-simulation for training and evaluation, yet most methods for dynamic driving scenes still rely on per-scene optimization, known camera calibration, or short frame windows, making them slow and impractical. We revisit this problem from

Cited by 0SourcecodeScholar
2026

Dissecting Post-Training: Uncovering the Complementary Roles of SFT and RL for Document Parsing

ICML 2026poster

Document parsing, the task of extracting diverse content from PDFs while preserving their structural integrity, has been significantly advanced by Multimodal Large Language Models (MLLMs). These models have achieved remarkable success, largely driven by extensive post-training on massive datasets. T…

Cited by 0SourceScholar
2026

DriveWorld-VLA: Unified Latent-Space World Modeling with Vision–Language–Action for Autonomous Driving

ICML 2026poster

End-to-end (E2E) autonomous driving has recently attracted increasing interest in unifying Vision–Language–Action (VLA) with World Models to enhance decision-making and forward-looking imagination. However, existing methods fail to effectively unify future scene evolution and action planning within …

Cited by 17SourceScholar
2026

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

ICLR 2026poster

The differing representation spaces required for visual understanding and generation pose a challenge in unifying them within the autoregressive paradigm of large language models. A vision tokenizer trained for reconstruction excels at capturing low-level visual appearance, making it well-suited for…

Cited by 0SourcecodeScholar
2026

Efficient Transformer Attention for SNNs via Hadamard Simplification

ICML 2026poster

Spiking Neural Networks (SNNs) offer low-power, brain-inspired computation, but Transformer-based SNNs face deployment challenges on neuromorphic hardware due to complex operations and high communication overhead. We propose hardware-efficient attention mechanisms, \textbf{Simplified Spiking Attenti…

Cited by 0SourceScholar
2026

Enhancing Diffusion Policies with Distribution-Matching Generator in Offline Reinforcement Learning

AAAI 2026technical

Offline reinforcement learning (RL) can learn policies from pre-collected offline datasets without interacting with the environment, but it suffers from the issue of out-of-distribution (OOD). Recent methods use the generative adversarial paradigm to learn policies, but easily fail to handle the con

Cited by 0SourcePDFScholar
2026

Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark

ICML 2026poster

Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, handwritten single-variable calculus work from a major U.S. public research university. Using OCR-conditioned large language…

Cited by 0SourceScholar
2026

FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing

CVPR 2026

With the surge of pre-trained text-to-image flow matching models, text-based image editing performance has gained remarkable improvement, especially for **simple editing** that only contains a single editing target. However, to satisfy the exploding editing requirements, the **complex editing** that

Cited by 0SourceScholar
2026

GIR-Bench: Versatile Benchmark for Generating Images with Reasoning

ICLR 2026poster

Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal intelligence. However, the community still lacks a rigorous reasoning-centric benchmark to systematically evaluate the align…

Cited by 0SourcecodeScholar
2026

Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation

CVPR 2026

Vision-Language Models (VLMs), leveraging their powerful visual perception and reasoning capabilities, have been widely applied in Unmanned Aerial Vehicle (UAV) tasks.However, the spatial intelligence capabilities of existing VLMs in UAV scenarios remain largely unexplored, raising concerns about th

Cited by 0SourcecodeScholar
2026

LAS: Loss-less ANN-SNN Conversion for Fully Spike-Driven Large Language Models

AAAI 2026technical

Spiking Large Language Models (LLMs) have emerged as an energy-efficient alternative to conventional LLMs through their event-driven computation. To effectively obtain spiking LLMs, researchers develop different ANN-to-SNN conversion methods by leveraging pre-trained ANN parameters while inheriting

Cited by 0SourcePDFScholar
2026

LLM-Guided Diagnostic Evidence Alignment for Medical Vision–Language Pretraining under Limited Pairing

ICML 2026poster

Most existing CLIP-style medical vision--language pretraining methods rely on global or local alignment with substantial paired data. However, global alignment is easily dominated by non-diagnostic information, while local alignment fails to integrate key diagnostic evidence. As a result, learning r…

Cited by 0SourceScholar
2026

MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving

CVPR 2026

Generative models have shown great potential in trajectory planning. Recent studies demonstrate that anchor-guided generative models are effective in modeling the uncertainty of driving behaviors and improving overall performance. However, these methods rely on discrete anchor vocabularies that must

Cited by 0SourcecodeScholar
2026

Path-Decoupled Hyperbolic Flow Matching for Few-Shot Adaptation

ICML 2026poster

Recent advances in cross-modal few-shot adaptation treat visual-semantic alignment as a continuous feature transport problem via Flow Matching (FM). However, we argue that Euclidean-based FM overlooks fundamental limitations of flat geometry, where polynomial volume growth fails to accommodate diver…

Cited by 0SourceScholar
2026

PerlAD: Towards Enhanced Closed-Loop End-to-End Autonomous Driving With Pseudo-Simulation-Based Reinforcement Learning

RA-L 2026

End-to-end autonomous driving policies based on Imitation Learning (IL) often struggle in closed-loop execution due to the misalignment between inadequate open-loop training objectives and real driving requirements. While Reinforcement Learning (RL) offers a solution by directly optimizing driving g

Cited by 1SourceScholar
2026

Personalize Your Gaussian: Consistent 3D Scene Personalization from a Single Image

AAAI 2026technical

Personalizing 3D scenes from a single reference image enables intuitive user-guided editing, which requires achieving both multi-view consistency across perspectives and referential consistency with the input image. However, these goals are particularly challenging due to the viewpoint bias caused b

Cited by 0SourcePDFScholar
2026

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

CVPR 2026

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-scale or complex dynamics. This limitation arises primarily because existing approa

Cited by 0SourceScholar
2026

Property Enhanced Instruction Tuning for Multi-Task Molecule Generation with Large Language Models

IJCAI 2026

Large language models (LLMs) are widely applied in various natural language processing tasks such as question answering and machine translation. However, due to the lack of labeled data and the difficulty of manual annotation for biochemical properties, the performance for molecule generation tasks

Cited by 0Scholar
2026

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

ICLR 2026poster

Recent studies have explored leveraging the world knowledge and cognitive capabilities of Vision-Language Models (VLMs) to address the long-tail problem in end-to-end autonomous driving. However, existing methods typically formulate trajectory planning as a language modeling task, where physical act…

Cited by 0SourcecodeScholar
2026

Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension

AAAI 2026technical

Recent advances in multi-modal large language models (MLLMs) have significantly improved object-level grounding and region captioning. However, they remain limited in visual relation understanding, struggling even with binary relation detection, let alone N-ary relations involving multiple semantic

Cited by 0SourcePDFScholar
2026

Rendering Multi-Human and Multi-Object with 3D Gaussian Splatting

ICRA 2026poster

Reconstructing dynamic scenes with multiple interacting humans and objects from sparse-view inputs is a critical yet challenging task, essential for creating high-fidelity digital twins for robotics and VR/AR. This problem, which we term Multi-Human Multi-Object (MHMO) rendering, presents two signif…

2026

Robustifying Vision-Language Models via Test-Time Prompt Adaptation

ICML 2026poster

Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing test-time adaptation methods typically rely on sample-level confidence heuristics, overlooking the intrinsic distributional…

Cited by 0SourceScholar
2026

SEF-MAP: Subspace-Decomposed Expert Fusion for Robust Multimodal HD Map Prediction

ICRA 2026poster

High-definition (HD) maps are essential for autonomous driving, yet multi-modal fusion often suffers from inconsistency between camera and LiDAR modalities, leading to performance degradation under low-light conditions, occlusions, or sparse point clouds. To address this, we propose SEF-MAP, a Subsp…

2026

SimScale: Learning to Drive via Real-World Simulation at Scale

CVPR 2026

Achieving fully autonomous driving systems requires learning rational decisions in a wide span of scenarios, including safety-critical and out-of-distribution ones. However, such cases are underrepresented in real-world corpus collected by human experts. To complement for the lack of data diversity,

Cited by 0SourcecodeScholar
2026

Spatial-Frequency Spiking Neural Network for Underwater Object Detection

AAAI 2026technical

Underwater object detection presents significant challenges due to the unique visual degradations in underwater environments, such as low contrast, poor visibility, and blurry object boundaries. While ANNs have achieved impressive detection accuracy, their high computational cost and power consumpti

Cited by 0SourcePDFScholar
2026

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

ICML 2026oral

Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D priors or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial …

Cited by 0SourceScholar
2026

SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks

ICML 2026poster

Vision-Language-Action (VLA) models have become a central paradigm for embodied intelligence. However, most existing approaches are built on large-scale Transformers, resulting in substantial inference latency and energy consumption that limit their practical deployment in low-power, real-time scena…

Cited by 0SourceScholar
2026

Towards a Foundation Model for Crowdsourced Label Aggregation

ICLR 2026poster

Inferring ground truth from noisy, crowdsourced labels is a fundamental challenge in machine learning. For decades, the dominant paradigm has relied on dataset-specific parameter estimation, a non-scalable method that fails to transfer knowledge. Recent efforts toward universal aggregation models do…

Cited by 0SourceScholar
2026

VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving

CVPR 2026

The significance of cross-view 3D geometric modeling capabilities for autonomous driving is self-evident, yet existing Vision-Language Models (VLMs) inherently lack this capability, resulting in their mediocre performance. While some promising approaches attempt to mitigate this by constructing Q&A

Cited by 0SourcecodeScholar
2026

VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness

AAAI 2026technical

The safe deployment of autonomous driving (AD) systems is fundamentally hindered by the long-tail problem, where rare yet critical driving scenarios are severely underrepresented in real-world data. Existing solutions including safety-critical scenario generation and closed-loop learning often rely

Cited by 0SourcePDFScholar
2026

What You See Is What You Reach: Towards Spatial Navigation with High-Level Human Instructions

AAAI 2026technical

Embodied navigation is a fundamental capability that enables embodied agents to effectively interact with the physical world in various complex environments. However, a significant gap remains between current embodied navigation tasks and real-world requirements, as existing methods often struggle t

Cited by 0SourcePDFScholar
2025

3D Annotation-Free Learning by Distilling 2D Open-Vocabulary Segmentation Models for Autonomous Driving

AAAI 2025technical

Point cloud data labeling is considered a time-consuming and expensive task in autonomous driving, whereas annotation-free learning training can avoid it by learning point cloud representations from unannotated data. In this paper, we propose AFOV, a novel 3D Annotation-Free framework assisted by 2D…

2025

Accelerated Over-Relaxation Heavy-Ball Method: Achieving Global Accelerated Convergence with Broad Generalization

ICLR 2025poster

The heavy-ball momentum method accelerates gradient descent with a momentum term but lacks accelerated convergence for general smooth strongly convex problems. This work introduces the Accelerated Over-Relaxation Heavy-Ball (AOR-HB) method, the first variant with provable global and accelerated conv…

Cited by 0SourcePDFScholar
2025

Adjacent-view Transformers for Supervised Surround-view Depth Estimation

IROS 2025

Depth estimation has been widely studied and serves as the fundamental step of 3D perception for robotics and autonomous driving. Though significant progress has been made in monocular depth estimation in the past decades, these attempts are mainly conducted on the KITTI benchmark with only front-vi

Cited by 5SourcecodeScholar
2025

Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing

ICML 2025poster

With the advance of diffusion models, today's video generation has achieved impressive quality. To extend the generation length and facilitate real-world applications, a majority of video diffusion models (VDMs) generate videos in an autoregressive manner, i.e., generating subsequent clips condition…

2025

Co-Fix3D: Enhancing 3D Object Detection With Collaborative Refinement

RA-L 2025

3D object detection in driving scenarios is particularly challenging due to factors such as sensor noise, occlusions, and the inherent sparsity of LiDAR point clouds, which can lead to the loss or incompleteness of key features, in turn affecting perception performance. To address these challenges,

Cited by 0SourcecodeScholar
2025

CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

CVPR 2025highlight

Interleaved image-text generation has emerged as a crucial multimodal task, aiming at creating sequences of interleaved visual and textual content given a query. Despite notable advancements in recent multimodal large language models (MLLMs), generating integrated image-text sequences that exhibit n…

2025

Cross-lingual Multimodal Sentiment Analysis for Low-Resource Languages via Language Family Disentanglement and Rethinking Transfer

ACL 2025finding

Existing multimodal sentiment analysis (MSA) methods have achieved significant success, leveraging cross-modal large-scale models (LLMs) and extensive pre-training data. However, these methods struggle to handle MSA tasks in low-resource languages. While multilingual LLMs enable cross-lingual transf…

2025

Cyclic Contrastive Knowledge Transfer for Open-Vocabulary Object Detection

ICLR 2025poster

In pursuit of detecting unstinted objects that extend beyond predefined categories, prior arts of open-vocabulary object detection (OVD) typically resort to pretrained vision-language models (VLMs) for base-to-novel category generalization. However, to mitigate the misalignment between upstream imag…

2025

Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models

NeurIPS 2025poster

Although multimodal large language models (MLLMs) exhibit remarkable reasoning capabilities on complex multimodal understanding tasks, they still suffer from the notorious 'hallucination' issue: generating outputs misaligned with obvious visual or factual evidence. Currently, training-based solution…

Cited by 0SourceScholar
2025

DisPose: Disentangling Pose Guidance for Controllable Human Image Animation

ICLR 2025poster

Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleton pose), recent works have attempted to introduce additional dense conditions (e.g., depth map) to ensure motion alignme…

2025

EANS: Reducing Energy Consumption for UAV with an Environmental Adaptive Navigation Strategy

IROS 2025

Unmanned Aerial Vehicles (Uavs) are limited by the onboard energy. Refinement of the navigation strategy directly affects both the flight velocity and the trajectory based on the adjustment of key parameters in the Uavs pipeline, thus reducing energy consumption. However, existing techniques tend to

Cited by 0SourceScholar
2025

Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning

CVPR 2025poster

Visual In-Context Learning (VICL) enables adaptively solving vision tasks by leveraging pixel demonstrations, mimicking human-like task completion through analogy. Prompt selection is critical in VICL, but current methods assume the existence of a single "ideal" prompt in a pool of candidates, which…

2025

Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

ICCV 2025poster

Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos…

2025

Interaction-Centric Knowledge Infusion and Transfer for Open Vocabulary Scene Graph Generation

NeurIPS 2025poster

Open-vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge from pre-trained large-scale models. Existing OVSGG methods always adopt a two-stage pipeline: 1) Infusing knowledge into large…

Cited by 0SourceScholar
2025

Inversion Circle Interpolation: Diffusion-based Image Augmentation for Data-scarce Classification

CVPR 2025poster

Data Augmentation (DA), i.e., synthesizing faithful and diverse samples to expand the original training set, is a prevalent and effective strategy to improve the performance of various data-scarce tasks. With the powerful image generation ability, diffusion-based DA has shown strong performance gain…

2025

IterIS: Iterative Inference-Solving Alignment for LoRA Merging

CVPR 2025poster

Low-rank adaptations (LoRA) are widely used to fine-tune large models across various domains for specific downstream tasks. While task-specific LoRAs are often available, concerns about data privacy and intellectual property can restrict access to training data, limiting the acquisition of a multi-t…

2025

KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification

ICASSP 2025accepted

Fine-tuning pre-trained vision models for specific tasks is a common practice in computer vision. However, this process becomes more expensive and resource-intensive as models grow larger. Recently, parameter-efficient fine-tuning (PEFT) methods have emerged as a popular solution to improve training…

Cited by 0SourceScholar
2025

Learning Causal Transition Matrix for Instance-dependent Label Noise

AAAI 2025technical

Noisy labels are both inevitable and problematic in machine learning methods, as they negatively impact models' generalization ability by causing overfitting. In the context of learning with noise, the transition matrix plays a crucial role in the design of statistically consistent algorithms. Howev…

Cited by 0SourcePDFScholar
2025

Learning Perceptive Humanoid Locomotion over Challenging Terrain

IROS 2025

Humanoid robots are engineered to navigate terrains akin to those encountered by humans, which necessitates human-like locomotion and perceptual abilities. Currently, the most reliable controllers for humanoid motion rely exclusively on proprioception, a reliance that becomes both dangerous and unre

Cited by 23SourceScholar
2025

Lightstereo: Channel Boost is All You Need for Efficient 2D Cost Aggregation

ICRA 2025

We present LightStereo, a cutting-edge stereomatching network crafted to accelerate the matching process. Departing from conventional methodologies that rely on aggregating computationally intensive 4D costs, LightStereo adopts the 3D cost volume as a lightweight alternative. While similar approache

Cited by 37SourcecodeScholar
2025

Modeling Uncertainty in Composed Image Retrieval via Probabilistic Embeddings

ACL 2025long

Composed Image Retrieval (CIR) enables users to search for images using multimodal queries that combine text and reference images. While metric learning methods have shown promise, they rely on deterministic point embeddings that fail to capture the inherent uncertainty in the input data, in which u…

2025

Multi-Resolution Decomposable Diffusion Model for Non-Stationary Time Series Anomaly Detection

ICLR 2025poster

Recently, generative models have shown considerable promise in unsupervised time series anomaly detection. Nonetheless, the task of effectively capturing complex temporal patterns and minimizing false alarms becomes increasingly challenging when dealing with non-stationary time series, characterized…

Cited by 0SourcePDFScholar
2025

Nautilus: Locality-aware Autoencoder for Scalable Mesh Generation

ICCV 2025poster

Triangle meshes are fundamental to 3D applications. Current automatic mesh generation methods typically rely on intermediate representations that lack the continuous surface quality inherent to meshes. Converting these representations into meshes produces dense, suboptimal outputs. Although recent a…

Cited by 0SourcePDFScholar
2025

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

EMNLP 2025

Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences. Typically, a reward model is trained or supplied to act as a proxy for humans in evaluating generated responses during the reinforcement training phase. Howe

2025

ReSim: Reliable World Simulation for Autonomous Driving

NeurIPS 2025spotlight

How can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusively on real-world driving data composed mainly of safe expert trajectories, struggle to follow hazardous or non-expert behaviors, which are rare in such d…

Cited by 0SourceScholar
2025

SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

NeurIPS 2025poster

Accurate spatial reasoning in outdoor environments—covering geometry, object pose, and inter-object relationships—is fundamental to downstream tasks such as mapping, motion forecasting, and high-level planning in autonomous driving. We introduce SURDS, a large-scale benchmark designed to systematica…

Cited by 0SourcecodeScholar
2025

SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment

CVPR 2025highlight

Integrating large language models (LLMs) into autonomous driving has attracted significant attention with the hope of improving generalization and explainability. However, existing methods often focus on either driving or vision-language understanding but achieving both high driving performance and…

Cited by 1SourcePDFScholar
2025

Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

CVPR 2025poster

Diffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and corresponding text prompts. To tackle this issue, reinforcement learning (RL) has been considered for diffusion model fin…

2025

WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation

IROS 2025

Object Goal Navigation-requiring an agent to locate a specific object in an unseen environment-remains a core challenge in embodied AI. Although recent progress in Vision-Language Model (VLM)-based agents has demonstrated promising perception and decision-making abilities through prompting, none has

Cited by 38SourcecodeScholar
2024

$\text{Di}^2\text{Pose}$: Discrete Diffusion Model for Occluded 3D Human Pose Estimation

NeurIPS 2024poster

Diffusion models have demonstrated their effectiveness in addressing the inherent uncertainty and indeterminacy in monocular 3D human pose estimation (HPE). Despite their strengths, the need for large search spaces and the corresponding demand for substantial training data make these models prone t…

Cited by 0SourcePDFScholar
2024

AdvDiffuser: Generating Adversarial Safety-Critical Driving Scenarios via Guided Diffusion

IROS 2024poster

Safety-critical scenarios are infrequent in natural driving environments but hold significant importance for the training and testing of autonomous driving systems. The prevailing approach involves generating safety-critical scenarios automatically in simulation by introducing adversarial adjustment…

Cited by 6SourceScholar
2024

An Efficient and Effective Transformer Decoder-Based Framework for Multi-Task Visual Grounding

ECCV 2024poster

"Most advanced visual grounding methods rely on Transformers for visual-linguistic feature fusion. However, these Transformer-based approaches encounter a significant drawback: the computational costs escalate quadratically due to the self-attention mechanism in the Transformer Encoder, particularly…

2024

Beyond Grounding: Extracting Fine-Grained Event Hierarchies across Modalities

AAAI 2024technical

Events describe happenings in our world that are of importance. Naturally, understanding events mentioned in multimedia content and how they are related forms an important way of comprehending our world. Existing literature can infer if events across textual and visual (video) domains are identical…

2024

BotanicGarden: A High-Quality Dataset for Robot Navigation in Unstructured Natural Environments

RA-L 2024

The rapid developments of mobile robotics and autonomous navigation over the years are largely empowered by public datasets for testing and upgrading, such as sensor odometry and SLAM tasks. Impressive demos and benchmark scores have arisen, which may suggest the maturity of existing navigation tech

Cited by 60SourcecodeScholar
2024

ClothPPO: A Proximal Policy Optimization Enhancing Framework for Robotic Cloth Manipulation with Observation-Aligned Action Spaces

IJCAI 2024poster

Vision-based robotic cloth unfolding has made great progress recently. However, prior works predominantly rely on value learning and have not fully explored policy-based techniques. Recently, the success of reinforcement learning on the large language model has shown that the policy gradient algorit…

2024

D-PBS: Dueling Priority-Based Search for Multiple Nonholonomic Robots Motion Planning in Congested Environments

RA-L 2024

This letter focuses on the multiple nonholonomic robots motion planning (MRMP) problem in congested and complex environments, where the complexity escalates dramatically with the increase in the number of robots, frequently leading to deadlocks. We present the <italic xmlns:mml="http://www.w3.org/19

Cited by 11SourceScholar
2024

Distributionally Generative Augmentation for Fair Facial Attribute Classification

CVPR 2024poster

Facial Attribute Classification (FAC) holds substantial promise in widespread applications. However FAC models trained by traditional methodologies can be unfair by exhibiting accuracy inconsistencies across varied data subpopulations. This unfairness is largely attributed to bias in data where some…

2024

Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving

ICRA 2024poster

Large Language Models (LLMs) have shown promise in the autonomous driving sector, particularly in generalization and interpretability. We introduce a unique objectlevel multimodal LLM architecture that merges vectorized numeric modalities with a pre-trained LLM to improve context understanding in dr…

Cited by 229SourcecodeScholar
2024

Generative End-to-End Autonomous Driving

ECCV 2024poster

"Directly producing planning results from raw sensors has been a long-desired solution for autonomous driving and has attracted increasing attention recently. Most existing end-to-end autonomous driving methods factorize this problem into perception, motion prediction, and planning. However, we argu…

2024

LLMs Can Evolve Continually on Modality for $\mathbb{X}$-Modal Reasoning

NeurIPS 2024poster

Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily on extensive modal-specific pretraining and joint-modal tuning, leading to significant computational burdens when expand…

2024

LingoQA: Video Question Answering for Autonomous Driving

ECCV 2024poster

"We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capab…

2024

MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding

EMNLP 2024main

Improving user experience and providing personalized search results in E-commerce platforms heavily rely on understanding purchase intention. However, existing methods for acquiring large-scale intentions bank on distilling large language models with human annotation for verification. Such an approa…

2024

Mrtnet: Multi-Resolution Temporal Network for Video Sentence Grounding

ICASSP 2024accepted

Video sentence grounding locates a specific moment in a video based on a text query. Existing methods focus on single temporal resolution, ignoring multi-scale temporal consistency. We introduce MRTNet, a multi-resolution grounding network with four key components: a feature encoder, a Multi-Resolut…

Cited by 0SourceScholar
2024

Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning

EMNLP 2024main

Reinforcement learning from human feedback (RLHF) and AI-generated feedback (RLAIF) have become prominent techniques that significantly enhance the functionality of pre-trained language models (LMs). These methods harness feedback, sourced either from humans or AI, as direct rewards or to shape rewa…

2024

RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter

ACL 2024findings

Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries. To date, most of the state-of-the-art TVR methods learn image-to-video transfer learning based on the large-scale pre-trained vision-language models (e.g., CLIP). However, fully fine-tuning these pre-train…

Cited by 17SourcePDFScholar
2024

SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos

ICLR 2024poster

We study the problem of procedure planning in instructional videos, which aims to make a goal-oriented sequence of action steps given partial visual state observations. The motivation of this problem is to learn a structured and plannable state and action space. Recent works succeeded in sequence mo…

Cited by 15SourcePDFScholar
2024

SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning

ECCV 2024poster

"Parameter-efficient transfer learning (PETL) has emerged as a flourishing research field for adapting large pre-trained models to downstream tasks, greatly reducing trainable parameters while grappling with memory challenges during fine-tuning. To address it, memory-efficient series (METL) avoid ba…

2024

SoundCount: Sound Counting from Raw Audio with Dyadic Decomposition Neural Network

AAAI 2024technical

In this paper, we study an underexplored, yet important and challenging problem: counting the number of distinct sounds in raw audio characterized by a high degree of polyphonicity. We do so by systematically proposing a novel end-to-end trainable neural network~(which we call DyDecNet, consisting o…

Cited by 2SourcePDFScholar
2024

Towards efficient deep spiking neural networks construction with spiking activity based pruning

ICML 2024poster

The emergence of deep and large-scale spiking neural networks (SNNs) exhibiting high performance across diverse complex datasets has led to a need for compressing network models due to the presence of a significant number of redundant structural units, aiming to more effectively leverage their low-p…

Cited by 9SourcePDFScholar
2024

Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion

ICASSP 2024accepted

We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM). Experiments on the Switchboard human-human conversation dataset demonstrate that our approach consistently outperforms…

Cited by 26SourceScholar
2024

UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory

CVPR 2024poster

Parameter-efficient transfer learning (PETL) i.e. fine-tuning a small portion of parameters is an effective strategy for adapting pre-trained models to downstream domains. To further reduce the memory demand recent PETL works focus on the more valuable memory-efficient characteristic. In this paper…

2024

View-Consistent 3D Editing with Gaussian Splatting

ECCV 2024poster

"The advent of 3D Gaussian Splatting (3DGS) has revolutionized 3D editing, offering efficient, high-fidelity rendering and enabling precise local manipulations. Currently, diffusion-based 2D editing models are harnessed to modify multi-view rendered images, which then guide the editing of 3DGS model…

Cited by 25SourcePDFScholar
2023

Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language Models

EMNLP 2023long findings

The age of social media is rife with memes. Understanding and detecting harmful memes pose a significant challenge due to their implicit meaning that is not explicitly conveyed through the surface text and image. However, existing harmful meme detection approaches only recognize superficial harm-ind…

Cited by 0SourcecodeScholar
2023

Compositional Feature Augmentation for Unbiased Scene Graph Generation

ICCV 2023poster

Scene Graph Generation (SGG) aims to detect all the visual relation triplets <sub, pred, obj> in a given image. With the emergence of various advanced techniques for better utilizing both the intrinsic and extrinsic information in each relation triplet, SGG has achieved great progress over the recen…

Cited by 44PDFcodeScholar
2023

Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection

ICLR 2023poster

Prompt tuning with large-scale pretrained vision-language models empowers open-vocabulary prediction trained on limited base categories, e.g., object classification and detection. In this paper, we propose compositional prompt tuning with motion cues: an extended prompt tuning paradigm for compositi…

2023

Cross-Utterance ASR Rescoring with Graph-Based Label Propagation

ICASSP 2023accepted

We propose a novel approach for ASR N-best hypothesis rescoring with graph-based label propagation by leveraging cross-utterance acoustic similarity. In contrast to conventional neural language model (LM) based ASR rescoring/reranking models, our approach focuses on acoustic information and conducts…

Cited by 0SourceScholar
2023

Dataset Bias Mitigation in Multiple-Choice Visual Question Answering and Beyond

EMNLP 2023long findings

Vision-language (VL) understanding tasks evaluate models' comprehension of complex visual scenes through multiple-choice questions. However, we have identified two dataset biases that models can exploit as shortcuts to resolve various VL tasks correctly without proper understanding. The first type o…

Cited by 0SourceScholar
2023

Discrepancy-Guided Reconstruction Learning for Image Forgery Detection

IJCAI 2023poster

In this paper, we propose a novel image forgery detection paradigm for boosting the model learning capacity on both forgery-sensitive and genuine compact visual patterns. Compared to the existing methods that only focus on the discrepant-specific patterns (\eg, noises, textures, and frequencies), ou…

Cited by 20SourcePDFScholar
2023

Dual-Stage Graph Convolution Network With Graph Learning For Traffic Prediction

ICASSP 2023accepted

Robust and accurate traffic forecasting is a key issue in intelligent transportation systems. Existing studies usually employ pre-defined spatial graph or learned fixed adjacency graph and design models to capture spatial and temporal features. However, pre-defined or fixed graph can not accurately…

Cited by 0SourceScholar
2023

Enhanced Chart Understanding via Visual Language Pre-training on Plot Table Pairs

ACL 2023findings

Building cross-model intelligence that can understand charts and communicate the salient information hidden behind them is an appealing challenge in the vision and language (V+L) community. The capability to uncover the underlined table data of chart figures is a critical key to automatic chart unde…

Cited by 0SourcePDFScholar
2023

Fairness-aware Contrastive Learning with Partially Annotated Sensitive Attributes

ICLR 2023poster

Learning high-quality representation is important and essential for visual recognition. Unfortunately, traditional representation learning suffers from fairness issues since the model may learn information of sensitive attributes. Recently, a series of studies have been proposed to improve fairness…

Cited by 35SourcePDFScholar
2023

IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models

EMNLP 2023long findings

The field of vision-and-language (VL) understanding has made unprecedented progress with end-to-end large pre-trained VL models (VLMs). However, they still fall short in zero-shot reasoning tasks that require multi-step inferencing. To achieve this goal, previous works resort to a divide-and-conquer…

Cited by 0SourcecodeScholar
2023

Iterative Proposal Refinement for Weakly-Supervised Video Grounding

CVPR 2023poster

Weakly-Supervised Video Grounding (WSVG) aims to localize events of interest in untrimmed videos with only video-level annotations. To date, most of the state-of-the-art WSVG methods follow a two-stage pipeline, i.e., firstly generating potential temporal proposals and then grounding with these prop…

2023

Progressive Deep Multi-View Comprehensive Representation Learning

AAAI 2023technical

Multi-view Comprehensive Representation Learning (MCRL) aims to synthesize information from multiple views to learn comprehensive representations of data items. Prevalent deep MCRL methods typically concatenate synergistic view-specific representations or average aligned view-specific representation…

2023

TempCLR: Temporal Alignment Representation with Contrastive Learning

ICLR 2023poster

Video representation learning has been successful in video-text pre-training for zero-shot transfer, where each sentence is trained to be close to the paired video clips in a common feature space. For long videos, given a paragraph of description where the sentences describe different segments of th…

2023

Two Heads are Better Than One: A Simple Exploration Framework for Efficient Multi-Agent Reinforcement Learning

NeurIPS 2023poster

Exploration strategy plays an important role in reinforcement learning, especially in sparse-reward tasks. In cooperative multi-agent reinforcement learning~(MARL), designing a suitable exploration strategy is much more challenging due to the large state space and the complex interaction among agent…

Cited by 3SourcePDFScholar
2023

Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language Models

NeurIPS 2023poster

Pretrained vision-language models, such as CLIP, have demonstrated strong generalization capabilities, making them promising tools in the realm of zero-shot visual recognition. Visual relation detection (VRD) is a typical task that identifies relationship (or interaction) types between object pairs…

2022

Broadband Sound Source Localisation via Non-Synchronous Measurements for Service Robots: A Tensor Completion Approach

RA-L 2022

Constraint by the physical geometry, the lower and upper frequency bound and the scale of the scanning area of a microphone array are limited. Owing to its movable feature, for the service robots, achieving a wider working frequency range with a global view requires a virtually larger and denser arr

Cited by 16SourceScholar
2022

Classification-Then-Grounding: Reformulating Video Scene Graphs As Temporal Bipartite Graphs

CVPR 2022poster

Today's VidSGG models are all proposal-based methods, i.e., they first generate numerous paired subject-object snippets as proposals, and then conduct predicate classification for each proposal. In this paper, we argue that this prevalent proposal-based framework has three inherent drawbacks: 1) The…

Cited by 42PDFcodeScholar
2022

CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention

ICLR 2022poster

Transformers have made great progress in dealing with computer vision tasks. However, existing vision transformers have not yet possessed the ability of building the interactions among features of different scales, which is perceptually important to visual inputs. The reasons are two-fold: (1) Input…

2022

Deconfounded Value Decomposition for Multi-Agent Reinforcement Learning

ICML 2022spotlight

Value decomposition (VD) methods have been widely used in cooperative multi-agent reinforcement learning (MARL), where credit assignment plays an important role in guiding the agents’ decentralized execution. In this paper, we investigate VD from a novel perspective of causal inference. We first sho…

Cited by 23SourcePDFScholar
2022

Few-Shot Object Detection With Fully Cross-Transformer

CVPR 2022oral

Few-shot object detection (FSOD), with the aim to detect novel objects using very few training examples, has recently attracted great research interest in the community. Metric-learning based methods have been demonstrated to be effective for this task using a two-branch based siamese network, and c…

Cited by 193PDFcodeScholar
2022

OpenFEAT: Improving Speaker Identification by Open-Set Few-Shot Embedding Adaptation with Transformer

ICASSP 2022accepted

Household speaker identification with few enrollment utterances is an important yet challenging problem, especially when household members share similar voice characteristics and room acoustics. A common embedding space learned from a large number of speakers is not universally applicable for the op…

Cited by 0SourceScholar
2022

Rethinking Multi-Modal Alignment in Multi-Choice VideoQA from Feature and Sample Perspectives

EMNLP 2022main

Reasoning about causal and temporal event relations in videos is a new destination of Video Question Answering (VideoQA). The major stumbling block to achieve this purpose is the semantic gap between language and video since they are at different levels of abstraction. Existing efforts mainly focus…

Cited by 6SourcePDFScholar
2022

Rethinking the Two-Stage Framework for Grounded Situation Recognition

AAAI 2022technical

Grounded Situation Recognition (GSR), i.e., recognizing the salient activity (or verb) category in an image (e.g.,buying) and detecting all corresponding semantic roles (e.g.,agent and goods), is an essential step towards “human-like” event understanding. Since each verb is associated with a specifi…

2022

The Devil Is in the Labels: Noisy Label Correction for Robust Scene Graph Generation

CVPR 2022oral

Unbiased SGG has achieved significant progress over recent years. However, almost all existing SGG models have overlooked the ground-truth annotation qualities of prevailing SGG datasets, i.e., they always assume: 1) all the manually annotated positive samples are equally correct; 2) all the un-anno…

Cited by 119PDFcodeScholar
2022

ULSM: Underground Localization and Semantic Mapping with Salient Region Loop Closure under Perceptually-Degraded Environment

IROS 2022poster

Simultaneous Localization and Mapping (SLAM) has greatly assisted in exploring perceptually-degraded underground environments, such as human-made tunnels, mine tunnels, and caves. However, the recurring sensor failures and spurious loop closures in these scenes bring significant challenges to applyi…

Cited by 3SourceScholar
2022

Weakly-Supervised Temporal Article Grounding

EMNLP 2022main

Given a long untrimmed video and natural language queries, video grounding (VG) aims to temporally localize the semantically-aligned video segments. Almost all existing VG work holds two simple but unrealistic assumptions: 1) All query sentences can be grounded in the corresponding video. 2) All que…

2021

Accelerate CNNs from Three Dimensions: A Comprehensive Pruning Framework

ICML 2021spotlight

Most neural network pruning methods, such as filter-level and layer-level prunings, prune the network model along one dimension (depth, width, or resolution) solely to meet a computational budget. However, such a pruning policy often leads to excessive reduction of that dimension, thus inducing a hu…

Cited by 77SourcePDFScholar
2021

Boundary Proposal Network for Two-stage Natural Language Video Localization

AAAI 2021technical

We aim to address the problem of Natural Language Video Localization (NLVL) — localizing the video segment corresponding to a natural language description in a long and untrimmed video. State-of-the-art NLVL methods are almost in one-stage fashion, which can be typically grouped into two categories:…

Cited by 187SourcePDFScholar
2021

FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention

NeurIPS 2021poster

We propose FMMformers, a class of efficient and flexible transformers inspired by the celebrated fast multipole method (FMM) for accelerating interacting particle simulation. FMM decomposes particle-particle interaction into near-field and far-field components and then performs direct and coarse-gra…

Cited by 33SourcePDFScholar
2021

Human-Like Controllable Image Captioning With Verb-Specific Semantic Roles

CVPR 2021poster

Controllable Image Captioning (CIC) -- generating image descriptions following designated control signals -- has received unprecedented attention over the last few years. To emulate the human ability in controlling caption generation, current CIC studies focus exclusively on control signals concerni…

Cited by 85PDFcodeScholar
2021

Natural Language Video Localization with Learnable Moment Proposals

EMNLP 2021main

Given an untrimmed video and a natural language query, Natural Language Video Localization (NLVL) aims to identify the video moment described by query. To address this task, existing methods can be roughly grouped into two groups: 1) propose-and-rank models first define a set of hand-designed moment…

2021

On Pursuit of Designing Multi-modal Transformer for Video Grounding

EMNLP 2021main

Video grounding aims to localize the temporal segment corresponding to a sentence query from an untrimmed video. Almost all existing video grounding methods fall into two frameworks: 1) Top-down model: It predefines a set of segment candidates and then conducts segment classification and regression.…

Cited by 91SourcePDFScholar
2021

Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding

AAAI 2021technical

The prevailing framework for solving referring expression grounding is based on a two-stage process: 1) detecting proposals with an object detector and 2) grounding the referent to one of the proposals. Existing two-stage solutions mostly focus on the grounding step, which aims to align the expressi…

2021

SimNet: Learning Reactive Self-driving Simulations from Real-world Observations

ICRA 2021poster

In this work we present a simple end-to-end trainable machine learning system capable of realistically simulating driving experiences. This can be used for verification of self-driving system performance without relying on expensive and time-consuming road testing. In particular, we frame the simula…

Cited by 115SourceScholar
2021

What data do we need for training an AV motion planner?

ICRA 2021poster

We investigate what grade of sensor data is required for training an imitation-learning-based AV planner on human expert demonstration. Machine-learned planners [1] are very hungry for training data, which is usually collected using vehicles equipped with the same sensors used for autonomous operati…

Cited by 15SourceScholar
2020

Counterfactual Samples Synthesizing for Robust Visual Question Answering

CVPR 2020poster

Despite Visual Question Answering (VQA) has realized impressive progress over the last few years, today's VQA models tend to capture superficial linguistic correlations in the train set and fail to generalize to the test set with different QA distributions. To reduce the language biases, several rec…

Cited by 401PDFcodeScholar
2020

Cross-View Tracking for Multi-Human 3D Pose Estimation at Over 100 FPS

CVPR 2020poster

Estimating 3D poses of multiple humans in real-time is a classic but still challenging task in computer vision. Its major difficulty lies in the ambiguity in cross-view association of 2D poses and the huge state space when there are multiple people in multiple views. In this paper, we present a nove…

Cited by 115PDFcodeScholar
2020

GOSMatch: Graph-of-Semantics Matching for Detecting Loop Closures in 3D LiDAR data

IROS 2020poster

Detecting loop closures in 3D Light Detection and Ranging (LiDAR) data is a challenging task since point-level methods always suffer from instability. This paper presents a semantic-level approach named GOSMatch to perform reliable place recognition. Our method leverages novel descriptors, which are…

Cited by 74SourceScholar
2020

Imitative Reinforcement Learning Fusing Vision and Pure Pursuit for Self-driving

ICRA 2020poster

Autonomous urban driving navigation is still an open problem and has ample room for improvement in unknown complex environments and terrible weather conditions. In this paper, we propose a two-stage framework, called IPP-RL, to handle these problems. IPP means an Imitation learning method fusing vis…

Cited by 14SourceScholar
2020

One Thousand and One Hours: Self-driving Motion Prediction Dataset

CoRL 2020

Motivated by the impact of large-scale datasets on ML systems we present the largest self-driving dataset for motion prediction to date, containing over 1,000 hours of data. This was collected by a fleet of 20 autonomous vehicles along a fixed route in Palo Alto, California, over a four-month period

2020

Trading Personalization for Accuracy: Data Debugging in Collaborative Filtering

NeurIPS 2020poster

Collaborative filtering has been widely used in recommender systems. Existing work has primarily focused on improving the prediction accuracy mainly via either building refined models or incorporating additional side information, yet has largely ignored the inherent distribution of the input rating…

2019

Counterfactual Critic Multi-Agent Training for Scene Graph Generation

ICCV 2019oral

Scene graphs --- objects as nodes and visual relationships as edges --- describe the whereabouts and interactions of objects in an image for comprehensive scene understanding. To generate coherent scene graphs, almost all existing methods exploit the fruitful visual context by modeling message passi…

Cited by 200PDFScholar
2019

Real-Time Vehicle Detection from Short-range Aerial Image with Compressed MobileNet

ICRA 2019poster

Vehicle detection from short-range aerial image faces challenges including vehicle blocking, irrelevant object interference, motion blurring, color variation etc., leading to the difficulty to achieve high detection accuracy and real-time detection speed. In this paper, benefiting from the recent de…

Cited by 19SourceScholar
2018

Real-Time Learning of Efficient Lift Generation on a Dynamically Scaled Flapping Wing Using Policy Search

ICRA 2018poster

In this work, we present a successful application of a policy search algorithm to a real-time robotic learning problem, where the goal is to maximize the efficiency of lift generation on a dynamically scaled flapping robotic wing. The robotic wing has two degrees-of-freedom, i.e., stroke and pitch,…

Cited by 10SourceScholar
2018

Zero-Shot Visual Recognition Using Semantics-Preserving Adversarial Embedding Networks

CVPR 2018poster

We propose a novel framework called Semantics-Preserving Adversarial Embedding Network (SP-AEN) for zero-shot visual recognition (ZSL), where test images and their classes are both unseen during training. SP-AEN aims to tackle the inherent problem — semantic loss — in the prevailing family of embedd…

Cited by 370SourcePDFScholar
2017

RGB-T SLAM: A flexible SLAM framework by combining appearance and thermal information

ICRA 2017poster

Visual SLAM in low illumination scenes remains a considerably challenging task since the available amount of appearance information frequently stays insufficient. To tackle with this problem, we propose a novel SLAM framework by using both appearance information and thermal information, which posses…

Cited by 77SourceScholar
2017

SCA-CNN: Spatial and Channel-Wise Attention in Convolutional Networks for Image Captioning

CVPR 2017poster

Visual attention has been successfully applied in structural prediction tasks such as visual captioning and question answering. Existing visual attention models are generally spatial, i.e., the attention is modeled as spatial probabilities that re-weight the last conv-layer feature map of a CNN enco…

Cited by 2297PDFcodeScholar