← Search

Jingwen Wang

13 accepted papers

2025

Dynamic planning and assembly for constructing mortar-joint multi-leaf stone masonry walls with a robotic arm

IROS 2025

The construction industry faces growing challenges in sustainability, labor shortages, and efficiency. Stone masonry, known for its durability and low environmental impact, has declined in modern construction due to high labor costs and slow building processes. This work presents a robotic system fo

Cited by 1SourceScholar
2025

ERCI: An Explainable Experience Replay Approach with Causal Inference for Deep Reinforcement Learning

AAAI 2025technical

Deep reinforcement learning (DRL) has gained significant attention in autonomous systems, yet its black-box nature and lack of explainability hinder user trust in safety-critical domains such as autonomous driving. Existing experience replay approaches enhance sample efficiency but often fail to cap…

2025

MRMT-PR: A Multi-Scale Reverse-View Mamba-Transformer for LiDAR Place Recognition

IROS 2025

Place recognition is a fundamental technology of high relevance for autonomous robot navigation. Existing methods encounter significant challenges arising from scene variations (e.g., illumination changes, dynamic objects), view-point shifts, and difficulties in data fusion and alignment. These fact

Cited by 1SourceScholar
2025

Multi-Task Invariant Representation Imitation Learning for Autonomous Driving

ICRA 2025

Imitation learning is a promising approach to acquiring autonomous driving policies by mimicking human driver behaviors. However, a major drawback of existing driving policies derived from imitation learning is their proneness to capturing spurious correlations, owing to the lack of an explicit caus

Cited by 0SourceScholar
2024

MorpheuS: Neural Dynamic 360deg Surface Reconstruction from Monocular RGB-D Video

CVPR 2024poster

Neural rendering has demonstrated remarkable success in dynamic scene reconstruction. Thanks to the expressiveness of neural representations prior works can accurately capture the motion and achieve high-fidelity reconstruction of the target object. Despite this real-world video scenarios often feat…

Cited by 2SourcePDFScholar
2024

Self-supervised Preference Optimization: Enhance Your Language Model with Preference Degree Awareness

EMNLP 2024finding

Recently, there has been significant interest in replacing the reward model in Reinforcement Learning with Human Feedback (RLHF) methods for Large Language Models (LLMs), such as Direct Preference Optimization (DPO) and its variants. These approaches commonly use a binary cross-entropy mechanism on…

2023

Co-SLAM: Joint Coordinate and Sparse Parametric Encodings for Neural Real-Time SLAM

CVPR 2023poster

We present Co-SLAM, a neural RGB-D SLAM system based on a hybrid representation, that performs robust camera tracking and high-fidelity surface reconstruction in real time. Co-SLAM represents the scene as a multi-resolution hash-grid to exploit its high convergence speed and ability to represent hig…

2023

FCC: Feature Clusters Compression for Long-Tailed Visual Recognition

CVPR 2023poster

Deep Neural Networks (DNNs) are rather restrictive in long-tailed data, since they commonly exhibit an under-representation for minority classes. Various remedies have been proposed to tackle this problem from different perspectives, but they ignore the impact of the density of Backbone Features (BF…

2023

SeMLaPS: Real-Time Semantic Mapping With Latent Prior Networks and Quasi-Planar Segmentation

RA-L 2023

The availability of real-time semantics greatly improves the core geometric functionality of SLAM systems, enabling numerous robotic and AR/VR applications. We present a new methodology for real-time semantic mapping from RGB-D sequences that combines a 2D neural network and a 3D network based on a

Cited by 9SourcecodeScholar
2021

Integrating Semantics and Neighborhood Information with Graph-Driven Generative Models for Document Retrieval

ACL 2021long

With the need of fast retrieval speed and small memory footprint, document hashing has been playing a crucial role in large-scale information retrieval. To generate high-quality hashing code, both semantics and neighborhood information are crucial. However, most existing methods leverage only one of…

2019

Controllable Video Captioning With POS Sequence Guidance Based on Gated Fusion Network

ICCV 2019poster

In this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos. We construct a novel gated fusion network, with one particularly designed cross-gating (CG) block, to effectively encode and fus…

Cited by 232PDFcodeScholar
2019

Semantic Conditioned Dynamic Modulation for Temporal Sentence Grounding in Videos

NeurIPS 2019poster

Temporal sentence grounding in videos aims to detect and localize one target video segment, which semantically corresponds to a given sentence. Existing methods mainly tackle this task via matching and aligning semantics between a sentence and candidate video segments, while neglect the fact that th…

2018

Bidirectional Attentive Fusion With Context Gating for Dense Video Captioning

CVPR 2018poster

Dense video captioning is a newly emerging task that aims at both localizing and describing all events in a video. We identify and tackle two challenges on this task, namely, (1) how to utilize both past and future contexts for accurate event proposal predictions, and (2) how to construct informativ…

Cited by 272SourcePDFScholar