← Search

Wen Jiang

17 accepted papers

2026

Active Next-Best-View Optimization for Risk-Averse Path Planning

ICRA 2026poster

Safe navigation in uncertain environments requires planning methods that integrate risk aversion with active perception. In this work, we present a unified frame- work that refines a coarse reference path by construct- ing tail-sensitive risk maps from Average Value-at-Risk statistics on an online-u…

2026

CETUS: Causal Event-Driven Temporal Modeling with Unified Variable-Rate Scheduling

ICRA 2026poster

Event cameras capture asynchronous pixel-level brightness changes with microsecond temporal resolution, offering unique advantages for high-speed vision tasks. Existing methods often convert event streams into intermediate representations such as frames, voxel grids, or point clouds, which inevitabl…

2026

EventFlash: Towards Efficient MLLMs for Event-Based Vision

ICLR 2026poster

Event-based multimodal large language models (MLLMs) enable robust perception in high-speed and low-light scenarios, addressing key limitations of frame-based MLLMs. However, current event-based MLLMs often rely on dense image-like processing paradigms, overlooking the spatiotemporal sparsity of eve…

Cited by 0SourcecodeScholar
2026

Walking Further: Semantic-Aware Multimodal Gait Recognition Under Long-Range Conditions

AAAI 2026technical

Gait recognition is an emerging biometric technology that enables non-intrusive and hard-to-spoof human identification. However, most existing methods are confined to short-range, unimodal settings and fail to generalize to long-range and cross-distance scenarios under real-world conditions. To addr

Cited by 0SourcePDFScholar
2025

Coevolutionary Emergent Systems Optimization with Applications to Ultra-High-Dimensional Metasurface Design : OAM Wave Manipulation

UAI 2025

Optimization problems in electromagnetic wave manipulation and metasurface design are becoming increasingly high-dimensional, often involving thousands of variables that need precise control. Traditional optimization algorithms face significant challenges in maintaining both accuracy and computation

Cited by 0SourcePDFScholar
2025

Multimodal LLM Guided Exploration and Active Mapping using Fisher Information

ICCV 2025poster

We present an active mapping system that could plan for long-horizon exploration goals and short-term actions with a 3D Gaussian Splatting (3DGS) representation. Existing methods either did not take advantage of recent developments in multimodal Large Language Models (LLM) or did not consider challe…

Cited by 0SourcePDFScholar
2025

Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting

ICRA 2025

We propose a framework for active next best view and touch selection for robotic manipulators using 3D Gaussian Splatting (3DGS). 3DGS is emerging as a useful explicit 3D scene representation for robotics, as it has the ability to represent scenes in a both photorealistic and geometrically accurate

Cited by 11SourcecodeScholar
2024

Contrastive Token Learning with Similarity Decay for Repetition Suppression in Machine Translation

EMNLP 2024finding

For crosslingual conversation and trade, Neural Machine Translation (NMT) is pivotal yet faces persistent challenges with monotony and repetition in generated content. Traditional solutions that rely on penalizing text redundancy or token reoccurrence have shown limited efficacy, particularly for le…

Cited by 0SourcePDFScholar
2024

MoDULA: Mixture of Domain-Specific and Universal LoRA for Multi-Task Learning

EMNLP 2024main

The growing demand for larger-scale models in the development of Large Language Models (LLMs) poses challenges for efficient training within limited computational resources. Traditional fine-tuning methods often exhibit instability in multi-task learning and rely heavily on extensive training resour…

Cited by 1SourcePDFScholar
2024

Mutual Information Assisted Graph Convolution Network for Cold-Start Recommendation

ICASSP 2024accepted

To solve the cold-start issue that cold items have no historical interactions to obtain collaborative feature as their representation, existing methods often represent them totally based on content feature obtained from inherent content (i.e., image, video and attributes). However, these methods wil…

Cited by 0SourceScholar
2021

Recognition of Dynamic Hand Gesture Based on Mm-Wave Fmcw Radar Micro-Doppler Signatures

ICASSP 2021accepted

Radar-based sensors provide an attractive choice for hand gesture recognition (HGR). The very challenging problems in radar-based HGR are radar echo data preprocessing and recognition accuracy. In this paper, we propose a convolutional neural network (CNN) for dynamic HGR based on a millimeter-wave…

Cited by 0SourceScholar
2020

Coherent Reconstruction of Multiple Humans From a Single Image

CVPR 2020poster

In this work, we address the problem of multi-person 3D pose estimation from a single image. A typical regression approach in the top-down setting of this problem would first detect all humans and then reconstruct each one of them independently. However, this type of prediction suffers from incohere…

Cited by 208PDFcodeScholar
2019

Fast and Robust Multi-Person 3D Pose Estimation From Multiple Views

CVPR 2019poster

This paper addresses the problem of 3D pose estimation for multiple people in a few calibrated camera views. The main challenge of this problem is to find the cross-view correspondences among noisy and incomplete 2D pose predictions. Most previous methods address this challenge by directly reasoning…

Cited by 266PDFScholar