← Search

Min Wang

46 accepted papers

2026

D-Nav: End-to-End Dynamic UAV Navigation with Dual-Resolution Motion Awareness

RSS 2026poster

Autonomous navigation in dense, dynamic clutter remains a fundamental challenge for Unmanned Aerial Vehicles (UAVs) due to the heterogeneous obstacle scales and complex motion patterns. Existing methods often rely on fragile explicit tracking or noise-sensitive implicit flow estimation, both of whic…

Cited by 0SourceScholar
2026

Offline Meta-Reinforcement Learning with Flow-Based Task Inference and Adaptive Correction of Feature Overgeneralization

AAAI 2026technical

Offline meta-reinforcement learning (OMRL) combines the strengths of learning from diverse datasets in offline RL with the adaptability to new tasks of meta-RL, promising safe and efficient knowledge acquisition by RL agents. However, OMRL still suffers extrapolation errors due to out-of-distributio

Cited by 0SourcePDFScholar
2026

Primary-Fine Decoupling for Action Generation in Robotic Imitation

ICLR 2026poster

Multi-modal distribution in robotic manipulation action sequences poses critical challenges for imitation learning. To this end, existing approaches often model the action space as either a discrete set of tokens or a continuous, latent-variable distribution. However, both approaches present trade-…

Cited by 0SourceScholar
2026

Reflect-then-Correct: Rebalancing Task Optimization for Generalizable Meta-Reinforcement Learning via Distributional Value Error Reduction

ICML 2026poster

Meta-Reinforcement Learning (Meta-RL) faces significant challenges in non-parametric settings, where vastly different return scales across diverse tasks cause severe gradient interference. Existing categorical solutions attempt to normalize these scales but often fail due to rigid discretization and…

Cited by 0SourceScholar
2026

Scalable Multi-View Subspace Clustering with Tensorized Anchor Guidance

CVPR 2026

Anchor-based multi-view clustering methods have gained significant attention for their effectiveness in handling large-scale datasets in recent years. The performance of these methods is highly dependent on anchor quality. However, current methods neglect the interactive relationships among cross-vi

Cited by 0SourcecodeScholar
2026

Structural Action Transformer for 3D Dexterous Manipulation

CVPR 2026

Achieving human-level dexterity in robots via imitation learning from heterogeneous datasets is hindered by the challenge of cross-embodiment skill transfer, particularly for high-DoF robotic hands. Existing methods, often relying on 2D observations and temporal-centric action representation, strugg

Cited by 0SourcecodeScholar
2026

UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVs

CVPR 2026

Active search and tracking of arbitrary targets by Unmanned Aerial Vehicles (UAVs) in cluttered environments remains a highly challenging problem. Existing methods either construct complex modular pipelines, leading to substantial computational costs, or adopt end-to-end controllers that often fail

Cited by 0SourcecodeScholar
2026

Wavelet Predictive Representations for Non-Stationary Reinforcement Learning

ICLR 2026poster

The real world is inherently non-stationary, with ever-changing factors, such as weather conditions and traffic flows, making it challenging for agents to adapt to varying environmental dynamics. Non-Stationary Reinforcement Learning (NSRL) addresses this challenge by training agents to adapt rapidl…

Cited by 0SourceScholar
2025

Active Perception Meets Rule-Guided RL: A Two-Phase Approach for Precise Object Navigation in Complex Environments

ICCV 2025poster

Object Goal Navigation (ObjectNav) in unknown environments presents significant challenges, particularly in Open-Vocabulary Mobile Manipulation (OVMM), where robots must efficiently explore large spaces, locate small objects, and accurately position themselves for subsequent manipulation. Existing a…

2025

Bright-NeRF: Brightening Neural Radiance Field with Color Restoration from Low-Light RAW Images

AAAI 2025technical

Neural Radiance Fields (NeRF) have demonstrated prominent performance in novel view synthesis tasks. However, their input heavily relies on image acquisition under normal light conditions, making it challenging to learn accurate scene contents in low-light environments where images typically exhibit…

Cited by 0SourcePDFScholar
2025

Chain-of-Focus Prompting: Leveraging Sequential Visual Cues to Prompt Large Autoregressive Vision Models

ICLR 2025poster

In-context learning (ICL) has revolutionized natural language processing by enabling models to adapt to diverse tasks with only a few illustrative examples. However, the exploration of ICL within the field of computer vision remains limited. Inspired by Chain-of-Thought (CoT) prompting in the langua…

Cited by 0SourcePDFScholar
2025

DP-Habitat: Bridging the Gap Between Simulation and Reality for Visual Navigation in Dynamic Pedestrian Environments

ICRA 2025

Visual navigation in dynamic environments poses a considerable challenge, particularly in scenarios with diverse pedestrian behaviors. Traditional simulators primarily focus on static scenes, while existing dynamic pedestrian simulators often suffer limitations such as monotonous pedestrian models,

Cited by 0SourcecodeScholar
2025

Geometric Logit Decoupling for Energy-Based Graph Out-of-distribution Detection

NeurIPS 2025poster

GNNs have achieved remarkable performance across a range of tasks, but their reliability under distribution shifts remains a significant challenge. In particular, energy-based OOD detection methods—which compute energy scores from GNN logits—suffer from unstable performance due to a fundamental coup…

Cited by 0SourceScholar
2025

LastingBench: Defend Benchmarks Against Knowledge Leakage

EMNLP 2025

The increasing size and complexity of large language models (LLMs) raise concerns about their ability to “cheat” on standard Question Answering (QA) benchmarks by memorizing task-specific data. This undermines the validity of benchmark evaluations, as they no longer reflect genuine model capabilitie

Cited by 0SourcePDFScholar
2025

Optimizing for the Shortest Path in Denoising Diffusion Model

CVPR 2025highlight

In this research, we propose a novel denoising diffusion model based on shortest-path modeling that optimizes residual propagation to enhance both denoising efficiency and quality. Drawing on Denoising Diffusion Implicit Models (DDIM) and insights from graph theory, our model, termed the Shortest Pa…

2024

Automated Non-invasive Analysis of Motile Sperms Using Cross-scale Guidance Network

ICRA 2024poster

Unbiased measurement of sperm morphometric and motility parameters is essential for assessing fertility potential and guiding visual feedback for microrobotic manipulation. Automated analysis of multiple sperms and selection of an optimal sperm is crucial for in vitro fertilisation treatment such as…

Cited by 0SourceScholar
2024

Image2Sentence based Asymmetrical Zero-shot Composed Image Retrieval

ICLR 2024spotlight

The task of composed image retrieval (CIR) aims to retrieve images based on the query image and the text describing the users' intent. Existing methods have made great progress with the advanced large vision-language (VL) model in CIR task, however, they generally suffer from two main issues: lack…

Cited by 12SourcePDFScholar
2024

Instance-aware Exploration-Verification-Exploitation for Instance ImageGoal Navigation

CVPR 2024poster

As a new embodied vision task Instance ImageGoal Navigation (IIN) aims to navigate to a specified object depicted by a goal image in an unexplored environment. The main challenge of this task lies in identifying the target object from different viewpoints while rejecting similar distractors. Existin…

2024

MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics Components

AAAI 2024technical

Meta-Reinforcement Learning (Meta-RL) aims to reveal shared characteristics in dynamics and reward functions across diverse training tasks. This objective is achieved by meta-learning a policy that is conditioned on task representations with encoded trajectory data or context, thus allowing rapid ad…

Cited by 2SourcePDFScholar
2024

Moderate Message Passing Improves Calibration: A Universal Way to Mitigate Confidence Bias in Graph Neural Networks

AAAI 2024technical

Confidence calibration in Graph Neural Networks (GNNs) aims to align a model's predicted confidence with its actual accuracy. Recent studies have indicated that GNNs exhibit an under-confidence bias, which contrasts the over-confidence bias commonly observed in deep neural networks. However, our dee…

Cited by 2SourcePDFScholar
2024

Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective Approach

AAAI 2024technical

Text recognition methods are gaining rapid development. Some advanced techniques, e.g., powerful modules, language models, and un- and semi-supervised learning schemes, consecutively push the performance on public benchmarks forward. However, the problem of how to better optimize a text recognition…

2024

Superpixel-informed Implicit Neural Representation for Multi-Dimensional Data

ECCV 2024poster

"Recently, implicit neural representations (INRs) have attracted increasing attention for multi-dimensional data recovery. However, INRs simply map coordinates via a multi-layer perceptron (MLP) to corresponding values, ignoring the inherent semantic information of the data. To leverage semantic pri…

Cited by 2SourcePDFScholar
2024

Underwater Vibration Adhesion by Frequency-Controlled Rigid Disc for Underwater Robotics Grasping

RA-L 2024

Underwater adsorption is the key function of underwater robot operation, in order to complete the underwater wall fixation, underwater object capture and other operations. The underwater adhesion mechanism based on vibration control has the advantage of using the liquid viscosity to pull the object

Cited by 5SourceScholar
2023

HandNeRF: Neural Radiance Fields for Animatable Interacting Hands

CVPR 2023poster

We propose a novel framework to reconstruct accurate appearance and geometry with neural radiance fields (NeRF) for interacting hands, enabling the rendering of photo-realistic images and videos for gesture animation from arbitrary views. Given multi-view images of a single hand or interacting hands…

Cited by 29SourcePDFScholar
2023

WaveForM: Graph Enhanced Wavelet Learning for Long Sequence Forecasting of Multivariate Time Series

AAAI 2023technical

Multivariate time series (MTS) analysis and forecasting are crucial in many real-world applications, such as smart traffic management and weather forecasting. However, most existing work either focuses on short sequence forecasting or makes predictions predominantly with time domain features, which…

2022

A Universal PINNs Method for Solving Partial Differential Equations with a Point Source

IJCAI 2022poster

In recent years, deep learning technology has been used to solve partial differential equations (PDEs), among which the physics-informed neural networks (PINNs)method emerges to be a promising method for solving both forward and inverse PDE problems. PDEs with a point source that is expressed as a D…

Cited by 12SourcePDFScholar
2022

CMT: Context-Matching-Guided Transformer for 3D Tracking in Point Clouds

ECCV 2022poster

"How to effectively match the target template features with the search area is the core problem in point-cloud-based 3D single object tracking. However, in the literature, most of the methods focus on devising sophisticated matching modules at point-level, while overlooking the rich spatial context…

Cited by 27SourcePDFScholar
2022

Learning Token-Based Representation for Image Retrieval

AAAI 2022technical

In image retrieval, deep local features learned in a data-driven manner have been demonstrated effective to improve retrieval performance. To realize efficient retrieval on large image database, some approaches quantize deep local features with a large codebook and match images with aggregated match…

2022

Meta-Auto-Decoder for Solving Parametric Partial Differential Equations

NeurIPS 2022accept

Many important problems in science and engineering require solving the so-called parametric partial differential equations (PDEs), i.e., PDEs with different physical parameters, boundary conditions, shapes of computation domains, etc. Recently, building learning-based numerical solvers for parametr…

Cited by 44SourcePDFScholar
2022

Optimization over Disentangled Encoding: Unsupervised Cross-Domain Point Cloud Completion via Occlusion Factor Manipulation

ECCV 2022poster

"Recently, studies considering domain gaps in shape completion attracted more attention, due to the undesirable performance of supervised methods on real scans. They only noticed the gap in input scans, but ignored the gap in output prediction, which is specific for completion. In this paper, we dis…

2022

Rethinking Efficient Lane Detection via Curve Modeling

CVPR 2022poster

This paper presents a novel parametric curve-based method for lane detection in RGB images. Unlike state-of-the-art segmentation-based and point detection-based methods that typically require heuristics to either decode predictions or formulate a large sum of anchors, the curve-based methods can lea…

Cited by 202PDFcodeScholar
2022

Versatile Motion Generation of Magnetic Origami Spring Robots in the Uniform Magnetic Field

RA-L 2022

Magnetic soft robots have attracted widespread attention for their untethered, remotely operated, and compliant deformation characteristics. Earlier work has demonstrated magnetic origami robots' diverse locomotion capabilities. This letter will focus on the motion generation and open-loop control o

Cited by 22SourceScholar
2021

A Novel end-to-end Speech Emotion Recognition Network with Stacked Transformer Layers

ICASSP 2021accepted

Speech emotion recognition (SER) aims to automatically recognize emotional category for a given speech utterance. The performance of a SER system heavily relies on the effectiveness of global representation expressed at utterance level. To effectively extract such a global feature, the mainstream of…

Cited by 0SourceScholar
2021

A Trace-restricted Kronecker-Factored Approximation to Natural Gradient

AAAI 2021technical

Second-order optimization methods have the ability to accelerate convergence by modifying the gradient through the curvature matrix. There have been many attempts to use second-order optimization methods for training deep neural networks. In this work, inspired by diagonal approximations and factore…

Cited by 13SourcePDFScholar
2021

Contextual Similarity Aggregation with Self-attention for Visual Re-ranking

NeurIPS 2021poster

In content-based image retrieval, the first-round retrieval result by simple visual feature comparison may be unsatisfactory, which can be refined by visual re-ranking techniques. In image retrieval, it is observed that the contextual similarity among the top-ranked images is an important clue to di…

2021

Design and Experimental Evaluation of a Multi-Mode Mobile Robot Based on Eccentric Paddle Mechanism

RA-L 2021

In order to improve climbing performance of the robot in complex environments, the wheel-legged mechanism is gradually being widely used. In this letter, we simplify the previous eccentric paddle mechanism (ePaddle), and propose a 2-DOF ePaddle-based robot which integrates stability and maneuverabil

Cited by 4SourceScholar
2021

Dynamic tracking for microrobot with active magnetic sensor array

ICRA 2021poster

Accurate position feedback in a wide range is critical for medical microrobotics and robot-assisted examinations, such as colonoscopy, bronchoscopy and capsule endoscopy examination. Among the many modalities of positioning feedback, magnetic tracking is a preferable method due to the unique advanta…

Cited by 11SourceScholar
2021

Learning Deep Local Features With Multiple Dynamic Attentions for Large-Scale Image Retrieval

ICCV 2021poster

In image retrieval, learning local features with deep convolutional networks has been demonstrated effective to improve the performance. To discriminate deep local features, some research efforts turn to attention learning. However, existing attention-based methods only generate a single attention m…

Cited by 31PDFcodeScholar
2021

SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate Curvature

CVPR 2021poster

The bottleneck of computation burden limits the widespread use of the 2nd order optimization algorithms for training deep neural networks. In this paper, we present a computationally efficient approximation for natural gradient descent, named Swift Kronecker-Factored Approximate Curvature (SKFAC), w…

Cited by 31PDFScholar
2021

THOR, Trace-based Hardware-driven Layer-Oriented Natural Gradient Descent Computation

AAAI 2021technical

It is well-known that second-order optimizer can accelerate the training of deep neural networks, however, the huge computation cost of second-order optimization makes it impractical to apply in real practice. In order to reduce the cost, many methods have been proposed to approximate a second-order…

Cited by 9SourcePDFScholar
2021

Unimodal and Crossmodal Refinement Network for Multimodal Sequence Fusion

EMNLP 2021main

Effective unimodal representation and complementary crossmodal representation fusion are both important in multimodal representation learning. Prior works often modulate one modal feature to another straightforwardly and thus, underutilizing both unimodal and crossmodal representation refinements, w…

2017

Integration of multiple genomic imaging data for the study of schizophrenia using joint nonnegative matrix factorization

ICASSP 2017accepted

Schizophrenia (SZ) is a complex disease caused by a lot genetic variants, epigenetic and brain region abnormalities. In this study, we adopted a joint nonnegative matrix factorization method to integrate three datasets including single nucleotide polymorphism (SNP), brain activity measured by functi…

Cited by 0SourceScholar