← Search

Ning Wang

38 accepted papers

2026

Design of an Active Haptic Interface Using Proprioception Feedback for Continuous Endovascular Teleoperation

RA-L 2026

Force feedback is essential for safe endovascular teleoperation, yet typically constrained by complex sensor integration. This article presents a compact active haptic interface system designed for robotic catheterization. Leveraging the intrinsic proprioception of a Permanent Magnet Synchronous Mot

Cited by 0SourceScholar
2026

FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification

AAAI 2026technical

We introduce FIXME, the first end-to-end and large-scale benchmark for evaluating Large Language Models (LLMs) in hardware design functional verification (FV). Comprising 747 tasks derived from real-world hardware designs, FIXME spans five core FV sub-sets: specification comprehension, reference mod

Cited by 0SourcePDFScholar
2026

OpenTSLM: Time-Series Language Models for Reasoning over Multivariate Medical Text- and Time-Series Data

ICML 2026poster

Large Language Models (LLMs) have shown strong capability in interpreting multimodal data but remain limited in their ability to natively handle time-series data. Addressing this limitation could enable the translation of longitudinal and wearable sensing data into actionable insights and patient-fa…

Cited by 0SourceScholar
2026

PolygMap: A Perceptive Locomotion Framework for Humanoid Robot Stair Climbing

ICRA 2026poster

Recently, biped robot walking technology has been significantly developed; however, mainly in a bland walking scheme. To emulate human walking, robots need to step on the positions they see in unknown spaces accurately. In this paper, we present PolyMap, a perception-based locomotion planning framew…

2026

REAP: Enhancing RAG with Recursive Evaluation and Adaptive Planning for Multi-Hop Question Answering

AAAI 2026technical

Retrieval-augmented generation (RAG) has been extensively employed to mitigate hallucinations in large language models (LLMs). However, existing methods for multi-hop reasoning tasks often lack global planning, increasing the risk of falling into local reasoning impasses. Insufficient exploitation o

Cited by 0SourcePDFScholar
2026

SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition

CVPR 2026

Zero-shot skeleton-based action recognition aims to recognize unseen actions by transferring knowledge from seen categories through semantic descriptions. Most existing methods typically align skeleton features with textual embeddings within a shared latent space. However, the absence of contextual

Cited by 0SourcecodeScholar
2026

SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free Acceleration

CVPR 2026

Visual AutoRegressive (VAR) modeling has garnered significant attention for its innovative next-scale prediction paradigm. However, mainstream VAR paradigms attend to all tokens across historical scales at each autoregressive step. As the next scale resolution grows, the computational complexity of

Cited by 0SourcecodeScholar
2025

Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation

IJCAI 2025

Embodied agents exhibit immense potential across a multitude of domains, making the assurance of their behavioral safety a fundamental prerequisite for their widespread deployment. However, existing research predominantly concentrates on the security of general large language models, lacking special

2025

Human-Robot Cooperative Heavy Payload Manipulation based on Whole-Body Model Predictive Control

IROS 2025

Human-robot collaborative manipulation with mobile, multiple manipulators is crucial for expanding robotic applications, requiring precise handling of coupled force-position constraints between partners. Current systems, however, exhibit end-effector oscillations and instability during dynamic inter

Cited by 1SourceScholar
2025

Intervening in Black Box: Concept Bottleneck Model for Enhancing Human Neural Network Mutual Understanding

ICCV 2025poster

Recent advances in deep learning have led to increasingly complex models with deeper layers and more parameters, reducing interpretability and making their decisions harder to understand. While many methods explain black-box reasoning, most lack effective interventions or only operate at sample-leve…

2025

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency

ACL 2025long

Large Multimodal Models (LMMs) have recently demonstrated impressive performance on general video comprehension benchmarks. Nevertheless, for broader applications, the robustness of their temporal analysis capability needs to be thoroughly investigated yet predominantly ignored. Motivated by this, w…

Cited by 0SourcePDFScholar
2025

Prompt-guided Disentangled Representation for Action Recognition

NeurIPS 2025poster

Action recognition is a fundamental task in video understanding. Existing methods typically extract unified features to process all actions in one video, which makes it challenging to model the interactions between different objects in multi-action scenarios. To alleviate this issue, we explore dise…

Cited by 0SourcecodeScholar
2025

RobAVA: A Large-scale Dataset and Baseline Towards Video based Robotic Arm Action Understanding

ICCV 2025poster

Understanding the behaviors of robotic arms is essential for various robotic applications such as logistics management, precision agriculture, and automated manufacturing. However, the lack of large-scale and diverse datasets significantly hinders progress in video-based robotic arm action understan…

2025

Robot-Based Automatic Charging for Electric Vehicles Using Incremental Learning and Biomimetic Control

ICRA 2025

With the growing popularity of electric vehicles, the demand for robot-based unmanned automatic charging has become both urgent and challenging. Two key challenges need to be addressed: how to efficiently locate the charging port, and how to compliantly insert the connector into the port. In this pa

Cited by 0SourceScholar
2025

Unsupervised Multi-View Outlier Detection via Optimal Graph Filtering

ICASSP 2025accepted

Unsupervised multi-view outlier detection has garnered increasing attention in recent years, yet existing methods face persistent challenges. Many approaches rely predominantly on first-order neighborhood information, overlooking the richer insights offered by higher-order structures, which can degr…

Cited by 0SourceScholar
2025

VLIMNet: A Visible Light And Infrared Image Matching Network Based On Segment Anything Model And SuperPoint

ICASSP 2025accepted

This paper introduces a novel method for matching visible light and infrared images, termed the Visible Light and Infrared Image Matching Network (VLIMNet). In the image encoding stage, we incorporate a generative architecture-based modality transformation network after the SuperPoint encoder, enabl…

Cited by 0SourceScholar
2024

A Prompt-Based Method with Multi-View Optimization for Open Relation Extraction

ICASSP 2024accepted

Open Relation Extraction (OpenRE) is a task that involves discovering new relation types by referring to labeled instances. Existing methods mainly rely on large pre-trained models to obtain the relation representation of entity pairs, and then jointly train the supervised and unsupervised data usin…

Cited by 0SourceScholar
2024

Bridging the Gap: Sketch to Color Diffusion Model with Semantic Prompt Learning

ICASSP 2024accepted

Automatic anime sketch colorization aims to generate a color image from a sketch image, which is challenging due to limited structure and semantic understanding, leading to constrained style, and semantic color inconsistency. In this paper, we introduce a sketch to color diffusion model with semanti…

Cited by 0SourceScholar
2024

General Collaborative Framework between Large Language Model and Experts for Universal Information Extraction

EMNLP 2024finding

Recently, unified information extraction has garnered widespread attention from the NLP community, which aims to use a unified paradigm to perform various information extraction tasks. However, prevalent unified IE approaches inevitably encounter challenges such as noise interference, abstract label…

Cited by 0SourcePDFScholar
2024

Language Model Guided Interpretable Video Action Reasoning

CVPR 2024poster

Although neural networks excel in video action recognition tasks their "black-box" nature makes it challenging to understand the rationale behind their decisions. Recent approaches used inherently interpretable models to analyze video actions in a manner akin to human reasoning. However it has been…

2023

Efficient Image Captioning for Edge Devices

AAAI 2023technical

Recent years have witnessed the rapid progress of image captioning. However, the demands for large memory storage and heavy computational burden prevent these captioning models from being deployed on mobile devices. The main obstacles lie in the heavyweight visual feature extractors (i.e., object de…

Cited by 22SourcePDFScholar
2023

Vision-and-Force-Based Compliance Control for a Posterior Segment Ophthalmic Surgical Robot

RA-L 2023

In ophthalmic surgery, particularly in procedures involving the posterior segment, clinicians face significant challenges in maintaining precise control of hand-held instruments without damaging the fundus tissue. Typical targets of this type of surgery are the internal limiting membrane (ILM) and t

Cited by 7SourceScholar
2022

A 5-DOFs Robot for Posterior Segment Eye Microsurgery

RA-L 2022

In retinal surgery clinicians access the internal volume of the eyeball through small scale trocar ports, typically 0.65 mm in diameter, to treat vitreoretinal disorders like idiopathic epiretinal membrane and age-related macular holes. The treatment of these conditions involves the removal of thin

Cited by 17SourceScholar
2022

An Adaptive Fuzzy Control for Human-in-the-Loop Operations With Varying Communication Time Delays

RA-L 2022

Time delay, especially varying time delay, is always an important factor affecting the stability to the human-in-the-loop system. Previous research usually focuses on the performance of the internal signal transmission part, but rarely considers the whole system with human and environmental factors

Cited by 8SourceScholar
2022

Diff-Net: Image Feature Difference Based High-Definition Map Change Detection for Autonomous Driving

ICRA 2022poster

Up-to-date High-Definition (HD) maps are essential for self-driving cars. To achieve constantly updated HD maps, we present a deep neural network (DNN), Diff-Net, to detect changes in them. Compared to traditional methods based on object detectors, the essential design in our work is a parallel feat…

Cited by 8SourceScholar
2021

Contrastive Transformation for Self-supervised Correspondence Learning

AAAI 2021technical

In this paper, we focus on the self-supervised learning of visual correspondence using unlabeled videos in the wild. Our method simultaneously considers intra- and inter-video representation associations for reliable correspondence estimation. The intra-video learning transforms the image contents a…

2021

Joint Inductive and Transductive Learning for Video Object Segmentation

ICCV 2021poster

Semi-supervised video object segmentation is a task of segmenting the target object in a video sequence given only a mask annotation in the first frame. The limited information available makes it an extremely challenging task. Most previous best-performing methods adopt matching-based transductive r…

Cited by 122PDFcodeScholar
2021

Transformer Meets Tracker: Exploiting Temporal Context for Robust Visual Tracking

CVPR 2021poster

In video object tracking, there exist rich temporal contexts among successive frames, which have been largely overlooked in existing trackers. In this work, we bridge the individual video frames and explore the temporal contexts across them via a transformer architecture for robust object tracking.…

Cited by 865PDFcodeScholar
2020

Adaptive Informative Sampling with Environment Partitioning for Heterogeneous Multi-Robot Systems

IROS 2020poster

Multi-robot systems are widely used in environmental exploration and modeling, especially in hazardous environments. However, different types of robots are limited by different mobility, battery life, sensor type, etc. Heterogeneous robot systems are able to utilize various types of robots and provi…

Cited by 45SourceScholar
2020

NAS-FCOS: Fast Neural Architecture Search for Object Detection

CVPR 2020poster

The success of deep neural networks relies on significant architecture engineering. Recently neural architecture search (NAS) has emerged as a promise to greatly reduce manual effort in network design by automatically searching for optimal architectures, although typically such algorithms need an ex…

Cited by 281PDFScholar
2018

Multi-Cue Correlation Filters for Robust Visual Tracking

CVPR 2018poster

In recent years, many tracking algorithms achieve impressive performance via fusing multiple types of features, however, most of them fail to fully explore the context among the adopted multiple features and the strength of them. In this paper, we propose an efficient multi-cue analysis framework fo…

2017

A friction model with velocity, temperature and load torque effects for collaborative industrial robot joints

IROS 2017poster

In this paper, a comprehensive friction model for collaborative industrial robot joints is proposed which takes into account the velocity, temperature and load torque effects. The model indicates that the velocity and temperature have a strong influence on viscous friction nonlinearly, whereas load…

Cited by 40SourceScholar
2016

Control and modeling for direct teaching of industrial articulated robotic arms

IROS 2016poster

This paper presents an improved force-free control method based on current, which can be applied to industrial articulated robotic arms with large mass and large friction torque for direct teaching. Three kinds of torques that influence direct teaching are analyzed, and thus a calibration method and…

Cited by 8SourceScholar