← Search

Lin Zhang

75 accepted papers

2026

Continuous-Space Multi-Agent Path Finding via Enhanced Prioritized Search With Ackermann Kinematic Constraints

RA-L 2026

Multi-Agent Path Finding (MAPF) in complex environments remains challenging due to high computational complexity, frequent conflicts, and realistic motion constraints. Most existing methods focus on discrete spaces or idealized omnidirectional models, often neglecting or partially considering nonhol

Cited by 0SourceScholar
2026

HYBRID PRUNING: IN-SITU COMPRESSION OF SELF-SUPERVISED SPEECH MODELS FOR SPEAKER VERIFICATION AND ANTI-SPOOFING

ICASSP 2026oral

Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on resource-constrained devices. While structured pruning is a key technique for model compression, existing methods typica…

Cited by 0SourcePDFScholar
2026

Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

ICML 2026poster

In long-horizon tasks, recent agents based on Large Language Models (LLMs) face a significant challenge that sparse, outcome-based rewards make it difficult to assign credit to intermediate steps. Previous methods mainly focus on creating dense reward signals to guide learning, either through tradit…

Cited by 0SourceScholar
2026

RealRep: Generalized SDR-to-HDR Conversion via Attribute-Disentangled Representation Learning

AAAI 2026technical

High-Dynamic-Range Wide-Color-Gamut (HDR-WCG) technology is becoming increasingly widespread, driving a growing need for converting Standard Dynamic Range (SDR) content to HDR. Existing methods primarily rely on fixed tone mapping operators, which struggle to handle the diverse appearances and degra

Cited by 0SourcePDFScholar
2026

RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation

CVPR 2026

Visual-language grounding aims to establish semantic correspondences between natural language and visual entities, enabling models to accurately identify and localize target objects based on textual instructions. Existing VLG approaches focus on coarse-grained, object-level localization, while tradi

Cited by 0SourcecodeScholar
2026

SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images

AAAI 2026technical

Recent advances in generative AI have accelerated the production of ultra-high-resolution visual content. However, traditional image formats face significant limitations in efficient compression and real-time decoding, which restricts their applicability on end-user devices. Inspired by 3D Gaussian

Cited by 0SourcePDFScholar
2025

BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference

AAAI 2025technical

The Mixture-of-Experts (MoE) structure scales the Transformer-based large language models (LLMs) and improves their performance with only the sub-linear increase in computation resources. Recently, a fine-grained DeepSeekMoE structure is proposed, which can further improve the computing efficiency o…

2025

CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification

ICASSP 2025accepted

Self-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-hea…

Cited by 0SourceScholar
2025

Continual Unsupervised Domain Adaptation for Audio Deepfake Detection

ICASSP 2025accepted

Audio deepfake detection (ADD) aims to verify the authenticity of audio. However, its performance declines sharply when facing significant domain discrepancies caused by unknown datasets. Unsupervised domain adaptation (UDA) has been applied to mitigate domain mismatch. However, as generative models…

Cited by 0SourceScholar
2025

DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models

CVPR 2025poster

Upcycled Mixture-of-Experts (MoE) models have shown great potential in various tasks by converting the original Feed-Forward Network (FFN) layers in pre-trained dense models into MoE layers. However, these models still suffer from significant parameter inefficiency due to the introduction of multipl…

Cited by 1SourcePDFScholar
2025

Deep Reinforcement Learning-Based Trajectory Tracking Framework for 4WS Robots Considering Switch of Steering Modes

IROS 2025

The application scenarios of automated robots are undergoing a paradigm shift from structured environments to unstructured, complex settings. In highly constrained settings like factory inspections or disaster rescue, conventional steering systems show clear drawbacks. While the four-wheel independe

Cited by 1SourceScholar
2025

Dynamic Network Topology Analysis, Design, and Evaluation for Multi-Robot Vehicle Transfer in High-Density Storage Yards

IROS 2025

With the rapid advancement of intelligent manufacturing and the rise of emerging markets, global auto-mobile exports have surged, placing unprecedented demands on logistics infrastructure. Efficient coordination of multiple robots for vehicle autonomous transfer is essential in high-density storage

Cited by 0SourceScholar
2025

Dynamically Optimize MTD Strategy in Satellite Computing Systems Using A2C Reinforcement Learning

ICASSP 2025accepted

The Satellite Computing System (SCS) faces an increasing number of attacks. Although Moving Target Defense (MTD) can effectively mitigate attacks in ground networks, it is not well-suited for SCS due to the highly dynamic nature of both SCS traffic and attackers’ scanning behaviors. In this paper, w…

Cited by 0SourceScholar
2025

EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification

NeurIPS 2025poster

Understanding the internal mechanisms of transformer-based language models remains challenging. Mechanistic interpretability based on circuit discovery aims to reverse engineer neural networks by analyzing their internal processes at the level of computational subgraphs. In this paper, we revisit ex…

Cited by 0SourceScholar
2025

Efficient and Accurate Prompt Optimization: the Benefit of Memory in Exemplar-Guided Reflection

ACL 2025long

Automatic prompt engineering aims to enhance the generation quality of large language models (LLMs). Recent works utilize feedbacks generated from erroneous cases to guide the prompt optimization. During inference, they may further retrieve several semantically-related exemplars and concatenate them…

2025

Entrospect: Information-Theoretic Self-Reflection Elicits Better Response Refinement of Small Language Models

ACL 2025finding

Self-reflection helps de-hallucinate Large Language Models (LLMs). However, the effectiveness of self-reflection remains insufficiently validated in the context of Small Language Models (SLMs), which exhibit limited semantic capacities. In particular, we demonstrate that the conventional self-reflec…

2025

FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have shown impressive video content understanding capabilities but struggle with fine-grained motion comprehension. To comprehensively assess the motion understanding ability of existing MLLMs, we introduce FAVOR-Bench, which comprises 1,776 videos from both…

Cited by 0SourceScholar
2025

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

CVPR 2025poster

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth o…

2025

HFSENet: Hierarchical Fusion Semantic Enhancement Network for RGB-T Semantic Segmentation in Annealing Furnace Operation Area

IROS 2025

Regular temperature measurement of critical parts of an annealing furnace has always been a difficult task. Due to the harsh environment of high temperature, high noise, and darkness in the annealing furnace operation area, unmanned vehicles equipped with the RGB-T semantic segmentation model are us

Cited by 0SourceScholar
2025

LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation

ICASSP 2025accepted

Previous fake speech datasets were constructed from a defender’s perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created LlamaPartialSpoof, a 130-hour dataset that contains both fully and partially…

Cited by 0SourceScholar
2025

M3HG: Multimodal, Multi-scale, and Multi-type Node Heterogeneous Graph for Emotion Cause Triplet Extraction in Conversations

ACL 2025finding

Emotion Cause Triplet Extraction in Multimodal Conversations (MECTEC) has recently gained significant attention in social media analysis, aiming to extract emotion utterances, cause utterances, and emotion categories simultaneously. However, the scarcity of related datasets, with only one published…

2025

Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning

NAACL 2025findings

Transformer-based language models have achieved significant success; however, their internal mechanisms remain largely opaque due to the complexity of non-linear interactions and high-dimensional operations. While previous studies have demonstrated that these models implicitly embed reasoning trees,…

Cited by 0SourcePDFScholar
2025

OSDFace: One-Step Diffusion Model for Face Restoration

CVPR 2025poster

Diffusion models have demonstrated impressive performance in face restoration. Yet, their multi-step inference process remains computationally intensive, limiting their applicability in real-world scenarios. Moreover, existing methods often struggle to generate face images that are harmonious, reali…

2025

PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding

NeurIPS 2025poster

While Large Language Models (LLMs) demonstrate strong performance across domains, their long-context capabilities are limited by transient neural activations causing information decay and unstructured feed-forward network (FFN) weights leading to semantic fragmentation. Inspired by the brain’s worki…

Cited by 0SourceScholar
2025

Representing Sounds as Neural Amplitude Fields: A Benchmark of Coordinate-MLPs and a Fourier Kolmogorov-Arnold Framework

AAAI 2025technical

Although Coordinate-MLP-based implicit neural representations have excelled in representing radiance fields, 3D shapes, and images, their application to audio signals remains underexplored. To fill this gap, we investigate existing implicit neural representations, from which we extract 3 types of po…

2025

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning

ICCV 2025poster

We propose SC-Captioner, a reinforcement learning framework that enables the self-correcting capability of image caption models. Our crucial technique lies in the design of the reward function to incentivize accurate caption corrections. Specifically, the predicted and reference captions are decompo…

2025

The Missing Piece in Model Editing: A Deep Dive into the Hidden Damage Brought By Model Editing

ICASSP 2025accepted

Large Language Models have revolutionized numerous tasks with their remarkable efficacy. However, editing these models, crucial for rectifying outdated or erroneous information, often leads to a complex issue known as the ripple effect in the hidden space. While difficult to detect, this effect can…

Cited by 0SourceScholar
2025

Towards Audio-Visual Navigation in Noisy Environments: A Large-Scale Benchmark Dataset and an Architecture Considering Multiple Sound-Sources

AAAI 2025technical

Audio-visual navigation has received considerable attention in recent years. However, the majority of related investigations have focused on single sound-source scenarios. Studies in this field for multiple sound-source scenarios remain underexplored due to the limitations of two aspects. First, the…

2024

3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object Detection

NeurIPS 2024poster

Transformer-based architectures have been proven successful in detecting 3D objects from point clouds. However, the quadratic complexity of the attention mechanism struggles to encode rich information as point cloud resolution increases. Recently, state space models (SSM) such as Mamba have gained g…

Cited by 0SourcePDFScholar
2024

CPAUG: Refining Copy-Paste Augmentation for Speech Anti-Spoofing

ICASSP 2024accepted

Conventional copy-paste augmentations generate new training instances by concatenating existing utterances to increase the amount of data for neural network training. However, the direct application of copy-paste augmentation for anti-spoofing is problematic. This paper refines the copy-paste augmen…

Cited by 0SourceScholar
2024

DetectBench: Can Large Language Model Detect and Piece Together Implicit Evidence?

EMNLP 2024finding

Detecting evidence within the context is a key step in the process of reasoning task. Evaluating and enhancing the capabilities of LLMs in evidence detection will strengthen context-based reasoning performance. This paper proposes a benchmark called DetectBench for verifying the ability to detect an…

2024

The Control Strategy for Vehicle Transfer Robots in RO/RO Terminal Environments

IROS 2024poster

In the labor-intensive Roll-On/Roll-Off (RO/RO) terminal environment, research on vehicle transport robots with mobility, stability, and reliability is receiving increasing attention. This paper presents a novel control framework for a Straddle-Type Dual-Body vehicle transfer robot. Initially, fine…

Cited by 0SourceScholar
2024

Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation

AAAI 2024technical

New Natural Langauge Process~(NLP) benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present Xiezhi, the most comprehensive evaluation suite designed to assess holistic domain knowledge.Xiezhi comprises multiple-choice questions across 516 diverse…

2023

A Diffusion Model for Event Skeleton Generation

ACL 2023findings

Event skeleton generation, aiming to induce an event schema skeleton graph with abstracted event nodes and their temporal relations from a set of event instance graphs, is a critical step in the temporal complex event schema induction task. Existing methods effectively address this task from a graph…

2023

D2Match: Leveraging Deep Learning and Degeneracy for Subgraph Matching

ICML 2023poster

Subgraph matching is a fundamental building block for graph-based applications and is challenging due to its high-order combinatorial nature. Existing studies usually tackle it by combinatorial optimization or learning-based methods. However, they suffer from exponential computational costs or searc…

2023

Do Not Train It: A Linear Neural Architecture Search of Graph Neural Networks

ICML 2023poster

Neural architecture search (NAS) for Graph neural networks (GNNs), called NAS-GNNs, has achieved significant performance over manually designed GNN architectures. However, these methods inherit issues from the conventional NAS methods, such as high computational cost and optimization difficulty. Mor…

2023

Dynamic Obstacle Avoidance for Cable-Driven Parallel Robots With Mobile Bases via Sim-to-Real Reinforcement Learning

RA-L 2023

A Cable-Driven Parallel Robot (CDPR) with Mobile Bases (MBs) can modify its geometric architecture and is suitable for manipulation tasks in constrained environments. In manipulation tasks, a CDPR with MBs inevitably encounters obstacles, including dynamic obstacles. However, the high dimensional st

Cited by 31SourceScholar
2023

Feature Expansion for Graph Neural Networks

ICML 2023poster

Graph neural networks aim to learn representations for graph-structured data and show impressive performance in node classification. Recently, many methods have studied the representations of GNNs from the perspective of optimization goals and spectral graph theory. However, the feature space that d…

2023

LMR: A Large-Scale Multi-Reference Dataset for Reference-Based Super-Resolution

ICCV 2023poster

It is widely agreed that reference-based super-resolution (RefSR) achieves superior results by referring to similar high quality images, compared to single image super-resolution (SISR). Intuitively, the more references, the better performance. However, previous RefSR methods have all focused on sin…

Cited by 23PDFcodeScholar
2023

MAP: Multimodal Uncertainty-Aware Vision-Language Pre-Training Model

CVPR 2023poster

Multimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets. Such uncertainty is problematic for our interpretation, including inter- and intra-modal uncertainty. Little effort has studied the modeling of this uncertainty,…

2023

MVP-Tuning: Multi-View Knowledge Retrieval with Prompt Tuning for Commonsense Reasoning

ACL 2023long

Recent advances in pre-trained language models (PLMs) have facilitated the development ofcommonsense reasoning tasks. However, existing methods rely on multi-hop knowledgeretrieval and thus suffer low accuracy due toembedded noise in the acquired knowledge. In addition, these methods often attain hi…

2023

Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker Verification

ICASSP 2023accepted

Optimal transport (OT) can be applied to cross-domain adaptation in speaker verification (SV) by converting speakers' probability distributions from source to target domains. However, in scenarios involving over-massive categories (speakers) or difficult samples in discrimination, OT often has diffi…

Cited by 0SourceScholar
2023

Solving Math Word Problems via Cooperative Reasoning induced Language Models

ACL 2023long

Large-scale pre-trained language models (PLMs) bring new opportunities to challenging problems, especially those that need high-level intelligence, such as the math word problem (MWPs). However, directly applying existing PLMs to MWPs can fail as the generation process lacks sufficient supervision a…

2022

A General Framework For Incomplete Cross-Modal Retrieval With Missing Labels And Missing Modalities

ICASSP 2022accepted

Among various cross-modal retrieval methods, the supervised methods achieve the best performance by exploiting the semantic labels. However, in realistic applications, the data are not always complete with labels and full multi-modal data, which makes these methods hard to be used. In this paper, we…

Cited by 0SourceScholar
2022

CCRobot-V: A Silkworm-Like Cooperative Cable-Climbing Robotic System for Cable Inspection and Maintenance

ICRA 2022poster

This paper presents CCRobot-V, the fifth version of CCRobot, a cooperative serial multi-robot system for bridge cable inspection and maintenance that uses silkworm-like locomotion to climb the entire length of super-long stay cable at high speeds while carrying heavy inspection/maintenance equipment…

Cited by 13SourceScholar
2022

CS-REP: Making Speaker Verification Networks Embracing Re-Parameterization

ICASSP 2022accepted

Automatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and verification accuracy are essential. This study proposes cross-s…

Cited by 0SourceScholar
2022

Chunkfusion: A Learning-Based RGB-D 3D Reconstruction Framework Via Chunk-Wise Integration

ICASSP 2022accepted

Recent years have witnessed a growing interest in online RGB-D 3D reconstruction. On the premise of ensuring the reconstruction accuracy with noisy depth scans, making the system scalable to various environments is still challenging. In this paper, we devote our efforts to try to fill in this resear…

Cited by 0SourceScholar
2022

Exact Shape Correspondence via 2D graph convolution

NeurIPS 2022accept

For exact 3D shape correspondence (matching or alignment), i.e., the task of matching each point on a shape to its exact corresponding point on the other shape (or to be more specific, matching at geodesic error 0), most existing methods do not perform well due to two main problems. First, on nearly…

Cited by 6SourcePDFScholar
2022

Fine-Tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated Learning

CVPR 2022poster

Federated Learning (FL) is an emerging distributed learning paradigm under privacy constraint. Data heterogeneity is one of the main challenges in FL, which results in slow convergence and degraded performance. Most existing approaches only tackle the heterogeneity challenge by restricting the local…

Cited by 385PDFcodeScholar
2022

H2GNN: Hierarchical-Hops Graph Neural Networks for Multi-Robot Exploration in Unknown Environments

RA-L 2022

Multi-robot coarse-to-fine exploration in unknown environments makes great sense in many application fields like search and rescue. For different stages of the task, robots need to extract information from the environment discriminately, which can improve their decision-making capability. To this en

Cited by 50SourceScholar
2022

Learning From Temporal Gradient for Semi-Supervised Action Recognition

CVPR 2022poster

Semi-supervised video action recognition tends to enable deep neural networks to achieve remarkable performance even with very limited labeled data. However, existing methods are mainly transferred from current image-based methods (e.g., FixMatch). Without specifically utilizing the temporal dynamic…

Cited by 88PDFcodeScholar
2022

PCBERT: Parent and Child BERT for Chinese Few-shot NER

COLING 2022main

Achieving good performance on few-shot or zero-shot datasets has been a long-term challenge for NER. The conventional semantic transfer approaches on NER will decrease model performance when the semantic distribution is quite different, especially in Chinese few-shot NER. Recently, prompt-tuning has…

Cited by 13SourcePDFScholar
2022

RRSR:Reciprocal Reference-Based Image Super-Resolution with Progressive Feature Alignment and Selection

ECCV 2022poster

"Reference-based image super-resolution (RefSR) is a promising SR branch and has shown great potential in overcoming the limitations of single image super-resolution. While previous state-of-the-art RefSR methods mainly focus on improving the efficacy and robustness of reference feature transfer, it…

Cited by 19SourcePDFScholar
2022

Self-Augmented Unpaired Image Dehazing via Density and Depth Decomposition

CVPR 2022poster

To overcome the overfitting issue of dehazing models trained on synthetic hazy-clean image pairs, many recent methods attempted to improve models' generalization ability by training on unpaired data. Most of them simply formulate dehazing and rehazing cycles, yet ignore the physical properties of th…

Cited by 261PDFcodeScholar
2022

Towards Controllable and Physical Interpretable Underwater Scene Simulation

ICASSP 2022accepted

The realistic simulation of underwater scenes has important significance for many researches related to underwater vision, such as underwater image restoration, underwater moving object monitoring, etc. To date, however, the existing underwater scene simulation pipelines are either too complicated d…

Cited by 0SourceScholar
2022

Zero-Shot Learners for Natural Language Understanding via a Unified Multiple Choice Perspective

EMNLP 2022main

We propose a new paradigm for zero-shot learners that is format agnostic, i.e., it is compatible with any format and applicable to a list of language tasks, such as text classification, commonsense reasoning, coreference resolution, and sentiment analysis. Zero-shot learning aims to train a model on…

2021

Federated Learning for Non-IID Data via Unified Feature Learning and Optimization Objective Alignment

ICCV 2021poster

Federated Learning (FL) aims to establish a shared model across decentralized clients under the privacy-preserving constraint. Despite certain success, it is still challenging for FL to deal with non-IID (non-independent and identical distribution) client data, which is a general scenario in real-wo…

Cited by 100PDFScholar
2021

MT-ORL: Multi-Task Occlusion Relationship Learning

ICCV 2021poster

Retrieving occlusion relation among objects in a single image is challenging due to sparsity of boundaries in image. We observe two key issues in existing works: firstly, lack of an architecture which can exploit the limited amount of coupling in the decoder stage between the two subtasks, namely oc…

Cited by 8PDFcodeScholar
2021

Semi-Supervised Multimodal Image Translation for Missing Modality Imputation

ICASSP 2021accepted

Missing data is a common problem in multimodal and multi-view learning. It raises a critical challenge for most multimodal algorithms, which are unable to deal with incomplete datasets. Rather than discarding entries with missing modalities, this paper aims to reconstruct the complete image-based mu…

Cited by 0SourceScholar
2021

Towards Robust Autonomous Coverage Navigation for Carlike Robots

RA-L 2021

Thanks to their high carrying capacity and strong maneuverability, carlike robots which move with non-holonomic constraints, are frequently utilized in numerous coverage operation fields. In such fields, the robots need to complete the coverage task via autonomous path planning and tracking, which i

Cited by 5SourcecodeScholar
2019

A Self-Adaptive Motion Scaling Framework for Surgical Robot Remote Control

RA-L 2019

Master-slave control is a common form of human-robot interaction for robotic surgery. To ensure seamless and intuitive control, a mechanism of self-adaptive motion scaling during teleoperaton is proposed in this letter. The operator can retain precise control when conducting delicate or complex mani

Cited by 44SourceScholar
2019

Design and Fabrication of a 3-D Printed Metallic Flexible Joint for Snake-Like Surgical Robot

RA-L 2019

Snake-like robots have numerous applications in minimally invasive surgery. One important research topic of snake-like robots is the flexible joint mechanism and its actuation. This letter describes the design and fabrication of a new type of flexible joint mechanism that is enabled by metal powder

Cited by 68SourceScholar
2019

Design and Verification of A Portable Master Manipulator Based on an Effective Workspace Analysis Framework

IROS 2019poster

Master manipulators represent a key component of Robot-Assisted Minimally Invasive Surgery (RAMIS). In this paper, an Analytic Hierarchy Process (AHP) method is used to construct an effective workspace analysis framework, which can assist the configuration selection and design evaluation of a portab…

Cited by 20SourceScholar
2019

Designing, Prototyping, and Testing a Flexible Suturing Robot for Transanal Endoscopic Microsurgery

RA-L 2019

Suturing and knot tying in a confined space is a technically challenging yet clinically demanding task in minimally invasive surgery, which requires the use of highly articulated instruments passing through small incisions on the patient's body. Manually operating such instruments is usually very di

Cited by 23SourceScholar
2018

Cross-Scene Suture Thread Parsing for Robot Assisted Anastomosis based on Joint Feature Learning

IROS 2018poster

Task autonomy is an important consideration for the development of future surgical robots. For robot-assisted anastomosis, suture thread detection is a prerequisite for subsequent robot manipulation. Previous works on automatic thread detection are focused on the learning of the models with specific…

Cited by 11SourceScholar
2018

Depth Estimation of Optically Transparent Microrobots Using Convolutional and Recurrent Neural Networks

IROS 2018poster

Estimating the three-dimensional (3D) position of microrobots is necessary in order to develop closed-loop control techniques and to improve the user's 3D perception in the micro-scale. This paper describes a depth estimation method based on supervised learning for optically transparent microrobots…

Cited by 9SourceScholar
2017

Autonomous scanning for endomicroscopic mosaicing and 3D fusion

ICRA 2017poster

Robot-assisted minimally invasive surgery can benefit from the automation of common, repetitive or well-defined but ergonomically difficult tasks. One such task is the scanning of a pick-up endomicroscopy probe over a complex, undulating tissue surface to enhance the effective field-of-view through…

Cited by 54SourceScholar