← Search

Peng Li

159 accepted papers

2026

AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers

AAAI 2026technical

Visual autoregressive modeling (VAR) via next-scale prediction has emerged as a scalable image generation paradigm. While Key and Value (KV) caching in large language models (LLMs) has been extensively studied, next-scale prediction presents unique challenges, and KV caching design for next-scale ba

Cited by 0SourcePDFScholar
2026

GAUSS: Graph-Assisted Uncertainty Quantification using Structure and Semantics for Long-Form Generation in LLMs

ICML 2026poster

In critical domains like clinical reporting, legal analysis, and policy drafting, large language models (LLMs) are increasingly expected to produce extended, fact‑rich narratives rather than isolated sentences. Reliable uncertainty quantification in such long‑form outputs is crucial. Existing techni…

Cited by 0SourceScholar
2026

GDAs-OT: A Prediction Method of Gene-Disease Associations Based on Optimal Transport for Identifying Genes Related to Immune-Related Adverse Events

IJCAI 2026

Immune Checkpoint Inhibitors (ICIs) represent a cornerstone of modern cancer immunotherapy. However, their clinical application is frequently accompanied by immune-related Adverse Events (irAEs) of diverse severity. Predicting Gene-Disease Associations (GDAs) is crucial for identifying the related g

Cited by 0Scholar
2026

GraphFlow: A Graph-Based Workflow Management for Efficient LLM-Agent Serving

ICML 2026poster

Large Language Model (LLM)-based agents demonstrate strong reasoning and execution capabilities on complex tasks when guided by structured instructions, commonly referred to as workflows. However, existing workflow-assisted agent serving systems typically rely on predefined templates and shallow mat…

Cited by 0SourceScholar
2026

Khan-GCL: Kolmogorov–Arnold Network Based Graph Contrastive Learning with Hard Negatives

AAAI 2026technical

Graph contrastive learning (GCL) has demonstrated great promise for learning generalizable graph representations from unlabeled data. However, conventional GCL approaches face two critical limitations: (1) the restricted expressive capacity of multilayer perceptron (MLP) based encoders, and (2) subo

Cited by 0SourcePDFScholar
2026

Listening Through the Noise: Cauchy-Driven Diffusion Bridges for Robust Gastrointestinal Auscultation and Clinical Benchmarking

ICML 2026spotlight

Gastrointestinal (GI) motility assessment via bowel sounds (BS) offers a non-invasive alternative to resource-intensive clinical standards. However, the diagnostic utility of BS is often compromised by its spectral overlap with non-stationary speech interference. While generative models have advance…

Cited by 0SourceScholar
2026

MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly

CVPR 2026

Scaling artist-designed meshes to high triangle numbers remains challenging for autoregressive generative models. Existing transformer-based methods suffer from long-sequence bottlenecks and limited quantization resolution, primarily due to the large number of tokens required and constrained quantiz

Cited by 0SourcecodeScholar
2026

PRIME: A Decoupled Multi-agent Actor-Critic for Multi-view Clustering

IJCAI 2026

Deep multi-view clustering draws plentiful attention in various domains, owing to remarkable performance in learning patterns from complementary information of multi-view data. However, previous methods encounter two challenges. They utilize a single pre-defined clustering strategy to perceive diver

Cited by 0Scholar
2026

Tab-semiSL: Tabular Data-Driven Semi-Supervised Learning to Identify Factors Associated with Immune-Related Adverse Events

IJCAI 2026

Immune Checkpoint Inhibitors (ICIs) have become a major therapeutic strategy in cancer treatment. However, widespread ICIs use can cause mild-to-severe immune-related Adverse Events (irAEs). Identifying factors associated with irAEs is beneficial for assessing the risk of irAEs occurrence during ICI

Cited by 0Scholar
2026

UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Pass

CVPR 2026

We present UniSH, a unified, feed-forward framework for joint metric-scale 3D scene and human reconstruction. A key challenge in this domain is the scarcity of large-scale, annotated real-world data, forcing a reliance on synthetic datasets. This reliance introduces a significant sim-to-real domain

Cited by 0SourcecodeScholar
2026

VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning

CVPR 2026

Most of the multi-agent video understanding frameworks adopt static and non-learnable tool invocation mechanisms, which limit the discovery of diverse clues essential for robust perception and reasoning regarding temporally or spatially complex videos. To address this challenge, we propose a novel M

Cited by 0SourceScholar
2026

Visual-Friendly Concept Protection via Selective Adversarial Perturbations

AAAI 2026technical

Personalized concept generation by tuning diffusion models with a few images raises potential legal and ethical concerns regarding privacy and intellectual property rights. Researchers attempt to prevent malicious personalization using adversarial perturbations. However, previous efforts have mainly

Cited by 0SourcePDFScholar
2026

YuE: Scaling Open Foundation Models for Long-Form Music Generation

ICLR 2026poster

We tackle the task of long-form music generation, particularly the challenging \textbf{lyrics-to-song} problem, by introducing \textbf{YuE (乐)}, a family of open-source music generation foundation models. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while…

Cited by 0SourcecodeScholar
2025

A Novel Effective Loop Gait and Stabilizing Morphology Parameterization in Snake Robots

IROS 2025

Improving motion speed and efficiency remains a critical challenge in snake robots gait control. This paper introduces the Loop gait, a novel locomotion gait designed to enhance both speed and energy efficiency of snake robots without passive wheels. Compared to Crawler gait and S-pedal gait, which

Cited by 0SourceScholar
2025

ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models

ACL 2025long

Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimodal Large Language Models (MLLMs), active perception has been largely overlooked.…

2025

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization

CVPR 2025poster

Recently, model merging methods have demonstrated powerful strengths in combining abilities on various tasks from multiple Large Language Models (LLMs). While previous model merging methods mainly focus on merging homogeneous models with identical architecture, they meet challenges when dealing with…

2025

Adversarial Robust Memory-Based Continual Learner

ICCV 2025poster

Despite the remarkable advances that have been made in continual learning, the adversarial vulnerability of such methods has not been fully discussed. We delve into the adversarial robustness of memory-based continual learning algorithms and observe limited robustness improvement by directly applyin…

2025

Align3R: Aligned Monocular Depth Estimation for Dynamic Videos

CVPR 2025highlight

Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video diffusion model to generate video depth conditioned on the i…

Cited by 14SourcePDFScholar
2025

Bench4Merge: A Comprehensive Benchmark for Merging in Realistic Dense Traffic with Micro-Interactive Vehicles

IROS 2025

While the capabilities of autonomous driving have advanced rapidly, merging into dense traffic remains a significant challenge, many motion planning methods for this scenario have been proposed but it is hard to evaluate them. Most existing closed-loop simulators rely on rule-based controls for othe

Cited by 0SourcecodeScholar
2025

Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations

NeurIPS 2025poster

The growing scale of evaluation tasks has led to the widespread adoption of automated evaluation using LLMs, a paradigm known as “LLM-as-a-judge”. However, improving its alignment with human preferences without complex prompts or fine-tuning remains challenging. Previous studies mainly optimize base…

Cited by 0SourceScholar
2025

CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models

CVPR 2025poster

Vision-Language Models (VLMs) have recently witnessed significant progress in visual comprehension. As the permitting length of image context grows, VLMs can now comprehend a broader range of views and spaces. Current benchmarks provide insightful analysis of VLMs in tasks involving complex visual i…

2025

Contrastive Private Data Synthesis via Weighted Multi-PLM Fusion

ICML 2025poster

Substantial quantity and high quality are the golden rules of making a good training dataset with sample privacy protection equally important. Generating synthetic samples that resemble high-quality private data while ensuring Differential Privacy (DP), a formal privacy guarantee, promises scalabili…

2025

DongbaMIE: A Multimodal Information Extraction Dataset for Evaluating Semantic Understanding of Dongba Pictograms

EMNLP 2025

Dongba pictographic is the only pictographic script still in use in the world. Its pictorial ideographic features carry rich cultural and contextual information. However, due to the lack of relevant datasets, research on semantic understanding of Dongba hieroglyphs has progressed slowly. To this end

2025

Dual-AEB: Synergizing Rule-Based and Multimodal Large Language Models for Effective Emergency Braking

ICRA 2025

Automatic Emergency Braking (AEB) systems are a crucial component in ensuring the safety of passengers in autonomous vehicles. Conventional AEB systems primarily rely on closed-set perception modules to recognize traffic conditions and assess collision risks. To enhance the adaptability of AEB syste

Cited by 3SourcecodeScholar
2025

Dynamic-static Feature Fusion with Multi-scale Attention for Continuous Blood Glucose Prediction

ICASSP 2025accepted

Accurate continuous blood glucose prediction is an effective and direct method for treating type 2 diabetes mellitus. However, current methods are commonly single-domain single-scale blood glucose prediction models. That is, they only learn time correlations within constant time steps of continuous…

Cited by 0SourceScholar
2025

Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective

ICLR 2025poster

Direct Preference Optimization (DPO) has gained attention as an efficient alternative to reinforcement learning from human feedback (RLHF) for aligning large language models (LLMs) with human preferences. Despite its advantages, DPO suffers from a length bias, generating responses longer than those…

2025

Enabling In-Flight Metamorphosis in Multirotors with a Center-Driven Scissor Extendable Airframe for Adaptive Navigation

ICRA 2025

To address complex mission tasks, multirotors benefit from in-flight reconfiguration that enhances their morphological adaptability. This paper presents the Center-Driven Scissor Extendable Airframe (CDSEA), a novel one-degree-of-freedom (DOF) morphing airframe designed to replace traditional fixed-

Cited by 0SourceScholar
2025

From Learning to Mastery: Achieving Safe and Efficient Real-World Autonomous Driving with Human-in-the-Loop Reinforcement Learning

IROS 2025

Autonomous driving with reinforcement learning (RL) has significant potential. However, applying RL in real-world settings remains challenging due to the need for safe, efficient, and robust learning. Incorporating human expertise into the learning process can help overcome these challenges by reduc

Cited by 0SourcecodeScholar
2025

G2: Guided Generation for Enhanced Output Diversity in LLMs

EMNLP 2025

Large Language Models (LLMs) have demonstrated exceptional performance across diverse natural language processing tasks. However, these models exhibit a critical limitation in output diversity, often generating highly similar content across multiple attempts. This limitation significantly affects ta

2025

Gradient-guided Attention Map Editing: Towards Efficient Contextual Hallucination Mitigation

NAACL 2025findings

In tasks such as summarization and open-book question answering (QA), Large Language Models (LLMs) frequently experience “contextual hallucination”, where they generate irrelevant or incorrect responses despite having access to accurate information in the input. This issue often stems from the model…

2025

Hard Sample Aware Robust Contrastive Learning for Multi-View Clustering

ICASSP 2025accepted

Multi-view clustering aims to divide samples into several clusters, by mining and utilizing the consistency and complementarity of multi-view data. Recent years, numerous deep contrastive multi-view clustering methods have been proposed to address the false negative issue by using self-supervised in…

Cited by 0SourceScholar
2025

How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in An Extensible Escape Game

ICCV 2025poster

The rapid advancing of Multimodal Large Language Models (MLLMs) has spurred interest in complex multimodal reasoning tasks in the real-world and virtual environment, which require coordinating multiple abilities, including visual perception, visual reasoning, spatial awareness, and target deduction.…

2025

LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents

ICCV 2025poster

Existing MLLMs encounter significant challenges in modeling the temporal context within long videos. Currently, mainstream Agent-based methods use external tools to assist a single MLLM in answering long video questions. Despite such tool-based support, a solitary MLLM still offers only a partial un…

2025

Leveraging Language-based Representations for Better Solving Symbol-related Problems with Large Language Models

COLING 2025main

Symbols such as numerical sequences, chemical formulas, and table delimiters exist widely, playing important roles in symbol-related tasks such as abstract reasoning, chemical property prediction, and tabular question-answering. Compared to tasks based on natural language expressions, large language…

2025

MASTER: A Multi-granularity Invariant Structure Clustering Scheme for Multi-view Clustering

IJCAI 2025

Deep multi-view clustering has attracted increasing attention in the pattern mining of data. However, most of them perform self-learning mechanisms in a single space, ignoring the fruitful structural information hidden in different-level feature spaces. Meanwhile, they conduct the reconstruction con

Cited by 0SourcePDFScholar
2025

MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models

EMNLP 2025

Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. Due to their strong performance in image-text alignment, MLLMs can effectively understand image-text pairs with clear meanings. However, effectively resolving the inherent ambiguiti

2025

Meta-Reinforcement Learning With Evolving Gradient Regularization

RA-L 2025

Deep reinforcement learning (DRL) typically requires reinitializing training for new tasks, limiting its generalization due to isolated knowledge transfer. Meta-reinforcement learning (Meta-RL) addresses this by enabling rapid adaptation through prior task experiences, yet existing gradient-based me

Cited by 1SourceScholar
2025

Multimodal Point Cloud Registration Method Based on Centerline-Guided Expansion and Contraction: An Optimization Strategy Applied in Bronchial Lumen Map Building

IROS 2025

In this work, a multimodal point cloud registration method using CT and video frames is proposed to optimize the modeling of the bronchial cavity environment. Preoperative CT data improve the quality of point clouds acquired from intraoperative video frames. Initially, preoperative CT scans are used

Cited by 0SourceScholar
2025

PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models

EMNLP 2025

Text Image Machine Translation (TIMT) aims to translate texts embedded within an image into another language. Current TIMT studies primarily focus on providing translations for all the text within an image, while neglecting to provide bounding boxes and covering limited scenarios. In this work, we e

2025

PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing

CVPR 2025poster

Photorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, existing methods for monocular full-body reconstruction, typically relying on front and/or predicted back view, still struggle with satisfactory performance due to the ill-posed nature o…

2025

Perspective Transition of Large Language Models for Solving Subjective Tasks

ACL 2025finding

Large language models (LLMs) have revolutionized the field of natural language processing, enabling remarkable progress in various tasks. Different from objective tasks such as commonsense reasoning and arithmetic question-answering, the performance of LLMs on subjective tasks is still limited, wher…

2025

Rethinking Long Context Generation from the Continual Learning Perspective

COLING 2025main

Due to the limited context window, Large Language Models (LLMs) struggle with processing long contexts. Although fine-tuning can extend the context window, it incurs substantial computation costs. In contrast, recent tuning-free approaches reallocate the attention mechanism or incorporate temporary…

Cited by 1SourcePDFScholar
2025

Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models

COLING 2025main

State-of-the-art Large Multi-Modal Models (LMMs) have demonstrated exceptional capabilities in vision-language tasks. Despite their advanced functionalities, the performances of LMMs are still limited in challenging scenarios that require complex reasoning with multiple levels of visual information.…

2025

Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency

IJCAI 2025

Speculative decoding accelerates Large Language Model (LLM) inference by employing a small speculative model (SSM) to generate multiple candidate tokens and verify them using the LLM in parallel. This technique has been widely integrated into LLM inference serving systems. However, inference request

Cited by 0SourcePDFScholar
2025

SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human Reconstruction

NeurIPS 2025poster

Photorealistic 3D full-body human reconstruction from a single image is a critical yet challenging task for applications in films and video games due to inherent ambiguities and severe self-occlusions. While recent approaches leverage SMPL estimation and SMPL-conditioned image generative models to h…

Cited by 0SourceScholar
2025

TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels

NeurIPS 2025poster

Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall short in separating the camera motion from foreground dynamic…

Cited by 0SourceScholar
2024

Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization

ACL 2024long

Large Language Models (LLMs) exhibit robust problem-solving capabilities for diverse tasks. However, most LLM-based agents are designed as specific task solvers with sophisticated prompt engineering, rather than agents capable of learning and evolving through interactions. These task solvers necessi…

2024

Amend to Alignment: Decoupled Prompt Tuning for Mitigating Spurious Correlation in Vision-Language Models

ICML 2024poster

Fine-tuning the learnable prompt for a pre-trained vision-language model (VLM), such as CLIP, has demonstrated exceptional efficiency in adapting to a broad range of downstream tasks. Existing prompt tuning methods for VLMs do not distinguish spurious features introduced by biased training data from…

Cited by 4SourcePDFScholar
2024

Browse and Concentrate: Comprehending Multimodal Content via Prior-LLM Context Fusion

ACL 2024long

With the bloom of Large Language Models (LLMs), Multimodal Large Language Models (MLLMs) that incorporate LLMs with pre-trained vision models have recently demonstrated impressive performance across diverse vision-language tasks. However, they fall short to comprehend context involving multiple imag…

2024

CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language Models

ACL 2024long

Multimodal large language models (MLLMs) have demonstrated promising results in a variety of tasks that combine vision and language. As these models become more integral to research and applications, conducting comprehensive evaluations of their capabilities has grown increasingly important. However…

Cited by 8SourcePDFScholar
2024

DEEM: Dynamic Experienced Expert Modeling for Stance Detection

COLING 2024main

Recent work has made a preliminary attempt to use large language models (LLMs) to solve the stance detection task, showing promising results. However, considering that stance detection usually requires detailed background knowledge, the vanilla reasoning method may neglect the domain knowledge to ma…

2024

EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models

CVPR 2024highlight

Vision-language models (VLMs) have recently shown promising results in traditional downstream tasks. Evaluation studies have emerged to assess their abilities with the majority focusing on the third-person perspective and only a few addressing specific tasks from the first-person perspective. Howeve…

2024

Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages

ACL 2024long

While large language models (LLMs) have been pre-trained on multilingual corpora, their performance still lags behind in most languages compared to a few resource-rich languages. One common approach to mitigate this issue is to translate training data from resource-rich languages into other language…

2024

Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention

NeurIPS 2024poster

In this paper, we introduce **Era3D**, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resu…

Cited by 7SourcePDFScholar
2024

FuseGen: PLM Fusion for Data-generation based Zero-shot Learning

EMNLP 2024main

Data-generation based zero-shot learning, although effective in training Small Task-specific Models (STMs) via synthetic datasets generated by Pre-trained Language Models (PLMs), is often limited by the low quality of such synthetic datasets. Previous solutions have primarily focused on single PLM s…

2024

High-Dimensional Bayesian Optimization via Semi-Supervised Learning with Optimized Unlabeled Data Sampling

ICML 2024spotlight

We introduce a novel semi-supervised learning approach, named Teacher-Student Bayesian Optimization ($\texttt{TSBO}$), integrating the teacher-student paradigm into BO to minimize expensive labeled data queries for the first time. $\texttt{TSBO}$ incorporates a teacher model, an unlabeled data sampl…

2024

Human-Robot Interactive Creation of Artistic Portrait Drawings

ICRA 2024poster

In this paper, we present a novel system for Human-Robot Interactive Creation of Artworks (HRICA). Different from previous robot painters, HRICA allows a human user and a robot to alternately draw strokes on a canvas, to collaboratively create a portrait drawing through frequent interactions. The ke…

Cited by 0SourcecodeScholar
2024

Model Composition for Multimodal Large Language Models

ACL 2024long

Recent developments in Multimodal Large Language Models (MLLMs) have shown rapid progress, moving towards the goal of creating versatile MLLMs that understand inputs from various modalities. However, existing methods typically rely on joint training with paired multimodal instruction data, which is…

2024

Multi-agent Reinforcement Learning with Hybrid Action Space for Free Gait Motion Planning of Hexapod Robots

CoRL 2024poster

Legged robots are able to overcome challenging terrains through diverse gaits formed by contact sequences. However, environments characterized by discrete footholds present significant challenges. In this paper, we tackle the problem of free gait motion planning for hexapod robots walking in randoml…

Cited by 0SourceScholar
2024

Offline Meta-Reinforcement Learning with Evolving Gradient Agreement

IROS 2024poster

Meta-Reinforcement Learning (Meta-RL) is a machine learning paradigm aimed at learning reinforcement learning policies that can quickly adapt to unseen tasks with few-shot data. Nevertheless, applying Meta-RL to real-world applications faces challenges due to the cost of data acquisition. To address…

Cited by 0SourceScholar
2024

On Learning Discriminative Features from Synthesized Data for Self-Supervised Fine-Grained Visual Recognition

ECCV 2024poster

"Self-Supervised Learning (SSL) has become a prominent approach for acquiring visual representations across various tasks, yet its application in fine-grained visual recognition (FGVR) is challenged by the intricate task of distinguishing subtle differences between categories. To overcome this, we i…

Cited by 2SourcePDFScholar
2024

PANDA: Preference Adaptation for Enhancing Domain-Specific Abilities of LLMs

ACL 2024findings

While Large language models (LLMs) have demonstrated considerable capabilities across various natural language tasks, they often fall short of the performance achieved by domain-specific state-of-the-art models. One potential approach to enhance domain-specific capabilities of LLMs involves fine-tun…

2024

Pluggable Neural Machine Translation Models via Memory-augmented Adapters

COLING 2024main

Although neural machine translation (NMT) models perform well in the general domain, it remains rather challenging to control their generation behavior to satisfy the requirement of different users. Given the expensive training cost and the data scarcity challenge of learning a new model from scratc…

2024

Position: Towards Unified Alignment Between Agents, Humans, and Environment

ICML 2024poster

The rapid progress of foundation models has led to the prosperity of autonomous agents, which leverage the universal capabilities of foundation models to conduct reasoning, decision-making, and environmental interaction. However, the efficacy of agents remains limited when operating in intricate, re…

Cited by 4SourcePDFScholar
2024

Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models

ACL 2024long

Large Language Models (LLMs) have achieved remarkable performance in objective tasks such as open-domain question answering and mathematical reasoning, which can often be solved through recalling learned factual knowledge or chain-of-thought style reasoning. However, we find that the performance of…

2024

Semi-supervised Learning of Dynamical Systems with Neural Ordinary Differential Equations: A Teacher-Student Model Approach

AAAI 2024technical

Modeling dynamical systems is crucial for a wide range of tasks, but it remains challenging due to complex nonlinear dynamics, limited observations, or lack of prior knowledge. Recently, data-driven approaches such as Neural Ordinary Differential Equations (NODE) have shown promising results by leve…

Cited by 1SourcePDFScholar
2024

Set Prediction Guided by Semantic Concepts for Diverse Video Captioning

AAAI 2024technical

Diverse video captioning aims to generate a set of sentences to describe the given video in various aspects. Mainstream methods are trained with independent pairs of a video and a caption from its ground-truth set without exploiting the intra-set relationship, resulting in low diversity of generated…

Cited by 3SourcePDFScholar
2024

StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

ACL 2024findings

Large Language Models (LLMs) have witnessed remarkable advancements in recent years, prompting the exploration of tool learning, which integrates LLMs with external tools to address diverse real-world challenges. Assessing the capability of LLMs to utilise tools necessitates large-scale and stable b…

Cited by 36SourcePDFScholar
2024

Statler: State-Maintaining Language Models for Embodied Reasoning

ICRA 2024poster

There has been a significant research interest in employing large language models to empower intelligent robots with complex reasoning. Existing work focuses on harnessing their abilities to reason about the histories of their actions and observations. In this paper, we explore a new dimension in wh…

Cited by 41SourcecodeScholar
2024

ToolRerank: Adaptive and Hierarchy-Aware Reranking for Tool Retrieval

COLING 2024main

Tool learning aims to extend the capabilities of large language models (LLMs) with external tools. A major challenge in tool learning is how to support a large number of tools, including unseen tools. To address this challenge, previous studies have proposed retrieving suitable tools for the LLM bas…

2024

Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-speech Gesture Generation

CVPR 2024poster

Generating vivid and emotional 3D co-speech gestures is crucial for virtual avatar animation in human-machine interaction applications. While the existing methods enable generating the gestures to follow a single emotion label they overlook that long gesture sequence modeling with emotion transition…

2023

AHPA: Adaptive Horizontal Pod Autoscaling Systems on Alibaba Cloud Container Service for Kubernetes

AAAI 2023technical

The existing resource allocation policy for application instances in Kubernetes cannot dynamically adjust according to the requirement of business, which would cause an enormous waste of resources during fluctuations. Moreover, the emergence of new cloud services puts higher resource management requ…

Cited by 16SourcePDFScholar
2023

An Extensible Plug-and-Play Method for Multi-Aspect Controllable Text Generation

ACL 2023long

Recently, multi-aspect controllable text generation that controls the generated text in multiple aspects (e.g., sentiment, topic, and keywords) has attracted increasing attention. Although methods based on parameter efficient tuning like prefix-tuning could achieve multi-aspect controlling in a plug…

2023

AutoNF: Automated Architecture Optimization of Normalizing Flows with Unconstrained Continuous Relaxation Admitting Optimal Discrete Solution

AAAI 2023technical

Normalizing flows (NF) build upon invertible neural networks and have wide applications in probabilistic modeling. Currently, building a powerful yet computationally efficient flow model relies on empirical fine-tuning over a large design space. While introducing neural architecture search (NAS) to…

Cited by 2SourcePDFScholar
2023

Bridging the Gap between Decision and Logits in Decision-based Knowledge Distillation for Pre-trained Language Models

ACL 2023long

Conventional knowledge distillation (KD) methods require access to the internal information of teachers, e.g., logits. However, such information may not always be accessible for large pre-trained language models (PLMs). In this work, we focus on decision-based KD for PLMs, where only teacher decisio…

2023

CodeIE: Large Code Generation Models are Better Few-Shot Information Extractors

ACL 2023long

Large language models (LLMs) pre-trained on massive corpora have demonstrated impressive few-shot learning ability on many NLP tasks. A common practice is to recast the task into a text-to-text format such that generative LLMs of natural language (NL-LLMs) like GPT-3 can be prompted to solve it. How…

2023

Continual Knowledge Distillation for Neural Machine Translation

ACL 2023long

While many parallel corpora are not publicly accessible for data copyright, data privacy and competitive differentiation reasons, trained translation models are increasingly available on open platforms. In this work, we propose a method called continual knowledge distillation to take advantage of ex…

2023

Efficient and Hybrid Decoder for Local Map Construction in Bird'-Eye-View

ICRA 2023poster

High-definition maps are crucial perception elements for autonomous robot navigation systems, which can provide accurate scene layout and environment information for downstream motion prediction and planning control tasks. Traditional methods based on manual annotation or SLAM algorithms require mas…

Cited by 1SourceScholar
2023

Exploiting Contextual Objects and Relations for 3D Visual Grounding

NeurIPS 2023poster

3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, this task is challenging due to the necessity to capture 3D contextual information…

2023

Failures Pave the Way: Enhancing Large Language Models through Tuning-free Rule Accumulation

EMNLP 2023long main

Large Language Models (LLMs) have showcased impressive performance. However, due to their inability to capture relationships among samples, these frozen LLMs inevitably keep repeating similar mistakes. In this work, we propose our Tuning-free Rule Accumulation (TRAN) framework, which guides LLMs in…

Cited by 0SourcecodeScholar
2023

Filling the Image Information Gap for VQA: Prompting Large Language Models to Proactively Ask Questions

EMNLP 2023long findings

Large Language Models (LLMs) demonstrate impressive reasoning ability and the maintenance of world knowledge not only in natural language tasks, but also in some vision-language tasks such as open-domain knowledge-based visual question answering (OK-VQA). As images are invisible to LLMs, researchers…

Cited by 0SourcecodeScholar
2023

Geogcn: Geometric Dual-Domain Graph Convolution Network For Point Cloud Denoising

ICASSP 2023accepted

We propose GeoGCN, a novel geometric dual-domain graph convolution network for point cloud denoising (PCD). Beyond the traditional wisdom of PCD, to fully exploit the geometric information of point clouds, we define two kinds of surface normals, one is called Real Normal (RN), and the other is Virtu…

Cited by 0SourceScholar
2023

ISmallNet: Densely Nested Network with Label Decoupling for Infrared Small Target Detection

ICASSP 2023accepted

Small targets are often submerged in cluttered backgrounds of infrared images. Conventional detectors tend to generate false alarms, while CNN-based detectors lose small targets in deep layers. To this end, we propose iSmallNet, a multi-stream densely nested network with label decoupling for infrare…

Cited by 0SourceScholar
2023

Improving Adversarial Robustness of Deep Equilibrium Models with Explicit Regulations Along the Neural Dynamics

ICML 2023poster

Deep equilibrium (DEQ) models replace the multiple-layer stacking of conventional deep networks with a fixed-point iteration of a single-layer transformation. Having been demonstrated to be competitive in a variety of real-world scenarios, the adversarial robustness of general DEQs becomes increasin…

Cited by 8SourcePDFScholar
2023

Knowledge Transfer in Incremental Learning for Multilingual Neural Machine Translation

ACL 2023long

In the real-world scenario, a longstanding goal of multilingual neural machine translation (MNMT) is that a single model can incrementally adapt to new language pairs without accessing previous training data. In this scenario, previous studies concentrate on overcoming catastrophic forgetting while…

2023

Learn and Consolidate: Continual Adaptation for Zero-Shot and Multilingual Neural Machine Translation

EMNLP 2023long main

Although existing multilingual neural machine translation (MNMT) models have demonstrated remarkable performance to handle multiple translation directions in a single model and achieved zero-shot translation between language pairs unseen in training, they still suffer from relatively poor translatio…

Cited by 0SourceScholar
2023

Learning Visibility Field for Detailed 3D Human Reconstruction and Relighting

CVPR 2023poster

Detailed 3D reconstruction and photo-realistic relighting of digital humans are essential for various applications. To this end, we propose a novel sparse-view 3d human reconstruction framework that closely incorporates the occupancy field and albedo field with an additional visibility field--it not…

Cited by 19SourcePDFScholar
2023

Plug-and-Play Knowledge Injection for Pre-trained Language Models

ACL 2023long

Injecting external knowledge can improve the performance of pre-trained language models (PLMs) on various downstream NLP tasks. However, massive retraining is required to deploy new knowledge injection methods or knowledge bases for downstream tasks. In this work, we are the first to study how to im…

2023

Prompt-Guided Retrieval Augmentation for Non-Knowledge-Intensive Tasks

ACL 2023findings

Retrieval-augmented methods have received increasing attention to support downstream tasks by leveraging useful information from external resources. Recent studies mainly focus on exploring retrieval to solve knowledge-intensive (KI) tasks. However, the potential of retrieval for most non-knowledge-…

2023

Real-Time Whole-Body Collision Avoidance and Path Following of a Snake Robot Through MPC-based Optimization Strategies

IROS 2023poster

The work in this paper delves into the challenge of whole elongated body's obstacle avoidance during path following for a class of bionic snake robots. Currently, most studies focus solely on preventing the robot's head from colliding with obstacles through designed controllers. However, due to the…

Cited by 5SourceScholar
2023

Self-Knowledge Guided Retrieval Augmentation for Large Language Models

EMNLP 2023long findings

Large language models (LLMs) have shown superior performance without task-specific fine-tuning. Despite the success, the knowledge stored in the parameters of LLMs could still be incomplete and difficult to update due to the computational costs. As complementary, retrieval-based methods can offer no…

Cited by 0SourceScholar
2023

Tightly-Coupled Visual-DVL Fusion For Accurate Localization of Underwater Robots

IROS 2023poster

This paper proposes a tightly-coupled visual-Doppler-Velocity-Log (visual-DVL) fusion method for underwater robot localization through integrating the velocity measurements from a DVL into a visual odometry (VO). Considering that employing the DVL measurements in dead-reckoning systems easily leads…

Cited by 4SourceScholar
2023

TrOMR:Transformer-Based Polyphonic Optical Music Recognition

ICASSP 2023accepted

Optical Music Recognition (OMR) is an important technology in music and has been researched for a long time. Previous approaches for OMR are usually based on CNN for image understanding and RNN for music symbol classification. In this paper, we propose a transformer-based approach with excellent glo…

Cited by 0SourceScholar
2023

Unified Detoxifying and Debiasing in Language Generation via Inference-time Adaptive Optimization

ICLR 2023poster

Recently pre-trained language models (PLMs) have prospered in various natural language generation (NLG) tasks due to their ability to generate fairly fluent text. Nevertheless, these models are observed to capture and reproduce harmful contents in training corpora, typically toxic language and socia…

Cited by 38SourcePDFScholar
2023

Weakly Supervised Vision-and-Language Pre-training with Relative Representations

ACL 2023long

Weakly supervised vision-and-language pre-training (WVLP), which learns cross-modal representations with limited cross-modal supervision, has been shown to effectively reduce the data cost of pre-training while maintaining decent performance on downstream tasks. However, current WVLP methods use onl…

Cited by 2SourcePDFScholar
2023

ifUNet++: Iterative Feedback UNet++ for Infrared Small Target Detection

ICASSP 2023accepted

Small targets are often submerged in the cluttered backgrounds of infrared images. In this paper, we propose an iterative feedback UNet++ for infrared small target detection, dubbed ifUNet++. Unlike most of existing methods, ifU-Net++ enables to concentrate on small targets while weakening the inter…

Cited by 0SourceScholar
2022

A Closed-Loop Perception, Decision-Making and Reasoning Mechanism for Human-Like Navigation

IJCAI 2022poster

Reliable navigation systems have a wide range of applications in robotics and autonomous driving. Current approaches employ an open-loop process that converts sensor inputs directly into actions. However, these open-loop schemes are challenging to handle complex and dynamic real-world scenarios due…

2022

A Simple but Effective Pluggable Entity Lookup Table for Pre-trained Language Models

ACL 2022short

Pre-trained language models (PLMs) cannot well recall rich factual knowledge of entities exhibited in large-scale corpora, especially those rare entities. In this paper, we propose to build a simple but effective Pluggable Entity Lookup Table (PELT) on demand by aggregating the entity’s output repre…

2022

A Template-based Method for Constrained Neural Machine Translation

EMNLP 2022main

Machine translation systems are expected to cope with various types of constraints in many practical scenarios. While neural machine translation (NMT) has achieved strong performance in unconstrained cases, it is non-trivial to impose pre-specified constraints into the translation process of NMT mod…

2022

CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text Generation

ACL 2022long

Existing reference-free metrics have obvious limitations for evaluating controlled text generation models. Unsupervised metrics can only provide a task-agnostic evaluation result which correlates weakly with human judgments, whereas supervised ones may overfit task-specific data with poor generaliza…

2022

Do Pre-trained Models Benefit Knowledge Graph Completion? A Reliable Evaluation and a Reasonable Approach

ACL 2022findings

In recent years, pre-trained language models (PLMs) have been shown to capture factual knowledge from massive texts, which encourages the proposal of PLM-based knowledge graph completion (KGC) models. However, these models are still quite behind the SOTA KGC models in terms of performance. In this w…

2022

ELLE: Efficient Lifelong Pre-training for Emerging Data

ACL 2022findings

Current pre-trained language models (PLM) are typically trained with static data, ignoring that in real-world scenarios, streaming data of various sources may continuously grow. This requires PLMs to integrate the information from all the sources in a lifelong manner. Although this goal could be ach…

2022

End-to-End Unsupervised Vision-and-Language Pre-training with Referring Expression Matching

EMNLP 2022main

Recently there has been an emerging interest in unsupervised vision-and-language pre-training (VLP) that learns multimodal representations without parallel image-caption data. These pioneering works significantly reduce the cost of VLP on data collection and achieve promising results compared to sup…

Cited by 6SourcePDFScholar
2022

Entropy-Based Vocabulary Substitution for Incremental Learning in Multilingual Neural Machine Translation

EMNLP 2022main

In a practical real-world scenario, the longstanding goal is that a universal multilingual translation model can be incrementally updated when new language pairs arrive. Specifically, the initial vocabulary only covers some of the words in new languages, which hurts the translation quality for incre…

2022

Event-Triggered Tracking Control Scheme for Quadrotors with External Disturbances: Theory and Validations

ICRA 2022poster

This article studies the tracking control of a quadrotor unmanned aerial vehicle (UAV) under time-varying external disturbances. An event-triggered sliding mode control (SMC) strategy is proposed by introducing a new triggering condition form of desired trajectory, quadrotor position, and velocity.…

Cited by 6SourceScholar
2022

From Mimicking to Integrating: Knowledge Integration for Pre-Trained Language Models

EMNLP 2022finding

Investigating better ways to reuse the released pre-trained language models (PLMs) can significantly reduce the computational cost and the potential environmental side-effects. This paper explores a novel PLM reuse paradigm, Knowledge Integration (KI). Without human annotations available, KI aims to…

2022

I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection

AAAI 2022technical

Can you find me? By simulating how humans to discover the so-called 'perfectly'-camouflaged object, we present a novel boundary-guided separated attention network (call BSA-Net). Beyond the existing camouflaged object detection (COD) wisdom, BSA-Net utilizes two-stream separated attention modules to…

2022

Knowledge Inheritance for Pre-trained Language Models

NAACL 2022long

Recent explorations of large-scale pre-trained language models (PLMs) have revealed the power of PLMs with huge amounts of parameters, setting off a wave of training ever-larger PLMs. However, it requires tremendous computational resources to train a large-scale PLM, which may be practically unaffor…

2022

LF-VIO: A Visual-Inertial-Odometry Framework for Large Field-of-View Cameras with Negative Plane

IROS 2022poster

Visual-inertial-odometry has attracted extensive attention in the field of autonomous driving and robotics. The size of Field of View (FoV) plays an important role in Visual-Odometry (VO) and Visual-Inertial-Odometry (VIO), as a large FoV enables to perceive a wide range of surrounding scene element…

Cited by 21SourcecodeScholar
2022

LightPose: A Lightweight and Efficient Model with Transformer for Human Pose Estimation

ICASSP 2022accepted

The prediction of keypoints by generating high-resolution heatmaps has become a popular solution in human pose estimation. While this kind of method requires up-sampling or deconvolution operations, which would bring a great challenge to the acceleration of model inference. If performing keypoint pr…

Cited by 0SourceScholar
2022

MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction

EMNLP 2022main

The diverse relationships among real-world events, including coreference, temporal, causal, and subevent relations, are fundamental to understanding natural languages. However, two drawbacks of existing datasets limit event relation extraction (ERE) tasks: (1) Small scale. Due to the annotation comp…

2022

MoEfication: Transformer Feed-forward Layers are Mixtures of Experts

ACL 2022findings

Recent work has shown that feed-forward networks (FFNs) in pre-trained Transformers are a key component, storing various linguistic and factual knowledge. However, the computational patterns of FFNs are still unclear. In this work, we study the computational patterns of FFNs and observe that most in…

2022

On Transferability of Prompt Tuning for Natural Language Processing

NAACL 2022long

Prompt tuning (PT) is a promising parameter-efficient method to utilize extremely large pre-trained language models (PLMs), which can achieve comparable performance to full-parameter fine-tuning by only tuning a few soft prompts. However, PT requires much more training time than fine-tuning. Intuiti…

2022

ROSE: Robust Selective Fine-tuning for Pre-trained Language Models

EMNLP 2022main

Even though the large-scale language models have achieved excellent performances, they suffer from various adversarial attacks.A large body of defense methods has been proposed. However, they are still limited due to redundant attack search spaces and the inability to defend against various types of…

2022

Rethinking the Promotion Brought by Contrastive Learning to Semi-Supervised Node Classification

IJCAI 2022poster

Graph Contrastive Learning (GCL) has proven highly effective in promoting the performance of Semi-Supervised Node Classification (SSNC). However, existing GCL methods are generally transferred from other fields like CV or NLP, whose underlying working mechanism remains underexplored. In this work, w…

Cited by 5SourcePDFScholar
2022

Unsupervised Dependency Graph Network

ACL 2022long

Recent work has identified properties of pretrained self-attention models that mirror those of dependency parse structures. In particular, some self-attention heads correspond well to individual dependency types. Inspired by these developments, we propose a new competitive mechanism that encourages…

2022

Variable-Friction-Based In-Hand Manipulation of Fabrics Applied to Unfolding Operations

RA-L 2022

The automatic state recognition and handling of fabrics, which are typical deformable objects, is difficult to realize. The development of related technologies for fabric manipulation remains impractical. This research was inspired by the process by which humans handle fabrics using their fingers. W

Cited by 3SourceScholar
2021

Aspect-Level Sentiment-Controllable Review Generation with Mutual Learning Framework

AAAI 2021technical

Review generation, aiming to automatically generate review text according to the given information, is proposed to assist in the unappealing review writing. However, most of existing methods only consider the overall sentiments of reviews and cannot achieve aspect-level sentiment control. Even thoug…

Cited by 10SourcePDFScholar
2021

Backpropagated Neighborhood Aggregation for Accurate Training of Spiking Neural Networks

ICML 2021spotlight

While Backpropagation (BP) has been applied to spiking neural networks (SNNs) achieving encouraging results, a key challenge involved is to backpropagate a differentiable continuous-valued loss over layers of spiking neurons exhibiting discontinuous all-or-none firing activities. Existing methods de…

Cited by 25SourcePDFScholar
2021

CLEVE: Contrastive Pre-training for Event Extraction

ACL 2021long

Event extraction (EE) has considerably benefited from pre-trained language models (PLMs) by fine-tuning. However, existing pre-training methods have not involved modeling event characteristics, resulting in the developed EE models cannot take full advantage of large-scale unsupervised data. To this…

2021

CascadeBERT: Accelerating Inference of Pre-trained Language Models via Calibrated Complete Models Cascade

EMNLP 2021finding

Dynamic early exiting aims to accelerate the inference of pre-trained language models (PLMs) by emitting predictions in internal layers without passing through the entire model. In this paper, we empirically analyze the working mechanism of dynamic early exiting and find that it faces a performance…

2021

CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the Wild

EMNLP 2021main

Existing relation extraction (RE) methods typically focus on extracting relational facts between entity pairs within single sentences or documents. However, a large quantity of relational facts in knowledge bases can only be inferred across documents in practice. In this work, we present the problem…

2021

Context Tracking Network: Graph-based Context Modeling for Implicit Discourse Relation Recognition

NAACL 2021long

Implicit discourse relation recognition (IDRR) aims to identify logical relations between two adjacent sentences in the discourse. Existing models fail to fully utilize the contextual information which plays an important role in interpreting each local sentence. In this paper, we thus propose a nove…

Cited by 26SourcePDFScholar
2021

Deep Reinforcement Learning for Multi-contact Motion Planning of Hexapod Robots

IJCAI 2021poster

Legged locomotion in a complex environment requires careful planning of the footholds of legged robots. In this paper, a novel Deep Reinforcement Learning (DRL) method is proposed to implement multi-contact motion planning for hexapod robots moving on uneven plum-blossom piles. First, the motion of…

Cited by 15SourcePDFScholar
2021

Dynamic Knowledge Distillation for Pre-trained Language Models

EMNLP 2021main

Knowledge distillation (KD) has been proved effective for compressing large-scale pre-trained language models. However, existing methods conduct KD statically, e.g., the student model aligns its output distribution to that of a selected teacher model on the pre-defined training dataset. In this pape…

2021

ERICA: Improving Entity and Relation Understanding for Pre-trained Language Models via Contrastive Learning

ACL 2021long

Pre-trained Language Models (PLMs) have shown superior performance on various downstream Natural Language Processing (NLP) tasks. However, conventional pre-training objectives do not explicitly model relational facts in text, which are crucial for textual understanding. To address this issue, we pro…

2021

Guiding Non-Autoregressive Neural Machine Translation Decoding with Reordering Information

AAAI 2021technical

Non-autoregressive neural machine translation (NAT) generates each target word in parallel and has achieved promising inference acceleration. However, existing NAT models still have a big gap in translation quality compared to autoregressive neural machine translation models due to the multimodality…

2021

Learning to Navigate in a VUCA Environment: Hierarchical Multi-expert Approach

IROS 2021poster

Despite decades of efforts, robot navigation in a real scenario with volatility, uncertainty, complexity, and ambiguity (VUCA for short), remains a challenging topic. Inspired by the central nervous system (CNS), we propose a hierarchical multi-expert learning framework for autonomous navigation in…

Cited by 9SourceScholar
2021

RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models

EMNLP 2021main

Backdoor attacks, which maliciously control a well-trained model’s outputs of the instances with specific triggers, are recently shown to be serious threats to the safety of reusing deep neural networks (DNNs). In this work, we propose an efficient online defense mechanism based on robustness-aware…

2021

Rethinking Stealthiness of Backdoor Attack against NLP Models

ACL 2021long

Recent researches have shown that large natural language processing (NLP) models are vulnerable to a kind of security threat called the Backdoor Attack. Backdoor attacked models can achieve good performance on clean test sets but perform badly on those input sentences injected with designed trigger…

2021

Topology-Imbalance Learning for Semi-Supervised Node Classification

NeurIPS 2021poster

The class imbalance problem, as an important issue in learning node representations, has drawn increasing attention from the community. Although the imbalance considered by existing studies roots from the unequal quantity of labeled examples in different classes (quantity imbalance), we argue that g…

2020

Distributed Consensus Control of Multiple UAVs in a Constrained Environment

ICRA 2020poster

In this paper, we investigate the consensus problem of multiple unmanned aerial vehicles (UAVs) in the presence of environmental constraints under a general communication topology containing a directed spanning tree. First, based on a position transformation function, we propose a novel dynamic refe…

Cited by 15SourceScholar
2020

SNIAE-SSE Deformation Mechanism Enabled Scalable Multicopter: Design, Modeling and Flight Performance Validation

ICRA 2020poster

This paper focuses on designing, modeling and validating a novel scalable multicopter whose deformation mechanism, called SNIAE-SSE, relies on a combination of simple non-intersecting angulated elements (SNIAEs) and straight scissor-like elements (SSEs). The proposed SNIAE-SSE mechanism has the adva…

Cited by 14SourceScholar
2020

Temporal Spike Sequence Learning via Backpropagation for Deep Spiking Neural Networks

NeurIPS 2020spotlight

Spiking neural networks (SNNs) are well suited for spatio-temporal learning and implementations on energy-efficient event-driven neuromorphic processors. However, existing SNN error backpropagation (BP) methods lack proper handling of spiking discontinuities and suffer from low performance compared…

2019

Adaptive Vision-Based Control for Rope-Climbing Robot Manipulator

IROS 2019poster

While the mechanism of Rope-Climbing provides much flexibility, it opens up challenges to the development of the controller for Robotic Manipulator installed on Rope-Climbing robot(RCR), which is called Rope-Climbing Robot Manipulator(RCRM) here. In particular, the deformable nature of the rope resu…

Cited by 7SourceScholar
2019

An Approximation-Free Simple Control Scheme for Uncertain Quadrotor Systems: Theory and Validations

IROS 2019poster

In this paper, a simple tracking control scheme is proposed for quadrotor systems with uncertain dynamics. It precludes the necessity for prohibitive analytic computation of the derivatives of the desired (virtual) attitude that is typically employed in controlling quadrotor systems. Moreover, this…

Cited by 3SourceScholar
2019

Development of an Autonomous Sanding Robot with Structured-Light Technology

IROS 2019poster

Large demand for robotics and automation has been reflected in the sanding works, as current manual operations are labor-intensive, without consistent quality, and also subject to safety and health issues. While several machines have been developed to automate one or two steps in the sanding works,…

Cited by 8SourceScholar
2019

Self-modeling Tracking Control of Crawler Fire Fighting Robot Based on Causal Network

IROS 2019poster

In this paper, a self-modeling method based on a causal network is proposed for the tracking control of the Crawler Fire Fighting Robot (CFFR). The method mainly consists of two parts, one is a motion model, based on data driving, learning to establish the correspondence between control signal seque…

Cited by 0SourceScholar
2019

Spike-Train Level Backpropagation for Training Deep Recurrent Spiking Neural Networks

NeurIPS 2019poster

Spiking neural networks (SNNs) well support spatiotemporal learning and energy-efficient event-driven hardware neuromorphic processors. As an important class of SNNs, recurrent spiking neural networks (RSNNs) possess great computational power. However, the practical application of RSNNs is severely…

2018

A Synchronization Scheme for Position Control of Multiple Rope-Climbing Robots

ICRA 2018poster

The ability of rope-climbing robots in aloft operation is limited by its self-supporting and locomotion ability. In many applications, a given task is also too complex to be achieved by a single rope-climbing robot acting alone. The solution of multiple rope-climbing robots can overcome the limitati…

Cited by 5SourceScholar
2018

Hybrid Macro/Micro Level Backpropagation for Training Deep Spiking Neural Networks

NeurIPS 2018poster

Spiking neural networks (SNNs) are positioned to enable spatio-temporal information processing and ultra-low power event-driven neuromorphic hardware. However, SNNs are yet to reach the same performances of conventional deep artificial neural networks (ANNs), a long-standing challenge due to complex…

2016

Accelerating stochastic computation for binary classification applications

ICASSP 2016accepted

Stochastic computation is a non-conventional computation paradigm, which uses digital circuits to operate on stochastic bit streams. Although it has advantages such as strong fault tolerance and low hardware cost, its drawback is its long computation time. In this work, we target at stochastic compu…

Cited by 0SourceScholar
2015

A new robotic uterine positioner for laparoscopic hysterectomy with passive safety mechanisms: Design and experiments

IROS 2015poster

In this paper, we present a new robotic uterine positioner for total laparoscopic hysterectomy. The robot is designed to actively position the patient's uterus during surgery, a lengthy and tedious task that is traditionally performed by a human assistant. Safety is simply the most important concern…

Cited by 28SourceScholar
2015

Adaptive image-based positioning of RCM mechanisms using angle and distance features

IROS 2015poster

In this paper, we address the positioning problem of remote centre of motion (RCM) mechanisms with uncalibrated image feedback from a monocular camera. Nowadays, RCM mechanisms are widely used in minimally invasive robotic surgery due to their ability to distally rotate a tool around a fixed entry p…

Cited by 7SourceScholar
2015

Design and control of a novel multi-state compliant safe joint for robotic surgery

ICRA 2015poster

In this paper, we propose a novel design of compliant safe joint, which has flexibility when the work load exceeds a predefined threshold. The compliance is generated by a spring. We design a special transmission mechanism to convert axial motion into circumferential motion such that the linear comp…

Cited by 5SourceScholar
2015

Modeling, design and control of an endoscope manipulator for FESS

IROS 2015poster

This paper presents the development of an endoscope manipulator with passive and active structures for functional endoscopic sinus surgery (FESS). The 5-DoF passive structure has three translations and two rotations (T3R2) that allows the surgeon to manually place the endoscope near to the entry poi…

Cited by 19SourceScholar