← Search

Yu Sun

144 accepted papers

2026

A Cable-Driven Soft Robotic Hand with an In-Hand RGB-D Camera for Dexterous Grasping and Manipulation

ICRA 2026poster

The aspiration to replicate the capabilities of the human hand has driven innovations in the design of soft robotic hands. Despite these advancements, many existing designs of soft hands still lack effective in-hand vision and the ability for each finger to achieve active multidegree-of-freedom moti…

Cited by 1Scholar
2026

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

CVPR 2026

Multimodal large language models (MLLMs) have achieved remarkable progress on various vision-language tasks, yet their visual perception remains limited. Humans, in comparison, perceive complex scenes efficiently by dynamically scanning and focusing on salient regions in a sequential "blink-like" pr

Cited by 0SourceScholar
2026

DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter

ICML 2026poster

Bimanual dexterous manipulation relies on integrating multimodal inputs to perform complex real-world tasks. To address the challenges of effectively combining these modalities, we propose DECO, a decoupled multimodal diffusion transformer that disentangles vision, proprioception, and tactile signal…

Cited by 0SourceScholar
2026

Fine-Grained Classification for Depth Estimation From Monocular Microscopy for Robotic Micromanipulation of Motile Cells

RA-L 2026

Manipulation of motile cells is crucial for biological research and clinical applications. However, obtaining Z-axis visual feedback under monocular microscopy remains a challenge for robotic micromanipulation. Traditional depth-from-focus and depth-from-defocus methods fail to handle motile cells d

Cited by 0SourceScholar
2026

K-Track: A Kirigami-Inspired Tracked Robot With Negative Pressure Co-operative Adhesion for Wall-to-Wall Transition

RA-L 2026

This letter presents a novel wall-climbing robot design that combines kirigami-inspired adhesive tracks with negative pressure absorption, enabling reliable wall-to-wall transitions in complex environments. The kirigami-inspired tracks demonstrate a significant enhancement in adhesion performance, a

Cited by 0SourceScholar
2026

Learning to Discover at Test Time

ICML 2026spotlight

How can we use AI to discover a new state of the art for a scientific problem? Prior work in test-time scaling, such as AlphaEvolve, performs search by prompting a frozen LLM. We perform reinforcement learning at test time, so the LLM can continue to train, but now with experience specific to the te…

Cited by 0SourceScholar
2026

Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models

CVPR 2026

Recently, unified multimodal models (UMMs) have made remarkable progress in integrating visual understanding and generation, demonstrating strong potential for complex text-to-image (T2I) tasks. Despite their theoretical promise, a persistent capability gap exists: UMMs typically exhibit superior vi

Cited by 0SourcecodeScholar
2026

PRISM: PROBABILISTIC AND ROBUST INVERSE SOLVER WITH MEASUREMENT-CONDITIONED DIFFUSION PRIOR FOR BLIND INVERSE PROBLEMS

ICASSP 2026oral

Diffusion models are now commonly used to solve inverse problems in computational imaging. However, most diffusion-based inverse solvers require complete knowledge of the forward operator to be used. In this work, we introduce a novel probabilistic and robust inverse solver with measurement-conditio…

Cited by 0SourcePDFScholar
2026

Provably Accelerated Imaging with Restarted Inertia and Score-based Image Priors

ICLR 2026poster

Fast convergence and high-quality image recovery are two essential features of algorithms for solving ill-posed imaging inverse problems. Existing methods, such as regularization by denoising (RED), often focus on designing sophisticated image priors to improve reconstruction quality, while leaving…

Cited by 0SourcecodeScholar
2026

Reasoning Compartmentalization: Bridging the Concretization Gap via Abstraction-based Routing

ICML 2026poster

While previous research has documented the sensitivity of Large Language Models (LLMs) to surface-level performance degradation, the underlying impact on internal representations and learning dynamics remains under-explored. In this work, we study this question using a controlled setup with paired r…

Cited by 0SourceScholar
2026

Robotic Cell Manipulation at the Solid-Liquid Interface for Cryopreservation

ICRA 2026poster

Automating cell manipulation at a solid-liquid interface is a critical challenge for biomedical applications such as embryo cryopreservation. Unlike manipulation in a full liquid medium, the cell-substrate contact creates a significant static friction force that is not readily measurable with curren…

Cited by 0Scholar
2026

WMVLM: Evaluating Diffusion Model Image Watermarking via Vision-Language Models

ICML 2026poster

Digital watermarking is essential for securing generated images from diffusion models. Accurate watermark evaluation is critical for algorithm development, yet existing methods have significant limitations: they lack a unified framework for both residual and semantic watermarks, provide results with…

Cited by 0SourceScholar
2026

Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training

AAAI 2026technical

Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated datasets. Although some zero-shot methods attempt to trajectory control in the l

Cited by 0SourcePDFScholar
2025

Automated Video Object Detection of Motile Cells Under Microscopy

ICRA 2025

Video object detection (VOD) of motile cells (e.g., bacteria and sperm) under microscopy is challenging due to motion blur, sporadic out-of-focus, and pose variations. Compared with VOD in generic scenes, the lower contrast and smaller color space of microscopy imaging further introduce feature over

Cited by 0SourceScholar
2025

BeamLoRA: Beam-Constraint Low-Rank Adaptation

ACL 2025long

Due to the demand for efficient fine-tuning of large language models, Low-Rank Adaptation (LoRA) has been widely adopted as one of the most effective parameter-efficient fine-tuning methods. Nevertheless, while LoRA improves efficiency, there remains room for improvement in accuracy. Herein, we adop…

Cited by 0SourcePDFScholar
2025

Benchmarking Multi-Object Grasping

RA-L 2025

In this work, we describe a multi-object grasping benchmark to evaluate the grasping and manipulation capabilities of robotic systems in both pile and surface scenarios. The benchmark introduces three robot multi-object grasping benchmarking protocols designed to challenge different aspects of robot

Cited by 3SourceScholar
2025

Channel and space-based joint rate allocation algorithm

ICASSP 2025accepted

Rate control is a critical component for image and video compression Particularly under limited network bandwidth conditions, bitrate control is essential to ensure efficient image transmission by effectively allocation channel resources. In this research, since both Channel and Spatial have relatio…

Cited by 0SourceScholar
2025

Continuous Convolution for Automated Measurement of Sperm Flagella

ICRA 2025

Quantifying sperm flagellar beating behavior (e.g., beating amplitude, frequency, and wavelength) plays a crucial role in biological research, clinical diagnostics, and the design of sperm-inspired microrobots. However, existing computational methods struggle to accurately and efficiently analyze th

Cited by 0SourcecodeScholar
2025

CritiQ: Mining Data Quality Criteria from Human Preferences

ACL 2025long

Language model heavily depends on high-quality data for optimal performance. Existing approaches rely on manually designed heuristics, the perplexity of existing models, training classifiers, orcareful prompt engineering, which require significant expert experience and human annotation effort while…

2025

Curiosity-Driven Reinforcement Learning from Human Feedback

ACL 2025long

Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but often at the cost of reduced output diversity. This trade-off between diversity and alignment quality remains a significant challenge. Drawing inspiration from…

2025

Graph Structure Learning for Spatial-Temporal Imputation: Adapting to Node and Feature Scales

AAAI 2025technical

Spatial-temporal data collected across different geographic locations often suffer from missing values, posing challenges to data analysis. Existing methods primarily leverage fixed spatial graphs to impute missing values, which implicitly assume that the spatial relationship is roughly the same for…

2025

HFT: Half Fine-Tuning for Large Language Models

ACL 2025long

Large language models (LLMs) with one or more fine-tuning phases have become necessary to unlock various capabilities, enabling LLMs to follow natural language instructions and align with human preferences. However, it carries the risk of catastrophic forgetting during sequential training, the param…

2025

Image-Based Compliance Control for Robotic Steering of a Ferromagnetic Guidewire

ICRA 2025

Robotic steering of magnetic guidewires has shown great potential in accelerating endovascular interventions, enhancing the success rate of time-sensitive surgeries such as stroke treatment. Incomplete state feedback of the guidewire from 2D perspective images and unknown interactions with the surro

Cited by 0SourceScholar
2025

Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking

ACL 2025long

Large language models (LLMs) face inherent performance bottlenecks under parameter constraints, particularly in processing critical tokens that demand complex reasoning. Empirical analysis reveals challenging tokens induce abrupt gradient spikes across layers, exposing architectural stress points in…

2025

InverseBench: Benchmarking Plug-and-Play Diffusion Priors for Inverse Problems in Physical Sciences

ICLR 2025spotlight

Plug-and-play diffusion priors (PnPDP) have emerged as a promising research direction for solving inverse problems. However, current studies primarily focus on natural image restoration, leaving the performance of these algorithms in scientific inverse problems largely unexplored. To address this…

2025

Learning to (Learn at Test Time): RNNs with Expressive Hidden States

ICML 2025spotlight

Self-attention performs well in long context but has quadratic complexity. Existing RNN layers have linear complexity, but their performance in long context is limited by the expressive power of their hidden states. We present a practical framework for instantiating sequence modeling layers with lin…

2025

LoRO: Real-Time on-Device Secure Inference for LLMs via TEE-Based Low Rank Obfuscation

NeurIPS 2025poster

While Large Language Models (LLMs) have gained remarkable success, they are consistently at risk of being stolen when deployed on untrusted edge devices. As a solution, TEE-based secure inference has been proposed to protect valuable model property. However, we identify a statistical vulnerability i…

Cited by 0SourceScholar
2025

MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions

ICLR 2025poster

Reinforcement learning from human feedback (RLHF) has demonstrated effectiveness in aligning large language models (LLMs) with human preferences. However, token-level RLHF suffers from the credit assignment problem over long sequences, where delayed rewards make it challenging for the model to disce…

2025

Meta Guidance: Incorporating Inductive Biases into Deep Time Series Imputers

NeurIPS 2025poster

Missing values, frequently encountered in time series data, can significantly impair the effectiveness of analytical methods. While deep imputation models have emerged as the predominant approach due to their superior performance, explicitly incorporating inductive biases aligned with time-series ch…

Cited by 0SourceScholar
2025

Mixture of Hidden-Dimensions: Not All Hidden-States’ Dimensions are Needed in Transformer

ICML 2025poster

Transformer models encounter inefficiency when scaling hidden dimensions due to the uniform expansion of parameters. When delving into the sparsity of hidden dimensions, we observe that only a small subset of dimensions are highly activated, where some dimensions are commonly activated across tokens…

Cited by 0SourcePDFScholar
2025

One-Minute Video Generation with Test-Time Training

CVPR 2025poster

Transformers today still struggle to generate one-minute videos because self-attention layers are inefficient for long context. Alternatives such as Mamba layers struggle to produce coherent scenes because their hidden states are small and less expressive. We experiment with Test-Time Training (TTT)…

2025

PolyGuard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset

NeurIPS 2025poster

As large language models (LLMs) become widespread across diverse applications, concerns about the security and safety of LLM interactions have intensified. Numerous guardrail models and benchmarks have been developed to ensure LLM content safety. However, existing guardrail benchmarks are often buil…

Cited by 0SourceScholar
2025

PromptHMR: Promptable Human Mesh Recovery

CVPR 2025poster

Human pose and shape (HPS) estimation presents challenges in diverse scenarios such as crowded scenes, person-person interactions, and single-view reconstruction. Existing approaches lack mechanisms to incorporate auxiliary "side information" that could enhance reconstruction accuracy in such challe…

Cited by 0SourcePDFScholar
2025

ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning

EMNLP 2025

Reasoning-based large language models have excelled in mathematics and programming, yet their potential in knowledge-intensive medical question answering remains underexplored and insufficiently validated in clinical contexts. To bridge this gap, we introduce ReasonMed , the largest medical reasonin

2025

Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging

ACL 2025long

Mixture-of-Experts (MoE) shines brightly in large language models (LLMs) and demonstrates outstanding performance in plentiful natural language processing tasks. However, existing methods transforming LLMs from dense to MoE face significant data requirements and typically rely on large-scale post-tr…

2025

Whitened Score Diffusion: A Structured Prior for Imaging Inverse Problems

NeurIPS 2025poster

Conventional score-based diffusion models (DMs) may struggle with anisotropic Gaussian diffusion processes due to the required inversion of covariance matrices in the denoising score matching training objective \cite{vincent_connection_2011}. We propose Whitened Score (WS) diffusion models, a novel…

Cited by 0SourcecodeScholar
2024

Automated Sperm Immobilization with a Clinically-Compatible and Compact XYZ Stage

ICRA 2024poster

Automated positioning systems play a pivotal role in micro-scale cell manipulation. In clinical intracytoplasmic sperm injection (ICSI) of in vitro fertilization (IVF) treatment, a motile sperm needs to be immobilized by glass micropipette tapping for subsequent surgical steps. The process requires…

Cited by 0SourceScholar
2024

Automated Sperm Morphology Analysis Based on Instance-Aware Part Segmentation

ICRA 2024poster

Traditional sperm morphology analysis is based on tedious manual annotation. Automated morphology analysis of a high number of sperm requires accurate segmentation of each sperm part and quantitative morphology evaluation. State-of-the-art instance-aware part segmentation networks follow a "detect-t…

Cited by 2SourceScholar
2024

Autoregressive Pre-Training on Pixels and Texts

EMNLP 2024main

The integration of visual and textual information represents a promising direction in the advancement of language models. In this paper, we explore the dual modality of language—both visual and textual—within an autoregressive framework, pre-trained on both document images and texts. Our method empl…

2024

ChatPose: Chatting about 3D Human Pose

CVPR 2024poster

We introduce ChatPose a framework employing Large Language Models (LLMs) to understand and reason about 3D human poses from images or textual descriptions. Our work is motivated by the human ability to intuitively understand postures from a single image or a brief description a process that intertwi…

2024

DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion

NeurIPS 2024poster

Large language models (LLMs) with billions of parameters demonstrate impressive performance. However, the widely used Multi-Head Attention (MHA) in LLMs incurs substantial computational and memory costs during inference. While some efforts have optimized attention mechanisms by pruning heads or shar…

Cited by 4SourcePDFScholar
2024

F-Eval: Asssessing Fundamental Abilities with Refined Evaluation Methods

ACL 2024long

Large language models (LLMs) garner significant attention for their unprecedented performance, leading to an increasing number of researches evaluating LLMs. However, these evaluation benchmarks are limited to assessing the instruction-following capabilities, overlooking the fundamental abilities th…

2024

Frequency-aware Generative Models for Multivariate Time Series Imputation

NeurIPS 2024poster

Missing data in multivariate time series are common issues that can affect the analysis and downstream applications. Although multivariate time series data generally consist of the trend, seasonal and residual terms, existing works mainly focus on optimizing the modeling for the first two items. How…

Cited by 2SourcePDFScholar
2024

From Cooking Recipes to Robot Task Trees – Improving Planning Correctness and Task Efficiency by Leveraging LLMs with a Knowledge Network

ICRA 2024poster

Task planning for robotic cooking involves generating a sequence of actions for a robot to prepare a meal successfully. This paper introduces a novel task tree generation pipeline producing correct planning and efficient execution for cooking tasks. Our method first uses a large language model (LLM)…

Cited by 16SourceScholar
2024

GI-PIP: Do We Require Impractical Auxiliary Dataset for Gradient Inversion Attacks?

ICASSP 2024accepted

Deep gradient inversion attacks expose a serious threat to Federated Learning (FL) by accurately recovering private data from shared gradients. However, the state-of-the-art heavily relies on impractical assumptions to access excessive auxiliary data, which violates the basic data partitioning princ…

Cited by 0SourceScholar
2024

Generalizing End-To-End Autonomous Driving In Real-World Environments Using Zero-Shot LLMs

CoRL 2024poster

Traditional autonomous driving methods adopt modular design, decomposing tasks into sub-tasks, including perception, prediction, planning, and control. In contrast, end-to-end autonomous driving directly outputs actions from raw sensor data, avoiding error accumulation. However, training an end-to-e…

Cited by 5SourceScholar
2024

High-Order Contrastive Learning with Fine-grained Comparative Levels for Sparse Ordinal Tensor Completion

ICML 2024poster

Contrastive learning is a powerful paradigm for representation learning with prominent success in computer vision and NLP, but how to extend its success to high-dimensional tensors remains a challenge. This is because tensor data often exhibit high-order mode-interactions that are hard to profile an…

Cited by 0SourcePDFScholar
2024

Kresling Origami With Differentiation Flaw Design for Multidirectional Crawling Robot

RA-L 2024

Multidirectional motion ability is a significant factor for crawling robots. Inspired by the Kresling origami pattern, this work introduces a soft pneumatic actuator that can achieve a compound motion including twisting, contraction, and multidirectional bending under the control of one single air s

Cited by 4SourceScholar
2024

LEMON: Reviving Stronger and Smaller LMs from Larger LMs with Linear Parameter Fusion

ACL 2024long

In the new era of language models, small models (with billions of parameter sizes) are receiving increasing attention due to their flexibility and cost-effectiveness in deployment. However, limited by the model size, the performance of small models trained from scratch may often be unsatisfactory. L…

2024

LOCR: Location-Guided Transformer for Optical Character Recognition

EMNLP 2024finding

Academic documents are packed with texts, equations, tables, and figures, requiring comprehensive understanding for accurate Optical Character Recognition (OCR). While end-to-end OCR methods offer improved accuracy over layout-based approaches, they often grapple with significant repetition issues,…

2024

NACL: A General and Effective KV Cache Eviction Framework for LLM at Inference Time

ACL 2024long

Large Language Models (LLMs) have ignited an innovative surge of AI applications, marking a new era of exciting possibilities equipped with extended context windows. However, hosting these models is cost-prohibitive mainly due to the extensive memory consumption of KV Cache involving long-context mo…

2024

On Training Data Influence of GPT Models

EMNLP 2024main

Amidst the rapid advancements in generative language models, the investigation of how training data shapes the performance of GPT models is still emerging. This paper presents GPTfluence, a novel approach that leverages a featurized simulation to assess the impact of training examples on the trainin…

2024

Principled Probabilistic Imaging using Diffusion Models as Plug-and-Play Priors

NeurIPS 2024poster

Diffusion models (DMs) have recently shown outstanding capabilities in modeling complex image distributions, making them expressive image priors for solving Bayesian inverse problems. However, most existing DM-based methods rely on approximations in the generative process to be generic to different…

2024

Rigid-Soft Hybrid Suction Cups for Enhanced Anti-Torque and Energy-Efficient Attachment

RA-L 2024

In the realm of robotics, suction-based adhesion plays a pivotal role in applications ranging from object transfer to wall-climbing robots. To improve the sealing and attachment stability of suction cups, researchers have employed state-of-the-art techniques, including the use of soft materials with

Cited by 3SourceScholar
2024

TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation

CVPR 2024poster

We address the problem of regressing 3D human pose and shape from a single image with a focus on 3D accuracy. The current best methods leverage large datasets of 3D pseudo-ground-truth (p-GT) and 2D keypoints leading to robust performance. With such methods however we observe a paradoxical decline i…

2024

Weakly-Supervised Depth Completion during Robotic Micromanipulation from a Monocular Microscopic Image

ICRA 2024poster

Obtaining three-dimensional information, especially the z-axis depth information, is crucial for robotic micromanipulation. Due to the unavailability of depth sensors such as lidars in micromanipulation setups, traditional depth acquisition methods such as depth from focus or depth from defocus dire…

Cited by 0SourceScholar
2023

A Novel Mode Selection-Based Fast Intra Prediction Algorithm for Spatial SHVC

ICASSP 2023accepted

Due to multi-layer encoding and Inter-layer prediction, Spatial Scalable High-Efficiency Video Coding (SSHVC) has extremely high coding complexity. It is very crucial to improve its coding speed so as to promote widespread and cost-effective SSHVC applications. In this paper, we have proposed a nove…

Cited by 0SourceScholar
2023

An Embarrassingly Easy but Strong Baseline for Nested Named Entity Recognition

ACL 2023short

Named entity recognition (NER) is the task to detect and classify entity spans in the text. When entity spans overlap between each other, the task is named as nested NER. Span-based methods have been widely used to tackle nested NER. Most of these methods get a score matrix, where each entry corresp…

2023

ERNIE-Code: Beyond English-Centric Cross-lingual Pretraining for Programming Languages

ACL 2023findings

Software engineers working with the same programming language (PL) may speak different natural languages (NLs) and vice versa, erecting huge barriers to communication and working efficiency. Recent studies have demonstrated the effectiveness of generative pre-training in computer programs, yet they…

2023

ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model With Knowledge-Enhanced Mixture-of-Denoising-Experts

CVPR 2023highlight

Recent progress in diffusion models has revolutionized the popular technology of text-to-image generation. While existing approaches could produce photorealistic high-resolution images with text conditions, there are still several open problems to be solved, which limits the further improvement of i…

Cited by 140SourcePDFScholar
2023

End-to-End Pipeline for Trigger Detection on Hit and Track Graphs

AAAI 2023technical

There has been a surge of interest in applying deep learning in particle and nuclear physics to replace labor-intensive offline data analysis with automated online machine learning tasks. This paper details a novel AI-enabled triggering solution for physics experiments in Relativistic Heavy Ion Coll…

Cited by 2SourcePDFScholar
2023

Instance-wise Batch Label Restoration via Gradients in Federated Learning

ICLR 2023poster

Gradient inversion attacks have posed a serious threat to the privacy of federated learning. The attacks search for the optimal pair of input and label best matching the shared gradients and the search space of the attacks can be reduced by pre-restoring labels. Recently, label restoration technique…

2023

Pose-Oriented Transformer with Uncertainty-Guided Refinement for 2D-to-3D Human Pose Estimation

AAAI 2023technical

There has been a recent surge of interest in introducing transformers to 3D human pose estimation (HPE) due to their powerful capabilities in modeling long-term dependencies. However, existing transformer-based methods treat body joints as equally important inputs and ignore the prior knowledge of h…

Cited by 51SourcePDFScholar
2023

TRACE: 5D Temporal Regression of Avatars With Dynamic Cameras in 3D Environments

CVPR 2023poster

Although the estimation of 3D human pose and shape (HPS) is rapidly progressing, current methods still cannot reliably estimate moving humans in global coordinates, which is critical for many applications. This is particularly challenging when the camera is also moving, entangling human and camera m…

2023

UTC-IE: A Unified Token-pair Classification Architecture for Information Extraction

ACL 2023long

Information Extraction (IE) spans several tasks with different output structures, such as named entity recognition, relation extraction and event extraction. Previously, those tasks were solved with different models because of diverse task output structures. Through re-examining IE tasks, we find th…

2023

Unleashing the Power of Gradient Signal-to-Noise Ratio for Zero-Shot NAS

ICCV 2023poster

Neural Architecture Search (NAS) aims to automatically find optimal neural network architectures in an efficient way. Zero-Shot NAS is a promising technique that leverages proxies to predict the accuracy of candidate architectures without any training. However, we have observed that most existing pr…

Cited by 6PDFcodeScholar
2022

Clip-Tuning: Towards Derivative-free Prompt Learning with a Mixture of Rewards

EMNLP 2022finding

Derivative-free prompt learning has emerged as a lightweight alternative to prompt tuning, which only requires model inference to optimize the prompts. However, existing work did not take full advantage of the over-parameterized characteristics of large pre-trained language models (PLMs). In this pa…

Cited by 19SourcePDFScholar
2022

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

EMNLP 2022finding

Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and utilization of layout-centered knowledge, leading to sub-optimal performances. In this paper, we propose ERNIE-Layout, a…

2022

Learning Cross-Video Neural Representations for High-Quality Frame Interpolation

ECCV 2022poster

"This paper considers the problem of temporal video interpolation, where the goal is to synthesize a new video frame given its two neighbors. We propose Cross-Video Neural Representation (CURE) as the first video interpolation method based on neural fields (NF). NF refers to the recent class of meth…

2022

Learning Monocular Mesh Recovery of Multiple Body Parts Via Synthesis

ICASSP 2022accepted

In this paper, we focus on simultaneously recovering the 3D mesh of multiple body parts from a single RGB image. One of the main challenges is that available datasets with full-body 3D annotations are very limited. This results in poor generalization ability of existing learning-based methods. Exist…

Cited by 0SourceScholar
2022

Multi-Object Grasping - Efficient Robotic Picking and Transferring Policy for Batch Picking

IROS 2022poster

In a typical fulfillment center, the order fulfilling process is managed by a warehouse management system (WMS). For efficiency, WMS usually applies batch picking, also called multi-order picking, to collect the same items for multiple orders. Suppose an item appears in multiple orders, instead of r…

Cited by 14SourceScholar
2022

OCRTOC: A Cloud-Based Competition and Benchmark for Robotic Grasping and Manipulation

RA-L 2022

In this paper, we propose a cloud-based benchmark for robotic grasping and manipulation, called the OCRTOC benchmark. The benchmark focuses on the object rearrangement problem, specifically table organization tasks. We provide a set of identical real robot setups and facilitate remote experiments of

Cited by 58SourcecodeScholar
2022

Putting People in Their Place: Monocular Regression of 3D People in Depth

CVPR 2022poster

Given an image with multiple people, our goal is to directly regress the pose and shape of all the people as well as their relative depth. Inferring the depth of a person in an image, however, is fundamentally ambiguous without knowing their height. This is particularly problematic when the scene co…

Cited by 180PDFcodeScholar
2022

Research Challenges and Progress in Robotic Grasping and Manipulation Competitions

RA-L 2022

This paper discusses recent research progress in robotic grasping and manipulation in the light of the latest Robotic Grasping and Manipulation Competitions (RGMCs). We first provide an overview of past benchmarks and competitions related to the robotics manipulation field. Then, we discuss the meth

Cited by 57SourceScholar
2022

Robotic Cell Manipulation for Blastocyst Biopsy

ICRA 2022poster

Soft tissue cutting is used for incision, separation and removal of tissues or cells. Due to high deformation of soft tissues resulting from their viscosity and elasticity, it is challenging to accurately cut the tissue along a desired path and control the force applied to the tissue for reducing in…

Cited by 6SourceScholar
2022

Simple and Effective Relation-based Embedding Propagation for Knowledge Representation Learning

IJCAI 2022poster

Relational graph neural networks have garnered particular attention to encode graph context in knowledge graphs (KGs). Although they achieved competitive performance on small KGs, how to efficiently and effectively utilize graph context for large KGs remains an open problem. To this end, we propose…

2022

Unified Matrix Coding for NN Originated MIP in H.266/VVC

ICASSP 2022accepted

Matrix-based Intra Prediction (MIP) is an effective coding algorithm in H.266/Versatile Video Coding (VVC) which is originated by Neural Networks (NN). With the requirement of low complexity, MIP is conducted by a matrix-vector multiplication. To handle with the diversity of video content, 30 matric…

Cited by 0SourceScholar
2021

Async-RED: A Provably Convergent Asynchronous Block Parallel Stochastic Method using Deep Denoising Priors

ICLR 2021spotlight

Regularization by denoising (RED) is a recently developed framework for solving inverse problems by integrating advanced denoisers as image priors. Recent work has shown its state-of-the-art performance when combined with pre-trained deep denoisers. However, current RED algorithms are inadequate for…

Cited by 18SourcePDFScholar
2021

Automated End-Effector Alignment for Robotic Cell Manipulation

ICRA 2021poster

Cell manipulation is a key technology in many biomedical and clinical applications, in which end-effector alignment is a critical procedure. Presently, end-effector alignment is performed manually and suffers from large misalignment error and inconsistency. Manual alignment often undesirably moves t…

Cited by 3SourceScholar
2021

CVAE-based Re-anchoring for Implicit Discourse Relation Classification

EMNLP 2021finding

Training implicit discourse relation classifiers suffers from data sparsity. Variational AutoEncoder (VAE) appears to be the proper solution. It is because ideally VAE is capable of generating inexhaustible varying samples, and this facilitates selective data augmentation. However, our experiments s…

Cited by 13SourcePDFScholar
2021

ERNIE-Doc: A Retrospective Long-Document Modeling Transformer

ACL 2021long

Transformers are not suited for processing long documents, due to their quadratically increasing memory and time consumption. Simply truncating a long document or applying the sparse attention mechanism will incur the context fragmentation problem or lead to an inferior modeling capability against c…

2021

ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language Understanding

NAACL 2021long

Coarse-grained linguistic information, such as named entities or phrases, facilitates adequately representation learning in pre-training. Previous works mainly focus on extending the objective of BERT’s Masked Language Modeling (MLM) from masking individual tokens to contiguous sequences of n tokens…

2021

ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual Corpora

EMNLP 2021main

Recent studies have demonstrated that pre-trained cross-lingual models achieve impressive performance in downstream cross-lingual tasks. This improvement benefits from learning a large amount of monolingual and parallel corpora. Although it is generally acknowledged that parallel corpora are critica…

2021

ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene Graphs

AAAI 2021technical

We propose a knowledge-enhanced approach, ERNIE-ViL, which incorporates structured knowledge obtained from scene graphs to learn joint representations of vision-language. ERNIE-ViL tries to build the detailed semantic connections (objects, attributes of objects and relationships between objects) acr…

Cited by 425SourcePDFScholar
2021

Guest Editorial: Introduction to the Special Issue on Benchmarking Protocols for Robotic Manipulation

RA-L 2021

The papers in this special section focus on benchmarking protocols for robotic manipulation. Benchmarks are crucial for analyzing the effectiveness of an approach against a common basis, providing a quantitative means for interpreting performance. Carefully designed and widely recognized benchmarks

Cited by 3SourceScholar
2021

Masked Label Prediction: Unified Message Passing Model for Semi-Supervised Classification

IJCAI 2021poster

Graph neural network (GNN) and label propagation algorithm (LPA) are both message passing algorithms, which have achieved superior performance in semi-supervised classification. GNN performs feature propagation by a neural network to make predictions, while LPA uses label propagation across graph ad…

2021

Monocular, One-Stage, Regression of Multiple 3D People

ICCV 2021poster

This paper focuses on the regression of multiple 3D people from a single RGB image. Existing approaches predominantly follow a multi-stage pipeline that first detects people in bounding boxes and then independently regresses their 3D body meshes. In contrast, we propose to Regress all meshes in a On…

Cited by 327PDFcodeScholar
2021

Multi-Object Grasping – Estimating the Number of Objects in a Robotic Grasp

IROS 2021poster

A human hand can grasp a desired number of objects at once from a pile based solely on tactile sensing. To do so, a robot needs to make a grasp in a pile, sense the number of objects in the grasp before lifting, and predict how many will remain in the grasp after lifting. It is a very challenging pr…

Cited by 27SourceScholar
2021

Online Learning of Unknown Dynamics for Model-Based Controllers in Legged Locomotion

RA-L 2021

The performance of a model-based controller can severely suffer when its model inaccurately represents the real world dynamics. We propose to learn a time-varying, locally linear residual model along the robot's current trajectory, to compensate for the prediction errors of the controller's model. S

Cited by 65SourceScholar
2021

Self-Supervised Policy Adaptation during Deployment

ICLR 2021spotlight

In most real world scenarios, a policy trained by reinforcement learning in one environment needs to be deployed in another, potentially quite different environment. However, generalization across different environments is known to be hard. A natural solution would be to keep training after deployme…

2021

Stochastic Deep Unfolding for Imaging Inverse Problems

ICASSP 2021accepted

Deep unfolding networks are rapidly gaining attention for solving imaging inverse problems. However, the computational and memory complexity of existing deep unfolding networks scales with the size of the full measurement set, limiting their applicability to certain large-scale imaging inverse probl…

Cited by 0SourceScholar
2020

An SEM-Based Nanomanipulation System for Multi-Physical Characterization of Single InGaN/GaN Nanowires

IROS 2020poster

Functional nanomaterials possess exceptional multi-physical (e.g., mechanical, electrical and optical) properties compared with their bulk counterparts. To facilitate both synthesis and device applications of these nanomaterials, it is highly desired to characterize their multi-physical properties w…

Cited by 14SourceScholar
2020

Automated Eye-in-Hand Robot-3D Scanner Calibration for Low Stitching Errors

ICRA 2020poster

A 3D measurement system consisting of a 3D scanner and an industrial robot (eye-in-hand) is commonly used to scan large object under test (OUT) from multiple fieldof-views (FOVs) for complete measurement. A data stitching process is required to align multiple FOVs into a single coordinate system. Ma…

Cited by 10SourceScholar
2020

Benchmarking Protocols for Evaluating Small Parts Robotic Assembly Systems

RA-L 2020

This paper presents a set of performance metrics, test methods, and associated artifacts to help progress the development and deployment of robotic assembly systems. The designs for three task board artifacts that replicate small part insertion and fastening operations such as threading, snap fittin

Cited by 95SourceScholar
2020

Design and Control of a Piezo Drill for Robotic Piezo-Driven Cell Penetration

RA-L 2020

Cell penetration is an indispensable step in many cell surgery tasks. Conventionally, cell penetration is achieved by passively indenting and eventually puncturing the cell membrane, during which undesired large cell deformation is induced. Piezo drills have been developed to penetrate cells with le

Cited by 28SourceScholar
2020

ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning Framework for Natural Language Generation

IJCAI 2020poster

Current pre-training works in natural language generation pay little attention to the problem of exposure bias on downstream tasks. To address this issue, we propose an enhanced multi-flow sequence to sequence pre-training and fine-tuning framework named ERNIE-GEN, which bridges the discrepancy betw…

2020

Model-Based Robotic Cell Aspiration: Tackling Nonlinear Dynamics and Varying Cell Sizes

RA-L 2020

Aspirating a single cell from the outside to the inside of a micropipette is widely used for cell transfer and manipulation. Due to the small volume of a single cell (picoliter) and nonlinear dynamics involved in the aspiration process, it is challenging to accurately and quickly position a cell to

Cited by 13SourceScholar
2020

Robotic Control of a Magnetic Swarm for On-Demand Intracellular Measurement

ICRA 2020poster

In biology, fluorescent dyes are routinely used for biochemical measurements such as pH and ion concentrations. They, especially when used for detecting a low concentration of ions, suffer from low signal-to-noise ratios (SNR); and increasing the concentration of fluorescent dyes causes more sever c…

Cited by 6SourceScholar
2020

Robotic Swarm Control for Precise and On-Demand Embolization

ICRA 2020poster

Existing approaches for robotic control of magnetic swarms are not capable of generating magnetic aggregates precisely in an arbitrarily specified target region in a fluidic flow environment. Such a swarm control capability is demanded by medical applications such as clinical embolization (i.e., loc…

Cited by 8SourceScholar
2020

Test-Time Training with Self-Supervision for Generalization under Distribution Shifts

ICML 2020poster

In this paper, we propose Test-Time Training, a general approach for improving the performance of predictive models when training and test data come from different distributions. We turn a single unlabeled test sample into a self-supervised learning problem, on which we update the model parameters b…

Cited by 945SourcePDFScholar
2019

Accurate Pouring using Model Predictive Control Enabled by Recurrent Neural Network

IROS 2019poster

Humans perform the task of pouring often and in which exhibit consistent accuracy regardless of the complicated dynamics of the liquid. Model predictive control (MPC) appears to be a natural candidate solution for the task of accurate pouring considering its wide use in industrial applications. Howe…

Cited by 19SourceScholar
2019

Automated Aortic Pressure Regulation in ex vivo Heart Perfusion

ICRA 2019poster

This paper presents the first system for automated ex vivo perfusion of an isolated heart and regulating the heart's aortic pressure (AoP). An adaptive controller was developed for AoP regulation and maintained the heart's physiological aerobic metabolism. A mathematical model of the perfusion syste…

Cited by 0SourceScholar
2019

Automated Laser Ablation of Motile Sperm for Immobilization

RA-L 2019

Automated manipulation of single cells is required in both biological and clinical applications. In clinical infertility treatments, a single motile sperm is immobilized and inserted into an egg cell for in vitro fertilization. Sperm immobilization is essential to ease the ensuing pick-up procedure,

Cited by 5SourceScholar
2019

Doubly Robust Joint Learning for Recommendation on Data Missing Not at Random

ICML 2019oral

In recommender systems, usually the ratings of a user to most items are missing and a critical problem is that the missing ratings are often missing not at random (MNAR) in reality. It is widely acknowledged that MNAR ratings make it difficult to accurately predict the ratings and unbiasedly estimat…

Cited by 282SourcePDFScholar
2019

Human Mesh Recovery From Monocular Images via a Skeleton-Disentangled Representation

ICCV 2019poster

We describe an end-to-end method for recovering 3D human body mesh from single images and monocular videos. Different from the existing methods try to obtain all the complex 3D pose, shape, and camera parameters from one coupling feature, we propose a skeleton-disentangling based framework, which di…

Cited by 214PDFcodeScholar
2019

Image Restoration Using Total Variation Regularized Deep Image Prior

ICASSP 2019accepted

In the past decade, sparsity-driven regularization has led to significant improvements in image reconstruction. Traditional regularizers, such as total variation (TV), rely on analytical models of sparsity. However, increasingly the field is moving towards trainable models, inspired from deep learni…

Cited by 0SourceScholar
2019

Regularized Fourier Ptychography Using an Online Plug-and-play Algorithm

ICASSP 2019accepted

The plug-and-play priors (PnP) framework has been recently shown to achieve state-of-the-art results in regularized image reconstruction by leveraging a sophisticated denoiser within an iterative algorithm. In this paper, we propose a new online PnP algorithm for Fourier ptychographic microscopy (FP…

Cited by 0SourceScholar
2019

Robotic Orientation Control of Deformable Cells

ICRA 2019poster

Robotic manipulation of deformable objects (vs. rigid objects) has been a classic topic in robotics. Compared to deformable synthetic objects such as rubber balls and clothes, biological cells are highly deformable and more prone to damage. This paper presents robotic manipulation of deformable cell…

Cited by 8SourceScholar
2018

Automated Non-Invasive Measurement of Sperm Motility and Morphology Parameters

ICRA 2018poster

Measuring the motility and morphology parameters of motile cells is important for revealing their functional characteristics. This paper presents automation techniques that, for the first time, enable automated, non-invasive measurement of motility and morphology parameters of individual sperms. Com…

Cited by 4SourceScholar
2018

KDGAN: Knowledge Distillation with Generative Adversarial Networks

NeurIPS 2018poster

Knowledge distillation (KD) aims to train a lightweight classifier suitable to provide accurate inference with constrained resources in multi-label learning. Instead of directly consuming feature-label pairs, the classifier is trained by a teacher, i.e., a high-capacity model whose training may be r…

2018

Robotic Intracellular Manipulation: 3D Navigation and Measurement Inside a Single Cell

ICRA 2018poster

Magnetic micromanipulation is an untethered technique and has enabled numerous applications in the scale of millimeters to micrometers from the tissue level to cell level. However, existing systems are not capable of maneuvering a sub-micrometer object for precise force control, preventing the reali…

Cited by 8SourceScholar
2017

A model of vertebral motion and key point recognition of drilling with force in robot-assisted spinal surgery

IROS 2017poster

Pedicle drilling is a crucial and high-risk process in spinal surgery. Due to the respiration and cardiac cycle, the position of spine would fluctuate during operations, which result in an increase of the difficulty in state recognition of pedicle drilling. To guarantee the safety and validity, a mo…

Cited by 8SourceScholar
2017

Learning to pour

IROS 2017poster

Pouring is a simple task people perform daily. It is the second most frequently executed motion in cooking scenarios, after pick-and-place. We present a pouring trajectory generation approach, which uses force feedback from the cup to determine the future velocity of pouring. The approach uses recur…

Cited by 31SourceScholar
2017

Robotic Pick-And-Place of Multiple Embryos for Vitrification

RA-L 2017

Embryo vitrification is an essential cryopreservation technique in IVF (in vitro fertilization) clinics. Vitrification involves pick-and-place of an embryo in multiple types of cryoprotectant solutions for processing before placing the embryo on a vitrification straw for cryopreservation in liquid n

Cited by 37SourceScholar
2017

Three-dimensional robotic control of a 5-micrometer magnetic bead for intra-embryonic navigation and measurement

ICRA 2017poster

Magnetic micromanipulation has the advantage of untethered control, high precision, and biocompatibility and has recently undergone great advances. The magnetic micromanipulation task to tackle in this work is to three-dimensionally navigate a 5-micrometer magnetic bead inside a mouse embryo and per…

Cited by 1SourceScholar
2016

An automated system for investigating sperm orientation in fluid flow

ICRA 2016

Mammalian sperms reorient against fluid flow in the female reproductive tract, known as rheotaxis. Compared to chemotaxis that provides short-distance guidance, rheotaxis provides long-distance guidance for a sperm to find the egg cell. However, only a low number of sperms are capable of rheotaxis a

Cited by 6SourceScholar
2016

Functional object-oriented network for manipulation learning

IROS 2016poster

This paper presents a novel structured knowledge representation called the functional object-oriented network (FOON) to model the connectivity of the functional-related objects and their motions in manipulation tasks. The graphical model FOON is learned by observing object state change and human man…

Cited by 106SourceScholar
2016

Supervised Word Mover's Distance

NeurIPS 2016oral

Accurately measuring the similarity between text documents lies at the core of many real world applications of machine learning. These include web-search ranking, document recommendation, multi-lingual document matching, and article categorization. Recently, a new document metric, the word mover's d…

2015

Automated micro-aspiration of mouse embryo limb bud tissue

ICRA 2015poster

Mechanical force is an integral part of tissue morphogenesis and patterning. We have developed an automated micro-aspiration system to investigate how mouse limb bud tissue responds to extrinsic forces in order to understand whether tissue-generated forces can be a part of the mechanism causing orie…

Cited by 7SourceScholar