← Search

Pheng-Ann Heng

98 accepted papers

2026

Benchmarking Endoscopic Surgical Image Restoration and Beyond

CVPR 2026

In endoscopic surgery, a clear and high-quality visual field is critical for surgeons to make accurate intraoperative decisions. However, persistent visual degradation, including smoke generated by energy devices, lens fogging from thermal gradients, and lens contamination due to blood or tissue flu

Cited by 0SourcecodeScholar
2026

EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning

ICLR 2026poster

Robust 3D hand reconstruction is challenging in egocentric vision due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior works attempt to mitigate the challenges by scaling up training data or incorporating auxiliary cues, often falling short of effectively handling unse…

Cited by 0SourcecodeScholar
2026

From Manuals to Actions: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation

CVPR 2026

Vision-Language-Action (VLA) models have recently emerged, demonstrating strong generalization in robotic scene understanding and manipulation. However, when confronted with long-horizon tasks that require defined goal states, such as LEGO assembly or object rearrangement, existing VLA models still

Cited by 0SourceScholar
2026

HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

CVPR 2026

Human-product images, which showcase the integration of humans and products, play a vital role in advertising, e-commerce, and digital marketing. The essential challenge of generating such images lies in ensuring the high-fidelity preservation of product details. Among existing paradigms, reference-

Cited by 0SourcecodeScholar
2026

IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation

AAAI 2026technical

Recent visual generative models enable story generation with consistent characters from text, but human-centric story generation faces additional challenges, such as maintaining detailed and diverse human face consistency and coordinating multiple characters across different images. This paper prese

Cited by 0SourcePDFScholar
2026

KnowGuard: Knowledge-Driven Abstention for Multi-Round Clinical Reasoning

ICLR 2026poster

In clinical practice, physicians refrain from making decisions when patient information is insufficient. This behavior, known as abstention, is a critical safety mechanism preventing potentially harmful misdiagnoses. Recent investigations have reported the application of large language models (LLMs)…

Cited by 0SourceScholar
2026

Lost in Tokenization: Context as the Key to Unlocking Biomolecular Understanding in Scientific LLMs

ICLR 2026poster

Scientific Large Language Models (Sci-LLMs) have emerged as a promising frontier for accelerating biological discovery. However, these models face a fundamental challenge when processing raw biomolecular sequences: the tokenization dilemma. Whether treating sequences as a specialized language, riski…

Cited by 0SourcecodeScholar
2026

MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current emotional benchmarks remain limited, as it is still unknown: (a)…

Cited by 0SourcecodeScholar
2026

Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception

ICLR 2026poster

Fine-grained perception of multimodal information is critical for advancing human–AI interaction. With recent progress in audio–visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel, have emerged as a promising paradigm for achieving richer unde…

Cited by 0SourcecodeScholar
2026

ProteinAE: Protein Diffusion Autoencoders for Structure Encoding

ICLR 2026poster

Developing effective representations of protein structures is essential for advancing protein science, particularly for protein generative modeling. Current approaches often grapple with the complexities of the $\operatorname{SE}(3)$ manifold, rely on discrete tokenization, or the need for multiple…

Cited by 0SourcecodeScholar
2026

Rethinking Intermediate Representation for VLM-based Robot Manipulation

CVPR 2026

Vision-Language Model (VLM) is now an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate representation often needs a tradeoff between VLM-comprehensibility and generalizability. Inspired by context-free gramm

Cited by 0SourceScholar
2026

SurgPub-Video: A Comprehensive Surgical Video Framework for Enhanced Surgical Intelligence in Vision-Language Model

AAAI 2026technical

Vision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these challenges, we make the following contributions: (i) SurgPub-Video,

Cited by 0SourcePDFScholar
2026

Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos

CVPR 2026

Intraoperative bleeding in laparoscopic surgery causes rapid obscuration of the operative field to hinder the surgical process and increases the risk of postoperative complications. Intelligent detection of bleeding areas can quantify the blood loss to assist decision-making, while locating bleeding

Cited by 0SourcecodeScholar
2026

Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation

CVPR 2026

Recent advances in visual generation have increasingly explored the integration of reasoning capabilities. They incorporate textual reasoning, i.e., think, either before (as pre-planning) or after (as post-refinement) the generation process, yet they lack on-the-fly multimodal interaction during the

Cited by 0SourcecodeScholar
2026

Unifying Diffusion and Autoregression for Generalizable Vision-Language-Action Model

ICLR 2026poster

A central objective of manipulation policy design is to enable robots to comprehend human instructions and predict generalized actions in unstructured environments. Recent autoregressive vision-language-action (VLA) approaches discretize actions into bins to exploit the pretrained reasoning and gene…

Cited by 0SourceScholar
2026

Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework

CVPR 2026

To continuously enhance model adaptability in surgical video scene parsing, recent studies incrementally update it to progressively learn to segment an increasing number of surgical instruments over time. However, prior works constantly overlooked the potential of positive forward knowledge transfer

Cited by 0SourceScholar
2026

reAR: Rethinking Visual Autoregressive Models via Token-wise Consistency Regularization

ICLR 2026poster

Visual autoregressive (AR) generation offers a promising path toward unifying vision and language models, yet its performance remains suboptimal against diffusion models. Prior work often attributes this gap to tokenizer limitations and rasterization ordering. In this work, we identify a core bottle…

Cited by 0SourceScholar
2025

COS3D: Collaborative Open-Vocabulary 3D Segmentation

NeurIPS 2025poster

Open-vocabulary 3D segmentation is a fundamental yet challenging task, requiring a mutual understanding of both segmentation and language. However, existing Gaussian-splatting-based methods rely either on a single 3D language field, leading to inferior segmentation, or on pre-computed class-agnostic…

Cited by 0SourceScholar
2025

CellVerse: Do Large Language Models Really Understand Cell Biology?

NeurIPS 2025poster

Recent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis ta…

Cited by 0SourcecodeScholar
2025

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

NeurIPS 2025poster

Recent advancements underscore the significant role of Reinforcement Learning (RL) in enhancing the Chain-of-Thought (CoT) reasoning capabilities of large language models (LLMs). Two prominent RL algorithms, Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO), are cent…

Cited by 0SourcecodeScholar
2025

EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights

CVPR 2025poster

Traffic Anomaly Understanding (TAU) is essential for improving public safety and transportation efficiency by enabling timely detection and response to incidents. Beyond existing methods, which rely largely on visual data, we propose to consider audio cues, a valuable source that offers strong hints…

2025

Fast Image Super-Resolution via Consistency Rectified Flow

ICCV 2025poster

Diffusion models (DMs) have demonstrated remarkable success in real-world image super-resolution (SR), yet their reliance on time-consuming multi-step sampling largely hinders their practical applications. While recent efforts have introduced few- or single-step solutions, existing methods either in…

Cited by 0SourcePDFScholar
2025

Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning

NeurIPS 2025poster

Generalized policy and execution efficiency constitute the two critical challenges in robotic manipulation. While recent foundation policies benefit from the common-sense reasoning capabilities of internet-scale pretrained vision-language models (VLMs), they often suffer from low execution frequency…

Cited by 0SourcecodeScholar
2025

GLID$^2$E: A Gradient-Free Lightweight Fine-tune Approach for Discrete Biological Sequence Design

NeurIPS 2025poster

The design of biological sequences is essential for engineering functional biomolecules that contribute to advancements in human health and biotechnology. Recent advances in diffusion models, with their generative power and efficient conditional sampling, have made them a promising approach for sequ…

Cited by 0SourceScholar
2025

MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding

AAAI 2025technical

We introduce MM-Mixing, a multi-modal mixing alignment framework for 3D understanding. MM-Mixing applies mixing-based methods to multi-modal data, preserving and optimizing cross-modal connections while enhancing diversity and improving alignment across modalities. Our proposed two-stage training pi…

Cited by 0SourcePDFScholar
2025

MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models

IJCAI 2025

Text-to-image diffusion models can generate high-quality images but lack fine-grained control of visual concepts, limiting their creativity. Thus, we introduce component-controllable personalization, a new task that enables users to customize and reconfigure individual components within concepts. Th

Cited by 0SourcePDFScholar
2025

Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splatting

CVPR 2025poster

Lifting multi-view 2D instance segmentation to a radiance field has proven effective to enhance 3D understanding. Existing works rely on direct matching for end-to-end lifting, yielding inferior results, or employ a two-stage solution constrained by complex pre- or post-processing. In this work, we…

2025

SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency

NeurIPS 2025poster

Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial role of scenes in storytelling, which restricts their creativit…

Cited by 0SourceScholar
2025

SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems

ACL 2025finding

The rapid advancement of Large Multi-modal Models (LMMs) has enabled their application in scientific problem-solving, yet their fine-grained capabilities remain under-explored. In this paper, we introduce SciVerse, a multi-modal scientific evaluation benchmark to thoroughly assess LMMs across 5,735…

2025

Surgical Workflow Recognition and Blocking Effectiveness Detection in Laparoscopic Liver Resection with Pringle Maneuver

AAAI 2025technical

Pringle maneuver (PM) in laparoscopic liver resection aims to reduce blood loss and provide a clear surgical view by intermittently blocking blood inflow of the liver, whereas prolonged PM may cause ischemic injury. To comprehensively monitor this surgical procedure and provide timely warnings of in…

2025

T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

NeurIPS 2025poster

Recent advancements in large language models have demonstrated how chain-of-thought (CoT) and reinforcement learning (RL) can improve performance. However, applying such reasoning strategies to the visual generation domain remains largely unexplored. In this paper, we present **T2I-R1**, a novel rea…

Cited by 0SourcecodeScholar
2025

UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation

CVPR 2025poster

Estimating the 3D pose of hand and potential hand-held object from monocular images is a longstanding challenge. Yet, existing methods are specialized, focusing on either bare-hand or hand interacting with object. No method can flexibly handle both scenarios and their performance degrades when appli…

2025

What We Miss Matters: Learning from the Overlooked in Point Cloud Transformers

NeurIPS 2025poster

Point Cloud Transformers have become a cornerstone in 3D representation for their ability to model long-range dependencies via self-attention. However, these models tend to overemphasize salient regions while neglecting other informative regions, which limits feature diversity and compromises robust…

Cited by 0SourceScholar
2024

LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic Surgery

ICRA 2024poster

Visual question answering (VQA) can be fundamentally crucial for promoting robotic-assisted surgical education. In practice, the needs of trainees are constantly evolving, such as learning more surgical types and adapting to new surgical instruments/techniques. Therefore, continually updating the VQ…

Cited by 17SourcecodeScholar
2024

LoRAExit: Empowering Dynamic Modulation of LLMs in Resource-limited Settings using Low-rank Adapters

EMNLP 2024finding

Large Language Models (LLMs) have exhibited remarkable performance across various natural language processing tasks. However, deploying LLMs on resource-limited settings remains a challenge. While early-exit techniques offer an effective approach, they often require compromised training methods that…

Cited by 0SourcePDFScholar
2024

Neural P$^3$M: A Long-Range Interaction Modeling Enhancer for Geometric GNNs

NeurIPS 2024poster

Geometric graph neural networks (GNNs) have emerged as powerful tools for modeling molecular geometry. However, they encounter limitations in effectively capturing long-range interactions in large molecular systems. To address this challenge, we introduce **Neural P$^3$M**, a versatile enhancer of g…

2024

PCF-Lift: Panoptic Lifting by Probabilistic Contrastive Fusion

ECCV 2024poster

"Panoptic lifting is an effective technique to address the 3D panoptic segmentation task by unprojecting 2D panoptic segmentations from multi-views to 3D scene. However, the quality of its results largely depends on the 2D segmentations, which could be noisy and error-prone, so its performance often…

2024

PPN-Pack: Placement Proposal Network for Efficient Robotic Bin Packing

RA-L 2024

Robotic bin packing is a challenging task, requiring compactly packing objects in a container and also efficiently performing the computation, such that the robot arm need not wait too long before taking action. In this work, we introduce PPN-Pack, a novel learning-based approach to improve the effi

Cited by 6SourceScholar
2024

Sample-Efficient Multiagent Reinforcement Learning with Reset Replay

ICML 2024poster

The popularity of multiagent reinforcement learning (MARL) is growing rapidly with the demand for real-world tasks that require swarm intelligence. However, a noticeable drawback of MARL is its low sample efficiency, which leads to a huge amount of interactions with the environment. Surprisingly, fe…

Cited by 0SourcePDFScholar
2024

Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models

ECCV 2024poster

"This paper addresses the limitations of adverse weather image restoration approaches trained on synthetic data when applied to real-world scenarios. We formulate a semi-supervised learning framework employing vision-language models to enhance restoration performance across diverse adverse weather c…

2024

Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement Learning

NeurIPS 2024spotlight

As a marriage between offline RL and meta-RL, the advent of offline meta-reinforcement learning (OMRL) has shown great promise in enabling RL agents to multi-task and quickly adapt while acquiring knowledge safely. Among which, context-based OMRL (COMRL) as a popular paradigm, aims to learn a univer…

2024

Uncertainty-Aware Suction Grasping for Cluttered Scenes

RA-L 2024

In this work, we present a multi-stage pipeline that aims to accurately predict suction grasps for objects with varying properties in cluttered and complex scenes. Existing methods face difficulties in generalizing to unseen objects and effectively handling noisy depth/point cloud data, which often

Cited by 11SourcecodeScholar
2024

Unveiling the Generalization Power of Fine-Tuned Large Language Models

NAACL 2024long

While Large Language Models (LLMs) have demonstrated exceptional multitasking abilities, fine-tuning these models on downstream, domain-specific datasets is often necessary to yield superior performance on test sets compared to their counterparts without fine-tuning. However, the comprehensive effec…

2023

CauSSL: Causality-inspired Semi-supervised Learning for Medical Image Segmentation

ICCV 2023poster

Semi-supervised learning (SSL) has recently demonstrated great success in medical image segmentation, significantly enhancing data efficiency with limited annotations. However, despite its empirical benefits, there are still concerns in the literature about the theoretical foundation and explanation…

Cited by 57PDFcodeScholar
2023

Class-Conditional Sharpness-Aware Minimization for Deep Long-Tailed Recognition

CVPR 2023poster

It's widely acknowledged that deep learning models with flatter minima in its loss landscape tend to generalize better. However, such property is under-explored in deep long-tailed recognition (DLTR), a practical problem where the model is required to generalize equally well across all classes when…

2023

Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training

IJCAI 2023poster

Masked Autoencoders (MAE) have shown promising performance in self-supervised learning for both 2D and 3D computer vision. However, existing MAE-style methods can only learn from the data of a single modality, i.e., either images or point clouds, which neglect the implicit semantic and geometric cor…

Cited by 59SourcePDFScholar
2023

On Improving Boundary Quality of Instance Segmentation in Cluttered and Chaotic Scenarios

ICRA 2023poster

Instance segmentation is a long-standing task for supporting robotic bin picking. However, objects of diverse classes can be closely packed with occlusions in cluttered and chaotic scenes, hence, even recent methods could have difficulty in locating clear and precise boundaries to distinguish nearby…

Cited by 1SourceScholar
2023

On the Pitfall of Mixup for Uncertainty Calibration

CVPR 2023poster

By simply taking convex combinations between pairs of samples and their labels, mixup training has been shown to easily improve predictive accuracy. It has been recently found that models trained with mixup also perform well on uncertainty calibration. However, in this study, we found that mixup tra…

Cited by 16SourcePDFScholar
2023

Prototypical Variational Autoencoder for 3D Few-shot Object Detection

NeurIPS 2023poster

Few-Shot 3D Point Cloud Object Detection (FS3D) is a challenging task, aiming to detect 3D objects of novel classes using only limited annotated samples for training. Considering that the detection performance highly relies on the quality of the latent features, we design a VAE-based prototype learn…

Cited by 8SourcePDFScholar
2023

RepMode: Learning to Re-Parameterize Diverse Experts for Subcellular Structure Prediction

CVPR 2023highlight

In biological research, fluorescence staining is a key technique to reveal the locations and morphology of subcellular structures. However, it is slow, expensive, and harmful to cells. In this paper, we model it as a deep learning task termed subcellular structure prediction (SSP), aiming to predict…

2023

SDF-Pack: Towards Compact Bin Packing with Signed-Distance-Field Minimization

IROS 2023poster

Robotic bin packing is very challenging, especially when considering practical needs such as object variety and packing compactness. This paper presents SDF-Pack, a new approach based on signed distance field (SDF) to model the geometric condition of objects in a container and compute the object pla…

Cited by 11SourcecodeScholar
2023

Traj-MAE: Masked Autoencoders for Trajectory Prediction

ICCV 2023poster

Trajectory prediction has been a crucial task in building a reliable autonomous driving system by anticipating possible dangers. One key issue is to generate consistent trajectory predictions without colliding. To overcome the challenge, we propose an efficient masked autoencoder for trajectory pred…

Cited by 57PDFScholar
2023

Uncertainty Estimation by Fisher Information-based Evidential Deep Learning

ICML 2023poster

Uncertainty estimation is a key factor that makes deep learning reliable in practical applications. Recently proposed evidential neural networks explicitly account for different uncertainties by treating the network's outputs as evidence to parameterize the Dirichlet distribution, and achieve impres…

2023

Video Dehazing via a Multi-Range Temporal Alignment Network With Physical Prior

CVPR 2023poster

Video dehazing aims to recover haze-free frames with high visibility and contrast. This paper presents a novel framework to effectively explore the physical haze priors and aggregate temporal information. Specifically, we design a memory-based physical prior guidance module to encode the prior-relat…

2022

3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery

ICRA 2022poster

Automatic laparoscope motion control is fundamentally important for surgeons to efficiently perform operations. However, its traditional control methods based on tool tracking without considering information hidden in surgical scenes are not intelligent enough, while the latest supervised imitation…

Cited by 15SourceScholar
2022

A Sim-to-Real Object Recognition and Localization Framework for Industrial Robotic Bin Picking

RA-L 2022

We present a generic and robust sim-to-real deep-learning-based framework, namely S2R-Pick, for fast and accurate object recognition and localization in industrial robotic bin picking. Unlike existing works designed for general everyday environments, objects for industrial bin picking are often text

Cited by 59SourceScholar
2022

Acknowledging the Unknown for Multi-Label Learning with Single Positive Labels

ECCV 2022poster

"Due to the difficulty of collecting exhaustive multi-label annotations, multi-label datasets often contain partial labels. We consider an extreme of this weakly supervised learning problem, called single positive multi-label learning (SPML), where each multi-label training image has only one positi…

2022

Pseudo-label Guided Cross-video Pixel Contrast for Robotic Surgical Scene Segmentation with Limited Annotations

IROS 2022poster

Surgical scene segmentation is fundamentally crucial for prompting cognitive assistance in robotic surgery. However, pixel-wise annotating surgical video in a frame-by-frame manner is expensive and time consuming. To greatly reduce the labeling burden, in this work, we study semi-supervised scene se…

Cited by 6SourcecodeScholar
2022

RePFormer: Refinement Pyramid Transformer for Robust Facial Landmark Detection

IJCAI 2022poster

This paper presents a Refinement Pyramid Transformer (RePFormer) for robust facial landmark detection. Most facial landmark detectors focus on learning representative image features. However, these CNN-based feature representations are not robust enough to handle complex real-world scenarios due to…

Cited by 19SourcePDFScholar
2022

SESR: Self-Ensembling Sim-to-Real Instance Segmentation for Auto-Store Bin Picking

IROS 2022poster

Instance segmentation is an important task for supporting robotic grasping in auto-store scenarios. Accurate segmentation usually relies on the quantity and quality of available annotated training data. However, it requires tremendous cost to obtain these labels. In this work, without requiring any…

Cited by 2SourceScholar
2022

Single-Domain Generalization in Medical Image Segmentation via Test-Time Adaptation from Shape Dictionary

AAAI 2022technical

Domain generalization typically requires data from multiple source domains for model learning. However, such strong assumption may not always hold in practice, especially in medical field where the data sharing is highly concerned and sometimes prohibitive due to privacy issue. This paper studies th…

Cited by 47SourcePDFScholar
2022

Towards Robust Part-aware Instance Segmentation for Industrial Bin Picking

ICRA 2022poster

Industrial bin picking is a challenging task that requires accurate and robust segmentation of individual object instances. Particularly, industrial objects can have irregular shapes, that is, thin and concave, whereas in bin-picking scenarios, objects are often closely packed with strong occlusion.…

Cited by 15SourceScholar
2022

Transformer-based Working Memory for Multiagent Reinforcement Learning with Action Parsing

NeurIPS 2022accept

Learning in real-world multiagent tasks is challenging due to the usual partial observability of each agent. Previous efforts alleviate the partial observability by historical hidden states with Recurrent Neural Networks, however, they do not consider the multiagent characters that either the multia…

Cited by 20SourcePDFScholar
2021

Accurate Grid Keypoint Learning for Efficient Video Prediction

IROS 2021poster

Video prediction methods generally consume substantial computing resources in training and deployment, among which keypoint-based approaches show promising improvement in efficiency by simplifying dense image prediction to light keypoint prediction. However, keypoint locations are often modeled only…

Cited by 18SourcecodeScholar
2021

Beyond Class-Conditional Assumption: A Primary Attempt to Combat Instance-Dependent Label Noise

AAAI 2021technical

Supervised learning under label noise has seen numerous advances recently, while existing theoretical findings and empirical results broadly build up on the class-conditional noise (CCN) assumption that the noise is independent of input features given the true label. In this work, we present a theor…

2021

C3-SemiSeg: Contrastive Semi-Supervised Segmentation via Cross-Set Learning and Dynamic Class-Balancing

ICCV 2021poster

The semi-supervised semantic segmentation methods utilize the unlabeled data to increase the feature discriminative ability to alleviate the burden of the annotated data. However, the dominant consistency learning diagram is limited by a) the misalignment between features from labeled and unlabeled…

Cited by 98PDFScholar
2021

Domain Adaptive Robotic Gesture Recognition with Unsupervised Kinematic-Visual Data Alignment

IROS 2021poster

Automated surgical gesture recognition is of great importance in robot-assisted minimally invasive surgery. However, existing methods assume that training and testing data are from the same domain, which suffers from severe performance degradation when a domain gap exists, such as the simulator and…

Cited by 4SourceScholar
2021

FedDG: Federated Domain Generalization on Medical Image Segmentation via Episodic Learning in Continuous Frequency Space

CVPR 2021poster

Federated learning allows distributed medical institutions to collaboratively learn a shared prediction model with privacy protection. While at clinical deployment, the models trained in federated learning can still suffer from performance drop when applied to completely unseen hospitals outside the…

Cited by 586PDFcodeScholar
2021

Flattening Sharpness for Dynamic Gradient Projection Memory Benefits Continual Learning

NeurIPS 2021poster

The backpropagation networks are notably susceptible to catastrophic forgetting, where networks tend to forget previously learned skills upon learning new ones. To address such the 'sensitivity-stability' dilemma, most previous efforts have been contributed to minimizing the empirical risk with diff…

2021

Learning Semantic Context from Normal Samples for Unsupervised Anomaly Detection

AAAI 2021technical

Unsupervised anomaly detection aims to identify data samples that have low probability density from a set of input samples, and only the normal samples are provided for model training. The inference of abnormal regions on the input image requires an understanding of the surrounding semantic context.…

Cited by 179SourcePDFScholar
2021

Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence Learning

ICCV 2021poster

This paper presents a self-supervised method for learning reliable visual correspondence from unlabeled videos. We formulate the correspondence as finding paths in a joint space-time graph, where nodes are grid patches sampled from frames, and are linked by two type of edges: (i) neighbor relations…

Cited by 21PDFScholar
2021

Noise against noise: stochastic label noise helps combat inherent label noise

ICLR 2021spotlight

The noise in stochastic gradient descent (SGD) provides a crucial implicit regularization effect, previously studied in optimization by analyzing the dynamics of parameter updates. In this paper, we are interested in learning with noisy labels, where we have a collection of samples with potential mi…

Cited by 48SourcePDFScholar
2021

One to Many: Adaptive Instrument Segmentation via Meta Learning and Dynamic Online Adaptation in Robotic Surgical Video

ICRA 2021poster

Surgical instrument segmentation in robot-assisted surgery (RAS) - especially that using learning-based models - relies on the assumption that training and testing videos are sampled from the same domain. However, it is impractical and expensive to collect and annotate sufficient data from every new…

Cited by 26SourceScholar
2021

Robustness of Accuracy Metric and its Inspirations in Learning with Noisy Labels

AAAI 2021technical

For multi-class classification under class-conditional label noise, we prove that the accuracy metric itself can be robust. We concretize this finding's inspiration in two essential aspects: training and validation, with which we address critical issues in learning with noisy labels. For training, w…

2021

Single-Stage Instance Shadow Detection With Bidirectional Relation Learning

CVPR 2021poster

Instance shadow detection aims to find shadow instances paired with the objects that cast the shadows. The previous work adopts a two-stage framework to first predict shadow instances, object instances, and shadow-object associations from the region proposals, then leverage a post-processing to matc…

Cited by 36PDFScholar
2021

SurRoL: An Open-source Reinforcement Learning Centered and dVRK Compatible Platform for Surgical Robot Learning

IROS 2021poster

Autonomous surgical execution relieves tedious routines and surgeon’s fatigue. Recent learning-based methods, especially reinforcement learning (RL) based methods, achieve promising performance for dexterous manipulation, which usually requires the simulation to collect data efficiently and reduce t…

Cited by 95SourcecodeScholar
2020

A Learning-Driven Framework with Spatial Optimization For Surgical Suture Thread Reconstruction and Autonomous Grasping Under Multiple Topologies and Environmental Noises

IROS 2020poster

Surgical knot tying is one of the most fundamental and important procedures in surgery, and a high-quality knot can significantly benefit the postoperative recovery of the patient. However, a longtime operation may easily cause fatigue to surgeons, especially during the tedious wound closure task. I…

Cited by 17SourceScholar
2020

A Multi-Task Mean Teacher for Semi-Supervised Shadow Detection

CVPR 2020poster

Existing shadow detection methods suffer from an intrinsic limitation in relying on limited labeled datasets, and they may produce poor results in some complicated situations. To boost the shadow detection performance, this paper presents a multi-task mean teacher model for semi-supervised shadow de…

Cited by 190PDFcodeScholar
2020

Automatic Gesture Recognition in Robot-assisted Surgery with Reinforcement Learning and Tree Search

ICRA 2020poster

Automatic surgical gesture recognition is fundamental for improving intelligence in robot-assisted surgery, such as conducting complicated tasks of surgery surveillance and skill evaluation. However, current methods treat each frame individually and produce the outcomes without effective considerati…

Cited by 69SourceScholar
2020

Learning from Extrinsic and Intrinsic Supervisions for Domain Generalization

ECCV 2020poster

The generalization capability of neural networks across domains is crucial for real-world applications. We argue that a generalized object recognition system should well understand the relationships among different images and also the images themselves at the same time. To this end, we present a new…

Cited by 238SourcePDFScholar
2020

PointAugment: An Auto-Augmentation Framework for Point Cloud Classification

CVPR 2020oral

We present PointAugment, a new auto-augmentation framework that automatically optimizes and augments point cloud samples to enrich the data diversity when we train a classification network. Different from existing auto-augmentation methods for 2D images, PointAugment is sample-aware and takes an adv…

Cited by 230PDFcodeScholar
2019

Deep Multi-Model Fusion for Single-Image Dehazing

ICCV 2019poster

This paper presents a deep multi-model fusion network to attentively integrate multiple models to separate layers and boost the performance in single-image dehazing. To do so, we first formulate the attentional feature integration module to maximize the integration of the convolutional neural networ…

Cited by 146PDFScholar
2018

Bidirectional Feature Pyramid Network with Recurrent Attention Residual Modules for Shadow Detection

ECCV 2018poster

This paper presents a network to detect shadows by exploring and combining global context in deep layers and local context in shallow layers of a deep convolutional neural network (CNN). There are two technical contributions in our network design. First, we formulate the recurrent attention residual…

2018

Direction-Aware Spatial Context Features for Shadow Detection

CVPR 2018poster

Shadow detection is a fundamental and challenging task, since it requires an understanding of global image semantics and there are various backgrounds around shadows. This paper presents a novel network for shadow detection by analyzing image context in a direction-aware manner. To achieve this, we…

Cited by 484SourcePDFScholar
2018

EC-Net: an Edge-aware Point set Consolidation Network

ECCV 2018poster

Point clouds obtained from 3D scans are typically sparse, irregular, and noisy, and required to be consolidated. In this paper, we present the first deep learning based {em edge-aware} technique to facilitate the consolidation of point clouds. We design our network to process points grouped in local…

Cited by 340SourcePDFScholar
2018

PU-Net: Point Cloud Upsampling Network

CVPR 2018poster

Learning and analyzing 3D point clouds with deep networks is challenging due to the sparseness and irregularity of the data. In this paper, we present a data-driven point cloud upsampling technique. The key idea is to learn multi-level features per point and expand the point set via a multi-branch c…

2017

Cascaded Feature Network for Semantic Segmentation of RGB-D Images

ICCV 2017poster

Fully convolutional network (FCN) has been successfully applied in semantic segmentation of scenes represented with RGB images. Images augmented with depth channel provide more understanding of the geometric information of the scene in the image. The question is how to best exploit this additional i…

Cited by 175PDFScholar