← Search

Hang Su

152 accepted papers

2026

4DP-QA: Scalable QA for 4D Perception in Vision Language Models

CVPR 2026

Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D scene, challenging in itself, is further complicated by two factors. First, VLMs observe motion indirectly via its projection onto 2D images. Second

Cited by 0SourceScholar
2026

AVO-QP: Task-Adaptive Real-Time Obstacle Avoidance for Redundant Manipulators on Edge Platforms

RA-L 2026

We address real-time obstacle avoidance for redundant manipulators where tracking and safety constraints can conflict and render quadratic programs (QPs) infeasible. We propose AVO-QP, a sensor-guided velocity-level planner that fuses RGB-D depth with learned detection to maintain situational awaren

Cited by 0SourceScholar
2026

Benchmarking Trustworthiness in Multimodal LLMs for Video Understanding

AAAI 2026technical

Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability.

Cited by 0SourcePDFScholar
2026

Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual Segmentation

CVPR 2026

Audio-Visual Segmentation (AVS) aims to accurately segment sounding objects in video frames by leveraging audio-visual correspondence cues. However, it remains challenging due to the intrinsic semantic incompleteness within a single modality and the semantic gap between audio and visual representati

Cited by 0SourceScholar
2026

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Video Generation

ICML 2026poster

To achieve real-time video generation, current approaches distill pretrained bidirectional video diffusion models into few-step autoregressive (AR) models. This process involves an *architectural gap*, as it converts full attention into causal attention. In this paper, we demonstrate that existing m…

Cited by 0SourceScholar
2026

Cross-task Calibration for Asynchronous Federated Continual Learning

ICML 2026poster

Federated Continual Learning (FCL) aims to empower distributed devices to learn a sequence of tasks over time. However, existing FCL research largely relies on the impractical assumption of synchronous new task arrival. This overlooks the reality of asynchronous user behavior and system latencies, f…

Cited by 0SourceScholar
2026

DiffusionNFT: Online Diffusion Reinforcement with Forward Process

ICLR 2026oral

Online reinforcement learning (RL) has been central to post-training language models, but its extension to diffusion models remains challenging due to intractable likelihoods. Recent works discretize the reverse sampling process to enable GRPO-style training, yet they inherit fundamental drawbacks,…

Cited by 0SourcecodeScholar
2026

Dual-Seed Evolutionary Algorithm for Noise Optimization in Diffusion Models

AAAI 2026technical

Diffusion models have emerged as state-of-the-art generative methods, particularly excelling in conditional tasks such as prompt-driven image synthesis. While recent research emphasizes the pivotal role of noise seeds in enhancing text-image alignment and generating human-preferred outputs,these wor

Cited by 0SourcePDFScholar
2026

Exploratory Diffusion Model for Unsupervised Reinforcement Learning

ICLR 2026oral

Unsupervised reinforcement learning (URL) pre-trains agents by exploring diverse states in reward-free environments, aiming to enable efficient adaptation to various downstream tasks. Without extrinsic rewards, prior methods rely on intrinsic objectives, but heterogeneous exploration data demand str…

Cited by 0SourcecodeScholar
2026

FedCD: Towards Consolidated Distillation for Heterogeneous Federated Learning

AAAI 2026technical

Knowledge Distillation (KD) serves as an effective approach to addressing heterogeneity issues in Federated Learning (FL), leveraging additional datasets to align local and global models better. There are two primary distillation paradigms: feature-based distillation, which utilizes intermediate-lay

Cited by 0SourcePDFScholar
2026

H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation

AAAI 2026technical

Imitation learning for robotic manipulation faces a fundamental challenge: the scarcity of large-scale, high-quality robot demonstration data. Recent robotic foundation models often pre-train on cross-embodiment robot datasets to increase data scale, while they face significant limitations as the di

Cited by 0SourcePDFScholar
2026

Helix: Evolutionary Reinforcement Learning for Open-Ended Scientific Problem Solving

ICLR 2026poster

Large language models (LLMs) with reasoning abilities have demonstrated growing promise for tackling complex scientific problems. Yet such tasks are inherently domain-specific, unbounded and open-ended, demanding exploration across vast and flexible solution spaces. Existing approaches, whether pure…

Cited by 0SourceScholar
2026

Lightweight Federated Incremental Learning via Decoupled Replay

ICML 2026poster

Federated Incremental Learning (FIL) aims to learn streaming tasks across distributed clients without catastrophic forgetting while preserving privacy. Most existing methods focus on sample-based replay techniques, which mitigate forgetting by replaying historical data samples. However, such methods…

Cited by 0SourceScholar
2026

MePo: Meta Post-Refinement for Rehearsal-Free General Continual Learning

ICML 2026poster

To cope with uncertain changes of the external world, intelligent systems must continually learn from complex, evolving environments and respond in real time. This ability, collectively known as general continual learning (GCL), encapsulates practical challenges such as online datastreams and blurry…

Cited by 0SourceScholar
2026

Motus: A Unified Latent Action World Model

CVPR 2026

While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and hinders learning from large-scale, heterogeneous data. In this pa

Cited by 0SourcecodeScholar
2026

RC-FCL: Combating Asynchronous Concept Drift in Federated Continual Learning via Retrospective Calibration

ICML 2026poster

Federated Continual Learning (FCL) enables the continuous acquisition of knowledge from streaming tasks, but inherently struggles with the temporal dynamics of client data distributions. These dynamics naturally induce asynchronous concept drift, where distribution shifts occur independently across …

Cited by 0SourceScholar
2026

RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization

ICML 2026poster

Vision-Language-Action (VLA) models hold promise for generalist robotics but currently struggle with data scarcity, architectural inefficiencies, and the inability to generalize across different hardware platforms. We introduce RDT2, a robotic foundation model built upon a 7B parameter VLM designed …

Cited by 0SourceScholar
2026

ReflexDiffusion: Reflection-Enhanced Trajectory Planning for High-lateral-acceleration Scenarios in Autonomous Driving

AAAI 2026technical

Generating safe and reliable trajectories for autonomous vehicles in long-tail scenarios remains a significant challenge, particularly for High-lateral-acceleration maneuvers such as sharp turns that represent critical safety situations. Existing trajectory planners exhibit systematic failures in th

Cited by 0SourcePDFScholar
2026

Symmetry-Aware Skill Transfer with Energy-Tank Passive Control for Ankle Exoskeletons

ICRA 2026poster

This paper presents a unified framework that combines symmetry-aware skill transfer with energy-tank passive control to achieve safe and adaptive ankle exoskeleton assistance. Subject-specific ankle references are first extracted from wearable IMU data : Dynamic Time Warping (DTW) aligns gait cycles…

Cited by 0Scholar
2026

Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention

ICLR 2026poster

Although Large Reasoning Models (LRMs) have progressed in solving complex problems, their chain-of-thought (CoT) reasoning often contains harmful content that can persist even when the final responses appear safe. We show that this issue still remains in existing methods which overlook the unique si…

Cited by 0SourceScholar
2025

A Linear N-Point Solver for Structure and Motion from Asynchronous Tracks

ICCV 2025poster

Structure and continuous motion estimation from point correspondences is a fundamental problem in computer vision that has been powered by well-known algorithms such as the familiar 5-point or 8-point algorithm. However, despite their acclaim, these algorithms are limited to processing point corresp…

Cited by 0SourcePDFScholar
2025

Accelerating PDE-Constrained Optimization by the Derivative of Neural Operators

ICML 2025poster

PDE-Constrained Optimization (PDECO) problems can be accelerated significantly by employing gradient-based methods with surrogate models like neural operators compared to traditional numerical solvers. However, this approach faces two key challenges: (1) **Data inefficiency**: Lack of efficient dat…

Cited by 0SourcePDFScholar
2025

AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?

ICCV 2025poster

Vision Language Models (VLMs) have exhibited remarkable generalization capabilities, yet their robustness in dynamic real-world scenarios remains largely unexplored. To systematically evaluate VLMs' robustness to real-world 3D variations, we propose AdvDreamer, the first framework capable of generat…

Cited by 0SourcePDFScholar
2025

AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization via Multi-LLMs

NAACL 2025findings

Recent studies show that large language models (LLMs) are vulnerable to jailbreak attacks, which can bypass their defense mechanisms. However, existing jailbreak research often exhibits limitations in universality, validity, and efficiency. Therefore, we rethink jailbreaking LLMs and define three ke…

2025

Exploring the Generalizability of Factual Hallucination Mitigation via Enhancing Precise Knowledge Utilization

EMNLP 2025

Large Language Models (LLMs) often struggle to align their responses with objective facts, resulting in the issue of factual hallucinations , which can be difficult to detect and mislead users without relevant knowledge. Although post-training techniques have been employed to mitigate the issue, exi

2025

Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation

NAACL 2025long

Faithfulness evaluators based on Large Language Models (LLMs) are often fooled by the fluency of the text and struggle with identifying errors in the summaries, usually leading to high false negative rate. We propose an approach to summary faithfulness evaluation in which multiple LLM-based agents a…

2025

Learning to Summarize from LLM-generated Feedback

NAACL 2025long

Developing effective text summarizers remains a challenge due to issues like hallucinations, key information omissions, and verbosity in LLM-generated summaries. This work explores using LLM-generated feedback to improve summary quality by aligning the summaries with human preferences for faithfulne…

Cited by 4SourcePDFScholar
2025

MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization

EMNLP 2025

Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods are essential for reducing both cost and human effort. However, such methods require a strong meta-evaluation benchmark g

Cited by 0SourcePDFScholar
2025

Personalized Question Answering with User Profile Generation and Compression

EMNLP 2025

Large language models (LLMs) offer a novel and convenient avenue for humans to acquire knowledge. However, LLMs are prone to providing “midguy” answers regardless of users’ knowledge background, thereby failing to meet each user’s personalized needs. To tackle the problem, we propose to generate per

2025

RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

ICLR 2025poster

Bimanual manipulation is essential in robotics, yet developing foundation models is extremely challenging due to the inherent complexity of coordinating two robot arms (leading to multi-modal action distributions) and the scarcity of training data. In this paper, we present the Robotics Diffusion Tr…

2025

RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation

IROS 2025

Visual augmentation has become a crucial technique for enhancing the visual robustness of imitation learning. However, existing methods are often limited by prerequisites such as camera calibration or the need for controlled environments (e.g., green screen setups). In this work, we introduce RoboEn

Cited by 36SourcecodeScholar
2025

Self-Consistent Model-based Adaptation for Visual Reinforcement Learning

IJCAI 2025

Visual reinforcement learning agents typically face serious performance declines in real-world applications caused by visual distractions. Existing methods rely on fine-tuning the policy's representations with hand-crafted augmentations. In this work, we propose Self-Consistent Model-based Adaptatio

Cited by 0SourcePDFScholar
2025

Toward Guidance-Free AR Visual Generation via Condition Contrastive Alignment

ICLR 2025oral

Classifier-Free Guidance (CFG) is a critical technique for enhancing the sample quality of visual generative models. However, in autoregressive (AR) multi-modal generation, CFG introduces design inconsistencies between language and visual content, contradicting the design philosophy of unifying diff…

Cited by 2SourcePDFScholar
2025

Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages

ACL 2025long

Evaluation frameworks for text summarization have evolved in terms of both domain coverage and metrics. However, existing benchmarks still lack domain-specific assessment criteria, remain predominantly English-centric, and face challenges with human annotation due to the complexity of reasoning. To…

2025

Understanding and Improving Information Preservation in Prompt Compression for LLMs

EMNLP 2025

Recent advancements in large language models (LLMs) have enabled their successful application to a broad range of tasks. However, in information-intensive tasks, the prompt length can grow fast, leading to increased computational requirements, performance degradation, and induced biases from irrelev

2025

Zero-Shot Monocular Scene Flow Estimation in the Wild

CVPR 2025award

Large models have shown generalization across datasets for many low-level vision tasks, like depth estimation, but no such general models exist for scene flow.Even though scene flow prediction has wide potential, its practical use is limited because of the lack of generalization of current predictiv…

Cited by 1SourcePDFScholar
2024

Aligning Diffusion Behaviors with Q-functions for Efficient Continuous Control

NeurIPS 2024poster

Drawing upon recent advances in language model alignment, we formulate offline Reinforcement Learning as a two-stage optimization problem: First pretraining expressive generative policies on reward-free behavior datasets, then finetuning these policies to align with task-specific annotations like Q-…

2024

An N-Point Linear Solver for Line and Motion Estimation with Event Cameras

CVPR 2024poster

Event cameras respond primarily to edges---formed by strong gradients---and are thus particularly well-suited for line-based motion estimation. Recent work has shown that events generated by a single line each satisfy a polynomial constraint which describes a manifold in the space-time volume. Multi…

Cited by 9SourcePDFScholar
2024

CERET: Cost-Effective Extrinsic Refinement for Text Generation

NAACL 2024long

Large Language Models (LLMs) are powerful models for generation tasks, but they may not generate good quality outputs in their first attempt. Apart from model fine-tuning, existing approaches to improve prediction accuracy and quality typically involve LLM self-improvement / self-reflection that inc…

2024

CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model

ECCV 2024poster

"Feed-forward 3D generative models like the Large Reconstruction Model (LRM) [?] have demonstrated exceptional generation speed. However, the transformer-based methods do not leverage the geometric priors of the triplane component in their architecture, often leading to sub-optimal quality given the…

2024

Can Your Model Tell a Negation from an Implicature? Unravelling Challenges With Intent Encoders

ACL 2024long

Conversational systems often rely on embedding models for intent classification and intent clustering tasks. The advent of Large Language Models (LLMs), which enable instructional embeddings allowing one to adjust semantics over the embedding space using prompts, are being viewed as a panacea for th…

Cited by 2SourcePDFScholar
2024

Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights

ECCV 2024poster

"(CIC) evolves traditional image captioning into a more complex domain, necessitating the ability for multimodal reasoning. It aims to generate image captions given specific contextual information. This paper further introduces a novel domain of (). Unlike CIC, which solely relies on broad context,…

2024

Controllable Navigation Instruction Generation with Chain of Thought Prompting

ECCV 2024poster

"Instruction generation is a vital and multidisciplinary research area with broad applications. Existing instruction generation models are limited to generating instructions in a single style from a particular dataset, and the style and content of generated instructions cannot be controlled. Moreove…

2024

DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-Training

ICML 2024poster

Pre-training has been investigated to improve the efficiency and performance of training neural operators in data-scarce settings. However, it is largely in its infancy due to the inherent complexity and diversity, such as long trajectories, multiple scales and varying dimensions of partial differen…

2024

Diffusion Models are Certifiably Robust Classifiers

NeurIPS 2024poster

Generative learning, recognized for its effective modeling of data distributions, offers inherent advantages in handling out-of-distribution instances, especially for enhancing robustness to adversarial attacks. Among these, diffusion classifiers, utilizing powerful diffusion models, have demonstrat…

2024

Embodied Active Defense: Leveraging Recurrent Feedback to Counter Adversarial Patches

ICLR 2024poster

The vulnerability of deep neural networks to adversarial patches has motivated numerous defense strategies for boosting model robustness. However, the prevailing defenses depend on single observation or pre-established adversary information to counter adversarial patches, often failing to be confron…

Cited by 3SourcePDFScholar
2024

Exploring the Transferability of Visual Prompting for Multimodal Large Language Models

CVPR 2024highlight

Although Multimodal Large Language Models (MLLMs) have demonstrated promising versatile capabilities their performance is still inferior to specialized models on downstream tasks which makes adaptation necessary to enhance their utility. However fine-tuning methods require independent training for e…

2024

FineSurE: Fine-grained Summarization Evaluation using LLMs

ACL 2024long

Automated evaluation is crucial for streamlining text summarization benchmarking and model development, given the costly and time-consuming nature of human evaluation. Traditional methods like ROUGE do not correlate well with human judgment, while recently proposed LLM-based metrics provide only sum…

2024

Fourier Controller Networks for Real-Time Decision-Making in Embodied Learning

ICML 2024poster

Transformer has shown promise in reinforcement learning to model time-varying features for obtaining generalized low-level robot policies on diverse robotics datasets in embodied learning. However, it still suffers from the issues of low data efficiency and high inference latency. In this paper, we…

2024

Full-Distance Evasion of Pedestrian Detectors in the Physical World

NeurIPS 2024poster

Many studies have proposed attack methods to generate adversarial patterns for evading pedestrian detection, alarming the computer vision community about the need for more attention to the robustness of detectors. However, adversarial patterns optimized by these methods commonly have limited perform…

2024

Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

ECCV 2024poster

"In this paper, we develop an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category names or referring expressions. The key solution of open-set object detection i…

2024

Human-Robot Collaboration Through a Multi-Scale Graph Convolution Neural Network With Temporal Attention

RA-L 2024

Collaborative robots sensing and understanding the movements and intentions of their human partners are crucial for realizing human-robot collaboration. Human skeleton sequences are widely recognized as a kind of data with great application potential in human action recognition. In this letter, a mu

Cited by 19SourceScholar
2024

Improved Operator Learning by Orthogonal Attention

ICML 2024spotlight

This work presents orthogonal attention for constructing neural operators to serve as surrogates to model the solutions of a family of Partial Differential Equations (PDEs). The motivation is that the kernel integral operator, which is usually at the core of neural operators, can be reformulated wit…

2024

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

ECCV 2024poster

"This paper presents (), a general-purpose multimodal assistant trained using an end-to-end approach that systematically expands the capabilities of large multimodal models (LMMs). maintains a skill repository that contains a wide range of vision and vision-language pre-trained models (tools), and i…

2024

Layer-Aware Analysis of Catastrophic Overfitting: Revealing the Pseudo-Robust Shortcut Dependency

ICML 2024poster

Catastrophic overfitting (CO) presents a significant challenge in single-step adversarial training (AT), manifesting as highly distorted deep neural networks (DNNs) that are vulnerable to multi-step adversarial attacks. However, the underlying factors that lead to the distortion of decision boundari…

2024

MAGID: An Automated Pipeline for Generating Synthetic Multi-modal Datasets

NAACL 2024long

Development of multimodal interactive systems is hindered by the lack of rich, multimodal (text, images) conversational data, which is needed in large quantities for LLMs. Previous approaches augment textual dialogues with retrieved images, posing privacy, diversity, and quality constraints. In this…

2024

Machine Vision Therapy: Multimodal Large Language Models Can Enhance Visual Robustness via Denoising In-Context Learning

ICML 2024poster

Although pre-trained models such as Contrastive Language-Image Pre-Training (CLIP) show impressive generalization results, their robustness is still limited under Out-of-Distribution (OOD) scenarios. Instead of undesirably leveraging human annotation as commonly done, it is possible to leverage the…

Cited by 15SourcePDFScholar
2024

Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood Discrepancy

NeurIPS 2024poster

Text-to-image diffusion models have achieved tremendous success in the field of controllable image generation, while also coming along with issues of privacy leakage and data copyrights. Membership inference arises in these contexts as a potential auditing method for detecting unauthorized data usag…

2024

MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models

NeurIPS 2024poster

Despite the superior capabilities of Multimodal Large Language Models (MLLMs) across diverse tasks, they still face significant trustworthiness challenges. Yet, current literature on the assessment of trustworthy MLLMs remains limited, lacking a holistic evaluation to offer thorough insights into fu…

Cited by 5SourcecodeScholar
2024

Noise Contrastive Alignment of Language Models with Explicit Rewards

NeurIPS 2024poster

User intentions are typically formalized as evaluation rewards to be maximized when fine-tuning language models (LMs). Existing alignment methods, such as Direct Preference Optimization (DPO), are mainly tailored for pairwise preference data where rewards are implicitly defined rather than explicitl…

2024

Omniview-Tuning: Boosting Viewpoint Invariance of Vision-Language Pre-training Models

ECCV 2024oral

"Vision-Language Pre-training (VLP) models like CLIP have achieved remarkable success in computer vision and particularly demonstrated superior robustness to distribution shifts of 2D images. However, their robustness under 3D viewpoint variations is still limited, which can hinder the development f…

2024

PEAC: Unsupervised Pre-training for Cross-Embodiment Reinforcement Learning

NeurIPS 2024poster

Designing generalizable agents capable of adapting to diverse embodiments has achieved significant attention in Reinforcement Learning (RL), which is critical for deploying RL agents in various real-world applications. Previous Cross-Embodiment RL approaches have focused on transferring knowledge ac…

2024

PINNacle: A Comprehensive Benchmark of Physics-Informed Neural Networks for Solving PDEs

NeurIPS 2024poster

While significant progress has been made on Physics-Informed Neural Networks (PINNs), a comprehensive comparison of these methods across a wide range of Partial Differential Equations (PDEs) is still lacking. This study introduces PINNacle, a benchmarking tool designed to fill this gap. PINNacle pro…

2024

Reference Neural Operators: Learning the Smooth Dependence of Solutions of PDEs on Geometric Deformations

ICML 2024poster

For partial differential equations on domains of arbitrary shapes, existing works of neural operators attempt to learn a mapping from geometries to solutions. It often requires a large dataset of geometry-solution pairs in order to obtain a sufficiently accurate neural operator. However, for many in…

Cited by 2SourcePDFScholar
2024

Rethinking Model Ensemble in Transfer-based Adversarial Attacks

ICLR 2024poster

It is widely recognized that deep learning models lack robustness to adversarial examples. An intriguing property of adversarial examples is that they can transfer across different models, which enables black-box attacks without any knowledge of the victim model. An effective strategy to improve the…

2024

Robust Classification via a Single Diffusion Model

ICML 2024poster

Diffusion models have been applied to improve adversarial robustness of image classifiers by purifying the adversarial noises or generating realistic data for adversarial training. However, diffusion-based purification can be evaded by stronger adaptive attacks while adversarial training does not pe…

Cited by 65SourcePDFScholar
2024

Score Regularized Policy Optimization through Diffusion Behavior

ICLR 2024poster

Recent developments in offline reinforcement learning have uncovered the immense potential of diffusion modeling, which excels at representing heterogeneous behavior policies. However, sampling from diffusion policies is considerably slow because it necessitates tens to hundreds of iterative inferen…

2024

Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection

NAACL 2024long

Semi-supervised dialogue summarization (SSDS) leverages model-generated summaries to reduce reliance on human-labeled data and improve the performance of summarization models. While addressing label noise, previous works on semi-supervised learning primarily focus on natural language understanding t…

2024

TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization

NAACL 2024long

Single document news summarization has seen substantial progress on faithfulness in recent years, driven by research on the evaluation of factual consistency, or hallucinations. We ask whether these advances carry over to other text summarization domains. We propose a new evaluation benchmark on top…

2024

Towards Transferable Targeted 3D Adversarial Attack in the Physical World

CVPR 2024poster

Compared with transferable untargeted attacks transferable targeted adversarial attacks could specify the misclassification categories of adversarial samples posing a greater threat to security-critical tasks. In the meanwhile 3D adversarial samples due to their potential of multi-view robustness ca…

2024

UniSumEval: Towards Unified, Fine-grained, Multi-dimensional Summarization Evaluation for LLMs

EMNLP 2024finding

Existing benchmarks for summarization quality evaluation often lack diverse input scenarios, focus on narrowly defined dimensions (e.g., faithfulness), and struggle with subjective and coarse-grained annotation schemes. To address these shortcomings, we create UniSumEval benchmark, which extends the…

2023

A 5-Point Minimal Solver for Event Camera Relative Motion Estimation

ICCV 2023oral

Event-based cameras are ideal for line-based motion estimation, since they predominantly respond to edges in the scene. However, accurately determining the camera displacement based on events continues to be an open problem. This is because line feature extraction and dynamics estimation are tightly…

Cited by 11PDFScholar
2023

All Are Worth Words: A ViT Backbone for Diffusion Models

CVPR 2023poster

Vision transformers (ViT) have shown promise in various vision tasks while the U-Net based on a convolutional neural network (CNN) remains dominant in diffusion models. We design a simple and general ViT-based architecture (named U-ViT) for image generation with diffusion models. U-ViT is characteri…

2023

Benchmarking Robustness of 3D Object Detection to Common Corruptions

CVPR 2023poster

3D object detection is an important task in autonomous driving to perceive the surroundings. Despite the excellent performance, the existing 3D detectors lack the robustness to real-world corruptions caused by adverse weathers, sensor noises, etc., provoking concerns about the safety and reliability…

2023

Bi-level Physics-Informed Neural Networks for PDE Constrained Optimization using Broyden's Hypergradients

ICLR 2023poster

Deep learning based approaches like Physics-informed neural networks (PINNs) and DeepONets have shown promise on solving PDE constrained optimization (PDECO) problems. However, existing methods are insufficient to handle those PDE constraints that have a complicated or nonlinear dependency on optim…

Cited by 19SourcePDFScholar
2023

COCO-O: A Benchmark for Object Detectors under Natural Distribution Shifts

ICCV 2023poster

Practical object detection application can lose its effectiveness on image inputs with natural distribution shifts. This problem leads the research community to pay more attention on the robustness of detectors under Out-Of-Distribution (OOD) inputs. Existing works construct datasets to benchmark th…

Cited by 23PDFcodeScholar
2023

Contrastive Energy Prediction for Exact Energy-Guided Diffusion Sampling in Offline Reinforcement Learning

ICML 2023poster

Guided sampling is a vital approach for applying diffusion models in real-world tasks that embeds human-defined guidance during the sampling procedure. This paper considers a general setting where the guidance is defined by an (unnormalized) energy function. The main challenge for this setting is th…

2023

DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

ICLR 2023poster

We present DINO (DETR with Improved deNoising anchOr boxes), a strong end-to-end object detector. DINO improves over previous DETR-like models in performance and efficiency by using a contrastive way for denoising training, a look forward twice scheme for box prediction, and a mixed query selection…

2023

DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding

AAAI 2023technical

In this paper, we study the problem of visual grounding by considering both phrase extraction and grounding (PEG). In contrast to the previous phrase-known-at-test setting, PEG requires a model to extract phrases from text and locate objects from image simultaneously, which is a more practical setti…

2023

Detection Transformer with Stable Matching

ICCV 2023poster

This paper is concerned with the matching stability problem across different decoder layers in DEtection TRansformers (DETR). We point out that the unstable matching in DETR is caused by a multi-optimization path problem, which is highlighted by the one-to-one matching design in DETR. To address thi…

Cited by 46PDFcodeScholar
2023

Dynamic Speech Endpoint Detection with Regression Targets

ICASSP 2023accepted

Interactive voice assistants have been widely used as input interfaces in various scenarios, e.g. on smart home devices, wearables and on AR devices. Detecting the end of a speech query, i.e. speech end-pointing, is an important task for voice assistants to interact with users. Traditionally, speech…

Cited by 0SourceScholar
2023

Enhancing Abstractiveness of Summarization Models through Calibrated Distillation

EMNLP 2023long findings

In this paper, we propose a novel approach named DisCal to enhance the level of abstractiveness (measured by n-gram overlap) without sacrificing the informativeness (measured by ROUGE) of generated summaries. DisCal exposes diverse pseudo summaries with two supervision to the student model. Firstly,…

Cited by 0SourceScholar
2023

GNOT: A General Neural Operator Transformer for Operator Learning

ICML 2023poster

Learning partial differential equations' (PDEs) solution operators is an essential problem in machine learning. However, there are several challenges for learning operators in practical applications like the irregular mesh, multiple input functions, and complexity of the PDEs' solution. To address t…

2023

Hierarchical Decomposition of Prompt-Based Continual Learning: Rethinking Obscured Sub-optimality

NeurIPS 2023spotlight

Prompt-based continual learning is an emerging direction in leveraging pre-trained knowledge for downstream continual learning, and has almost reached the performance pinnacle under supervised pre-training. However, our empirical research reveals that the current strategies fall short of their full…

2023

Meta-Reinforcement Learning Based on Self-Supervised Task Representation Learning

AAAI 2023technical

Meta-reinforcement learning enables artificial agents to learn from related training tasks and adapt to new tasks efficiently with minimal interaction data. However, most existing research is still limited to narrow task distributions that are parametric and stationary, and does not consider out-of-…

Cited by 16SourcePDFScholar
2023

MultiAdam: Parameter-wise Scale-invariant Optimizer for Multiscale Training of Physics-informed Neural Networks

ICML 2023poster

Physics-informed Neural Networks (PINNs) have recently achieved remarkable progress in solving Partial Differential Equations (PDEs) in various fields by minimizing a weighted sum of PDE loss and boundary loss. However, there are several critical challenges in the training of PINNs, including the la…

Cited by 22SourcePDFScholar
2023

NUNO: A General Framework for Learning Parametric PDEs with Non-Uniform Data

ICML 2023poster

The neural operator has emerged as a powerful tool in learning mappings between function spaces in PDEs. However, when faced with real-world physical data, which are often highly non-uniformly distributed, it is challenging to use mesh-based techniques such as the FFT. To address this, we introduce…

2023

Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

ICLR 2023poster

In offline reinforcement learning, weighted regression is a common method to ensure the learned policy stays close to the behavior policy and to prevent selecting out-of-sample actions. In this work, we show that due to the limited distributional expressivity of policy models, previous methods might…

2023

On the Reuse Bias in Off-Policy Reinforcement Learning

IJCAI 2023poster

Importance sampling (IS) is a popular technique in off-policy evaluation, which re-weights the return of trajectories in the replay buffer to boost sample efficiency. However, training with IS can be unstable and previous attempts to address this issue mainly focus on analyzing the variance of IS. I…

2023

One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

ICML 2023poster

This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is -- learning diffusion models for marginal, conditional, and joint distributions can be unified as predicting the noise in the perturb…

2023

Overcoming Recency Bias of Normalization Statistics in Continual Learning: Balance and Adaptation

NeurIPS 2023poster

Continual learning entails learning a sequence of tasks and balancing their knowledge appropriately. With limited access to old training samples, much of the current work in deep neural networks has focused on overcoming catastrophic forgetting of old tasks in gradient-based optimization. However, t…

2023

ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation

NeurIPS 2023spotlight

Score distillation sampling (SDS) has shown great promise in text-to-3D generation by distilling pretrained large-scale text-to-image diffusion models, but suffers from over-saturation, over-smoothing, and low-diversity problems. In this work, we propose to model the 3D parameter as a random variabl…

2023

Towards Effective Adversarial Textured 3D Meshes on Physical Face Recognition

CVPR 2023highlight

Face recognition is a prevailing authentication solution in numerous biometric applications. Physical adversarial attacks, as an important surrogate, can identify the weaknesses of face recognition systems and evaluate their robustness before deployed. However, most existing physical attacks are eit…

2023

Towards Viewpoint-Invariant Visual Recognition via Adversarial Training

ICCV 2023poster

Visual recognition models are not invariant to viewpoint changes in the 3D world, as different viewing directions can dramatically affect the predictions given the same object. Although many efforts have been devoted to making neural networks invariant to 2D image translations and rotations, viewpoi…

Cited by 11PDFScholar
2022

A Multitask Learning Framework for Speaker Change Detection with Content Information from Unsupervised Speech Decomposition

ICASSP 2022accepted

Speaker Change Detection (SCD) is a task of determining the time boundaries between speech segments of different speakers. SCD system can be applied to many tasks, such as speaker diarization, speaker tracking, and transcribing audio with multiple speakers. Recent advancements in deep learning lead…

Cited by 0SourceScholar
2022

A Unified Hard-Constraint Framework for Solving Geometrically Complex PDEs

NeurIPS 2022accept

We present a unified hard-constraint framework for solving geometrically complex PDEs with neural networks, where the most commonly used Dirichlet, Neumann, and Robin boundary conditions (BCs) are considered. Specifically, we first introduce the "extra fields'' from the mixed finite element method t…

2022

Boosting Transferability of Targeted Adversarial Examples via Hierarchical Generative Networks

ECCV 2022poster

"Transfer-based adversarial attacks can evaluate model robustness in the black-box setting. Several methods have demonstrated impressive untargeted transferability, however, it is still challenging to efficiently produce targeted transferability. To this end, we develop a simple yet effective framew…

2022

Cluster Attack: Query-based Adversarial Attacks on Graph with Graph-Dependent Priors

IJCAI 2022poster

While deep neural networks have achieved great success in graph analysis, recent work has shown that they are vulnerable to adversarial attacks. Compared with adversarial attacks on image classification, performing adversarial attacks on graphs is more challenging because of the discrete and non-dif…

Cited by 18SourcePDFScholar
2022

DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

ICLR 2022poster

We present in this paper a novel query formulation using dynamic anchor boxes for DETR (DEtection TRansformer) and offer a deeper understanding of the role of queries in DETR. This new formulation directly uses box coordinates as queries in Transformer decoders and dynamically updates them layer by…

2022

Exploring Memorization in Adversarial Training

ICLR 2022poster

Deep learning models have a propensity for fitting the entire training set even with random labels, which requires memorization of every training sample. In this paper, we explore the memorization effect in adversarial training (AT) for promoting a deeper understanding of model capacity, convergence…

2022

GSmooth: Certified Robustness against Semantic Transformations via Generalized Randomized Smoothing

ICML 2022spotlight

Certified defenses such as randomized smoothing have shown promise towards building reliable machine learning systems against $\ell_p$ norm bounded attacks. However, existing methods are insufficient or unable to provably defend against semantic transformations, especially those without closed-form…

Cited by 31SourcePDFScholar
2022

Human-Robot Shared Control for Surgical Robot Based on Context-Aware Sim-to-Real Adaptation

ICRA 2022poster

Human-robot shared control, which integrates the advantages of both humans and robots, is an effective approach to facilitate efficient surgical operation. Learning from demonstration (LfD) techniques can be used to automate some of the surgical sub tasks for the construction of the shared control m…

Cited by 43SourceScholar
2022

Policy Learning for Robust Markov Decision Process with a Mismatched Generative Model

AAAI 2022technical

In high-stake scenarios like medical treatment and auto-piloting, it's risky or even infeasible to collect online experimental data to train the agent. Simulation-based training can alleviate this issue, but may suffer from its inherent mismatches from the simulator and real environment. It is there…

Cited by 8SourcePDFScholar
2022

Towards Safe Reinforcement Learning via Constraining Conditional Value-at-Risk

IJCAI 2022poster

Though deep reinforcement learning (DRL) has obtained substantial success, it may encounter catastrophic failures due to the intrinsic uncertainty of both transition and observation. Most of the existing methods for safe reinforcement learning can only handle transition disturbance or observation di…

2022

Two Coupled Rejection Metrics Can Tell Adversarial Examples Apart

CVPR 2022poster

Correctly classifying adversarial examples is an essential but challenging requirement for safely deploying machine learning models. As reported in RobustBench, even the state-of-the-art adversarially trained models struggle to exceed 67% robust test accuracy on CIFAR-10, which is far from practical…

Cited by 24PDFcodeScholar
2022

ViewFool: Evaluating the Robustness of Visual Recognition to Adversarial Viewpoints

NeurIPS 2022accept

Recent studies have demonstrated that visual recognition models lack robustness to distribution shift. However, current work mainly considers model robustness to 2D image transformations, leaving viewpoint changes in the 3D world less explored. In general, viewpoint changes are prevalent in various…

2021

Accumulative Poisoning Attacks on Real-time Data

NeurIPS 2021poster

Collecting training data from untrusted sources exposes machine learning services to poisoning adversaries, who maliciously manipulate training data to degrade the model accuracy. When trained on offline datasets, poisoning adversaries have to inject the poisoned data in advance before training, and…

2021

Beta Distribution Guided Aspect-aware Graph for Aspect Category Sentiment Analysis with Affective Knowledge

EMNLP 2021main

In this paper, we investigate the Aspect Category Sentiment Analysis (ACSA) task from a novel perspective by exploring a Beta Distribution guided aspect-aware graph construction based on external knowledge. That is, we are no longer entangled about how to laboriously search the sentiment clues for c…

2021

Black-Box Detection of Backdoor Attacks With Limited Information and Data

ICCV 2021poster

Although deep neural networks (DNNs) have made rapid progress in recent years, they are vulnerable in adversarial environments. A malicious backdoor could be embedded in a model by poisoning the training dataset, whose intention is to make the infected model give wrong predictions during inference w…

Cited by 142PDFScholar
2021

Combining Tree Search and Action Prediction for State-of-the-Art Performance in DouDiZhu

IJCAI 2021poster

AlphaZero has achieved superhuman performance on various perfect-information games, such as chess, shogi and Go. However, directly applying AlphaZero to imperfect-information games (IIG) is infeasible, due to the fact that traditional MCTS methods cannot handle missing information of other players.…

2021

Learning Task-Distribution Reward Shaping with Meta-Learning

AAAI 2021technical

Reward shaping is one of the most effective methods to tackle the crucial yet challenging problem of credit assignment and accelerate Reinforcement Learning. However, designing shaping functions usually requires rich expert knowledge and hand-engineering, and the difficulties are further exacerbated…

Cited by 21SourcePDFScholar
2021

LiBRe: A Practical Bayesian Approach to Adversarial Detection

CVPR 2021poster

Despite their appealing flexibility, deep neural networks (DNNs) are vulnerable against adversarial examples. Various adversarial defense strategies have been proposed to resolve this problem, but they typically demonstrate restricted practicability owing to unsurmountable compromise on universality…

Cited by 78PDFcodeScholar
2021

QAIR: Practical Query-Efficient Black-Box Attacks for Image Retrieval

CVPR 2021poster

We study the query-based attack against image retrieval to evaluate its robustness against adversarial examples under the black-box setting, where the adversary only has query access to the top-k ranked unlabeled images from the database. Compared with query attacks in image classification, which pr…

Cited by 64PDFcodeScholar
2021

Sensor Fusion-based Anthropomorphic Control of Under-Actuated Bionic Hand in Dynamic Environment

IROS 2021poster

Under-actuated bionic hands have achieved tremendous popularity in many fields because of their advantages of lightweight, budget-friendly, satisfactory flexibility, and adaptability. Except for the bionic mechanical design, various anthropomorphic control strategies have been proposed and investiga…

Cited by 15SourceScholar
2021

Towards Face Encryption by Generating Adversarial Identity Masks

ICCV 2021poster

As billions of personal data being shared through social media and network, the data privacy and security have drawn an increasing attention. Several attempts have been made to alleviate the leakage of identity information from face photos, with the aid of, e.g., image obfuscation techniques. Howeve…

Cited by 120PDFcodeScholar
2021

Unsupervised Part Segmentation Through Disentangling Appearance and Shape

CVPR 2021poster

We study the problem of unsupervised discovery and segmentation of object parts, which, as an intermediate local representation, are capable of finding intrinsic object structure and providing more explainable recognition results. Recent unsupervised methods have greatly relaxed the dependency on an…

Cited by 45PDFScholar
2020

Adversarial Distributional Training for Robust Deep Learning

NeurIPS 2020poster

Adversarial training (AT) is among the most effective techniques to improve model robustness by augmenting training data with adversarial examples. However, most existing AT methods adopt a specific attack to craft adversarial examples, leading to the unreliable robustness against other unseen attac…

2020

Benchmarking Adversarial Robustness on Image Classification

CVPR 2020oral

Deep neural networks are vulnerable to adversarial examples, which becomes one of the most important research problems in the development of deep learning. While a lot of efforts have been made in recent years, it is of great significance to perform correct and complete evaluations of the adversaria…

Cited by 354PDFcodeScholar
2020

Bi-level Score Matching for Learning Energy-based Latent Variable Models

NeurIPS 2020poster

Score matching (SM) provides a compelling approach to learn energy-based models (EBMs) by avoiding the calculation of partition function. However, it remains largely open to learn energy-based latent variable models (EBLVMs), except some special cases. This paper presents a bi-level score matching (…

2020

Bilateral Teleoperation Control of a Redundant Manipulator with an RCM Kinematic Constraint

ICRA 2020poster

In this paper, a bilateral teleoperation control of a serial robot manipulator, which guarantees a Remote Center of Motion (RCM) constraint in its kinematic level, is developed. A two-layered approach based on the energy tank model is proposed to achieve haptic feedback on the end effector with a pe…

Cited by 35SourceScholar
2020

Boosting Adversarial Training with Hypersphere Embedding

NeurIPS 2020poster

Adversarial training (AT) is one of the most effective defenses against adversarial attacks for deep learning models. In this work, we advocate incorporating the hypersphere embedding (HE) mechanism into the AT procedure by regularizing the features onto compact manifolds, which constitutes a lightw…

2020

Deep Neural Network Approach in Robot Tool Dynamics Identification for Bilateral Teleoperation

RA-L 2020

For bilateral teleoperation, the haptic feedback demands the availability of accurate force information transmitted from the remote site. Nevertheless, due to the limitation of the size, the force sensor is usually attached outside of the patient's abdominal cavity for the surgical operation. Hence,

Cited by 138SourceScholar
2020

Defense Against Adversarial Attacks via Controlling Gradient Leaking on Embedded Manifolds

ECCV 2020poster

Deep neural networks are vulnerable to adversarial attacks. Though various attempts have been made, it is still largely open to fully understand the existence of adversarial samples and thereby develop effective defense strategies. In this paper, we present a new perspective, namely gradient leaking…

Cited by 26SourcePDFScholar
2020

Hierarchical optimization Control of Redundant Manipulator for Robot-assisted Minimally Invasive Surgery

IROS 2020poster

For the time varying optimization problem, the tracking error cannot converge to zero at the finite time because of the optimal solution changing over time. This paper proposes a novel varying parameter recurrent neural network (VPRNN) based hierarchical optimization of a 7-DoF surgical manipulator…

Cited by 9SourceScholar
2020

Improving Motion Planning for Surgical Robot with Active Constraints

IROS 2020poster

In this paper, an improved motion planning scheme is proposed for surgical robot control with multiple active constraints, including joint constraints, joint velocity constraints and remote center of motion constraints. It introduces an improved recurrent neural network (RNN) to optimize the online…

Cited by 6SourceScholar
2020

Internet of Things (IoT)-based Collaborative Control of a Redundant Manipulator for Teleoperated Minimally Invasive Surgeries

ICRA 2020poster

In this paper, an Internet of Things-based human-robot collaborative control scheme is developed in Robot-assisted Minimally Invasive Surgery scenario. A hierarchical operational space formulation is designed to exploit the redundancies of the 7-DoFs redundant manipulator to handle multiple operatio…

Cited by 67SourceScholar
2020

Reinforcement Learning Based Manipulation Skill Transferring for Robot-assisted Minimally Invasive Surgery

ICRA 2020poster

The complexity of surgical operation can be released significantly if surgical robots can learn the manipulation skills by imitation from complex tasks demonstrations such as puncture, suturing, and knotting, etc.. This paper proposes a reinforcement learning algorithm based manipulation skill trans…

Cited by 28SourceScholar
2020

Training Interpretable Convolutional Neural Networks by Differentiating Class-specific Filters

ECCV 2020poster

Convolutional neural networks (CNNs) have been successfully used in a range of tasks. However, CNNs are often viewed as ""black-box"" and lack of interpretability. One main reason is due to the filter-class entanglement -- an intricate many-to-many correspondence between filters and classes. Most ex…

2019

Efficient Decision-Based Black-Box Adversarial Attacks on Face Recognition

CVPR 2019poster

Face recognition has obtained remarkable progress in recent years due to the great improvement of deep convolutional neural networks (CNNs). However, deep CNNs are vulnerable to adversarial examples, which can cause fateful consequences in real-world face recognition applications with security-sensi…

Cited by 516PDFScholar
2019

Evading Defenses to Transferable Adversarial Examples by Translation-Invariant Attacks

CVPR 2019oral

Deep neural networks are vulnerable to adversarial examples, which can mislead classifiers by adding imperceptible perturbations. An intriguing property of adversarial examples is their good transferability, making black-box attacks feasible in real-world applications. Due to the threat of adversari…

Cited by 1086PDFcodeScholar
2019

Improved Human-Robot Collaborative Control of Redundant Robot for Teleoperated Minimally Invasive Surgery

RA-L 2019

An improved human-robot collaborative control scheme is proposed in a teleoperated minimally invasive surgery scenario, based on a hierarchical operational space formulation of a seven-degree-of-freedom redundant robot. Redundancy is exploited to guarantee a remote center of motion (RCM) constraint

Cited by 194SourceScholar
2019

Improving Black-box Adversarial Attacks with a Transfer-based Prior

NeurIPS 2019poster

We consider the black-box adversarial setting, where the adversary has to generate adversarial perturbations without access to the target models to compute gradients. Previous methods tried to approximate the gradient either by using a transfer gradient of a surrogate white-box model, or based on th…

2019

Manipulability Optimization Control of a Serial Redundant Robot for Robot-assisted Minimally Invasive Surgery

ICRA 2019poster

This paper proposes a manipulability optimization control of a 7-DoF robot manipulator for Robot-Assisted Minimally Invasive Surgery (RAMIS), which at the same time guarantees a Remote Center of Motion (RCM). The first degree of redundancy of the manipulator is used to achieve an RCM constraint, the…

Cited by 62SourceScholar
2019

Pixel-Adaptive Convolutional Neural Networks

CVPR 2019poster

Convolutions are the fundamental building blocks of CNNs. The fact that their weights are spatially shared is one of the main reasons for their widespread use, but it is also a major limitation, as it makes convolutions content-agnostic. We propose a pixel-adaptive convolution (PAC) operation, a sim…

Cited by 383PDFcodeScholar
2018

Boosting Adversarial Attacks With Momentum

CVPR 2018poster

Deep neural networks are vulnerable to adversarial examples, which poses security concerns on these algorithms due to the potentially severe consequences. Adversarial attacks serve as an important surrogate to evaluate the robustness of deep learning models before they are deployed. However, most of…

Cited by 3515SourcePDFScholar
2018

Interpret Neural Networks by Identifying Critical Data Routing Paths

CVPR 2018poster

Interpretability of a deep neural network aims to explain the rationale behind its decisions and enable the users to understand the intelligent agents, which has become an important issue due to its importance in practical applications. To address this issue, we develop a Distillation Guided Routing…

Cited by 122SourcePDFScholar
2018

Recognizing Minimal Facial Sketch by Generating Photorealistic Faces With the Guidance of Descriptive Attributes

ICASSP 2018accepted

Cross-modal sketch-photo recognition is of vital importance in law enforcement and public security. Most existing methods are dedicated to bridging the gap between the low-level visual features of sketches and photo images, which is limited due to intrinsic differences in pixel values. In this paper…

Cited by 0SourceScholar
2018

SPLATNet: Sparse Lattice Networks for Point Cloud Processing

CVPR 2018poster

We present a network architecture for processing point clouds that directly operates on a collection of points represented as a sparse set of samples in a high-dimensional lattice. Naively applying convolutions on this lattice scales poorly, both in terms of memory and computational cost, as the siz…

2018

Safety-Enhanced Human-Robot Interaction Control of Redundant Robot for Teleoperated Minimally Invasive Surgery

ICRA 2018poster

In this paper, a teleoperation control of a 7-DoF robot manipulator for Minimally Invasive Surgery (MIS), which guarantees a safety-enhanced compliant behavior in the null space, is described. The redundancy of the manipulator is exploited to provide a flexible workspace for nurses or other staff (a…

Cited by 62SourceScholar
2018

Textbook Question Answering Under Instructor Guidance With Memory Networks

CVPR 2018poster

Textbook Question Answering (TQA) is a task to choose the most proper answers by reading a multi-modal context of abundant essays and images. TQA serves as a favorable test bed for visual and textual reasoning. However, most of the current methods are incapable of reasoning over the long contexts an…

2017

End-To-End Face Detection and Cast Grouping in Movies Using Erdos-Renyi Clustering

ICCV 2017spotlight

We present an end-to-end system for detecting and clustering faces by identity in full-length movies. Unlike works that start with a predefined set of detected faces, we consider the end-to-end problem of detection and clustering together. We make three separate contributions. First, we combine a st…

Cited by 50PDFScholar
2016

Joint instance and feature importance re-weighting for person reidentification

ICASSP 2016accepted

Person reidentification refers to the task of recognizing the same person under different non-overlapping camera views. Presently, person reidentification based on metric learning is proved to be effective among various techniques, which exploits the labeled data to learn a subspace that maximizes t…

Cited by 0SourceScholar
2015

Active Sample Selection and Correction Propagation on a Gradually-Augmented Graph

CVPR 2015poster

When data have a complex manifold structure or the characteristics of data evolve over time, it is unrealistic to expect a graph-based semi-supervised learning method to achieve flawless classification given a small number of initial annotations. To address this issue with minimal human intervention…

Cited by 15SourcePDFScholar
2015

Improvements on transducing syllable lattice to word lattice for keyword search

ICASSP 2015accepted

This paper investigates a weighted finite state transducer (WFST) based syllable decoding and transduction method for keyword search (KWS), and compares it with sub-word search and phone confusion methods in detail. Acoustic context dependent phone models are trained from word forced alignments and…

Cited by 0SourceScholar
2015

Multi-View Convolutional Neural Networks for 3D Shape Recognition

ICCV 2015poster

A longstanding question in computer vision concerns the representation of 3D shapes for recognition: should 3D shapes be represented with descriptors operating on their native 3D formats, such as voxel grid or polygon mesh, or can they be effectively represented with view-based descriptors? We addre…

Cited by 4470PDFScholar