← Search

Jian Liu

91 accepted papers

2026

Can We Build a Monolithic Model for Fake Image Detection? SICA: Semantic-Induced Constrained Adaptation for Unified-Yet-Discriminative Artifact Feature Space Reconstruction

ICML 2026poster

Fake Image Detection (FID), aiming at unified detection across four image forensic subdomains, is critical in real-world forensic scenarios. Compared with ensemble approaches, monolithic FID models are theoretically more promising, but to date, consistently yield inferior performance in practice. In…

Cited by 0SourceScholar
2026

CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human

ICRA 2026poster

In this work, we present CollabVLA, a self-reflective vision-language-action framework that transforms a standard visuomotor policy into a collaborative assistant. CollabVLA tackles key limitations of prior VLAs, including domain overfitting, non-interpretable reasoning, and the high latency of auxi…

2026

Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation

ICLR 2026poster

Generating realistic 3D Human-Human Interaction (HHI) requires coherent modeling of the physical plausibility of the agents and their interaction semantics. Existing methods compress all motion information into a single latent representation, limiting their ability to capture fine-grained actions an…

Cited by 0SourcecodeScholar
2026

GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation

ICLR 2026poster

Recent attempts to transfer features from 2D Vision–Language Models (VLMs) to 3D semantic segmentation expose a persistent trade-off. Directly projecting 2D features into 3D yields noisy and fragmented predictions, whereas enforcing geometric coherence necessitates costly training pipelines and larg…

Cited by 0SourcecodeScholar
2026

MA-RWG: A Multi-Agent Framework for Thematically Structuring and Generation of Related Work

IJCAI 2026

AI-driven survey generation has advanced rapidly, yet related work generation (RWG) remains relatively underexplored. Unlike surveys that provide broad literature overviews, RWG synthesizes prior studies for a single focal paper, requiring contextual fit, cross-paper comparison, and accurate attribu

Cited by 0Scholar
2026

Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models

ICLR 2026poster

Large Audio Language Models (LALMs) represent an important frontier in multimodal AI, addressing diverse audio tasks. Recently, post-training of LALMs has received increasing attention due to significant performance improvements over foundation models. While single-stage post-training such as reinfo…

Cited by 0SourcecodeScholar
2026

Mesh-Pro: Asynchronous Advantage-guided Ranking Preference Optimization for Artist-style Quadrilateral Mesh Generation

CVPR 2026

Reinforcement learning (RL) has demonstrated remarkable success in text and image generation, yet its potential in 3D generation remains largely unexplored. Existing attempts typically rely on offline direct preference optimization (DPO) method, which suffers from low training efficiency and limited

Cited by 0SourceScholar
2026

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

CVPR 2026

Omnimodal large language models (OmniLLMs) have attracted increasing research attention of late towards unified audio-video understanding. However, the high computational cost of processing longer joint audio-video token sequences has become a key bottleneck. Existing token compression methods have

Cited by 0SourcecodeScholar
2026

QuadGPT: Native Quadrilateral Mesh Generation with Autoregressive Models

ICLR 2026poster

The generation of quadrilateral-dominant meshes is a cornerstone of professional 3D content creation. However, existing generative models generate quad meshes by first generating triangle meshes and then merging triangles into quadrilaterals with some specific rules, which typically produces quad m…

Cited by 0SourceScholar
2026

ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization

IJCAI 2026

Offline-to-online reinforcement learning harnesses the stability of offline pretraining and the flexibility of online fine-tuning. A key challenge lies in the non-stationary distribution shift between offline datasets and the evolving online policy. Common approaches often rely on static mixing rati

Cited by 0Scholar
2026

Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance

ICLR 2026poster

Fine-tuning safety-aligned large language models (LLMs) can substantially compromise their safety. Previous approaches require many safety samples or calibration sets, which not only incur significant computational overhead during realignment but also lead to noticeable degradation in model utility.…

Cited by 0SourceScholar
2026

Suppress and Diversify: Refining Robust Pathways for Corruption Robustness

ICML 2026spotlight

Model robustness against natural image corruptions is essential for safety-critical applications. While existing methods primarily focus on implicit representation learning, we provide the first systematic exploration of computational pathways to explicitly characterize internal robustness. We ident…

Cited by 0SourceScholar
2026

TextShield-R1: Reinforced Reasoning for Tampered Text Detection

AAAI 2026technical

The growing prevalence of tampered images poses serious security threats, highlighting the urgent need for reliable detection methods. Multimodal large language models (MLLMs) demonstrate strong potential in analyzing tampered images and generating interpretations. However, they still struggle with

Cited by 0SourcePDFScholar
2026

Towards Trustworthy and Identifiable Virtual Face Generation

ICML 2026poster

Identifiable virtual face (IVF) generation aims to transform a user's original face into a virtual face for high utility privacy protection. The IVF is visually and statistically different from the original face, which can still be used for recognizing the user's identity. Despite the advantage, the…

Cited by 0SourceScholar
2026

Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence

ICML 2026poster

The impressive performance of generalist large language models (LLMs) such as GPT-4 and Claude in healthcare raises a critical question: will domain-specific medical specialist models become obsolete? We argue that the future of medical artificial intelligence (AI) lies not in building monolithic me…

Cited by 0SourceScholar
2025

Adaptive Merchant-Centric Risk Control via Unbiased Decision-Making and Dynamic Optimization in E-Commerce

AAAI 2025technical

In the domain of merchant-oriented risk control decisions within e-commerce, balancing the effectiveness of risk management with merchant satisfaction remains a critical challenge. Strict risk control strategies, while effectively mitigating risks, often lead to increased merchant dissatisfaction. C…

Cited by 0SourcePDFScholar
2025

Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference Optimization

NeurIPS 2025poster

We introduce Auto-Connect, a novel approach for automatic rigging that explicitly preserves skeletal connectivity through a connectivity-preserving tokenization scheme. Unlike previous methods that predict bone positions represented as two joints or first predict points before determining connectivi…

Cited by 0SourceScholar
2025

Backdooring Self-Supervised Contrastive Learning by Noisy Alignment

ICCV 2025poster

Self-supervised contrastive learning (CL) effectively learns transferable representations from unlabeled data containing images or image-text pairs but suffers vulnerability to data poisoning backdoor attacks (DPCLs). An adversary can inject poisoned images into pretraining datasets, causing comprom…

2025

EmoRLTalk: Speech-Driven Emotional Facial Animation With Offline Reinforcement Learning

IROS 2025

In recent years, significant breakthroughs have been made in audio-guided 3D facial animation. However, existing methods mainly focus on lip shape and audio consistency and still face key challenges to achieve alignment between facial emotions and speech emotions. To overcome this limitation, we int

Cited by 0SourceScholar
2025

FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies

NeurIPS 2025poster

The increasing realism of synthetic images generated by advanced models such as VAEs, GANs, and LDMs poses significant challenges for synthetic image detection. To address this issue, we explore two artifact types introduced during the generation process: (1) latent distribution deviations and (2) d…

Cited by 0SourcecodeScholar
2025

FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization

EMNLP 2025

The rapid advancement of large language models (LLMs) has exacerbated the memory bottleneck due to the widening gap between model parameter scaling and hardware capabilities. While post-training quantization techniques effectively reduce memory overhead, existing methods predominantly rely on static

Cited by 0SourcePDFScholar
2025

ForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and Localization

NeurIPS 2025poster

The field of Fake Image Detection and Localization (FIDL) is highly fragmented, encompassing four domains: deepfake detection (Deepfake), image manipulation detection and localization (IMDL), artificial intelligence-generated image detection (AIGC), and document image manipulation localization (Doc)…

Cited by 0SourcecodeScholar
2025

FreeMesh: Boosting Mesh Generation with Coordinates Merging

ICML 2025poster

The next-coordinate prediction paradigm has emerged as the de facto standard in current auto-regressive mesh generation methods. Despite their effectiveness, there is no efficient measurement for the various tokenizers that serialize meshes into sequences. In this paper, we introduce a new metric P…

Cited by 0SourcePDFScholar
2025

G-DexGrasp: Generalizable Dexterous Grasping Synthesis Via Part-Aware Prior Retrieval and Prior-Assisted Generation

ICCV 2025poster

Recent advances in dexterous grasping synthesis have demonstrated significant progress in producing reasonable and plausible grasps for many task purposes. But it remains challenging to generalize to unseen object categories and diverse task instructions. In this paper, we propose G-DexGrasp, a retr…

Cited by 0SourcePDFScholar
2025

KGCRR: An Effective Metric-Driven Knowledge Graph Completion Framework by Designing a Novel Upper Bound Function with Adaptive Approximation to Reciprocal Rank

AAAI 2025technical

Knowledge Graph Embedding (KGE) methods have achieved great success in predicting missing links in knowledge graphs, a task also known as Knowledge Graph Completion (KGC). Under this task, the Reciprocal Rank (RR) of ground-truth items serve as a key indicator for evaluating the method’s performance…

2025

LanCOPE: Language-Guided Category-Level Object Pose Estimation From a Single RGB Image

RA-L 2025

Monocular RGB-based category-level object pose estimation is more practical and cost-effective for robotics. However, existing methods do not fully exploit the rich semantic and contextual information in multimodal data (e.g. language) that provides additional object attributes to guide the model in

Cited by 1SourceScholar
2025

MambaML: Exploring State Space Models for Multi-Label Image Classification

ICCV 2025poster

Mamba, a selective state-space model, has recently seen widespread application across various visual tasks due to its exceptional ability to capture long-range dependencies. While promising results have been demonstrated in image classification, its potential in multi-label image classification rema…

Cited by 0SourcePDFScholar
2025

Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning

NeurIPS 2025spotlight

Existing pretrained models for 3D mesh generation often suffer from data biases and produce low-quality results, while global reinforcement learning (RL) methods rely on object-level rewards that struggle to capture local structure details. To address these challenges, we present $\textbf{Mesh-RFT}$…

Cited by 0SourceScholar
2025

MoFRR: Mixture of Diffusion Models for Face Retouching Restoration

ICCV 2025poster

The widespread use of face retouching on social media platforms raises concerns about the authenticity of face images. While existing methods focus on detecting face retouching, how to accurately recover the original faces from the retouched ones has yet to be answered. This paper introduces Face Re…

Cited by 0SourcePDFScholar
2025

MonoDiff9D: Monocular Category-Level 9D Object Pose Estimation via Diffusion Model

ICRA 2025

Object pose estimation is a core means for robots to understand and interact with their environment. For this task, monocular category-level methods are attractive as they require only a single RGB camera. However, current methods rely on shape priors or CAD models of the intra-class known objects.

Cited by 5SourcecodeScholar
2025

Multi-Stage LLM Fine-Tuning with a Continual Learning Setting

NAACL 2025findings

In recent years, large language models (LLMs) have made significant progress in knowledge-intensive applications. However, when adapting them to specific domains, we may encounter a multi-stage continuous learning scenario, especially in cases where domain knowledge evolves rapidly.This issue severe…

Cited by 1SourcePDFScholar
2025

OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM

ICCV 2025poster

Despite the remarkable progress of multimodal large language models (MLLMs), they continue to face challenges in achieving competitive performance on ordinal regression (OR; a.k.a. ordinal classification). To address this issue, this paper presents OrderChain, a novel and general prompting paradigm…

2025

PSBench: a large-scale benchmark for estimating the accuracy of protein complex structural models

NeurIPS 2025poster

Predicting protein complex structures is essential for protein function analysis, protein design, and drug discovery. While AI methods like AlphaFold can predict accurate structural models for many protein complexes, reliably estimating the quality of these predicted models (estimation of model accu…

Cited by 0SourcecodeScholar
2025

PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object Detection

AAAI 2025technical

Single-Domain Generalized Object Detection (S-DGOD) aims to train on a single source domain for robust performance across a variety of unseen target domains by taking advantage of an object detector. Existing S-DGOD approaches often rely on data augmentation strategies, including a composition of vi…

2025

RGB-Based Category-Level Object Pose Estimation via Depth Recovery and Adaptive Refinement

RA-L 2025

Category-level pose estimation methods have received widespread attention as they can be generalized to intra-class unseen objects. Although RGB-D-based category-level methods have made significant progress, reliance on depth image limits practical application. RGB-based methods offer a more practic

Cited by 3SourceScholar
2025

Scalable Autoregressive Monocular Depth Estimation

CVPR 2025poster

This paper proposes a new autoregressive model as an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with an autoregressive prediction paradigm, based on two core designs. First, our depth autoregressive model (DAR) treats the…

2025

Scaling Mesh Generation via Compressive Tokenization

CVPR 2025poster

We propose a compressive yet effective mesh tokenization, Blocked and Patchified Tokenization (BPT), facilitating the generation of meshes exceeding 8k faces. BPT compresses mesh sequences by employing block-wise indexing and patch aggregation, reducing their length by approximately 75% compared to…

2025

TAET: Two-Stage Adversarial Equalization Training on Long-Tailed Distributions

CVPR 2025poster

Adversarial robustness remains a significant challenge in deploying deep neural networks for real-world applications. While adversarial training is widely acknowledged as a promising defense strategy, most existing studies primarily focus on balanced datasets, neglecting the fact that real-world dat…

2025

WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models

ICASSP 2025accepted

Although Large Language Models (LLMs) excel in NLP tasks, they still need external tools to extend their ability. Current research on tool learning with LLMs often assumes mandatory tool use, which does not always align with real-world situations, where the necessity for tools is uncertain, and inco…

Cited by 0SourceScholar
2024

AFBench: A Large-scale Benchmark for Airfoil Design

NeurIPS 2024poster

Data-driven generative models have emerged as promising approaches towards achieving efficient mechanical inverse design. However, due to prohibitively high cost in time and money, there is still lack of open-source and large-scale benchmarks in this field. It is mainly the case for airfoil inverse…

2024

Adaptive Shape Servoing of Elastic Rods Using Parameterized Regression Features and Auto-Tuning Motion Controls

RA-L 2024

The robotic manipulation of deformable linear objects has shown great potential in a wide range of real-world applications. However, it presents many challenges due to the objects' non-linear properties and high-dimensional geometric configuration. In this letter, we propose an efficient shape servo

Cited by 33SourceScholar
2024

Divide-and-Aggregate Learning for Evaluating Performance on Unlabeled Data

AAAI 2024technical

Artificial Intelligence (AI) models have become an integral part of modern society, significantly improving human lives. However, ensuring the reliability and safety of these models is of paramount importance. One critical aspect is the continuous monitoring and verification of model performance to…

2024

EAB-FL: Exacerbating Algorithmic Bias through Model Poisoning Attacks in Federated Learning

IJCAI 2024poster

Federated Learning (FL) is a technique that allows multiple parties to train a shared model collaboratively without disclosing their private data. It has become increasingly popular due to its distinct privacy advantages. However, FL models can suffer from biases against certain demographic groups (…

2024

Osprey: Pixel Understanding with Visual Instruction Tuning

CVPR 2024poster

Multimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning. However current MLLMs primarily focus on image-level or box-level understanding falling short in achieving fine-grained vision-language alignment…

2024

Position: TrustLLM: Trustworthiness in Large Language Models

ICML 2024poster

Large language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLM…

Cited by 95SourcePDFScholar
2024

Strengthening Layer Interaction via Dynamic Layer Attention

IJCAI 2024poster

In recent years, employing layer attention to enhance interaction among hierarchical layers has proven to be a significant advancement in building network structures. In this paper, we delve into the distinction between layer attention and the general attention mechanism, noting that existing layer…

2024

SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment

ICML 2024poster

Multimodal alignment between language and vision is the fundamental topic in current vision-language model research. Contrastive Captioners (CoCa), as a representative method, integrates Contrastive Language-Image Pretraining (CLIP) and Image Caption (IC) into a unified framework, resulting in impre…

Cited by 4SourcePDFScholar
2024

Towards Multi-Relational Multi-Hop Reasoning over Dense Temporal Knowledge Graphs

ACL 2024findings

Temporal knowledge graph reasoning has emerged as a crucial task for answering time-dependent questions within a knowledge graph (KG).Despite tremendous progress, the present research is impeded by the sparsity of a temporal KG and an over-reliance on simple single-relational reasoning patterns. To…

2024

Video Event Extraction with Multi-View Interaction Knowledge Distillation

AAAI 2024technical

Video event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, whi…

Cited by 2SourcePDFScholar
2023

A Multi-modal Debiasing Model with Dynamical Constraint for Robust Visual Question Answering

ACL 2023findings

Recent studies have pointed out that many well-developed Visual Question Answering (VQA) systems suffer from bias problem. Despite the remarkable performance gained on In-Distribution (ID) datasets, the VQA model might merely capture the superficial correlation from question to answer rather than sh…

Cited by 10SourcePDFScholar
2023

A Quality-based Syntactic Template Retriever for Syntactically-Controlled Paraphrase Generation

EMNLP 2023long main

Existing syntactically-controlled paraphrase generation (SPG) models perform promisingly with human-annotated or well-chosen syntactic templates. However, the difficulty of obtaining such templates actually hinders the practical application of SPG models. For one thing, the prohibitive cost makes it…

Cited by 0SourcecodeScholar
2023

Addressing NER Annotation Noises with Uncertainty-Guided Tree-Structured CRFs

EMNLP 2023long main

Real-world named entity recognition (NER) datasets are notorious for their noisy nature, attributed to annotation errors, inconsistencies, and subjective interpretations. Such noises present a substantial challenge for traditional supervised learning methods. In this paper, we present a new and unif…

Cited by 0SourceScholar
2023

AffordPose: A Large-Scale Dataset of Hand-Object Interactions with Affordance-Driven Hand Pose

ICCV 2023poster

How human interact with objects depends on the functional roles of the target objects, which introduces the problem of affordance-aware hand-object interaction. It requires a large number of human demonstrations for the learning and understanding of plausible and appropriate hand-object interactions…

Cited by 47PDFcodeScholar
2023

Document-Level Event Argument Extraction With a Chain Reasoning Paradigm

ACL 2023long

Document-level event argument extraction aims to identify event arguments beyond sentence level, where a significant challenge is to model long-range dependencies. Focusing on this challenge, we present a new chain reasoning paradigm for the task, which can generate decomposable first-order logic ru…

Cited by 17SourcePDFScholar
2023

Exploring the Mutual Influence Between Self-Supervised Single-Frame and Multi-Frame Depth Estimation

RA-L 2023

Although both self-supervised single-frame and multi-frame depth estimation methods only require unlabeled monocular videos for training, the information they leverage varies because single-frame methods mainly rely on appearance-based features while multi-frame methods focus on geometric cues. Cons

Cited by 8SourcecodeScholar
2023

Generating Transferable 3D Adversarial Point Cloud via Random Perturbation Factorization

AAAI 2023technical

Recent studies have demonstrated that existing deep neural networks (DNNs) on 3D point clouds are vulnerable to adversarial examples, especially under the white-box settings where the adversaries have access to model parameters. However, adversarial 3D point clouds generated by existing white-box me…

2023

Label-efficient Segmentation via Affinity Propagation

NeurIPS 2023poster

Weakly-supervised segmentation with label-efficient sparse annotations has attracted increasing research attention to reduce the cost of laborious pixel-wise labeling process, while the pairwise affinity modeling techniques play an essential role in this task. Most of the existing approaches focus o…

2023

Novel Slot Detection With an Incremental Setting

EMNLP 2023long findings

Current dialogue systems face diverse user requests and rapid change domains, making quickly adapt to scenarios with previous unseen slot types become a major challenge. Recently, researchers have introduced novel slot detection (NSD) to discover potential new types. However, dialogue system with NS…

Cited by 0SourceScholar
2023

Point2Mask: Point-supervised Panoptic Segmentation via Optimal Transport

ICCV 2023poster

Weakly-supervised image segmentation has recently attracted increasing research attentions, aiming to avoid the expensive pixel-wise labeling. In this paper, we present an effective method, namely Point2Mask, to achieve high-quality panoptic prediction using only a single random point annotation per…

Cited by 27PDFcodeScholar
2023

Scene-Aware Label Graph Learning for Multi-Label Image Classification

ICCV 2023poster

Multi-label image classification refers to assigning a set of labels for an image. One of the main challenges of this task is how to effectively capture the correlation among labels. Existing studies on this issue mostly rely on the statistical label co-occurrence or semantic similarity of labels. H…

Cited by 31PDFScholar
2023

Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation

ACL 2023long

Knowledge distillation (KD) is a promising technique for model compression in neural machine translation. However, where the knowledge hides in KD is still not clear, which may hinder the development of KD. In this work, we first unravel this mystery from an empirical perspective and show that the k…

2022

Auditing Privacy Defenses in Federated Learning via Generative Gradient Leakage

CVPR 2022poster

Federated Learning (FL) framework brings privacy benefits to distributed learning systems by allowing multiple clients to participate in a learning task under the coordination of a central server without exchanging their private data. However, recent studies have revealed that private information ca…

Cited by 146PDFcodeScholar
2022

Conditional Bilingual Mutual Information Based Adaptive Training for Neural Machine Translation

ACL 2022long

Token-level adaptive training approaches can alleviate the token imbalance problem and thus improve neural machine translation, through re-weighting the losses of different target tokens based on specific statistical metrics (e.g., token frequency or mutual information). Given that standard translat…

2022

Invisible and Efficient Backdoor Attacks for Compressed Deep Neural Networks

ICASSP 2022accepted

Compressed deep neural network (DNN) models have been widely deployed in many resource-constrained platforms and devices. However, the security issue of the compressed models, especially their vulnerability against backdoor attacks, is not well explored yet. In this paper, we study the feasibility o…

Cited by 0SourceScholar
2022

Iterative Constrained Back-Translation for Unsupervised Domain Adaptation of Machine Translation

COLING 2022main

Back-translation has been proven to be effective in unsupervised domain adaptation of neural machine translation (NMT). However, the existing back-translation methods mainly improve domain adaptability by generating in-domain pseudo-parallel data that contains sentence-structural knowledge, paying l…

2022

Label Relation Graphs Enhanced Hierarchical Residual Network for Hierarchical Multi-Granularity Classification

CVPR 2022poster

Hierarchical multi-granularity classification (HMC) assigns hierarchical multi-granularity labels to each object and focuses on encoding the label hierarchy, e.g., ["Albatross", "Laysan Albatross"] from coarse-to-fine levels. However, the definition of what is fine-grained is subjective, and the ima…

Cited by 62PDFcodeScholar
2022

RIBAC: Towards Robust and Imperceptible Backdoor Attack against Compact DNN

ECCV 2022poster

"Recently backdoor attack has become an emerging threat to the security of deep neural network (DNN) models. To date, most of the existing studies focus on backdoor attack against the uncompressed model; while the vulnerability of compressed DNNs, which are widely used in the practical applications,…

2022

Saliency as Evidence: Event Detection with Trigger Saliency Attribution

ACL 2022long

Event detection (ED) is a critical subtask of event extraction that seeks to identify event triggers of certain types in texts. Despite significant advances in ED, existing methods typically follow a “one model fits all types” approach, which sees no differences between event types and often results…

2022

Visual Attention-Based Self-Supervised Absolute Depth Estimation Using Geometric Priors in Autonomous Driving

RA-L 2022

Although existing monocular depth estimation methods have made great progress, predicting an accurate absolute depth map from a single image is still challenging due to the limited modeling capacity of networks and the scale ambiguity issue. In this paper, we introduce a fully Visual Attention-based

Cited by 27SourceScholar
2022

Watermark Vaccine: Adversarial Attacks to Prevent Watermark Removal

ECCV 2022poster

"As a common security tool, visible watermarking has been widely applied to protect copyrights of digital images. However, recent works have shown that visible watermarks can be removed by DNNs without damaging their host images. Such watermark-removal techniques pose a great threat to the ownership…

2021

Cross-Domain Slot Filling as Machine Reading Comprehension

IJCAI 2021poster

With task-oriented dialogue systems being widely applied in everyday life, slot filling, the essential component of task-oriented dialogue systems, is required to be quickly adapted to new domains that contain domain-specific slots with few or no training data. Previous methods for slot filling usua…

2021

Discourse-Level Event Temporal Ordering with Uncertainty-Guided Graph Completion

IJCAI 2021poster

Learning to order events at discourse-level is a crucial text understanding task. Despite many efforts for this task, the current state-of-the-art methods rely heavily on manually designed features, which are costly to produce and are often specific to tasks/domains/datasets. In this paper, we pro…

2021

Enabling Fast and Universal Audio Adversarial Attack Using Generative Model

AAAI 2021technical

Recently, the vulnerability of deep neural network (DNN)-based audio systems to adversarial attacks has obtained increasing attention. However, the existing audio adversarial attacks allow the adversary to possess the entire user's audio input as well as granting sufficient time budget to generate t…

Cited by 79SourcePDFScholar
2021

Improving Stylized Neural Machine Translation with Iterative Dual Knowledge Transfer

IJCAI 2021poster

Stylized neural machine translation (NMT) aims to translate sentences of one style into sentences of another style, which is essential for the application of machine translation in a real-world scenario. However, a major challenge in this task is the scarcity of high-quality parallel data which is s…

2021

Machine Reading Comprehension as Data Augmentation: A Case Study on Implicit Event Argument Extraction

EMNLP 2021main

Implicit event argument extraction (EAE) is a crucial document-level information extraction task that aims to identify event arguments beyond the sentence level. Despite many efforts for this task, the lack of enough training data has long impeded the study. In this paper, we take a new perspective…

Cited by 71SourcePDFScholar
2021

Natural Language Inference in Context – Investigating Contextual Reasoning over Long Texts

AAAI 2021technical

Natural language inference (NLI) is a fundamental NLP task, investigating the entailment relationship between two texts.
 Popular NLI datasets present the task at sentence-level. While adequate for testing semantic representations, they fall short for testing contextual reasoning over long texts, wh…

2021

Saliency-based Multi-View Mixed Language Training for Zero-shot Cross-lingual Classification

EMNLP 2021finding

Recent multilingual pre-trained models, like XLM-RoBERTa (XLM-R), have been demonstrated effective in many cross-lingual tasks. However, there are still gaps between the contextualized representations of similar words in different languages. To solve this problem, we propose a novel framework named…

Cited by 8SourcePDFScholar
2021

Solving Aspect Category Sentiment Analysis as a Text Generation Task

EMNLP 2021main

Aspect category sentiment analysis has attracted increasing research attention. The dominant methods make use of pre-trained language models by learning effective aspect category-specific representations, and adding specific output layers to its pre-trained representation. We consider a more direct…

2020

Designing A Dummy Skin by Evaluating Contacts between A Human Hand and A Robot End Tip

IROS 2020poster

Many manufacturing industries have a high demand for the construction of collaborative operation systems using industrial robots. Although there is a preexisting set of safety verification data in ISO/TS 15066 for collaborative operations, there is no established testing method for safety validation…

Cited by 10SourceScholar
2020

Graph-Based Knowledge Integration for Question Answering over Dialogue

COLING 2020main

Question answering over dialogue, a specialized machine reading comprehension task, aims to comprehend a dialogue and to answer specific questions. Despite many advances, existing approaches for this task did not consider dialogue structure and background knowledge (e.g., relationships between speak…

2020

Knowledge Enhanced Event Causality Identification with Mention Masking Generalizations

IJCAI 2020poster

Identifying causal relations of events is a crucial language understanding task. Despite many efforts for this task, existing methods lack the ability to adopt background knowledge, and they typically generalize poorly to new, previously unseen data. In this paper, we present a new method for event…

2020

LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

IJCAI 2020poster

Machine reading is a fundamental task for testing the capability of natural language understand- ing, which is closely related to human cognition in many aspects. With the rising of deep learning techniques, algorithmic models rival human performances on simple QA, and thus increasingly challenging…

2020

Real-Time, Universal, and Robust Adversarial Attacks Against Speaker Recognition Systems

ICASSP 2020accepted

As the popularity of voice user interface (VUI) exploded in recent years, speaker recognition system has emerged as an important medium of identifying a speaker in many security-required applications and services. In this paper, we propose the first real-time, universal, and robust adversarial attac…

Cited by 0SourceScholar
2018

Caging Loops in Shape Embedding Space: Theory and Computation

ICRA 2018poster

We propose to synthesize feasible caging grasps for a target object through computing Caging Loops, a closed curve defined in the shape embedding space of the object. Different from the traditional methods, our approach decouples caging loops from the surface geometry of target objects through worki…

Cited by 5SourceScholar