← Search

Jin Wang

80 accepted papers

2026

Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency

AAAI 2026technical

Cross-modal Knowledge Distillation has demonstrated promising performance on paired modalities with strong semantic connections, referred to as Symmetric Cross-modal Knowledge Distillation (SCKD). However, implementing SCKD becomes exceedingly constrained in real-world scenarios due to the limited a

Cited by 0SourcePDFScholar
2026

Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement Learning

CVPR 2026

Large Vision-Language Models (LVLMs) often omit or misrepresent critical visual content in generated image captions. Minimizing such information loss will force LVLMs to focus on image details to generate precise descriptions. However, measuring information loss during modality conversion is inheren

Cited by 0SourceScholar
2026

Diffusion-Based mmWave Radar Point Cloud Enhancement Driven by Range Images

RA-L 2026

Millimeter-wave (mmWave) radar has attracted significant attention in robotics and autonomous driving due to its robustness in harsh environments. However, the radar point clouds are typically sparse and noisy, which limits its futher development. Traditional mmWave radar enhancement approaches ofte

Cited by 6SourceScholar
2026

From Denoising to De-Channeling: Integrating Physical Channel Priors into Diffusion Models for Radio Signal Understanding

ICML 2026spotlight

In recent years, wireless signal recognition (WSR), which leverages artificial intelligence (AI) to identify properties of passively received radio signals, has garnered significant attention due to its broad applications, such as spectrum management. Existing WSR methods typically learn directly fr…

Cited by 0SourceScholar
2026

HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models

AAAI 2026technical

3D understanding has drawn significant attention recently, leveraging Vision-Language Models (VLMs) to enable multi-modal reasoning between point cloud and text data. Current 3D-VLMs directly embed the 3D point clouds into 3D tokens, following large 2D-VLMs with powerful reasoning capabilities. Howe

Cited by 0SourcePDFScholar
2026

Koopman-Assisted Trajectory Synthesis: A Data Augmentation Framework for Offline Imitation Learning

ICLR 2026poster

Data augmentation plays a pivotal role in offline imitation learning (IL) by alleviating covariate shift, yet existing methods remain constrained. Single-step techniques frequently violate underlying system dynamics, whereas trajectory-level approaches are plagued by compounding errors or scalabilit…

Cited by 0SourceScholar
2026

LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language Models

AAAI 2026technical

Aligning Large Language Models (LLMs) with human preferences is critical, yet traditional fine-tuning methods are computationally expensive and inflexible. While test-time alignment offers a promising alternative, existing approaches often rely on distorted trajectory-level signals or inefficient sa

Cited by 0SourcePDFScholar
2026

Quantum Machine Learning and Grover’s Algorithm for Quantum Optimization of Robotic Manipulators

ICRA 2026poster

Optimizing high-degree-of-freedom robotic manipulators requires searching complex, high-dimensional configuration spaces, a task that is computationally challenging for classical methods. This paper introduces a quantum-native framework that integrates Quantum Machine Learning (QML) with Grover's al…

2026

SAPO: Self-Adaptive Process Optimization Makes Small Reasoners Stronger

AAAI 2026technical

Existing self-evolution methods overlook the influence of fine-grained reasoning steps, which leads to the reasoner-verifier gap. The computational inefficiency of Monte Carlo (MC) process supervision further exacerbates the difficulty in mitigating the gap. Motivated by the Error-Related Negativity

Cited by 0SourcePDFScholar
2026

Step-GRPO: Enhancing Reasoning Quality and Efficiency via Structured PRM-Based Reinforcement Learning

AAAI 2026technical

Large reasoning models (LRMs) improve performance at test time by thinking longer, but this often leads to overthinking and high computational cost. To address this, recent reinforcement learning (RL) methods adopt outcome-level rewards, such as rule- or prompt-based signals, that favor shorter corr

Cited by 0SourcePDFScholar
2026

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

ICML 2026poster

Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for text-to-image and text-to-video generation. However, we find that directly applying these techniques to image-to-video (I2V) models often fails to yield …

Cited by 0SourceScholar
2026

TS-DDAE: A novel Temporal-Spectral Denoising Diffusion AutoEncoder for Wireless Signal Recognition Model Pre-training

ICLR 2026poster

Wireless Signal Recognition (WSR) aims to identify the property of received signals using Artificial Intelligence (AI) without any prior knowledge, which has been widely used in civil and military radios. The current AI trend of pre-training and fine-tuning has shown great performance, and the exist…

Cited by 0SourcecodeScholar
2025

Batch Selection for Multi-Label Classification Guided by Uncertainty and Dynamic Label Correlations

AAAI 2025technical

The accuracy of deep neural networks is significantly influenced by the effectiveness of mini-batch construction during training. In single-label scenarios, such as binary and multi-class classification tasks, it has been demonstrated that batch selection algorithms preferring samples with higher un…

2025

CapsuleBot: A Novel Hybrid Aerial-Ground Bi-Copter Robot With Two Actuated-Wheel-Rotors

RA-L 2025

This paper presents the design, modeling, and experimental validation of CapsuleBot, a novel hybrid aerial-ground bi-copter robot designed for long-endurance and low-noise operations. CapsuleBot combines the maneuverability of a bi-copter in the air with the low power consumption and low noise of a

Cited by 10SourceScholar
2025

CoDynTrust: Robust Asynchronous Collaborative Perception via Dynamic Feature Trust Modulus

ICRA 2025

Collaborative perception, fusing information from multiple agents, can extend perception range so as to improve perception performance. However, temporal asynchrony in real-world environments, caused by communication delays, clock misalignment, or sampling configuration differences, can lead to info

Cited by 5SourcecodeScholar
2025

Data-Free Black-Box Federated Learning via Zeroth-Order Gradient Estimation

AAAI 2025technical

Federated learning (FL) enables decentralized clients to collaboratively train a global model under the orchestration of a central server without exposing their individual data. However, the iterative exchange of model parameters between the server and clients imposes heavy communication burdens, ri…

2025

Divide-Solve-Combine: An Interpretable and Accurate Prompting Framework for Zero-shot Multi-Intent Detection

AAAI 2025technical

Zero-shot multi-intent detection is capable of capturing multiple intents within a single utterance without any training data, which gains increasing attention. Building on the success of large language models (LLM), dominant approaches in the literature explore prompting techniques to enable zero-s…

2025

Dual-Path Contrastive Short Text Clustering with High-order Random Walk

ICASSP 2025accepted

In recent years, several robust contrastive text clustering methods have been proposed. While these methods have achieved significant performances, two issues remain. First, the false negative problem is still not fully resolved, and the false positive issue also arises because all in-neighborhood a…

Cited by 0SourceScholar
2025

Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network Approach

AAAI 2025technical

Event cameras have recently been introduced into image semantic segmentation, owing to their high temporal resolution and other advantageous properties. However, existing event-based semantic segmentation methods often fail to fully exploit the complementary information provided by frames and events…

2025

Event-Enhanced Blurry Video Super-Resolution

AAAI 2025technical

In this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insu…

2025

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities

NeurIPS 2025spotlight

The rapid progress of large language models (LLMs) has catalyzed the emergence of multimodal large language models (MLLMs) that unify visual understanding and image generation within a single framework. However, most existing MLLMs rely on autoregressive (AR) architectures, which impose inherent lim…

Cited by 0SourceScholar
2025

Forensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language Models

CVPR 2025poster

Recently, the rapid development of AIGC has significantly boosted the diversities of fake media spread in the Internet, posing unprecedented threats to social security, politics, law, and etc.To detect the ever-increasingly **diverse** malicious fake media in the new era of AIGC, recent studies have…

2025

Hop-level Direct Preference Optimization for Knowledge Graph Reasoning with Trees

ICASSP 2025accepted

Recent advancements in knowledge graph question answering (KGQA) have shown promise, yet existing methods often fail to align with human reasoning patterns. This study proposes HD-PORT (hop-level direct preference optimization for knowledge graph reasoning with trees), a novel approach that combines…

Cited by 0SourceScholar
2025

INSTINCT: Instance-Level Interaction Architecture for Query-Based Collaborative Perception

ICCV 2025poster

Collaborative perception systems overcome single-vehicle limitations in long-range detection and occlusion scenarios by integrating multi-agent sensory data, improving accuracy and safety. However, frequent cooperative interactions and real-time requirements impose stringent bandwidth constraints. P…

2025

INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling

ICCV 2025poster

Hallucinations in large vision-language models (LVLMs) pose significant challenges for real-world applications, as LVLMs may generate responses that appear plausible yet remain inconsistent with the associated visual content. This issue rarely occurs in human cognition. We argue that this discrepanc…

2025

ISP2HRNet: Learning to Reconstruct High Resolution Image from Irregularly Sampled Pixels via Hierarchical Gradient Learning

ICCV 2025poster

While image signals are typically defined on a regular 2D grid, there are scenarios where they are only available at irregular positions. In such cases, reconstructing a complete image on regular grid is essential. This paper introduces ISP2HRNet, an end-to-end network designed to reconstruct high r…

2025

Jury-and-Judge Chain-of-Thought for Uncovering Toxic Data in 3D Visual Grounding

NeurIPS 2025poster

3D Visual Grounding (3DVG) faces persistent challenges due to coarse scene-level observations and logically inconsistent annotations, which introduce ambiguities that compromise data quality and hinder effective model supervision. To address these challenges, we introduce Refer-Judge, a novel framew…

Cited by 0SourcecodeScholar
2025

Learning to Reason via Self-Iterative Process Feedback for Small Language Models

COLING 2025main

Small language models (SLMs) are more efficient, cost-effective, and customizable than large language models (LLMs), though they often underperform in specific areas like reasoning. Past methods for enhancing SLMs’ reasoning, such as supervised fine-tuning and distillation, often depend on costly ex…

2025

MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

ICLR 2025poster

The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their evaluation has not kept pace with their development. To fill this ga…

2025

MegActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer

AAAI 2025technical

Diffusion models have demonstrated superior performance in portrait animation. However, current approaches relied on either visual or audio modality to control character movements, failing to exploit the potential of mixed-modal control. This challenge arises from the difficulty in balancing the we…

2025

Multi-Attribute Multi-Grained Adaptation of Pre-Trained Language Models for Text Understanding from Bayesian Perspective

AAAI 2025technical

Current neural networks often employ multi-domain-learning or attribute-injecting mechanisms to incorporate non-independent and identically distributed (non-IID) information for text understanding tasks by capturing individual characteristics and the relationships among samples. However, the extent…

2025

Quantum Machine Learning and Grover's Algorithm for Quantum Optimization of Robotic Manipulators

RA-L 2025

Optimizing high-degree-of-freedom robotic manipulators requires searching complex, high-dimensional configuration spaces, a task that is computationally challenging for classical methods. This paper introduces a quantum-native framework that integrates Quantum Machine Learning (QML) with Grover's al

Cited by 0SourceScholar
2025

RALAD: Bridging the Real-to-Sim Domain Gap in Autonomous Driving with Retrieval-Augmented Learning

IROS 2025

As end-to-end autonomous driving advances toward real-world deployment, ensuring the safety of autonomous vehicles (AVs) has become a critical requirement for their commercial viability. While rule-based AVs have traditionally undergone rigorous testing in both real-world and simulated environments

Cited by 2SourcecodeScholar
2025

Reasoning with Trees: Faithful Question Answering over Knowledge Graph

COLING 2025main

Recent advancements in large language models (LLMs) have shown remarkable progress in reasoning capabilities, yet they still face challenges in complex, multi-step reasoning tasks. This study introduces Reasoning with Trees (RwT), a novel framework that synergistically integrates LLMs with knowledge…

Cited by 0SourcePDFScholar
2025

RoboNurse-VLA: Robotic Scrub Nurse System based on Vision-Language-Action Model

IROS 2025

In modern healthcare, the demand for autonomous robotic assistants has grown significantly, particularly in the operating room, where surgical tasks require precision and reliability. Robotic scrub nurses have emerged as a promising solution to improve efficiency and reduce human error during surger

Cited by 28SourcecodeScholar
2025

Sample-aware Adaptive Structured Pruning for Large Language Models

AAAI 2025technical

Large language models (LLMs) have achieved outstanding performance in natural language processing, but enormous model sizes and high computational costs limit their practical deployment. Structured pruning can effectively reduce the resource demands for deployment by removing redundant model paramet…

2025

TEST-V: TEst-time Support-set Tuning for Zero-shot Video Classification

IJCAI 2025

Recently, adapting Vision Language Models (VLMs) to zero-shot visual classification by tuning class embedding with a few prompts (Test-time Prompt Tuning, TPT) or replacing class names with generated visual samples (support-set) has shown promising results. However, TPT cannot avoid the semantic gap

Cited by 0SourcePDFScholar
2025

Tailless Flapping-Wing Robot With Bio-Inspired Elastic Passive Legs for Multi-Modal Locomotion

RA-L 2025

Flapping-wing robots offer significant versatility; however, achieving efficient multi-modal locomotion remains challenging. This paper presents the design, modeling, and experimentation of a novel tailless flapping-wing robot with three independently actuated pairs of wings. Inspired by the leg mor

Cited by 4SourceScholar
2025

Topology-of-Question-Decomposition: Enhancing Large Language Models with Information Retrieval for Knowledge-Intensive Tasks

COLING 2025main

Large language models (LLMs) are increasingly deployed for general problem-solving across various domains yet remain constrained to chaining immediate reasoning steps and depending solely on parametric knowledge. Integrating an information retrieval system directly into the reasoning process of LLMs…

2025

Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain Adaptation

AAAI 2025technical

Conventional multi-source domain few-shot adaptation (MFDA) faces the challenge of further reducing the load on edge-side devices in low-resource scenarios. Considering the native language-supervised advantage of CLIP and the plug-and-play nature of prompt to transfer CLIP efficiently, this paper in…

2024

Autonomous Behavior Planning For Humanoid Loco-manipulation Through Grounded Language Model

IROS 2024poster

Enabling humanoid robots to perform autonomously loco-manipulation in unstructured environments is crucial and highly challenging for achieving embodied intelligence. This involves robots being able to plan their actions and behaviors in long-horizon tasks while using multi-modality to perceive devi…

Cited by 4SourceScholar
2024

Boosting the Adversarial Robustness of Graph Neural Networks: An OOD Perspective

ICLR 2024poster

Current defenses against graph attacks often rely on certain properties to eliminate structural perturbations by identifying adversarial edges from normal edges. However, this dependence makes defenses vulnerable to adaptive (white-box) attacks from adversaries with the same knowledge. Adversarial t…

2024

Continuous Heatmap Regression for Pose Estimation via Implicit Neural Representation

NeurIPS 2024poster

Heatmap regression has dominated human pose estimation due to its superior performance and strong generalization. To meet the requirements of traditional explicit neural networks for output form, existing heatmap-based methods discretize the originally continuous heatmap representation into 2D pixel…

2024

Diagnosing the Compositional Knowledge of Vision Language Models from a Game-Theoretic View

ICML 2024poster

Compositional reasoning capabilities are usually considered as fundamental skills to characterize human perception. Recent studies show that current Vision Language Models (VLMs) surprisingly lack sufficient knowledge with respect to such capabilities. To this end, we propose to thoroughly diagnose…

2024

DuetSim: Building User Simulator with Dual Large Language Models for Task-Oriented Dialogues

COLING 2024main

User Simulators play a pivotal role in training and evaluating task-oriented dialogue systems. Traditional user simulators typically rely on human-engineered agendas, resulting in generated responses that often lack diversity and spontaneity. Although large language models (LLMs) exhibit a remarkabl…

2024

Enhancing Semantics in Multimodal Chain of Thought via Soft Negative Sampling

COLING 2024main

Chain of thought (CoT) has proven useful for problems requiring complex reasoning. Many of these problems are both textual and multimodal. Given the inputs in different modalities, a model generates a rationale and then uses it to answer a question. Because of the hallucination issue, the generated…

2024

Event-assisted Low-Light Video Object Segmentation

CVPR 2024poster

In the realm of video object segmentation (VOS) the challenge of operating under low-light conditions persists resulting in notably degraded image quality and compromised accuracy when comparing query and memory frames for similarity computation. Event cameras characterized by their high dynamic ran…

2024

Event-based Head Pose Estimation: Benchmark and Method

ECCV 2024poster

"Head pose estimation (HPE) is crucial for various applications, including human-computer interaction, augmented reality, and driver monitoring. However, traditional RGB-based methods struggle in challenging conditions like sudden movement and extreme lighting. Event cameras, as a neuromorphic senso…

2024

HYPERmotion: Learning Hybrid Behavior Planning for Autonomous Loco-manipulation

CoRL 2024poster

Enabling robots to autonomously perform hybrid motions in diverse environments can be beneficial for long-horizon tasks such as material handling, household chores, and work assistance. This requires extensive exploitation of intrinsic motion capabilities, extraction of affordances from rich environ…

Cited by 6SourcecodeScholar
2024

Improving Personalized Sentiment Representation with Knowledge-enhanced and Parameter-efficient Layer Normalization

COLING 2024main

Existing studies on personalized sentiment classification consider a document review as an overall text unit and incorporate backgrounds (i.e., user and product information) to learn sentiment representation. However, it is difficult when these methods meet the current pretrained language models (PL…

2024

Instruction Tuning with Retrieval-based Examples Ranking for Aspect-based Sentiment Analysis

ACL 2024findings

Aspect-based sentiment analysis (ABSA) identifies sentiment information related to specific aspects and provides deeper market insights to businesses and organizations. With the emergence of large language models (LMs), recent studies have proposed using fixed examples for instruction tuning to refo…

2024

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

ICML 2024poster

Large Vision-Language Models (LVLMs) show significant strides in general-propose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in t…

Cited by 84SourcePDFScholar
2024

Personalized LoRA for Human-Centered Text Understanding

AAAI 2024technical

Effectively and efficiently adapting a pre-trained language model (PLM) for human-centered text understanding (HCTU) is challenging since user tokens are million-level in most personalized applications and do not have concrete explicit semantics. A standard and parameter-efficient approach (e.g., Lo…

2024

Rethinking Prior Information Generation with CLIP for Few-Shot Segmentation

CVPR 2024poster

Few-shot segmentation remains challenging due to the limitations of its labeling information for unseen classes. Most previous approaches rely on extracting high-level feature maps from the frozen visual encoder to compute the pixel-wise similarity as a key prior guidance for the decoder. However su…

2024

SoftMCL: Soft Momentum Contrastive Learning for Fine-grained Sentiment-aware Pre-training

COLING 2024main

The pre-training for language models captures general language understanding but fails to distinguish the affective impact of a particular context to a specific word. Recent works have sought to introduce contrastive learning (CL) for sentiment-aware pre-training in acquiring affective information.…

2024

Sparse Beats Dense: Rethinking Supervision in Radar-Camera Depth Completion

ECCV 2024poster

"It is widely believed that sparse supervision is worse than dense supervision in the field of depth completion, but the underlying reasons for this are rarely discussed. To this end, we revisit the task of radar-camera depth completion and present a new method with sparse LiDAR supervision to outpe…

2024

UC-NERF: Neural Radiance Field for Under-Calibrated Multi-View Cameras in Autonomous Driving

ICLR 2024poster

Multi-camera setups find widespread use across various applications, such as autonomous driving, as they greatly expand sensing capabilities. Despite the fast development of Neural radiance field (NeRF) techniques and their wide applications in both indoor and outdoor scenes, applying NeRF to multi…

Cited by 9SourcePDFScholar
2024

Wrong-of-Thought: An Integrated Reasoning Framework with Multi-Perspective Verification and Wrong Information

EMNLP 2024finding

Chain-of-Thought (CoT) has become a vital technique for enhancing the performance of Large Language Models (LLMs), attracting increasing attention from researchers. One stream of approaches focuses on the iterative enhancement of LLMs by continuously verifying and refining their reasoning outputs fo…

2024

Zero-Shot Cross-Domain Dialogue State Tracking via Dual Low-Rank Adaptation

ACL 2024long

Zero-shot dialogue state tracking (DST) seeks to enable dialogue systems to transition to unfamiliar domains without manual annotation or extensive retraining. Prior research has approached this objective by embedding prompts into language models (LMs). Common methodologies include integrating promp…

2023

Birder: Communication-Efficient 1-bit Adaptive Optimizer for Practical Distributed DNN Training

NeurIPS 2023poster

Various gradient compression algorithms have been proposed to alleviate the communication bottleneck in distributed learning, and they have demonstrated effectiveness in terms of high compression ratios and theoretical low communication complexity. However, when it comes to practically training mod…

Cited by 1SourcePDFScholar
2023

Domain Generalization via Switch Knowledge Distillation for Robust Review Representation

ACL 2023findings

Applying neural models injected with in-domain user and product information to learn review representations of unseen or anonymous users incurs an obvious obstacle in content-based recommender systems. For the generalization of the in-domain classifier, most existing models train an extra plain-text…

2023

FedID: Federated Interactive Distillation for Large-Scale Pretraining Language Models

EMNLP 2023long main

The growing concerns and regulations surrounding the protection of user data privacy have necessitated decentralized training paradigms. To this end, federated learning (FL) is widely studied in user-related natural language processing (NLP). However, it suffers from several critical limitations inc…

Cited by 0SourcecodeScholar
2023

Implicit Identity Leakage: The Stumbling Block to Improving Deepfake Detection Generalization

CVPR 2023poster

In this paper, we analyse the generalization ability of binary classifiers for the task of deepfake detection. We find that the stumbling block to their generalization is caused by the unexpected learned identity representation on images. Termed as the Implicit Identity Leakage, this phenomenon has…

2023

Joint Multimodal Entity-Relation Extraction Based on Edge-Enhanced Graph Alignment Network and Word-Pair Relation Tagging

AAAI 2023technical

Multimodal named entity recognition (MNER) and multimodal relation extraction (MRE) are two fundamental subtasks in the multimodal knowledge graph construction task. However, the existing methods usually handle two tasks independently, which ignores the bidirectional interaction between them. This p…

2023

Learning to Memorize Entailment and Discourse Relations for Persona-Consistent Dialogues

AAAI 2023technical

Maintaining engagement and consistency is particularly important in dialogue systems. Existing works have improved the performance of dialogue systems by intentionally learning interlocutor personas with sophisticated network structures. One issue with this approach is that it requires more personal…

2023

Roller-Quadrotor: A Novel Hybrid Terrestrial/Aerial Quadrotor with Unicycle-Driven and Rotor-Assisted Turning

IROS 2023poster

The Roller-Quadrotor is a novel quadrotor that combines the maneuverability of aerial drones with the endurance of ground vehicles. This work focuses on the design, modeling, and experimental validation of the Roller-Quadrotor. Flight capabilities are achieved through a quadrotor config-uration, wit…

Cited by 5SourceScholar
2023

Supervised Contrastive Few-Shot Learning for High-Frequency Time Series

AAAI 2023technical

Significant progress has been made in representation learning, especially with recent success on self-supervised contrastive learning. However, for time series with less intuitive or semantic meaning, sampling bias may be inevitably encountered in unsupervised approaches. Although supervised contras…

2023

Towards Understanding the Generalization of Deepfake Detectors from a Game-Theoretical View

ICCV 2023poster

This paper aims to explain the generalization of deepfake detectors from the novel perspective of multi-order interactions among visual concepts. Specifically, we propose three hypotheses: 1. Deepfake detectors encode multi-order interactions among visual concepts, in which the low-order interacti…

Cited by 17PDFScholar
2022

Accelerating Inference for Pretrained Language Models by Unified Multi-Perspective Early Exiting

COLING 2022main

Conditional computation algorithms, such as the early exiting (EE) algorithm, can be applied to accelerate the inference of pretrained language models (PLMs) while maintaining competitive performance on resource-constrained devices. However, this approach is only applied to the vertical architecture…

2022

End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video Grounding

ACL 2022long

Natural language spatial video grounding aims to detect the relevant objects in video frames with descriptive sentences as the query. In spite of the great advances, most existing methods rely on dense video frame annotations, which require a tremendous amount of human effort. To achieve effective g…

Cited by 41SourcePDFScholar
2022

Explaining Deepfake Detection by Analysing Image Matching

ECCV 2022poster

"This paper aims to interpret how deepfake detection models learn artifact features of images when just supervised by binary labels. To this end, three hypotheses from the perspective of image matching are proposed as follows. 1. Deepfake detection models indicate real/fake images based on visual co…

2022

Incrementally Stochastic and Accelerated Gradient Information Mixed Optimization for Manipulator Motion Planning

RA-L 2022

This paper introduces a novel motion planner, incrementally stochastic and accelerated gradient information mixed optimization (iSAGO), for robotic manipulators in a narrow workspace. Primarily, we propose the overall scheme of iSAGO informed by the mixed momenta for an efficient constrained optimiz

Cited by 7SourceScholar
2022

Interpretable Generative Adversarial Networks

AAAI 2022technical

Learning a disentangled representation is still a challenge in the field of the interpretability of generative adversarial networks (GANs). This paper proposes a generic method to modify a traditional GAN into an interpretable GAN, which ensures that filters in an intermediate layer of the generator…

2022

Knowledge Distillation with Reptile Meta-Learning for Pretrained Language Model Compression

COLING 2022main

The billions, and sometimes even trillions, of parameters involved in pre-trained language models significantly hamper their deployment in resource-constrained devices and real-time applications. Knowledge distillation (KD) can transfer knowledge from the original model (i.e., teacher) into a compac…

2022

Locally Aggregated Feature Attribution on Natural Language Model Understanding

NAACL 2022long

With the growing popularity of deep-learning models, model understanding becomes more important. Much effort has been devoted to demystify deep neural networks for better explainability. Some feature attribution methods have shown promising results in computer vision, especially the gradient-based m…

2021

Virtual-Fixture Based Drilling Control for Robot-Assisted Craniotomy: Learning From Demonstration

RA-L 2021

One of the promising solutions for drilling craniotomy is robot-assisted surgery with human guidance. The present study deals with a piecewise collaborative drilling task assisted by a robot while containing aligning and drilling. It can enable surgeons to complete the operation more efficiently and

Cited by 37SourceScholar
2020

Discovering Subsequence Patterns for Next POI Recommendation

IJCAI 2020poster

Next Point-of-Interest (POI) recommendation plays an important role in location-based services. State-of-the-art methods learn the POI-level sequential patterns in the user's check-in sequence but ignore the subsequence patterns that often represent the socio-economic activities or coherence of pref…

Cited by 0SourcePDFScholar