← Search

peng zhang

122 accepted papers

2026

AdAEM: An Adaptively and Automated Extensible Evaluation Method of LLMs' Value Difference

ICLR 2026oral

Assessing Large Language Models (LLMs)' underlying value differences enables comprehensive comparison of their misalignment, cultural adaptability, and biases. Nevertheless, current value measurement methods face the informativeness challenge: with often outdated, contaminated, or generic test quest…

Cited by 0SourcecodeScholar
2026

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO

ICML 2026poster

Deploying GRPO on Flow Matching models has proven effective for text-to-image generation. However, existing paradigms typically propagate an outcome-based reward to all preceding denoising steps without distinguishing the local effect of each step. Moreover, current group-wise ranking mainly compare…

Cited by 0SourceScholar
2026

DarkBench+: An Extended Benchmark for Evaluating Dark Patterns in Large Language Models

AAAI 2026technical

With the widespread deployment of large language models (LLMs) in human-computer interaction, dark patterns have extended from traditional visual interfaces to conversational AI systems. While existing research has confirmed the prevalence of dark patterns in LLMs, current evaluation benchmarks face

Cited by 0SourcePDFScholar
2026

Disentangling Consensus and Value-Specific Representations for Controllable Pluralistic Value Alignment of LLMs

ICML 2026poster

With the widespread deployment of large language models (LLMs), aligning model outputs with pluralistic human values has become an important research problem. Recent approaches that train task-specific experts and merge them through parameter aggregation have shown promise for pluralistic alignment.…

Cited by 0SourceScholar
2026

FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation

AAAI 2026technical

Recent unified models have demonstrated that the reasoning capacity of Multimodal Large Language Models (MLLMs) can be leveraged to facilitate diffusion-based image generation with impressive flexibility and performance. However, approaches that rely heavily on MLLMs for high-level semantic encoding

Cited by 0SourcePDFScholar
2026

FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention

AAAI 2026technical

Medical vision-language pre-training (VLP) offers significant potential for advancing medical image understanding by leveraging paired image-report data. However, existing methods are limited by False Negatives (FaNe) induced by semantically similar texts and insufficient fine-grained cross-modal al

Cited by 0SourcePDFScholar
2026

GOCM: Single-Step Graph Outlier Synthesis via Origin Consistency Model

ICML 2026poster

Supervised Graph Outlier Detection has long been constrained by severe class imbalance, and although recent diffusion-based augmentation methods have improved sample quality, their practical utility is hindered by the high computational costs of multi-step iterative sampling and the stochasticity of…

Cited by 0SourceScholar
2026

IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization

AAAI 2026technical

Trained on various human-authored corpora, Large Language Models (LLMs) have demonstrated a certain capability of reflecting specific human-like traits (e.g., personality or values) by prompting, benefiting applications like personalized LLMs and social simulations. However, existing methods suffer

Cited by 0SourcePDFScholar
2026

MERLIN: Building Low-SNR Robust Multimodal LLMs for Electromagnetic Signals

CVPR 2026

The paradigm of Multimodal Large Language Models (MLLMs) offers a promising blueprint for advancing the electromagnetic (EM) domain. However, prevailing approaches often deviate from the native MLLM paradigm, instead using task-specific or pipelined architectures that lead to fundamental limitations

Cited by 0SourcecodeScholar
2026

MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models

ICLR 2026poster

Quantization significantly accelerates inference in large language models (LLMs) by replacing original high-precision matrices with low-precision counterparts. Recent advances in weight-activation quantization have primarily focused on mapping both weights and activations to the INT4 format. Althoug…

Cited by 0SourcecodeScholar
2026

MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions

AAAI 2026technical

Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despite substantial work investigating the trustworthiness of language models, MMLMs

Cited by 0SourcePDFScholar
2026

Ophiuchus: Incentivizing Tool-augmented ''Think with Images'' for Joint Medical Segmentation, Understanding and Reasoning

ICML 2026poster

Recent medical MLLMs have made significant progress in generating step by step textual reasoning chains. However, they still struggle with complex clinical tasks that necessitate dynamic and iterative focusing on fine-grained visual regions. To close this gap, we introduce Ophiuchus, a versatile, to…

Cited by 0SourceScholar
2026

Phased One-Step Adversarial Equilibrium for Video Diffusion Models

AAAI 2026technical

Video diffusion generation suffers from critical sampling efficiency bottlenecks, particularly for large-scale models and long contexts. Existing video acceleration methods, adapted from image-based techniques, lack a single-step distillation ability for large-scale video models and task generalizat

Cited by 0SourcePDFScholar
2026

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment

CVPR 2026

Existing facial reenactment methods struggle with a trade-off between expressiveness and fine-grained controllability. Holistic facial reenactment models often sacrifice granular control for expressiveness, while methods designed for control may struggle with fidelity and robust disentanglement. Ins

Cited by 0SourceScholar
2026

Unified Interaction Consistency Learning for Single-Source Domain-Generalized Object Detection in Urban Scene

AAAI 2026technical

Domain generalization remains a critical challenge for deploying neural networks, particularly in out-of-distribution object detection. The distributional discrepancy between training (e.g., daytime-sunny) and the realistic condition (e.g., night-rainy) inevitably produces imprecise localization and

Cited by 0SourcePDFScholar
2025

3D-RPE: Enhancing Long-Context Modeling Through 3D Rotary Position Encoding

AAAI 2025technical

An essential component in Large Language Models (LLMs) is Rotary Position Encoding (RoPE) , which efficiently manages positional dependencies in long-context modeling. However, when the number of input tokens surpasses the pretrained capacity of LLMs, their ability to process and generate text is ma…

2025

Advancing Generalizable Tumor Segmentation with Anomaly-Aware Open-Vocabulary Attention Maps and Frozen Foundation Diffusion Models

CVPR 2025poster

We explore Generalizable Tumor Segmentation, aiming to train a single model for zero-shot tumor segmentation across diverse anatomical regions. Existing methods face limitations related to segmentation quality, scalability, and the range of applicable imaging modalities. In this paper, we uncover th…

2025

Aerodynamic Coefficients Prediction via Cross-Attention Fusion and Physical-Informed Training

AAAI 2025technical

Aerodynamic coefficient prediction is pivotal in aircraft and vehicles' design, performance evaluation, and motion control. Integrating artificial neural networks into aerodynamic coefficient prediction offers a promising alternative to traditional numerical methods burdened by extensive computation…

Cited by 0SourcePDFScholar
2025

Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

ICCV 2025poster

Recent character image animation methods based on diffusion models, such as Animate Anyone, have made significant progress in generating consistent and generalizable character animations. However, these approaches fail to produce reasonable associations between characters and their environments. To…

Cited by 0SourcePDFScholar
2025

CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCR

CVPR 2025poster

Scene Text Retrieval (STR) seeks to identify all images containing a given query string. Existing methods typically rely on an explicit Optical Character Recognition (OCR) process of text spotting or localization, which is susceptible to complex pipelines and accumulated errors. To settle this, we r…

Cited by 0SourcePDFScholar
2025

Conformal Anomaly Detection in Event Sequences

ICML 2025poster

Anomaly detection in continuous-time event sequences is a crucial task in safety-critical applications. While existing methods primarily focus on developing a superior test statistic, they fail to provide guarantees regarding the false positive rate (FPR), which undermines their reliability in pract…

Cited by 0SourcePDFScholar
2025

Consistency Rating of Semantic Transparency: an Evaluation Method for Metaphor Competence in Idiom Understanding Tasks

COLING 2025main

Idioms condense complex semantics into fixed phrases, and their meaning is often not directly connected to the literal meaning of their constituent words, making idiom comprehension a test of metaphor competence. Metaphor, as a cognitive process in human beings, has not yet found an effective evalua…

Cited by 0SourcePDFScholar
2025

Controllable and Expressive One-Shot Video Head Swapping

ICCV 2025poster

In this paper, we propose a novel diffusion-based multi-condition controllable framework for video head swapping, which seamlessly transplant a human head from a static image into a dynamic video, while preserving the original body and background of target video, and further allowing to tweak head e…

Cited by 0SourcePDFScholar
2025

Don’t Half-listen: Capturing Key-part Information in Continual Instruction Tuning

ACL 2025long

Instruction tuning for large language models (LLMs) can drive them to produce results consistent with human goals in specific downstream tasks. However, the process of continual instruction tuning (CIT) for LLMs may bring about the catastrophic forgetting (CF) problem, where previously learned abili…

2025

ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering

ICCV 2025poster

Precisely evaluating semantic alignment between text prompts and generated videos remains a challenge in Text-to-Video (T2V) Generation. Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without fine-grained alignment details, failing to align with human pre…

Cited by 0SourcePDFScholar
2025

Equilibrium Policy Generalization: A Reinforcement Learning Framework for Cross-Graph Zero-Shot Generalization in Pursuit-Evasion Games

NeurIPS 2025poster

Equilibrium learning in adversarial games is an important topic widely examined in the fields of game theory and reinforcement learning (RL). Pursuit-evasion game (PEG), as an important class of real-world games from the fields of robotics and security, requires exponential time to be accurately sol…

Cited by 0SourceScholar
2025

Exploring Timeline Control for Facial Motion Generation

CVPR 2025poster

This paper introduces a new control signal for facial motion generation: timeline control. Compared to audio and text signals, timelines provide more fine-grained control, such as generating specific facial motions with precise timing. Users can specify a multi-track timeline of facial actions arran…

Cited by 0SourcePDFScholar
2025

GLoCIM: Global-view Long Chain Interest Modeling for news recommendation

COLING 2025main

Accurately recommending candidate news articles to users has always been the core challenge of news recommendation system. News recommendations often require modeling of user interest to match candidate news. Recent efforts have primarily focused on extracting local subgraph information in a global…

Cited by 1SourcePDFScholar
2025

Gamma Distribution PCA-Enhanced Feature Learning for Angle-Robust SAR Target Recognition

ICML 2025poster

Scattering characteristics of synthetic aperture radar (SAR) targets are typically related to observed azimuth and depression angles. However, in practice, it is difficult to obtain adequate training samples at all observation angles, which probably leads to poor robustness of deep networks. In thi…

2025

Graph Anomaly Detection via Multi-Scale Reconstruction of Graph Encoder-Decoder Networks

ICASSP 2025accepted

Existing unsupervised graph anomaly detection (GAD) methods can be categorized into reconstruction based methods and contrastive learning based methods. The principle of reconstruction methods is to capture anomalous nodes based on data reconstruction errors. However, existing reconstruction methods…

Cited by 0SourceScholar
2025

Large Language Models Enhanced Personalized Graph Neural Architecture Search in Federated Learning

AAAI 2025technical

Personalized federated learning (PFL) on graphs is an emerging field focusing on the collaborative development of architectures across multiple clients, each with distinct graph data distributions while adhering to strict privacy standards. This area often requires extensive expert intervention in m…

2025

Leveraging SD Map to Augment HD Map-based Trajectory Prediction

CVPR 2025poster

Latest trajectory prediction models in real-world autonomous driving systems often rely on online High-Definition (HD) maps to understand the road environment.However, online HD maps suffer from perception errors and feature redundancy, which hinder the performance of HD map-based trajectory predict…

Cited by 0SourcePDFScholar
2025

NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding

NeurIPS 2025poster

Underwater exploration offers critical insights into our planet and attracts increasing attention for its broader applications in resource exploration, national security, etc. We study the underwater scene understanding methods, which aim to achieve automated underwater exploration. The underwater s…

Cited by 0SourcecodeScholar
2025

Node-Centric Meta Structure Search in Heterogeneous Graphs

ICASSP 2025accepted

Heterogeneous graphs are increasingly used to represent complex real-world scenarios with diverse entities and interactions by meta structures. Recently, the search of meta structures is combined with graph neural architecture search to automatically extract the semantic knowledge for various tasks…

Cited by 0SourceScholar
2025

NovPhy: A Physical Reasoning Benchmark for Open-World AI Systems Author Links Open Overlay Panel (Abstract Reprint)

IJCAI 2025

Due to the emergence of AI systems that interact with the physical environment, there is an increased interest in incorporating physical reasoning capabilities into those AI systems. But is it enough to only have physical reasoning capabilities to operate in a real physical environment? In the real

2025

OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking

NeurIPS 2025poster

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video content from input text while emulating the target identity'…

Cited by 0SourceScholar
2025

SIGraph: Saliency Image-Graph Network for Retinal Disease Classification in Fundus Image

AAAI 2025technical

An efficient and precise diagnosis of retinal diseases is a fundamental goal for auxiliary diagnostic systems in ophthalmology. Inspired by the importance of scattered subtle lesions in manual retinal disease diagnosis, recent research has achieved state-of-the-art performance by mining information…

Cited by 0SourcePDFScholar
2025

Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark

CVPR 2025poster

Document image retrieval (DIR) aims to retrieve document images from a gallery according to a given query. Existing DIR methods are primarily based on image queries that retrieve documents within the same coarse semantic category, e.g., newspapers or receipts. However, these methods struggle to effe…

2025

Tracing Copied Pixels and Regularizing Patch Affinity in Copy Detection

ICCV 2025poster

Image Copy Detection (ICD) aims to identify manipulated content between image pairs through robust feature representation learning. While self-supervised learning (SSL) has advanced ICD systems, existing view-level contrastive methods struggle with sophisticated edits due to insufficient fine-graine…

Cited by 0SourcePDFScholar
2024

A Decision-Making Algorithm for Robotic Breast Ultrasound High-Quality Imaging via Broad Reinforcement Learning From Demonstration

RA-L 2024

Robotic breast ultrasound (RBUS) aims to standardize breast ultrasonography, reduce the workload of sonographers, and provide high-quality ultrasound (US) images for subsequent diagnosis. In the process of RBUS screening, adjusting the US probe correctly and efficiently to acquire high-quality US im

Cited by 11SourceScholar
2024

A Quantum-Inspired Matching Network with Linguistic Theories for Metaphor Detection

COLING 2024main

Enabling machines with the capability to recognize and comprehend metaphors is a crucial step toward achieving artificial intelligence. In linguistic theories, metaphor can be identified through Metaphor Identification Procedure (MIP) or Selectional Preference Violation (SPV), both of which are typi…

2024

BoolQuestions: Does Dense Retrieval Understand Boolean Logic in Language?

EMNLP 2024finding

Dense retrieval, which aims to encode the semantic information of arbitrary text into dense vector representations or embeddings, has emerged as an effective and efficient paradigm for text retrieval, consequently becoming an essential component in various natural language processing systems. These…

2024

CcDPM: A Continuous Conditional Diffusion Probabilistic Model for Inverse Design

AAAI 2024technical

Engineering design methods aim to generate new designs that meet desired performance requirements. Past work has directly introduced conditional Generative Adversarial Networks (cGANs) into this field and achieved promising results in single-point design problems(one performance requirement under on…

Cited by 4SourcePDFScholar
2024

DENEVIL: TOWARDS DECIPHERING AND NAVIGATING THE ETHICAL VALUES OF LARGE LANGUAGE MODELS VIA INSTRUCTION LEARNING

ICLR 2024poster

Large Language Models (LLMs) have made unprecedented breakthroughs, yet their increasing integration into everyday life might raise societal risks due to generated unethical content. Despite extensive study on specific issues like bias, the intrinsic values of LLMs remain largely unexplored from a m…

Cited by 13SourcePDFScholar
2024

Deciphering Rumors: A Multi-Task Learning Approach with Intent-aware Hierarchical Contrastive Learning

EMNLP 2024main

Social networks are rife with noise and misleading information, presenting multifaceted challenges for rumor detection. In this paper, from the perspective of human cognitive subjectivity, we introduce the mining of individual latent intentions and propose a novel multi-task learning framework, the…

Cited by 1SourcePDFScholar
2024

DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction

CVPR 2024poster

Audio-visual saliency prediction can draw support from diverse modality complements but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies denoising diffusion models have shown more promising in unifying task fra…

Cited by 6SourcePDFScholar
2024

Distilling Causal Effect of Data in Continual Few-shot Relation Learning

COLING 2024main

Continual Few-Shot Relation Learning (CFRL) aims to learn an increasing number of new relational patterns from a data stream. However, due to the limited number of samples and the continual training mode, this method frequently encounters the catastrophic forgetting issues. The research on causal in…

2024

Force-Position Hybrid Control for Robot Assisted Thoracic-Abdominal Puncture With Respiratory Movement

RA-L 2024

Percutaneous puncture is a widely used procedure in the diagnosis and therapy of cancer such as biopsy and ablation operations, while the organs in the thoracic and abdominal cavities are significantly affected by patients' respiratory movement. In this study, a robotic puncture system with respirat

Cited by 7SourceScholar
2024

LA-UCL: LLM-Augmented Unsupervised Contrastive Learning Framework for Few-Shot Text Classification

COLING 2024main

The few-shot tasks require the model to have the ability to generalize from a few samples. However, due to the lack of cognitive ability, the current works cannot fully utilize limited samples to expand the sample space and still suffer from overfitting issues. To address the problems, we propose a…

Cited by 11SourcePDFScholar
2024

MHPS: Multimodality-Guided Hierarchical Policy Search for Knowledge Graph Reasoning

ICASSP 2024accepted

Recently, path inference-based knowledge graph reasoning (KGR) methods have attracted great attention due to their good performance and interpretability. However, as the number of hops increases, the search space grows exponentially, making the reward sparse and the process of reasoning difficult. T…

Cited by 0SourceScholar
2024

Meta Structure Search for Link Weight Prediction in Heterogeneous Graphs

ICASSP 2024accepted

Recently link weight prediction has attracted an increasing research interest due to its merits in quantifying the strength between nodes within a graph. Nonetheless, current link weight prediction methods focus solely on graph topology, disregarding node feature information embedded in graphs. In r…

Cited by 0SourceScholar
2024

MuseChat: A Conversational Music Recommendation System for Videos

CVPR 2024highlight

Music recommendation for videos attracts growing interest in multi-modal research. However existing systems focus primarily on content compatibility often ignoring the users' preferences. Their inability to interact with users for further refinements or to provide explanations leads to a less satisf…

2024

Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization

EMNLP 2024finding

Large language models (LLMs) have revolutionized the role of AI, yet pose potential social risks. To steer LLMs towards human preference, alignment technologies have been introduced and gained increasing attention. Nevertheless, existing methods heavily rely on high-quality positive-negative trainin…

2024

Neural Jump-Diffusion Temporal Point Processes

ICML 2024spotlight

We present a novel perspective on temporal point processes (TPPs) by reformulating their intensity processes as solutions to stochastic differential equations (SDEs). In particular, we first prove the equivalent SDE formulations of several classical TPPs, including Poisson processes, Hawkes processe…

Cited by 5SourcePDFScholar
2024

On Hardware-efficient Inference in Probabilistic Circuits

UAI 2024poster

Probabilistic circuits (PCs) offer a promising avenue to perform embedded reasoning under uncertainty. They support efficient and exact computation of various probabilistic inference tasks by design. Hence, hardware-efficient computation of PCs is highly interesting for edge computing applications.…

Cited by 0SourcePDFScholar
2024

On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models

IJCAI 2024poster

Big models have achieved revolutionary breakthroughs in the field of AI, but they also pose potential ethical and societal risks to humans. Addressing such problems, alignment technologies were introduced to make these models conform to human preferences and values. Despite the considerable advancem…

Cited by 12SourcePDFScholar
2024

PuLID: Pure and Lightning ID Customization via Contrastive Alignment

NeurIPS 2024poster

We propose Pure and Lightning ID customization (PuLID), a novel tuning-free ID customization method for text-to-image generation. By incorporating a Lightning T2I branch with a standard diffusion one, PuLID introduces both contrastive alignment loss and accurate ID loss, minimizing disruption to the…

2024

Quantum Topic Model: Topic Modeling Using Variational Quantum Circuits

ICASSP 2024accepted

Quantum machine learning aims to leverage quantum computing to enhance the computing capabilities and storage efficiency of classical machine learning. Among them, quantum generative models, as a type of unsupervised machine learning, have been proven to learn distributions that are outside of class…

Cited by 0SourceScholar
2024

Quantum-Inspired Neural Network with Runge-Kutta Method

AAAI 2024technical

In recent years, researchers have developed novel Quantum-Inspired Neural Network (QINN) frameworks for the Natural Language Processing (NLP) tasks, inspired by the theoretical investigations of quantum cognition. However, we have found that the training efficiency of QINNs is significantly lower th…

Cited by 2SourcePDFScholar
2024

Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning

ICML 2024poster

Trustworthiness is an essential prerequisite for the real-world application of large language models. In this paper, we focus on the trustworthiness of language models with respect to retrieval augmentation. Despite being supported with external evidence, retrieval-augmented generation still suffers…

2024

Urban Waterlogging Detection: A Challenging Benchmark and Large-Small Model Co-Adapter

ECCV 2024poster

"Urban waterlogging poses a major risk to public safety and infrastructure. Conventional methods using water-level sensors need high-maintenance to hardly achieve full coverage. Recent advances employ surveillance camera imagery and deep learning for detection, yet these struggle amidst scarce data…

2023

An Intent-based and Annotation-free Method for Duplicate Question Detection in CQA Forums

EMNLP 2023long findings

With the advent of large language models (LLMs), Community Question Answering (CQA) forums offer well-curated questions and answers that can be utilized for instruction-tuning, effectively training LLMs to be aligned with human intents. However, the issue of duplicate questions arises as the volume…

Cited by 0SourceScholar
2023

Are Intermediate Layers and Labels Really Necessary? A General Language Model Distillation Method

ACL 2023findings

The large scale of pre-trained language models poses a challenge for their deployment on various devices, with a growing emphasis on methods to compress these models, particularly knowledge distillation. However, current knowledge distillation methods rely on the model’s intermediate layer features…

2023

CASP-Net: Rethinking Video Saliency Prediction From an Audio-Visual Consistency Perceptual Perspective

CVPR 2023poster

Incorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP methods are capable of exploiting semantic correlation between vision and audio modalitie…

2023

Domain-specific Attention with Distributional Signatures for Multi-Domain End-to-end Task-Oriented Dialogue

ACL 2023findings

The end-to-end task-oriented dialogue system has achieved great success in recent years. Most of these dialogue systems need to accommodate multi-domain dialogue in real-world scenarios. However, due to the high cost of dialogue data annotation and the scarcity of labeled dialogue data, existing met…

2023

GKD: A General Knowledge Distillation Framework for Large-scale Pre-trained Language Model

ACL 2023industry

Currently, the reduction in the parameter scale of large-scale pre-trained language models (PLMs) through knowledge distillation has greatly facilitated their widespread deployment on various devices. However, the deployment of knowledge distillation systems faces great challenges in real-world indu…

2023

GLM-130B: An Open Bilingual Pre-trained Model

ICLR 2023poster

We introduce GLM-130B, a bilingual (English and Chinese) pre-trained language model with 130 billion parameters. It is an attempt to open-source a 100B-scale model as good as GPT-3 (davinci) and unveil how models of such a scale can be successfully pre-trained. Over the course of this effort, we fac…

2023

Learning Expressive And Generalizable Motion Features For Face Forgery Detection

ICASSP 2023accepted

Previous face forgery detection methods mainly focus on appearance features, which may be easily attacked by sophisticated manipulation. Considering the majority of current face manipulation methods generate fake faces based on a single frame, which do not take frame consistency and coordination int…

Cited by 0SourceScholar
2023

LightFormer: Light-weight Transformer Using SVD-based Weight Transfer and Parameter Sharing

ACL 2023findings

Transformer has become an important technique for natural language processing tasks with great success. However, it usually requires huge storage space and computational cost, making it difficult to be deployed on resource-constrained edge devices. To compress and accelerate Transformer, we propose…

2023

ReGANIE: Rectifying GAN Inversion Errors for Accurate Real Image Editing

AAAI 2023technical

The StyleGAN family succeed in high-fidelity image generation and allow for flexible and plausible editing of generated images by manipulating the semantic-rich latent style space. However, projecting a real image into its latent space encounters an inherent trade-off between inversion quality and e…

Cited by 7SourcePDFScholar
2023

Rumor Detection on Social Media with Crowd Intelligence and ChatGPT-Assisted Networks

EMNLP 2023long main

In the era of widespread dissemination through social media, the task of rumor detection plays a pivotal role in establishing a trustworthy and reliable information environment. Nonetheless, existing research on rumor detection confronts several challenges: the limited expressive power of text encod…

Cited by 0SourceScholar
2023

SpeedDETR: Speed-aware Transformers for End-to-end Object Detection

ICML 2023poster

Vision Transformers (ViTs) have continuously achieved new milestones in object detection. However, the considerable computation and memory burden compromise their efficiency and generalization of deployment on resource-constraint devices. Besides, efficient transformer-based detectors designed by ex…

Cited by 3SourcePDFScholar
2023

UGC: Unified GAN Compression for Efficient Image-to-Image Translation

ICCV 2023poster

Recent years have witnessed the prevailing progress of Generative Adversarial Networks (GANs) in image-to-image translation. However, the success of these GAN models hinges on ponderous computational costs and labor-expensive training data. Current efficient GAN learning techniques often fall into t…

Cited by 4PDFcodeScholar
2023

XDailyDialog: A Multilingual Parallel Dialogue Corpus

ACL 2023long

High-quality datasets are significant to the development of dialogue models. However, most existing datasets for open-domain dialogue modeling are limited to a single language. The absence of multilingual open-domain dialog datasets not only limits the research on multilingual or cross-lingual trans…

2022

A Sequential Flow Control Framework for Multi-hop Knowledge Base Question Answering

EMNLP 2022main

One of the key challenges of knowledge base question answering (KBQA) is the multi-hop reasoning. Since in different hops, one attends to different parts of question, it is important to dynamically represent the question semantics for each hop. Existing methods, however, (i) infer the dynamic questi…

2022

ACENet: Attention Guided Commonsense Reasoning on Hybrid Knowledge Graph

EMNLP 2022main

Augmenting pre-trained language models (PLMs) with knowledge graphs (KGs) has demonstrated superior performance on commonsense reasoning. Given a commonsense based QA context (question and multiple choices), existing approaches usually estimate the plausibility of candidate choices separately based…

2022

ClusterFormer: Neural Clustering Attention for Efficient and Effective Transformer

ACL 2022long

Recently, a lot of research has been carried out to improve the efficiency of Transformer. Among them, the sparse pattern-based method is an important branch of efficient Transformers. However, some existing sparse methods usually use fixed patterns to select words, without considering similarities…

2022

DART: Articulated Hand Model with Diverse Accessories and Rich Textures

NeurIPS 2022accept

Hand, the bearer of human productivity and intelligence, is receiving much attention due to the recent fever of digital twins. Among different hand morphable models, MANO has been widely used in vision and graphics community. However, MANO disregards textures and accessories, which largely limits it…

2022

DoSEA: A Domain-specific Entity-aware Framework for Cross-Domain Named Entity Recogition

COLING 2022main

Cross-domain named entity recognition aims to improve performance in a target domain with shared knowledge from a well-studied source domain. The previous sequence-labeling based method focuses on promoting model parameter sharing among domains. However, such a paradigm essentially ignores the domai…

2022

Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation

EMNLP 2022main

Transformer has been demonstrated effective in Neural Machine Translation (NMT). However, it is memory-consuming and time-consuming in edge devices, resulting in some difficulties for real-time feedback. To compress and accelerate Transformer, we propose a Hybrid Tensor-Train (HTT) decomposition, wh…

Cited by 13SourcePDFScholar
2022

Learning Common Dependency Structure for Unsupervised Cross-Domain Ner

ICASSP 2022accepted

Unsupervised cross-domain NER task aims to solve the issues when data in a new domain are fully-unlabeled. It leverages labeled data from source domain to predict entities in unlabeled target domain. Since training models on large domain corpus is time-consuming, in this paper, we consider an altern…

Cited by 0SourceScholar
2022

Learning Multiple Explainable and Generalizable Cues for Face Anti-Spoofing

ICASSP 2022accepted

Although previous CNN based face anti-spoofing methods have achieved promising performance under intra-dataset testing, they suffer from poor generalization under cross-dataset testing. The main reason is that they learn the network with only binary supervision, which may learn arbitrary cues overfi…

Cited by 0SourceScholar
2022

Medical Ultrasound Image Quality Assessment for Autonomous Robotic Screening

RA-L 2022

Autonomous ultrasound scanning robots have attracted the attention of researchers, and the real-time quality assessment of ultrasound images is the key technology of them. Existing robot systems usually use pixel-level feature statistical methods such as grayscale, confidence map, etc. However, in c

Cited by 18SourceScholar
2022

MorphTE: Injecting Morphology in Tensorized Embeddings

NeurIPS 2022accept

In the era of deep learning, word embeddings are essential when dealing with text tasks. However, storing and accessing these embeddings requires a large amount of space. This is not conducive to the deployment of these models on resource-limited devices. Combining the powerful compression capabilit…

2022

Novel Multi-Criteria Sustainable Evaluation for Production Scheduling Based on Fuzzy Analytic Network Process and Cumulative Prospect Theory-Enhanced VIKOR

RA-L 2022

Achieving sustainability is currently an important development direction for the manufacturing industry. With the consideration of all pillars of sustainability, valid evaluation of sustainability during scheduling optimization has become an emerging critical issue in the production scheduling loops

Cited by 6SourceScholar
2022

Parameter-free Dynamic Graph Embedding for Link Prediction

NeurIPS 2022accept

Dynamic interaction graphs have been widely adopted to model the evolution of user-item interactions over time. There are two crucial factors when modelling user preferences for link prediction in dynamic interaction graphs: 1) collaborative relationship among users and 2) user personalized interact…

2022

Personalized Image Aesthetics Assessment With Rich Attributes

CVPR 2022poster

Personalized image aesthetics assessment (PIAA) is challenging due to its highly subjective nature. People's aesthetic tastes depend on diversified factors, including image characteristics and subject characters. The existing PIAA databases are limited in terms of annotation diversity, especially th…

Cited by 80PDFScholar
2022

QaDialMoE: Question-answering Dialogue based Fact Verification with Mixture of Experts

EMNLP 2022finding

Fact verification is an essential tool to mitigate the spread of false information online, which has gained a widespread attention recently. However, a fact verification in the question-answering dialogue is still underexplored. In this paper, we propose a neural network based approach called questi…

Cited by 5SourcePDFScholar
2022

Subgraph Neighboring Relations Infomax for Inductive Link Prediction on Knowledge Graphs

IJCAI 2022poster

Inductive link prediction for knowledge graph aims at predicting missing links between unseen entities, those not shown in training stage. Most previous works learn entity-specific embeddings of entities, which cannot handle unseen entities. Recent several methods utilize enclosing subgraph to obtai…

2022

Syntax-Based Graph Matching for Knowledge Base Question Answering

ICASSP 2022accepted

Semantic parsing is a mainstream method of knowledge base question answering task that first generates a set of logical forms according to question and knowledge base (KB), and then selects the most matching one to get answers. However, existing selection methods are usually based on word-level matc…

Cited by 0SourceScholar
2022

Towards Explainable Action Recognition by Salient Qualitative Spatial Object Relation Chains

AAAI 2022technical

In order to be trusted by humans, Artificial Intelligence agents should be able to describe rationales behind their decisions. One such application is human action recognition in critical or sensitive scenarios, where trustworthy and explainable action recognizers are expected. For example, reliable…

Cited by 6SourcePDFScholar
2021

Continuous Self-Attention Models with Neural ODE Networks

AAAI 2021technical

Stacked self-attention models receive widespread attention, due to its ability of capturing global dependency among words. However, the stacking of many layers and components generates huge parameters, leading to low parameter efficiency. In response to this issue, we propose a lightweight architect…

Cited by 21SourcePDFScholar
2021

HIP Network: Historical Information Passing Network for Extrapolation Reasoning on Temporal Knowledge Graph

IJCAI 2021poster

In recent years, temporal knowledge graph (TKG) reasoning has received significant attention. Most existing methods assume that all timestamps and corresponding graphs are available during training, which makes it difficult to predict future events. To address this issue, recent works learn to infer…

2021

Hybrid Adaptive Control Strategy for Continuum Surgical Robot Under External Load

RA-L 2021

Natural orifice transluminal endoscopic surgery (NOTES) has received significant attentions due to its minimal incision trauma compared with traditional multi-port robot assisted surgery. Continuum robot can be used in NOTES due to its high flexibility which can adapt to circuitous paths. However, t

Cited by 45SourceScholar
2021

Learning Position and Target Consistency for Memory-Based Video Object Segmentation

CVPR 2021poster

This paper studies the problem of semi-supervised video object segmentation(VOS). Multiple works have shown that memory-based approaches can be effective for video object segmentation. They are mostly based on pixel-level matching, both spatially and temporally. The main shortcoming of memory-based…

Cited by 134PDFScholar
2021

MDNN: A Multimodal Deep Neural Network for Predicting Drug-Drug Interaction Events

IJCAI 2021poster

The interaction of multiple drugs could lead to serious events, which causes injuries and huge medical costs. Accurate prediction of drug-drug interaction (DDI) events can help clinicians make effective decisions and establish appropriate therapy programs. Recently, many AI-based techniques have bee…

2021

Multiphish: Multi-Modal Features Fusion Networks for Phishing Detection

ICASSP 2021accepted

Phishing is an increasingly serious cybercrime. Phishers create phishing websites by mimicking legitimate websites to confuse users and steal their personal information. The proliferation of phishing websites and more advanced camouflage techniques are problems faced by most existing methods. In thi…

Cited by 0SourceScholar
2021

Natural Language Processing Meets Quantum Physics: A Survey and Categorization

EMNLP 2021main

Recent research has investigated quantum NLP, designing algorithms that process natural language in quantum computers, and also quantum-inspired algorithms that improve NLP performance on classical computers. In this survey, we review representative methods at the intersection of NLP and quantum phy…

Cited by 22SourcePDFScholar
2021

Wase: Learning When to Attend for Speaker Extraction in Cocktail Party Environments

ICASSP 2021accepted

In the speaker extraction problem, it is found that additional information from the target speaker contributes to the tracking and extraction of the target speaker, which includes voiceprint, lip movement, facial expression, and spatial information. However, no one cares for the cue of sound onset,…

Cited by 0SourceScholar
2020

Encoding word order in complex embeddings

ICLR 2020spotlight

Sequential word order is important when processing text. Currently, neural networks (NNs) address this by modeling word position using position embeddings. The problem is that position embeddings capture the position of individual words, but not the ordered relationship (e.g., adjacency or precedenc…

Cited by 148SourcecodeScholar
2020

KoGuN: Accelerating Deep Reinforcement Learning via Integrating Human Suboptimal Knowledge

IJCAI 2020poster

Reinforcement learning agents usually learn from scratch, which requires a large number of interactions with the environment. This is quite different from the learning process of human. When faced with a new task, human naturally have the common sense and use the prior knowledge to derive an initial…

Cited by 0SourcePDFScholar
2020

Overcoming Language Priors with Self-supervised Learning for Visual Question Answering

IJCAI 2020poster

Most Visual Question Answering (VQA) models suffer from the language prior problem, which is caused by inherent data biases. Specifically, VQA models tend to answer questions (e.g., what color is the banana?) based on the high-frequency answers (e.g., yellow) ignoring image contents. Existing approa…

2020

Texture and Shape Biased Two-Stream Networks for Clothing Classification and Attribute Recognition

CVPR 2020poster

Clothes category classification and attribute recognition have achieved distinguished success with the development of deep learning. People have found that landmark detection plays a positive role in these tasks. However, little research is committed to analyzing these tasks from the perspective of…

Cited by 66PDFScholar
2019

A Tensorized Transformer for Language Modeling

NeurIPS 2019poster

Latest development of neural models has connected the encoder and decoder through a self-attention mechanism. In particular, Transformer, which is solely based on self-attention, has led to breakthroughs in Natural Language Processing (NLP) tasks. However, the multi-head attention mechanism, as a ke…

2019

Predicting Tongue Motion in Unlabeled Ultrasound Videos Using Convolutional Lstm Neural Networks

ICASSP 2019accepted

A challenge in speech production research is to predict future tongue movements based on a short period of past tongue movements. This study tackles speaker-dependent tongue motion prediction problem in unlabeled ultrasound videos with convolutional long short-term memory (ConvLSTM) networks. The mo…

Cited by 0SourceScholar
2018

Hierarchical Bilinear Pooling for Fine-Grained Visual Recognition

ECCV 2018poster

Fine-grained visual recognition is challenging because it highly relies on the modeling of various semantic parts and fine-grained feature learning. Bilinear pooling based models have been shown to be effective at fine-grained recognition, while most previous approaches neglect the fact that inter-l…

2016

Yin and Yang: Balancing and Answering Binary Visual Questions

CVPR 2016poster

The complex compositional structure of language makes problems at the intersection of vision and language challenging. But language also provides a strong prior that can result in good superficial performance, without the underlying models truly understanding the visual content. This can hinder pro…

Cited by 439PDFScholar