← Search

Bin Wu

31 accepted papers

2026

CTX-Coder: Cross-Attention Architectures Empower LLMs for Long-Context Vulnerability Detection

AAAI 2026technical

Software vulnerabilities have increased sharply, underscoring the growing urgency for effective detection methods. Although large language model (LLM) based methods have shown promise in this task, current state-of-the-art LLM approaches struggle with functions that have long contexts. In this pape

Cited by 1SourcePDFScholar
2026

CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style Transfer

ICLR 2026poster

Transferring visual style between images while preserving semantic correspondence between similar objects remains a central challenge in computer vision. While existing methods have made great strides, most of them operate at global level but overlook region-wise and even pixel-wise semantic corresp…

Cited by 0SourceScholar
2026

From Interaction Trajectories to Prompt Rules: Credit Assignment for Multi-Agent Prompt Optimization

ICML 2026poster

Large language model (LLM)-based multi-agent systems commonly rely on natural-language prompts to specify agent behavior, yet optimizing these prompts remains challenging when agent roles and interaction structures are fixed by design. In such systems, behaviors emerge over long, noisy interaction t…

Cited by 0SourceScholar
2026

InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem

ICML 2026poster

The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation. The fundamental nature of scientific evaluation needs knowledgeable grounding, collective deliberation, and multi-criteri…

Cited by 0SourceScholar
2026

SAFE: Semantic- and Frequency-Enhanced Curriculum for Cross-Domain Deepfake Detection

AAAI 2026technical

Driven by advances in GANs and diffusion models, deepfake content has reached an unprecedented level of photorealism, causing detectors to deteriorate once they leave their training domain. Most prior studies adopt CLIP as the backbone of an image-level binary classifier, yet overlook CLIP’s core st

Cited by 0SourcePDFScholar
2026

TGCA-LLM: Time-Aware Graph-Text Contrastive Alignment for Enhancing LLMs in Temporal Knowledge Graph Completion

AAAI 2026technical

Temporal Knowledge Graph Completion (TKGC) aims to infer missing facts by modeling historical events and latent temporal dependencies in Temporal Knowledge Graphs (TKGs). Recently, TKGC methods that integrate graph embeddings into Large Language Models (LLMs) have shown great promise by leveraging t

Cited by 0SourcePDFScholar
2025

A Joint Optimization Framework for Enhancing Efficiency of Tool Utilization in LLM Agents

ACL 2025finding

Large Language Models (LLMs) augmented with external tools have demonstrated remarkable capabilities in complex problem solving. Existing efforts for tool utilization typically involve an LLM agent that contains instructions on using the description of the available tools to determine and call the t…

2025

Boosting LLM’s Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning

ACL 2025long

Molecular structure elucidation involves deducing a molecule’s structure from various types of spectral data, which is crucial in chemical experimental analysis. While large language models (LLMs) have shown remarkable proficiency in analyzing and reasoning through complex tasks, they still encounte…

Cited by 0SourcePDFScholar
2025

Causality-Guided Context-Aware Multimodal Public Speaking Anxiety Detection for Out-of-Distribution Generalization

ICASSP 2025accepted

Public Speaking Anxiety Detection (PSAD) is a complex and challenging task that involves detecting anxiety through diverse multimodal cues. While deep neural networks have achieved remarkable success in this task, their performance tends to degrade significantly under distribution shifts, especially…

Cited by 0SourceScholar
2025

Entropy-Based Decoding for Retrieval-Augmented Large Language Models

NAACL 2025long

Augmenting Large Language Models (LLMs) with retrieved external knowledge has proven effective in improving the factual accuracy of generated responses. Despite their success, retrieval-augmented LLMs still face the distractibility issue, where the generated responses are negatively influenced by no…

Cited by 2SourcePDFScholar
2025

Fuel-Optimal Operational Speed Planning for Autonomous Trucking on Highways

ICRA 2025

The rapid advancement of autonomous driving technology, particularly in autonomous trucking on highways, shows great value for enhancing efficiency and reducing costs in the logistics industry. In this work, we define the full-trip speed planning problem for autonomous trucks under delivery time and

Cited by 0SourceScholar
2025

Interesting Culture: Social Relation Recognition from Videos via Culture De-confounding

EMNLP 2025

Social relationship recognition, as one of the fundamental tasks in video understanding, contributes to the construction and application of multi-modal knowledge graph. Previous works have mainly focused on two aspects: generating character graphs and multi-modal fusion. However, they often overlook

Cited by 0SourcePDFScholar
2025

Learning Causally Disentangled Representations for Fair Personality Detection

IJCAI 2025

Personality detection aims to identify the personality traits implied in social posts. Existing methods mainly focus on learning the mapping between user-generated posts and personality trait labels but inevitably suffer from potential harm caused by individual bias, as these posts are written by au

Cited by 0SourcePDFScholar
2025

Synergizing Multimodal Temporal Knowledge Graphs and Large Language Models for Social Relation Recognition

EMNLP 2025

Recent years have witnessed remarkable advances in Large Language Models (LLMs). However, in the task of social relation recognition, Large Language Models (LLMs) encounter significant challenges due to their reliance on sequential training data, which inherently restricts their capacity to effectiv

2025

Synthetic Data is an Elegant GIFT for Continual Vision-Language Models

CVPR 2025poster

Pre-trained Vision-Language Models (VLMs) require Continual Learning (CL) to efficiently update their knowledge and adapt to various downstream tasks without retraining from scratch. However, for VLMs, in addition to the loss of knowledge previously learned from downstream tasks, pre-training knowle…

2025

TEACH: A Contrastive Knowledge Adaptive Distillation Framework for Classical Chinese Understanding

ACL 2025long

Traditional methods for processing classical Chinese typically segment language understanding into discrete tasks, which overlook crucial background information and reduce user engagement. Large language models (LLMs) provide integrated solutions, yet they entail high computational costs and risks o…

2025

Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework

NeurIPS 2025poster

Graph Transformers (GTs) have emerged as a powerful paradigm for graph representation learning due to their ability to model diverse node interactions. However, existing GTs often rely on intricate architectural designs tailored to specific interactions, limiting their flexibly. To address this, we…

Cited by 0SourceScholar
2025

Video-Poetry Retrieval with Multimodal Knowledge Graph Guided Unsupervised Pre-training

ICASSP 2025accepted

Classical Chinese poetry, with its rich cultural heritage, holds immense artistic value. Recently, research on multi-modal approaches to classical poetry has garnered attention. Current research mainly focuses on poetry and images, but images cannot fully capture dynamic scenes like sunrises. Howeve…

Cited by 0SourceScholar
2024

AC-EVAL: Evaluating Ancient Chinese Language Understanding in Large Language Models

EMNLP 2024finding

Given the importance of ancient Chinese in capturing the essence of rich historical and cultural heritage, the rapid advancements in Large Language Models (LLMs) necessitate benchmarks that can effectively evaluate their understanding of ancient contexts. To meet this need, we present AC-EVAL, an in…

2024

Data Augmented Graph Neural Networks for Personality Detection

AAAI 2024technical

Personality detection is a fundamental task for user psychology research. One of the biggest challenges in personality detection lies in the quantitative limitation of labeled data collected by completing the personality questionnaire, which is very time-consuming and labor-intensive. Most of the ex…

Cited by 7SourcePDFScholar
2024

Exploring Question Guidance and Answer Calibration for Visually Grounded Video Question Answering

EMNLP 2024finding

Video Question Answering (VideoQA) tasks require not only correct answers but also visual evidence. The “localize-then-answer” strategy, while enhancing accuracy and interpretability, faces challenges due to the lack of temporal localization labels in VideoQA datasets. Existing methods often train t…

Cited by 0SourcePDFScholar
2024

Instruction Tuning With Loss Over Instructions

NeurIPS 2024poster

Instruction tuning plays a crucial role in shaping the outputs of language models (LMs) to desired styles. In this work, we propose a simple yet effective method, Instruction Modelling (IM), which trains LMs by applying a loss function to the instruction and prompt part rather than solely to the out…

2023

Adaptive Compositional Continual Meta-Learning

ICML 2023poster

This paper focuses on continual meta-learning, where few-shot tasks are heterogeneous and sequentially available. Recent works use a mixture model for meta-knowledge to deal with the heterogeneity. However, these methods suffer from parameter inefficiency caused by two reasons: (1) the underlying as…

Cited by 15SourcePDFScholar
2023

Graph Sampling-based Meta-Learning for Molecular Property Prediction

IJCAI 2023poster

Molecular property is usually observed with a limited number of samples, and researchers have considered property prediction as a few-shot problem. One important fact that has been ignored by prior works is that each molecule can be recorded with several different properties simultaneously. To effec…

2022

A Multi-Modal Knowledge Graph for Classical Chinese Poetry

EMNLP 2022finding

Classical Chinese poetry has a long history and is a precious cultural heritage of humankind. Displaying the classical Chinese poetry in a visual way, helps to cross cultural barriers in different countries, making it enjoyable for all the people. In this paper, we construct a multi-modal knowledge…

2022

Contrastive Graph Transformer Network for Personality Detection

IJCAI 2022poster

Personality detection is to identify the personality traits underlying social media posts. Most of the existing work is mainly devoted to learning the representations of posts based on labeled data. Yet the ground-truth personality traits are collected through time-consuming questionnaires. Thus, on…

2021

Litesing: Towards Fast, Lightweight and Expressive Singing Voice Synthesis

ICASSP 2021accepted

LiteSing proposed in this paper is a high-quality singing voice synthesis (SVS) system, which is fast, lightweight and expressive. This model mainly stacks several non-autoregressive WaveNet blocks in the encoder and decoder under a generative adversarial architecture, predicts full conditions from…

Cited by 0SourceScholar
2019

Low-cost Measurement of Industrial Shock Signals via Deep Learning Calibration

ICASSP 2019accepted

Special high-end sensors with expensive hardware are usually needed to measure shock signals with high accuracy. In this paper, we show that cheap low-end sensors calibrated by deep neural networks are also capable to measure high-g shocks accurately. Firstly we perform drop shock tests to collect a…

Cited by 0SourceScholar