← Search

Bo Wu

36 accepted papers

2026

ClimateAR: Multi-Scale Autoregressive Generative Modeling for Seasonal-to-Interannual Climate Forecasting

ICML 2026poster

Accurate seasonal‑to‑interannual climate forecasting provides critical support for decision-making in agriculture, energy, and disaster preparedness. Current deterministic models often fail to capture climate uncertainty, while existing generative approaches oversimplify the system by neglecting key…

Cited by 0SourceScholar
2026

From Semantics to Spectrum: A New Lens on Graph Augmentation Strategy

AAAI 2026technical

Graph augmentation is a cornerstone of effective graph contrastive learning, yet existing methods often rely on random designed perturbations, which may distort latent semantics and impair representation quality. In this work, we argue that semantic consistency can be effectively approximated by low

Cited by 0SourcePDFScholar
2026

Read the Room: Video Social Reasoning with Mental-Physical Causal Chains

ICLR 2026poster

``Read the room,'' or the ability to infer others' mental states from subtle social cues, is a hallmark of human social intelligence but remains a major challenge for current AI systems. Existing social reasoning datasets are limited in complexity, scale, and coverage of mental states, falling short…

Cited by 0SourcecodeScholar
2025

Integrating Large Language Models and Möbius Group Transformations for Temporal Knowledge Graph Embedding on the Riemann Sphere

AAAI 2025technical

The significance of Temporal Knowledge Graphs (TKGs) in Artificial Intelligence (AI) lies in their capacity to incorporate time-dimensional information, support complex reasoning and prediction, optimize decision-making processes, enhance the accuracy of recommendation systems, promote multimodal da…

Cited by 0SourcePDFScholar
2025

MalImgDA: Diffusion-based Data Augmentation for Long-tailed Malware Family Classification

ICASSP 2025accepted

With the rapid improvement of machine learning technology, leveraging machine learning methods for malware classification has emerged as a viable approach. However, under real-world circumstance, the imbalanced or long-tailed distribution among various malware families, poses a critical challenge to…

Cited by 0SourceScholar
2025

Retrieval-Augmented Multilingual Citation Generation

ICASSP 2025accepted

Retrieval-augmented citation generation (RACG) helps users trust the large language model output by retrieving evidence from reliable sources. However, most current RACG research focuses on single-language tasks, particularly in English, and overlooks the need for cross-lingual evidence retrieval an…

Cited by 0SourceScholar
2025

SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation

CVPR 2025poster

Vision-language temporal alignment is a crucial capability for human dynamic recognition and cognition in real-world scenarios. While existing research focuses on capturing vision-language relevance, it faces limitations due to biased temporal distributions, imprecise annotations, and insufficient c…

Cited by 0SourcePDFScholar
2024

Hypergraph-Based Session Modeling: A Multi-Collaborative Self-Supervised Approach for Enhanced Recommender Systems

COLING 2024main

Session-based recommendation (SBR) is a challenging task that involves predicting a user’s next item click based on their recent session history. Presently, many state-of-the-art methodologies employ graph neural networks to model item transitions. Notwithstanding their impressive performance, graph…

Cited by 4SourcePDFScholar
2024

Improving Robustness of GNN-based Anomaly Detection by Graph Adversarial Training

COLING 2024main

Graph neural networks (GNNs) play a fundamental role in anomaly detection, excelling at the identification of node anomalies by aggregating information from neighboring nodes. Nonetheless, they exhibit vulnerability to attacks, with even minor alterations in the graph structure or node attributes re…

Cited by 7SourcePDFScholar
2024

SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge

CVPR 2024poster

Reasoning from visual dynamics scenes has many real world applications. However existing video reasoning benchmarks are still inadequate since they were mainly designed for factual or situated reasoning and rarely involve broader knowledge in the real world. Our work aims to delve deeper into reason…

Cited by 12SourcePDFScholar
2024

Selective Prompting Tuning for Personalized Conversations with LLMs

ACL 2024findings

In conversational AI, personalizing dialogues with persona profiles and contextual understanding is essential. Despite large language models’ (LLMs) improved response coherence, effective persona integration remains a challenge. In this work, we first study two common approaches for personalizing LL…

2024

Uncertainty-Aware Deployment of Pre-trained Language-Conditioned Imitation Learning Policies

IROS 2024poster

Large-scale robotic policies trained on data from diverse tasks and robotic platforms hold great promise for enabling general-purpose robots; however, reliable generalization to new environment conditions remains a major challenge. Toward addressing this challenge, we propose a novel approach for un…

Cited by 1SourcecodeScholar
2023

Cross-Modal Matching and Adaptive Graph Attention Network for RGB-D Scene Recognition

ICASSP 2023accepted

Despite the significant advances in RGB-D scene recognition, there are several major limitations that need further investigation. For example, simply extracting modal-specific features neglects the complex relationships among multiple modalities of features. Moreover, cross-modal features have not b…

Cited by 0SourceScholar
2023

Distance-Based Propagation for Efficient Knowledge Graph Reasoning

EMNLP 2023long main

Knowledge graph completion (KGC) aims to predict unseen edges in knowledge graphs (KGs), resulting in the discovery of new facts. A new class of methods have been proposed to tackle this problem by aggregating path information. These methods have shown tremendous ability in the task of KGC. However…

Cited by 0SourcecodeScholar
2023

Enhancing Dynamic GCN for Node Attribute Forecasting with Meta Spatial-Temporal Learning (Student Abstract)

AAAI 2023technical

Node attribute forecasting has recently attracted considerable attention. Recent attempts have thus far utilize dynamic graph convolutional network (GCN) to predict future node attributes. However, few prior works have notice that the complex spatial and temporal interaction between nodes, which wil…

Cited by 0SourcePDFScholar
2023

Exploiting High-Order Interaction Relations to Explore User Intent (Student Abstract)

AAAI 2023technical

This paper studies the problem of exploring the user intent for session-based recommendations. Its challenges come from the uncertainty of user behavior and limited information. However, current endeavors cannot fully explore the mutual interactions among sessions and do not explicitly model the com…

Cited by 1SourcePDFScholar
2023

Intent Does Matter! Propagating High-Order Relations for Exploring Interest Preferences

ICASSP 2023accepted

Session-based recommendation (SBR) aims to predict the user’s action at the next timestamp according to an anonymous yet short interaction sequence (i.e., session). Almost all the existing SBR solutions for user preference are only based on the current session without exploiting the high-order relat…

Cited by 0SourceScholar
2023

Learning Situation Hyper-Graphs for Video Question Answering

CVPR 2023poster

Answering questions about complex situations in videos requires not only capturing of the presence of actors, objects, and their relations, but also the evolution of these relationships over time. A situation hyper-graph is a representation that describes situations as scene sub-graphs for video fra…

2023

Learning from Children: Improving Image-Caption Pretraining via Curriculum

ACL 2023findings

Image-caption pretraining has been quite successfully used for downstream vision tasks like zero-shot image classification and object detection. However, image-caption pretraining is still a hard problem – it requires multiple concepts (nouns) from captions to be aligned to several objects in images…

2023

Personalized Dialogue Generation with Persona-Adaptive Attention

AAAI 2023technical

Persona-based dialogue systems aim to generate consistent responses based on historical context and predefined persona. Unlike conventional dialogue generation, the persona-based dialogue needs to consider both dialogue context and persona, posing a challenge for coherent training. Specifically, thi…

2023

Select The Best: Enhancing Graph Representation with Adaptive Negative Sample Selection

ICASSP 2023accepted

Graph contrastive learning (GCL) has emerged as a powerful tool to address real-world widespread label scarcity problems and has achieved impressive success in the graph learning domain. Albeit their remarkable performance, most current works mainly focus on designing sample augmentation methods, wh…

Cited by 0SourceScholar
2022

Eureka: Neural Insight Learning for Knowledge Graph Reasoning

COLING 2022main

The human recognition system has presented the remarkable ability to effortlessly learn novel knowledge from only a few trigger events based on prior knowledge, which is called insight learning. Mimicking such behavior on Knowledge Graph Reasoning (KGR) is an interesting and challenging research pro…

Cited by 0SourcePDFScholar
2022

Improving Dynamic Graph Convolutional Network with Fine-Grained Attention Mechanism

ICASSP 2022accepted

Graph convolutional network (GCN) is a novel framework that utilizes a pre-defined Laplacian matrix to learn graph data effectively. With its powerful nonlinear fitting ability, GCN can produce high-quality node embedding. However, generalized GCN can only handle static graphs, whereas a large numbe…

Cited by 0SourceScholar
2021

A Joint Training Framework of Multi-Look Separator and Speaker Embedding Extractor for Overlapped Speech

ICASSP 2021accepted

In multi-talker cases, overlapped speech degrades the speaker verification (SV) performance dramatically. To tackle this challenging problem, speech separation with multi-channel techniques can be adopted to extract each speaker’s signals to improve the SV performance. In this paper, a joint trainin…

Cited by 0SourceScholar
2021

Advice-Guided Reinforcement Learning in a non-Markovian Environment

AAAI 2021technical

We study a class of reinforcement learning tasks in which the agent receives its reward for complex, temporally-extended behaviors sparsely. For such tasks, the problem is how to augment the state-space so as to make the reward function Markovian in an efficient way. While some existing solutions as…

Cited by 47SourcePDFScholar
2021

Decentralized Classification with Assume-Guarantee Planning

IROS 2021poster

We study the problem of decentralized classification conducted over a network of mobile sensors. We model the multiagent classification task as a hypothesis testing problem where each sensor has to almost surely find the true hypothesis from a finite set of candidate hypotheses. Each sensor makes no…

Cited by 0SourceScholar
2021

NxMTransformer: Semi-Structured Sparsification for Natural Language Understanding via ADMM

NeurIPS 2021poster

Natural Language Processing (NLP) has recently achieved great success by using huge pre-trained Transformer networks. However, these models often contain hundreds of millions or even billions of parameters, bringing challenges to online deployment due to latency constraints. Recently, hardware manuf…

Cited by 21SourcePDFScholar
2021

STAR: A Benchmark for Situated Reasoning in Real-World Videos

NeurIPS 2021poster

Reasoning in the real world is not divorced from situations. How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence. This paper introduces a new benchmark that evaluates the situated reasoning ability via…

Cited by 195SourceScholar
2021

Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning Tasks

AAAI 2021technical

In continuing tasks, average-reward reinforcement learning may be a more appropriate problem formulation than the more common discounted reward formulation. As usual, learning an optimal policy in this setting typically requires a large amount of training experiences. Reward shaping is a common appr…

Cited by 64SourcePDFScholar
2020

Audio-Visual Recognition of Overlapped Speech for the LRS2 Dataset

ICASSP 2020accepted

Automatic recognition of overlapped speech remains a highly challenging task to date. Motivated by the bimodal nature of human speech perception, this paper investigates the use of audio-visual technologies for overlapped speech recognition. Three issues associated with the construction of audio-vis…

Cited by 82SourceScholar
2020

Enhancing Neural Models with Vulnerability via Adversarial Attack

COLING 2020main

Natural Language Sentence Matching (NLSM) serves as the core of many natural language processing tasks. 1) Most previous work develops a single specific neural model for NLSM tasks. 2) There is no previous work considering adversarial attack to improve the performance of NLSM tasks. 3) Adversarial a…

2020

Improving Reverberant Speech Training Using Diffuse Acoustic Simulation

ICASSP 2020accepted

We present an efficient and realistic geometric acoustic simulation approach for generating and augmenting training data in speech-related machine learning tasks. Our physically-based acoustic simulation method is capable of modeling occlusion, specular and diffuse reflections of sound in complicate…

Cited by 0SourceScholar
2020

Learning the Compositional Visual Coherence for Complementary Recommendations

IJCAI 2020poster

Complementary recommendations, which aim at providing users product suggestions that are supplementary and compatible with their obtained items, have become a hot topic in both academia and industry in recent years. Existing work mainly focused on modeling the co-purchased relations between two item…

Cited by 0SourcePDFScholar