← Search

Wenjie Zhang

30 accepted papers

2026

BOLT: Decision‑Aligned Distillation and Budget-Aware Routing for Constrained Multimodal QA on Robots

ICLR 2026poster

Robotic systems can require multimodal reasoning under stringent constraints of latency, memory, and energy. Standard instruction tuning and token-level distillation fail to deliver decision quality, reliability, and interpretability under these constraints. We introduce BOLT, a decision-aligned dis…

Cited by 0SourceScholar
2026

DHG-Bench: A Comprehensive Benchmark for Deep Hypergraph Learning

ICLR 2026poster

Deep graph models have achieved great success in network representation learning. However, their focus on pairwise relationships restricts their ability to learn pervasive higher-order interactions in real-world systems, which can be naturally modeled as hypergraphs. To tackle this issue, Hypergraph…

Cited by 0SourcecodeScholar
2026

EntRAG: Entity-Centric Retrieval-Augmented Generation for Knowledge-based Visual Question Answering

ICML 2026poster

Knowledge-based Visual Question Answering (KB-VQA) remains a challenging task, particularly when queries require precise identification and grounding of fine-grained entities within large-scale knowledge base. Existing methods often treat visual and textual signals in isolation and rely heavily on i…

Cited by 0SourceScholar
2026

Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation Learning

ICML 2026poster

Vector Quantization (VQ) has recently emerged as a promising approach for learning discrete representations of graph-structured data. However, a fundamental challenge, i.e., codebook collapse, remains underexplored in the graph domain, significantly limiting the expressiveness and generalization of …

Cited by 0SourceScholar
2026

Graph of States: Solving Abductive Tasks with Large Language Models

ICML 2026poster

Logical reasoning encompasses deduction, induction, and abduction. However, while Large Language Models (LLMs) have effectively mastered the former two, abductive reasoning remains significantly underexplored. Existing frameworks, predominantly designed for static deductive tasks, fail to generalize…

Cited by 0SourceScholar
2026

SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement Learning

ICML 2026poster

Actor-critic (AC) methods are a cornerstone of reinforcement learning (RL) but offer limited interpretability. Current explainable RL methods seldom use *state attributions* to assist training. Rather, they treat all state features equally, thereby neglecting the heterogeneous impacts of individual …

Cited by 0SourceScholar
2026

SMAP: Semantic Route Planning with Map-Grounded Multimodal Alignment

CVPR 2026

Semantic route planning involves generating itineraries that align with user intent while respecting real-world spatial constraints. However, text-only large language models (LLMs) often hallucinate geographically implausible routes due to poor spatial grounding. Inspired by how humans use maps for

Cited by 0SourcecodeScholar
2026

Towards Generative Graph Matching for Graph Edit Distance Computation

ICML 2026poster

Graph Edit Distance (GED), which aims to find an edit path with minimum number of edit operations to transform one graph into another, is a fundamental NP-hard problem and a widely used graph similarity measure. Recent matching-based hybrid approaches have demonstrated better scalability than A* sea…

Cited by 0SourceScholar
2026

Unlocking Multi-Modal Potentials for Link Prediction on Dynamic Text-Attributed Graphs

AAAI 2026technical

Dynamic Text-Attributed Graphs (DyTAGs) are a novel graph paradigm that captures evolving temporal events (edges) alongside rich textual attributes. Existing studies can be broadly categorized into TGNN-driven and LLM-driven approaches, both of which encode textual attributes and temporal structures

Cited by 0SourcePDFScholar
2025

EchoDiffusion: Waveform Conditioned Diffusion Models for Echo-Based Depth Estimation

AAAI 2025technical

To extract spatial information, depth estimation using conventional echo-based methods typically employs models with encoder-decoder architectures, such as UNet. However, these methods may face challenges in extracting fine details from echo waveforms and handling multi-scale feature extraction with…

2025

HydraRAG: Structured Cross-Source Enhanced Large Language Model Reasoning

EMNLP 2025

Retrieval-augmented generation (RAG) enhances large language models (LLMs) by incorporating external knowledge. Current hybrid RAG system retrieves evidence from both knowledge graphs (KGs) and text documents to support LLM reasoning. However, it faces challenges like handling multi-hop reasoning, m

2025

MMAPG: A Training-Free Framework for Multimodal Multi-hop Question Answering via Adaptive Planning Graphs

EMNLP 2025

Multimodal Multi-hop question answering requires integrating information from diverse sources, such as images and texts, to derive answers. Existing methods typically rely on sequential retrieval and reasoning, where each step builds on the previous output. However, this single-path paradigm makes t

Cited by 0SourcePDFScholar
2025

Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning

ICRA 2025

This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the l

Cited by 1SourcecodeScholar
2025

R-DTI: Drug Target Interaction Prediction Based on Second-Order Relevance Exploration

AAAI 2025technical

Drug Target Interaction (DTI) prediction has witnessed promising performance boosts accompanied by advanced multimodal feature extraction. However, existing approaches suffer from two main difficulties. First, the complex protein structures cannot be well represented by current protein-sequence-base…

2025

SML: A Backdoor Defense for Non-Intrusive Speech Quality Assessment via Semi-Supervised and Multi-Task Learning

ICASSP 2025accepted

Non-intrusive speech quality assessment (NISQA) is widely used in speech downstream tasks due to its ability to predict the quality of speech without a reference speech. However, few researchers have focused on the backdoor security of NISQA. Despite the backdoor defenses have been extensively studi…

Cited by 0SourceScholar
2025

Sample-Efficient Tabular Self-Play for Offline Robust Reinforcement Learning

NeurIPS 2025poster

Multi-agent reinforcement learning (MARL), as a thriving field, explores how multiple agents independently make decisions in a shared dynamic environment. Due to environmental uncertainties, policies in MARL must remain robust to tackle the sim-to-real gap. We focus on robust two-player zero-sum Mar…

Cited by 0SourceScholar
2025

Speed Master: Quick or Slow Play to Attack Speaker Recognition

AAAI 2025technical

Backdoor attacks pose a significant threat during the model's training phase. Attackers craft pre-defined triggers to break deep neural networks, ensuring the model accurately classifies clean samples during inference yet erroneously classifies samples added with these triggers. Recent studies have…

Cited by 0SourcePDFScholar
2025

Towards Unsupervised Training of Matching-based Graph Edit Distance Solver via Preference-aware GAN

NeurIPS 2025poster

Graph Edit Distance (GED) is a fundamental graph similarity metric widely used in various applications. However, computing GED is an NP-hard problem. Recent state-of-the-art hybrid GED solver has shown promising performance by formulating GED as a bipartite graph matching problem, then leveraging a…

Cited by 0SourceScholar
2024

A Novel Quasi-Passive Non-Anthropomorphic Lower Limb Exoskeleton for Load-Bearing

RA-L 2024

Load-bearing walking is an important requirement for defense and military, logistics and transportation. In this letter, a novel quasi-passive non-anthropomorphic lower limb exoskeleton is proposed to augment human load-bearing capacities. The exoskeleton employs a rigid leg rod that transmits loads

Cited by 3SourceScholar
2024

Construction and Application of Materials Knowledge Graph in Multidisciplinary Materials Science via Large Language Model

NeurIPS 2024poster

Knowledge in materials science is widely dispersed across extensive scientific literature, posing significant challenges for efficient discovery and integration of new materials. Traditional methods, often reliant on costly and time-consuming experimental approaches, further complicate rapid innovat…

Cited by 4SourcePDFScholar
2024

Hypergraph Self-supervised Learning with Sampling-efficient Signals

IJCAI 2024poster

Self-supervised learning (SSL) provides a promising alternative for representation learning on hypergraphs without costly labels. However, existing hypergraph SSL models are mostly based on contrastive methods with the instance-level discrimination strategy, suffering from two significant limitation…

2024

QUAR-VLA: Vision-Language-Action Model for Quadruped Robots

ECCV 2024poster

"The important manifestation of robot intelligence is the ability to naturally interact and autonomously make decisions. Traditional quadruped robot learning typically handles language interaction and visual autonomous perception separately, which, while simplifying system design, also limits the sy…

Cited by 19SourcePDFScholar
2022

Generalized Equivariance and Preferential Labeling for GNN Node Classification

AAAI 2022technical

Existing graph neural networks (GNNs) largely rely on node embeddings, which represent a node as a vector by its identity, type, or content. However, graphs with unattributed nodes widely exist in real-world applications (e.g., anonymized social networks). Previous GNNs either assign random labels t…

2022

Lyra: A Benchmark for Turducken-Style Code Generation

IJCAI 2022poster

Recently, neural techniques have been used to generate source code automatically. While promising for declarative languages, these approaches achieve much poorer performance on datasets for imperative languages. Since a declarative language is typically embedded in an imperative language (i.e., the…

2020

NLocalSAT: Boosting Local Search with Solution Prediction

IJCAI 2020poster

The Boolean satisfiability problem (SAT) is a famous NP-complete problem in computer science. An effective way for solving a satisfiable SAT problem is the stochastic local search (SLS). However, in this method, the initialization is assigned in a random manner, which impacts the effectiveness of SL…

2020

Residual Feature Aggregation Network for Image Super-Resolution

CVPR 2020poster

Recently, very deep convolutional neural networks (CNNs) have shown great power in single image super-resolution (SISR) and achieved significant improvements against traditional methods. Among these CNN-based methods, the residual connections play a critical role in boosting the network performance.…

Cited by 645PDFScholar
2018

TextSnake: A Flexible Representation for Detecting Text of Arbitrary Shapes

ECCV 2018poster

Driven by deep neural networks and large scale datasets, scene text detection methods have progressed substantially over the past years, continuously refreshing the performance records on various standard benchmarks. However, limited by the representations (axis-aligned rectangles, rotated rectangle…

Cited by 707SourcePDFScholar