← Search

Wei Pang

19 accepted papers

2026

Do Vision and Text Cues Exhibit Evidential Coupling? UFO: A Benchmark for Compositional Multimodal Reasoning in Unified Models

ICML 2026poster

Unified Foundation Models (UFMs), which support interleaved multimodal generation and understanding, have been proposed as a promising paradigm for reasoning about dynamic world states, yet it remains unclear whether the visual content they generate functions as grounded evidence for subsequent reas…

Cited by 0SourceScholar
2026

GraphOmni: A Comprehensive and Extensible Benchmark Framework for Large Language Models on Graph-theoretic Tasks

ICLR 2026poster

This paper introduces GraphOmni, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs on graph-theoretic tasks articulated in natural language. GraphOmni spans diverse graph types, serialization formats, and prompting schemes, substantially extending upon prior efforts i…

Cited by 0SourcecodeScholar
2026

SPARD: Single-step Inference with Adaptive Sampling in Residual Diffusion for Human Motion Prediction

AAAI 2026technical

The task of stochastic human motion prediction has attracted significant attention in recent years due to its wide-ranging applications in robotics, animation, and human-computer interaction. While diffusion models have demonstrated promising progress in this domain, they remain hindered by two crit

Cited by 0SourcePDFScholar
2026

SVAgent: Storyline-guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration

CVPR 2026

Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. Although recent advances have introduced various approaches for video understanding, most existing methods still rely on

Cited by 0SourceScholar
2026

The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward

ICLR 2026poster

A central paradox in fine-tuning Large Language Models (LLMs) with Reinforcement Learning with Verifiable Reward (RLVR) is the frequent degradation of multi-attempt performance (Pass@k) despite improvements in single-attempt accuracy (Pass@1). This is often accompanied by catastrophic forgetting, wh…

Cited by 0SourceScholar
2026

When Vision Meets Graphs: A Survey on Graph Reasoning and Learning

IJCAI 2026

Graphs are a fundamental data structure underlying many problems in the natural and social sciences. Over the past decade, Graph Neural Networks (GNNs) have dominated graph machine learning, supported by solid theoretical foundations. Yet scientists often understand graph structure through vision: c

Cited by 0Scholar
2025

Asymptotically Stable Quaternion-valued Hopfield-structured Neural Network with Periodic Projection-based Supervised Learning Rules

NeurIPS 2025poster

Motivated by the geometric advantages of quaternions in representing rotations and postures, we propose a quaternion-valued supervised learning Hopfield-structured neural network (QSHNN) with a fully connected structure inspired by the classic Hopfield neural network (HNN). Starting from a continuou…

Cited by 0SourceScholar
2025

Boosting Short Text Classification with Multi-Source Information Exploration and Dual-Level Contrastive Learning

AAAI 2025technical

Short text classification, as a research subtopic in natural language processing, is more challenging due to its semantic sparsity and insufficient labeled samples in practical scenarios. We propose a novel model named MI-DELIGHT for short text classification in this work. Specifically, it first per…

2025

DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models

NAACL 2025long

Recent progress in video-text retrieval has been driven largely by advancements in model architectures and training strategies. However, the representation learning capabilities of video-text retrieval models remain constrained by low-quality and limited training data annotations. To address this is…

Cited by 0SourcePDFScholar
2025

Graph Few-Shot Learning via Adaptive Spectrum Experts and Cross-Set Distribution Calibration

NeurIPS 2025poster

Graph few-shot learning has attracted increasing attention due to its ability to rapidly adapt models to new tasks with only limited labeled nodes. Despite the remarkable progress made by existing graph few-shot learning methods, several key limitations remain. First, most current approaches rely on…

Cited by 0SourceScholar
2025

MERMAID: Multi-perspective Self-reflective Agents with Generative Augmentation for Emotion Recognition

EMNLP 2025

Multimodal large language models (MLLMs) have demonstrated strong performance across diverse multimodal tasks, achieving promising outcomes. However, their application to emotion recognition in natural images remains underexplored. MLLMs struggle to handle ambiguous emotional expressions and implici

Cited by 0SourcePDFScholar
2025

Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers

NeurIPS 2025poster

Academic poster generation is a crucial yet challenging task in scientific communication, requiring the compression of long-context interleaved documents into a single, visually coherent page. To address this challenge, we introduce Paper2Poster, the first benchmark and metric suite for poster gene…

Cited by 0SourcecodeScholar
2025

The Underappreciated Power of Vision Models for Graph Structural Understanding

NeurIPS 2025poster

Graph Neural Networks operate through bottom-up message-passing, fundamentally differing from human visual perception, which intuitively captures global structures first. We investigate the underappreciated potential of vision models for graph understanding, finding they achieve performance comparab…

Cited by 0SourceScholar
2024

ClavaDDPM: Multi-relational Data Synthesis with Cluster-guided Diffusion Models

NeurIPS 2024poster

Recent research in tabular data synthesis has focused on single tables, whereas real-world applications often involve complex data with tens or hundreds of interconnected tables. Previous approaches to synthesizing multi-relational (multi-table) data fall short in two key aspects: scalability for la…

Cited by 8SourcePDFScholar
2024

Phased Instruction Fine-Tuning for Large Language Models

ACL 2024findings

Instruction Fine-Tuning, a method enhancing pre-trained language models’ capabilities from mere next-word prediction to complex instruction following, often employs a one-off training approach on diverse instruction dataset. However, this method may not effectively enhance models’ adherence to instr…

2023

WaveForM: Graph Enhanced Wavelet Learning for Long Sequence Forecasting of Multivariate Time Series

AAAI 2023technical

Multivariate time series (MTS) analysis and forecasting are crucial in many real-world applications, such as smart traffic management and weather forecasting. However, most existing work either focuses on short sequence forecasting or makes predictions predominantly with time domain features, which…

2020

Acoustofluidic Tweezers for the 3D Manipulation of Microparticles

ICRA 2020poster

Non-contact manipulation is of great importance in the actuation of micro-robotics. It is challenging to contactless manipulate micro-scale objects over large spatial distance in fluid. Here, we describe a novel approach for the dynamic position control of microparticles in three-dimensional (3D) sp…

Cited by 10SourceScholar