← Search

Ping Wei

29 accepted papers

2026

Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification

CVPR 2026

As the harm caused by fake news grows, the task of detecting and grounding multi-modal media manipulation (DGM4) is gaining more attention. Existing multimodal methods overlook fine-grained semantic alignment between visual and textual modalities, thereby limiting their ability to detect sophisticat

Cited by 0SourceScholar
2026

Read the Room: Video Social Reasoning with Mental-Physical Causal Chains

ICLR 2026poster

``Read the room,'' or the ability to infer others' mental states from subtle social cues, is a hallmark of human social intelligence but remains a major challenge for current AI systems. Existing social reasoning datasets are limited in complexity, scale, and coverage of mental states, falling short…

Cited by 0SourcecodeScholar
2025

Alchemy: Amplifying Theorem-Proving Capability Through Symbolic Mutation

ICLR 2025poster

Formal proofs are challenging to write even for experienced experts. Recent progress in Neural Theorem Proving (NTP) shows promise in expediting this process. However, the formal corpora available on the Internet are limited compared to the general text, posing a significant data scarcity challenge…

2025

Beyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and Harmonization

CVPR 2025poster

Anomaly detection is a significant task for its application and research value. While existing methods have made impressive progress within the same modality, cross-modal anomaly detection remains an open and challenging problem. In this paper, we propose a cross-modal anomaly detection model that i…

2025

CFDONEval: A Comprehensive Evaluation of Operator-Learning Neural Network Models for Computational Fluid Dynamics

IJCAI 2025

In this paper, we introduce CFDONEval, a comprehensive evaluation of 12 operator-learning-based neural network (ON) models to simulate 7 benchmark fluid dynamics problems. These problems cover a range of 2D scenarios, including Darcy flow, two-phase flow, Taylor-Green vortex, lid-driven cavity flow,

2025

Decentralized but Not Compromised: Modular Architecture with Refined Observation for Multi-Agent Model-Based Reinforcement Learning

IROS 2025

Multi-agent adversarial tasks such as swarm robotics and autonomous vehicle coordination, demand efficient decentralized collaboration under partial observability. While model-free multi-agent RL (MF-MARL) methods suffer from necessitating extensive environment interactions, most existing multi-agen

Cited by 0SourceScholar
2025

I2-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting

ICCV 2025poster

Forecasting the evolution of 3D scenes and generating unseen scenarios through occupancy-based world models offers substantial potential to enhance the safety of autonomous driving systems. While tokenization has revolutionized image and video generation, efficiently tokenizing complex 3D scenes rem…

2025

OTIAS: OcTree Implicit Adaptive Sampling for Multispectral and Hyperspectral Image Fusion

AAAI 2025technical

Implicit Neural Representation (INR) methods have demonstrated great potential in arbitrary-scale super-resolution tasks. This success is primarily due to their ability to continuously represent images using coordinates. In the task of remote sensing image fusion, INR methods have also shown promisi…

2025

STCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Prediction

CVPR 2025poster

3D occupancy and scene flow offer a detailed and dynamic representation of 3D scene. Recognizing the sparsity and complexity of 3D space, previous vision-centric methods have employed implicit learning-based approaches to model spatial and temporal information. However, these approaches struggle to…

2025

Stochastic-Aware Mamba Diffusion for Pedestrian Trajectory Prediction

ICASSP 2025accepted

Pedestrian trajectory prediction plays a crucial role in understanding human behavior and intentions. Due to the inherent randomness in human movement, current research constructs trajectories in stochastic space and uses diffusion models to reverse the denoising process. The commonly used denoising…

Cited by 0SourceScholar
2025

TOTP: Transferable Online Pedestrian Trajectory Prediction with Temporal-Adaptive Mamba Latent Diffusion

ICCV 2025poster

Pedestrian trajectory prediction is crucial for many intelligent tasks. While existing methods predict future trajectories from fixed-frame historical observations, they are limited by the observational perspective and the need for extensive historical information, resulting in prediction delays and…

Cited by 0SourcePDFScholar
2025

Unveiling Multi-View Anomaly Detection: Intra-view Decoupling and Inter-view Fusion

AAAI 2025technical

Anomaly detection has garnered significant attention for its extensive industrial application value. Most existing methods focus on single-view scenarios and fail to detect anomalies hidden in blind spots, leaving a gap in addressing the demands of multi-view detection in practical applications. Ens…

2024

PDENNEval: A Comprehensive Evaluation of Neural Network Methods for Solving PDEs

IJCAI 2024poster

The rapid development of neural network (NN) methods for solving partial differential equations (PDEs) has created an urgent need for evaluation and comparison of these methods. In this study, we propose PDENNEval, a comprehensive and systematic evaluation of 12 NN methods for PDEs. These methods…

2024

Relation DETR: Exploring Explicit Position Relation Prior for Object Detection

ECCV 2024oral

"This paper presents a general scheme for enhancing the convergence and performance of DETR (DEtection TRansformer). We investigate the slow convergence problem in transformers from a new perspective, suggesting that it arises from the self-attention that introduces no structural bias over inputs. T…

2024

Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement

CVPR 2024poster

DETR-like methods have significantly increased detection performance in an end-to-end manner. The mainstream two-stage frameworks of them perform dense self-attention and select a fraction of queries for sparse cross-attention which is proven effective for improving performance but also introduces a…

2024

Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection

CVPR 2024poster

Video moment retrieval and highlight detection are two highly valuable tasks in video understanding but until recently they have been jointly studied. Although existing studies have made impressive advancement recently they predominantly follow the data-driven bottom-up paradigm. Such paradigm overl…

2022

Asymmetric Relation Consistency Reasoning for Video Relation Grounding

ECCV 2022poster

"Video relation grounding has attracted growing attention in the fields of video understanding and multimodal learning. While the past years have witnessed remarkable progress in this issue, the difficulties of multi-instance and complex temporal reasoning make it still a challenging task. In this p…

Cited by 5SourcePDFScholar
2022

Image Steganalysis with Convolutional Vision Transformer

ICASSP 2022accepted

Recent research has shown that deep learning based methods offer more accurate detection for image steganalysis than the traditional detection paradigm based on rich media models. Existing network architectures based on deep learning, however, stack more and more convolutional layers to increase loc…

Cited by 0SourceScholar
2022

Joint Learning for Addressee Selection and Response Generation in Multi-Party Conversation

ICASSP 2022accepted

A large number of multi-party conversation scenarios exist in social networks, which have been seldom studied in the field of human-machine conversation. In this paper, we study a novel task of joint learning for addressee selection and response generation in multi-party conversations. Systems are e…

Cited by 0SourceScholar
2018

Where and Why Are They Looking? Jointly Inferring Human Attention and Intentions in Complex Tasks

CVPR 2018poster

This paper addresses a new problem - jointly inferring human attention, intentions, and tasks from videos. Given an RGB-D video where a human performs a task, we answer three questions simultaneously: 1) where the human is looking - attention prediction; 2) why the human is looking there - intention…

Cited by 82SourcePDFScholar