← Search

Jia Wang

35 accepted papers

2026

A Hybrid Space Model for Misaligned Multi-modality Image Fusion

AAAI 2026technical

Infrared and visible image fusion aims to integrate complementary information, such as thermal saliency from infrared imagery and fine-grained texture details from visible imagery. However, real-world multi-modal misalignment and geometric deformation often introduce severe artifacts. Most existing

Cited by 1SourcePDFScholar
2026

Bidirectional Noise Injection: Enhancing Diffusion Models via Coordinated Input-Output Perturbation

AAAI 2026technical

Diffusion models have demonstrated remarkable success in image generation, yet a persistent challenge remains: the bias between model predictions and the target distribution. In this paper, we propose a Bidirectional Noise Injection framework for enhancing diffusion models, implemented via Coordinat

Cited by 0SourcePDFScholar
2026

Breaking Down Market Barriers: Distilled Prompt-Tuning Approach for Cross-Market Recommendation

AAAI 2026technical

Cross-market recommendation (CMR) faces severe challenges from distribution shifts between data-rich source markets and sparse target markets. Existing methods rely on a pre-training and fine-tuning paradigm for knowledge transfer, yet suffer from two key limitations: i) the objective gap between pr

Cited by 0SourcePDFScholar
2026

C^2FG: Control Classifier-Free Guidance via Score Discrepancy Analysis

CVPR 2026

Classifier-Free Guidance (CFG) is a cornerstone of modern conditional diffusion models, yet its reliance on the fixed or heuristic dynamic guidance weight is predominantly empirical and overlooks the inherent dynamics of the diffusion process. In this paper, we provide a rigorous theoretical analysi

Cited by 0SourceScholar
2026

From IDs to Semantics: A Generative Framework for Cross-Domain Recommendation with Adaptive Semantic Tokenization

AAAI 2026technical

Cross-domain recommendation (CDR) is crucial for improving recommendation accuracy and generalization, yet traditional methods are often hindered by the reliance on shared user/item IDs, which are unavailable in most real-world scenarios. Consequently, many efforts have focused on learning disentang

Cited by 0SourcePDFScholar
2026

Mitigating Gradient Pathology in PINNs through Aligned Constraint

ICML 2026poster

While Physics-Informed Neural Networks (PINNs) are powerful for solving Partial Differential Equations (PDEs), their training is often paralyzed by gradient pathology. The gradients from PDE residuals and boundary constraints oppose each other, trapping the model in local minima. Current solutions, …

Cited by 0SourceScholar
2026

PU-BENCH: A UNIFIED BENCHMARK FOR RIGOROUS AND REPRODUCIBLE PU LEARNING

ICLR 2026poster

Positive-Unlabeled (PU) learning, a challenging paradigm for training binary classifiers from only positive and unlabeled samples, is fundamental to many applications. While numerous PU learning methods have been proposed, the research is systematically hindered by the lack of a standardized and com…

Cited by 0SourcecodeScholar
2026

Proactive Risk-Aware Trajectory Planning for Autonomous Driving in Unstructured Environments Via Reinforcement Learning with Adaptive Reward Design

ICRA 2026poster

Trajectory planning for autonomous driving in dynamic unstructured traffic remains a fundamental challenge. Existing methods are often reactive, i.e., they only respond to observed situations without explicitly anticipating future risks. Moreover, most reinforcement learning based approaches rely on…

Cited by 0Scholar
2026

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference

ICML 2026poster

DeepSeek-OCR leverages visual–text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundant textual and structural information. Moreover, current token pruning methods for conventional vision–language models (VLMs) fail to preserve textua…

Cited by 0SourceScholar
2025

DPGP: A Hybrid 2D-3D Dual Path Potential Ghost Probe Zone Prediction Framework for Safe Autonomous Driving

IROS 2025

Modern robots must coexist with humans in dense urban environments. A key challenge is the ghost probe problem, where pedestrians or objects unexpectedly rush into traffic paths. This issue affects both autonomous vehicles and human drivers. Existing works propose vehicle-to-everything (V2X) strateg

Cited by 2SourceScholar
2025

DiffusionREC: Diffusion Model with Adaptive Condition for Referring Expression Comprehension

AAAI 2025technical

The objective of referring expression comprehension (REC) is to accurately identify the object in an image described by a given expression. Existing REC methods, including transformer-based and graph-based approaches among others, have shown robust performance in REC tasks. In this study, we present…

Cited by 0SourcePDFScholar
2025

GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning

NeurIPS 2025poster

With the rapid development of Large Vision Language Models, the focus of Graphical User Interface (GUI) agent tasks shifts from single-screen tasks to complex screen navigation challenges. However, real-world GUI environments, such as PC software and mobile Apps, are often complex and proprietary,…

Cited by 0SourceScholar
2025

GaussianEnhancer: A General Rendering Enhancer for Gaussian Splatting

ICASSP 2025accepted

Gaussian Splatting (GS) methods, including 3DGS and 2DGS, have demonstrated exceptional performance in real-time novel view synthesis (NVS), emerging as a transformative technology in the fields of explicit rendering and computer graphics. However, GS-based methods still face challenges in rendering…

Cited by 0SourceScholar
2025

Hybrid Spatial-Frequency Attention Network For Fine-Grained Skeleton-Based Action Recognition

ICASSP 2025accepted

Recently, Transformer-based methods have gained popularity in skeleton-based action recognition due to their advantages in modeling long-range dependencies. However, Transformer lacks the inductive biases towards skeletal topology and tends to capture salient features, potentially overlooking subtle…

Cited by 0SourceScholar
2025

Information Density Principle for MLLM Benchmarks

ICCV 2025poster

With the emergence of Multimodal Large Language Models (MLLMs), hundreds of benchmarks have been developed to ensure the reliability of MLLMs in downstream tasks. However, the evaluation mechanism itself may not be reliable. For developers of MLLMs, questions remain about which benchmark to use and…

2025

MMET: A Multi-Input and Multi-Scale Transformer for Efficient PDEs Solving

IJCAI 2025

Partial Differential Equations (PDEs) are fundamental for modeling physical systems, yet solving them in a generic and efficient manner using machine learning-based approaches remains challenging due to limited multi-input and multi-scale generalization capabilities, as well as high computational co

2025

On Designing General and Expressive Quantum Graph Neural Networks with Applications to MILP Instance Representation

ICLR 2025poster

Graph-structured data is ubiquitous, and graph learning models have recently been extended to address complex problems like mixed-integer linear programming (MILP). However, studies have shown that the vanilla message-passing based graph neural networks (GNNs) suffer inherent limitations in learning…

Cited by 1SourcePDFScholar
2025

Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning

NeurIPS 2025poster

The remarkable reasoning capability of large language models (LLMs) stems from cognitive behaviors that emerge through reinforcement with verifiable rewards. This work investigates how to transfer this principle to Multimodal LLMs (MLLMs) to unlock advanced visual reasoning. We introduce a two-stage…

Cited by 0SourceScholar
2025

PScalpel: A Machine Learning-based Guider for Protein Phase-Separating Behaviour Alteration

AAAI 2025technical

Missense mutations could affect the Liquid-Liquid Phase Separation (LLPS) propensity of proteins and lead to aberrant phase-separating behaviours, which are recently found to be associated with many diseases including Alzheimer's and cancer. However, the regulatory role of mutations in LLPS remains…

2025

Pruning for Sparse Diffusion Models Based on Gradient Flow

ICASSP 2025accepted

Diffusion Models (DMs) have impressive capabilities among generation models, but are limited to slower inference speeds and higher computational costs. Previous works utilize one-shot structure pruning to derive lightweight DMs from pre-trained ones, but this approach often leads to a significant dr…

Cited by 0SourceScholar
2025

Reinforcement Active Client Selection for Federated Heterogeneous Graph Learning

AAAI 2025technical

Carefully selecting clients to participate in aggregation can assist the global model in achieving better performance. However, existing research on federated heterogeneous graph learning (FHGL) has shown limited attention to the client selection (CS) problem. Current CS algorithms face challenges i…

Cited by 0SourcePDFScholar
2025

Role-aware Multi-agent Reinforcement Learning for Coordinated Emergency Traffic Control

NeurIPS 2025poster

Emergency traffic control presents an increasingly critical challenge, requiring seamless coordination among emergency vehicles, regular vehicles, and traffic lights to ensure efficient passage for all vehicles. Existing models primarily only focus on traffic light control, leaving emergency and reg…

Cited by 0SourceScholar
2025

SILM: A Subjective Intent Based Low-Latency Framework for Multiple Traffic Participants Joint Trajectory Prediction

IROS 2025

Trajectory prediction is a fundamental technology for advanced autonomous driving systems and represents one of the most challenging problems in the field of cognitive intelligence. Accurately predicting the future trajectories of each traffic participant is a prerequisite for building high safety a

Cited by 0SourceScholar
2025

SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models

EMNLP 2025

Large language models (LLMs) show strong potential for simulating human social behaviors and interactions, yet lack large-scale, systematically constructed benchmarks for evaluating their alignment with real-world social attitudes. To bridge this gap, we introduce SocioBench—a comprehensive benchmar

2025

TableDreamer: Progressive and Weakness-guided Data Synthesis from Scratch for Table Instruction Tuning

ACL 2025finding

Despite the commendable progress of recent LLM-based data synthesis methods, they face two limitations in generating table instruction tuning data. First, they can not thoroughly explore the vast input space of table understanding tasks, leading to limited data diversity. Second, they ignore the wea…

2024

AttentionLUT: Attention Fusion-Based Canonical Polyadic LUT for Real-Time Image Enhancement

ICASSP 2024accepted

Recently, many algorithms have employed image-adaptive lookup tables (LUTs) to achieve real-time image enhancement. Nonetheless, a prevailing trend among existing methods has been the employment of linear combinations of basic LUTs to formulate image-adaptive LUTs, which limits the generalization ab…

Cited by 0SourceScholar
2024

Bias-aware Boolean Matrix Factorization Using Disentangled Representation Learning

UAI 2024poster

Boolean matrix factorization (BMF) has been widely utilized in fields such as recommendation systems, graph learning, text mining, and -omics data analysis. Traditional BMF methods decompose a binary matrix into the Boolean product of two lower-rank Boolean matrices plus homoscedastic random errors.…

2024

Document Set Expansion with Positive-Unlabeled Learning Using Intractable Density Estimation

COLING 2024main

The Document Set Expansion (DSE) task involves identifying relevant documents from large collections based on a limited set of example documents. Previous research has highlighted Positive and Unlabeled (PU) learning as a promising approach for this task. However, most PU methods rely on the unreali…

2024

FAMIM: A Novel Frequency-Domain Augmentation Masked Image Model Framework for Domain Generalizable Face Anti-Spoofing

ICASSP 2024accepted

While existing face anti-spoofing (FAS) methods have achieved high performance on in-domain datasets, good generalization is crucial for their real-world application. Previous domain generalizable FAS methods have attempted to identify common features of live samples from different domains in the sp…

Cited by 0SourceScholar
2024

Fine-Grained Discrepancy Contrastive Learning for Robust Fake News Detection

ICASSP 2024accepted

In recent years, fake news on social media has become a significant threat to societal security, elevating fake news detection to a research priority. Among various strategies, fact-checking detection methods stand out for their accuracy, leveraging evidence from dedicated fact databases. However, t…

Cited by 0SourceScholar
2024

GeneFormer: Learned Gene Compression using Transformer-Based Context Modeling

ICASSP 2024accepted

The development of gene sequencing technology sparks an explosive growth of gene data. Thus, the storage of gene data has become an important issue. Recently, researchers begin to investigate deep learning-based gene data compression, which outperforms general traditional methods. In this paper, we…

Cited by 0SourceScholar
2024

Practical Privacy-Preserving MLaaS: When Compressive Sensing Meets Generative Networks

AAAI 2024technical

The Machine-Learning-as-a-Service (MLaaS) framework allows one to grab low-hanging fruit of machine learning techniques and data science, without either much expertise for this sophisticated sphere or provision of specific infrastructures. However, the requirement of revealing all training data to t…

Cited by 1SourcePDFScholar
2023

Certifiable Out-of-Distribution Generalization

AAAI 2023technical

Machine learning methods suffer from test-time performance degeneration when faced with out-of-distribution (OoD) data whose distribution is not necessarily the same as training data distribution. Although a plethora of algorithms have been proposed to mitigate this issue, it has been demonstrated t…

2021

Bounds all around: training energy-based models with bidirectional bounds

NeurIPS 2021poster

Energy-based models (EBMs) provide an elegant framework for density estimation, but they are notoriously difficult to train. Recent work has established links to generative adversarial networks, where the EBM is trained through a minimax game with a variational value function. We propose a bidirecti…

Cited by 19SourcePDFScholar