← Search

Zhihua Wang

27 accepted papers

2026

MDS-VQA: Model-Informed Data Selection for Video Quality Assessment

CVPR 2026

Learning-based video quality assessment (VQA) has advanced rapidly, yet progress is increasingly constrained by a disconnect between model design and dataset curation. Model-centric approaches often iterate on fixed benchmarks, while data-centric efforts collect new human labels without systematical

Cited by 0SourcecodeScholar
2026

PhysInOne: Visual Physics Learning and Reasoning in One Suite

CVPR 2026

We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike existing datasets limited to merely hundreds or thousands of examples, PhysInOne provides 2 million videos across 153,810 dynamic 3D scenes, covering 71

Cited by 0SourcecodeScholar
2026

RoboMatch: A Unified Mobile-Manipulation Teleoperation Platform with Auto-Matching Network Architecture for Long-Horizon Tasks

ICRA 2026poster

This paper presents RoboMatch, a novel unified teleoperation platform for mobile manipulation with an auto-matching network architecture, designed to tackle long-horizon tasks in dynamic environments. Our system enhances teleoperation performance, data collection efficiency, task accuracy, and opera…

2025

Deep Opinion-Unaware Blind Image Quality Assessment by Learning and Adapting from Multiple Annotators

IJCAI 2025

Existing deep neural network (DNN)-based blind image quality assessment (BIQA) methods primarily rely on human-rated datasets for training. However, collecting human labels is extremely time-consuming and labor-intensive, posing a significant bottleneck for practical applications. To address this ch

2025

Enhancing Low-Light Images: A Synthetic Data Perspective on Practical and Generalizable Solutions

AAAI 2025technical

Recently, deep neural networks (DNNs) have emerged as the leading approach for low-light image enhancement (LLIE). However, training these models generally requires large-scale paired datasets, which are challenging to obtain due to the labor-intensive and time-consuming nature of real-world data co…

2025

ICAA-Mamba: Vision Mamba for Image Color Aesthetics Assessment

ICASSP 2025accepted

Image Color Aesthetics Assessment (ICAA) focuses on evaluating the aesthetic quality of color composition within images. This task involves analyzing and quantifying the visual appeal of color arrangements, taking into account factors such as harmony, contrast, and balance, to provide an objective a…

Cited by 0SourceScholar
2025

Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models

ACL 2025long

Interpretation is critical for disease diagnosis, but existing models struggle to balance predictive accuracy with human-understandable rationales. While large language models (LLMs) offer strong reasoning abilities, their clinical use is limited by high computational costs and restricted multimodal…

2025

MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal Guidance

AAAI 2025technical

Neural Representations for Videos (NeRV) has emerged as a promising implicit neural representation (INR) approach for video analysis, which represents videos as neural networks with frame indexes as inputs. However, NeRV-based methods are time-consuming when adapting to a large number of diverse v…

2025

Multi-view Evidential Learning-based Medical Image Segmentation

AAAI 2025technical

Medical image segmentation provides useful information about the shape and size of organs, which is beneficial for improving diagnosis, analysis, and treatment. Despite traditional deep learning-based models can extract domain-specific knowledge, they face a generalization bottleneck due to the limi…

Cited by 0SourcePDFScholar
2025

ProMedTS: A Self-Supervised, Prompt-Guided Multimodal Approach for Integrating Medical Text and Time Series

ACL 2025finding

Large language models (LLMs) have shown remarkable performance in vision-language tasks, but their application in the medical field remains underexplored, particularly for integrating structured time series data with unstructured clinical notes. In clinical practice, dynamic time series data, such a…

Cited by 0SourcePDFScholar
2025

Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition

ACL 2025long

The past years have witnessed a proliferation of large language models (LLMs). Yet, reliable evaluation of LLMs is challenging due to the inaccuracy of standard metrics in human perception of text quality and the inefficiency in sampling informative test examples for human evaluation. This paper pre…

2024

Active Retrosynthetic Planning Aware of Route Quality

ICLR 2024poster

Retrosynthetic planning is a sequential decision-making process of identifying synthetic routes from the available building block materials to reach a desired target molecule. Though existing planning approaches show promisingly high solving rates and low costs, the trivial route cost evaluation via…

Cited by 3SourcePDFScholar
2024

Attention Beats Linear for Fast Implicit Neural Representation Generation

ECCV 2024poster

"Implicit Neural Representation (INR) has gained increasing popularity as a data representation method, serving as a prerequisite for innovative generation models. Unlike gradient-based methods, which exhibit lower efficiency in inference, the adoption of hyper-network for generating parameters in M…

2024

Multimodal Representation Distribution Learning for Medical Image Segmentation

IJCAI 2024poster

Medical image segmentation is one of the most critical tasks in medical image analysis. However, the performance of existing methods is limited by the lack of high-quality labeled data due to the expensive data annotation. To alleviate this limitation, we propose a novel multi-modal learning method…

2024

Multiscale Sliced Wasserstein Distances as Perceptual Color Difference Measures

ECCV 2024poster

"Contemporary color difference (CD) measures for photographic images typically operate by comparing co-located pixels, patches in a “perceptually uniform” color space, or features in a learned latent space. Consequently, these measures inadequately capture the human color perception of misaligned im…

2024

RetroOOD: Understanding Out-of-Distribution Generalization in Retrosynthesis Prediction

AAAI 2024technical

Machine learning-assisted retrosynthesis prediction models have been gaining widespread adoption, though their performances oftentimes degrade significantly when deployed in real-world applications embracing out-of-distribution (OOD) molecules or reactions. Despite steady progress on standard benchm…

Cited by 4SourcePDFScholar
2023

Federated Intelligent Terminals Facilitate Stuttering Monitoring

ICASSP 2023accepted

Stuttering is a complicated language disorder. The most common form of stuttering is developmental stuttering, which begins in childhood. Early monitoring and intervention are essential for the treatment of children with stuttering. Automatic speech recognition technology has shown its great potenti…

Cited by 0SourceScholar
2023

Learning Chemical Rules of Retrosynthesis with Pre-training

AAAI 2023technical

Retrosynthesis aided by artificial intelligence has been a very active and bourgeoning area of research, for its critical role in drug discovery as well as material science. Three categories of solutions, i.e., template-based, template-free, and semi-template methods, constitute mainstream solutions…

Cited by 15SourcePDFScholar
2023

Learning Instrumental Variable from Data Fusion for Treatment Effect Estimation

AAAI 2023technical

The advent of the big data era brought new opportunities and challenges to draw treatment effect in data fusion, that is, a mixed dataset collected from multiple sources (each source with an independent treatment assignment mechanism). Due to possibly omitted source labels and unmeasured confounders…

2023

Learning a Deep Color Difference Metric for Photographic Images

CVPR 2023poster

Most well-established and widely used color difference (CD) metrics are handcrafted and subject-calibrated against uniformly colored patches, which do not generalize well to photographic images characterized by natural scene complexities. Constructing CD formulae for photographic images is still an…

2023

PTADisc: A Cross-Course Dataset Supporting Personalized Learning in Cold-Start Scenarios

NeurIPS 2023poster

The focus of our work is on diagnostic tasks in personalized learning, such as cognitive diagnosis and knowledge tracing. The goal of these tasks is to assess students' latent proficiency on knowledge concepts through analyzing their historical learning records. However, existing research has been l…

2022

ConfounderGAN: Protecting Image Data Privacy with Causal Confounder

NeurIPS 2022accept

The success of deep learning is partly attributed to the availability of massive data downloaded freely from the Internet. However, it also means that users' private data may be collected by commercial organizations without consent and used to train their models. Therefore, it's important and necess…

Cited by 5SourcePDFScholar
2022

Improving Deep Embedded Clustering via Learning Cluster-level Representations

COLING 2022main

Driven by recent advances in neural networks, various Deep Embedding Clustering (DEC) based short text clustering models are being developed. In these works, latent representation learning and text clustering are performed simultaneously. Although these methods are becoming increasingly popular, the…

Cited by 1SourcePDFScholar
2022

The Role of Deconfounding in Meta-learning

ICML 2022spotlight

Meta-learning has emerged as a potent paradigm for quick learning of few-shot tasks, by leveraging the meta-knowledge learned from meta-training tasks. Well-generalized meta-knowledge that facilitates fast adaptation in each task is preferred; however, recent evidence suggests the undesirable memori…

2020

RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds

CVPR 2020oral

We study the problem of efficient semantic segmentation for large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we i…

Cited by 2143PDFcodeScholar
2018

DEFO-NET: Learning Body Deformation Using Generative Adversarial Networks

ICRA 2018poster

Modelling the physical properties of everyday objects is a fundamental prerequisite for autonomous robots. We present a novel generative adversarial network (DEFO-NET), able to predict body deformations under external forces from a single RGB-D image. The network is based on an invertible conditiona…

Cited by 10SourceScholar