← Search

Tao Tan

17 accepted papers

2026

CDMIQA: A Cross-Domain Perceptual Method and Benchmark Dataset for Medical Image Quality Assessment

IJCAI 2026

Medical image quality assessment (IQA) serves as a critical safeguard for precise clinical diagnosis and treatment. However, existing methods still face challenges arising from data scarcity and heterogeneity across imaging domains, which confine solutions to domain-specific designs and limit their

Cited by 0Scholar
2026

FLOW: Optimal Transport-Driven Feature Warping for Generalized Remote Physiological Measurement

CVPR 2026

Remote photoplethysmography (rPPG) enables non-contact physiological measurement from facial videos but often suffers from severe performance degradation under domain shifts. Traditional STMap-based methods [??] rely on predefined spatio-temporal representations that offer engineered robustness but

Cited by 0SourceScholar
2026

HiFi-Mesh: High-Fidelity Efficient 3D Mesh Generation via Compact Autoregressive Dependence

AAAI 2026technical

High-fidelity 3D meshes can be tokenized into one-dimension (1D) sequences and directly modeled using autoregressive approaches for faces and vertices. However, existing methods suffer from insufficient resource utilization, resulting in slow inference and the ability to handle only small-scale sequ

Cited by 0SourcePDFScholar
2026

LUMIN: A Longitudinal Multi-modal Knowledge Decomposition Network for Predicting Breast Cancer Recurrence

AAAI 2026technical

Accurate prediction of breast cancer recurrence after treatment is essential for improving long-term outcomes. However, existing models are limited by three key challenges: (1) they typically rely on single-modal data, missing cross-modal interactions; (2) they analyze static snapshots, failing to c

Cited by 0SourcePDFScholar
2026

Multiple-play Stochastic Bandits with Prioritized Arm Capacity Sharing

AAAI 2026technical

This paper proposes a variant of multiple-play stochastic bandits tailored to resource allocation problems arising from LLM applications, edge intelligence, etc. The model is composed of finite number of arms and plays. Each arm has a stochastic number of capacities, and each unit of capacity is ass

Cited by 0SourcePDFScholar
2026

PHASE-Net: Physics-Grounded Harmonic Attention System for Efficient Remote Photoplethysmography Measurement

CVPR 2026

Remote photoplethysmography (rPPG) measurement enables non-contact physiological monitoring but suffers from accuracy degradation under head motion and illumination changes. Existing deep learning methods are mostly heuristic and lack theoretical grounding, limiting robustness and interpretability.

Cited by 0SourcecodeScholar
2026

PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing

ICLR 2026poster

Remote photoplethysmography (rPPG) enables non-contact physiological measurement but remains highly susceptible to illumination changes, motion artifacts, and limited temporal modeling. Large Language Models (LLMs) excel at capturing long-range dependencies, offering a potential solution but struggl…

Cited by 0SourceScholar
2026

SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition

AAAI 2026technical

Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However, when treating LLM as a recognizer, two questions arise: 1) How can LLMs underst

Cited by 0SourcePDFScholar
2025

A 3D Attenuation Coefficient based Degradation Estimation for Real Non-Homogeneous Dehazing

ICASSP 2025accepted

The end-to-end image dehazing network relies on paired training data. However, there is limited training data available for real non-homogeneous dehazing, which limits the performance of dehazing networks on real non-homogeneous hazy images. To overcome this limitation, we propose a method to augmen…

Cited by 0SourceScholar
2025

Contrastive Learning via Randomly Generated Deep Supervision

ICASSP 2025accepted

Unsupervised visual representation learning has gained significant attention in the computer vision community, driven by recent advancements in contrastive learning. Most existing contrastive learning frameworks rely on instance discrimination as a pretext task, treating each instance as a distinct…

Cited by 0SourceScholar
2025

MoEdit: On Learning Quantity Perception for Multi-object Image Editing

CVPR 2025poster

Multi-object images are widely present in the real world, spanning various areas of daily life. Efficient and accurate editing of these images is crucial for applications such as augmented reality, advertisement design, and medical imaging. Stable Diffusion (SD) has ushered in a new era of high-qual…

2024

A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image Classifications

ICASSP 2024accepted

Although current data augmentation methods are successful to alleviate the data insufficiency, conventional augmentation are primarily intra-domain while advanced generative adversarial networks (GANs) generate images remaining uncertain, particularly in small-scale datasets. In this paper, we propo…

Cited by 0SourceScholar
2024

Adaptive Order Q-learning

IJCAI 2024poster

This paper revisits the estimation bias control problem of Q-learning, motivated by the fact that the estimation bias is not always evil, i.e., some environments benefit from overestimation bias or underestimation bias, while others suffer from these biases. Different from previous coarse-grained b…

Cited by 0SourcePDFScholar
2024

Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents

ACL 2024long

With the remarkable advancements of large language models (LLMs), LLM-based agents have become a research hotspot in human-computer interaction.However, there is a scarcity of benchmarks available for LLM-based mobile agents.Benchmarking these agents generally faces three main challenges:(1) The ine…

2024

MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding

EMNLP 2024finding

Recently, mobile AI agents based on VLMs have been gaining increasing attention. These works typically utilize VLM as a foundation, fine-tuning it with instruction-based mobile datasets. However, these VLMs are typically pre-trained on general-domain data, which often results in a lack of fundamenta…

2019

Local to Global Learning: Gradually Adding Classes for Training Deep Neural Networks

CVPR 2019poster

We propose a new learning paradigm, Local to Global Learning (LGL), for Deep Neural Networks (DNNs) to improve the performance of classification problems. The core of LGL is to learn a DNN model from fewer categories (local) to more categories (global) gradually within the entire training set. LGL i…

Cited by 16PDFcodeScholar