← Search

Tao Zhou

29 accepted papers

2026

Bidirectional Channel-selective Semantic Interaction for Semi-Supervised Medical Segmentation

AAAI 2026technical

Semi-supervised medical image segmentation is an effective method for addressing scenarios with limited labeled data. Existing methods mainly rely on frameworks such as mean teacher and dual-stream consistency learning. These approaches often face issues like error accumulation and model structural

Cited by 0SourcePDFScholar
2026

Coarse-to-Fine Latent Guidance: A Multi-Scale Diffusion Transformer for Traffic Flow Forecasting

IJCAI 2026

Accurate traffic flow prediction is fundamental to Intelligent Transportation Systems (ITS). However, traffic dynamics exhibit inherent multi-scale heterogeneity, where stable global trends are often masked by stochastic local fluctuations. Existing methods struggle to reconcile these conflicting re

Cited by 0Scholar
2026

Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization

CVPR 2026

Generating high-fidelity audio that is both semantically meaningful and temporally synchronized with silent videos remains a challenging problem in video-to-audio generation. Existing approaches often fail to capture fine-grained temporal correspondence between visual events and audio dynamics, lead

Cited by 0SourcecodeScholar
2026

MIMIC: Mask-Injected Manipulation Video Generation with Interaction Control

ICLR 2026poster

Embodied intelligence faces a fundamental bottleneck from limited large-scale interaction data. Video generation offers a scalable alternative, but manipulation videos remain particularly challenging, as they require capturing subtle, contact-rich dynamics. Despite recent advances, video diffusion m…

Cited by 0SourceScholar
2026

MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration

CVPR 2026

Medical vision-language pretraining (VLP) models have recently been investigated for their generalization to diverse downstream tasks. However, current medical VLP methods typically force the model to learn simple and complex concepts simultaneously. This anti-cognitive process leads to suboptimal f

Cited by 0SourcecodeScholar
2026

Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models

CVPR 2026

The rapid progress of Multi-Modal Large Language Models (MLLMs) has significantly advanced downstream applications. However, this progress also exposes serious transferable adversarial vulnerabilities. In general, existing adversarial attacks against MLLMs typically rely on surrogate models trained

Cited by 0SourcecodeScholar
2026

SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentation

CVPR 2026

Semi-supervised learning addresses label scarcity and high annotation costs in medical image segmentation by exploiting the latent information in unlabeled data to enhance model performance. Traditional discriminative segmentation relies on segmentation masks, neglecting feature-level distribution c

Cited by 0SourcecodeScholar
2026

Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic Space

ICML 2026poster

Infrared and visible image fusion aims to integrate complementary information from both modalities. However, most existing methods rely on Euclidean representations, which inherently impose geometric constraints that hinder effective semantic modelling. Specifically, Euclidean geometry imposes rigid…

Cited by 0SourceScholar
2026

Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction

CVPR 2026

With the rapid advancement and widespread application of vision-language pre-training (VLP) models, their vulnerability to adversarial attacks has become a critical concern. In general, the adversarial examples can typically be designed to exhibit transferable power, attacking not only different mod

Cited by 0SourcecodeScholar
2026

VELR: Efficient Video Reward Feedback via Ensemble Latent Reward Models

ICML 2026poster

Reward feedback learning (ReFL) is effective for both text-to-image (T2I) and text-to-video (T2V) generation with image reward models (RMs). However, image RMs are misaligned with temporal objectives of T2V, motivating ReFL with video reward models. Nevertheless, directly deploying video RMs is impr…

Cited by 0SourceScholar
2025

A Correlation Manifold Self-Attention Network for EEG Decoding

IJCAI 2025

Riemannian neural networks, which generalize the deep learning paradigm to non-Euclidean geometries, have garnered widespread attention across diverse applications in artificial intelligence. Among these, the representative attention models have been studied on various non-Euclidean spaces to geomet

2025

Bridging the User-side Knowledge Gap in Knowledge-aware Recommendations with Large Language Models

AAAI 2025technical

In recent years, knowledge graphs have been integrated into recommender systems as item-side auxiliary information, enhancing recommendation accuracy. However, constructing and integrating structural user-side knowledge remains a significant challenge due to the improper granularity and inherent sca…

2025

Compositional Condition Question Answering in Tabular Understanding

ICML 2025poster

Multimodal Large Language Models (MLLMs) for tabular understanding have made significant progress in tasks such as financial report analysis and public data tests. However, our comprehensive analysis shows that these models are still limited in certain simple scenarios, particularly when handling co…

2024

Enhancing Event Sequence Modeling with Contrastive Relational Inference

ICASSP 2024accepted

Neural temporal point processes(TPPs) have shown promise for modeling continuous-time event sequences. However, capturing the interactions between events is challenging yet critical for performing inference tasks like forecasting on event sequence data. Existing TPP models have focused on parameteri…

Cited by 0SourceScholar
2024

Memory-Assisted Sub-Prototype Mining for Universal Domain Adaptation

ICLR 2024poster

Universal domain adaptation aims to align the classes and reduce the feature gap between the same category of the source and target domains. The target private category is set as the unknown class during the adaptation process, as it is not included in the source domain. However, most existing metho…

Cited by 2SourcePDFScholar
2024

TPGP: Temporal-Parametric Optimization with Deep Grasp Prior for Dexterous Motion Planning

ICRA 2024poster

Grasping motion planning aims to find a feasible grasping trajectory in the configuration space given an input target grasp. While optimizing grasp motion with two or three-fingered grippers has been well studied, the study on natural grasp motion planning with a dexterous hand remains a very challe…

Cited by 2SourceScholar
2023

F&F Attack: Adversarial Attack against Multiple Object Trackers by Inducing False Negatives and False Positives

ICCV 2023poster

Multi-object tracking (MOT) aims to build moving trajectories for number-agnostic objects. Modern multi-object trackers commonly follow the tracking-by-detection strategy. Therefore, fooling detectors can be an effective solution but it usually requires attacks in multiple successive frames, resulti…

Cited by 10PDFScholar
2023

Prompting Neural Machine Translation with Translation Memories

AAAI 2023technical

Improving machine translation (MT) systems with translation memories (TMs) is of great interest to practitioners in the MT community. However, previous approaches require either a significant update of the model architecture and/or additional training efforts to make the models well-behaved when TMs…

Cited by 14SourcePDFScholar
2022

ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation

ACL 2022long

Residual networks are an Euler discretization of solutions to Ordinary Differential Equations (ODE). This paper explores a deeper relationship between Transformer and numerical ODE methods. We first show that a residual block of layers in Transformer can be described as a higher-order solution to OD…

2022

On Vision Features in Multimodal Machine Translation

ACL 2022long

Previous work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is on the quality of vision models. In this work, we investigate the impact of vision models on MMT. Given the fact that Transformer is becoming popular…

2021

CCT-Net: Category-Invariant Cross-Domain Transfer for Medical Single-to-Multiple Disease Diagnosis

ICCV 2021poster

A medical imaging model is usually explored for the diagnosis of a single disease. However, with the expanding demand for multi-disease diagnosis in clinical applications, multi-function solutions need to be investigated. Previous works proposed to either exploit different disease labels to conduct…

Cited by 11PDFScholar
2021

Context-aware Cross-level Fusion Network for Camouflaged Object Detection

IJCAI 2021poster

Camouflaged object detection (COD) is a challenging task due to the low boundary contrast between the object and its surroundings. In addition, the appearance of camouflaged objects varies significantly, e.g., object size and shape, aggravating the difficulties of accurate COD. In this paper, we pro…

2021

Specificity-Preserving RGB-D Saliency Detection

ICCV 2021poster

RGB-D saliency detection has attracted increasing attention, due to its effectiveness and the fact that depth cues can now be conveniently captured. Existing works often focus on learning a shared representation through various fusion strategies, with few methods explicitly considering how to preser…

Cited by 256PDFcodeScholar
2021

Visual-Textual Attentive Semantic Consistency for Medical Report Generation

ICCV 2021poster

Diagnosing diseases from medical radiographs and writing reports requires professional knowledge and is time-consuming. To address this, automatic medical report generation approaches have recently gained interest. However, identifying diseases as well as correctly predicting their corresponding siz…

Cited by 24PDFScholar
2020

Deep Learning of Neuromuscular and Visuomotor Control of a Biomimetic Simulated Humanoid

RA-L 2020

We present a biomimetic framework for human neuromuscular and visuomotor control that promises to be of value to researchers developing humanoid robots. Our framework features a biomechanically simulated human musculoskeletal model, actuated by numerous skeletal muscles, with realistic eyes driven b

Cited by 2SourceScholar
2020

Multi-Mutual Consistency Induced Transfer Subspace Learning for Human Motion Segmentation

CVPR 2020poster

Human motion segmentation based on transfer subspace learning is a rising interest in action-related tasks. Although progress has been made, there are still several issues within the existing methods. First, existing methods transfer knowledge from source data to target tasks by learning domain-inva…

Cited by 43PDFScholar
2019

PPR-Net:Point-wise Pose Regression Network for Instance Segmentation and 6D Pose Estimation in Bin-picking Scenarios

IROS 2019poster

Accurate object 6D pose estimation is a core task for robot bin-picking applications, especially when objects are randomly stacked with heavy occlusion. To address this problem, this paper proposes a simple but novel Point-wise Pose Regression Network (PPR-Net). For each point in the point cloud, th…

Cited by 87SourceScholar
2016

Robust visual tracking via inverse nonnegative matrix factorization

ICASSP 2016accepted

The establishment of robust target appearance model over time is an overriding concern in visual tracking. In this paper, we propose an inverse nonnegative matrix factorization (NMF) method for robust appearance modeling. Rather than using a linear combination of nonnegative basis vectors for each t…

Cited by 0SourceScholar
2015

Small target detection using an optimization-based filter

ICASSP 2015accepted

Small target detection is a critical problem in the Infrared Search And Track (IRST) system. Although it has been studied for years, there are some challenges remained, e.g. cloud edges and horizontal lines are likely to cause false alarms. This paper proposes a novel method using an optimization-ba…

Cited by 0SourceScholar