← Search

Qian Wu

9 accepted papers

2026

SpiralDiff: Spiral Diffusion with LoRA for RGB-to-RAW Conversion Across Cameras

CVPR 2026

RAW images preserve superior fidelity and rich scene information compared to RGB, making them essential for tasks in challenging imaging conditions. To alleviate the high cost of data collection, recent RGB-to-RAW conversion methods aim to synthesize RAW images from RGB. However, they overlook two k

Cited by 0SourcecodeScholar
2025

DDxTutor: Clinical Reasoning Tutoring System with Differential Diagnosis-Based Structured Reasoning

ACL 2025long

Clinical diagnosis education requires students to master both systematic reasoning processes and comprehensive medical knowledge. While recent advances in Large Language Models (LLMs) have enabled various medical educational applications, these systems often provide direct answers that could reduce…

2025

HealthCards: Exploring Text-to-Image Generation as Visual Aids for Healthcare Knowledge Democratizing and Education

EMNLP 2025

The evolution of text-to-image (T2I) generation techniques has introduced new capabilities for information visualization, with the potential to advance knowledge democratization and education. In this paper, we investigate how T2I models can be adapted to generate educational health knowledge conten

2024

Class-Agnostic Object Counting with Text-to-Image Diffusion Model

ECCV 2024poster

"Class-agnostic object counting aims to count objects of arbitrary classes with limited information (, a few exemplars or the class names) provided. It requires the model to effectively acquire the characteristics of the target objects and accurately perform counting, which can be challenging. In th…

Cited by 7SourcePDFScholar
2024

Differentiable Resolution Compression and Alignment for Efficient Video Classification and Retrieval

ICASSP 2024accepted

Optimizing video inference efficiency has become increasingly important with the growing demand for video analysis in various fields. Some existing methods achieve high efficiency by explicit discard of spatial or temporal information, which poses challenges in fast-changing and fine-grained scenari…

Cited by 0SourceScholar
2024

HaltingVT: Adaptive Token Halting Transformer for Efficient Video Recognition

ICASSP 2024accepted

Action recognition in videos poses a challenge due to its high computational cost, especially for Joint Space-Time video transformers (Joint VT). Despite their effectiveness, the excessive number of tokens in such architectures significantly limits their efficiency. In this paper, we propose Halting…

Cited by 0SourceScholar
2022

Motion Sensitive Contrastive Learning for Self-Supervised Video Representation

ECCV 2022poster

"Contrastive learning has shown great potential in video representation learning. However, existing approaches fail to sufficiently exploit short-term motion dynamics, which are crucial to various down-stream video understanding tasks. In this paper, we propose Motion Sensitive Contrastive Learning…

Cited by 20SourcePDFScholar
2021

End-to-End Human Object Interaction Detection With HOI Transformer

CVPR 2021poster

We propose HOI Transformer to tackle human object interaction (HOI) detection in an end-to-end manner. Current approaches either decouple HOI task into separated stages of object detection and interaction classification or introduce surrogate interaction problem. In contrast, our method, named HOI T…

Cited by 266PDFcodeScholar
2016

A new haze image database with detailed air quality information and a novel no-reference image quality assessment method for haze images

ICASSP 2016accepted

In this paper, we propose a new standard haze image database with nearly all kinds of haze situations. Our database includes haze-free images as well as different levels and situations of haze images, such as snowy and extremely serious haze images. Our database also records the related weather and…

Cited by 0SourceScholar