← Search

Dahua Gao

9 accepted papers

2025

A Transmitter-Model Unaware Generative Image Compression Framework for Semantic Communication

ICASSP 2025accepted

Unlike traditional bit-level data transmission methods, semantic communication focuses on conveying the meaning behind the data. Though promising results have been achieved, existing end-to-end learning-based semantic communication frameworks often require a synchronization of deep models between th…

Cited by 0SourceScholar
2025

FreeAlign: Superior Text-Image Alignment by Modulating Prompt Attention

ICASSP 2025accepted

In recent years, Text-to-Image (T2I) models have made remarkable advancements, yet accurate accurate association of attributes remains a key challenge. This paper presents FreeAlign, a novel training-free framework designed to enhance attribute alignment in T2I generation. By modulating attention an…

Cited by 0SourceScholar
2025

Stable Control Visual AutoRegressive Model: Precise and Efficient Image Generation via Scale Alignment

ICASSP 2025accepted

Although diffusion models advance condition-based visual generation, they suffer from speed and cost issues, unlike faster AutoRegressive methods that are limited in performance. To address these, we introduce the Stable Control Visual AutoRegressive Model (SCVAR). SCVAR ensures stable control by al…

Cited by 0SourceScholar
2025

Synthesizing Efficient Trajectory-Controllable Co-Speech Gesture with Latent Consistency Model

ICASSP 2025accepted

Diffusion models excel at audio-driven gesture generation. However, the reverse denoising process is computationally intensive and time-redundant. Additionally, existing methods predominantly focus on the quality and diversity of upper-body gestures, without considering the movement trajectories of…

Cited by 0SourceScholar
2024

Mosic: Multimodal Semantic Integrated Communication for Health Monitoring in Iot Scenarios

ICASSP 2024accepted

Monitoring multimodal signals provides a more comprehensive understanding of health conditions compared to singlemode monitoring. In the face of the significant volumes of multimodal signals, existing IoT health monitoring systems primarily focus on high-fidelity signal transmission by encoding mult…

Cited by 0SourceScholar
2024

SG2SC: A Generative Semantic Communication Framework for Scene Understanding-Oriented Image Transmission

ICASSP 2024accepted

In recent years, semantic communication based on deep learning for source-channel joint encoding has garnered significant attention. It utilizes network models trained end-to-end to represent signals as embedding vectors and has demonstrated superior performance compared to traditional methods. Howe…

Cited by 0SourceScholar
2023

Residual Degradation Learning Unfolding Framework With Mixing Priors Across Spectral and Spatial for Compressive Spectral Imaging

CVPR 2023poster

To acquire a snapshot spectral image, coded aperture snapshot spectral imaging (CASSI) is proposed. A core problem of the CASSI system is to recover the reliable and fine underlying 3D spectral cube from the 2D measurement. By alternately solving a data subproblem and a prior subproblem, deep unfold…

2015

High-Speed Hyperspectral Video Acquisition With a Dual-Camera Architecture

CVPR 2015poster

We propose a novel dual-camera design to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. Our work has two key technical contributions. First, we build a dual-camera system that simultaneously captures a panchromatic video at a high frame rate and a hypers…

Cited by 116SourcePDFScholar