← Search

Jingjia Huang

8 accepted papers

2025

Accelerated Diffusion via High-Low Frequency Decomposition for Pan-Sharpening

AAAI 2025technical

Pan-sharpening aims to preserve the spectral information of the multi-spectral (MS) image while leveraging the high-frequency details from the guided high-resolution panchromatic (PAN) image to enhance its spatial resolution. The key challenge is how to preserve the spectral information from the MS…

Cited by 0SourcePDFScholar
2025

Sp3ctralMamba: Physics-Driven Joint State Space Model for Hyperspectral Image Reconstruction

AAAI 2025technical

Hyperspectral image (HSI) reconstruction aims to restore the original 3D HSIs from the 2D hyperspectral snapshot compressive images (SCIs). The key to high-fidelity HSI reconstruction lies in designing refined spatial and spectral attention mechanisms, which are crucial for generating fine-grained r…

Cited by 0SourcePDFScholar
2024

Progressive High-Frequency Reconstruction for Pan-Sharpening with Implicit Neural Representation

AAAI 2024technical

Pan-sharpening aims to leverage the high-frequency signal of the panchromatic (PAN) image to enhance the resolution of its corresponding multi-spectral (MS) image. However, deep neural networks (DNNs) tend to prioritize learning the low-frequency components during the training process, which limits…

Cited by 11SourcePDFScholar
2024

Stitching Segments and Sentences towards Generalization in Video-Text Pre-training

AAAI 2024technical

Video-language pre-training models have recently achieved remarkable results on various multi-modal downstream tasks. However, most of these models rely on contrastive learning or masking modeling to align global features across modalities, neglecting the local associations between video frames and…

Cited by 6SourcePDFScholar
2023

Causality Compensated Attention for Contextual Biased Visual Recognition

ICLR 2023poster

Visual attention does not always capture the essential object representation desired for robust predictions. Attention modules tend to underline not only the target object but also the common co-occurring context that the module thinks helpful in the training. The problem is rooted in the confoundin…

Cited by 23SourcePDFScholar
2023

Clover: Towards a Unified Video-Language Alignment and Fusion Model

CVPR 2023poster

Building a universal video-language model for solving various video understanding tasks (e.g., text-video retrieval, video question answering) is an open challenge to the machine learning field. Towards this goal, most recent works build the model by stacking uni-modal and cross-modal feature encode…

2023

Revisiting Temporal Modeling for CLIP-Based Image-to-Video Knowledge Transferring

CVPR 2023poster

Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention for their potential to improve visual representation learning in the video domain. In this paper, based on the CLIP model…

2019

AttPool: Towards Hierarchical Feature Representation in Graph Convolutional Networks via Attention Mechanism

ICCV 2019poster

Graph convolutional networks (GCNs) are potentially short of the ability to learn hierarchical representation for graph embedding, which holds them back in the graph classification task. Here, we propose AttPool, which is a novel graph pooling module based on attention mechanism, to remedy the probl…

Cited by 87PDFcodeScholar