← Search

Baoquan Zhang

18 accepted papers

2026

Improved Masked Image Generation with Knowledge-Augmented Token Representations

AAAI 2026technical

Masked image generation (MIG) has demonstrated remarkable efficiency and high-fidelity images by enabling parallel token prediction. Existing methods typically rely solely on the model itself to learn semantic dependencies among visual token sequences. However, directly learning such semantic depend

Cited by 0SourcePDFScholar
2026

S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum Domain

CVPR 2026

Parameter Efficient Fine-Tuning (PEFT) is a key technique for adapting a large pretrained model to downstream tasks by fine-tuning only a small number of parameters. Recent methods based on Fourier transforms have further reduced the fine-tuned parameters scale by only fine-tuning a few spectral coe

Cited by 0SourceScholar
2026

SJD-SV: Speculative Jacobi Decoding with Semantics Verification for Autoregressive Image Generation

ICML 2026poster

Speculative Jacobi Decoding (SJD) is an important approach for accelerating autoregressive image generation. Although SJD has shown superior performance, recent studies point out that it usually suffers from a token ambiguity issue during token verification but its reason can not be well explained. …

Cited by 0SourceScholar
2026

Satellite-Text-Prompted Large Language Model for Photovoltaic Power Forecasting

AAAI 2026technical

Photovoltaic (PV) power forecasting is critical for the operation of solar power plants and the coordination of energy within power grids. This work aims to predict future PV power time series by leveraging multimodal data. While recent studies have incorporated numerical modalities such as satellit

Cited by 0SourcePDFScholar
2025

AlphaPre: Amplitude-Phase Disentanglement Model for Precipitation Nowcasting

CVPR 2025poster

Precipitation nowcasting involves using current radar observation sequences to predict future radar sequences and determine future precipitation distribution, which is crucial for disaster warning, traffic planning, and agricultural production. Despite numerous advancements, challenges persist in ac…

2025

AsyncDSB: Schedule-Asynchronous Diffusion Schrödinger Bridge for Image Inpainting

AAAI 2025technical

Image inpainting is an important image generation task, which aims to restore corrupted image from partial visible area. Recently, diffusion Schrödinger bridge methods effectively tackle this task by modeling the translation between corrupted and target images as a diffusion Schrödinger bridge proce…

Cited by 0SourcePDFScholar
2025

Sensitivity-Aware Efficient Fine-Tuning via Compact Dynamic-Rank Adaptation

CVPR 2025poster

Parameter-Efficient Fine-Tuning (PEFT) is a fundamental research problem in computer vision, which aims to tune a few of parameters for efficient storage and adaptation of pre-trained vision models. Recently, sensitivity-aware parameter efficient fine-tuning method (SPT) addresses this problem by id…

Cited by 0SourcePDFScholar
2025

Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long Text

CVPR 2025highlight

Image quantization is a crucial technique in image generation, aimed at learning a codebook that encodes an image into a discrete token sequence. Recent advancements have seen researchers exploring learning multi-modal codebook (i.e., text-aligned codebook) by utilizing image caption semantics, aimi…

Cited by 0SourcePDFScholar
2024

A Challenge Dataset and Effective Models for Conversational Stance Detection

COLING 2024main

Previous stance detection studies typically concentrate on evaluating stances within individual instances, thereby exhibiting limitations in effectively modeling multi-party discussions concerning the same specific topic, as naturally transpire in authentic social media interactions. This constraint…

2024

Codebook Transfer with Part-of-Speech for Vector-Quantized Image Modeling

CVPR 2024poster

Vector-Quantized Image Modeling (VQIM) is a fundamental research problem in image synthesis which aims to represent an image with a discrete token sequence. Existing studies effectively address this problem by learning a discrete codebook from scratch and in a code-independent manner to quantize con…

Cited by 11SourcePDFScholar
2024

DiffCast: A Unified Framework via Residual Diffusion for Precipitation Nowcasting

CVPR 2024poster

Precipitation nowcasting is an important spatio-temporal prediction task to predict the radar echoes sequences based on current observations which can serve both meteorological science and smart city applications. Due to the chaotic evolution nature of the precipitation systems it is a very challeng…

2024

LG-VQ: Language-Guided Codebook Learning

NeurIPS 2024poster

Vector quantization (VQ) is a key technique in high-resolution and high-fidelity image synthesis, which aims to learn a codebook to encode an image with a sequence of discrete codes and then generate an image in an auto-regression manner. Although existing methods have shown superior performance,…

Cited by 3SourcePDFScholar
2024

MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot Learning

AAAI 2024technical

Equipping a deep model the ability of few-shot learning (FSL) is a core challenge for artificial intelligence. Gradient-based meta-learning effectively addresses the challenge by learning how to learn novel tasks. Its key idea is learning a deep model in a bi-level optimization manner, where the out…

Cited by 53SourcePDFScholar
2023

PCR: Proxy-Based Contrastive Replay for Online Class-Incremental Continual Learning

CVPR 2023poster

Online class-incremental continual learning is a specific task of continual learning. It aims to continuously learn new classes from data stream and the samples of data stream are seen only once, which suffers from the catastrophic forgetting issue, i.e., forgetting historical knowledge of old class…

2022

Hyperbolic Knowledge Transfer with Class Hierarchy for Few-Shot Learning

IJCAI 2022poster

Few-shot learning (FSL) aims to recognize a novel class with very few instances, which is a challenging task since it suffers from a data scarcity issue. One way to effectively alleviate this issue is introducing explicit knowledge summarized from human past experiences to achieve knowledge transfer…

Cited by 18SourcePDFScholar
2022

MetaNODE: Prototype Optimization as a Neural ODE for Few-Shot Learning

AAAI 2022technical

Few-Shot Learning (FSL) is a challenging task, i.e., how to recognize novel classes with few examples? Pre-training based methods effectively tackle the problem by pre-training a feature extractor and then predicting novel classes via a cosine nearest neighbor classifier with mean-based prototypes.…

2022

Sentiment Interpretable Logic Tensor Network for Aspect-Term Sentiment Analysis

COLING 2022main

Aspect-term sentiment analysis (ATSA) is an important task that aims to infer the sentiment towards the given aspect-terms. It is often required in the industry that ATSA should be performed with interpretability, computational efficiency and high accuracy. However, such an ATSA method has not yet b…

Cited by 18SourcePDFScholar
2021

Prototype Completion With Primitive Knowledge for Few-Shot Learning

CVPR 2021poster

Few-shot learning is a challenging task, which aims to learn a classifier for novel classes with few examples. Pre-training based meta-learning methods effectively tackle the problem by pre-training a feature extractor and then fine-tuning it through the nearest centroid based meta-learning. However…

Cited by 161PDFcodeScholar