← Search

Jian Jin

13 accepted papers

2026

Hashed Watermark as a Filter: A Unified Defense Against Forging and Overwriting Attacks in Neural Network Watermarking

AAAI 2026technical

As valuable digital assets, deep neural networks necessitate robust ownership protection, positioning neural network watermarking (NNW) as a promising solution. Among various NNW approaches, weight-based methods are favored for their simplicity and practicality; however, they remain generally vulne

Cited by 0SourcePDFScholar
2026

QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality Assessment

CVPR 2026

No-Reference Point Cloud Quality Assessment (NR-PCQA) still struggles with generalization, primarily due to the scarcity of annotated point cloud datasets. Since the Human Visual System (HVS) drives perceptual quality assessment independently of media types, prior knowledge on quality learned from i

Cited by 0SourcecodeScholar
2026

R4-CGQA: Retrieval-based Vision Language Models for Computer Graphics Image Quality Assessment

CVPR 2026

Immersive Computer Graphics (CGs) rendering has become ubiquitous in modern daily life. However, comprehensively evaluating CG quality remains challenging for two reasons: (1) existing CG datasets lack systematic descriptions of rendering quality; and (2) existing CG quality assessment methods canno

Cited by 0SourcecodeScholar
2026

The Last Byte: Learning Just Enough for Machine-Oriented Image Compression

AAAI 2026technical

Just recognizable distortion (JRD) has been introduced for image compression for machines, aiming to quantify the maximum coding distortion that can be tolerated by a specific perception model, thereby defining the upper bound of machine vision redundancy (MVR). However, existing JRD-based redundanc

Cited by 0SourcePDFScholar
2025

Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization

ICML 2025poster

Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established in NLP, such as large transformer models, extensive pre-training, and the auto-regression paradigm, among others. In thi…

2025

LaTexBlend: Scaling Multi-concept Customized Generation with Latent Textual Blending

CVPR 2025highlight

Customized text-to-image generation renders user-specified concepts into novel contexts based on textual prompts. Scaling the number of concepts in customized generation meets a broader demand for user creation, whereas existing methods face challenges with generation quality and computational effic…

Cited by 1SourcePDFScholar
2025

Twofold Debiasing Enhances Fine-Grained Learning with Coarse Labels

AAAI 2025technical

The Coarse-to-Fine Few-Shot (C2FS) task is designed to train models using only coarse labels, then leverages a limited number of subclass samples to achieve fine-grained recognition capabilities. This task presents two main challenges: coarse-grained supervised pre-training suppresses the extraction…

2024

Customized Generation Reimagined: Fidelity and Editability Harmonized

ECCV 2024poster

"Customized generation aims to incorporate a novel concept into a pre-trained text-to-image model, enabling new generations of the concept in novel contexts guided by textual prompts. However, customized generation suffers from an inherent trade-off between concept fidelity and editability, i.e., be…

2024

Rethinking Guidance Information to Utilize Unlabeled Samples: A Label Encoding Perspective

ICML 2024poster

Empirical Risk Minimization (ERM) is fragile in scenarios with insufficient labeled samples. A vanilla extension of ERM to unlabeled samples is Entropy Minimization (EntMin), which employs the soft-labels of unlabeled samples to guide their learning. However, EntMin emphasizes prediction discriminab…

2024

Semantic Lens: Instance-Centric Semantic Alignment for Video Super-resolution

AAAI 2024technical

As a critical clue of video super-resolution (VSR), inter-frame alignment significantly impacts overall performance. However, accurate pixel-level alignment is a challenging task due to the intricate motion interweaving in the video. In response to this issue, we introduce a novel paradigm for VSR n…

2024

Towards Surveillance Video-and-Language Understanding: New Dataset Baselines and Challenges

CVPR 2024poster

Surveillance videos are important for public security. However current surveillance video tasks mainly focus on classifying and localizing anomalous events. Existing methods are limited to detecting and classifying the predefined events with unsatisfactory semantic understanding although they have o…

Cited by 18SourcePDFScholar
2023

A Lightweight Fourier Convolutional Attention Encoder for Multi-Channel Speech Enhancement

ICASSP 2023accepted

Beamforming weights prediction via deep neural networks has been one of the main methods in multi-channel speech enhancement tasks. The spectral-spatial cues are crucial in beamforming weights estimation, however, many existing works fail to optimally predict the beamforming weights with an absence…

Cited by 0SourceScholar
2022

CUP: Curriculum Learning based Prompt Tuning for Implicit Event Argument Extraction

IJCAI 2022poster

Implicit event argument extraction (EAE) aims to identify arguments that could scatter over the document. Most previous work focuses on learning the direct relations between arguments and the given trigger, while the implicit relations with long-range dependency are not well studied. Moreover, recen…