← Search

Feng Yang

34 accepted papers

2026

Domain-Aware Multi-View Contrastive Representation Learning for Protein Subcellular Localization Prediction

AAAI 2026technical

Protein subcellular localization prediction is essential for understanding protein function and cellular organization. However, existing methods exhibit two major limitations: (1) they overlook the critical role of evolutionarily conserved protein domains, which are fundamental functional and struct

Cited by 0SourcePDFScholar
2026

Dynamic Geometric Equivariant Network for Full-Atom Antibody Design

AAAI 2026technical

Antibody design is critically important in biomedical and therapeutic contexts but remains extremely challenging due to the complexity of antibody sequence–structure relationships and stringent antigen specificity requirements. Traditional computational approaches rely on multi-stage pipelines and o

Cited by 0SourcePDFScholar
2026

Injection Without Distortion: Geometrically Constrained Knowledge Enhancement for Vision-Language Models

AAAI 2026technical

Vision-Language Models (VLMs) are widely used in tasks like Open-Vocabulary Object Detection and zero-shot Classification, owing to their powerful generalization. However, recent research reveals that VLMs exhibit significant performance instability when tasked with recognizing concepts at varying g

Cited by 0SourcePDFScholar
2026

MPL: Match-guided Prototype Learning for Few-shot Action Recognition

CVPR 2026

Current few-shot action recognition methods achieve impressive performance by learning representative prototypes and designing diverse video matching strategies. However, these approaches typically face two critical limitations: i) prototypes learned through implicit sample interactions lack clear s

Cited by 0SourcecodeScholar
2026

OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

AAAI 2026technical

We present OpenDriveVLA, a Vision-Language Action (VLA) model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially-grounded driving actions by leveraging multimodal inputs, including both 2D and 3D instance-aware visual representa

Cited by 0SourcePDFScholar
2026

Unlearning without Forgetting: Securely Removing Targeted Concepts from Large-Scale Vision-Language Open-Vocabulary Detectors

CVPR 2026

Open-vocabulary detectors (OvOD) inherit tightly coupled cross-modal knowledge from web-scale pretraining, creating privacy, copyright, and compliance risks. Existing machine unlearning methods face geometric entanglement interference in OvOD: forgetting updates inevitably distort preserved knowledg

Cited by 0SourceScholar
2026

WaTeRFlow: Watermark Temporal Robustness via Flow Consistency

CVPR 2026

Image watermarking supports authenticity and provenance, yet many schemes are still easy to bypass with various distortions and powerful generative edits. Deep learning-based watermarking has improved robustness to diffusion-based image editing, but a gap remains when a watermarked image is converte

Cited by 0SourceScholar
2025

3D-GSW: 3D Gaussian Splatting for Robust Watermarking

CVPR 2025poster

As 3D Gaussian Splatting (3D-GS) gains significant attention and its commercial usage increases, the need for watermarking technologies to prevent unauthorized use of the 3D-GS models and rendered images has become increasingly important. In this paper, we introduce a robust watermarking method for…

2025

Calibrated Multi-Preference Optimization for Aligning Diffusion Models

CVPR 2025poster

Aligning text-to-image (T2I) diffusion models with prefer-ence optimization is valuable for human-annotated datasets, but the heavy cost of manual data collection limits scalability. Using reward models offers an alternative, however, current preference optimization methods fall short in exploiting…

Cited by 5SourcePDFScholar
2025

Collaborative Dual-Branch Spatial-Frequency Enhancement Network for Low-Light Images

ICASSP 2025accepted

Low-light images are commonly present due to imaging factors such as insufficient light, night shooting and back lit. Existing low-light image enhancement (LLIE) methods typically rely on a low-light input image for enhancement, which seldom leverage information contained in its high-light counterpa…

Cited by 0SourceScholar
2025

CounterPC: Counterfactual Feature Realignment for Unsupervised Domain Adaptation on Point Clouds

ICCV 2025poster

Understanding real-world 3D point clouds is challenging due to domain shifts, causing geometric variations like density changes, noise, and occlusions. The key challenge is disentangling domain-invariant semantics from domain-specific geometric variations, as point clouds exhibit local inconsistency…

Cited by 0SourcePDFScholar
2025

Cropper: Vision-Language Model for Image Cropping through In-Context Learning

CVPR 2025poster

The goal of image cropping is to identify visually appealing crops in an image. Conventional methods are trained on specific datasets and fail to adapt to new requirements. Recent breakthroughs in large vision-language models (VLMs) enable visual in-context learning without explicit training. Howeve…

Cited by 2SourcePDFScholar
2025

Debiased Prototype Evolving for Point Cloud Domain Adaptation via 3D Foundation Models

ICASSP 2025accepted

Domain adaptation in point cloud data is essential for improving downstream tasks in autonomous driving, robotics, and 3D modeling. 3D Foundation models, driven by scaling laws, have significantly advanced point cloud applications by embedding rich semantic knowledge of geometric structures. However…

Cited by 0SourceScholar
2025

Design and Development of a GPR-Equipped Robot for Full-space External Diseases Detection in Drainage Pipelines*

IROS 2025

Soil diseases around drainage pipelines are a major factor in road collapse. Robots designed to detect these diseases face multiple challenges, including harsh internal environments, size limitations, difficulties in achieving full external space coverage, and the impact of pose misalignment on dise

Cited by 0SourceScholar
2025

Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation

CVPR 2025highlight

Text-to-image (T2I) generation has made significant advances in recent years, but challenges still remain in the generation of perceptual artifacts, misalignment with complex prompts, and safety. The prevailing approach to address these issues involves collecting human feedback on generated images,…

Cited by 2SourcePDFScholar
2025

Motion Artifact Removal in Pixel-Frequency Domain via Alternate Masks and Diffusion Model

AAAI 2025technical

Motion artifacts present in magnetic resonance imaging (MRI) can seriously interfere with clinical diagnosis. Removing motion artifacts is a straightforward solution and has been extensively studied. However, paired data are still heavily relied on in recent works and the perturbations in k-space (f…

2024

Optical Diffusion Models for Image Generation

NeurIPS 2024poster

Diffusion models generate new samples by progressively decreasing the noise from the initially provided random distribution. This inference procedure generally utilizes a trained neural network numerous times to obtain the final output, creating significant latency and energy consumption on digital…

Cited by 2SourcePDFScholar
2024

Rich Human Feedback for Text-to-Image Generation

CVPR 2024poster

Recent Text-to-Image (T2I) generation models such as Stable Diffusion and Imagen have made significant progress in generating high-resolution images based on text descriptions. However many generated images still suffer from issues such as artifacts/implausibility misalignment with text descriptions…

2024

WateRF: Robust Watermarks in Radiance Fields for Protection of Copyrights

CVPR 2024poster

The advances in the Neural Radiance Fields (NeRF) research offer extensive applications in diverse domains but protecting their copyrights has not yet been researched in depth. Recently NeRF watermarking has been considered one of the pivotal solutions for safely deploying NeRF-based 3D representati…

2023

Re-mine, Learn and Reason: Exploring the Cross-modal Semantic Correlations for Language-guided HOI detection

ICCV 2023poster

Human-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predict <human, action, object> triplets. Despite the challenges posed by the numerous interaction combinations, t…

Cited by 31PDFScholar
2023

SVDiff: Compact Parameter Space for Diffusion Fine-Tuning

ICCV 2023poster

Recently, diffusion models have achieved remarkable success in text-to-image generation, enabling the creation of high-quality images from text prompts and various conditions. However, existing methods for customizing these models are limited by handling multiple personalized subjects and the risk o…

Cited by 271PDFScholar
2023

VILA: Learning Image Aesthetics From User Comments With Vision-Language Pretraining

CVPR 2023poster

Assessing the aesthetics of an image is challenging, as it is influenced by multiple factors including composition, color, style, and high-level semantics. Existing image aesthetic assessment (IAA) methods primarily rely on human-labeled rating scores, which oversimplify the visual aesthetic informa…

2022

Deep 3D-to-2D Watermarking: Embedding Messages in 3D Meshes and Extracting Them From 2D Renderings

CVPR 2022poster

Digital watermarking is widely used for copyright protection. Traditional 3D watermarking approaches or commercial software are typically designed to embed messages into 3D meshes, and later retrieve the messages directly from distorted/undistorted watermarked 3D meshes. However, in many cases, user…

Cited by 41PDFScholar
2022

MAXIM: Multi-Axis MLP for Image Processing

CVPR 2022oral

Recent progress on Transformers and multi-layer perceptron (MLP) models provide new network architectural designs for computer vision tasks. Although these models proved to be effective in many vision tasks such as image recognition, there remain challenges in adapting them for low-level vision. The…

Cited by 624PDFcodeScholar
2022

MaxViT: Multi-axis Vision Transformer

ECCV 2022poster

"Transformers have recently gained significant attention in the computer vision community. However, the lack of scalability of self-attention mechanisms with respect to image size has limited their wide adoption in state-of-the-art vision backbones. In this paper we introduce an efficient and scalab…

2021

Adversarially Adaptive Normalization for Single Domain Generalization

CVPR 2021poster

Single domain generalization aims to learn a model that performs well on many unseen domains with only one domain data for training. Existing works focus on studying the adversarial domain augmentation (ADA) to improve the model's generalization capability. The impact on domain generalization from t…

Cited by 161PDFScholar
2021

COMISR: Compression-Informed Video Super-Resolution

ICCV 2021poster

Most video super-resolution methods focus on restoring high-resolution video frames from low-resolution videos without taking into account compression. However, most videos on the web or mobile devices are compressed, and the compression can be severe when the bandwidth is limited. In this paper, we…

Cited by 48PDFcodeScholar
2021

Rich Features for Perceptual Quality Assessment of UGC Videos

CVPR 2021poster

Video quality assessment for User Generated Content (UGC) is an important topic in both industry and academia. Most existing methods only focus on one aspect of the perceptual quality assessment, such as technical quality or compression artifacts. In this paper, we create a large scale dataset to co…

Cited by 105PDFScholar
2016

Cluster-based dictionary learning and locality-constrained sparse reconstruction for trajectory classification

ICASSP 2016accepted

Trajectory classification has been extensively investigated in recent years, however, the problems about automatically modeling unlabeled and incomplete trajectories are far from being solved. In this paper, we propose a Cluster-based Dictionary Learning (CDL) approach that firstly constructs an ini…

Cited by 0SourceScholar
2016

Fast intra mode decision and block matching for HEVC screen content compression

ICASSP 2016accepted

Screen content coding (SCC) is the latest extension of the High-Efficiency Video Coding (HEVC) aiming to improve the compression efficiency of screen content video. With newly developed tools such as intra block copy (IntraBC) and palette (PLT) mode, SCC has been able to compress the desktop screens…

Cited by 0SourceScholar