← Search

Adams Wai-Kin Kong

21 accepted papers

2026

**TandemFoilSet**: Datasets for Flow Field Prediction of Tandem-Airfoil Through the Reuse of Single Airfoils

ICLR 2026poster

Accurate simulation of flow fields around tandem geometries is critical for engineering design but remains computationally intensive. Existing machine learning approaches typically focus on simpler cases and lack evaluation on multi-body configurations. To support research in this area, we present *…

Cited by 0SourceScholar
2026

All That Glitters Is Not Gold: Key-Secured 3D Secrets within 3D Gaussian Splatting

ICLR 2026poster

Recent advances in 3D Gaussian Splatting (3DGS) have revolutionized scene reconstruction, opening new possibilities for 3D steganography by hiding 3D secrets within 3D covers. The key challenge in steganography is ensuring imperceptibility while maintaining high-fidelity reconstruction. However, exi…

Cited by 0SourcecodeScholar
2026

Does FLUX Already Know How to Perform Physically Plausible Image Composition?

ICLR 2026poster

Image composition aims to seamlessly insert a user-specified object into a new scene, but existing models struggle with complex lighting (e.g., accurate shadows, water reflections) and diverse, high-resolution inputs. Modern text-to-image diffusion models (e.g., SD3.5, FLUX) already encode essential…

Cited by 0SourcecodeScholar
2026

DragFlow: Unleashing DiT Priors with Region-Based Supervision for Drag Editing

ICLR 2026poster

Drag-based image editing has long suffered from distortions in the target region, largely because the priors of earlier base models, Stable Diffusion, are insufficient to project optimized latents back onto the natural image manifold. With the shift from UNet-based DDPMs to more scalable DiT with fl…

Cited by 0SourcecodeScholar
2026

MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic Alignment

AAAI 2026technical

Propelled by the breakthrough in deep generative models, audio-to-image generation has emerged as a pivotal cross-modal task that converts complex auditory signals into rich visual representations. However, previous works only focus on single-source audio inputs for image generation, ignoring the mu

Cited by 0SourcePDFScholar
2026

Test-Time Adaptation without Source Data for Out-of-Domain Bioactivity Prediction

ICLR 2026poster

Accurate prediction of protein-ligand bioactivity is a cornerstone of modern drug discovery, yet current deep learning methods often struggle with out-of-domain (OOD) generalization. The existing methods rely on access to source data, making them impractical in scenarios where data cannot be accesse…

Cited by 0SourceScholar
2025

Decreasing Word Error Rates in Paragraph Handwritten Text Recognition with Synthetic Data

ICASSP 2025accepted

Handwritten Text Recognition (HTR) faces a persistent challenge with the scarcity of data at the paragraph level, arising from the difficulty of acquiring diverse, cost-efficient, and cleanly labeled datasets for training. As such, works in HTR leverage segmentation, regularization techniques, and l…

Cited by 0SourceScholar
2025

Easing Training Process of Rectified Flow Models Via Lengthening Inter-Path Distance

ICLR 2025spotlight

Recent research pinpoints that different diffusion methods and architectures trained on the same dataset produce similar results for the same input noise. This property suggests that they have some preferable noises for a given sample. By visualizing the noise-sample pairs of rectified flow model…

Cited by 0SourcePDFScholar
2025

Enhancing Bioactivity Prediction via Spatial Emptiness Representation of Protein-ligand Complex and Union of Multiple Pockets

NeurIPS 2025poster

Predicting the bioactivity of candidate ligands remains a central challenge in drug discovery. Ligands and endogenous substrates often compete for the same binding sites on target proteins, and the extent to which a ligand can modulate protein function depends not only on its binding but also on how…

Cited by 0SourceScholar
2025

Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances

ICLR 2025poster

Current image watermarking methods are vulnerable to advanced image editing techniques enabled by large-scale text-to-image models. These models can distort embedded watermarks during editing, posing significant challenges to copyright protection. In this work, we introduce W-Bench, the first compre…

2025

When and Where do Data Poisons Attack Textual Inversion?

ICCV 2025poster

Poisoning attacks pose significant challenges to the robustness of diffusion models (DMs). In this paper, we systematically analyze when and where poisoning attacks textual inversion (TI), a widely used personalization technique for DMs. We first introduce Semantic Sensitivity Maps, a novel method f…

2024

Finite Volume Features, Global Geometry Representations, and Residual Training for Deep Learning-based CFD Simulation

ICML 2024spotlight

Computational fluid dynamics (CFD) simulation is an irreplaceable modelling step in many engineering designs, but it is often computationally expensive. Some graph neural network (GNN)-based CFD methods have been proposed. However, the current methods inherit the weakness of traditional numerical si…

2024

MACE: Mass Concept Erasure in Diffusion Models

CVPR 2024poster

The rapid expansion of large-scale text-to-image diffusion models has raised growing concerns regarding their potential misuse in creating harmful or misleading content. In this paper we introduce MACE a finetuning framework for the task of mass concept erasure. This task aims to prevent models from…

2023

Audio-Visual Deception Detection: DOLOS Dataset and Parameter-Efficient Crossmodal Learning

ICCV 2023poster

Deception detection in conversations is a challenging yet important task, having pivotal applications in many fields such as credibility assessment in business, multimedia anti-frauds, and custom security. Despite this, deception detection research is hindered by the lack of high-quality deception d…

Cited by 12PDFcodeScholar
2022

Exploiting the Relationship Between Kendall's Rank Correlation and Cosine Similarity for Attribution Protection

NeurIPS 2022accept

Model attributions are important in deep neural networks as they aid practitioners in understanding the models, but recent studies reveal that attributions can be easily perturbed by adding imperceptible noise to the input. The non-differentiable Kendall's rank correlation is a key performance index…

Cited by 14SourcePDFScholar
2022

Pure Transformer with Integrated Experts for Scene Text Recognition

ECCV 2022poster

"Scene text recognition (STR) involves the task of reading text in cropped images of natural scenes. Conventional models in STR employ convolutional neural network (CNN) followed by recurrent neural network in an encoder-decoder framework. In recent times, the transformer architecture is being widel…

Cited by 29SourcePDFScholar