← Search

Kang Zhang

15 accepted papers

2026

A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers

ICLR 2026poster

Diffusion Transformers have achieved state-of-the-art performance in class-conditional and multimodal generation, yet the structure of their learned conditional embeddings remains poorly understood. In this work, we present the first systematic study of these embeddings and uncover a notable redunda…

Cited by 0SourceScholar
2025

FinLLM-B: When Large Language Models Meet Financial Breakout Trading

NAACL 2025industry

Trading range breakout is a key method in the technical analysis of financial trading, widely employed by traders in financial markets such as stocks, futures, and foreign exchange. However, distinguishing between true and false breakout and providing the correct rationale cause significant challeng…

Cited by 0SourcePDFScholar
2025

Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation

NeurIPS 2025poster

We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guid…

Cited by 0SourcecodeScholar
2025

Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision

NeurIPS 2025poster

Distinguishing visually similar objects by their motion remains a critical challenge in computer vision. Although supervised trackers show promise, contemporary self-supervised trackers struggle when visual cues become ambiguous, limiting their scalability and generalization without extensive labele…

Cited by 0SourceScholar
2024

BI-MDRG: Bridging Image History in Multimodal Dialogue Response Generation

ECCV 2024poster

"Multimodal Dialogue Response Generation (MDRG) is a recently proposed task where the model needs to generate responses in texts, images, or a blend of both based on the dialogue context. Due to the lack of a large-scale dataset specifically for this task and the benefits of leveraging powerful pre-…

2024

Cross-view Masked Diffusion Transformers for Person Image Synthesis

ICML 2024poster

We present X-MDPT ($\underline{Cross}$-view $\underline{M}$asked $\underline{D}$iffusion $\underline{P}$rediction $\underline{T}$ransformers), a novel diffusion model designed for pose-guided human image generation. X-MDPT distinguishes itself by employing masked diffusion transformers that operate…

2023

Flow-Guided Deformable Alignment Network with Self-Supervision for Video Inpainting

ICASSP 2023accepted

Video inpainting aims to utilize plausible contents to fill missing regions in the video. State-of-the-art video inpainting methods typically generate the missing contents of the target frame (current frame) by aggregating the temporal information of reference frames (neighboring frames) aligned usi…

Cited by 0SourceScholar
2023

Infrared and Visible Image Fusion by Using Multi-Scale Transformation and Fractional-Order Gradient Information

ICASSP 2023accepted

The fusion of infrared and visible images is hard due to their different modalities. Different from existing methods using the integer-order gradient, we design an optimization model to fuse infrared and visible images using fractional-order gradient information. In this way, the complementary infor…

Cited by 0SourceScholar
2023

Semi-Supervised Video Inpainting With Cycle Consistency Constraints

CVPR 2023poster

Deep learning-based video inpainting has yielded promising results and gained increasing attention from researchers. Generally, these methods usually assume that the corrupted region masks of each frame are known and easily obtained. However, the annotation of these masks are labor-intensive and exp…

Cited by 18SourcePDFScholar
2022

Decoupled Adversarial Contrastive Learning for Self-Supervised Adversarial Robustness

ECCV 2022poster

"\textit{Adversarial training} (AT) for robust representation learning and \textit{self-supervised learning} (SSL) for unsupervised representation learning are two active research fields. Integrating AT into SSL, multiple prior works have accomplished a highly significant yet challenging task: learn…

2022

Dual Temperature Helps Contrastive Learning Without Many Negative Samples: Towards Understanding and Simplifying MoCo

CVPR 2022poster

Contrastive learning (CL) is widely known to require many negative samples, 65536 in MoCo for instance, for which the performance of a dictionary-free framework is often inferior because the negative sample size (NSS) is limited by its mini-batch size (MBS). To decouple the NSS from the MBS, a dynam…

Cited by 59PDFcodeScholar
2022

How Does SimSiam Avoid Collapse Without Negative Samples? A Unified Understanding with Self-supervised Contrastive Learning

ICLR 2022poster

To avoid collapse in self-supervised learning (SSL), a contrastive loss is widely used but often requires a large number of negative samples. Without negative samples yet achieving competitive performance, a recent work~\citep{chen2021exploring} has attracted significant attention for providing a mi…

Cited by 98SourcePDFScholar
2022

Investigating Top-k White-Box and Transferable Black-Box Attack

CVPR 2022poster

Existing works have identified the limitation of top-1 attack success rate (ASR) as a metric to evaluate the attack strength but exclusively investigated it in the white-box setting, while our work extends it to a more practical black-box setting: transferable attack. It is widely reported that stro…

Cited by 51PDFcodeScholar
2015

3D Fragment Reassembly Using Integrated Template Guidance and Fracture-Region Matching

ICCV 2015poster

This paper studies matching of fragmented objects to recompose their original geometry. Solving this geometric reassembly problem has direct applications in archaeology and forensic investigation in the computer-aided restoration of damaged artifacts and evidence. We develop a new algorithm to effec…

Cited by 86PDFScholar