← Search

Dong Liang

33 accepted papers

2026

Continuous Exposure-Time Modeling for Realistic Atmospheric Turbulence Synthesis

CVPR 2026

Atmospheric turbulence significantly degrades long-range imaging by introducing geometric warping and exposure-time-dependent blur, which adversely affects both visual quality and the performance of high-level vision tasks. Existing methods for synthesizing turbulence effects often oversimplify the

Cited by 0SourcecodeScholar
2026

D$^2$O: A Dual Debiasing Operator for Training-Free Test-Time Adaptation of Vision–Language Models

ICML 2026poster

Training-free test-time adaptation (TTA) for vision-language models (VLMs) can boost zero-shot classification under mild shifts but often collapses under severe environment/style shifts. We identify two shared failure modes: (i) retrieval confounding, where feature similarity is dominated by style a…

Cited by 0SourceScholar
2026

Elucidating the Design Space of Arbitrary-Noise-Based Diffusion Models

CVPR 2026

Although EDM aims to unify the design space of diffusion models, its reliance on fixed Gaussian noise prevents it from explaining emerging flow-based methods that diffuse arbitrary noise. Moreover, our study reveals that EDM's forcible injection of Gaussian noise has adverse effects on image restora

Cited by 0SourcecodeScholar
2026

Hamiltonian Asymmetric Fusion: One-Way Safe Directed Refinement under Modality Imbalance

ICML 2026poster

Multimodal fusion is commonly implemented via symmetric token interaction, implicitly allowing information to flow in both directions. Under *modality imbalance*---when an auxiliary stream is substantially noisier than a designated primary stream---such symmetry creates a *backflow channel* that inj…

Cited by 0SourceScholar
2026

LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review Platform

ICML 2026poster

Literature reviews are essential to reflect the landscape of research fields. Large language models, especially deep research agents, have recently shown strong capabilities in automated literature review generation. However, it remains a challenging task to rigorously evaluate the scientific value …

Cited by 0SourceScholar
2026

MMRQA: SIGNAL-ENHANCED MULTIMODAL LARGE LANGUAGE MODELS FOR MRI QUALITY ASSESSMENT

ICASSP 2026poster

Magnetic resonance imaging (MRI) quality assessment is crucial for clinical decision-making, yet remains challenging due to data scarcity and protocol variability. Traditional approaches face fundamental trade-offs: signal-based methods like MRIQC provide quantitative metrics but lack semantic under…

Cited by 0SourcePDFScholar
2026

MultiMedBench: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQA

AAAI 2026technical

Knowledge editing (KE) provides a scalable approach for updating factual knowledge in large language models without full retraining. While previous studies have demonstrated effectiveness in general domains and medical QA tasks, little attention has been paid to KE in multimodal medical scenarios. U

Cited by 0SourcePDFScholar
2026

Resolving the Timestep Scaling Paradox in Spiking Neural Networks with a Timestep-Scalable Neuron Model

ICML 2026poster

Spiking Neural Networks (SNNs) have garnered increasing attention for their biological plausibility, energy efficiency, and temporal modeling capability. Due to the non-differentiability of spike generation, a widely used supervised training method for SNNs is backpropagation through time with surro…

Cited by 0SourceScholar
2026

World-Shaper: A Unified Framework for 360° Panoramic Editing

ICML 2026poster

Being able to edit panoramic images is crucial for creating realistic 360° visual experiences. However, existing perspective-based image editing methods fail to model the spatial structure of panoramas. Conventional cube-map decompositions attempt to overcome this problem but inevitably break global…

Cited by 0SourceScholar
2026

Zero-shot Implicit Neural Manifold Representation (INMR) for Ultra-high Temporal Resolution Dynamic MRI

AAAI 2026technical

Capturing accurate dynamic information of moving organs is essential for functional assessment using non-invasive imaging modalities. Achieving high temporal resolution visualization of physiological processes remains a critical challenge in dynamic magnetic resonance imaging (MRI) when reconstructi

Cited by 0SourcePDFScholar
2025

DualCnst: Enhancing Zero-Shot Out-of-Distribution Detection via Text-Image Consistency in Vision-Language Models

NeurIPS 2025poster

Pretrained vision-language models (VLMs), such as CLIP, have shown promising zero-shot out-of-distribution (OOD) detection capabilities by leveraging semantic similarities between input images and textual labels. However, most existing approaches focus solely on expanding the label space in the text…

Cited by 0SourceScholar
2025

Finding Local Diffusion Schrodinger Bridge using Kolmogorov-Arnold Network

CVPR 2025poster

In image generation, Schrodinger Bridge (SB)-based methods theoretically enhance the efficiency and quality compared to the diffusion models by finding the least costly path between two distributions. However, they are computationally expensive and time-consuming when applied to complex image data.…

2025

InterGSEdit: Interactive 3D Gaussian Splatting Editing with 3D Geometry-Consistent Attention Prior

ICCV 2025poster

3D Gaussian Splatting based 3D editing has demonstrated impressive performance in recent years. However, the multi-view editing often exhibits significant local inconsistency, especially in areas of non-rigid deformation, which lead to local artifacts, texture blurring, or semantic variations in edi…

Cited by 0SourcePDFScholar
2025

LoD: Loss-difference OOD Detection by Intentionally Label-Noisifying Unlabeled Wild Data

IJCAI 2025

Using unlabeled wild data containing both in-distribution (ID) and out-of-distribution (OOD) data to improve the safety and reliability of models has recently received increasing attention. Existing methods either design customized losses for labeled ID and unlabeled wild data then perform joint opt

2025

StructSR: Refuse Spurious Details in Real-World Image Super-Resolution

AAAI 2025technical

Diffusion-based models have shown great promise in real-world image super-resolution (Real-ISR), but often generate content with structural errors and spurious texture details due to the empirical priors and illusions of these models. To address this issue, we introduce StructSR, a simple, effective…

2024

Coalition Formation Game Approach for Task Allocation in Heterogeneous Multi-Robot Systems under Resource Constraints

IROS 2024poster

This paper studies a case of the multi-robot task allocation (MRTA) problem, where each unmanned aerial vehicle (UAV) is endowed with multiple but limited resources. Completing each task necessitates UAVs to combine different resources through coalition formation, which will incur various costs incl…

Cited by 0SourceScholar
2024

G2L-CariGAN: Caricature Generation from Global Structure to Local Features

AAAI 2024technical

Existing GAN-based approaches to caricature generation mainly focus on exaggerating a character’s global facial structure. This often leads to the failure in highlighting significant facial features such as big eyes and hook nose. To address this limitation, we propose a new approach termed as G2L-C…

Cited by 0SourcePDFScholar
2024

Theoretical Investigations and Practical Enhancements on Tail Task Risk Minimization in Meta Learning

NeurIPS 2024poster

Meta learning is a promising paradigm in the era of large models and task distributional robustness has become an indispensable consideration in real-world scenarios. Recent advances have examined the effectiveness of tail task risk minimization in fast adaptation robustness improvement \citep{wang…

2023

ALL-E: Aesthetics-guided Low-light Image Enhancement

IJCAI 2023poster

Evaluating the performance of low-light image enhancement (LLE) is highly subjective, thus making integrating human preferences into image enhancement a necessity. Existing methods fail to consider this and present a series of potentially valid heuristic criteria for training enhancement models. In…

2023

Improving Lens Flare Removal with General-Purpose Pipeline and Multiple Light Sources Recovery

ICCV 2023poster

When taking images against strong light sources, the resulting images often contain heterogeneous flare artifacts. These artifacts can importantly affect image visual quality and downstream computer vision tasks. While collecting real data pairs of flare-corrupted/flare-free images for training flar…

Cited by 26PDFcodeScholar
2023

MSFORMER: Multi-Scale Transformer with Neighborhood Consensus for Feature Matching

ICASSP 2023accepted

Existing feature matching methods tend to extract feature descriptors by feeding down-sampled feature maps into a Transformer that is unable to extend feature scales, leading to false correspondences between small-size objects. This paper proposes MSFormer, which uses Transformers situated in differ…

Cited by 0SourceScholar
2023

Spammer Detection on Short Video Applications: A new Challenge and Baselines

ICASSP 2023accepted

Users can interact with the advertisements and share their impressions through the review system on short video applications. However, spammers may post false or malicious comments to mislead normal users due to profit-driven reasons, damaging the community’s positive atmosphere. In this paper, we i…

Cited by 0SourceScholar
2022

I Can Find You! Boundary-Guided Separated Attention Network for Camouflaged Object Detection

AAAI 2022technical

Can you find me? By simulating how humans to discover the so-called 'perfectly'-camouflaged object, we present a novel boundary-guided separated attention network (call BSA-Net). Beyond the existing camouflaged object detection (COD) wisdom, BSA-Net utilizes two-stream separated attention modules to…

2022

MBA-RainGAN: A Multi-Branch Attention Generative Adversarial Network for Mixture of Rain Removal

ICASSP 2022accepted

Rain severely degrades the visibility of scene objects, especially when images are captured through the glass under rainy weather. We observe three intriguing phenomena: 1) rain is a mixture of raindrops, rain streaks and rainy haze; 2) the depth from the camera determines the degree of object visib…

Cited by 0SourceScholar
2022

Semantically Contrastive Learning for Low-Light Image Enhancement

AAAI 2022technical

Low-light image enhancement (LLE) remains challenging due to the unfavorable prevailing low-contrast and weak-visibility problems of single RGB images. In this paper, we respond to the intriguing learning-related question -- if leveraging both accessible unpaired over/underexposed images and high-le…

2021

Cross Scene Video Foreground Segmentation Via Co-Occurrence Probability Oriented Supervised and Unsupervised Model Interaction

ICASSP 2021accepted

Using only one deep model for cross scene video foreground segmentation is still very challenging because existing methods are scene-dependent, which restricts the consistent segmentation. In this paper, we propose a cross scene video foreground segmentation framework to extend the generalization ca…

Cited by 0SourceScholar
2021

Nlkd: Using Coarse Annotations For Semantic Segmentation Based on Knowledge Distillation

ICASSP 2021accepted

Modern supervised learning relies on a large amount of training data, yet there are many noisy annotations in real datasets. For semantic segmentation tasks, pixel-level annotation noise is typically located at the edge of an object, while pixels within objects are fine-annotated. We argue the coars…

Cited by 0SourceScholar
2021

Robust Spatial-Temporal Correlation Model for Background Initialization in Severe Scene

ICASSP 2021accepted

Scene background initialization is an important step as one low-layer method for high-layer applications in computer vision. However, this process is always affected by practical challenges such as illumination changes, back-ground motion, camera jitter, intermittent movement and bad weather outdoor…

Cited by 0SourceScholar
2019

Score-specific Non-maximum Suppression and Coexistence Prior for Multi-scale Face Detection

ICASSP 2019accepted

Face detection is an ultimate component to support various visual facial related tasks. However, detecting faces with extremely low resolution or high occlusion is still an open problem. In this paper, we propose a two-step general approach to refine the performance of modern face detectors accordin…

Cited by 0SourceScholar
2017

Tensor RPCA by Bayesian CP Factorization With Complex Noise

ICCV 2017poster

The RPCA model has achieved good performances in various applications. However, two defects limit its effectiveness. Firstly, it is designed for dealing with data in matrix form, which fails to exploit the structure information of higher order tensor data in some pratical situations. Secondly, it ad…

Cited by 23PDFScholar