← Search

Dingkang Yang

51 accepted papers

2026

Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control

CVPR 2026

In computational pathology, understanding and generation have evolved along disparate paths: advanced understanding models already exhibit diagnostic-level competence, whereas generative models largely simulate pixels. Progress remains hindered by three coupled factors: the scarcity of large, high-q

Cited by 0SourcecodeScholar
2026

Diffusion Probe: Generated Image Result Prediction Using CNN Probes

CVPR 2026

Text-to-image (T2I) diffusion models currently lack an efficient mechanism for early quality assessment, forcing costly random trial-and-error in scenarios requiring multiple generations (e.g., iterating on prompts, agent-based image generation, flow-grpo). To address this, we first reveal a strong

Cited by 0SourceScholar
2026

Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models

CVPR 2026

Multimodal biomedical Vision-Language Models (VLMs) exhibit immense potential in the field of Continual Learning (CL). However, they confront a core dilemma: how to preserve fine-grained intra-modality features while bridging the significant domain gap across different modalities. To address this ch

Cited by 0SourcecodeScholar
2026

Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology

ICLR 2026poster

Recent years have witnessed remarkable progress in multimodal learning within computational pathology. Existing models primarily rely on vision and language modalities; however, language alone lacks molecular specificity and offers limited pathological supervision, leading to representational bottle…

Cited by 0SourcecodeScholar
2026

LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have shown great promise but require substantial computational resources during inference. Attackers can exploit this by inducing excessive output, leading to resource exhaustion and service degradation. Prior energy-latency attacks aim to increase generation…

Cited by 0SourceScholar
2026

MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement

AAAI 2026technical

Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers pro

Cited by 0SourcePDFScholar
2026

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-turn Dialogue

ICML 2026poster

Multimodal Large Language Models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermined by {hallucination snowballing}: a phenomenon where initial errors amplify across conversational turns, leading to a collapse in coherence. This f…

Cited by 0SourceScholar
2026

ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation (VLN) requires agents to accurately perceive complex visual environments and reason over navigation instructions and histories. However, existing methods passively process redundant visual inputs and treat all historical contexts indiscriminately, resulting in ineffici

Cited by 0SourceScholar
2026

Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding

CVPR 2026

Document understanding is a long-standing practical task. Vision-Language Models (VLMs) have gradually become a primary approach in this domain, demonstrating effective performance on single-page tasks. However, their effectiveness diminishes when handling long documents. In such scenarios, clues ar

Cited by 0SourceScholar
2026

SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension

AAAI 2026technical

Satire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particularly in purely visual forms, remains a challenging task for current vision-langu

Cited by 0SourcePDFScholar
2026

TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

CVPR 2026

Visual Text Rendering (VTR) remains a critical challenge in text-to-image generation, where even advanced models frequently produce text with structural anomalies such as distortion, blurriness, and misalignment. However, we find that leading MLLMs and specialist OCR models largely fail to perceive

Cited by 0SourcecodeScholar
2026

UniMGS: Unifying Mesh and 3D Gaussian Splatting with Single-Pass Rasterization and Proxy-Based Deformation

AAAI 2026technical

Joint rendering and deformation of mesh and 3D Gaussian Splatting (3DGS) have significant value as both representations offer complementary advantages for graphics applications. However, due to differences in representation and rendering pipelines, existing studies render meshes and 3DGS separately,

Cited by 0SourcePDFScholar
2025

AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining

NeurIPS 2025poster

Learning rate is widely regarded as crucial for effective foundation model pretraining. Recent research explores and demonstrates the transferability of learning rate configurations across varying model and dataset sizes, etc. Nevertheless, these approaches are constrained to specific training scen…

Cited by 0SourceScholar
2025

BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation

AAAI 2025technical

With the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all necessary structures. Many current 3D scene generation methods…

2025

Boosting Adversarial Transferability with Spatial Adversarial Alignment

NeurIPS 2025poster

Deep neural networks are vulnerable to adversarial examples that exhibit transferability across various models. Numerous approaches are proposed to enhance the transferability of adversarial examples, including advanced optimization, data augmentation, and model modifications. However, these methods…

Cited by 0SourceScholar
2025

Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models

EMNLP 2025

Multi-modal keyphrase prediction (MMKP) aims to advance beyond text-only methods by incorporating multiple modalities of input information to produce a set of conclusive phrases. Traditional multi-modal approaches have been proven to have significant limitations in handling the challenging absence a

2025

CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation

ICASSP 2025accepted

Automatic medical report generation (MRG), which possesses significant research value as it can aid radiologists in clinical diagnosis and report composition, has garnered increasing attention. Despite recent progress, generating accurate reports remains arduous due to the requirement for precise cl…

Cited by 0SourceScholar
2025

DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding

NeurIPS 2025poster

We introduce DanmakuTPPBench, a comprehensive benchmark designed to advance multi-modal Temporal Point Process (TPP) modeling in the era of Large Language Models (LLMs). While TPPs have been widely studied for modeling temporal event sequences, existing datasets are predominantly unimodal, hinderin…

Cited by 0SourcecodeScholar
2025

Debiased Multimodal Understanding for Human Language Sequences

AAAI 2025technical

Human multimodal language understanding (MLU) is an indispensable component of expression analysis (e.g., sentiment or humor) from heterogeneous modalities, including visual postures, linguistic contents, and acoustic behaviours. Existing works invariably focus on designing sophisticated structures…

Cited by 1SourcePDFScholar
2025

Guiding Inter-domain Class Balancing With Salient Features For Domain Adaptive Object Detection

ICASSP 2025accepted

Although multi-scale alignment has improved domain adaptive object detection by addressing data distribution differences and annotation challenges, little attention has been given to class distribution differences between domains. Additionally, the utilization of feature information across different…

Cited by 0SourceScholar
2025

Improving Factuality in Large Language Models via Decoding-Time Hallucinatory and Truthful Comparators

AAAI 2025technical

Despite their remarkable capabilities, Large Language Models (LLMs) are prone to generate responses that contradict verifiable facts, i.e., unfaithful hallucination content. Existing efforts generally focus on optimizing model parameters or editing semantic representations, which compromise the inte…

2025

MAFD: Fine-Grained Motion Style Transfer with Adaptive Signal Fusion

ICASSP 2025accepted

Motion style transfer allows for the swift switching of different styles within the same motion for virtual avatars, offering significant efficiency gains and enhanced motion diversity compared to traditional motion capture methods. However, many existing methods struggle with controlling fine detai…

Cited by 0SourceScholar
2025

MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation

CVPR 2025poster

Diffusion models have shown excellent performance in text-to-image generation. However, existing methods often suffer from performance bottlenecks when dealing with complex prompts involving multiple objects, characteristics, and relations. Therefore, we propose a Multi-agent Collaboration-based Co…

Cited by 0SourcePDFScholar
2025

MMPF: Multi-Modal Perception Framework for Abnormal Medical Condition Detection

AAAI 2025technical

As the global population ages and the incidence of chronic diseases increases, the demand for early detection of abnormal medical conditions is increasing. Traditional health monitoring methods often require significant resources and specialized personnel, limiting their widespread use. Leveraging a…

Cited by 0SourcePDFScholar
2024

A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing Modalities

AAAI 2024technical

Multimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of modality completeness. However, many inevitable factors in real-world scenarios lead to uncertain missing modalities, which invalidate the fixed multimodal fusion…

Cited by 13SourcePDFScholar
2024

Align before Collaborate: Mitigating Feature Misalignment for Robust Multi-Agent Perception

ECCV 2024oral

"Collaborative perception has received widespread attention recently since it enhances the perception ability of autonomous vehicles via inter-agent information sharing. However, the performance of existing systems is hindered by the unavoidable collaboration noises, which induce feature-level spati…

Cited by 0SourcePDFScholar
2024

CPR-Coach: Recognizing Composite Error Actions based on Single-class Training

CVPR 2024poster

Fine-grained medical action analysis plays a vital role in improving medical skill training efficiency but it faces the problems of data and algorithm shortage. Cardiopulmonary Resuscitation (CPR) is an essential skill in emergency treatment. Currently the assessment of CPR skills mainly depends on…

2024

Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities

CVPR 2024poster

Multimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However in real-world applications some practical factors cause uncertain modality missingness which drastically degrades the model's…

Cited by 15SourcePDFScholar
2024

De-confounded Data-free Knowledge Distillation for Handling Distribution Shifts

CVPR 2024poster

Data-Free Knowledge Distillation (DFKD) is a promising task to train high-performance small models to enhance actual deployment without relying on the original training data. Existing methods commonly avoid relying on private data by utilizing synthetic or sampled data. However a long-overlooked iss…

Cited by 6SourcePDFScholar
2024

HybridOcc: NeRF Enhanced Transformer-Based Multi-Camera 3D Occupancy Prediction

RA-L 2024

Vision-based 3D semantic scene completion (SSC) describes autonomous driving scenes through 3D volume representations. However, the occlusion of invisible voxels by scene surfaces poses challenges to current SSC methods in hallucinating refined 3D geometry. This letter proposes HybridOcc, a hybrid 3

Cited by 17SourceScholar
2024

Out of Thin Air: Exploring Data-Free Adversarial Robustness Distillation

AAAI 2024technical

Adversarial Robustness Distillation (ARD) is a promising task to solve the issue of limited adversarial robustness of small capacity models while optimizing the expensive computational costs of Adversarial Training (AT). Despite the good robust performance, the existing ARD methods are still impract…

Cited by 9SourcePDFScholar
2024

Pathology-knowledge Enhanced Multi-instance Prompt Learning for Few-shot Whole Slide Image Classification

ECCV 2024poster

"Current multi-instance learning algorithms for pathology image analysis often require a substantial number of Whole Slide Images for effective training but exhibit suboptimal performance in scenarios with limited learning data. In clinical settings, restricted access to pathology slides is inevitab…

Cited by 8SourcePDFScholar
2024

PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications

NeurIPS 2024poster

Developing intelligent pediatric consultation systems offers promising prospects for improving diagnostic efficiency, especially in China, where healthcare resources are scarce. Despite recent advances in Large Language Models (LLMs) for Chinese medicine, their performance is sub-optimal in pediatri…

2024

Robust Emotion Recognition in Context Debiasing

CVPR 2024poster

Context-aware emotion recognition (CAER) has recently boosted the practical applications of affective computing techniques in unconstrained environments. Mainstream CAER methods invariably extract ensemble representations from diverse contexts and subject-centred characteristics to perceive the targ…

Cited by 23SourcePDFScholar
2024

Self-Cooperation Knowledge Distillation for Novel Class Discovery

ECCV 2024poster

"Novel Class Discovery (NCD) aims to discover unknown and novel classes in an unlabeled set by leveraging knowledge already learned about known classes. Existing works focus on instance-level or class-level knowledge representation and build a shared representation space to achieve performance impro…

Cited by 4SourcePDFScholar
2024

Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning

NeurIPS 2024poster

Multimodal Sentiment Analysis (MSA) is an important research area that aims to understand and recognize human sentiment through multiple modalities. The complementary information provided by multimodal fusion promotes better sentiment analysis compared to utilizing only a single modality. Neverthele…

Cited by 1SourcePDFScholar
2024

Towards Multimodal Sentiment Analysis Debiasing via Bias Purification

ECCV 2024poster

"Multimodal Sentiment Analysis (MSA) aims to understand human intentions by integrating emotion-related clues from diverse modalities, such as visual, language, and audio. Unfortunately, the current MSA task invariably suffers from unplanned dataset biases, particularly multimodal utterance-level la…

Cited by 18SourcePDFScholar
2023

A Novel Efficient Multi-View Traffic-Related Object Detection Framework

ICASSP 2023accepted

With the rapid development of intelligent transportation system applications, a tremendous amount of multi-view video data has emerged to enhance vehicle perception. However, performing video analytics efficiently by exploiting the spatial-temporal redundancy from video data remains challenging. Acc…

Cited by 0SourceScholar
2023

AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception

ICCV 2023poster

Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIs…

Cited by 54PDFcodeScholar
2023

Adversarial Contrastive Distillation with Adaptive Denoising

ICASSP 2023accepted

Adversarial Robustness Distillation (ARD) is a novel method to boost the robustness of small models. Unlike general adversarial training, its robust knowledge transfer can be less easily restricted by the model capacity. However, the teacher model that provides the robustness of knowledge does not a…

Cited by 0SourceScholar
2023

Context De-Confounded Emotion Recognition

CVPR 2023poster

Context-Aware Emotion Recognition (CAER) is a crucial and challenging task that aims to perceive the emotional states of the target person with contextual information. Recent approaches invariably focus on designing sophisticated architectures or mechanisms to extract seemingly meaningful representa…

2023

D-CONFORMER: Deformable Sparse Transformer Augmented Convolution for Voxel-Based 3D Object Detection

ICASSP 2023accepted

Although CNN-based and Transformer-based detectors have made impressive improvements in 3D object detection, these two network paradigms suffer from the interference of insufficient receptive field and local detail weakening, which significantly limits the feature extraction performance of the backb…

Cited by 0SourceScholar
2023

Efficient Decision-based Black-box Patch Attacks on Video Recognition

ICCV 2023poster

Although Deep Neural Networks (DNNs) have demonstrated excellent performance, they are vulnerable to adversarial patches that introduce perceptible and localized perturbations to the input. Generating adversarial patches on images has received much attention, while adversarial patches on videos have…

Cited by 23PDFScholar
2023

How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent Perception

NeurIPS 2023poster

Multi-agent collaborative perception has recently received widespread attention as an emerging application in driving scenarios. Despite the advancements in previous efforts, challenges remain due to various noises in the perception procedure, including communication redundancy, transmission delay,…

2023

Improving Generalization in Visual Reinforcement Learning via Conflict-aware Gradient Agreement Augmentation

ICCV 2023poster

Learning a policy with great generalization to unseen environments remains challenging but critical in visual reinforcement learning. Despite the success of augmentation combination in the supervised learning generalization, naively applying it to visual RL algorithms may damage the training efficie…

Cited by 25PDFScholar
2023

MSN-net: Multi-Scale Normality Network for Video Anomaly Detection

ICASSP 2023accepted

Existing unsupervised video anomaly detection methods often suffer from performance degradation due to the overgeneralization of deep models. In this paper, we propose a simple yet effective Multi-Scale Normality network (MSN-net) that uses hierarchical memories to learn multi-level prototypical spa…

Cited by 0SourceScholar
2023

Spatio-Temporal Domain Awareness for Multi-Agent Collaborative Perception

ICCV 2023poster

Multi-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this em…

Cited by 68PDFcodeScholar
2023

Towards Simultaneous Segmentation Of Liver Tumors And Intrahepatic Vessels Via Cross-Attention Mechanism

ICASSP 2023accepted

Accurate visualization of liver tumors and their surrounding blood vessels is essential for noninvasive diagnosis and prognosis prediction of tumors. In medical image segmentation, there is still a lack of in-depth research on the simultaneous segmentation of liver tumors and peritumoral blood vesse…

Cited by 0SourceScholar
2022

CA-SpaceNet: Counterfactual Analysis for 6D Pose Estimation in Space

IROS 2022poster

Reliable and stable 6D pose estimation of un-cooperative space objects plays an essential role in on-orbit servicing and debris removal missions. Considering that the pose estimator is sensitive to background interference, this paper proposes a counterfactual analysis framework named CA-SpaceNet to…

Cited by 21SourcecodeScholar
2022

Emotion Recognition for Multiple Context Awareness

ECCV 2022poster

"Understanding emotion in context is a rising hotspot in the computer vision community. Existing methods lack reliable context semantics to mitigate uncertainty in expressing emotions and fail to model multiple context representations complementarily. To alleviate these issues, we present a context-…

2022

Robust Adversarial Reinforcement Learning with Dissipation Inequation Constraint

AAAI 2022technical

Robust adversarial reinforcement learning is an effective method to train agents to manage uncertain disturbance and modeling errors in real environments. However, for systems that are sensitive to disturbances or those that are difficult to stabilize, it is easier to learn a powerful adversary than…

Cited by 20SourcePDFScholar