← Search

Dongbao Yang

12 accepted papers

2026

Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach

ICLR 2026poster

Recently, Multimodal Large Language Models (MLLMs) have achieved exceptional performance across diverse tasks, continually surpassing previous expectations regarding their capabilities. Nevertheless, their proficiency in perceiving emotions from images remains debated, with studies yielding divergen…

Cited by 0SourcecodeScholar
2026

Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning

ICLR 2026poster

The maturation of Large Audio Language Models (LALMs) has raised growing expectations for them to comprehend complex audio much like humans. Current efforts primarily replicate text-based reasoning by contextualizing audio content through a one-time encoding, which introduces a critical information…

Cited by 0SourcecodeScholar
2026

Focus, Align, and Sustain: Counteracting Gradient Dilution in Incremental Object Detection

ICML 2026poster

Adapting Detection Transformers to Incremental Object Detection (IOD) poses a systemic challenge, as set-based optimization is inherently destabilized by sequential learning. In this work, we identify Gradient Dilution as the root cause of performance degradation, wherein optimization signals requir…

Cited by 0SourceScholar
2025

An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability

ICML 2025poster

The advancements in Multimodal Large Language Models (MLLMs) have enabled various multimodal tasks to be addressed under a zero-shot paradigm. This paradigm sidesteps the cost of model fine-tuning, emerging as a dominant trend in practical application. Nevertheless, Multimodal Sentiment Analysis (MS…

Cited by 0SourcePDFScholar
2025

Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance

AAAI 2025technical

Scene text spotting has attracted the enthusiasm of relative researchers in recent years. Most existing scene text spotters follow the detection-then-recognition paradigm, where the vanilla detection module hardly determines the reading order and leads to failure recognition. After rethinking the au…

Cited by 2SourcePDFScholar
2025

DCA: Dividing and Conquering Amnesia in Incremental Object Detection

AAAI 2025technical

Incremental object detection (IOD) aims to cultivate an object detector that can continuously localize and recognize novel classes while preserving its performance on previous classes. Existing methods achieve certain success by improving knowledge distillation and exemplar replay for transformer-ba…

2025

Specifying What You Know or Not for Multi-Label Class-Incremental Learning

AAAI 2025technical

Existing class incremental learning is mainly designed for single-label classification task, which is ill-equipped for multi-label scenarios due to the inherent contradiction of learning objectives for samples with incomplete labels. We argue that the main challenge to overcome this contradiction in…

2025

The Role of Video Generation in Enhancing Data-Limited Action Understanding

IJCAI 2025

Video action understanding tasks in real-world scenarios often suffer from data limitations. In this paper, we address the data-limited action understanding problem by bridging data scarcity. We propose a novel method that leverages a text-to-video diffusion transformer to generate annotated data fo

Cited by 0SourcePDFScholar
2024

Accurate and Robust Scene Text Recognition via Adversarial Training

ICASSP 2024accepted

Adversarial training (AT) is a methodology that utilizes adversarial examples in the training process to enhance a model’s resistance to adversarial attacks and improve generalization. Despite its efficacy in several non-sequential computer vision tasks such as classification and object detection, i…

Cited by 0SourceScholar
2024

TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control

NeurIPS 2024spotlight

Centred on content modification and style preservation, Scene Text Editing (STE) remains a challenging task despite considerable progress in text-to-image synthesis and text-driven image manipulation recently. GAN-based STE methods generally encounter a common issue of model generalization, while Di…

2023

One-Shot Replay: Boosting Incremental Object Detection via Retrospecting One Object

AAAI 2023technical

Modern object detectors are ill-equipped to incrementally learn new emerging object classes over time due to the well-known phenomenon of catastrophic forgetting. Due to data privacy or limited storage, few or no images of the old data can be stored for replay. In this paper, we design a novel One-S…

Cited by 7SourcePDFScholar
2020

SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text Recognition

CVPR 2020poster

Scene text recognition is a hot research topic in computer vision. Recently, many recognition methods based on the encoder-decoder framework have been proposed, and they can handle scene texts of perspective distortion and curve shape. Nevertheless, they still face lots of challenges like image blur…

Cited by 340PDFcodeScholar