← Search

Yanjie Wang

13 accepted papers

2026

Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning

ICML 2026poster

Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is effective, it raises concerns regarding privacy and storage. Thus, generative replay has emerged as a viable alternative, synthesizing old data usin…

Cited by 0SourceScholar
2026

Learnability-Driven Knowledge Assimilation for Class-Incremental Semantic Segmentation

ICML 2026poster

Class-incremental semantic segmentation learns new classes while retaining old ones without access to past data. Although existing methods alleviate catastrophic forgetting on old classes, new-class performance remains limited. We identify the key bottleneck arises from low-margin regions, where the…

Cited by 0SourceScholar
2025

A Bounding Box is Worth One Token - Interleaving Layout and Text in a Large Language Model for Document Understanding

ACL 2025finding

Recently, many studies have demonstrated that exclusively incorporating OCR-derived text and spatial layouts with large language models (LLMs) can be highly effective for document understanding tasks. However, existing methods that integrate spatial layouts with text have limitations, such as produc…

2025

Advancing Sequential Numerical Prediction in Autoregressive Models

ACL 2025short

Autoregressive models have become the de facto choice for sequence generation tasks, but standard approaches treat digits as independent tokens and apply cross-entropy loss, overlooking the coherent structure of numerical sequences. This paper introduces Numerical Token Integrity Loss(NTIL) to addre…

2025

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

ICCV 2025poster

The application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in high-quality image-text datasets for fine-tuning image understanding, but there is still a lack of comparable datasets…

2025

High-dimension Prototype is a Better Incremental Object Detection Learner

ICLR 2025poster

Incremental object detection (IOD), surpassing simple classification, requires the simultaneous overcoming of catastrophic forgetting in both recognition and localization tasks, primarily due to the significantly higher feature space complexity. Integrating Knowledge Distillation (KD) would mitigate…

Cited by 0SourcePDFScholar
2025

MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

ACL 2025finding

Text-Centric Visual Question Answering (TEC-VQA) in its proper format not only facilitates human-machine interaction in text-centric visual environments but also serves as a de facto gold proxy to evaluate AI models in the domain of text-centric scene understanding. Nonetheless, most existing TEC-VQ…

2024

A Soft Crawling Robot with Multi-Modal Locomotion Inspired by the Movement Mechanism of Snake Scales

RA-L 2024

The existing soft crawling robots usually have single-motion mode, which results in poor motion adaptability and significantly restricts the application field of the soft crawling robots. To further increase the motion adaptability of the soft crawling robots and expand their application space, in t

Cited by 5SourceScholar
2024

Elysium: Exploring Object-level Perception in Videos through Semantic Integration Using MLLMs

ECCV 2024poster

"Multi-modal Large Language Models (MLLMs) have demonstrated their ability to perceive objects in still images, but their application in video-related tasks, such as object tracking, remains understudied. This lack of exploration is primarily due to two key challenges. Firstly, extensive pretraining…

2024

PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity Recognition

NeurIPS 2024poster

In this study, we aim to reduce generation latency for Named Entity Recognition (NER) with Large Language Models (LLMs). The main cause of high latency in LLMs is the sequential decoding process, which autoregressively generates all labels and mentions for NER, significantly increase the sequence le…

2024

SNIDA: Unlocking Few-Shot Object Detection with Non-linear Semantic Decoupling Augmentation

CVPR 2024poster

Once only a few-shot annotated samples are available the performance of learning-based object detection would be heavily dropped. Many few-shot object detection (FSOD) methods have been proposed to tackle this issue by adopting image-level augmentations in linear manners. Nevertheless those handcraf…

Cited by 9SourcePDFScholar
2023

Blue Hand: A Novel Type of Soft Anthropomorphic Hand Based on Pneumatic Series-Parallel Mechanism

RA-L 2023

Hand dexterity is tremendously valuable to robots for task-dependent manipulation and interacting with the world. In this work, we present a novel soft pneumatic dexterous hand, which demonstrates highly dexterous and versatile anthropomorphic properties. Inspired by human hand, the proposed hand po

Cited by 15SourceScholar