← Search

Yiming Wu

15 accepted papers

2026

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning

ICML 2026poster

Tool-augmented reasoning has emerged as a promising direction for enhancing the reasoning capabilities of multimodal large language models (MLLMs). However, existing studies mainly focus on enabling models to perform tool invocation, while neglecting the necessity of invoking tools. We argue that to…

Cited by 0SourceScholar
2026

Mass Concept Erasure in Diffusion Models with Concept Hierarchy

AAAI 2026technical

The success of diffusion models has raised concerns about the generation of unsafe or harmful content, prompting concept erasure approaches that fine-tune modules to suppress specific concepts while preserving general generative capabilities. However, as the number of erased concepts grows, these me

Cited by 0SourcePDFScholar
2026

SUPERVISED MAKEUP TRANSFER WITH A CURATED DATASET: DECOUPLING IDENTITY AND MAKEUP FEATURES FOR ENHANCED TRANSFORMATION

ICASSP 2026poster

Diffusion models have recently shown strong progress in generative tasks, offering a more stable alternative to GAN-based approaches for makeup transfer. Existing methods often suffer from limited datasets, poor disentanglement between identity and makeup features, and weak controllability. To addre…

Cited by 0SourcePDFScholar
2026

Towards Human-Imperceptible Backdoor Attacks on Text-to-Image Diffusion Models

CVPR 2026

Deep learning models are well known to be susceptible to backdoor attacks, and text-to-image generation models are no exception. When a specific trigger is embedded in the input, a backdoored model can be manipulated to perform attacker-defined malicious behaviors, such as generating harmful or inap

Cited by 0SourceScholar
2025

On-Device Diffusion Transformer Policy for Efficient Robot Manipulation

ICCV 2025poster

Diffusion Policies have significantly advanced robotic manipulation tasks via imitation learning, but their application on resource-constrained mobile platforms remains challenging due to computational inefficiency and extensive memory footprint. In this paper, we propose LightDP, a novel framework…

Cited by 0SourcePDFScholar
2024

Mrtnet: Multi-Resolution Temporal Network for Video Sentence Grounding

ICASSP 2024accepted

Video sentence grounding locates a specific moment in a video based on a text query. Existing methods focus on single temporal resolution, ignoring multi-scale temporal consistency. We introduce MRTNet, a multi-resolution grounding network with four key components: a feature encoder, a Multi-Resolut…

Cited by 0SourceScholar
2024

Panoptic Scene Graph Generation with Semantics-Prototype Learning

AAAI 2024technical

Panoptic Scene Graph Generation (PSG) parses objects and predicts their relationships (predicate) to connect human language and visual scenes. However, different language preferences of annotators and semantic overlaps between predicates lead to biased predicate annotations in the dataset, i.e. diff…

2024

Progressive Classifier and Feature Extractor Adaptation for Unsupervised Domain Adaptation on Point Clouds

ECCV 2024poster

"Unsupervised domain adaptation (UDA) is a critical challenge in the field of point cloud analysis. Previous works tackle the problem either by feature extractor adaptation to enable a shared classifier to distinguish domain-invariant features, or by classifier adaptation to evolve the classifier to…

2024

Self-Distilled Dynamic Fusion Network for Language-Based Fashion Retrieval

ICASSP 2024accepted

In the domain of language-based fashion image retrieval, pinpointing the desired fashion item using both a reference image and its accompanying textual description is an intriguing challenge. Existing approaches lean heavily on static fusion techniques, intertwining image and text. Despite their com…

Cited by 0SourceScholar
2022

Difficulty-Aware Neural Band-to-Piano Score Arrangement based on Note- and Statistic-Level Criteria

ICASSP 2022accepted

This paper describes a neural music arrangement method that converts a given band score into a piano score with an elementary or advanced level. The major challenge of this task lies in its ill-posed nature, i.e., various piano arrangements are plausible for a band score. In this paper, we take a sc…

Cited by 0SourceScholar
2020

BANet: Bidirectional Aggregation Network With Occlusion Handling for Panoptic Segmentation

CVPR 2020oral

Panoptic segmentation aims to perform instance segmentation for foreground instances and semantic segmentation for background stuff simultaneously. The typical top-down pipeline concentrates on two key issues: 1) how to effectively model the intrinsic interaction between semantic segmentation and in…

Cited by 92PDFcodeScholar
2019

ChamNet: Towards Efficient Network Design Through Platform-Aware Model Adaptation

CVPR 2019poster

This paper proposes an efficient neural network (NN) architecture design methodology called Chameleon that honors given resource constraints. Instead of developing new building blocks or using computationally-intensive reinforcement learning algorithms, our approach leverages existing efficient netw…

Cited by 341PDFcodeScholar
2019

FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search

CVPR 2019oral

Designing accurate and efficient ConvNets for mobile devices is challenging because the design space is combinatorially large. Due to this, previous neural architecture search (NAS) methods are computationally expensive. ConvNet architecture optimality depends on factors such as input resolution and…

Cited by 1699PDFcodeScholar