← Search

Nemo Chen

4 accepted papers

2026

IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning

ICLR 2026poster

Large Vision-Language Models (LVLMs) achieve impressive performance across multiple tasks. A significant challenge, however, is their prohibitive inference cost when processing high-resolution visual inputs. While visual token pruning has emerged as a promising solution, existing methods that primar…

Cited by 0SourceScholar
2025

CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection

NeurIPS 2025poster

With the exponential growth of data, traditional object detection methods are increasingly struggling to handle vast vocabulary object detection tasks effectively. We analyze two key limitations of classification-based detectors: positive gradient dilution, where rare positive categories receive ins…

Cited by 0SourcecodeScholar
2025

DynamicFace: High-Quality and Consistent Face Swapping for Image and Video using Composable 3D Facial Priors

ICCV 2025poster

Face swapping transfers the identity of a source face to a target face while retaining the attributes like expression, pose, hair, and background of the target face. Advanced face swapping methods have achieved attractive results. However, these methods often inadvertently transfer identity informat…

2025

InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention

NeurIPS 2025poster

Diffusion models have demonstrated remarkable capabilities in generating high-quality images. Recent advancements in Layout-to-Image (L2I) generation have leveraged positional conditions and textual descriptions to facilitate precise and controllable image synthesis. Despite overall progress, curren…

Cited by 0SourcecodeScholar