← Search

Fangyi Chen

10 accepted papers

2026

MetaVLA: Unified Meta Co-Training for Efficient Embodied Adaptation

ICLR 2026poster

Vision–Language–Action (VLA) models show promise in embodied reasoning, yet remain far from true generalists—they often require task-specific fine-tuning, incur high compute costs, and generalize poorly to unseen tasks. We propose MetaVLA, a unified, backbone-agnostic post-training framework for eff…

Cited by 0SourceScholar
2026

STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision

AAAI 2026technical

Vision-language models (VLMs) have made significant strides in reasoning, yet they often struggle with complex multimodal tasks and tend to generate overly verbose outputs. A key limitation is their reliance on chain-of-thought (CoT) reasoning, despite many tasks benefiting from alternative topologi

Cited by 0SourcePDFScholar
2025

Masked Autoencoders Are Effective Tokenizers for Diffusion Models

ICML 2025spotlight

Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that i…

Cited by 8SourcePDFScholar
2025

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer

CVPR 2025poster

Efficient image tokenization with high compression ratios remains a critical challenge for training generative models.We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple codewords into each latent token, substantially increasing the re…

2023

Enhanced Training of Query-Based Object Detection via Selective Query Recollection

CVPR 2023poster

This paper investigates a phenomenon where query-based object detectors mispredict at the last decoding stage while predicting correctly at an intermediate stage. We review the training process and attribute the overlooked phenomenon to two limitations: lack of training emphasis and cascading errors…

Cited by 61SourcePDFScholar
2022

"Unitail: Detecting, Reading, and Matching in Retail Scene"

ECCV 2022poster

"To make full use of computer vision technology in stores, it is required to consider the actual needs that fit the characteristics of the retail scene. Pursuing this goal, we introduce the United Retail Datasets (Unitail), a large-scale benchmark of basic visual tasks on products that challenges al…

2021

Semantic Relation Reasoning for Shot-Stable Few-Shot Object Detection

CVPR 2021poster

Few-shot object detection is an imperative and long-lasting problem due to the inherent long-tail distribution of real-world data. Its performance is largely affected by the data scarcity of novel classes. But the semantic relation between the novel classes and the base classes is constant regardles…

Cited by 247PDFScholar
2020

Solving Missing-Annotation Object Detection with Background Recalibration Loss

ICASSP 2020accepted

This paper focuses on a novel and challenging detection scenario: A majority of true objects/instances is unlabeled in the datasets, so these missing-labeled areas will be regarded as the background during training. Previous art [1] on this problem has proposed to use soft sampling to re-weight the…

Cited by 0SourceScholar