← Search

Yi-Ping Phoebe Chen

10 accepted papers

2026

Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human Preferences

CVPR 2026

Narrative image generation aims to create images featuring multiple distinct characters while capturing their interrelationships, posing significant challenges for current text-to-image diffusion models. As a result, general personalized methods often suffer from poor semantic alignment, identity bl

Cited by 0SourceScholar
2026

Hypergraph-State Collaborative Reasoning for Multi-Object Tracking

CVPR 2026

Motion reasoning serves as the cornerstone of multi-object tracking (MOT), as it enables consistent association of targets across frames. However, existing motion estimation approaches face two major limitations: (1) instability caused by noisy or probabilistic predictions, and (2) vulnerability und

Cited by 0SourcecodeScholar
2026

Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imagery

CVPR 2026

Vision-language foundation models (VLFMs) promise zero-shot and retrieval understanding for Earth observation. While operational satellite systems often lack full multi-spectral coverage, making RGB-only inference highly desirable for scalable deployment, the adoption of VLFMs for satellite imagery

Cited by 0SourcecodeScholar
2025

GA-S3: Comprehensive Social Network Simulation with Group Agents

ACL 2025finding

Social network simulation is developed to provide a comprehensive understanding of social networks in the real world, which can be leveraged for a wide range of applications such as group behavior emergence, policy optimization, and business strategy development. However, billions of individuals and…

2025

SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding

CVPR 2025poster

Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding,…

Cited by 1SourcePDFScholar
2025

Temporal Coherent Object Flow for Multi-Object Tracking

AAAI 2025technical

Multi-object tracking is a challenging vision task that requires simultaneous reasoning about object detection and object association. Conventional solutions use frame as the basic unit and typically rely on a motion predictor that exploits the appearance features to associate detected candidates, l…

Cited by 0SourcePDFScholar
2023

Compact Transformer Tracker with Correlative Masked Modeling

AAAI 2023technical

Transformer framework has been showing superior performances in visual object tracking for its great strength in information aggregation across the template and search image with the well-known attention mechanism. Most recent advances focus on exploring attention mechanism variants for better infor…

2022

Balanced Contrastive Learning for Long-Tailed Visual Recognition

CVPR 2022poster

Real-world data typically follow a long-tailed distribution, where a few majority categories occupy most of the data while most minority categories contain a limited number of samples. Classification models minimizing cross-entropy struggle to represent and classify the tail classes. Although the pr…

Cited by 259PDFcodeScholar
2021

Multi-Directional Convolution Networks with Spatial-Temporal Feature Pyramid Module for Action Recognition

ICASSP 2021accepted

Recent attempts show that factorizing 3D convolutional filters into separate spatial and temporal components brings impressive improvement in action recognition. However, traditional temporal convolution operating along the temporal dimension will aggregate unrelated features, since the feature maps…

Cited by 0SourceScholar