← Search

Delong Chen

12 accepted papers

2026

TV2TV: A Unified Framework for Interleaved Language and Video Generation

CVPR 2026

Video generation models are rapidly advancing, but can still struggle with complex video outputs that require significant semantic branching or repeated high-level reasoning about what should happen next. In this paper, we introduce a new class of omni video-text models that integrate ideas from rec

Cited by 0SourceScholar
2026

VL-JEPA: Joint Embedding Predictive Architecture for Vision-language

ICLR 2026poster

We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA predicts continuous embeddings of the target texts. By learning in an abstract representation space, the model can focu…

Cited by 0SourceScholar
2025

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions

EMNLP 2025

While densely annotated image captions significantly facilitate the learning of robust vision-language alignment, methodologies for systematically optimizing human annotation efforts remain underexplored. We introduce Chain-of-Talkers (CoTalk), an AI-in-the-loop methodology designed to maximize the

Cited by 0SourcePDFScholar
2025

High-Dimension Human Value Representation in Large Language Models

NAACL 2025long

The widespread application of Large Language Models (LLMs) across various tasks and fields has necessitated the alignment of these models with human values and preferences. Given various approaches of human value alignment, such as Reinforcement Learning with Human Feedback (RLHF), constitutional le…

2025

Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models

COLING 2025main

We introduce a novel analysis that leverages linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs). By measuring the similarity between LLM activation differences across minimal pairs, we quantify the linguistic similarity and gain insight into the…

2025

Making Large Vision Language Models to Be Good Few-Shot Learners

AAAI 2025technical

Few-shot classification (FSC) is a fundamental yet challenging task in computer vision that involves recognizing novel classes from limited data. While previous methods have focused on enhancing visual features or incorporating additional modalities, Large Vision Language Models (LVLMs) offer a prom…

2025

Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing Images

ICASSP 2025accepted

The Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where semantic contour extraction—such as identifying buildings, road networks, and coastlines-holds significant practical va…

Cited by 6SourceScholar
2025

Subobject-level Image Tokenization

ICML 2025poster

Patch-based image tokenization ignores the morphology of the visual world, limiting effective and efficient learning of image understanding. Inspired by subword tokenization, we introduce subobject-level adaptive token segmentation and explore several approaches, including superpixel, SAM, and a pro…

2024

Measuring Political Bias in Large Language Models: What Is Said and How It Is Said

ACL 2024long

We propose to measure political bias in LLMs by analyzing both the content and style of their generated content regarding political issues. Existing benchmarks and measures focus on gender and racial biases. However, political bias exists in LLMs and can lead to polarization and other harms in downs…

Cited by 30SourcePDFScholar
2023

Few-shot Classification via Ensemble Learning with Multi-Order Statistics

IJCAI 2023poster

Transfer learning has been widely adopted for few-shot classification. Recent studies reveal that obtaining good generalization representation of images on novel classes is the key to improving the few-shot classification accuracy. To address this need, we prove theoretically that leveraging ensembl…

Cited by 9SourcePDFScholar