← Search

Chang Wen Chen

17 accepted papers

2026

VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning

ICLR 2026poster

Videos, with their unique temporal dimension, demand precise grounded understanding, where answers are directly linked to visual, interpretable evidence. Despite significant breakthroughs in text-based reasoning with large language models, multi-modal reasoning - especially for videos - remains limi…

Cited by 9SourcecodeScholar
2025

SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control

AAAI 2025technical

Autonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data for autonomous driving applications and present SubjectDrive, the first model proven to scale generative data production…

Cited by 11SourcePDFScholar
2025

UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning

NeurIPS 2025poster

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention has been given to scaling fine-grained pixel-level understa…

Cited by 0SourceScholar
2024

E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

NeurIPS 2024poster

Recent advances in Video Large Language Models (Video-LLMs) have demonstrated their great potential in general-purpose video understanding. To verify the significance of these models, a number of benchmarks have been proposed to diagnose their capabilities in different scenarios. However, existing b…

2024

SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer

CVPR 2024poster

Diffusion Transformer (DiT) has emerged as the new trend of generative diffusion models on image generation. In view of extremely slow convergence in typical DiT recent breakthroughs have been driven by mask strategy that significantly improves the training efficiency of DiT with additional intra-im…

Cited by 27SourcePDFScholar
2023

GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular Video

ICCV 2023poster

3D human pose estimation has been researched for decades with promising fruits. 3D human pose lifting is one of the promising research directions toward the task where both estimated pose and ground truth pose data are used for training. Existing pose lifting works mainly focus on improving the perf…

Cited by 80PDFcodeScholar
2021

Improving Contrastive Learning by Visualizing Feature Transformation

ICCV 2021poster

Contrastive learning, which aims at minimizing the distance between positive pairs while maximizing that of negative ones, has been widely and successfully applied in unsupervised feature learning, where the design of positive and negative (pos/neg) pairs is one of its keys. In this paper, we attemp…

Cited by 104PDFcodeScholar
2019

AVT: Unsupervised Learning of Transformation Equivariant Representations by Autoencoding Variational Transformations

ICCV 2019poster

The learning of Transformation-Equivariant Representations (TERs), which is introduced by Hinton et al. [??], has been considered as a principle to reveal visual structures under various transformations. It contains the celebrated Convolutional Neural Networks (CNNs) as a special case that only equi…

Cited by 49PDFScholar
2018

DA-GAN: Instance-Level Image Translation by Deep Attention Generative Adversarial Networks

CVPR 2018poster

Unsupervised image translation, which aims in translating two independent sets of images, is challenging in discovering the correct correspondences without paired data. Existing works build upon Generative Adversarial Networks (GANs) such that the distribution of the translated images are indistingu…

Cited by 182SourcePDFScholar
2017

A-Lamp: Adaptive Layout-Aware Multi-Patch Deep Convolutional Neural Network for Photo Aesthetic Assessment

CVPR 2017poster

Deep convolutional neural networks (CNN) have recently been shown to generate promising results for aesthetics assessment. However, the performance of these deep CNN methods is often compromised by the constraint that the neural network only takes the fixed-size input. To accommodate this requiremen…

Cited by 260PDFScholar
2017

Fast Haze Removal for Nighttime Image Using Maximum Reflectance Prior

CVPR 2017poster

In this paper, we address a haze removal problem from a single nighttime image, even in the presence of varicolored and non-uniform illumination. The core idea lies in a novel maximum reflectance prior. We first introduce the nighttime hazy imaging model, which includes a local ambient illumination…

Cited by 234PDFScholar