← Search

Kefan Chen

9 accepted papers

2026

Do Vision and Text Cues Exhibit Evidential Coupling? UFO: A Benchmark for Compositional Multimodal Reasoning in Unified Models

ICML 2026poster

Unified Foundation Models (UFMs), which support interleaved multimodal generation and understanding, have been proposed as a promising paradigm for reasoning about dynamic world states, yet it remains unclear whether the visual content they generate functions as grounded evidence for subsequent reas…

Cited by 0SourceScholar
2025

FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation

CVPR 2025highlight

Despite remarkable progress in image generation models, generating realistic hands remains a persistent challenge due to their complex articulation, varying viewpoints, and frequent occlusions. We present FoundHand, a large-scale domain-specific diffusion model for synthesizing single and dual hand…

Cited by 0SourcePDFScholar
2025

InteractAvatar: Modeling Hand-Face Interaction in Photorealistic Avatars with Deformable Gaussians

ICCV 2025poster

With the rising interest from the community in digital avatars coupled with the importance of expressions and gestures in communication, modeling natural avatar behavior remains an important challenge across many industries such as teleconferencing, gaming, and AR/VR. Human hands are the primary too…

Cited by 0SourcePDFScholar
2025

UVGS: Reimagining Unstructured 3D Gaussian Splatting using UV Mapping

CVPR 2025poster

3D Gaussian Splatting (3DGS) has demonstrated superior quality in modeling 3D objects and scenes. However, generating 3DGS remains challenging due to their discrete, unstructured, and permutation-invariant nature. In this work, we present a simple yet effective method to overcome these challenges. W…

Cited by 2SourcePDFScholar
2024

DiVa-360: The Dynamic Visual Dataset for Immersive Neural Fields

CVPR 2024highlight

Advances in neural fields are enabling high-fidelity capture of the shape and appearance of dynamic 3D scenes. However their capabilities lag behind those offered by conventional representations such as 2D videos because of algorithmic challenges and the lack of large-scale multi-view real-world dat…

Cited by 6SourcePDFScholar
2024

MANUS: Markerless Grasp Capture using Articulated 3D Gaussians

CVPR 2024poster

Understanding how we grasp objects with our hands has important applications in areas like robotics and mixed reality. However this challenging problem requires accurate modeling of the contact between hands and objects.To capture grasps existing methods use skeletons meshes or parametric models tha…

Cited by 12SourcePDFScholar
2023

FACE: Evaluating Natural Language Generation with Fourier Analysis of Cross-Entropy

NeurIPS 2023poster

Measuring the distance between machine-produced and human language is a critical open problem. Inspired by empirical findings from psycholinguistics on the periodicity of entropy in language, we propose FACE, a set of metrics based on Fourier Analysis of the estimated Cross-Entropy of language, for…

2020

An Analysis of SVD for Deep Rotation Estimation

NeurIPS 2020poster

Symmetric orthogonalization via SVD, and closely related procedures, are well-known techniques for projecting matrices onto O(n) or SO(n). These tools have long been used for applications in computer vision, for example optimal 3D alignment problems solved by orthogonal Procrustes, rotation averagin…