← Search

Bowen Pan

14 accepted papers

2026

LoG3D: Ultra-High-Resolution 3D Shape Modeling via Local-to-Global Partitioning

CVPR 2026

Generating high-fidelity 3D contents remains a fundamental challenge due to the complexity of representing arbitrary topologies--such as open surfaces and intricate internal structures--while preserving geometric details. Prevailing methods based on signed distance fields (SDFs) are hampered by cost

Cited by 0SourceScholar
2026

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

ICLR 2026poster

Unified multimodal Large Language Models (LLMs) that can both understand and generate visual content hold immense potential. However, existing open-source models often suffer from a performance trade-off between these capabilities. We present Manzano, a simple and scalable unified framework that sub…

Cited by 0SourceScholar
2026

VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement

CVPR 2026

We propose VIAFormer, a Voxel-Image Alignment transFormer model designed for Multi-view Conditioned Voxel Refinement--the task of repairing incomplete noisy voxels using calibrated multi-view images as guidance. Its effectiveness stems from a synergistic design: an Image Index that provides explicit

Cited by 1SourceScholar
2024

Brain Netflix: Scaling Data to Reconstruct Videos from Brain Signals

ECCV 2024poster

"The field of brain-to-stimuli reconstruction has seen significant progress in the last few years, but techniques continue to be subject-specific and are usually tested on a single dataset. In this work, we present a novel technique to reconstruct videos from functional Magnetic Resonance Imaging (f…

Cited by 2SourcePDFScholar
2024

IntrinsicAnything: Learning Diffusion Priors for Inverse Rendering Under Unknown Illumination

ECCV 2024poster

"† Corresponding author. This paper aims to recover object materials from posed images captured under an unknown static lighting condition. Recent methods solve this task by optimizing material parameters through differentiable physically based rendering. However, due to the coupling between object…

2024

LangNav: Language as a Perceptual Representation for Navigation

NAACL 2024findings

We explore the use of language as a perceptual representation for vision-and-language navigation (VLN), with a focus on low-data settings. Our approach uses off-the-shelf vision systems for image captioning and object detection to convert an agent’s egocentric panoramic view at each time step into n…

2023

HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World

ICCV 2023poster

Building an interactive AI assistant that can perceive, reason, and collaborate with humans in the real world has been a long-standing pursuit in the AI community. This work is part of a broader research effort to develop intelligent agents that can interactively guide humans through performing task…

Cited by 55PDFcodeScholar
2022

CageNeRF: Cage-based Neural Radiance Field for Generalized 3D Deformation and Animation

NeurIPS 2022accept

While implicit representations have achieved high-fidelity results in 3D rendering, it remains challenging to deforming and animating the implicit field. Existing works typically leverage data-dependent models as deformation priors, such as SMPL for human body animation. However, this dependency on…

Cited by 59SourcePDFScholar
2021

Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting

NeurIPS 2021poster

We introduce Argoverse 2 (AV2) — a collection of three datasets for perception and forecasting research in the self-driving domain. The annotated Sensor Dataset contains 1,000 sequences of multimodal data, encompassing high-resolution imagery from seven ring cameras, and two stereo cameras in additi…

Cited by 722SourcecodeScholar
2021

IA-RED$^2$: Interpretability-Aware Redundancy Reduction for Vision Transformers

NeurIPS 2021poster

The self-attention-based model, transformer, is recently becoming the leading backbone in the field of computer vision. In spite of the impressive success made by transformers in a variety of vision tasks, it still suffers from heavy computation and intensive memory costs. To address this limitation…

Cited by 183SourcePDFScholar
2021

VA-RED$^2$: Video Adaptive Redundancy Reduction

ICLR 2021poster

Performing inference on deep learning models for videos remains a challenge due to the large amount of computational resources required to achieve robust recognition. An inherent property of real-world videos is the high correlation of information across frames which can translate into redundancy in…

Cited by 20SourcePDFScholar
2020

Cross-View Semantic Segmentation for Sensing Surroundings

RA-L 2020

Sensing surroundings plays a crucial role in human spatial perception, as it extracts the spatial configuration of objects as well as the free space from the observations. To facilitate the robot perception with such a surrounding sensing capability, we introduce a novel visual task called Cross-vie

Cited by 317SourcecodeScholar
2018

Recurrent Residual Module for Fast Inference in Videos

CVPR 2018poster

Deep convolutional neural networks (CNNs) have made impressive progress in many video recognition tasks such as video pose estimation and video object detection. However, running CNN inference on video requires numerous computation and is usually slow. In this work, we propose a framework called Rec…

Cited by 46SourcePDFScholar