← Search

Wenbo Hu

41 accepted papers

2026

ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping

ICLR 2026poster

Recent advances in multimodal large reasoning models (MLRMs) have substantially improved their ability to solve complex textual and visual tasks. However, these models tend to *overthink* on simple problems, producing unnecessarily lengthy reasoning traces, while *under-exploring* on challenging one…

Cited by 20SourcecodeScholar
2026

Benchmarking Trustworthiness in Multimodal LLMs for Video Understanding

AAAI 2026technical

Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability.

Cited by 0SourcePDFScholar
2026

Concept Bottleneck Models for Explainable Decision Making: A Survey of Progress, Taxonomy, and Future Directions

IJCAI 2026

Deep neural networks deliver strong performance but remain opaque, limiting their use in high-stakes domains that require transparency and human oversight. Concept Bottleneck Models (CBMs) address this gap by introducing a human-interpretable concept layer that mediates inputs and decisions, enablin

Cited by 0Scholar
2026

G$^2$TAM: Geometry Grounded Track Anything Model

ICML 2026poster

Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current video segmentation models depend on explicit object appearance memory banks for instance tracking, yet they remain vulnera…

Cited by 0SourceScholar
2026

G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning

CVPR 2026

Vision-Language Models (VLMs) still lack robustness in spatial intelligence, demonstrating poor performance on spatial understanding and reasoning tasks. We attribute this gap to the absence of a visual geometry learning process capable of reconstructing 3D space from 2D images. We present G^2VLM, a

Cited by 0SourcecodeScholar
2026

IC-Custom: Diverse Image Customization via In-Context Learning

ICLR 2026poster

Image customization, a crucial technique for industrial media production, aims to generate content that is consistent with reference images. However, current approaches conventionally separate image customization into position-aware and position-free customization paradigms and lack a universal fram…

Cited by 0SourcecodeScholar
2026

Interleaving Reasoning for Better Text-to-Image Generation

ICLR 2026poster

Unified multimodal understanding and generation models recently have achieve significant improvement in image generation capability, yet a large gap remains in instruction following and detail preservation compared to systems that tightly couple comprehension with generation such as GPT-4o. Motivate…

Cited by 0SourcecodeScholar
2026

MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE

CVPR 2026

We present MotionCrafter, a framework that leverages video generators to jointly reconstruct 4D geometry and estimate dense motion from a monocular video. The key idea is a joint representation of dense 3D point maps and 3D scene flows in a shared coordinate system, together with a 4D VAE tailored t

Cited by 0SourceScholar
2026

ProSafePrune: Projected Safety Pruning for Mitigating Over-Refusal in LLMs

ICLR 2026poster

Large Language Models (LLMs) excel in various domains, but their safe deployment faces the challenge of balancing safety and utility. Existing alignment strategies often strengthen refusal mechanisms to reduce harmful outputs, but harmless instructions with superficial risky words are mistakenly rej…

Cited by 0SourcecodeScholar
2026

Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers

CVPR 2026

Recent breakthroughs in 3D generative modeling have yielded remarkable progress in static shape synthesis, yet truly dynamic 4D generation remains elusive, hindered by temporal artifacts and prohibitive computational demand. We present Sculpt4D, a native 4D generative framework that seamlessly integ

Cited by 2SourcecodeScholar
2026

Sparse-Scale Transformer with Bidirectional Awareness for Time Series Forecasting

AAAI 2026technical

Time series forecasting (TSF) plays a crucial role in many real-world applications, such as weather prediction and economic planning. While Transformer-based models have shown strong capabilities in modeling long-range dependencies, effectively capturing the multi-scale temporal dynamics inherent in

Cited by 0SourcePDFScholar
2026

VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

CVPR 2026

Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as videos inherently capture dynamics in the projected 2D image plane. To bridge this gap, we introduce VerseCrafter, a geome

Cited by 0SourcecodeScholar
2025

3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model

NeurIPS 2025poster

Humans excel at performing complex tasks by leveraging long-term memory across temporal and spatial experiences. In contrast, current Large Language Models (LLMs) struggle to effectively plan and act in dynamic, multi-room 3D environments. We posit that part of this limitation is due to the lack of…

Cited by 0SourceScholar
2025

DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

CVPR 2025highlight

Estimating video depth in open-world scenarios is challenging due to the diversity of videos in appearance, content motion, camera movement, and length. We present DepthCrafter, an innovative method for generating temporally consistent long depth sequences with intricate details for open-world video…

2025

GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors

ICCV 2025poster

Despite remarkable advancements in video depth estimation, existing methods fall short in geometric fidelity due to their affine-invariant predictions, restricting their applicability in reconstruction and other metrically grounded downstream tasks. We propose a novel point map Variational Autoencod…

Cited by 0SourcePDFScholar
2025

MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models

ICLR 2025poster

Existing multimodal retrieval benchmarks primarily focus on evaluating whether models can retrieve and utilize external textual knowledge for question answering. However, there are scenarios where retrieving visual information is either more beneficial or easier to access than textual data. In this…

Cited by 9SourcePDFScholar
2025

Mani-GS: Gaussian Splatting Manipulation with Triangular Mesh

CVPR 2025poster

Neural 3D representations, such as Neural Radiation Fields (NeRF), excel at producing photorealistic rendering results but lack the flexibility for manipulation and editing which is crucial for content creation. However, manipulating NeRF is not highly controllable and requires a long training and i…

Cited by 10SourcePDFScholar
2025

NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images

CVPR 2025poster

Recent advancements in generative models have significantly improved novel view synthesis (NVS) from multi-view data. However, existing methods depend on external multi-view alignment processes, such as explicit pose estimation or pre-reconstruction, which limits their flexibility and accessibility,…

Cited by 1SourcePDFScholar
2025

NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors

ICCV 2025poster

Surface normal estimation serves as a cornerstone for a spectrum of computer vision applications. While numerous efforts have been devoted to static image scenarios, ensuring temporal coherence in video-based normal estimation remains a formidable challenge. Instead of merely augmenting existing met…

2025

SATA: A Paradigm for LLM Jailbreak via Simple Assistive Task Linkage

ACL 2025finding

Large language models (LLMs) have made significant advancements across various tasks, but their safety alignment remains a major concern. Exploring jailbreak prompts can expose LLMs’ vulnerabilities and guide efforts to secure them. Existing methods primarily design sophisticated instructions for th…

2025

SURE: Safety Understanding and Reasoning Enhancement for Multimodal Large Language Models

EMNLP 2025

Multimodal large language models (MLLMs) demonstrate impressive capabilities by integrating visual and textual information. However, the incorporation of visual modalities also introduces new and complex safety risks, rendering even the most advanced models vulnerable to sophisticated jailbreak atta

2025

TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models

ICCV 2025poster

We present TrajectoryCrafter, a novel approach to redirect camera trajectories for monocular videos. By disentangling deterministic view transformations from stochastic content generation, our method achieves precise control over user-specified camera trajectories. We propose a novel dual-stream con…

Cited by 0SourcePDFScholar
2025

Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models

COLING 2025main

Multimodal large language models (MLLMs) combine visual and textual data for tasks like image captioning and visual question answering. Proper uncertainty calibration is crucial but challenging for reliable use in areas like healthcare and autonomous driving. This paper investigates several MLLMs, f…

2025

Verbalized Representation Learning for Interpretable Few-Shot Generalization

ICCV 2025poster

Humans recognize objects after observing only a few examples, a remarkable capability enabled by their inherent language understanding of the real-world environment. Developing verbalized and interpretable representation can significantly improve model generalization in low-data settings. In this wo…

2024

Analytic-Splatting: Anti-Aliased 3D Gaussian Splatting via Analytic Integration

ECCV 2024oral

"3D Gaussian Splatting (3DGS) recently gained popularity by combining the advantages of both primitive-based and volumetric 3D representations, resulting in improved quality and efficiency for 3D scene rendering. However, 3DGS is not alias-free and still produces severe blurring or jaggies when rend…

2024

BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions

AAAI 2024technical

Vision Language Models (VLMs), which extend Large Language Models (LLM) by incorporating visual understanding capability, have demonstrated significant advancements in addressing open-ended visual question-answering (VQA) tasks. However, these models cannot accurately interpret images infused with t…

2024

CV-VAE: A Compatible Video VAE for Latent Generative Video Models

NeurIPS 2024poster

Spatio-temporal compression of videos, utilizing networks such as Variational Autoencoders (VAE), plays a crucial role in OpenAI's SORA and numerous other video generative models. For instance, many LLM-like video models learn the distribution of discrete tokens derived from 3D VAEs within the VQVAE…

2024

HiFi-123: Towards High-fidelity One Image to 3D Content Generation

ECCV 2024poster

"Recent advances in diffusion models have enabled 3D generation from a single image. However, current methods often produce suboptimal results for novel views, with blurred textures and deviations from the reference image, limiting their practical applications. In this paper, we introduce HiFi-123,…

Cited by 26SourcePDFScholar
2024

Inverse Rendering of Glossy Objects via the Neural Plenoptic Function and Radiance Fields

CVPR 2024poster

Inverse rendering aims at recovering both geometry and materials of objects. It provides a more compatible reconstruction for conventional rendering engines compared with the neural radiance fields (NeRFs). On the other hand existing NeRF-based inverse rendering methods cannot handle glossy objects…

Cited by 7SourcePDFScholar
2024

Matryoshka Query Transformer for Large Vision-Language Models

NeurIPS 2024poster

Large Vision-Language Models (LVLMs) typically encode an image into a fixed number of visual tokens (e.g., 576) and process these tokens with a language model. Despite their strong performance, LVLMs face challenges in adapting to varying computational constraints. This raises the question: can we a…

2024

Spotting the Unseen: Reciprocal Consensus Network Guided by Visual Archetypes

AAAI 2024technical

Humans often require only a few visual archetypes to spot novel objects. Based on this observation, we present a strategy rooted in ``spotting the unseen" by establishing dense correspondences between potential query image regions and a visual archetype, and we propose the Consensus Network (CoNet).…

2024

Texture-GS: Disentangle the Geometry and Texture for 3D Gaussian Splatting Editing

ECCV 2024poster

"3D Gaussian splatting, emerging as a groundbreaking approach, has drawn increasing attention for its capabilities of high-fidelity reconstruction and real-time rendering. However, it couples the appearance and geometry of the scene within the Gaussian attributes, which hinders the flexibility of ed…

Cited by 16SourcePDFScholar
2024

VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models

ACL 2024findings

Large Vision-Language Models (LVLMs) suffer from hallucination issues, wherein the models generate plausible-sounding but factually incorrect outputs, undermining their reliability. A comprehensive quantitative evaluation is necessary to identify and understand the extent of hallucinations in these…

2023

Tri-MipRF: Tri-Mip Representation for Efficient Anti-Aliasing Neural Radiance Fields

ICCV 2023oral

Despite the tremendous progress in neural radiance fields (NeRF), we still face a dilemma of the trade-off between quality and efficiency, e.g., MipNeRF presents fine-detailed and anti-aliased renderings but takes days for training, while Instant-ngp can accomplish the reconstruction in a few minute…

Cited by 142PDFcodeScholar
2021

Bidirectional Projection Network for Cross Dimension Scene Understanding

CVPR 2021poster

2D image representations are in regular grids and can be processed efficiently, whereas 3D point clouds are unordered and scattered in 3D space. The information inside these two visual domains is well complementary, e.g., 2D images have fine-grained texture while 3D point clouds contain plentiful ge…

Cited by 147PDFcodeScholar
2021

Sparse Needlets for Lighting Estimation With Spherical Transport Loss

ICCV 2021poster

Accurate lighting estimation is challenging yet critical to many computer vision and computer graphics tasks such as high-dynamic-range (HDR) relighting. Existing approaches model lighting in either frequency domain or spatial domain which is insufficient to represent the complex lighting conditions…

Cited by 112PDFScholar
2021

Two Birds with One Stone: Series Saliency for Accurate and Interpretable Multivariate Time Series Forecasting

IJCAI 2021poster

It is important yet challenging to perform accurate and interpretable time series forecasting. Though deep learning methods can boost forecasting accuracy, they often sacrifice interpretability. In this paper, we present a new scheme of series saliency to boost both accuracy and interpretability. By…

Cited by 36SourcePDFScholar