← Search

Guang Shi

4 accepted papers

2026

Depth Anything 3: Recovering the Visual Space from Any Views

ICLR 2026oral

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINOv2 encoder) is sufficient…

Cited by 0SourcecodeScholar
2026

Generative Universal Verifier as Multimodal Meta-Reasoner

ICLR 2026oral

We introduce *Generative Universal Verifier*, a novel concept and plugin designed for next-generation multimodal reasoning in vision-language models and unified multimodal models, providing the fundamental capability of reflection and refinement on visual outcomes during the reasoning and generation…

Cited by 0SourcecodeScholar
2025

DeltaFormer: Unlock the state space of Transformer

NeurIPS 2025poster

In recent years, large language models with Transformer architecture as the core have made breakthrough progress in many fields. At the same time, there are also some weaknesses in the large language model that have prompted people to reflect, among which the most fundamental one is the reflection o…

Cited by 0SourceScholar
2025

GMMCL: Adaptive Concept Drift in Data Streams with Gaussian Mixture Models based on Contrastive Learning

ICASSP 2025accepted

Classical classification methods often fail in dynamic environments where data distributions shift over time, known as concept drift. Applications like flight delay prediction and weather forecasting require handling such dynamic data streams. Concept drift can be either virtual, affecting unconditi…

Cited by 0SourceScholar