← Search

Michael Bi Mi

21 accepted papers

2026

AnchorSplat: Feed-Forward 3D Gaussian Splatting With 3D Geometric Priors

CVPR 2026

Recent feed-forward Gaussian reconstruction models adopt a pixel-aligned formulation that maps each 2D pixel to a 3D Gaussian, entangling Gaussian representations tightly with the input images. In this paper, we propose AnchorSplat, a novel feed-forward 3DGS framework for scene-level reconstruction

Cited by 0SourceScholar
2025

Semantic Segmentation on Raindrop Degraded Images Using Two-Stage Dual Teacher-Student Learning

AAAI 2025technical

Existing semantic segmentation methods face challenges when processing input images degraded by raindrops on the lens or windshield. Unlike other adverse conditions such as fog and nighttime, which degrade visual quality, raindrops not only impair visual appearances but also introduce misleading occ…

Cited by 0SourcePDFScholar
2024

DreamDrone: Text-to-Image Diffusion Models are Zero-shot Perpetual View Generators

ECCV 2024poster

"We introduce DreamDrone, a novel zero-shot and training-free pipeline for generating unbounded flythrough scenes from textual prompts. Different from other methods that focus on warping images frame by frame, we advocate explicitly warping the intermediate latent code of the pre-trained text-to-ima…

2024

End-to-End Video Semantic Segmentation in Adverse Weather using Fusion Blocks and Temporal-Spatial Teacher-Student Learning

NeurIPS 2024poster

Adverse weather conditions can significantly degrade the video frames, causing existing video semantic segmentation methods to produce erroneous predictions. In this work, we target adverse weather conditions and introduce an end-to-end domain adaptation strategy that leverages a fusion block, tempo…

Cited by 1SourcePDFScholar
2024

Enhancing Video Super-Resolution via Implicit Resampling-based Alignment

CVPR 2024highlight

In video super-resolution it is common to use a frame-wise alignment to support the propagation of information over time. The role of alignment is well-studied for low-level enhancement in video but existing works overlook a critical step -- resampling. We show through extensive experiments that for…

Cited by 15SourcePDFScholar
2024

HEAP: Unsupervised Object Discovery and Localization with Contrastive Grouping

AAAI 2024technical

Unsupervised object discovery and localization aims to detect or segment objects in an image without any supervision. Recent efforts have demonstrated a notable potential to identify salient foreground objects by utilizing self-supervised transformer features. However, their scopes only build upon p…

Cited by 3SourcePDFScholar
2024

Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching

NeurIPS 2024poster

Diffusion Transformers have recently demonstrated unprecedented generative capabilities for various tasks. The encouraging results, however, come with the cost of slow inference, since each denoising step requires inference on a transformer model with a large scale of parameters. In this study, we m…

2024

MetaISP: Efficient RAW-to-sRGB Mappings with Merely 1M Parameters

IJCAI 2024poster

State-of-the-art deep ISP models alleviate the dilemma of limited generalization capabilities across heterogeneous inputs by increasing the size and complexity of the network, which inevitably leads to considerable growth in parameter counts and FLOPs. To address this challenge, this paper presents…

Cited by 0SourcePDFScholar
2024

MotionMix: Weakly-Supervised Diffusion for Controllable Motion Generation

AAAI 2024technical

Controllable generation of 3D human motions becomes an important topic as the world embraces digital transformation. Existing works, though making promising progress with the advent of diffusion models, heavily rely on meticulously captured and annotated (e.g., text) high-quality motion corpus, a re…

2024

Semantic Segmentation in Multiple Adverse Weather Conditions with Domain Knowledge Retention

AAAI 2024technical

Semantic segmentation's performance is often compromised when applied to unlabeled adverse weather conditions. Unsupervised domain adaptation is a potential approach to enhancing the model's adaptability and robustness to adverse weather. However, existing methods encounter difficulties when sequent…

Cited by 4SourcePDFScholar
2024

Vista3D: unravel the 3d darkside of a single image

ECCV 2024poster

"We embark on the age-old quest: unveiling the hidden dimensions of objects from mere glimpses of their visible parts. To address this, we present Vista3D, a framework that realizes swift and consistent 3D generation within a mere 5 minutes. At the heart of Vista3D lies a two-phase approach: the coa…

2023

DepGraph: Towards Any Structural Pruning

CVPR 2023poster

Structural pruning enables model acceleration by removing structurally-grouped parameters from neural networks. However, the parameter-grouping patterns vary widely across different models, making architecture-specific pruners, which rely on manually-designed grouping schemes, non-generalizable to n…

2023

Improving Deep Regression with Ordinal Entropy

ICLR 2023poster

In computer vision, it is often observed that formulating regression problems as a classification task yields better performance. We investigate this curious phenomenon and provide a derivation to show that classification, with the cross-entropy loss, outperforms regression with a mean squared error…

2023

Overcoming the Trade-Off Between Accuracy and Plausibility in 3D Hand Shape Reconstruction

CVPR 2023poster

Direct mesh fitting for 3D hand shape reconstruction estimates highly accurate meshes. However, the resulting meshes are prone to artifacts and do not appear as plausible hand shapes. Conversely, parametric models like MANO ensure plausible hand shapes but are not as accurate as the non-parametric m…

Cited by 9SourcePDFScholar
2023

PARTNER: Level up the Polar Representation for LiDAR 3D Object Detection

ICCV 2023poster

Recently, polar-based representation has shown promising properties in perceptual tasks. In addition to Cartesian-based approaches, which separate point clouds unevenly, representing point clouds as polar grids has been recognized as an alternative due to (1) its advantage in robust performance unde…

Cited by 10PDFcodeScholar
2023

Priority-Centric Human Motion Generation in Discrete Latent Space

ICCV 2023poster

Text-to-motion generation is a formidable task, aiming to produce human motions that align with the input text while also adhering to human capabilities and physical laws. While there have been advancements in diffusion models, their application in discrete spaces remains underexplored. Current meth…

Cited by 52PDFScholar
2023

TM2D: Bimodality Driven 3D Dance Generation via Music-Text Integration

ICCV 2023poster

We propose a novel task for generating 3D dance movements that simultaneously incorporate both text and music modalities. Unlike existing works that generate dance movements using a single modality such as music, our goal is to produce richer dance movements guided by the instructive information pro…

Cited by 69PDFcodeScholar
2022

ONCE-3DLanes: Building Monocular 3D Lane Detection

CVPR 2022poster

We present ONCE-3DLanes, a real-world autonomous driving dataset with lane layout annotation in 3D space. Conventional 2D lane detection from a monocular image yields poor performance of following planning and control tasks in autonomous driving due to the case of uneven road. Predicting the 3D lane…

Cited by 76PDFcodeScholar
2022

Point2Seq: Detecting 3D Objects As Sequences

CVPR 2022poster

We present a simple and effective framework, named Point2Seq, for 3D object detection from point clouds. In contrast to previous methods that normally predict attributes of 3D objects all at once, we expressively model the interdependencies between attributes of 3D objects, which in turn enables a b…

Cited by 18PDFcodeScholar
2022

PoseTriplet: Co-Evolving 3D Human Pose Estimation, Imitation, and Hallucination Under Self-Supervision

CVPR 2022oral

Existing self-supervised 3D human pose estimation schemes have largely relied on weak supervisions like consistency loss to guide the learning, which, inevitably, leads to inferior results in real-world scenarios with unseen poses. In this paper, we propose a novel self-supervised approach that allo…

Cited by 57PDFcodeScholar