CVPR 2026 Accepted Papers
The full list of 4,068 papers accepted at CVPR 2026 (IEEE/CVF Conference on Computer Vision and Pattern Recognition). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
accepted: 4,068
- Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generationaccepted
- Where, What, Why: Toward Explainable 3D-GS Watermarkingaccepted
- Which Concepts to Forget and How to Refuse? Decomposing Concepts for Continual Unlearning in Large Vision-Language Modelsaccepted
- WhisperNet: A Scalable Solution for Bandwidth-Efficient Collaborationaccepted
- White-Balance First, Adjust Later: Cross-Camera Color Constancy via Vision-Language Evaluationaccepted
- Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Trainingaccepted
- Why Not Hyperparameter-Friendly Optimisation? A Monotonic Adaptive Norm Rescaling Approach For Long-Tailed Recognitionaccepted
- WiTTA-Bench: Benchmarking Test-Time Adaptation for WiFi Sensingaccepted
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMsaccepted
- WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognitionaccepted
- WildCap: Facial Albedo Capture in the Wild via Hybrid Inverse Renderingaccepted
- WildPose: A Unified Framework for Robust Pose Estimation in the Wildaccepted
- WildRayZer: Self-supervised Large View Synthesis in Dynamic Environmentsaccepted
- Will Multimodal Models Be Dazzled by Multi-Image Visual Puzzles?accepted
- WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editingaccepted
- WonderZoom: Multi-Scale 3D World Generationaccepted
- World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Modelsaccepted
- WorldGen: From Text to Traversable and Interactive 3D Worldsaccepted
- WorldLens: Full-Spectrum Evaluations of Driving World Models in Real Worldaccepted
- WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoningaccepted
- WorldReel: 4D Video Generation with Consistent Geometry and Motion Modelingaccepted
- WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memoriesaccepted
- Write Where It Matters: Policy-Guided Watermarks for 3D Gaussian Splattingaccepted
- X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detectionaccepted
- X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosisaccepted
- X-Part: High Fidelity And Structure Coherent Shape Decomposition And Completionaccepted
- X-WIN: Building Chest Radiograph World Model via Predictive Sensingaccepted
- X-band Radar Non-Line-of-Sight Imagingaccepted
- XPaintNet: An eXtreme Lightweight Framework for Stereoscopic Conversion without Inpainting Networkaccepted
- XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screeningaccepted
- YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detectionaccepted
- YOLO-ULM: Ultra-Lightweight Models for Real-Time Object Detectionaccepted
- YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removalaccepted
- YieldSAT: A Multimodal Benchmark Dataset for High-Resolution Crop Yield Predictionaccepted
- Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansionaccepted
- You Only Erase Once: Erasing Anything without Bringing Unexpected Contentaccepted
- Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and Generationaccepted
- Your Dissimilarities Define You: Complementary Learning Exploiting Class Diversitiesaccepted
- Your Latent Mask is Wrong: Pixel-Equivalent Latent Compositing for Diffusion Modelsaccepted
- Your One-Stop Solution for AI-Generated Video Detectionaccepted
- Yume1.5: A Text-Controlled Interactive World Generation Modelaccepted
- Z-Order Transformer for Feed-Forward Gaussian Splattingaccepted
- ZINA: Multimodal Fine-grained Hallucination Detection and Editingaccepted
- ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language Modelsaccepted
- Zero-Shot Depth Completion with Vision-Language Modelaccepted
- Zero-Shot Image Denoising via Hybrid Prior-Guided Pseudo Sample Generationaccepted
- Zero-Shot Reconstruction of Animatable 3D Avatars with Cloth Dynamics from a Single Imageaccepted
- Zero-shot Detection of AI-Generated Image via RAW-RGB Alignmentaccepted
- ZeroIDIR: Zero-Reference Illumination Degradation Image Restoration with Perturbed Consistency Diffusion Modelsaccepted
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Trainingaccepted
- Zoo3D: Zero-Shot 3D Object Detection at Scene Levelaccepted
- ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasksaccepted
- b-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignmentaccepted
- cryoSENSE: Compressive Sensing Enables High-throughput Microscopy with Sparse and Generative Priors on the Protein Cryo-EM Image Manifoldaccepted
- dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Modelsaccepted
- eRetinexGS: Retinex Modeling for Low-Light Scene Enhancement via Event Streams and 3D Gaussian Splattingaccepted
- fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understandingaccepted
- gQIR: Generative Quanta Image Reconstructionaccepted
- iLRM: An Iterative Large 3D Reconstruction Modelaccepted
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generationaccepted
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perceptionaccepted
- iSplat: Iterative Learning for Fine-Grained Gaussian Splattingaccepted
- mVLM: A Vision Language Model for mNPUsaccepted
- mmWaveFlow: Unified Enhancement and Generation of mmWave Human Point Cloudsaccepted
- pH-Strips for Selective Forgetting: A Blunt but Fast Diagnostic Baseline for Machine Unlearningaccepted
- rPPG-VQA: A Video Quality Assessment Framework for Unsupervised rPPG Trainingaccepted
- tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstructionaccepted
- x^2-Fusion: Cross-Modality and Cross-Dimension Flow Estimation in Event Edge Spaceaccepted
CVPR accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.