← All conferences

CVPR 2026 Accepted Papers

The full list of 4,068 papers accepted at CVPR 2026 (IEEE/CVF Conference on Computer Vision and Pattern Recognition). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

accepted: 4,068
  1. Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generationaccepted
  2. Where, What, Why: Toward Explainable 3D-GS Watermarkingaccepted
  3. Which Concepts to Forget and How to Refuse? Decomposing Concepts for Continual Unlearning in Large Vision-Language Modelsaccepted
  4. WhisperNet: A Scalable Solution for Bandwidth-Efficient Collaborationaccepted
  5. White-Balance First, Adjust Later: Cross-Camera Color Constancy via Vision-Language Evaluationaccepted
  6. Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Trainingaccepted
  7. Why Not Hyperparameter-Friendly Optimisation? A Monotonic Adaptive Norm Rescaling Approach For Long-Tailed Recognitionaccepted
  8. WiTTA-Bench: Benchmarking Test-Time Adaptation for WiFi Sensingaccepted
  9. Widget2Code: From Visual Widgets to UI Code via Multimodal LLMsaccepted
  10. WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognitionaccepted
  11. WildCap: Facial Albedo Capture in the Wild via Hybrid Inverse Renderingaccepted
  12. WildPose: A Unified Framework for Robust Pose Estimation in the Wildaccepted
  13. WildRayZer: Self-supervised Large View Synthesis in Dynamic Environmentsaccepted
  14. Will Multimodal Models Be Dazzled by Multi-Image Visual Puzzles?accepted
  15. WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editingaccepted
  16. WonderZoom: Multi-Scale 3D World Generationaccepted
  17. World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Modelsaccepted
  18. WorldGen: From Text to Traversable and Interactive 3D Worldsaccepted
  19. WorldLens: Full-Spectrum Evaluations of Driving World Models in Real Worldaccepted
  20. WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoningaccepted
  21. WorldReel: 4D Video Generation with Consistent Geometry and Motion Modelingaccepted
  22. WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memoriesaccepted
  23. Write Where It Matters: Policy-Guided Watermarks for 3D Gaussian Splattingaccepted
  24. X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detectionaccepted
  25. X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosisaccepted
  26. X-Part: High Fidelity And Structure Coherent Shape Decomposition And Completionaccepted
  27. X-WIN: Building Chest Radiograph World Model via Predictive Sensingaccepted
  28. X-band Radar Non-Line-of-Sight Imagingaccepted
  29. XPaintNet: An eXtreme Lightweight Framework for Stereoscopic Conversion without Inpainting Networkaccepted
  30. XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screeningaccepted
  31. YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detectionaccepted
  32. YOLO-ULM: Ultra-Lightweight Models for Real-Time Object Detectionaccepted
  33. YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removalaccepted
  34. YieldSAT: A Multimodal Benchmark Dataset for High-Resolution Crop Yield Predictionaccepted
  35. Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansionaccepted
  36. You Only Erase Once: Erasing Anything without Bringing Unexpected Contentaccepted
  37. Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and Generationaccepted
  38. Your Dissimilarities Define You: Complementary Learning Exploiting Class Diversitiesaccepted
  39. Your Latent Mask is Wrong: Pixel-Equivalent Latent Compositing for Diffusion Modelsaccepted
  40. Your One-Stop Solution for AI-Generated Video Detectionaccepted
  41. Yume1.5: A Text-Controlled Interactive World Generation Modelaccepted
  42. Z-Order Transformer for Feed-Forward Gaussian Splattingaccepted
  43. ZINA: Multimodal Fine-grained Hallucination Detection and Editingaccepted
  44. ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language Modelsaccepted
  45. Zero-Shot Depth Completion with Vision-Language Modelaccepted
  46. Zero-Shot Image Denoising via Hybrid Prior-Guided Pseudo Sample Generationaccepted
  47. Zero-Shot Reconstruction of Animatable 3D Avatars with Cloth Dynamics from a Single Imageaccepted
  48. Zero-shot Detection of AI-Generated Image via RAW-RGB Alignmentaccepted
  49. ZeroIDIR: Zero-Reference Illumination Degradation Image Restoration with Perturbed Consistency Diffusion Modelsaccepted
  50. ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Trainingaccepted
  51. Zoo3D: Zero-Shot 3D Object Detection at Scene Levelaccepted
  52. ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasksaccepted
  53. b-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignmentaccepted
  54. cryoSENSE: Compressive Sensing Enables High-throughput Microscopy with Sparse and Generative Priors on the Protein Cryo-EM Image Manifoldaccepted
  55. dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Modelsaccepted
  56. eRetinexGS: Retinex Modeling for Low-Light Scene Enhancement via Event Streams and 3D Gaussian Splattingaccepted
  57. fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understandingaccepted
  58. gQIR: Generative Quanta Image Reconstructionaccepted
  59. iLRM: An Iterative Large 3D Reconstruction Modelaccepted
  60. iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generationaccepted
  61. iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perceptionaccepted
  62. iSplat: Iterative Learning for Fine-Grained Gaussian Splattingaccepted
  63. mVLM: A Vision Language Model for mNPUsaccepted
  64. mmWaveFlow: Unified Enhancement and Generation of mmWave Human Point Cloudsaccepted
  65. pH-Strips for Selective Forgetting: A Blunt but Fast Diagnostic Baseline for Machine Unlearningaccepted
  66. rPPG-VQA: A Video Quality Assessment Framework for Unsupervised rPPG Trainingaccepted
  67. tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstructionaccepted
  68. x^2-Fusion: Cross-Modality and Cross-Dimension Flow Estimation in Event Edge Spaceaccepted

Looking for submission deadlines instead? See the conference deadline calendar.