← All conferences

CVPR 2026 Accepted Papers

The full list of 4,068 papers accepted at CVPR 2026 (IEEE/CVF Conference on Computer Vision and Pattern Recognition). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

accepted: 4,068
  1. $L^{2}DGS$: Low-Light Dynamic Gaussian Splattingaccepted
  2. $\alpha$Matte4K & $\mu$Matting: Dataset and Model for Ultra-Micro Precision Alpha Video Mattingaccepted
  3. $\oslash$ Source Models Leak What They Shouldn't $\nrightarrow$: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial Optimizationaccepted
  4. $\phi$-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Modelsaccepted
  5. 2-Shots in the Dark: Low-Light Denoising with Minimal Data Acquisitionaccepted
  6. 240FPS Stereo Vision from Monocular Mixed Spikesaccepted
  7. 2D-LFM: Lifting Foundation Model without 3D Supervisionaccepted
  8. 2ndMatch: Finetuning Pruned Diffusion Models via Second-Order Jacobian Matchingaccepted
  9. 3D Gaussian Splatting at Arbitrary Resolutions with Compact Proxy Anchorsaccepted
  10. 3D Gaussian Splatting from Unposed Spike Streamaccepted
  11. 3D Gaussian Splatting with Self-Constrained Priors for High Fidelity Surface Reconstructionaccepted
  12. 3D Space as a Scratchpad for Editable Text-to-Image Generationaccepted
  13. 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Cloudsaccepted
  14. 3D-Aware Implicit Motion Control for View-Adaptive Human Video Generationaccepted
  15. 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understandingaccepted
  16. 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Imageaccepted
  17. 3D-IDE: 3D Implicit Depth Emergentaccepted
  18. 3D-LATTE: Latent Space 3D Editing from Textual Instructionsaccepted
  19. 3D-Object Perception Transformer (3PT)accepted
  20. 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decodingaccepted
  21. 3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objectsaccepted
  22. 3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experienceaccepted
  23. 3M-TI: High-Quality Mobile Thermal Imaging via Calibration-free Multi-Camera Cross-Modal Diffusionaccepted
  24. 4C4D: 4 Camera 4D Gaussian Splattingaccepted
  25. 4D Local Modeling Toward Dynamic Global Perception for Ambiguity-free Rotation-Invariant Point Cloud Analysisaccepted
  26. 4D Primitive-Mache: Glueing Primitives for Persistent 4D Scene Reconstructionaccepted
  27. 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillationaccepted
  28. 4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Videoaccepted
  29. 4DP-QA: Scalable QA for 4D Perception in Vision Language Modelsaccepted
  30. 4DSurf: High-Fidelity Dynamic Scene Surface Reconstructionaccepted
  31. 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Modelsaccepted
  32. A Bit is All You Need! Efficient Video Capture via Single Bit Imagingaccepted
  33. A Causal Marriage between VLM and IRM from Understanding to Reasoningaccepted
  34. A Closed-Form Solution for Debiasing Vision-Language Models with Utility Guarantees Across Modalities and Tasksaccepted
  35. A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helpsaccepted
  36. A Combination of Noise and Bilateral Filters Achieve Supralinear and Scalable Adversarial Robustness in CNNsaccepted
  37. A Cross-view Fusion Framework for Robust 6-DoF Grasp Pose Estimationaccepted
  38. A Debiased Reconstruction-based Framework for Training-Free Detection of AI-Generated Imagesaccepted
  39. A Difference-in-Difference Approach to Detecting AI-Generated Imagesaccepted
  40. A Faster Path to Continual Learningaccepted
  41. A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokensaccepted
  42. A Geometric Algebra-Informed 3DGS Framework for Wireless Channel Predictionaccepted
  43. A Mixed Diet Makes DINO An Omnivorous Vision Encoderaccepted
  44. A More Word-like Image Tokenization for MLLMsaccepted
  45. A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoningaccepted
  46. A Polynomial Chaos Framework for Causal Discovery in Nonlinear Uncertain Systemsaccepted
  47. A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language Modelsaccepted
  48. A Sanity Check for Multi-In-Domain Face Forgery Detection in the Real Worldaccepted
  49. A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generationaccepted
  50. A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detectionaccepted
  51. A Stitch in Time: Learning Procedural Workflow via Self-Supervised Plackett-Luce Rankingaccepted
  52. A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Spaceaccepted
  53. A Supervised Multi-task Framework for Joint cryo-ET Restoration Enabled by Generative Physical Simulationaccepted
  54. A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image Generationaccepted
  55. A Training-Free Style-Personalization via SVD-Based Feature Decompositionaccepted
  56. A Unified Framework for Knowledge Transfer in Bidirectional Model Scalingaccepted
  57. A Unified Perspective on Adversarial Membership Manipulation in Vision Modelsaccepted
  58. A2GC: Asymmetric Aggregation with Geometric Constraints for Locally Aggregated Descriptorsaccepted
  59. A3: Towards Advertising Aesthetic Assessmentaccepted
  60. ACE-Merging: Data-Free Model Merging with Adaptive Covariance Estimationaccepted
  61. ACPV-Net: All-Class Polygonal Vectorization for Seamless Vector Map Generation from Aerial Imageryaccepted
  62. ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Modelsaccepted
  63. AD-GBC: Anisotropic Granular-Ball Skip-Connection Refiner for UNet-Based Medical Image Segmentationaccepted
  64. ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoningaccepted
  65. AE2VID: Event-based Video Reconstruction via Aperture Modulationaccepted
  66. AERGS-SLAM: Auto-Exposure-Robust Stereo 3D Gaussian Splatting SLAMaccepted
  67. AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Modelsaccepted
  68. AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructionsaccepted
  69. AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Modelsaccepted
  70. AGiLe: Learning Robust Long-Horizon Manipulation via Affordance-Grounded Bidirectional Latent Planningaccepted
  71. AHS: Adaptive Head Synthesis via Synthetic Data Augmentationsaccepted
  72. AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimationaccepted
  73. AKCMamba-YOLO: Selective State Space Models For Real-Time Object Detectionaccepted
  74. ALLNet: Multi-task Dense Prediction for Degraded Imagesaccepted
  75. AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backendaccepted
  76. AMap: Distilling Future Priors for Ahead-Aware Online HD Map Constructionaccepted
  77. AMusE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understandingaccepted
  78. ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoningaccepted
  79. APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigationaccepted
  80. APPO: Attention-guided Perception Policy Optimization for Video Reasoningaccepted
  81. AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videosaccepted
  82. ARC Is a Vision Problem!accepted
  83. AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidanceaccepted
  84. ARES: Unifying Asymmetric RGB-Event Stereo for Probabilistic Scene Flow Estimationaccepted
  85. ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavioraccepted
  86. ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoningaccepted
  87. ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generationaccepted
  88. ART: Articulated Reconstruction Transformeraccepted
  89. AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Modelsaccepted
  90. AToken: A Unified Tokenizer for Visionaccepted
  91. AURA: Multi-modal Shared Autonomy for Urban Navigationaccepted
  92. AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMsaccepted
  93. AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Modelsaccepted
  94. AVA-VLA: Improving Vision-Language-Action models with Active Visual Attentionaccepted
  95. AVATAR: Reinforcement Learning to See, Hear, and Reason Over Videoaccepted
  96. AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMsaccepted
  97. AVGGT: Rethinking Global Attention for Accelerating VGGTaccepted
  98. AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Networkaccepted
  99. AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision-Language Modelsaccepted
  100. Abstract 3D Perception for Spatial Intelligence in Vision-Language Modelsaccepted
  101. AcTTA: Rethinking Test-Time Adaptation via Dynamic Activationaccepted
  102. Accelerating Autoregressive Video Diffusion via History-Guided Cache and Residual Correctionaccepted
  103. Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based Perspectiveaccepted
  104. Accelerating Diffusion via Hybrid Data-Pipeline Parallelism Based on Conditional Guidance Schedulingaccepted
  105. Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestepaccepted
  106. Accelerating Streaming Video Large Language Models via Hierarchical Token Compressionaccepted
  107. AceTone: Bridging Words and Colors for Conditional Image Gradingaccepted
  108. Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoningaccepted
  109. Act2See: Emergent Active Visual Perception for Video Reasoningaccepted
  110. ActAvatar: Temporally-Aware Precise Action Control for Talking Avatarsaccepted
  111. Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movementsaccepted
  112. Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulationaccepted
  113. Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic Manipulationaccepted
  114. ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusionaccepted
  115. Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Modelsaccepted
  116. Active Inference for Micro-Gesture Recognition: EFE-Guided Temporal Sampling and Adaptive Learningaccepted
  117. Active Intelligence in Video Avatars via Closed-loop World Modelingaccepted
  118. Active Perceptual Inference: A Corticothalamic-Inspired Dynamic Nested Recurrent Network for Multimodal Sentiment Analysis with Incomplete Dataaccepted
  119. ActiveAD: Planning-Oriented Active Learning for End-to-End Autonomous Drivingaccepted
  120. ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Modelaccepted
  121. ActivePolicy: Active Gaussian Reconstruction and Optimization Strategy Based on Global-Local Information Gainaccepted
  122. ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulationaccepted
  123. ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videosaccepted
  124. AdaBet: Gradient-free Layer Selection for Efficient Training of Deep Neural Networksaccepted
  125. AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generationaccepted
  126. AdaDexTrack: Dynamic Modulation for Adaptive and Generalizable Dexterous Manipulation Trackingaccepted
  127. AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLMaccepted
  128. AdaPrior: Bayesian-Inspired Adaptive Prior Correction for Long-Tailed Continual Learningaccepted
  129. AdaRadar: Rate Adaptive Spectral Compression for Radar-based Perceptionaccepted
  130. AdaSFormer: Adaptive Serialized Transformers for Monocular Semantic Scene Completion from Indoor Environmentsaccepted
  131. AdaSVD: Singular Value Decomposition with Adaptive Mechanisms for Large Multimodal Modelsaccepted
  132. AdaSpark: Adaptive Sparsity for Efficient Long-Video Understandingaccepted
  133. AdaSpot: Spend Resolution Where It Matters for Precise Event Spottingaccepted
  134. AdapAction: Adaptive Target Action Backdoor Attack against GUI Agentsaccepted
  135. AdapTok: Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Spaceaccepted
  136. AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisitionaccepted
  137. Adapter Shield: A Unified Framework with Built-in Authentication for Preventing Unauthorized Zero-Shot Image-to-Image Generationaccepted
  138. Adapting In-context Generation for Enhanced Composed Image Retrievalaccepted
  139. Adapting Lightweight Image-based Counting Models for Video Crowd Countingaccepted
  140. Adapting Point Cloud Analysis via Multimodal Bayesian Distribution Learningaccepted
  141. Adapting a Pre-trained Single-Cell Foundation Model to Spatial Gene Expression Generation from Histology Imagesaccepted
  142. Adaptive 3D Perception for Small Aerial Targets Under Sparse Sampling via Reinforcement Learningaccepted
  143. Adaptive Action Chunking at Inference-time for Vision-Language-Action Modelsaccepted
  144. Adaptive Anisotropic Gaussian Splatting for Multi-contrast MRI Arbitrary-Scale Super-Resolution with Anatomy Guidanceaccepted
  145. Adaptive Auxiliary Prompt Blending for Target-Faithful Diffusion Generationaccepted
  146. Adaptive Bayesian Early-Exit Networks for Efficient Non-Transferable Learningaccepted
  147. Adaptive Capacity Autoregressive Visual Trackingaccepted
  148. Adaptive Confidence Regularization for Multimodal Failure Detectionaccepted
  149. Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern Recognitionaccepted
  150. Adaptive Depth Lightweight RGB-T Tracking with Holistic Token Routingaccepted
  151. Adaptive Learned Image Compression with Graph Neural Networksaccepted
  152. Adaptive Spatial-Temporal Window: Unlocking the Potential of Event Cameras in Heterogeneous Velocity Scenariosaccepted
  153. Adaptive Spectral Feature Forecasting for Diffusion Sampling Accelerationaccepted
  154. Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generationaccepted
  155. Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learningaccepted
  156. AdvFM: Lookahead Flow-Matching Velocity-Field Attacks for Imperceptible and Transferable Adversarial Examplesaccepted
  157. Advancing Cancer Prognosis with Hierarchical Fusion of Genomic, Proteomic and Pathology Imaging Data from a Systems Biology Perspectiveaccepted
  158. Advancing Image Classification with Discrete Diffusion Classification Modelingaccepted
  159. Adversarial Style Optimization: Enhancing VLM Jailbreaks by GRPO-based Stylistic Triggers Optimizationaccepted
  160. AeroAgent: A Vision-Physics-Decision Framework for Aerodynamic Vehicle Designaccepted
  161. AeroDGS: Physically Consistent Dynamic Gaussian Splatting for Single-Sequence Aerial 4D Reconstructionaccepted
  162. AeroGS: Scale-Aware Gaussian Splatting for Pose-Free Dynamic UAV Scene Reconstructionaccepted
  163. Aesthetic Camera Viewpoint Suggestion with 3D Aesthetic Fieldaccepted
  164. Affine Perspective-Three-Point Problemaccepted
  165. AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondenceaccepted
  166. AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesisaccepted
  167. AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiersaccepted
  168. Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulationaccepted
  169. Affordance-First Decomposition for Continual Learning in Video-Language Understandingaccepted
  170. Affostruction: 3D Affordance Grounding with Generative Reconstructionaccepted
  171. Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detectionaccepted
  172. AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detectionaccepted
  173. Agentic Retoucher for Text-To-Image Generationaccepted
  174. Agentic Video Summarization via Self-Reflecting Multimodal Understandingaccepted
  175. Agile Deliberation: Concept Deliberation for Subjective Visual Classificationaccepted
  176. Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrievalaccepted
  177. AirSim360: A Panoramic Simulation Platform within Drone Viewaccepted
  178. AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generationaccepted
  179. Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly Detectionaccepted
  180. Align Images Before You Generateaccepted
  181. Align Once to Explain: Feature Alignment for Scalable B-cosification of Foundational Vision Transformersaccepted
  182. Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied Agentsaccepted
  183. AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignmentaccepted
  184. Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human Preferencesaccepted
  185. Aligning Text, Images and 3D Structure Token-by-Tokenaccepted
  186. Aligning What Vision-Language Models See and Perceive with Adaptive Information Flowaccepted
  187. All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Modelsaccepted
  188. All Vehicles Can Lie: Efficient Adversarial Defense in Fully Untrusted-Vehicle Collaborative Perception via Pseudo-Random Bayesian Inferenceaccepted
  189. All in One: Unifying Deepfake Detection, Tampering Localization, and Source Tracing with a Robust Landmark-Identity Watermarkaccepted
  190. All-in-One Slider for Attribute Manipulation in Diffusion Modelsaccepted
  191. An Efficient Token Compression Framework for Visual Object Trackingaccepted
  192. An Empirical Study on How Video-LLMs Answer Video Questionsaccepted
  193. An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Drivingaccepted
  194. An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learningaccepted
  195. Anatomica: Localized Control over Geometric and Topological Properties for Anatomical Diffusion Modelsaccepted
  196. Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose Predictionaccepted
  197. Anchor-Guided Gradient Alignment for Incomplete Multimodal Learningaccepted
  198. AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flowsaccepted
  199. AnchorSplat: Feed-Forward 3D Gaussian Splatting With 3D Geometric Priorsaccepted
  200. Anchoring and Rescaling Attention for Semantically Coherent Inbetweeningaccepted
  201. Anchoring the Mind of Multimodal Reasoners: Cognitive Bias as a Vector for Jailbreak Attacksaccepted
  202. Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Samplingaccepted
  203. AniMimic: Imitating 3D Animation from Video Priorsaccepted
  204. Animator-Centric Skeleton Generation on Objects with Fine-Grained Detailsaccepted
  205. Annotation-Efficient Coreset Selection for Context-dependent Segmentationaccepted
  206. Anomaly as Non-Conformity via Training-Free Graph Laplacian Energy Minimizationaccepted
  207. Anomaly-Related Residual Fields for Cross-domain Anomaly Detectionaccepted
  208. AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectorsaccepted
  209. AnthroTAP: Learning Point Tracking with Real-World Motionaccepted
  210. Anti-Degradation Lifelong Multi-View Clusteringaccepted
  211. Anti-I2V: Safeguarding your Photos from Malicious Image-to-video Generationaccepted
  212. AntiStyler: Defending Object Detection Models Against Adversarial Patch Attacks Using Style Removalaccepted
  213. Any Resolution Any Geometry: From Multi-View To Multi-Patchaccepted
  214. Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Studyaccepted
  215. Any4D: Unified Feed-Forward Metric 4D Reconstructionaccepted
  216. AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement Optimizationaccepted
  217. AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual Referencesaccepted
  218. AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusionaccepted
  219. AnyPcc: Compressing Any Point Cloud with a Single Universal Modelaccepted
  220. ApET: Approximation-Error Guided Token Compression for Efficient VLMsaccepted
  221. Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generationaccepted
  222. Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learningaccepted
  223. ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wildaccepted
  224. Archon: A Unified Multimodal Model for Holistic Digital Human Generationaccepted
  225. Are Image-to-Video Models Good Zero-Shot Image Editors?accepted
  226. Are We Ready for RL in Text-to-3D Generation? A Progressive Investigationaccepted
  227. ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactionsaccepted
  228. ArtLLM: Generating Articulated Assets via 3D LLMaccepted
  229. ArtPro: Self-Supervised Articulated Object Reconstruction with Adaptive Integration of Mobility Proposalsaccepted
  230. ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understandingaccepted
  231. Artiverse: A Diverse and Physically Grounded Dataset for Articulated Objectsaccepted
  232. Asking like Socrates: Socrates helps VLMs understand remote sensing imagesaccepted
  233. AssemblyBench: Physics-Aware Assembly of Complex Industrial Objectsaccepted
  234. Assignment-Driven Hash Learning in a Hyper-Semantic Space for On-the-Fly Category Discoveryaccepted
  235. AstraNav-Memory: Contexts Compression for Long Memoryaccepted
  236. AsymLoc: Towards Asymmetric Feature Matching for Efficient Visual Localizationaccepted
  237. Asynchronous Temporal Modeling with Two-Agent Framework for Streaming Dense Video Captioningaccepted
  238. AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robotsaccepted
  239. Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything Modelaccepted
  240. Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazingaccepted
  241. Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformeraccepted
  242. Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Modelsaccepted
  243. Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decodingaccepted
  244. Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swappingaccepted
  245. Attribution as Retrieval: Model-Agnostic AI-Generated Image Attributionaccepted
  246. Attribution-Guided Model Rectification of Unreliable Neural Network Behaviorsaccepted
  247. Audio-sync Video Instance Editing with Granularity-Aware Mask Refineraccepted
  248. AudioAvatar: Personalized Audio-driven Whole-body Talking Avatarsaccepted
  249. AudioStory: Generating Long-Form Narrative Audio with Large Language Modelsaccepted
  250. Authorize-on-Demand: Dynamic Authorization with Legality-Aware Intellectual Property Protection for VLMsaccepted
  251. AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generationaccepted
  252. AutoDebias: An Automated Framework for Detecting and Mitigating Backdoor Biases in Text-to-Image Modelsaccepted
  253. AutoRegressive Generation with B-rep Holistic Token Sequence Representationaccepted
  254. AutoTraces: Autoregressive Trajectory Forecasting via Multimodal Large Language Modelsaccepted
  255. Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversationaccepted
  256. AvatarPointillist: AutoRegressive 4D Gaussian Avatarizationaccepted
  257. AviaSafe: A Physics-Informed Data-Driven Model for Aviation Safety-Critical Cloud Forecastsaccepted
  258. AwareVLN: Reasoning with Self-awareness for Vision-Language Navigationaccepted
  259. B$^3$-Seg: Camera-Free, Training-Free 3DGS Segmentation via Analytic EIG and Beta-Bernoulli Bayesian Updatesaccepted
  260. BA-GS: Bayesian Adaptive Gaussian Splatting for SFM-Free 3D Reconstructionaccepted
  261. BALM: A Model-Agnostic Framework for Balanced Multimodal Learning under Imbalanced Missing Ratesaccepted
  262. BAMI: Training-Free Bias Mitigation in GUI Groundingaccepted
  263. BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Modelsaccepted
  264. BD-Merging: Bias-Aware Dynamic Model Merging with Evidence-Guided Contrastive Learningaccepted
  265. BDNet:Bio-Inspired Dual-Backbone Small Object Detection Networkaccepted
  266. BEA-GS: BEyond RAdiance Supervision in 3DGS for Precise Object Extractionaccepted
  267. BEV-CAR: Enhancing Monocular Bird's Eye View Segmentation with Context-Aware Rasterizationaccepted
  268. BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Imagesaccepted
  269. BHCast: Unlocking Black Hole Plasma Dynamics from a Single Blurry Image with Long-Term Forecastingaccepted
  270. BIT: Matching-based Bi-directional Interaction Transformation Network for Visible-Infrared Person Re-Identificationaccepted
  271. BOP-ASK: Object-Interaction Reasoning for Vision-Language Modelsaccepted
  272. BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship Detectionaccepted
  273. BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Modelsaccepted
  274. Back to Basics: Let Denoising Generative Models Denoiseaccepted
  275. Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detectionaccepted
  276. Back to Source: Open-Set Continual Test-Time Adaptation via Domain Compensationaccepted
  277. Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanationsaccepted
  278. BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentationaccepted
  279. Balanced Dataset Distillation via Modeling Multiple Visual Pattern Distributionaccepted
  280. Balanced Hierarchical Contrastive Learning with Decoupled Queries for Fine-grained Object Detection in Remote Sensing Imagesaccepted
  281. BarbieGait: An Identity-Consistent Synthetic Human Dataset with Versatile Cloth-Changing for Gait Recognitionaccepted
  282. Basis-Oriented Low-rank Transfer for Few-Shot and Test-Time Adaptationaccepted
  283. Batch Loss Score for Dynamic Data Pruningaccepted
  284. Batman: Benign Knowledge Alignment Through Malicious Null Space in Federated Backdoor Attackaccepted
  285. Bayesian Decomposition and Semantic Completion for Few-shot Semantic Segmentationaccepted
  286. BeautyGRPO: Aesthetic Alignment for Face Retouching via Dynamic Path Guidance and Fine-Grained Preference Modelingaccepted
  287. Benchmarking Endoscopic Surgical Image Restoration and Beyondaccepted
  288. Benchmarking PhD-Level Coding in 3D Geometric Computer Visionaccepted
  289. Benchmarking Single-Factor Physical Video-to-Audio Generationaccepted
  290. Best Segmentation Buddies for Image-Shape Correspondenceaccepted
  291. Better than Average: Spatially-Aware Aggregation of Segmentation Uncertainty Improves Downstream Performanceaccepted
  292. Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Predictionaccepted
  293. Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoningaccepted
  294. Beyond Appearance: Camouflaged Object Detection via Geometric Structureaccepted
  295. Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchorsaccepted
  296. Beyond Caption-Based Queries in Video Moment Retrievalaccepted
  297. Beyond Duality: A Hybrid Framework of Leveraging Shared and Private Features for RGB-Event Object Detectionaccepted
  298. Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extractionaccepted
  299. Beyond Euclidean Gossip: KL-Barycentric Consensus on Heterogeneous and Imbalanced Imagesaccepted
  300. Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object Trackingaccepted
  301. Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Modelsaccepted
  302. Beyond Geometry: Artistic Disparity Synthesis for Immersive 2D-to-3Daccepted
  303. Beyond Global Similarity: Multi-Conditional Retrieval for Fine-Grained Cross-Modal Understandingaccepted
  304. Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapteraccepted
  305. Beyond Ground-Truth: Leveraging Image Quality Priors for Real-World Image Restorationaccepted
  306. Beyond Heuristic Prompting: A Concept-Guided Bayesian Framework for Zero-Shot Image Recognitionaccepted
  307. Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language Modelsaccepted
  308. Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigationaccepted
  309. Beyond Mimicry: Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrationsaccepted
  310. Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognitionaccepted
  311. Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFTaccepted
  312. Beyond Myopic Alignment: Lookahead Optimization for Online Class-Incremental Learningaccepted
  313. Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classificationaccepted
  314. Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generationaccepted
  315. Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMsaccepted
  316. Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Controlaccepted
  317. Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object Detectionaccepted
  318. Beyond Reassembly: Fractured Object Recovery with Missing Partsaccepted
  319. Beyond Rule-Based Agents: Active Markov Games for Realistic Multi-Agent Interaction in Autonomous Drivingaccepted
  320. Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic Scenesaccepted
  321. Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrievalaccepted
  322. Beyond Sequential Tools: A Unified VLM Agent System for Photographic Post-Processing via Dynamic Multi-Expert Fusionaccepted
  323. Beyond Single Images: A Comprehensive Benchmark for Album-Level Vision-Language Understandingaccepted
  324. Beyond Single Solution: Multi-Hypothesis Deep Unfolding Network for Image Compressive Sensingaccepted
  325. Beyond Single-View Sufficiency: CVBench for Cross-View Human Understandingaccepted
  326. Beyond Soft Label: Dataset Distillation via Orthogonal Gradient Matchingaccepted
  327. Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimationaccepted
  328. Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image Fusionaccepted
  329. Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Interventionaccepted
  330. Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaborationaccepted
  331. Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic Segmentationaccepted
  332. Beyond Tie Points: Satellite Image Block Adjustment based on Dense Feature Consistencyaccepted
  333. Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretabilityaccepted
  334. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detectionaccepted
  335. Beyond What's Shared: Recovering Lost Unique Information from Intermediate Layers to Boost Multimodal Geo-Foundation Modelsaccepted
  336. Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake Detectionaccepted
  337. Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinationsaccepted
  338. Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Trainingaccepted
  339. Beyond the Ground Truth: Enhanced Supervision for Image Restorationaccepted
  340. Beyond the Static World: Continual Category Discovery under Visual Driftaccepted
  341. Beyond the Static-World: Lifelong Learning for All-in-One Medical Image Restorationaccepted
  342. Bezier Degradation Modeling for LiDAR-based Human Motion Captureaccepted
  343. Bi-Bridge: Bidirectional Diffusion Bridges for Low-Light Image Enhancementaccepted
  344. Bi-directional Autoregressive Diffusion for Large Complex Motion Interpolationaccepted
  345. BiEvLight: Bi-level Learning of Task-Aware Event Refinement for Low-Light Image Enhancementaccepted
  346. BiFM: Bidirectional Flow Matching for Few-Step Image Editing and Generationaccepted
  347. BiGMINT: Biologically-guided Hierarchical Multimodal Integration for Modeling Multiple Compound Activities in Drug Discoveryaccepted
  348. BiGain: Unified Token Compression for Joint Generation and Classificationaccepted
  349. BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generationaccepted
  350. BiOTPrompt: Bidirectional Optimal Transport Guided Prompting for Disease Evolution-aware Radiology Report Generationaccepted
  351. BiPA: Bilevel Prompt Adaptation for Underwater Instance Segmentationaccepted
  352. BiPreManip: Learning Affordance-Based Bimanual Preparatory Manipulation through Anticipatory Collaborationaccepted
  353. BiProLoRA: Bilevel Prompt LoRA for Real Scene Recoveryaccepted
  354. Bias In, Bias Out? Finding Unbiased Subnetworks in Vanilla Modelsaccepted
  355. Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Modelsaccepted
  356. Bias at the End of the Scoreaccepted
  357. Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereoaccepted
  358. Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detectionaccepted
  359. Bidirectional Normalizing Flow: From Data to Noise and Backaccepted
  360. Bidirectional Query-Driven Generation of Parametric CAD Sketchaccepted
  361. Bilevel Layer-Positioning LoRA for Real Image Dehazingaccepted
  362. BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformersaccepted
  363. BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignmentaccepted
  364. BiomedCCPL: Causal Conditional Prompt Learning for Biomedical Vision-Language Modelsaccepted
  365. Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compressionaccepted
  366. Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Modelsaccepted
  367. BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviationaccepted
  368. Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understandingaccepted
  369. Block-Sparse Global Attention for Efficient Multi-View Geometry Transformersaccepted
  370. Block-based Learned Image Compression without Blocking Artifactsaccepted
  371. BluRef: Unsupervised Image Deblurring with Dense-Matching Referencesaccepted
  372. BoostSLT: Boosting Sign Language Translation via a Plug-and-Play Diffusion-Based Semantic Enhanceraccepted
  373. Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processingaccepted
  374. Boosting Quantitive and Spatial Awareness for Zero-Shot Object Countingaccepted
  375. Boosting Reasoning in Large Multimodal Models via Activation Replayaccepted
  376. Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learningaccepted
  377. Boosting Vision-Language Models Towards Cross-Domain Incremental Object Detectionaccepted
  378. Boosting Vision-Language-Action Finetuning with Feasible Action Neighborhood Prioraccepted
  379. Boosting Visual Reprogramming for CLIP with Dual Granularity Alignmentaccepted
  380. Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learningaccepted
  381. Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual Segmentationaccepted
  382. Bootstrapping Multi-view Learning for Test-time Noisy Correspondenceaccepted
  383. Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptationaccepted
  384. Boundary-Responsive Differentiable Gating for Superpixel-Based Segmentationaccepted
  385. Breaking Multimodal LLM Safety via Video-Driven Promptingaccepted
  386. Breaking Semantic Boundaries: Distribution-Guided Semantic Exploration for Creative Generationaccepted
  387. Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditionsaccepted
  388. Breaking Spurious Correlations: Uncertainty-Driven Causal Transformers for AU Detectionaccepted
  389. Breaking the 3D Dataset Bottleneck: Fast Scalable Generation of Aligned 3D Assets from Scratch for Category 6D Pose Estimation and Robotic Graspingaccepted
  390. Breaking the Continuum: Discrete Distribution Learning for Structural MRI Reconstructionaccepted
  391. Breaking the Illusion: When Positive Meets Negative in Multimodal Decodingaccepted
  392. Breaking the Regional Perception Bottleneck of Multimodal Large Language Models via External Reasoning Frameworkaccepted
  393. Breaking the Scalability Limit of Multi-Projector Calibration with Embedded Camerasaccepted
  394. BrepGaussian: CAD reconstruction from Multi-View Images with Gaussian Splattingaccepted
  395. BrepVGAE: Variational Graph Autoencoder with Unified Latent Representation for B-repaccepted
  396. Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observationaccepted
  397. BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality Assessmentaccepted
  398. BrickNet: Graph-Backed Generative Brick Assemblyaccepted
  399. Bridge: Basis-Driven Causal Inference Marries VFMs for Domain Generalizationaccepted
  400. BridgeEQA: Virtual Embodied Agents for Real Bridge Inspectionsaccepted
  401. Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstructionaccepted
  402. Bridging Domain Expertise and Generalization for Performance Estimationaccepted
  403. Bridging Domains through Subspace-Aware Model Mergingaccepted
  404. Bridging Facial Understanding and Animation via Language Modelsaccepted
  405. Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-Resolutionaccepted
  406. Bridging Human Evaluation to Infrared and Visible Image Fusionaccepted
  407. Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verificationaccepted
  408. Bridging Privacy and Provenance: Traceable Virtual Identity Generationaccepted
  409. Bridging RGB and Hematoxylin Components: An Interleaved Guidance and Fusion Framework for Point Supervised Nuclei Segmentationaccepted
  410. Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigationaccepted
  411. Bridging the Modality Gap in Compositional Zero-Shot Learning via Sparse Alignment and Unimodal Memory Bankaccepted
  412. Bridging the Perception Gap in Image Super-Resolution Evaluationaccepted
  413. Bringing Your Portrait to 3D Presenceaccepted
  414. BuildAnyPoint: 3D Building Structured Abstraction from Diverse Point Cloudsaccepted
  415. Building Robust Vision Encoders for Cross-Dataset Evaluation in Immunofluorescent Microscopyaccepted
  416. Building a Precise Video Language with Human-AI Oversightaccepted
  417. BuildingGPT: Auto-Regressive Building Wireframe Reconstruction Model with Reinforcement Learningaccepted
  418. Bulk RNA-seq Guided Multi-modal Detection of Anomalous Regions in Human Cancer via Spatial Transcriptomicsaccepted
  419. BulletTime: Decoupled Control of Time and Camera Pose for Video Generationaccepted
  420. Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal Transportaccepted
  421. C-GenReg: Training-Free 3D Point Cloud Registration by Multi-View-Consistent Geometry-to-Image Generation with Probabilistic Modalities Fusionaccepted
  422. C-LaV: Conditional Latent Velocity Field Denoising for Weather-Robust LiDAR Place Recognitionaccepted
  423. CAD-Refiner: A Unified Framework for CAD Generation and Iterative Editingaccepted
  424. CADC: Content Adaptive Diffusion-Based Generative Image Compressionaccepted
  425. CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Modelsaccepted
  426. CAPT: Confusion-Aware Prompt Tuning for Reducing Vision-Language Misalignmentaccepted
  427. CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Modelaccepted
  428. CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topographyaccepted
  429. CARD: Correlation Aware Restoration with Diffusionaccepted
  430. CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoningaccepted
  431. CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editingaccepted
  432. CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysisaccepted
  433. CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interactionaccepted
  434. CARLoS: Retrieval via Concise Assessment Representation of LoRAs at Scaleaccepted
  435. CASPA: Graph-Structured Concept Anchors for Modality-Agnostic Adaptation in Vision-Language Modelsaccepted
  436. CASR: A Robust Cyclic Framework for Arbitrary Large-Scale Super-Resolution with Distribution Alignment and Self-Similarity Awarenessaccepted
  437. CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image Retrievalaccepted
  438. CATNet: Collaborative Alignment and Transformation Network for Cooperative Perceptionaccepted
  439. CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answeringaccepted
  440. CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioningaccepted
  441. CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detectionaccepted
  442. CD-Buffer: Complementary Dual-Buffer Framework for Test-Time Adaptation in Adverse Weather Object Detectionaccepted
  443. CDICS: Delving Into Fine-Grained Attribute for In-Context Segmentation via Compositional Prompts and Phased Decouplingaccepted
  444. CF-IPT: Cross-Modal Fusion Interactive Prompt Tuning of Vision-Language Pre-Trained Model for Multisource Remote Sensing Data Classificationaccepted
  445. CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidanceaccepted
  446. CG-Floor: Centroid-Guided Diffusion for Large-Scale Floorplan Generationaccepted
  447. CG-Reasoner: Centroid-Guided Positional Reasoning Segmentation for Medical Imaging with a Robust Visual-Text Consistency Metricaccepted
  448. CGHair: Compact Gaussian Hair Reconstruction with Card Clusteringaccepted
  449. CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuningaccepted
  450. CGU-Bayes: Causal Graph Uncertainty-Guided Bayesian Inference for Domain Generalizationaccepted
  451. CHAL: Causal-guided Hierarchical Anomaly-aware Learning for Moving Infrared Small Target Detectionaccepted
  452. CHEEM: Continual Learning by Reuse, New, Adapt and Skip - A Hierarchical Exploration-Exploitation Approachaccepted
  453. CHIPS: Efficient CLIP Adaptation via Curvature-aware Hybrid Influence-based Data Selectionaccepted
  454. CHIRP dataset: towards long-term, individual-level, behavioral monitoring of bird populations in the wildaccepted
  455. CI-VID: A Coherent Interleaved Text-Video Datasetaccepted
  456. CICA: Coupling Confidence-Aware Pretraining with Confidence-Informed Attention for Robust Multimodal Sentiment Analysisaccepted
  457. CIGMA: Causal Information-Gain Mechanistic Attribution of Attention Heads in Vision Transformersaccepted
  458. CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimationaccepted
  459. CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Spaceaccepted
  460. CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learningaccepted
  461. CLEP: Contrastive Language-Pose Pretrainingaccepted
  462. CLEX: Complementary Label Exchange Learning for Noisy Facial Expression Recognitionaccepted
  463. CLIP Is Shortsighted: Paying Attention Beyond the First Sentenceaccepted
  464. CLIP-like Model as a Foundational Density Ratio Estimatoraccepted
  465. CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptationaccepted
  466. CLP: A Real-World Dataset of Contaminated Lens Protectors for Robust Semantic Segmentationaccepted
  467. CLaD: Planning with Grounded Foresight via Cross-Modal Latent Dynamicsaccepted
  468. CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoningaccepted
  469. CME-CAD: Heterogeneous Collaborative Multi-Expert Reinforcement Learning for CAD Code Generationaccepted
  470. CMR-RD: Long-Tailed Adaptive VLM for Explainable CMR Diagnosisaccepted
  471. COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimationaccepted
  472. COPE: Consistent Occlusion and Prompt Enhancement Network for Occluded Person Re-identificationaccepted
  473. COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMsaccepted
  474. COPYLENS: Towards Copyrighted Characters Infringement Detection via Copyright-Aware Prompt Learningaccepted
  475. CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMsaccepted
  476. COT-FM: Cluster-wise Optimal Transport Flow Matchingaccepted
  477. CRAFT-LoRA: Content-Style Personalization via Rank-Constrained Adaptation and Training-Free Fusionaccepted
  478. CRAFT: Aligning Diffusion Models with Fine-Tuning Is Easier Than You Thinkaccepted
  479. CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructionsaccepted
  480. CREward: A Type-Specific Creativity Reward Modelaccepted
  481. CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registrationaccepted
  482. CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop Reasoningaccepted
  483. CROWn: A Unified Framework for Anti-Aliased Downsampling and Phase-Calibrated Fusion in 3D Medical Segmentationaccepted
  484. CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Modelsaccepted
  485. CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibrationaccepted
  486. CUBic: Coordinated Unified Bimanual Perception and Control Frameworkaccepted
  487. CUE: Concept-Aware Multi-Label Expansion to Mitigate Concept Confusion in Long-Tailed Learningaccepted
  488. CUPID: Generative 3D Reconstruction via Joint Object and Pose Modelingaccepted
  489. CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generationaccepted
  490. CURVE: A Benchmark for Cultural and Multilingual Long Video Reasoningaccepted
  491. CVA: Context-aware Video-text Alignment for Video Temporal Groundingaccepted
  492. C^2FG: Control Classifier-Free Guidance via Score Discrepancy Analysisaccepted
  493. CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusionaccepted
  494. CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answeringaccepted
  495. CaT-GS: Efficient 3DGS Rendering for Large-Scale Scenes with Inter-frame Caching and Tile Schedulingaccepted
  496. CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenizationaccepted
  497. CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generationaccepted
  498. CamDirector: Towards Long-Term Coherent Video Trajectory Editingaccepted
  499. CamPI: Physical Adversarial Examples through Camera Power Signal Injectionaccepted
  500. Camera Control for Text-to-Image Generation via Learning Viewpoint Tokensaccepted
  501. Camouflage-aware Image-Text Retrieval via Expert Collaborationaccepted
  502. Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling?accepted
  503. Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matchingaccepted
  504. Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformersaccepted
  505. Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detectionaccepted
  506. CanonCGT: Reference-Based Color Grading via Canonical Pivot Representationaccepted
  507. CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigationaccepted
  508. Captain Safari: A World Engine with Pose-Aligned 3D Memoryaccepted
  509. CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objectsaccepted
  510. CaptionQA: Is Your Caption as Useful as the Image Itself?accepted
  511. CaricHarmony: Contrastive Diffusion Paths for Identity-Preserving Caricature Synthesisaccepted
  512. Catalyst4D: High-Fidelity 3D-to-4D Scene Editing via Dynamic Propagationaccepted
  513. Causal Motion Diffusion Models for Autoregressive Motion Generationaccepted
  514. CausalLens: Sensitivity-Guided Multi-Head Causal Intervention for Hallucination Mitigation in Large Vision-Language Modelsaccepted
  515. CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Interventionaccepted
  516. Causality in Video Diffusers is Separable from Denoisingaccepted
  517. Cell-Type Prototype-Informed Neural Network for Gene Expression Estimation from Pathology Imagesaccepted
  518. ChArtist: Generating Pictorial Charts with Unified Spatial and Subject Controlaccepted
  519. Chain of Event-Centric Causal Thought for Physically Plausible Video Generationaccepted
  520. Chain of World: World Model Thinking in Latent Motionaccepted
  521. Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoningaccepted
  522. Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Modelsaccepted
  523. Chain-of-Thought Guided Multi-Modal Object Re-Identificationaccepted
  524. ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Senisngaccepted
  525. Changes in Real Time: Online Scene Change Detection with Multi-View Fusionaccepted
  526. Charge: A Comprehensive Novel View Synthesis Benchmark and Dataset to Bind Them Allaccepted
  527. Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Chartsaccepted
  528. ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understandingaccepted
  529. ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question Answeringaccepted
  530. ChimeraLoRA: Multi-Head LoRA-Guided Synthetic Datasetsaccepted
  531. ChordEdit: One-Step Low-Energy Transport for Image Editingaccepted
  532. Choreographing a World of Dynamic Objectsaccepted
  533. Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encodingaccepted
  534. ChronoGS: Disentangling Invariants and Changes in Multi-Period Scenesaccepted
  535. CineBrain: A Large-Scale Multi-Modal Audiovisual Brain Dataset for Brain-Conditioned Video Generationaccepted
  536. CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarizationaccepted
  537. CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generationaccepted
  538. Cinematic Audio Source Separation Using Visual Cuesaccepted
  539. Circuit Mechanisms for Spatial Relation Generation in Diffusion Transformersaccepted
  540. Circular-DPO: Aligning Multi-Stage 3D Generative Models via Preference Feedback Loopaccepted
  541. Clair Obscur: an Illumination-Aware Method for Real-World Image Vectorizationaccepted
  542. Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation Modelingaccepted
  543. Cleaning the Pool: Progressive Filtering of Unlabeled Pools in Deep Active Learningaccepted
  544. ClimaOoD: Improving Anomaly Segmentation via Physically Realistic Synthetic Dataaccepted
  545. Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosisaccepted
  546. ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene Reconstructionaccepted
  547. Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Modelsaccepted
  548. Closed-Form Concept Erasure via Double Projectionsaccepted
  549. Clothe and Poseaccepted
  550. Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Modelsaccepted
  551. Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretrainingaccepted
  552. Cluster-aware Anchor Learning for Multi-View Clusteringaccepted
  553. ClusterMark: Towards Robust Watermarking for Autoregressive Image Generators with Visual Token Clusteringaccepted
  554. Co-Me: Confidence Guided Token Merging for Visual Geometric Transformersaccepted
  555. CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detectionaccepted
  556. CoD: A Diffusion Foundation Model for Image Compressionaccepted
  557. CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inferenceaccepted
  558. CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detectionaccepted
  559. CoIn: Coverage and Informativeness-Guided Token Reduction for Efficient Large Multimodal Modelsaccepted
  560. CoLC: Communication-Efficient Collaborative Perception with LiDAR Completionaccepted
  561. CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generationaccepted
  562. CoLoR: The Devil is in Scene Coordinate Regression for Large-Scale Visual Localizationaccepted
  563. CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learningaccepted
  564. CoRiM: Conflict-driven Risk Minimization for Dynamic Multimodal Fusionaccepted
  565. CoRoGS: Contextual Gaussian Splatting for Robust Large-Deviation View Synthesisaccepted
  566. CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial Modelingaccepted
  567. CoT-Edit: Let CoT Guide Instruction Video Editingaccepted
  568. CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics Priorityaccepted
  569. CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Modelsaccepted
  570. CoWTracker: Tracking by Warping instead of Correlationaccepted
  571. CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoningaccepted
  572. CodeMMR: Bridging Natural Language, Code, and Image for Unified Retrievalaccepted
  573. CodePercept: Code-Grounded Visual STEM Perception for MLLMsaccepted
  574. CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimizationaccepted
  575. Coded-E2LF: Coded Aperture Light Field Imaging from Eventsaccepted
  576. CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Drivingaccepted
  577. CogniEdit: Dense Gradient Flow Optimization for Fine-Grained Image Editingaccepted
  578. CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoningaccepted
  579. ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Drivingaccepted
  580. Collaborative Multi-Mode Pruning for Vision-Language Modelsaccepted
  581. Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensingaccepted
  582. Color-Encoded Illumination for High-Speed Volumetric Scene Reconstructionaccepted
  583. ColorFLUX: A Structure-Color Decoupling Framework for Old Photo Colorizationaccepted
  584. ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimationaccepted
  585. Common Inpainted Objects In-N-Out of Contextaccepted
  586. CompBench: Benchmarking Complex Instruction-guided Image Editingaccepted
  587. CompetitorFormer: Mitigating Query Conflicts for 3D Instance Segmentation via Competitive Strategyaccepted
  588. Complementary Prototype Mapping for Efficient Multimodal Anomaly Detectionaccepted
  589. Complet4R: Geometric Complete 4D Reconstructionaccepted
  590. Composing Concepts from Images and Videos via Concept-prompt Bindingaccepted
  591. Composite-Attribute Person Re-Identification via Pose-Guided Disentanglementaccepted
  592. Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimizationaccepted
  593. Compositional Transformation Reasoning for Composed Video Retrievalaccepted
  594. Compressed-Domain-Aware Online Video Super-Resolutionaccepted
  595. Computation and Communication Efficient Federated Unlearning via On-server Gradient Conflict Mitigation and Expressionaccepted
  596. Computational Speckle Pattern Interferometryaccepted
  597. Computer Vision with a Superpixelation Cameraaccepted
  598. Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidenceaccepted
  599. Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approachaccepted
  600. Concept-Aware Batch Sampling Improves Language-Image Pretrainingaccepted
  601. Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generationaccepted
  602. Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustnessaccepted
  603. ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectorsaccepted
  604. ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token Optimizationaccepted
  605. Condensed Test-Time Adaptation of VLMs for Action Recognitionaccepted
  606. Conditional Factuality Controlled LLMs with Generalization Certificates via Conformal Samplingaccepted
  607. ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrievalaccepted
  608. Confidence-Guided Multi-Scale Aggregation for Sparse-View High-Resolution 3D Gaussian Splattingaccepted
  609. Conflict-Aware Adaptive Cross-Reconstruction for Multimodal Sentiment Analysisaccepted
  610. Confusion-Aware Spectral Regularizer for Long-Tailed Recognitionaccepted
  611. ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generationaccepted
  612. Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCRaccepted
  613. Consensus vs. Controversy: Mapping the Decision Space Where Architectures Divergeaccepted
  614. ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulationaccepted
  615. ConsistCompose: Unified Multimodal Layout Control for Image Compositionaccepted
  616. Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency Learningaccepted
  617. Consistent Instance Field for Dynamic Scene Understandingaccepted
  618. Contact-Aware Neural Dynamicsaccepted
  619. Content-Adaptive Hierarchical Hyperprior for Neural Video Codingaccepted
  620. Content-Aware Dynamic Patchification for Efficient Video Diffusionaccepted
  621. Content-Aware Frequency Encoding for Implicit Neural Representations with Fourier-Chebyshev Featuresaccepted
  622. Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigationaccepted
  623. Continual Distillation of Teachers from Different Domainsaccepted
  624. Continual Learning for fMRI-Based Brain Disorder Diagnosis via Functional Connectivity Matrices Generative Replayaccepted
  625. Continuous Exposure-Time Modeling for Realistic Atmospheric Turbulence Synthesisaccepted
  626. Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classificationaccepted
  627. Controllable Federated Prompt Learning at Test Timeaccepted
  628. Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervisionaccepted
  629. Convexity-Aware Noise Calibration: A Self-Supervised Framework for Noise-Level-Unknown Image Denoisingaccepted
  630. Convolutional Neural Networks Driven by Content Similarityaccepted
  631. CoopDiff: A Diffusion-Guided Approach for Cooperation under Corruptionsaccepted
  632. CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generationaccepted
  633. Coordinate Denoising for Non-Equilibrium Molecular Representation Learningaccepted
  634. Copy-Transform-Paste: Zero-Shot Object-Object Alignment Guided by Vision-Language and Geometric Constraintsaccepted
  635. Correspondence-Attention Alignment for Multi-View Diffusion Modelsaccepted
  636. CountGD++: Generalized Prompting for Open-World Countingaccepted
  637. Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoningaccepted
  638. Coupled Diffusion Sampling for Training-Free Multi-View Image Editingaccepted
  639. Coupling Liquid Time-Constant Encoders with Modern Hopfield Memoryaccepted
  640. Cov2Pose: Leveraging Spatial Covariance for Direct Manifold-aware 6-DoF Object Pose Estimationaccepted
  641. Coverage Optimization for Camera View Selectionaccepted
  642. CrackSSM: Reviving SSMs for Crack Segmentation via Dynamic Scanningaccepted
  643. CraftMesh: High-Fidelity Generative Mesh Manipulation via Poisson Seamless Fusionaccepted
  644. Critical Patch-Aware Sparse Prompting with Decoupled Training for Continual Learning on the Edgeaccepted
  645. Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigationaccepted
  646. Cross-Architecture Adaptation: Cloud-Edge Continual Test-Time Adaptation with Dynamic Sampling and Heterogeneous Distillationaccepted
  647. Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editingaccepted
  648. Cross-Domain Demo-to-Code via Neurosymbolic Counterfactual Reasoningaccepted
  649. Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptationaccepted
  650. Cross-Hand Latent Representation for Vision-Language-Action Modelsaccepted
  651. Cross-Instance Gaussian Splatting Registration via Geometry-Aware Feature-Guided Alignmentaccepted
  652. Cross-Modal Attention Calibration for LVLM Hallucination Mitigationaccepted
  653. Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Videoaccepted
  654. Cross-Modal Guided Visual Synthesis for Data-Efficient Multimodal Depression Recognitionaccepted
  655. Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmarkaccepted
  656. Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inferenceaccepted
  657. Cross-Subject EEG-to-Video Reconstruction and Beyondaccepted
  658. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classificationaccepted
  659. Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Imagesaccepted
  660. Cross-domain Dual-stream Feature Disentanglement for Brain Disorder Prediction with Sparsely Labeled PETaccepted
  661. Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmarkaccepted
  662. Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement Learningaccepted
  663. Cross-modal Representation Learning for Diffusion-generated Image Detectionaccepted
  664. CrossEarth-Gate: Fisher-Guided Adaptive Tuning Engine for Efficient Adaptation of Cross-Domain Remote Sensing Semantic Segmentationaccepted
  665. CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methodsaccepted
  666. CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction Reconstructionaccepted
  667. CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detectionaccepted
  668. CrowdGaussian: Reconstructing High-Fidelity 3D Gaussians for Human Crowd from a Single Imageaccepted
  669. CryoHype: Reconstructing a thousand cryo-EM structures with transformer-based hypernetworksaccepted
  670. CryoKRAQEN: Kernel-Regularized Annealing for Quantized Embedding Networks in Cryo-EM Heterogeneous Reconstructionaccepted
  671. CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective Videoaccepted
  672. Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokensaccepted
  673. Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generationaccepted
  674. Curvature-Aware Captioning: Leveraging Geodesic Attention for 3D Scene Understandingaccepted
  675. Curvature-Aware Zeroth-Order Optimization for Memory-Efficient Test-Time Adaptationaccepted
  676. CustomTex: High-fidelity Indoor Scene Texturing via Multi-Reference Customizationaccepted
  677. Customized Fusion: A Closed-Loop Dynamic Network for Adaptive Multi-Task-Aware Infrared-Visible Image Fusionaccepted
  678. Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Eventsaccepted
  679. Cycle-Consistent Tuning for Layered Image Decompositionaccepted
  680. CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird's-Eye-View Semantic Segmentationaccepted
  681. CycleManip: Enabling Cycle-based Manipulation via Effective History Perception and Understandingaccepted
  682. D$^2$-FOSA: Dual-Diffusion Guided EEG-to-Image Reconstruction with Frequency-Oriented Semantic Alignmentaccepted
  683. D-Convexity: A Unified Differentiable Convex Shape Prior via Quasi-Concavity for Data-driven Image Segmentationaccepted
  684. D-Prism: Differentiable Primitives for Structured Dynamic Modelingaccepted
  685. D2Cache: Second-Order Delta Caching for Higher Video Diffusion Accelerationaccepted
  686. D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image Dewarpingaccepted
  687. D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation Networkaccepted
  688. D2T2 - Multimodal Automated Planning for Brachytherapyaccepted
  689. D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigationaccepted
  690. DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detectionaccepted
  691. DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignmentaccepted
  692. DABO: Difficulty-Aware Bayesian Optimization with Diffusion-Learned Priorsaccepted
  693. DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimationaccepted
  694. DARC: Dual Adjustment Reasoning with Counterfactuals for Trustworthy Chest X-ray Classificationaccepted
  695. DASH: A Meta-Attack Framework for Synthesizing Effective and Stealthy Adversarial Examplesaccepted
  696. DBMSolver: A Training-free Diffusion Bridge Sampler for High-Quality Image-to-Image Translationaccepted
  697. DC-Merge: Improving Model Merging with Directional Consistencyaccepted
  698. DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generationaccepted
  699. DDSF: Robust Few-Shot Learning via Disentangled Subspaces with Determinantal Point Processaccepted
  700. DDT: Decoupled Diffusion Transformeraccepted
  701. DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformersaccepted
  702. DENALI: A Dataset Enabling Non-Line-of-Sight Spatial Reasoning with Low-Cost LiDARsaccepted
  703. DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learningaccepted
  704. DEVA: Fine-tuning Multimodal Large Language Models for Visual Perception Tasksaccepted
  705. DFD-HR: Generalizable Deepfake Detection via Hierarchical Routing Learningaccepted
  706. DF^2-VB: Dual-level Fuzzy Fusion with View-specific Boosting for Multi-view Multi-label Classificationaccepted
  707. DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Imagesaccepted
  708. DGS: Dual Gradient and Semantic-Shift Guided Low-Rank Adaptation for Class Incremental Learningaccepted
  709. DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spacesaccepted
  710. DIMOS: Disentangling Instance-level Moving Object Segmentationaccepted
  711. DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrievalaccepted
  712. DK-DDIL: Adaptive Knowledge Retention for Dynamic Domain-Incremental Learning in Medical Imagingaccepted
  713. DLVP-CLIP: Enhancing Fine-Grained Zero-Shot Anomaly Detection via Dynamic Local Visual Promptingaccepted
  714. DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Drivingaccepted
  715. DMAligner: Enhancing Image Alignment via Diffusion Model Based View Synthesisaccepted
  716. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Modelsaccepted
  717. DNF-SR: Dual-Input and Negative-Aware Feature Fine-Tuning for Real-World Image Super-Resolutionaccepted
  718. DP-FedAdamW: An Efficient Optimizer for Differentially Private Federated Large Modelsaccepted
  719. DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generationaccepted
  720. DPGF-Net: Dual-Prior Guided Fusion Network for Joint Assessment of Perceptual Quality and Semantic Consistency in AI-Generated Imagesaccepted
  721. DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalitiesaccepted
  722. DRAMA: Next-Gen Dynamic Orchestration for Resilient Multi-Agent Ecosystems in Fluxaccepted
  723. DREAM: Document Recognition with Explicit Adaptive Memoryaccepted
  724. DRM: Diffusion-based Reward Model With Step-wise Guidanceaccepted
  725. DROID-SLAM in the Wildaccepted
  726. DRS-GUI: Dynamic Region Search for Training-Free GUI Groundingaccepted
  727. DRiffusion: Draft-and-Refine Process Parallelizes Diffusion Models with Easeaccepted
  728. DSCA: Dynamic Subspace Concept Alignment for Lifelong VLM Editingaccepted
  729. DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDARaccepted
  730. DSFlash: Comprehensive Panoptic Scene Graph Generation in Realtimeaccepted
  731. DSO: Direct Steering Optimization for Bias Mitigationaccepted
  732. DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolutionaccepted
  733. DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inferenceaccepted
  734. DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolutionaccepted
  735. DVAR: Dynamic Visual Autoregressive Modeling for Image Super-Resolutionaccepted
  736. DVGT: Driving Visual Geometry Transformeraccepted
  737. D^3FER: Dual Channel and Dual Branch Network for Robust Facial Expression Recognition under Dual Challengesaccepted
  738. Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style Bridgingaccepted
  739. Dark3R: Learning Structure from Motion in the Darkaccepted
  740. DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action Recognitionaccepted
  741. DarkShake-DVS: Event-based Human Action Recognition under Low-light and Shaking Camera Conditionsaccepted
  742. Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasetsaccepted
  743. Data-Centric Meta-Learning for Robust Few-Shot Generalizationaccepted
  744. Dataset Distillation by Influence Matchingaccepted
  745. DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Rolesaccepted
  746. DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generationaccepted
  747. DeDelayed: Deleting Remote Inference Delay via On-Device Correctionaccepted
  748. DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentationaccepted
  749. DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representationsaccepted
  750. Debiased Sample Selection for Learning with Noisy Labelsaccepted
  751. Deciphering Genotype-Phenotype Mechanisms from High-Content Profiling via Knowledge-Guided Multi-modal Graph Learningaccepted
  752. Decision Boundary-aware Generation for Long-tailed Learningaccepted
  753. DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigationaccepted
  754. Decoding 3D Perception via BrainSSD: Synergistic Fusion of EEG Representations from Static and Dynamic Visual Streamsaccepted
  755. Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detectionaccepted
  756. Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compressionaccepted
  757. Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar Diagnosisaccepted
  758. Decouple Your Discovery and Memory in Continual Generalized Category Discoveryaccepted
  759. Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoningaccepted
  760. Decoupled Generative Modeling for Human-Object Interaction Synthesisaccepted
  761. Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translationaccepted
  762. Decoupled and Reusable Adaptation for Efficient Cross-Modal Transferaccepted
  763. Decoupling Bias, Aligning Distributions: Synergistic Fairness Optimization for Deepfake Detectionaccepted
  764. Decoupling Defense Strategies for Robust Image Watermarkingaccepted
  765. Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptationaccepted
  766. Decoupling Vision and Language: Codebook Anchored Visual Adaptationaccepted
  767. Deep Feature Deformation Weightsaccepted
  768. DeepAlign: Mitigating Modality Conflict through Modality-Specific Alignmentaccepted
  769. DeepProtect: Proactive Face-Swapping Defense using Identity Blending and Attribute Distortionaccepted
  770. DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Modelsaccepted
  771. Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Modelsaccepted
  772. DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake Detectionaccepted
  773. Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detectionaccepted
  774. Defending Unauthorized Model Merging via Dual-Stage Weight Protectionaccepted
  775. Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillationaccepted
  776. Deformation-based In-Context Learning for Point Cloud Understandingaccepted
  777. Degradation-Consistent Test-Time Adaptation for All-in-One Image Restorationaccepted
  778. Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenariosaccepted
  779. Dehallu3D: Hallucination-Mitigated 3D Generation from a Single Image via Cyclic View Consistency Refinementaccepted
  780. Dejavu: Towards Experience Feedback Learning for Embodied Intelligenceaccepted
  781. Delta Rectified Flow Sampling for Text-to-Image Editingaccepted
  782. DeltaQuant: 4-bit Video Diffusion Models with Spatiotemporal Delta Smoothingaccepted
  783. Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Modelsaccepted
  784. Demo2Tutorial: From Human Experience to Multimodal Software Tutorialsaccepted
  785. DemoFunGrasp: Universal Dexterous Functional Grasping via Demonstration-Editing Reinforcement Learningaccepted
  786. Den-TP: A Density-Balanced Data Curation and Evaluation Framework for Trajectory Predictionaccepted
  787. Denoise and Align: Towards Source-Free UDA for Robust Panoramic Semantic Segmentationaccepted
  788. Denoising as Path Planning: Training-Free Acceleration of Diffusion Models with DPCacheaccepted
  789. Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generationaccepted
  790. Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensorsaccepted
  791. Depth Any Endoscopy: Towards Self-Supervised Generalizable Depth Estimation in Monocular Endoscopyaccepted
  792. Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimationaccepted
  793. Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimationaccepted
  794. Depth Peeling for High-Fidelity Gaussian-Enhanced Surfel Renderingaccepted
  795. DepthFocus: Controllable Depth Estimation for See-Through Scenesaccepted
  796. Describe Anything Anywhere At Any Momentaccepted
  797. Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Modelsaccepted
  798. Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkageaccepted
  799. Designing to Forget: Deep Semi-parametric Models for Unlearningaccepted
  800. DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Videoaccepted
  801. Detect Any AI-Counterfeited Text Imageaccepted
  802. Detect Anything via Next Point Predictionaccepted
  803. DetectSCI: Toward Object-Guided ROI Reconstruction for High-Resolution Video Snapshot Compressive Imagingaccepted
  804. Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplificationaccepted
  805. Detecting Compressed AI-Generated Images via Phase Spectrum Robustnessaccepted
  806. Detecting Unknown Objects via Energy-based Separation for Open World Object Detectionaccepted
  807. DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoningaccepted
  808. Dexterous World Modelsaccepted
  809. DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Modelsaccepted
  810. DiGraphHal-Bench: Evaluating Multimodal Large Language Models on Complex Directed Graphsaccepted
  811. DiP: Taming Diffusion Models in Pixel Spaceaccepted
  812. DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum Knowledgeaccepted
  813. DiT-IC: Aligned Diffusion Transformer for Efficient Image Compressionaccepted
  814. DiT360: High-Fidelity Panoramic Image Generation via Hybrid Trainingaccepted
  815. Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbolsaccepted
  816. Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projectionaccepted
  817. Diagram2Structure: Unlocking LLMs' Diagram Comprehension through DiagramDiff, a Framework for Structuring Offline Diagramsaccepted
  818. DialogueVPR: Towards Conversational Visual Place Recognitionaccepted
  819. Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMsaccepted
  820. Diff-SemiER: Transparency-Aware Adaptive Fusion Diffusion Model with Generative Prior for Semi-Transparent Eyeglasses Removalaccepted
  821. Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene Generationaccepted
  822. DiffBMP: Differentiable Rendering with Bitmap Primitivesaccepted
  823. DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformersaccepted
  824. DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generationaccepted
  825. DiffSoup: Direct Differentiable Rasterization of Triangle Soup for Extreme Radiance Field Simplificationaccepted
  826. Differences That Matter: Auditing Models for Capability Gap Discovery and Rectificationaccepted
  827. Differentiable Adaptive 4D Structured Illumination for Joint Capture of Shape and Reflectanceaccepted
  828. Differentiable Laplacian Matrix Guided Superpixel Segmentationaccepted
  829. Differentiable Stroke Planning with Dual Parameterization for Efficient and High-Fidelity Painting Creationaccepted
  830. Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image Compressionaccepted
  831. Differentially Private 2D Human Pose Estimationaccepted
  832. DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic Manipulationaccepted
  833. Diffusion Forcing Planner: History-Annealed Planning with Time-Dependent Guidance for Autonomous Drivingaccepted
  834. Diffusion Guided Chain-of-Vision for Large Autoregressive Vision Modelsaccepted
  835. Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)accepted
  836. Diffusion Mental Averagesaccepted
  837. Diffusion Probe: Generated Image Result Prediction Using CNN Probesaccepted
  838. Diffusion Sampling Path Tells More: An Efficient Plug-and-Play Strategy for Sample Filteringaccepted
  839. Diffusion with a Linguistic Compass: Steering the Generation of Clinically Plausible Future sMRI Representations for Early MCI Conversion Predictionaccepted
  840. Diffusion-Based Makeup Transfer with Facial Region-Aware Makeup Featuresaccepted
  841. Diffusion-Based Native Adversarial Synthesis for Enhanced Medical Segmentation Generalizationaccepted
  842. Diffusion-Based sRGB Real Noise Generation via Prompt-Driven Noise Representation Learningaccepted
  843. DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localizationaccepted
  844. DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhanceraccepted
  845. Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentationaccepted
  846. DirectFisheye-GS: Enabling Native Fisheye Input in Gaussian Splatting with Cross-View Joint Optimizationaccepted
  847. Direction-aware 3D Large Multimodal Modelsaccepted
  848. DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Cachingaccepted
  849. Disco-GS: Gaussian Splatting in Dynamic Color Lightingaccepted
  850. Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentationaccepted
  851. Discovering Adaptive Task Dependencies for Efficient Multi-Task Representation Compressionaccepted
  852. Discriminative Perception via Anchored Description for Reasoning Segmentationaccepted
  853. Disentangle-then-Align: Non-Iterative Hybrid Multimodal Image Registration via Cross-Scale Feature Disentanglementaccepted
  854. Disentangled Textual Priors for Diffusion-based Image Super-Resolutionaccepted
  855. Disentanglement-wise Image Dehazing through Cross-Domain Manifold Consensusaccepted
  856. Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generationaccepted
  857. Distilling Balanced Knowledge from a Biased Teacheraccepted
  858. Distilling Quasi-Conformal Mapping: A Generalizable and Efficient Solution for Wide-Angle Correctionaccepted
  859. Distilling Unsigned Distance Function for Surface Reconstruction from 3D Gaussian Splattingaccepted
  860. Distributed Image Compression with Multimodal Side Information at Extremely Low Bitratesaccepted
  861. Distribution-Aligned Multimodal Fusion for Robust Object Detectionaccepted
  862. Diverse Video Generation with Determinantal Point Process-Guided Policy Optimizationaccepted
  863. DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformersaccepted
  864. DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPOaccepted
  865. Diversity over Uniformity: Rethinking Representation in Generated Image Detectionaccepted
  866. Divide, Conquer, and Aggregate: Asymmetric Experts for Class-Imbalanced Semi-Supervised Medical Image Segmentationaccepted
  867. Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understandingaccepted
  868. Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?accepted
  869. Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusionsaccepted
  870. Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacksaccepted
  871. Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBenchaccepted
  872. Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Controlaccepted
  873. Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answeringaccepted
  874. DocPrune: Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruningaccepted
  875. DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understandingaccepted
  876. Does YOLO Really Need to See Every Training Image in Every Epoch?accepted
  877. Domain Sensitive Federated Learning with Fisher-Informed Pruningaccepted
  878. Domain-Skewed Federated Learning with Feature Decoupling and Calibrationaccepted
  879. Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programsaccepted
  880. Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Modelsaccepted
  881. Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Designaccepted
  882. Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Drivingaccepted
  883. Draft and Refine with Visual Expertsaccepted
  884. Drainage: A Unifying Framework for Addressing Class Uncertaintyaccepted
  885. DreamOmni2: Multimodal Instruction-based Generation and Editingaccepted
  886. DreamSAC: Learning Hamiltonian World Models via Symmetry Explorationaccepted
  887. DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion Transformeraccepted
  888. DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prioraccepted
  889. DreamStereo: Towards Real-Time Stereo Inpainting for HD Videosaccepted
  890. DreamStyle: A Unified Framework for Video Stylizationaccepted
  891. DreamingComics: A Story Visualization Pipeline via Subject and Layout Customized Generation using Video Modelsaccepted
  892. Drift-Resilient Temporal Priors for Visual Trackingaccepted
  893. Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Drivingaccepted
  894. DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous Drivingaccepted
  895. DriveLaW: Unifying Planning and Video Generation in a Latent Driving Worldaccepted
  896. DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Drivingaccepted
  897. DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planningaccepted
  898. DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generationaccepted
  899. DriveVLN: Towards Mapless Vision-and-Language Navigation in Autonomous Drivingaccepted
  900. DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidanceaccepted
  901. Driving on Registersaccepted
  902. Dropping Anchor and Spherical Harmonics for Sparse-view Gaussian Splattingaccepted
  903. Dual Ascent Diffusion for Inverse Problemsaccepted
  904. Dual Band Thermal Videography: Separating Time-Varying Reflection and Emission Near Ambient Conditionsaccepted
  905. Dual Graph Regularized Deep Unfolding Network for Guided Depth Map Super-resolutionaccepted
  906. Dual-Agent Reinforcement Learning for Adaptive and Cost-Aware Visual-Inertial Odometryaccepted
  907. Dual-Estimator: Decoupling Global and Local Semantic Shift for Drift Compensation in Class-Incremental Learningaccepted
  908. Dual-Granularity Memory for Efficient Video Generationaccepted
  909. Dual-Level Confidence based Implicit Self-Refinement for Medical Visual Question Answeringaccepted
  910. Dual-Level Hypergraph Generation for Addressing Feature Scarcity in Whole-Slide Image Classificationaccepted
  911. Dual-Prototype-Guided Multi-task Learning for Unsupervised Anomaly Detection and Classificationaccepted
  912. Dual-branch Distilled Transformer for Efficient Asymmetric UAV Trackingaccepted
  913. Dual-level Adaptation for Multi-Object Tracking: Building Test-Time Calibration from Experience and Intuitionaccepted
  914. Dual-level Adapter Boosting Prompt-free Curvilinear Structure Segmentationaccepted
  915. DualMirage: Hunting Stealthy Multimodal LLM Agents via CAPTCHAs with Contour and Adversarial Illusionsaccepted
  916. DualPrim: Compact 3D Reconstruction with Positive and Negative Primitivesaccepted
  917. DualReg: Dual-Space Filtering and Reinforcement for Rigid Registrationaccepted
  918. DualSplat: Robust 3D Gaussian Splatting via Pseudo-Mask Bootstrapping from Reconstruction Failuresaccepted
  919. Duala: Dual-Level Alignment of Subjects and Stimuli for Cross-Subject fMRI Decodingaccepted
  920. DuetMerging: Synergizing Dynamic and Static Strategies for Mitigating Task Interference in Model Mergingaccepted
  921. DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidanceaccepted
  922. DuoGen: Towards Autonomous Interleaved Multimodal Generationaccepted
  923. DuoMo: Dual Motion Diffusion for World-Space Human Reconstructionaccepted
  924. DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detectionaccepted
  925. DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generationaccepted
  926. DynBridge: Bridging Imagination and Control through Interaction Dynamics for Robot Manipulationaccepted
  927. DynFusion: Rethinking Condition Fusion for Adaptive Multi-Conditional Text-to-Image Generationaccepted
  928. Dynamic Black-hole Emission Tomography with Physics-informed Neural Fieldsaccepted
  929. Dynamic Exposure Burst Image Restorationaccepted
  930. Dynamic Important Example Mining for Reinforcement Finetuningaccepted
  931. Dynamic Label Noise Suppression with Optimal Teacher Pool for Facial Expression Recognitionaccepted
  932. Dynamic Logits Adjustment and Exploration for Test-Time Adaptation in Vision Language Modelsaccepted
  933. Dynamic Magic: Unleashing Restricted Knowledge for Lifelong Person Re-Identificationaccepted
  934. Dynamic Momentum Recalibration in Online Gradient Learningaccepted
  935. Dynamic Stream Network for Combinatorial Explosion Problem in Deformable Medical Image Registrationaccepted
  936. Dynamic Token Reweighting for Robust Vision-Language Modelsaccepted
  937. Dynamic Visual SLAM using a General 3D Prioraccepted
  938. Dynamic-Static Decomposition for Novel View Synthesis of Dynamic Scenes with Spiking Neuronsaccepted
  939. Dynamic-eDiTor: Training-Free Text-Driven 4D Scene Editing with Multimodal Diffusion Transformeraccepted
  940. DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAsaccepted
  941. DynamicTree: Interactive Real Tree Animation via Sparse Voxel Spectrumaccepted
  942. DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Drivingaccepted
  943. Dynamics-Aware Preference Optimization for Vision-Language Modelsaccepted
  944. Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videosaccepted
  945. DynamicsBoost: Dynamic Plausible Video Generation via Annotation-Free Continuation Preference Optimizationaccepted
  946. E$^2$-SCI: Elastic Edge-Cloud Speculative Decoding via Credit Inertiaaccepted
  947. E-3DPSM: A State Machine for Event-based Egocentric 3D Human Pose Estimationaccepted
  948. E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-trainingaccepted
  949. E-comIQ-ZH: A Human-Aligned Dataset and Benchmark for Fine-Grained Evaluation of E-commerce Posters with Chain-of-Thoughtaccepted
  950. E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstructionaccepted
  951. E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Drivingaccepted
  952. EDGS: Eliminating Densification for Efficient Convergence of 3DGSaccepted
  953. EE-RL: Vision Language Guided Reinforcement Learning with Explorer and Expert model for End-to-End Autonomous Drivingaccepted
  954. EEGiT: Teaching Vision Transformers to Understand the EEG signalaccepted
  955. EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual Groundingaccepted
  956. EI-Part:Explode for Completion and Implode for Refinementaccepted
  957. ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and Test-time Generative Adaptationaccepted
  958. ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video Understandingaccepted
  959. ELVIS: Enhance Low-Light for Video Instance Segmentation in the Darkaccepted
  960. ELiC: Efficient LiDAR Geometry Compression via Cross-Bit-depth Feature Propagation and Bag-of-Encodersaccepted
  961. EMAD: Evidence-Centric Grounded Multimodal Diagnosis for Alzheimer's Diseaseaccepted
  962. EMGauss: Continuous Slice-to-3D Reconstruction via Dynamic Gaussian Modeling in Volume Electron Microscopyaccepted
  963. EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categoriesaccepted
  964. EMMA: Extracting Multiple physical parameters from Multimodal Dataaccepted
  965. EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Modelsaccepted
  966. EMR-Diff: Edge-aware Multimodal Residual Diffusion Model for Hyperspectral Image Super-resolutionaccepted
  967. ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understandingaccepted
  968. ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specializationaccepted
  969. EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detectionaccepted
  970. ESAM++: Efficient Online 3D Perception on the Edgeaccepted
  971. EV-CGNet: Co-visible Focused 3D-guided 2D Event Keypoint Detection Networkaccepted
  972. EVA: Efficient Reinforcement Learning for End-to-End Video Agentaccepted
  973. EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generationaccepted
  974. EVLF: Early Vision-Language Fusion for Generative Dataset Distillationaccepted
  975. EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformeraccepted
  976. EXOTIC: External Vision-driven Incomplete Multi-view Classificationaccepted
  977. EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrievalaccepted
  978. EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligenceaccepted
  979. EarlyTom: Early Token Compression Completes Fast Video Understandingaccepted
  980. Easy2Hard: From Partially to Fully Unmatched Modalities as Negative Samples in Contrastive Learningaccepted
  981. Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel Flowaccepted
  982. EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decompositioaccepted
  983. EasyV2V: A High-quality Instruction-based Video Editing Frameworkaccepted
  984. EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generationaccepted
  985. EchoPOSE: 6D Pose Estimation of Sparse Echocardiograms for Left-Ventricular 3D Shape Reconstructionaccepted
  986. EchoVDiff: Cardiac-Cycle Echocardiography Video Generation from Arbitrary Single Frameaccepted
  987. Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Modelsaccepted
  988. Echoes of Ownership: Adversarial-Guided Dual Injection for Copyright Protection in MLLMsaccepted
  989. EcoAlign: An Economically Rational Framework for Efficient LVLM Alignmentaccepted
  990. EcoSplat: Efficiency-controllable Feed-forward 3D Gaussian Splatting from Multi-view Imagesaccepted
  991. Edge-Focused Super-Resolution for Omnidirectional Images with Spherical Geometric Augmentationaccepted
  992. Edge-RecViT: Efficient Vision Transformer via Semantic-Refined Dynamic Recursionaccepted
  993. Edges Compete for Trust: Group Relative Edge Optimization for Building Reconstruction from Point Cloudsaccepted
  994. Edit-As-Act: Goal-Regressive Planning for Open-Vocabulary 3D Indoor Scene Editingaccepted
  995. Edit-aware RAW reconstructionaccepted
  996. Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceiversaccepted
  997. EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editingaccepted
  998. EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editingaccepted
  999. Editprint: General Digital Image Forensics via Editing Fingerprint with Self-Augmentation Trainingaccepted
  1000. EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal Modelsaccepted

Looking for submission deadlines instead? See the conference deadline calendar.

CVPR 2026 Accepted Papers · Full List of 4,068 Papers