← All conferences

CVPR 2026 Accepted Papers

The full list of 4,068 papers accepted at CVPR 2026 (IEEE/CVF Conference on Computer Vision and Pattern Recognition). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

accepted: 4,068
  1. SAM2Text: Towards Prompt-Free and Multi-Resolution Video Scene Text Segmentationaccepted
  2. SAME: Sparse and Anchored Model Editing for Heterogeneous Incremental Learning under Limited Dataaccepted
  3. SAMIX: Reinforcing SAM2 with Semantic Adapter and Reference Selecting Policy for Mix-Supervised Segmentationaccepted
  4. SAMTok: Representing Any Mask with Two Wordsaccepted
  5. SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment Anythingaccepted
  6. SANER: Switchable Adapter with Non-parametric Enhanced Routing for Person De-Reidentificationaccepted
  7. SAQN: Semantic-based Adaptive Query Network for 3D Referring Expression Segmentationaccepted
  8. SAR2Net: Learning Spatially Anchored Representations for Retrieval-Guided Cross-Stain Alignmentaccepted
  9. SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video Groundingaccepted
  10. SARMAE: Masked Autoencoder for SAR Representation Learningaccepted
  11. SASNet: Spatially-Adaptive Sinusoidal Networks for INRsaccepted
  12. SAT-RRG: LLM-Guided Self-Adaptive Training for Radiology Report Generation with Token-Level Push-Pull Optimizationaccepted
  13. SATTC: Structure-Aware Label-Free Test-Time Calibration for Cross-Subject EEG-to-Image Retrievalaccepted
  14. SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusionaccepted
  15. SAVE: Speech-Aware Video Representation Learning for Video-Text Retrievalaccepted
  16. SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observationaccepted
  17. SCE-Depth: A Spherical Compound Eye Framework for Wide FOV Depth Estimationaccepted
  18. SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate Embeddingsaccepted
  19. SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal Modelsaccepted
  20. SCoRe: Salience-Coverage Reduction for Vision Token Pruning in Vision-Language Modelsaccepted
  21. SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentationaccepted
  22. SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detectionaccepted
  23. SDGS: Spatial Difference Guided Gaussian Splatting for Simultaneous Localization and 3D Reconstructionaccepted
  24. SDTrack: A Baseline for Event-based Tracking via Spiking Neural Networksaccepted
  25. SDUIE: Semi-Supervised Diffusion for Underwater Image Enhancement with Quant-Text Dual Controlaccepted
  26. SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose Estimationaccepted
  27. SEA-Flow3D: Simplified, Efficient, and Accurate Scene Flow via Spatial Vector Sampling and Multi-scale Refinementaccepted
  28. SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asiaaccepted
  29. SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answeringaccepted
  30. SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decodingaccepted
  31. SEATrack: Simple, Efficient, and Adaptive Multimodal Trackeraccepted
  32. SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learningaccepted
  33. SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learningaccepted
  34. SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detectionaccepted
  35. SG-LoRA: Semantic-guided LoRA Parameters Generationaccepted
  36. SGAD-SLAM: Splatting Gaussians at Adjusted Depth for Better Radiance Fields in RGBD SLAMaccepted
  37. SGDE: Self-supervised Geometry Degradation Estimation Framework for Coded Aperture Compressive Spectral Imagingaccepted
  38. SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Drivingaccepted
  39. SGI: Structured 2D Gaussians for Efficient and Compact Large Image Representationaccepted
  40. SGS-Intrinsic: Semantic-Invariant Gaussian Splatting for Sparse-View Indoor Inverse Renderingaccepted
  41. SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signalsaccepted
  42. SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentationaccepted
  43. SHARP: Short-Window Streaming for Accurate and Robust Prediction in Motion Forecastingaccepted
  44. SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wildaccepted
  45. SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Trainingaccepted
  46. SIF: Semantically In-Distribution Fingerprints for Large Vision-Language Modelsaccepted
  47. SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Imagesaccepted
  48. SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokensaccepted
  49. SIMPACT: Simulation-Enabled Action Planning using Vision-Language Modelsaccepted
  50. SIMPLEPOSTER: A SIMPLE BASELINE FOR PRODUCT POSTER GENERATIONaccepted
  51. SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarkingaccepted
  52. SIR: Structured Image Representations for Explainable Robot Learningaccepted
  53. SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuationaccepted
  54. SLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic Scenesaccepted
  55. SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generationaccepted
  56. SMAP: Semantic Route Planning with Map-Grounded Multimodal Alignmentaccepted
  57. SMRABooth: Subject and Motion Representation Alignment for Customized Video Generationaccepted
  58. SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognitionaccepted
  59. SMVRT: Implicit Human 3D Modeling Using Sparse Multi-View Volumetric Reconstruction with Transformer Fusionaccepted
  60. SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth Estimationaccepted
  61. SO-Bench: A Structural Output Evaluation of Multimodal LLMaccepted
  62. SODA: Sensitivity-Oriented Dynamic Acceleration for Diffusion Transformeraccepted
  63. SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation Modelsaccepted
  64. SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt Contextsaccepted
  65. SPAN: Spatial-Projection Alignment for Monocular 3D Object Detectionaccepted
  66. SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentationaccepted
  67. SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledgeaccepted
  68. SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMsaccepted
  69. SPDMark: Selective Parameter Displacement for Robust Video Watermarkingaccepted
  70. SPE-MVS: Spatial Position Encoding Enhanced Multi-View Stereo with Monocular Depth Priorsaccepted
  71. SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understandingaccepted
  72. SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentationaccepted
  73. SPOT: Spatiotemporal Prompt Optimization for Motion-Stabilized MLLM-Guided Video Segmentationaccepted
  74. SPREAD: Spatial-Physical REasoning via geometry Aware Diffusionaccepted
  75. SR3R: Rethinking Super-Resolution 3D Reconstruction With Feed-Forward Gaussian Splattingaccepted
  76. SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion Trainingaccepted
  77. SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Namesaccepted
  78. SRGCD: Stability-Driven Region Growth Framework for 3D Change Detectionaccepted
  79. SRPO: Self-Referential Policy Optimization for Vision-Language-Action Modelsaccepted
  80. SSM-Aware Token-Efficient VMamba via Adaptive Patch Pruning and Merging for Person Re-Identificationaccepted
  81. ST4R-Splat: Spatio-Temporal Referring Segmentation in 4D Gaussian Splattingaccepted
  82. STAC: Plug-and-Play Spatio-Temporal Aware Cache Compression for Streaming 3D Reconstructionaccepted
  83. STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrativeaccepted
  84. STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMsaccepted
  85. STAR: Test-Time Adaptation Can Enhance Universal Prompt Learning for Vision-Language Modelsaccepted
  86. STARFlow-V: End-to-End Video Generative Modeling with Autoregressive Normalizing Flowsaccepted
  87. STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars Reconstructionaccepted
  88. STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolutionaccepted
  89. STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecastingaccepted
  90. STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregationaccepted
  91. STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detectionaccepted
  92. STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrievalaccepted
  93. SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detectionaccepted
  94. SURF: Signature-Retained Fast Video Generationaccepted
  95. SV-GS: Sparse View 4D Reconstruction with Skeleton-Driven Gaussian Splattingaccepted
  96. SVAgent: Storyline-guided Long Video Understanding via Cross-Modal Multi-Agent Collaborationaccepted
  97. SVBench: Evaluation of Video Generation Models on Social Reasoningaccepted
  98. SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Modelsaccepted
  99. SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attributionaccepted
  100. SaPaVe: Towards Active Perception and Manipulation in Vision-Language Action Models for Roboticsaccepted
  101. SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse Worldaccepted
  102. SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimizationaccepted
  103. SafeLogo: Turning Your Logos into Jailbreak Shields via Micro-Regional Adversarial Trainingaccepted
  104. SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow Transformersaccepted
  105. Saliency-Driven Token Merging for Vision Transformersaccepted
  106. Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learningaccepted
  107. Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Rewardaccepted
  108. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucinationaccepted
  109. Same Content, Different Answers: Cross-Modal Inconsistency in MLLMsaccepted
  110. Same or Not? Enhancing Visual Perception in Vision-Language Modelsaccepted
  111. Sampling-Aware Quantization for Diffusion Modelsaccepted
  112. Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Editsaccepted
  113. Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstructionaccepted
  114. Scalable Feature Matching via State Space Modeling and Sparse Correlationaccepted
  115. Scalable Multi-View Subspace Clustering with Tensorized Anchor Guidanceaccepted
  116. Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Modelsaccepted
  117. Scalable Trajectory Generation for Whole-Body Mobile Manipulationaccepted
  118. Scale Space Diffusionaccepted
  119. Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMsaccepted
  120. Scaling Dense Event-Stream Pretraining from Visual Foundation Modelsaccepted
  121. Scaling Instruction-Based Video Editing with a High-Quality Synthetic Datasetaccepted
  122. Scaling Multi-Identity Consistency for Image Customization via Multi-to-Multi Matching Paradigmaccepted
  123. Scaling Parallel Sequence Models to Vision Foundation Modelsaccepted
  124. Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformersaccepted
  125. Scaling Spatial Intelligence with Multimodal Foundation Modelsaccepted
  126. Scaling Test-Time Robustness of Vision-Language Models via Self-Critical Inference Frameworkaccepted
  127. Scaling Up AI-Generated Image Detection with Generator-Aware Prototypesaccepted
  128. Scaling View Synthesis Transformersaccepted
  129. Scaling Zero-Shot Reference-to-Video Generationaccepted
  130. Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanismaccepted
  131. Scaling-Aware Data Selection for End-to-End Autonomous Driving Systemsaccepted
  132. Scaling4D: Pushing the Frontier of Video Novel View Synthesis through Large-Scale Monocular Videosaccepted
  133. Scan Clusters, Not Pixels: A Cluster-Centric Paradigm for Efficient Ultra-high-definition Image Restorationaccepted
  134. SceMoS: Scene-Aware 3D Human Motion Synthesis by Planning with Geometry-Grounded Tokensaccepted
  135. ScenDi: 3D-to-2D Scene Diffusion Cascades for Urban Generationaccepted
  136. Scene Grounding in the Wildaccepted
  137. Scene Reconstruction as Mapping Priors for 3D Detectionaccepted
  138. Scene-Centric Unsupervised Video Panoptic Segmentationaccepted
  139. Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Modelsaccepted
  140. SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Modelaccepted
  141. SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotationsaccepted
  142. SceneTok: A Compressed, Diffusable Token Space for 3D Scenesaccepted
  143. Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understandingaccepted
  144. SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent Systemaccepted
  145. Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modelingaccepted
  146. Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoringaccepted
  147. Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformersaccepted
  148. SeD-UD: An Influence-Driven and Hierarchically-Decoupled Information Bottleneck for Multimodal Intent Recognitionaccepted
  149. SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion Modelsaccepted
  150. SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Drivingaccepted
  151. See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflectionaccepted
  152. See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsaccepted
  153. See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoningaccepted
  154. See Through the Noise: Improving Domain Generalization in Gaze Estimationaccepted
  155. See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understandingaccepted
  156. See What We Cannot See: A Geo-guided Reasoning Benchmark for Object Counting under Adverse Earth Observation Conditionsaccepted
  157. See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesisaccepted
  158. See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Togglesaccepted
  159. SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Groupingaccepted
  160. SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generationaccepted
  161. SeeU: Seeing the Unseen World via 4D Dynamics-aware Generationaccepted
  162. Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videosaccepted
  163. Seeing Beyond: Extrapolative Domain Adaptive Panoramic Segmentationaccepted
  164. Seeing Both Sides: Towards Bidirectional Semantic Alignment for Open-Vocabulary Camouflaged Object Segmentationaccepted
  165. Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindnessaccepted
  166. Seeing Conversations: Communication Context Identification in Egocentric Videoaccepted
  167. Seeing Depth Through Frequency and Motion: A Progressive Training Paradigm for Monocular Depth Estimationaccepted
  168. Seeing Motion Through Polarity for Event-based Action Recognitionaccepted
  169. Seeing Through Blur: Tackling Defocus in Spike-Based Imagingaccepted
  170. Seeing Through Touch: Tactile-Driven Visual Localization of Material Regionsaccepted
  171. Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspectiveaccepted
  172. Seeing Through the Shift: Causality-Inspired Robust Generalized Category Discoveryaccepted
  173. Seeing What Matters: A Training-Free Self-Guided Framework for Multimodal Detail Perception and Reasoningaccepted
  174. Seeing What Matters: Visual Preference Policy Optimization for Visual Generationaccepted
  175. Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understandingaccepted
  176. Seeing is Improving: Visual Feedback for Iterative Text Layout Refinementaccepted
  177. Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmarkaccepted
  178. Seeing through Light and Darkness: Sensor-Physics Grounded Deblurring HDR NeRF from Single-Exposure Images and Eventsaccepted
  179. Seeing through boxes: Non-Line-of-Sight 3D Reconstruction from Radar Signalsaccepted
  180. Seeing without Pixels: Perception from Camera Trajectoriesaccepted
  181. Seele: A Unified Acceleration Framework for Real-Time Gaussian Splatting on Mobile Devicesaccepted
  182. SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentationaccepted
  183. SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Imagesaccepted
  184. SegGBC: Justifiable Coarse-to-Fine Granular-Ball Computing for Enhancing Clustering Image Segmentationaccepted
  185. SegMo: Co-Designing Content-Aware Sparsity and Locally-Cohesive Segment Parallelism for Efficient VLM Inferenceaccepted
  186. SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentationaccepted
  187. SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Modelsaccepted
  188. SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMsaccepted
  189. Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoningaccepted
  190. Select, Hypothesize and Verify: Towards Verified Neuron Concept Interpretationaccepted
  191. Selection-as-Nonlinearity: Bridging Attention and Activation via a Joint Game-Decision Lens for Interpretable, Discriminative Visual Representationsaccepted
  192. Selective Amnesia using Contrastive Subnet Erasure for Class Level Unlearning in Vision Modelsaccepted
  193. Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentationaccepted
  194. Selectively Extracting and Injecting Visual Attributes into Text-to-Image Modelsaccepted
  195. Self-Attention Driven Tensor Representation for High-Order Data Recoveryaccepted
  196. Self-Consistency for LLM-Based Motion Trajectory Generation and Verificationaccepted
  197. Self-Corrected Image Generation with Explainable Latent Rewardsaccepted
  198. Self-Critical Distillation Network for Video-based Commonsense Captioningaccepted
  199. Self-Diffusion Driven Blind Imagingaccepted
  200. Self-Evaluation Unlocks Any-Step Text-to-Image Generationaccepted
  201. Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generationaccepted
  202. Self-guided Semantic Inspection for Zero-Shot Composed Image Retrievalaccepted
  203. Self-supervised Dynamic Heterogeneous Degradation Modeling for Unified Zero-Shot Image Restorationaccepted
  204. SelfHVD: Self-Supervised Handheld Video Deblurringaccepted
  205. Selfi: Self-improving Reconstruction Engine via 3D Geometric Feature Alignmentaccepted
  206. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detectionaccepted
  207. SemLayer: Semantic-aware Generative Segmentation and Layer Construction for Abstract Iconsaccepted
  208. SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidanceaccepted
  209. Semantic Alignment for Pose-Invariant Identity Preserving Diffusionaccepted
  210. Semantic Audio-Visual Navigation in Continuous Environmentsaccepted
  211. Semantic Context Matters: Improving Conditioning for Autoregressive Modelsaccepted
  212. Semantic Derivative Flow: Graph-Guided Diffusion for Controllable Instance Interactionsaccepted
  213. Semantic Foam: Unifying Spatial and Semantic Scene Decompositionaccepted
  214. Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learningaccepted
  215. Semantic Scale Space: A Framework for Controllable Image Abstractionaccepted
  216. Semantic-Adaptive Diffusion for Dynamic Spatiotemporal Fusionaccepted
  217. Semantic-Guided Global-Local Collaborative Prompt Learning for Few-Shot Class Incremental Learningaccepted
  218. SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action Planningaccepted
  219. Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusionaccepted
  220. Semi-Supervised Conformal Prediction With Unlabeled Nonconformity Scoreaccepted
  221. Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label Reforgingaccepted
  222. SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentationaccepted
  223. SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Cachingaccepted
  224. SenseSearch: Empowering Vision-Language Models with High-Resolution Agentic Search-Reasoning via Reinforcement Learningaccepted
  225. Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Drivingaccepted
  226. ShadowDraw: From Any Object to Shadow-Drawing Compositional Artaccepted
  227. Shape-of-You: Fused Gromov-Wasserstein Optimal Transport for Semantic Correspondence in-the-Wildaccepted
  228. ShapeAR: Generating Editable Shape Layers via Autoregressive Diffusionaccepted
  229. ShapeR: Robust Conditional 3D Shape Generation from Casual Capturesaccepted
  230. SharpTimeGS: Sharp and Stable Dynamic Gaussian Splatting via Lifespan Modulationaccepted
  231. Shedding Light on VLN Robustness: A Black-box Framework for Indoor Lighting-based Adversarial Attackaccepted
  232. ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimationaccepted
  233. ShiftLUT: Spatial Shift Enhanced Look-Up Tables for Efficient Image Restorationaccepted
  234. Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimationaccepted
  235. ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitionsaccepted
  236. ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinementaccepted
  237. ShowUI-p: Flow-based Generative Models as GUI Dexterous Handsaccepted
  238. ShreddingNet: Coarse-to-Fine Restoration for Multi-Source Shredded Manuscriptsaccepted
  239. SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Modelsaccepted
  240. SignPR: A Progressive Vector-Quantized Diffusion Framework for Sign Language Productionaccepted
  241. SimLBR: Learning to Detect Fake Images by Learning to Detect Real Imagesaccepted
  242. SimRecon: SimReady Compositional Scene Reconstruction from Real Videosaccepted
  243. SimScale: Learning to Drive via Real-World Simulation at Scaleaccepted
  244. Similarity-Consistent Likelihood Diffusion enables Hidden Person Detection from Wall Reflectionsaccepted
  245. Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learningaccepted
  246. Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimizationaccepted
  247. Simple but Effective Triplet-Based Compression Strategies for Compact Visual Localizationaccepted
  248. Simple-ViLMedSAM: Simple Text Prompts Meet Vision-Language Models for Medical Image Segmentationaccepted
  249. SinGeo: Unlock Single Model's Potential for Robust Cross-View Geo-Localizationaccepted
  250. SineProject: Machine Unlearning for Stable Vision-Language Alignmentaccepted
  251. Single-Round Scalable Analytic Federated Learningaccepted
  252. Single-step Diffusion-based Video Coding with Semantic-Temporal Guidanceaccepted
  253. SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognitionaccepted
  254. Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume Generationaccepted
  255. Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow Distillationaccepted
  256. SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawingaccepted
  257. SketchDeco: Training-Free Latent Composition for Precise Sketch Colourisationaccepted
  258. SketchFaceGS: Real-Time Sketch-Driven Face Editing and Generation with Gaussian Splattingaccepted
  259. SketchRevive: Fine-Grained Pixel-to-Vector Sketch Completion with Diffusion-Prior-Guided Multimodal LLMsaccepted
  260. SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and Moreaccepted
  261. SkillSight: Efficient First-Person Skill Assessment with Gazeaccepted
  262. Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal Predictionaccepted
  263. Sky2Ground: A Benchmark for Site Modeling under Varying Altitudeaccepted
  264. SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Designaccepted
  265. SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing Imageryaccepted
  266. Skyra: AI-Generated Video Detection via Grounded Artifact Reasoningaccepted
  267. SliderEdit: Continuous Image Editing with Fine-Grained Instruction Controlaccepted
  268. Small Object, Great Challenge: A Benchmark for Small Object Visual Groundingaccepted
  269. Smart Replay: Adaptive Scheduling of Memory Rehearsal for Computational Resource-Aware Incremental Learningaccepted
  270. SmokeSVD: Smoke Reconstruction from A Single View via Progressive Novel View Synthesis and Refinement with Diffusion Modelsaccepted
  271. Smoothing the Score Function to Enhance Generalization in Diffusion Modelsaccepted
  272. SoC: Semantic Orthogonal Calibration for Test-Time Prompt Tuningaccepted
  273. SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsaccepted
  274. SoccerMaster: A Vision Foundation Model for Soccer Understandingaccepted
  275. SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigationaccepted
  276. Socratic-Geo: Synthetic Data Generation and Cross-Modal Geometric Reasoning via Multi-Agent Interactionaccepted
  277. Soft Modality-Guided Expert Specialization in MoE-VLMsaccepted
  278. SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Modelsaccepted
  279. Solvability of the Viewing Graph Under the Affine Camera Modelaccepted
  280. Solving Minimal Problems Without Matrix Inversion Using FFT-Based Interpolationaccepted
  281. Solving a Nonlinear Blind Inverse Problem for Tagged MRI with Physics and Deep Generative Priorsaccepted
  282. SonoWorld: From One Image to a 3D Audio-Visual Sceneaccepted
  283. Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animationaccepted
  284. SounDiT: Geo-Contextual Soundscape-to-Landscape Generationaccepted
  285. Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Groupingaccepted
  286. SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Drivingaccepted
  287. SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Modelsaccepted
  288. SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Timeaccepted
  289. SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RLaccepted
  290. SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free Accelerationaccepted
  291. Sparse Spectral LoRA: Routed Experts for Medical VLMsaccepted
  292. Sparse Task Vector Mixup with Hypernetworks for Efficient Knowledge Transfer in Whole-Slide Image Prognosisaccepted
  293. Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Modelsaccepted
  294. Sparse-View Localization via Online Neural 3D Regressionaccepted
  295. SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse Camerasaccepted
  296. SparseOIT: Improving Order-Independent Transparency 3DGS via Active Set Methodaccepted
  297. SparseSplat: Towards Applicable Feed-Forward 3D Gaussian Splatting with Pixel-Unaligned Predictionaccepted
  298. SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Modelaccepted
  299. Sparsely Timing the Change: A Spiking Temporal Framework for Remote Sensing Interpretationaccepted
  300. Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detectionaccepted
  301. Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completionaccepted
  302. Spatia: Video Generation with Updatable Spatial Memoryaccepted
  303. SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Modelsaccepted
  304. Spatial Matters: Position-Guided 3D Referring Expression Segmentationaccepted
  305. Spatial Retrieval Augmented Autonomous Drivingaccepted
  306. Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videosaccepted
  307. Spatial-Frequency Collaborative Learning for Occluded Visible-Infrared Person Re-Identificationaccepted
  308. Spatial-SAM: Spatially Consistent 3D Electron Microscopy Segmentation with SDF Memory and Semi-Supervised Learningaccepted
  309. Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learningaccepted
  310. Spatial-Spectral Residuals Informed Diffusion Neural Operator for Pan-sharpeningaccepted
  311. SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modelingaccepted
  312. SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generationaccepted
  313. SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligenceaccepted
  314. SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoningaccepted
  315. SpatialTree: How Spatial Intelligence Branches Out in MLLMsaccepted
  316. SpatialVID: A Large-Scale Video Dataset with Spatial Annotationsaccepted
  317. Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Trackingaccepted
  318. Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensoraccepted
  319. Spatiotemporal Pyramid Flow Matching for Climate Emulationaccepted
  320. Spe-BEVHead: Rethinking the Detection Head Design for Bird's-Eye-View Object Detectionaccepted
  321. Specificity-aware reinforcement learning for fine-grained open-world classificationaccepted
  322. Spectral Conformal Risk Control: Distribution-Free Tail Guarantees via Bayesian Quadratureaccepted
  323. Spectral Mixture-of-Experts for Continual Learningaccepted
  324. Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action Segmentationaccepted
  325. Spectral Super-Resolution via Adversarial Unfolding and Data-Driven Spectrum Regularization: From Multispectral Satellite Data to NASA Hyperspectral Imageaccepted
  326. Spectral-Geometric Neural Fields for Pose-Free LiDAR View Synthesisaccepted
  327. Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imageryaccepted
  328. Spectrum from Defocus: Fast Spectral Imaging with Chromatic Focal Stackaccepted
  329. SpeeDe3DGS: Speedy Deformable 3D Gaussian Splatting with Temporal Pruning and Motion Groupingaccepted
  330. SpeeDiff: Scalable Pixel-Anchored End-to-End Latent Diffusion Modelaccepted
  331. Speeding Up the Learning of 3D Gaussians with Much Shorter Gaussian Listsaccepted
  332. Spherical Leech Quantization for Visual Tokenization and Generationaccepted
  333. Spherical Voronoi: Directional Appearance as a Differentiable Partition of the Sphereaccepted
  334. SpiderCam: Low-Power Snapshot Depth from Differential Defocusaccepted
  335. Spike-driven Discrete Aggregation for Event-based Object Detectionaccepted
  336. SpikeTrack: A Spike-driven Framework for Efficient Visual Trackingaccepted
  337. SpikeTrack: High-performance and Energy-efficient Event-Based Object Tracking with Spiking Neural Networkaccepted
  338. SpiralDiff: Spiral Diffusion with LoRA for RGB-to-RAW Conversion Across Camerasaccepted
  339. Spk2VidNet: A Hierarchical Recurrent Architecture for High-Fidelity Video Reconstruction from Long Spike-Camera Streamsaccepted
  340. Splat-Based Metal Artifact Reduction in Cone-Beam CT via Compact Attenuation Modelingaccepted
  341. SplatSuRe: Selective Super-Resolution for Multi-view Consistent 3D Gaussian Splattingaccepted
  342. Splatent: Splatting Diffusion Latents for Novel View Synthesisaccepted
  343. SplitFlux: Learning to Decouple Content and Style from a Single Imageaccepted
  344. Spot The Ball: A Benchmark for Visual Social Inferenceaccepted
  345. SpotEdit: Selective Region Editing in Diffusion Transformersaccepted
  346. StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representationaccepted
  347. StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answeringaccepted
  348. Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulationaccepted
  349. Stabilizing Feature Geometry in Noisy Pretrained Models for Robust Downstream Tasksaccepted
  350. Stabilizing Streaming Video Geometry via Dynamic Feature Normalizationaccepted
  351. Stable Mean Flow: Lyapunov-Inspired One-Step Flow Matchingaccepted
  352. Stable Spike: Dual Consistency Optimization via Bitwise AND Operations for Spiking Neural Networksaccepted
  353. Stable and Efficient Single-Rollout RL for Multimodal Reasoningaccepted
  354. StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasetsaccepted
  355. StableMaterials: Enhancing Diversity in Material Generation via Semi-Supervised Learningaccepted
  356. Stake the Points: Structure-Faithful Instance Unlearningaccepted
  357. Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generationaccepted
  358. Statistical Characteristic-Guided Denoising for Rapid High-Resolution Transmission Electron Microscopy Imagingaccepted
  359. Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioningaccepted
  360. Stealing Split Learning Bottom Models by Recovering Embedding Geometryaccepted
  361. Steering Where to Diffuse: Generative Modeling of Phenotypic Response Simulation with Steered Diffusion Bridgeaccepted
  362. Stepwise Credit Assignment for GRPO on Flow-Matching Modelsaccepted
  363. Stereo World Model: Camera-Guided Stereo Video Generationaccepted
  364. StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generationaccepted
  365. Stitch-a-Demo: Creating Video Demonstrations from Multistep Descriptionsaccepted
  366. Stochastic Ray Tracing for the Reconstruction of 3D Gaussian Splattingaccepted
  367. StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narrativesaccepted
  368. StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatarsaccepted
  369. StreamDiT: Real-Time Streaming Text-to-Video Generationaccepted
  370. StreamRAG: Enhancing Real-Time Video Understanding with Retrieval Augmentationaccepted
  371. StreamReady: Learning What to Answer and When in Long Streaming Videosaccepted
  372. StreamVLO: Streaming Visual-LiDAR Odometry with Cumulative Drift Compensationaccepted
  373. Streaming Diffusion Model for Fast Infrared and Visible Video Fusionaccepted
  374. Streaming Video Crime Anticipation with Spatio-Temporal Causal Reasoningaccepted
  375. Streaming Video Instruction Tuningaccepted
  376. StreamingTOM: Streaming Token Compression for Efficient Video Understandingaccepted
  377. Streamlined Knowledge Distillationaccepted
  378. Streamlined Open-Vocabulary Human-Object Interaction Detectionaccepted
  379. Stronger Normalization-Free Transformersaccepted
  380. StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cuesaccepted
  381. Structural Action Transformer for 3D Dexterous Manipulationaccepted
  382. Structural Graph Probing of Vision-Language Modelsaccepted
  383. Structure-Aware Representation Distillation for Tiny-Dense Object Segmentationaccepted
  384. Structure-to-Intensity Diffusion for Adverse-Weather LiDAR Generationaccepted
  385. Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward Modelingaccepted
  386. StyleDoctor: Towards Specialist Reward Model for Style-centric Generation Tasksaccepted
  387. StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image Referencesaccepted
  388. StyleTextGen: Style-Conditioned Multilingual Scene Text Generationaccepted
  389. SuP: Sub-cloud Driven Point Cloud Registrationaccepted
  390. Submodel Extraction for Efficient and Personalized Federated Learning via Optimal Transportaccepted
  391. Subspace Alignment for CLIP-based Continual Learning via Canonical Correlation Analysisaccepted
  392. SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modelingaccepted
  393. SunFaded: Illumination-Aware Gaussian Splatting for Dark Scenes with Camera-Mounted Active Lightingaccepted
  394. Superman: Unifying Skeleton and Vision for Human Motion Perception and Generationaccepted
  395. Suppressing Non-Semantic Noise in Masked Image Modeling Representationsaccepted
  396. SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmarkaccepted
  397. SwiftTailor: Efficient 3D Garment Generation with Geometry Image Representationaccepted
  398. SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overheadaccepted
  399. SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controlsaccepted
  400. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generationaccepted
  401. Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understandingaccepted
  402. SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perceptionaccepted
  403. SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generationaccepted
  404. SyncDreamer: Controllable and Expressive Avatar Generation Beyond the Talking Headaccepted
  405. SyncMos: Scalable Motion Synchronisation for Multi-Agent Scene Interactionaccepted
  406. Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videosaccepted
  407. SynthRGB-T: Language-Vision Guided Image Translation for Diversity Synthesisaccepted
  408. Synthesizing Visual Concepts as Vision-Language Programsaccepted
  409. Synthetic Curriculum Reinforces Compositional Text-to-Image Generationaccepted
  410. Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Groundingaccepted
  411. T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Groundingaccepted
  412. TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detectionaccepted
  413. TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Expertsaccepted
  414. TALO: Pushing 3D Vision Foundation Models Towards Globally Consistent Online Reconstructionaccepted
  415. TALON: Test-time Adaptive Learning for On-the-Fly Category Discoveryaccepted
  416. TAMER: A Tri-Modal Contrastive Alignment and Multi-Scale Embedding Refinement Framework for Zero-Shot ECG Diagnosisaccepted
  417. TANGO: Learning Distribution-wise Foundation Prior Consistency and Instance-wise Style Calibration for Medical Image Generalizationaccepted
  418. TANGO: Text-Anchored Guided Optimization for Robust Fine-tuning Vision-Language Models under Label Noiseaccepted
  419. TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion Accelerationaccepted
  420. TAPE: Task-Adaptive Prototype Evolution in Audio-Language Models for Fully Few-shot Class-incremental Audio Classificationaccepted
  421. TAR: Token-Aware Refinement for Fine-grained Generalized Category Discoveryaccepted
  422. TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Expertsaccepted
  423. TAlignDiff: Automatic Tooth Alignment assisted by Diffusion-based Transformation Learningaccepted
  424. TC-Pade: Trajectory-Consistent Pade Approximation for Diffusion Accelerationaccepted
  425. TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignmentaccepted
  426. TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Modelsaccepted
  427. TESO: Online Tracking of Essential Matrix by Stochastic Optimizationaccepted
  428. TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysisaccepted
  429. TEXTRIX: Latent Attribute Grid for Native Texture Generation and Beyondaccepted
  430. TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detectionaccepted
  431. TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detectionaccepted
  432. TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completionaccepted
  433. TGT: Text-Grounded Trajectories for Locally Controlled Video Generationaccepted
  434. TGTrack: Temporal Generative Learning for Unified Single Object Trackingaccepted
  435. THE MORE, THE MERRIER: CONTRASTIVE FUSION FOR HIGHER-ORDER MULTIMODAL ALIGNMENTaccepted
  436. TIACam: Text-Anchored Invariant Feature Learning with Auto-Augmentation for Camera-Robust Zero-Watermarkingaccepted
  437. TIGER: A Unified Framework for Time, Images and Geo-location Retrievalaccepted
  438. TIM: Temporal Decoupling with Iterative Mutual-Refinement Model for Longitudinal Radiology Report Generationaccepted
  439. TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Modelsaccepted
  440. TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignmentaccepted
  441. TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detectionaccepted
  442. TM-BSN: Triangular-Masked Blind-Spot Network for Real-World Self-Supervised Image Denoisingaccepted
  443. TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Dual-Level Scale-Oriented Contrastaccepted
  444. TRANSPORTER: Transferring Visual Semantics from VLM Manifoldsaccepted
  445. TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognitionaccepted
  446. TRIDENT: A Trimodal Cascade Generative Framework for Drug and RNA-Conditioned Cellular Morphology Synthesisaccepted
  447. TRM-VLA: Temporal-Aware Chain-of-Thought Reasoning and Memorization for Vision-Language-Action Modelsaccepted
  448. TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videosaccepted
  449. TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognitionaccepted
  450. TSTM: Temporal Segmentation for Task-relevant Mask in Visual Reinforcement Learning Generalizationaccepted
  451. TTAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Eventsaccepted
  452. TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Modelsaccepted
  453. TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Modelsaccepted
  454. TTRV: Test-Time Reinforcement Learning for Vision Language Modelsaccepted
  455. TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolutionaccepted
  456. TUNA: Taming Unified Visual Representations for Native Unified Multimodal Modelsaccepted
  457. TV2TV: A Unified Framework for Interleaved Language and Video Generationaccepted
  458. TVHighlights: LLM-Guided Human-Free Collaborative Training for Video Highlight Detection in Movies and TV Dramasaccepted
  459. TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummiesaccepted
  460. TWINGS: Thin Plate Splines Warp-aligned Initialization for Sparse-View Gaussian Splattingaccepted
  461. TableMix: Enhancing Multimodal Table Reasoning in MLLMs from a Data-Centric Perspectiveaccepted
  462. TacSIm: A Dataset and Benchmark for Football Tactical Style Imitationaccepted
  463. Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute Miningaccepted
  464. Tackling Model Bias via Game-theoretic Multi-agent Collaboration Framework for Hateful Meme Classificationaccepted
  465. TagSplat: Topology-Aware Gaussian Splatting for Dynamic Mesh Modeling and Trackingaccepted
  466. Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenesaccepted
  467. Talking Together: Synthesizing Co-Located 3D Conversations from Audioaccepted
  468. Taming Generative Diffusion Model for Task-Oriented Infrared Imagingaccepted
  469. Taming Noise-Induced Prototype Degradation for Privacy-Preserving Personalized Federated Fine-Tuningaccepted
  470. Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learningaccepted
  471. Taming Sampling Perturbations with Variance Expansion Loss for Latent Diffusion Modelsaccepted
  472. Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Controlaccepted
  473. Taming the Long Tail: Rebalancing Adversarial Training via Adaptive Perturbationaccepted
  474. Target-Aware Invertible Encoder with Reconstruction Guidance for Infrared Small Target Detectionaccepted
  475. Task-Aware Image Signal Processor for Advanced Visual Perceptionaccepted
  476. Task-Driven Implicit Representations for Automated Design of LiDAR Systemsaccepted
  477. Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentationaccepted
  478. TaskForce: Cooperative Multi-agent Reinforcement Learning for Multi-task Optimizationaccepted
  479. TaskIT: Memory-Efficient Fine-Tuning of Multi-LoRA LLMs via Cross-Task Importance Transferaccepted
  480. Tavatar: Topology-Aware Gaussian Attribute Derivation for Animatable Human Avatarsaccepted
  481. Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Modelsaccepted
  482. TeFlow: Enabling Multi-frame Supervision for Self-Supervised Feed-forward Scene Flow Estimationaccepted
  483. TeHOR: Text-Guided 3D Human and Object Reconstruction with Texturesaccepted
  484. Tea-Adapter: Teacher Adapter for Efficient Conditional Generationaccepted
  485. Teacher-Guided Routing for Sparse Vision Mixture-of-Expertsaccepted
  486. Teaching DINOv3 About Partial 3D Geometry: A Self-Supervised Geometry-Aware Approachaccepted
  487. TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Sizeaccepted
  488. Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attentionaccepted
  489. Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Modelaccepted
  490. TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learningaccepted
  491. TempoControl: Temporal Attention Guidance for Text-to-Video Modelsaccepted
  492. TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Predictionaccepted
  493. Temporal Equilibrium MeanFlow: Bridging the Scale Gap for One-Step Generationaccepted
  494. Temporal Imbalance of Positive and Negative Supervision in Class-Incremental Learningaccepted
  495. Temporal Interaction in Spiking Transformers with Multi-Delay Mixeraccepted
  496. Temporal Inversion for Learning Interval Change in Chest X-Raysaccepted
  497. Temporal Representation Enhancement (TRE): Learning to Forget Dominant Patterns for Enhanced Temporal Spiking Featuresaccepted
  498. TerraScope: Pixel-Grounded Visual Reasoning for Earth Observationaccepted
  499. TerraSeg: Self-Supervised Ground Segmentation for Any LiDARaccepted
  500. Test-Time 3D Occupancy Predictionaccepted
  501. Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisationaccepted
  502. Test-Time Attention Purification for Backdoored Large Vision Language Modelsaccepted
  503. Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative Modelingaccepted
  504. Test-Time Multi-Prompt Adaptation for Open-Vocabulary Remote Sensing Image Segmentationaccepted
  505. Test-Time Perturbation Tuning with Delayed Feedback for Vision-Language-Action Modelsaccepted
  506. Test-Time Training for LiDAR Semantic Segmentation under Corruption via Geometric Inlier Discriminationaccepted
  507. Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistencyaccepted
  508. Test-time Sparsity for Extreme Fast Action Diffusionaccepted
  509. Text-Driven 3D Hand Motion Generation from Sign Language Dataaccepted
  510. Text-Image Conditioned 3D Generationaccepted
  511. Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image Retrievalaccepted
  512. Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Modelsaccepted
  513. Text-guided Feature Disentanglement for Cross-modal Gait Recognitionaccepted
  514. TextFM: Robust Semi-dense Feature Matching with Language Guidanceaccepted
  515. TextOVSR: Text-Guided Real-World Opera Video Super-Resolutionaccepted
  516. TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Renderingaccepted
  517. Texvent: Asynchronous Event Data Simulation via Text Promptaccepted
  518. The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Modelsaccepted
  519. The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contextsaccepted
  520. The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignmentaccepted
  521. The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category Discoveryaccepted
  522. The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergyaccepted
  523. The Drift Kernel: Why Diffusion Models Change Even When Told Not Toaccepted
  524. The Geometry of Robustness: Optimizing Loss Landscape Curvature and Feature Manifold Alignment for Robust Finetuning of Vision-Language Modelsaccepted
  525. The Golden Subspace: Where Efficiency Meets Generalization in Continual Test-Time Adaptationaccepted
  526. The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generationaccepted
  527. The Invisible Gorilla Effect in Out-of-distribution Detectionaccepted
  528. The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognitionaccepted
  529. The Midas Touch for Metric Depthaccepted
  530. The Missing GAP: From Solving Square Jigsaw Puzzles to Handling Real World Archaeological Fragmentsaccepted
  531. The Missing Point in Vision Transformers for Universal Image Segmentationaccepted
  532. The Power of Decaying Steps: Enhancing Attack Stability and Transferability for Sign-based Optimizersaccepted
  533. The Power of Prior: Training-Free Open-Vocabulary Semantic Segmentation with LLaVAaccepted
  534. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detectionaccepted
  535. The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identificationaccepted
  536. The Surprising Effectiveness of Noise Pretraining for Implicit Neural Representationsaccepted
  537. The Universal Normal Embeddingaccepted
  538. The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injectionaccepted
  539. TherA: Thermal-Aware Visual-Language Prompting for Controllable RGB-to-Thermal Infrared Translationaccepted
  540. Thermal Diffusion Matters: Infrared Spatial-Temporal Video Super-Resolution through Heat Conduction Priorsaccepted
  541. Thermal is Always Wild: Characterizing and Addressing Challenges in Thermal-Only Novel View Synthesisaccepted
  542. Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detectionaccepted
  543. Thermally Activated Dual-Modal Adversarial Clothing against AI Surveillance Systemsaccepted
  544. Think 360deg: Beyond Depth: Evaluating the Width-centric Reasoning Capability of MLLMsaccepted
  545. Think Before You Drive: World Model-Inspired Multimodal Groundingaccepted
  546. Think Visually, Reason Textually: Vision-Language Synergy in Abstract Reasoningaccepted
  547. Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Viewsaccepted
  548. Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understandingaccepted
  549. Think-Then-Generate: Structural Chain-of-Thought Reasoning for Consistent 3D Generationaccepted
  550. Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Modelsaccepted
  551. ThinkGen: Generalized Thinking for Visual Generationaccepted
  552. Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMsaccepted
  553. Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Modelsaccepted
  554. Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoningaccepted
  555. Thinking in 360deg: Humanoid Visual Search in the Wildaccepted
  556. Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D Worldaccepted
  557. Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decodingaccepted
  558. Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understandingaccepted
  559. Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Modelaccepted
  560. Thinking with Programming Vision: Towards a Unified View for Thinking with Imagesaccepted
  561. Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigmaccepted
  562. Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generationaccepted
  563. ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inferenceaccepted
  564. Through the Frequency Lens: Cross-Domain Generalisable Gaze Estimation with Adaptive Modulationaccepted
  565. TiViBench: Benchmarking Think-in-Video Reasoning for Video Generationaccepted
  566. Time Blindness: Why Video-Language Models Can't See What Humans Can?accepted
  567. Time Without Time: Pseudo-Temporal Representation for Space-Time Super-Resolutionaccepted
  568. Time-Aware One Step Diffusion Network for Real-World Image Super-Resolutionaccepted
  569. Time-Specialized Event-Image Alignment for Blur-to-Video Decompositionaccepted
  570. TimeBridge: Self-Supervised Video Representation Learning via Start-End Joint Embedding and In-Between Frame Predictionaccepted
  571. TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMsaccepted
  572. TimeRipples: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Spaceaccepted
  573. TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understandingaccepted
  574. Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Modelsaccepted
  575. Token Warping Helps MLLMs Look from Nearby Viewpointsaccepted
  576. TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokensaccepted
  577. TokenHand: Discrete Token Representation for Efficient Hand Mesh Reconstructionaccepted
  578. TokenLight: Precise Lighting Control in Images using Attribute Tokensaccepted
  579. TokenSplat: Token-aligned 3D Gaussian Splatting for Feed-forward Pose-free Reconstructionaccepted
  580. TokenTrace: Multi-Concept Attribution through Watermarked Token Recoveryaccepted
  581. Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor Plansaccepted
  582. Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelityaccepted
  583. TopoCL: Topological Contrastive Learning for Medical Imagingaccepted
  584. TopoHR: Hierarchical Centerline Representation for Cyclic Topology Reasoning in Driving Scenes with Point-to-Instance Relationsaccepted
  585. TopoMA: Topology-Guided Multi-Agent Dense RGB 3D Reconstruction via Distributed Inferenceaccepted
  586. TopoMesh: High-Fidelity Mesh Autoencoding via Topological Unificationaccepted
  587. TopoSlide: Topologically-Informed Histopathology Whole Slide Image Representation Learningaccepted
  588. Topology-aware Feature Propagation for Unsupervised Non-rigid Point Cloud Correspondenceaccepted
  589. TouchDream: 3D Object Completion through Imagined Touchaccepted
  590. Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspectiveaccepted
  591. Toward Early Quality Assessment of Text-to-Image Diffusion Modelsaccepted
  592. Toward Generalizable Whole Brain Representations with High-Resolution Light-Sheet Dataaccepted
  593. Toward Low-Cost yet Effective Temporal Learning for UAV Trackingaccepted
  594. Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Datasetaccepted
  595. Towards Balanced Multi-Modal Learning in 3D Human Pose Estimationaccepted
  596. Towards Calibrating Prompt Tuning of Vision- Language Modelsaccepted
  597. Towards Cross-Modal Preservation, Consistency and Alignment for Privacy-Preserving Visible-Infrared Person Re-Identificationaccepted
  598. Towards Decompositional Human Motion Generation with Energy-Based Diffusion Modelsaccepted
  599. Towards Dynamic Modality Alignment in Multimodal Continual Learningaccepted
  600. Towards Efficient Medical Reasoning with Minimal Fine-Tuning Dataaccepted
  601. Towards Fine-Grained Attribution: Instance-Aware Preference Optimization for Aligning Diffusion Modelsaccepted
  602. Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Cloudsaccepted
  603. Towards GUI Agents: Vision-Language Diffusion Models for GUI Groundingaccepted
  604. Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learningaccepted
  605. Towards Generalized Multimodal Homography Estimationaccepted
  606. Towards Generalized Representations for Low-Light Understanding: When Signal Constancy Meets Semantic Enrichmentaccepted
  607. Towards High-Quality Image Segmentation: Improving Topology Accuracy by Penalizing Neighbor Pixelsaccepted
  608. Towards High-resolution and Disentangled Reference-based Sketch Colorizationaccepted
  609. Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interactionaccepted
  610. Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimizationaccepted
  611. Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion Transformersaccepted
  612. Towards Human-Imperceptible Backdoor Attacks on Text-to-Image Diffusion Modelsaccepted
  613. Towards Human-Like Robot Handwriting via Contour-Aware Generationaccepted
  614. Towards Intrinsic-Aware Monocular 3D Object Detectionaccepted
  615. Towards Knowledge-augmented Bayesian Deep Learning For Computer Visionaccepted
  616. Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid Robotsaccepted
  617. Towards Multimodal Domain Generalization with Few Labelsaccepted
  618. Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoningaccepted
  619. Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Datasetaccepted
  620. Towards Persistence: Learning Topological Constraints for Event-based Small Object Detectionaccepted
  621. Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Frameworkaccepted
  622. Towards Policy-Adaptive Image Guardrail: Benchmark and Methodaccepted
  623. Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Trainingaccepted
  624. Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priorsaccepted
  625. Towards Reasoning-Preserving Unlearning in Multimodal Large Language Modelsaccepted
  626. Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networksaccepted
  627. Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype Distillationaccepted
  628. Towards Robust Multimodal Large Language Models Against Jailbreak Attacksaccepted
  629. Towards Robust Sequential Decomposition for Complex Image Editingaccepted
  630. Towards Robust Vision Transformers: Path Dependency Analysis and a Simple Two-Stage Adversarial Trainingaccepted
  631. Towards Sparse Video Understanding and Reasoningaccepted
  632. Towards Stable Federated Continual Test-Time Adaptation in Wild Worldaccepted
  633. Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Videoaccepted
  634. Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approachaccepted
  635. Towards Storytelling Animations: Joint Synthesis of Human and Camera Motionsaccepted
  636. Towards Streaming Referring Video Segmentation via Large Language Modelaccepted
  637. Towards Training-free Scene Text Editingaccepted
  638. Towards Uncertainty-aware Unsupervised Domain Adaptation for Videos and Time-Series with Causal Optimal Transportaccepted
  639. Towards Unified Human Perception and Machine Understanding: Token Flow Guided Compression Frameworkaccepted
  640. Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysisaccepted
  641. Towards Visual Query Localization in the 3D Worldaccepted
  642. Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspectiveaccepted
  643. TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videosaccepted
  644. TrackMAE: Video Representation Learning via Track Mask and Predictaccepted
  645. Tracking by Predicting 3-D Gaussians Over Timeaccepted
  646. Tracking through Severe Occlusion via Event-Derived Transient Cuesaccepted
  647. Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animationaccepted
  648. TrafficAlign: Aligning Large Language Models for Traffic Scenario Generationaccepted
  649. Trainable Log-linear Sparse Attention for Efficient Diffusion Transformersaccepted
  650. Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automationaccepted
  651. Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learningaccepted
  652. Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid Promptaccepted
  653. Training-Only Heterogeneous Image-Patch-Text Graph Supervision for Advancing Few-Shot Learning Adaptersaccepted
  654. Training-free Detection of Generated Videos via Spatial-Temporal Likelihoodsaccepted
  655. Training-free Mixed-Resolution Latent Upsampling for Spatially Accelerated Diffusion Transformersaccepted
  656. Training-free Motion Factorization for Compositional Video Generationaccepted
  657. Training-free, Perceptually Consistent Low-Resolution Previews with High-Resolution Image for Efficient Workflows of Diffusion Modelsaccepted
  658. TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigationaccepted
  659. TrajTok: Learning Trajectory Tokens Enhances Video Understandingaccepted
  660. TransPrune: Token Transition Pruning for Efficient Large Vision-Language Modelaccepted
  661. Transform to Transfer: Boosting Adversarial Attack Transferability on Vision-Language Pre-training Modelsaccepted
  662. Transition Matching Distillation for Fast Video Generationaccepted
  663. Transition Models: Rethinking the Generative Learning Objectiveaccepted
  664. Translating Signals to Languages for sEMG-Based Activity Recognitionaccepted
  665. TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Explorationaccepted
  666. Tri-Modal Fusion Transformers for UAV-based Object Detectionaccepted
  667. Tri-Subspaces Disentanglement for Multimodal Sentiment Analysisaccepted
  668. TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detectionaccepted
  669. TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglementaccepted
  670. TriSim: Tri-Dimensional Similarity Modeling with Extreme Value Theory for False-Negative Mitigation in Remote Sensing Image-Text Retrievalaccepted
  671. TruckDrive: Long-Range Autonomous Highway Driving Datasetaccepted
  672. Trust-calibrated Collaborative Learning for Long-Tailed Visual Recognitionaccepted
  673. Tunable Soft Equivariance with Guaranteesaccepted
  674. Turbo-GS: Accelerating 3D Gaussian Fitting for High-Resolution Radiance Fieldsaccepted
  675. Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival Predictionaccepted
  676. Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detectionaccepted
  677. Twin-T & TwintVQA: A Reliable Structure-Detail Separating VLM and a Comprehensive Benchmark for Chart and Table Tasksaccepted
  678. U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generationaccepted
  679. U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequencesaccepted
  680. UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancementaccepted
  681. UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVsaccepted
  682. UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detectionaccepted
  683. UAVLight: A Benchmark for Illumination-Robust 3D Reconstruction in Unmanned Aerial Vehicle (UAV) Scenesaccepted
  684. UCAN: Unified Convolutional Attention Network for Expansive Receptive Fields in Lightweight Super-Resolutionaccepted
  685. UCMNet: Uncertainty-Aware Context Memory Network for Under-Display Camera Image Restorationaccepted
  686. UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimationaccepted
  687. UETrack: A Unified and Efficient Framework for Single Object Trackingaccepted
  688. UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modelingaccepted
  689. UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Modelsaccepted
  690. UI-Lens: Assessing General MLLMs' Potential to Automate UI Display Quality Assuranceaccepted
  691. UIKA: Fast Universal Head Avatar from Pose-Free Imagesaccepted
  692. ULF-Loc: Unbiased Landmark Feature for Robust Visual Localization with 3D Gaussian Splattingaccepted
  693. UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language Modelingaccepted
  694. UNICBench: UNIfied Counting Benchmark for MLLMaccepted
  695. UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attendersaccepted
  696. URICA: A Uniformity Region Affine Identifier Capture Algorithm for Arbitrary Region Retrieval in Pathology Imagesaccepted
  697. URScenes: A Multi-scenario Dataset for Unstructured Road Environmentsaccepted
  698. UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimationaccepted
  699. UTPTrack: Towards Simple and Unified Token Pruning for Visual Trackingaccepted
  700. UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigmaccepted
  701. UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language Conditionsaccepted
  702. U^2Flow: Uncertainty-Aware Unsupervised Optical Flow Estimationaccepted
  703. Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking from Sparse Inertial Sensors and Ranging-based Between-sensor Distancesaccepted
  704. Ultra-Fast Neural Video Compressionaccepted
  705. Ultra-Low Bitrate Perceptual Image Compression with Shallow Encoderaccepted
  706. UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratiosaccepted
  707. Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understandingaccepted
  708. UnReflectAnything: RGB-Only Highlight Removal by Rendering Synthetic Specular Supervisionaccepted
  709. Unblur-SLAM: Dense Neural SLAM for Blurry Inputsaccepted
  710. Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Modelsaccepted
  711. Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Modelsaccepted
  712. Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object Detectionaccepted
  713. Uncertainty-driven 3D Gaussian Splatting Active Mapping via Anisotropic Visibility Fieldaccepted
  714. Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Modelsaccepted
  715. Underground Plant Exploration: Non-Destructive 3D Root Assessment with GPR Based on Point Graph Neural Networkaccepted
  716. Understanding Counting Mechanisms in Large Language and Vision-Language Modelsaccepted
  717. Understanding Task Transfer in Vision-Language Modelsaccepted
  718. Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminabilityaccepted
  719. Understanding and Enforcing Weight Disentanglement in Task Arithmeticaccepted
  720. Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Modelsaccepted
  721. Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Modelsaccepted
  722. Understanding, Accelerating, and Improving MeanFlow Trainingaccepted
  723. Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generationaccepted
  724. Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing Modalitiesaccepted
  725. Uni-Hema: Unified Model for Digital Hematopathologyaccepted
  726. Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Imagesaccepted
  727. UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactionsaccepted
  728. UniChange: Unifying Change Detection with Multimodal Large Language Modelaccepted
  729. UniComp: Rethinking Video Compression Through Informational Uniquenessaccepted
  730. UniCompress: Token Compression for Unified Vision-Language Understanding and Generationaccepted
  731. UniCorrn: Unified Correspondence Transformer Across 2D and 3Daccepted
  732. UniDAC: Universal Metric Depth Estimation for Any Cameraaccepted
  733. UniDef: Universal Defense Against Unauthorized Image Manipulationaccepted
  734. UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videosaccepted
  735. UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifyingaccepted
  736. UniFusion: A Unified Image Fusion Framework with Robust Representation and Source-Aware Preservationaccepted
  737. UniGame: Turning a Unified Multimodal Model Into Its Own Adversaryaccepted
  738. UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in RLaccepted
  739. UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detectionaccepted
  740. UniGeoRS: A Unified Benchmark for Tri-view Geo-Localizationaccepted
  741. UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenesaccepted
  742. UniLDiff: Unlocking the Power of Diffusion Priors for All-in-One Image Restorationaccepted
  743. UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speakingaccepted
  744. UniLight: A Unified Representation for Lightingaccepted
  745. UniM: A Unified Any-to-Any Interleaved Multimodal Benchmarkaccepted
  746. UniMERNet: A Universal Network for Real-World Mathematical Expression Recognitionaccepted
  747. UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompressionaccepted
  748. UniPR: Unified Object-level Real-to-Sim Perception and Reconstruction from a Single Stereo Pairaccepted
  749. UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latentsaccepted
  750. UniPercept: A Unified Diffusion Model for Generalizable Visual Perceptionaccepted
  751. UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matchingaccepted
  752. UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimizationaccepted
  753. UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Registeraccepted
  754. UniSER: A Foundation Model for Unified Soft Effects Removalaccepted
  755. UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Passaccepted
  756. UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual Promptingaccepted
  757. UniT: Unified Multimodal Chain-of-Thought Test-time Scalingaccepted
  758. UniTEX: Universal High Fidelity Generative Texturing for 3D Shapesaccepted
  759. UniVBench: Towards Unified Evaluation for Video Foundation Modelsaccepted
  760. UniVerse: A Unified Modulation Framework for Segmentation-Free, Disentangled Multi-Concept Personalizationaccepted
  761. UniVerse: Empower Unified Generation with Reasoning and Knowledgeaccepted
  762. UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Editsaccepted
  763. Unified Camera Positional Encoding for Controlled Video Generationaccepted
  764. Unified Customized Generation by Disentangled Reward Modelingaccepted
  765. Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimizationaccepted
  766. Unified Latent Space for Understanding and Generation via Semantic Auto-encoderaccepted
  767. Unified Multimodal Models as Auto-Encodersaccepted
  768. Unified Number-Free Text-to-Motion Generation Via Flow Matchingaccepted
  769. Unified Personalized Understanding, Generating and Editingaccepted
  770. Unified Primitive Proxies for Structured Shape Completionaccepted
  771. Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retentionaccepted
  772. Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Cameraaccepted
  773. Unified Vector Floorplan Generation via Markup Representationaccepted
  774. Unifying Language-Action Understanding and Generation for Autonomous Drivingaccepted
  775. Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generationaccepted
  776. Unifying Precise Keyframes and Semantic Control via Multi-level Diffusionaccepted
  777. Unique Lives, Shared World: Learning from Single-Life Videosaccepted
  778. UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generationaccepted
  779. Universal 3D Shape Matching via Coarse-to-Fine Language Guidanceaccepted
  780. Universal Guideline-Driven Image Clustering via a Hybrid LLM Agentaccepted
  781. Universal-to-Specific: Dynamic Knowledge-Guided Multiple Instance Learning for Few-Shot Whole Slide Image Classificationaccepted
  782. Unlearning without Forgetting: Securely Removing Targeted Concepts from Large-Scale Vision-Language Open-Vocabulary Detectorsaccepted
  783. Unleashing Stealthy Backdoor Pandemic by Infecting a Single Diffusion Modelaccepted
  784. Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from Failuresaccepted
  785. Unleashing Vision-Language Semantics for Deepfake Video Detectionaccepted
  786. Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Modelsaccepted
  787. Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detectionaccepted
  788. Unlocking 3D Affordance Segmentation with 2D Semantic Knowledgeaccepted
  789. Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait Recognitionaccepted
  790. Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Frameworkaccepted
  791. Unlocking Pre-trained Weights: Parameter Inheritance for Zero-Shot Initializationaccepted
  792. Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methodsaccepted
  793. Unlocking Token Rewards via Training-Free Reward Attributionaccepted
  794. Unlocking the Power of Critical Factors for 3D Visual Geometry Estimationaccepted
  795. Unpaired Image Deraining Using Reward-Guided Self-Reinforcement Strategyaccepted
  796. Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Imagesaccepted
  797. Unsafe2Safe: Controllable Image Anonymization for Downstream Utilityaccepted
  798. Unstitching the Chimera: Frame-Level Risk and Train-Free Mitigation for Video Hallucinationaccepted
  799. Unsupervised 3d Motion Estimation Using Event Cameraaccepted
  800. Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priorsaccepted
  801. Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Modelaccepted
  802. Unsupervised Multi-agent and Single-agent Perception from Cooperative Viewsaccepted
  803. Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsamplingaccepted
  804. Urban-GS: A Unified 3D Gaussian Splatting Framework for Compact and High-Fidelity Aerial-to-Street Reconstructionaccepted
  805. V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMsaccepted
  806. V-DPM: 4D Video Reconstruction with Dynamic Point Mapsaccepted
  807. V-RGBX: Video Editing with Accurate Controls over Intrinsic Propertiesaccepted
  808. V2U4Real: A Real-world Large-scale Dataset for Vehicle-to-UAV Cooperative Perceptionaccepted
  809. VA-p: Variational Policy Alignment for Pixel-Aware Autoregressive Generationaccepted
  810. VABench: A Comprehensive Benchmark for Audio-Video Generationaccepted
  811. VAD-GS: Visibility-Aware Densification for 3D Gaussian Splatting in Dynamic Urban Scenesaccepted
  812. VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generationaccepted
  813. VAST: Video Ability-Stratified Taxonomy for Data-Efficient Video Reasoningaccepted
  814. VCP-Attack: Visual-Contrastive Projection for Transferable Black-Box Targeted Attacks on Large Vision-Language Modelsaccepted
  815. VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridgingaccepted
  816. VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimationaccepted
  817. VDFE: Difference-Aware 3D Scene Editing with Non-Intrusive Video Diffusion Priors for Multi-View Consistency and Efficiencyaccepted
  818. VDOT: Efficient Unified Video Creation via Optimal Transport Distillationaccepted
  819. VEMamba: Efficient Isotropic Reconstruction of Volume Electron Microscopy with Axial-Lateral Consistent Mambaaccepted
  820. VENI: Variational Encoder for Natural Illuminationaccepted
  821. VES-RFT: Rewarding Visual Evidence Sensitivity to Mitigate Hallucinations in Large Vision-Language Modelsaccepted
  822. VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluationaccepted
  823. VGA: Empowering Aerial-Ground Localization by Visual Geometry Alignmentaccepted
  824. VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scaleaccepted
  825. VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Drivingaccepted
  826. VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimationaccepted
  827. VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detectionaccepted
  828. VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentationaccepted
  829. VGGT-Ωaccepted
  830. VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Predictionaccepted
  831. VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinementaccepted
  832. VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Networkaccepted
  833. VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Datasetaccepted
  834. VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulationaccepted
  835. VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimationaccepted
  836. VIRO: Robust and Efficient Neuro-Symbolic Reasoning with Verification for Referring Expression Comprehensionaccepted
  837. VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentationaccepted
  838. VISTA: A Test-Time Self-Improving Video Generation Agentaccepted
  839. VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactionsaccepted
  840. VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessmentaccepted
  841. VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimizationaccepted
  842. VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Modelsaccepted
  843. VL-Eraser: Vacuum Distillation for Machine Unlearning in Vision-Language Modelsaccepted
  844. VL-RouterBench: A Benchmark for Vision-Language Model Routingaccepted
  845. VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modelingaccepted
  846. VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compressionaccepted
  847. VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstructionaccepted
  848. VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recoveryaccepted
  849. VLM-Loc: Localization in Point Cloud Maps via Vision-Language Modelsaccepted
  850. VLM-PTQ: Efficient Post-Training Quantization for Large Vision-Language Modelsaccepted
  851. VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigmaccepted
  852. VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detectionaccepted
  853. VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformationaccepted
  854. VMonarch: Efficient Video Diffusion Transformers with Structured Attentionaccepted
  855. VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillationaccepted
  856. VOSR: A Vision-Only Generative Model for Image Super-Resolutionaccepted
  857. VQ-VA World: Towards High-Quality Visual Question-Visual Answeringaccepted
  858. VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstructionaccepted
  859. VRCLIP: Multimodal Canonical Correlation Alignment for CLIP-Driven Vision-Radio Person Re-Identificationaccepted
  860. VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cuesaccepted
  861. VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environmentsaccepted
  862. VSRELL: A Simple Baseline for Video Super-Resolution and Enhancement in Low-Light Environmentaccepted
  863. VT-Intrinsic: Physics-Based Decomposition of Reflectance and Shading using a Single Visible-Thermal Image Pairaccepted
  864. VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangementaccepted
  865. VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skippingaccepted
  866. V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondenceaccepted
  867. Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervisionaccepted
  868. VarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAMaccepted
  869. Variation-aware Vision Token Dropping for Faster Large Vision-Language Modelsaccepted
  870. Variational Graph-based Normal Integrationaccepted
  871. VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inferenceaccepted
  872. VecGlypher: Unified Vector Glyph Generation with Language Modelsaccepted
  873. Vector Prism: Animating Vector Graphics by Stratifying Semantic Structureaccepted
  874. VectorArk: Learning Practical Image Vectorization with Rounded Polygon Representationaccepted
  875. Velox: Learning Representations of 4D Geometry and Appearanceaccepted
  876. Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Croppingaccepted
  877. Verifying Neural Network Robustness with Dual Perturbationsaccepted
  878. VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Controlaccepted
  879. VesMamba: 3D Pulmonary Vessel Segmentation from CT images via Mamba with Structural Perception and Scale-aware Filteringaccepted
  880. ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Bodyaccepted
  881. ViHOI: Human-Object Interaction Synthesis with Visual Priorsaccepted
  882. ViKey: Enhancing Temporal Understanding in Videos via Visual Promptingaccepted
  883. ViLearn: Accelerating Training Convergence of Image-to-3D Generation via Visibility Learningaccepted
  884. ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memoryaccepted
  885. ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunkingaccepted
  886. ViStoryBench: Comprehensive Benchmark Suite for Story Visualizationaccepted
  887. ViT$^3$: Unlocking Test-Time Training in Visionaccepted
  888. ViTPrompt: Training-Free Prompt Refinement with Visual Tokens for Open-Vocabulary Detectionaccepted
  889. Vibe Spaces for Creatively Connecting and Expressing Visual Conceptsaccepted
  890. VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generationsaccepted
  891. VidEoMT: Your ViT is Secretly Also a Video Segmentation Modelaccepted
  892. VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transferaccepted
  893. VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scaleaccepted
  894. Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learneraccepted
  895. Video Panels for Long Video Understandingaccepted
  896. Video-CoE: Reinforcing Video Event Prediction via Chain of Eventsaccepted
  897. Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Modelsaccepted
  898. Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPOaccepted
  899. Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot Learningaccepted
  900. VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understandingaccepted
  901. VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twiceaccepted
  902. VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learningaccepted
  903. VideoCoF: Unified Video Editing with Temporal Reasoneraccepted
  904. VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusionaccepted
  905. VideoITG: Multimodal Video Understanding with Instructed Temporal Groundingaccepted
  906. VideoMaMa: Mask-Guided Video Matting via Generative Prioraccepted
  907. VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognitionaccepted
  908. VideoRealBench: A Chain-of-Thought Realism Evaluation Benchmark for Generated Human-Centric Videosaccepted
  909. VideoSSR: Video Self-Supervised Reinforcement Learningaccepted
  910. VideoSeek: Long-Horizon Video Agent with Tool-Guided Seekingaccepted
  911. VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agentsaccepted
  912. VideoWorld 2: Learning Transferable Knowledge from Real-world Videosaccepted
  913. View-Aware Semantic Alignment for Aerial-Ground Person Re-Identificationaccepted
  914. VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QAaccepted
  915. Vinedresser3D: Towards Agentic Text-guided 3D Editingaccepted
  916. Virtual Full-stack Scanning of Brain MRI via Imputing Any Quantised Codeaccepted
  917. Virtual Immunohistochemistry Staining with Dual-Aligned Multi-Task Feature Guidanceaccepted
  918. Virtual Nodes Guided Dynamic Graph Neural Network for Brain Tumor Segmentation with Missing Modalitiesaccepted
  919. VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understandingaccepted
  920. VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Modelsaccepted
  921. VisPlay: Self-Evolving Vision-Language Modelsaccepted
  922. VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Modelsaccepted
  923. VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsaccepted
  924. VisiLock: Authorizing Instruction-based Image editing with Dual Score Distillationaccepted
  925. Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Modelsaccepted
  926. Vision Transformers Need More Than Registersaccepted
  927. Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identificationaccepted
  928. Vision-Language Model Guided Source-Free Domain Adaptation via Optimal Transportaccepted
  929. Vision-Oriented Lightweight Neural Architecture Search with Budget-Adaptive Evaluationaccepted
  930. Vision-Speech Models: Teaching Speech Models to Converse about Imagesaccepted
  931. VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesisaccepted
  932. VisionLeaf: Entropy-Guided Leaf-First Reasoning for Efficient and Accurate Think-with-Imageaccepted
  933. Vista4D: Video Reshooting with 4D Point Cloudsaccepted
  934. Visual Diffusion Models are Geometric Solversaccepted
  935. Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scalingaccepted
  936. Visual Grounding for Object Questionsaccepted
  937. Visual Personalization Turing Testaccepted
  938. Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detectionaccepted
  939. Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Modelsaccepted
  940. Visual-RRT: Finding Paths toward Visual-Goals via Differentiable Renderingaccepted
  941. VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformeraccepted
  942. VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenesaccepted
  943. ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videosaccepted
  944. VoDaSuRe: A Large-Scale Dataset Revealing Domain Shift in Volumetric Super-Resolutionaccepted
  945. Vocabulary Scaling Law: Tuning Open-vocabulary Predictors for Their Opennessaccepted
  946. Volumetric Functional Mapsaccepted
  947. VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentationaccepted
  948. Voxify3D: Pixel Art Meets Volumetric Renderingaccepted
  949. W2W: Language-Model-Based Trajectory Prediction with Reinforcement Learningaccepted
  950. WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Drivingaccepted
  951. WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generationaccepted
  952. WHU-MARS: A Multispectral Aerial-Ground Benchmark Towards Any-Scenario Person Re-Identificationaccepted
  953. WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrievalaccepted
  954. WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenariosaccepted
  955. WPT: World-to-Policy Transfer via Online World Model Distillationaccepted
  956. WRIVINDER: Towards Spatial Intelligence for Geo-locating Ground Images onto Satellite Imageryaccepted
  957. WaDi: Weight Direction-aware Distillation for One-step Image Synthesisaccepted
  958. WaTeRFlow: Watermark Temporal Robustness via Flow Consistencyaccepted
  959. WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigationaccepted
  960. Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Trainingaccepted
  961. Wanderland: Geometrically Grounded Simulation for Open-World Embodied AIaccepted
  962. Watch and Learn: Learning to Use Computers from Online Videosaccepted
  963. Wave-Former: Through-Occlusion 3D Reconstruction via Wireless Shape Completionaccepted
  964. Wavelet-Driven 3D Anomaly Detection under Pose-Agnostic and Sparse-Viewaccepted
  965. Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understandingaccepted
  966. WeDetect: Fast Open-Vocabulary Object Detection as Retrievalaccepted
  967. WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokensaccepted
  968. Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoningaccepted
  969. WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformationaccepted
  970. WeaveTime: Streaming from Earlier Frames into Emergent Memory in VideoLLMsaccepted
  971. WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Tracesaccepted
  972. WebGym: Scaling Training Environments for Long-Horizon Visual Web Agents with Realistic Tasksaccepted
  973. Weight Space Representation Learning via Neural Field Adaptationaccepted
  974. What Are You Doing? A Closer Look at Controllable Human Video Generationaccepted
  975. What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Modelsaccepted
  976. What Is It Like to Be a Noise? An Entropy-based Gaussian Noise Regularization for Diffusion Modelsaccepted
  977. What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is Rarely F1accepted
  978. What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?accepted
  979. What Matters in Practical Learned Image Compressionaccepted
  980. What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNsaccepted
  981. What's Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolutionaccepted
  982. When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapseaccepted
  983. When Anonymity Breaks: Identifying Models Behind Text-to-Image Leaderboardsaccepted
  984. When CLIP Sees More, It Fights Back Harder: Multi-View Guided Adaptive Counterattacks for Test-Time Adversarial Robustnessaccepted
  985. When Do Models Actually Decide? Mapping the Layer-Wise Decision Timeline in Pretrained Neural Networksaccepted
  986. When Lines Meet Textures: Spatial-Frequency Aligned Diffusion Features for Cross-Sparsity Correspondenceaccepted
  987. When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adaptersaccepted
  988. When Local Rules Create Global Order: Self-Organized Representation Learning for Latent Diffusion Modelsaccepted
  989. When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Modelsaccepted
  990. When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generatorsaccepted
  991. When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Modelsaccepted
  992. When Robots Should Say ''I Don't Know'': Benchmarking Abstention in Embodied Question Answeringaccepted
  993. When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety Guidanceaccepted
  994. When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMsaccepted
  995. When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detectionaccepted
  996. When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigmaccepted
  997. When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thoughtaccepted
  998. When to Think and When to Look: Uncertainty-Guided Lookbackaccepted
  999. Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generationaccepted
  1000. Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attentionaccepted

Looking for submission deadlines instead? See the conference deadline calendar.