← All conferences

CVPR 2025 Accepted Papers

The full list of 2,872 papers accepted at CVPR 2025 (IEEE/CVF Conference on Computer Vision and Pattern Recognition). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

Poster: 2,468Highlight: 388Award Candidate: 15
  1. Latent Space ImagingPoster
  2. LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand Diffusion.Poster
  3. Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric VideosPoster
  4. Learnable Infinite Taylor Gaussian for Dynamic View RenderingPoster
  5. Learned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus CuesPoster
  6. Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual SceneHighlight
  7. Learning Audio-guided Video Representation with Gated Attention for Video-Text RetrievalPoster
  8. Learning Class Prototypes for Unified Sparse-Supervised 3D Object DetectionHighlight
  9. Learning Compatible Multi-Prize Subnetworks for Asymmetric RetrievalPoster
  10. Learning Conditional Space-Time Prompt Distributions for Video Class-Incremental LearningHighlight
  11. Learning Dynamic Collaborative Network for Semi-supervised 3D Vessel SegmentationPoster
  12. Learning Endogenous Attention for Incremental Object DetectionPoster
  13. Learning Hazing to Dehazing: Towards Realistic Haze Generation for Real-World Image DehazingPoster
  14. Learning Heterogeneous Tissues with Mixture of Experts for Gigapixel Whole Slide ImagesPoster
  15. Learning Occlusion-Robust Vision Transformers for Real-Time UAV TrackingPoster
  16. Learning Partonomic 3D Reconstruction from Image CollectionsPoster
  17. Learning Person-Specific Animatable Face Models from In-the-Wild Images via a Shared Base ModelPoster
  18. Learning Phase Distortion with Selective State Space Models for Video Turbulence MitigationHighlight
  19. Learning Physics From Video: Unsupervised Physical Parameter Estimation for Continuous Dynamical SystemsPoster
  20. Learning Physics-Based Full-Body Human Reaching and Grasping from Brief Walking ReferencesPoster
  21. Learning Textual Prompts for Open-World Semi-Supervised LearningPoster
  22. Learning Visual Composition through Improved Semantic GuidancePoster
  23. Learning from Neighbors: Category Extrapolation for Long-Tail LearningPoster
  24. Learning from Streaming Video with Orthogonal GradientsPoster
  25. Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging ScenesPoster
  26. Learning on Model Weights using Tree ExpertsPoster
  27. Learning to Detect Objects from Multi-Agent LiDAR Scans without Manual LabelsPoster
  28. Learning to Filter Outlier Edges in Global SfMHighlight
  29. Learning to Highlight Audio by Watching MoviesPoster
  30. Learning to Normalize on the SPD Manifold under Bures-Wasserstein GeometryPoster
  31. Learning with Noisy Triplet Correspondence for Composed Image RetrievalPoster
  32. Learning-enabled Polynomial Lyapunov Function Synthesis via High-Accuracy Counterexample-Guided FrameworkPoster
  33. LesionLocator: Zero-Shot Universal Tumor Segmentation and Tracking in 3D Whole-Body ImagingPoster
  34. Less Attention is More: Prompt Transformer for Generalized Category DiscoveryPoster
  35. Less is More: Efficient Image Vectorization with Adaptive ParameterizationPoster
  36. Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual RecognitionHighlight
  37. Let Humanoids Hike! Integrative Skill Development on Complex TrailsPoster
  38. Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of SamplesPoster
  39. Let's Chorus: Partner-aware Hybrid Song-Driven 3D Head AnimationPoster
  40. Let's Verify and Reinforce Image Generation Step by StepPoster
  41. Leveraging 3D Geometric Priors in 2D Rotation Symmetry DetectionPoster
  42. Leveraging Global Stereo Consistency for Category-Level Shape and 6D Pose Estimation from Stereo ImagesPoster
  43. Leveraging Perturbation Robustness to Enhance Out-of-Distribution DetectionPoster
  44. Leveraging SD Map to Augment HD Map-based Trajectory PredictionPoster
  45. Leveraging Temporal Cues for Semi-Supervised Multi-View 3D Object DetectionPoster
  46. LiSu: A Dataset and Method for LiDAR Surface Normal EstimationPoster
  47. Libra-Merging: Importance-redundancy and Pruning-merging Trade-off for Acceleration Plug-in in Large Vision-Language ModelPoster
  48. LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer AttributionsPoster
  49. LidarGait++: Learning Local Features and Size Awareness from LiDAR Point Clouds for 3D Gait RecognitionPoster
  50. Lifelong Knowledge Editing for Vision Language Models with Low-Rank Mixture-of-ExpertsPoster
  51. Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic ManipulationPoster
  52. Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster InferencePoster
  53. Light Transport-aware Diffusion Posterior Sampling for Single-View Reconstruction of 3D VolumesHighlight
  54. LightLoc: Learning Outdoor LiDAR Localization at Light SpeedPoster
  55. Linear Attention Modeling for Learned Image CompressionPoster
  56. Link to the Past: Temporal Propagation for Fast 3D Human Reconstruction from Monocular VideoPoster
  57. Link-based Contrastive Learning for One-Shot Unsupervised Domain AdaptationPoster
  58. LiveCC: Learning Video LLM with Streaming Speech Transcription at ScalePoster
  59. LoKi: Low-dimensional KAN for Efficient Fine-tuning Image ModelsPoster
  60. LoRA Recycle: Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAsPoster
  61. LoRACLR: Contrastive Adaptation for Customization of Diffusion ModelsPoster
  62. LoTUS: Large-Scale Machine Unlearning with a Taste of UncertaintyPoster
  63. Locality-Aware Zero-Shot Human-Object Interaction DetectionPoster
  64. Locally Orderless Images for Optimization in Differentiable RenderingHighlight
  65. Logits DeConfusion with CLIP for Few-Shot LearningPoster
  66. LogoSP: Local-global Grouping of Superpoints for Unsupervised Semantic Segmentation of 3D Point CloudsPoster
  67. LongDiff: Training-Free Long Video Generation in One GoPoster
  68. LookingGlass: Generative Anamorphoses via Laplacian Pyramid WarpingPoster
  69. LotusFilter: Fast Diverse Nearest Neighbor Search via a Learned Cutoff TablePoster
  70. Low-Biased General Annotated Dataset GenerationPoster
  71. Low-Rank Adaptation in Multilinear Operator Networks for Security-Preserving Incremental LearningPoster
  72. Luminance-GS: Adapting 3D Gaussian Splatting to Challenging Lighting Conditions with View-Adaptive Curve AdjustmentPoster
  73. M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world SettingsPoster
  74. M3amba: Memory Mamba is All You Need for Whole Slide Image ClassificationPoster
  75. MAC-Ego3D: Multi-Agent Gaussian Consensus for Real-Time Collaborative Ego-Motion and Photorealistic 3D ReconstructionPoster
  76. MAD: Memory-Augmented Detection of 3D ObjectsPoster
  77. MAGE : Single Image to Material-Aware 3D via the Multi-View G-Buffer Estimation ModelPoster
  78. MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action AnticipationPoster
  79. MARBLE: Material Recomposition and Blending in CLIP-SpacePoster
  80. MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content CreationPoster
  81. MATCHA: Towards Matching AnythingHighlight
  82. MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse ViewsHighlight
  83. MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image GenerationPoster
  84. MDP: Multidimensional Vision Model Pruning with Latency ConstraintPoster
  85. MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh AttentionPoster
  86. MEET: Towards Memory-Efficient Temporal Sparse Deep Neural NetworksPoster
  87. MERGE: Multi-faceted Hierarchical Graph-based GNN for Gene Expression Prediction from Whole Slide Histopathology ImagesPoster
  88. MESC-3D:Mining Effective Semantic Cues for 3D Reconstruction from a Single ImagePoster
  89. METASCENES: Towards Automated Replica Creation for Real-world 3D ScansPoster
  90. MExD: An Expert-Infused Diffusion Model for Whole-Slide Image ClassificationPoster
  91. MFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and ForecastingPoster
  92. MI-DETR: An Object Detection Model with Multi-time Inquiries MechanismPoster
  93. MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal OutputPoster
  94. MIRE: Matched Implicit Neural RepresentationsPoster
  95. MITracker: Multi-View Integration for Visual Object TrackingHighlight
  96. MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio SynthesisPoster
  97. MNE-SLAM: Multi-Agent Neural SLAM for Mobile RobotsPoster
  98. MODA: Motion-Drift Augmentation for Inertial Human Motion AnalysisPoster
  99. MOS-Attack: A Scalable Multi-objective Adversarial Attack FrameworkPoster
  100. MOS: Modeling Object-Scene Associations in Generalized Category DiscoveryPoster
  101. MP-SfM: Monocular Surface Priors for Robust Structure-from-MotionPoster
  102. MTADiffusion: Mask Text Alignment Diffusion Model for Object InpaintingPoster
  103. MUST: The First Dataset and Unified Framework for Multispectral UAV Single Object TrackingPoster
  104. MV-SSM: Multi-View State Space Modeling for 3D Human Pose EstimationPoster
  105. MVDoppler-Pose: Multi-Modal Multi-View mmWave Sensing for Long-Distance Self-Occluded Human Walking Pose EstimationPoster
  106. MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion ModelPoster
  107. M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object SegmentationPoster
  108. MaDCoW: Marginal Distortion Correction for Wide-Angle Photography with Arbitrary ObjectsPoster
  109. MaSS13K: A Matting-level Semantic Segmentation BenchmarkPoster
  110. MagicArticulate: Make Your 3D Models Articulation-ReadyPoster
  111. Making Old Film Great Again: Degradation-aware State Space Model for Old Film RestorationPoster
  112. Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic SegmentationHighlight
  113. Mamba-Adaptor: State Space Model Adaptor for Visual RecognitionPoster
  114. Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space ModelsPoster
  115. MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training SmoothingPoster
  116. MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss AlpsHighlight
  117. MarkushGrapher: Joint Visual and Textual Recognition of Markush StructuresPoster
  118. Mask-Adapter: The Devil is in the Masks for Open-Vocabulary SegmentationPoster
  119. Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene UnderstandingPoster
  120. Masking meets Supervision: A Strong Learning AlliancePoster
  121. Matrix-Free Shared Intrinsics Bundle AdjustmentPoster
  122. Medusa: A Multi-Scale High-order Contrastive Dual-Diffusion Approach for Multi-View ClusteringPoster
  123. Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured MeshesPoster
  124. MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationHighlight
  125. Meta-Learning Hyperparameters for Parameter Efficient Fine-TuningHighlight
  126. MetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt TuningPoster
  127. MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural RepresentationHighlight
  128. Mimic In-Context Learning for Multimodal TasksPoster
  129. Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-MismatchPoster
  130. Mind the Gap: Detecting Black-box Adversarial Attacks in the Making through Query Update AnalysisPoster
  131. Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive JailbreakingHighlight
  132. Minding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion SegmentationPoster
  133. Minimal Interaction Seperated Tuning: A New Paradigm for Visual AdaptationPoster
  134. Minimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance SegmentationPoster
  135. Minority-Focused Text-to-Image Generation via Prompt OptimizationPoster
  136. MirrorVerse: Pushing Diffusion Models to Realistically Reflect the WorldPoster
  137. Mitigating Ambiguities in 3D Classification with Gaussian SplattingPoster
  138. Mixture of Submodules for Domain Adaptive Person SearchPoster
  139. MoEdit: On Learning Quantity Perception for Multi-object Image EditingPoster
  140. MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based DistillationPoster
  141. MoST: Efficient Monarch Sparse Tuning for 3D Representation LearningPoster
  142. MobileH2R: Learning Generalizable Human to Mobile Robot Handover Exclusively from Scalable and Diverse Synthetic DataPoster
  143. ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode ModelingPoster
  144. Model Diagnosis and Correction via Linguistic and Implicit Attribute EditingPoster
  145. Modeling Multiple Normal Action Representations for Error Detection in Procedural TasksPoster
  146. Mono2Stereo: A Benchmark and Empirical Study for Stereo ConversionPoster
  147. Mono3DVLT: Monocular-Video-Based 3D Visual Language TrackingPoster
  148. MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular DetectionPoster
  149. MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation ModelsPoster
  150. Monocular and Generalizable Gaussian Talking Head AnimationPoster
  151. Morpheus: Text-Driven 3D Gaussian Splat Shape and Color StylizationPoster
  152. Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph LearningPoster
  153. MotiF: Making Text Count in Image Animation with Motion Focal LossPoster
  154. MotionMap: Representing Multimodality in Human Pose ForecastingPoster
  155. MotionPRO: Exploring the Role of Pressure in Human MoCap and BeyondHighlight
  156. MotionPro: A Precise Motion Controller for Image-to-Video GenerationPoster
  157. Motions as Queries: One-Stage Multi-Person Holistic Human Motion CapturePoster
  158. Move-in-2D: 2D-Conditioned Human Motion GenerationPoster
  159. MuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image TranslationPoster
  160. Multi-Group Proportional Representations for Text-to-Image ModelsPoster
  161. Multi-Label Prototype Visual Spatial Search for Weakly Supervised Semantic SegmentationHighlight
  162. Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best PracticesPoster
  163. Multi-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEsPoster
  164. Multi-Modal Contrastive Masked Autoencoders: A Two-Stage Progressive Pre-training Approach for RGBD DatasetsPoster
  165. Multi-Modal Synergistic Implicit Image Enhancement for Efficient Optical Flow EstimationPoster
  166. Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual RepresentationPoster
  167. Multi-Scale Neighborhood Occupancy Masked Autoencoder for Self-Supervised Learning in LiDAR Point CloudsPoster
  168. Multi-View Pose-Agnostic Change Localization with Zero LabelsPoster
  169. Multi-focal Conditioned Latent Diffusion for Person Image SynthesisPoster
  170. Multi-modal Contrastive Learning with Negative Sampling Calibration for Phenotypic Drug DiscoveryPoster
  171. Multi-modal Knowledge Distillation-based Human Trajectory ForecastingPoster
  172. Multi-modal Medical Diagnosis via Large-small Model CollaborationPoster
  173. Multi-modal Topology-embedded Graph Learning for Spatially Resolved Genes Prediction from Pathology Images with Prior Gene Similarity InformationPoster
  174. Multi-modal Vision Pre-training for Medical Image AnalysisHighlight
  175. Multi-subject Open-set Personalization in Video GenerationPoster
  176. MultiMorph: On-demand Atlas ConstructionPoster
  177. MultimodalStudio: A Heterogeneous Sensor Dataset and Framework for Neural Rendering across Multiple Imaging ModalitiesPoster
  178. Multirate Neural Image Compression with Adaptive Lattice Vector QuantizationHighlight
  179. NADER: Neural Architecture Design via Multi-Agent CollaborationPoster
  180. NLPrompt: Noise-Label Prompt Learning for Vision-Language ModelsHighlight
  181. NN-Former: Rethinking Graph Structure in Neural Architecture RepresentationPoster
  182. NSD-Imagery: A Benchmark Dataset for Extending fMRI Vision Decoding Methods to Mental ImageryHighlight
  183. NTClick: Achieving Precise Interactive Segmentation With Noise-tolerant ClicksHighlight
  184. Navigating Image Restoration with VAR's Distribution Alignment PriorPoster
  185. Navigating the Unseen: Zero-shot Scene Graph Generation via Capsule-Based Equivariant FeaturesPoster
  186. NeISF++: Neural Incident Stokes Field for Polarized Inverse Rendering of Conductors and DielectricsPoster
  187. Nearly Zero-Cost Protection Against Mimicry by Personalized Diffusion ModelsPoster
  188. NeighborRetr: Balancing Hub Centrality in Cross-Modal RetrievalPoster
  189. Neural Hierarchical Decomposition for Single Image Plant ModelingPoster
  190. Neural Inverse Rendering from Propagating LightPoster
  191. Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light DiffusionPoster
  192. Neural Motion Simulator Pushing the Limit of World Models in Reinforcement LearningPoster
  193. Neural Video Compression with Context ModulationPoster
  194. NexusGS: Sparse View Synthesis with Epipolar Depth Priors in 3D Gaussian SplattingHighlight
  195. NightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene SegmentationPoster
  196. No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image RecognitionHighlight
  197. No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse WeatherPoster
  198. Noise Calibration and Spatial-Frequency Interactive Network for STEM Image EnhancementPoster
  199. Noise Diffusion for Enhancing Semantic Faithfulness in Text-to-Image SynthesisPoster
  200. Noise Modeling in One Hour: Minimizing Preparation Efforts for Self-supervised Low-Light RAW Image DenoisingPoster
  201. Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and SegmentationPoster
  202. Noise-Resistant Video Anomaly Detection via RGB Error-Guided Multiscale Predictive Coding and Dynamic MemoryPoster
  203. NoiseCtrl: A Sampling-Algorithm-Agnostic Conditional Generation Method for Diffusion ModelsPoster
  204. Non-Natural Image Understanding with Advancing Frequency-based Vision EncodersPoster
  205. Nonisotropic Gaussian Diffusion for Realistic 3D Human Motion PredictionPoster
  206. Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation AbilityPoster
  207. Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation ModelsPoster
  208. Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language ModelsHighlight
  209. Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question AnsweringPoster
  210. O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language ModelsHighlight
  211. OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation TriadPoster
  212. ODA-GAN: Orthogonal Decoupling Alignment GAN Assisted by Weakly-supervised Learning for Virtual Immunohistochemistry StainingPoster
  213. ODHSR: Online Dense 3D Reconstruction of Humans and Scenes from Monocular VideosPoster
  214. OFER: Occluded Face Expression ReconstructionPoster
  215. ONDA-Pose: Occlusion-Aware Neural Domain Adaptation for Self-Supervised 6D Object Pose EstimationPoster
  216. OODD: Test-time Out-of-Distribution Detection with Dynamic DictionaryPoster
  217. OPTICAL: Leveraging Optimal Transport for Contribution Allocation in Dataset DistillationHighlight
  218. ORIDa: Object-centric Real-world Image Composition DatasetPoster
  219. OW-OVD: Unified Open World and Open Vocabulary Object DetectionPoster
  220. Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic ManipulationPoster
  221. Object-aware Sound Source Localization via Audio-Visual Scene UnderstandingPoster
  222. Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object RecognitionPoster
  223. Odd-One-Out: Anomaly Detection by Comparing with NeighborsPoster
  224. OffsetOPT: Explicit Surface Reconstruction without NormalsPoster
  225. OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video ContextsPoster
  226. OmniSplat: Taming Feed-Forward 3D Gaussian Splatting for Omnidirectional Images with Editable CapabilitiesHighlight
  227. OmniStereo: Real-time Omnidireactional Depth Estimation with Multiview Fisheye CamerasPoster
  228. OmniStyle: Filtering High Quality Style Transfer Data at ScalePoster
  229. Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric VideosPoster
  230. Omnidirectional Multi-Object TrackingPoster
  231. On Denoising Walking Videos for Gait RecognitionPoster
  232. On the Out-Of-Distribution Generalization of Large Multimodal ModelsPoster
  233. On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free ApproachPoster
  234. On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only EventsPoster
  235. Once-Tuning-Multiple-Variants: Tuning Once and Expanded as Multiple Vision-Language Model VariantsPoster
  236. One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image FusionPoster
  237. One is Plenty: A Polymorphic Feature Interpreter for Immutable Heterogeneous Collaborative PerceptionPoster
  238. One-Step Event-Driven High-Speed AutofocusHighlight
  239. One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion ModelsPoster
  240. One-shot 3D Object Canonicalization based on Geometric and Semantic ConsistencyHighlight
  241. One2Any: One-Reference 6D Pose Estimation for Any ObjectPoster
  242. Online Task-Free Continual Learning via Dynamic Expansionable Memory DistributionPoster
  243. Online Video Understanding: OVBench and VideoChat-OnlinePoster
  244. OnlineAnySeg: Online Zero-Shot 3D Segmentation by Visual Foundation Model Guided 2D Mask MergingPoster
  245. Open Ad-hoc Categorization with Contextualized Feature LearningPoster
  246. Open Set Label Shift with Test Time Out-of-Distribution ReferencePoster
  247. Open-Canopy: Towards Very High Resolution Forest MonitoringHighlight
  248. OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text GenerationPoster
  249. OpenMIBOOD: Open Medical Imaging Benchmarks for Out-Of-Distribution DetectionPoster
  250. OpenSDI: Spotting Diffusion-Generated Images in the Open WorldPoster
  251. Opportunistic Single-Photon Time of FlightPoster
  252. OpticalNet: An Optical Imaging Dataset and Benchmark Beyond the Diffraction LimitHighlight
  253. Optimal Transport-Guided Source-Free Adaptation for Face Anti-SpoofingPoster
  254. Optimizing for the Shortest Path in Denoising Diffusion ModelHighlight
  255. Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned PolicyPoster
  256. OralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware MiningPoster
  257. OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic KernelsPoster
  258. Overcoming Shortcut Problem in VLM for Robust Out-of-Distribution DetectionHighlight
  259. PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language ModelsPoster
  260. PAVE: Patching and Adapting Video Large Language ModelsPoster
  261. PBR-NeRF: Inverse Rendering with Physics-Based Neural FieldsPoster
  262. PCM : Picard Consistency Model for Fast Parallel Sampling of Diffusion ModelsPoster
  263. PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic ManipulationPoster
  264. PEER Pressure: Model-to-Model Regularization for Single Source Domain GeneralizationPoster
  265. PGC: Physics-Based Gaussian Cloth from a Single PoseHighlight
  266. PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric VideosPoster
  267. PI-HMR: Towards Robust In-bed Temporal Human Shape Reconstruction with Contact Pressure SensingPoster
  268. PIAD: Pose and Illumination agnostic Anomaly DetectionPoster
  269. PICD: Versatile Perceptual Image Compression with Diffusion RenderingPoster
  270. PIDLoc: Cross-View Pose Optimization Network Inspired by PID ControllersPoster
  271. PIDSR: Complementary Polarized Image Demosaicing and Super-ResolutionPoster
  272. PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba AdapterPoster
  273. PMNI: Pose-free Multi-view Normal Integration for Reflective and Textureless Surface ReconstructionPoster
  274. POMP: Physics-consistent Motion Generative Model through Phase ManifoldsPoster
  275. POT: Prototypical Optimal Transport for Weakly Supervised Semantic SegmentationPoster
  276. POp-GS: Next Best View in 3D-Gaussian Splatting with P-OptimalityPoster
  277. PRaDA: Projective Radial Distortion AveragingPoster
  278. PS-Diffusion: Photorealistic Subject-Driven Image Editing with Disentangled Control and AttentionPoster
  279. PS-EIP: Robust Photometric Stereo Based on Event Interval ProfilePoster
  280. PSA-SSL: Pose and Size-aware Self-Supervised Learning on LiDAR Point CloudsPoster
  281. PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit RemeshingPoster
  282. PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion ModelPoster
  283. PURA: Parameter Update-Recovery Test-Time Adaption for RGB-T TrackingPoster
  284. PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language ModelsPoster
  285. PanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial AugmentationPoster
  286. PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene UnderstandingPoster
  287. Parallel Sequence Modeling via Generalized Spatial Propagation NetworkPoster
  288. Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear TransformationPoster
  289. Parameterized Blur Kernel Prior Learning for Local Motion DeblurringPoster
  290. Parametric Point Cloud Completion for Polygonal Surface ReconstructionPoster
  291. PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction ModelPoster
  292. PatchDEMUX: A Certifiably Robust Framework for Multi-label Classifiers Against Adversarial PatchesPoster
  293. PatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo AnomaliesPoster
  294. PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-ResolutionPoster
  295. Pattern Analogies: Learning to Perform Programmatic Image Edits by AnalogyPoster
  296. Percept, Memory, and Imagine: World Feature Simulating for Open-Domain Unknown Object DetectionPoster
  297. Perceptual Inductive Bias Is What You Need Before Contrastive LearningPoster
  298. Perceptual Video Compression with Neural WrappingPoster
  299. Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation MetricsHighlight
  300. Period-LLM: Extending the Periodic Capability of Multimodal Large Language ModelPoster
  301. Person De-reidentification: A Variation-guided Identity Shift ModelingPoster
  302. PersonaHOI: Effortlessly Improving Face Personalization in Human-Object Interaction GenerationPoster
  303. Perturb-and-Revise: Flexible 3D Editing with Generative TrajectoriesPoster
  304. Phoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action CorrectionPoster
  305. PhyS-EdiT: Physics-aware Semantic Image Editing with Text DescriptionPoster
  306. Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-TuningPoster
  307. Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot VisionPoster
  308. PlanarSplatting: Accurate Planar Surface Reconstruction in 3 MinutesHighlight
  309. Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept ControlPoster
  310. Plug-and-Play PPO: An Adaptive Point Prompt Optimizer Making SAM GreaterPoster
  311. Plug-and-Play Versatile Compressed Video EnhancementPoster
  312. Point Cloud Upsampling Using Conditional Diffusion Module with Adaptive Noise SuppressionPoster
  313. Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud AnalysisPoster
  314. Point-to-Region Loss for Semi-Supervised Point-Based Crowd CountingHighlight
  315. PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud LearningPoster
  316. PointSR: Self-Regularized Point Supervision for Drone-View Object DetectionPoster
  317. PolarFree: Polarization-based Reflection-Free ImagingPoster
  318. PolarNeXt: Rethink Instance Segmentation with Polar RepresentationPoster
  319. Polarized Color Screen MattingHighlight
  320. Poly-Autoregressive Prediction for Modeling InteractionsPoster
  321. Population Normalization for Federated LearningPoster
  322. Pos3R: 6D Pose Estimation for Unseen Objects Made EasyPoster
  323. Pose-Guided Temporal Enhancement for Robust Low-Resolution Hand ReconstructionPoster
  324. PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose EstimationPoster
  325. PoseTraj: Pose-Aware Trajectory Control in Video DiffusionPoster
  326. PosterO: Structuring Layout Trees to Enable Language Models in Generalized Content-Aware Layout GenerationPoster
  327. Potential Field Based Deep Metric LearningPoster
  328. Practical Solutions to the Relative Pose of Three Calibrated CamerasPoster
  329. Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class ImbalancePoster
  330. PreciseCam: Precise Camera Control for Text-to-Image GenerationPoster
  331. Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image EditingPoster
  332. Preserving Clusters in Prompt Learning for Unsupervised Domain AdaptationPoster
  333. Prior-free 3D Object TrackingHighlight
  334. ProHOC: Probabilistic Hierarchical Out-of-Distribution Classification via Multi-Depth NetworksPoster
  335. Probabilistic Prompt Distribution Learning for Animal Pose EstimationPoster
  336. ProbeSDF: Light Field Probes For Neural Surface ReconstructionPoster
  337. Probing the Mid-level Vision Capabilities of Self-Supervised LearningPoster
  338. Prof. Robot: Differentiable Robot Rendering Without Static and Self-CollisionsPoster
  339. Progress-Aware Video Frame CaptioningPoster
  340. Progressive Correspondence Regenerator for Robust 3D RegistrationPoster
  341. Progressive Focused Transformer for Single Image Super-ResolutionPoster
  342. Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D DataPoster
  343. ProjAttacker: A Configurable Physical Adversarial Attack for Face Recognition via ProjectorPoster
  344. Project-Probe-Aggregate: Efficient Fine-Tuning for Group RobustnessHighlight
  345. Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained AnalysisPoster
  346. Prompt2Perturb (P2P): Text-Guided Diffusion-Based Adversarial Attack on Breast Ultrasound ImagesPoster
  347. PromptHMR: Promptable Human Mesh RecoveryPoster
  348. Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context LearningPoster
  349. Proximal Algorithm Unrolling: Flexible and Efficient Reconstruction Networks for Single-Pixel ImagingPoster
  350. ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual GroundingPoster
  351. Pseudo Visible Feature Fine-Grained Fusion for Thermal Object DetectionPoster
  352. Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose InteractionPoster
  353. PyTorchGeoNodes: Enabling Differentiable Shape Programs for 3D Shape ReconstructionPoster
  354. Q-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction RegressionPoster
  355. QuCOOP: A Versatile Framework for Solving Composite and Binary-Parametrised Problems on Quantum AnnealersHighlight
  356. Quad-Pixel Image Defocus Deblurring: A New Benchmark and ModelPoster
  357. Quaffure: Real-Time Quasi-Static Neural Hair SimulationPoster
  358. QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the EdgePoster
  359. Query Efficient Black-Box Visual Prompting with Subspace LearningPoster
  360. Question-Aware Gaussian Experts for Audio-Visual Question AnsweringHighlight
  361. R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual LocalizationPoster
  362. R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt TuningPoster
  363. R2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action PlannerPoster
  364. RAEncoder: A Label-Free Reversible Adversarial Examples Encoder for Dataset Intellectual Property ProtectionPoster
  365. RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-EmbeddingsPoster
  366. RASP: Revisiting 3D Anamorphic Art for Shadow-Guided Packing of Irregular ObjectsPoster
  367. RC-AutoCalib: An End-to-End Radar-Camera Automatic Calibration NetworkPoster
  368. RCP-Bench: Benchmarking Robustness for Collaborative Perception Under Diverse CorruptionsPoster
  369. RDD: Robust Feature Detector and Descriptor using Deformable TransformerPoster
  370. REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity ConditioningPoster
  371. RGBAvatar: Reduced Gaussian Blendshapes for Online Modeling of Head AvatarsHighlight
  372. RICCARDO: Radar Hit Prediction and Convolution for Camera-Radar 3D Object DetectionPoster
  373. RL-RC-DoT: A Block-level RL agent for Task-Aware Video CompressionPoster
  374. ROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language ModelsPoster
  375. ROLL: Robust Noisy Pseudo-label Learning for Multi-View Clustering with Noisy CorrespondenceHighlight
  376. RORem: Training a Robust Object Remover with Human-in-the-LoopPoster
  377. RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera FusionPoster
  378. RaSS: Improving Denoising Diffusion Samplers with Reinforced Active Sampling SchedulerPoster
  379. Radio Frequency Ray Tracing with Neural Object Representation for Enhanced RF ModelingPoster
  380. Random Conditioning for Diffusion Model Compression with Distillation
  381. Random Conditioning with Distillation for Data-Efficient Diffusion Model CompressionPoster
  382. Rate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty EstimationPoster
  383. Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion ModelPoster
  384. ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence LearningPoster
  385. ReDiffDet: Rotation-equivariant Diffusion Model for Oriented Object DetectionPoster
  386. RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal ReperformancePoster
  387. ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data StreamsPoster
  388. ReWind: Understanding Long Videos with Instructed Learnable MemoryPoster
  389. Real-IAD D3: A Real-World 2D/Pseudo-3D/3D Dataset for Industrial Anomaly DetectionPoster
  390. Real-time High-fidelity Gaussian Human Avatars with Position-based Interpolation of Spatially Distributed MLPsHighlight
  391. Realistic Test-Time Adaptation of Vision-Language ModelsHighlight
  392. Reanimating Images using Neural Representations of Dynamic StimuliPoster
  393. ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and ReasoningPoster
  394. Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance GroundingPoster
  395. Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems ApproachHighlight
  396. Recognition-Synergistic Scene Text EditingPoster
  397. Reconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in DualPoster
  398. Reconstructing Animals and the WildPoster
  399. Reconstructing Close Human Interaction with Appearance and Proxemics ReasoningPoster
  400. Reconstructing Humans with a Biomechanically Accurate SkeletonPoster
  401. Reconstructing In-the-Wild Open-Vocabulary Human-Object InteractionsPoster
  402. Reconstructing People, Places, and CamerasHighlight
  403. Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal TransportPoster
  404. Rectification-specific Supervision and Constrained Estimator for Online Stereo RectificationPoster
  405. Recurrent Feature Mining and Keypoint Mixup Padding for Category-Agnostic Pose EstimationPoster
  406. Reducing Class-wise Confusion for Incremental Learning with Disentangled ManifoldsPoster
  407. RefPose: Leveraging Reference Geometric Correspondences for Accurate 6D Pose Estimation of Unseen ObjectsPoster
  408. Relation-Rich Visual Document Generator for Visual Information ExtractionPoster
  409. Relation3D : Enhancing Relation Modeling for Point Cloud Instance SegmentationPoster
  410. RelationField: Relate Anything in Radiance FieldsPoster
  411. Remote Photoplethysmography in Real-World and Extreme Lighting ScenariosPoster
  412. Reproducible Vision-Language Models Meet Concepts Out of Pre-TrainingPoster
  413. ResCLIP: Residual Attention for Training-free Dense Vision-language InferencePoster
  414. Resilient Sensor Fusion Under Adverse Sensor Failures via Multi-Modal Expert FusionPoster
  415. RestorGS: Depth-aware Gaussian Splatting for Efficient 3D Scene RestorationPoster
  416. Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher DistillationPoster
  417. Rethinking Decoder Design: Improving Biomarker Segmentation Using Depth-to-Space Restoration and Residual Linear AttentionPoster
  418. Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked AutoregressionPoster
  419. Rethinking Epistemic and Aleatoric Uncertainty for Active Open-Set Annotation: An Energy-Based ApproachPoster
  420. Rethinking Few-Shot Adaptation of Vision-Language Models in Two StagesPoster
  421. Rethinking Lanes and Points in Complex Scenarios for Monocular 3D Lane DetectionPoster
  422. Rethinking Noisy Video-Text Retrieval via Relation-aware AlignmentPoster
  423. Rethinking Personalized Aesthetics Assessment: Employing Physique Aesthetics Assessment as An ExemplificationHighlight
  424. Rethinking Query-based Transformer for Continual Image SegmentationPoster
  425. Rethinking Reconstruction and Denoising in the Dark: New Perspective, General Architecture and BeyondPoster
  426. Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking TransformersPoster
  427. Rethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level TasksPoster
  428. Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense GamePoster
  429. Reversing Flow for Image RestorationPoster
  430. Revisiting Audio-Visual Segmentation with Vision-Centric TransformerPoster
  431. Revisiting Backdoor Attacks against Large Vision-Language Models from Domain ShiftPoster
  432. Revisiting Fairness in Multitask Learning: A Performance-Driven Approach for Variance ReductionPoster
  433. Revisiting Generative Replay for Class Incremental Object DetectionPoster
  434. Revisiting Source-Free Domain Adaptation: Insights into Representativeness, Generalization, and VarietyPoster
  435. Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate RewardPoster
  436. RigGS: Rigging of 3D Gaussians for Modeling Articulated Objects in VideosPoster
  437. RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and SafetyPoster
  438. RivuletMLP: An MLP-based Architecture for Efficient Compressed Video Quality EnhancementPoster
  439. RoGSplat: Learning Robust Generalizable Human Gaussian Splatting from Sparse Multi-View ImagesPoster
  440. RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video NarrativesPoster
  441. RobSense: A Robust Multi-modal Foundation Model for Remote Sensing with Static, Temporal, and Incomplete Data AdaptabilityPoster
  442. RoboGround: Robotic Manipulation with Grounded Vision-Language PriorsPoster
  443. Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent AlignmentPoster
  444. Robust Message Embedding via Attention Flow-Based SteganographyPoster
  445. Robust Multi-Object 4D Generation for In-the-wild VideosPoster
  446. Robust Multimodal Survival Prediction with Conditional Latent Differentiation Variational AutoEncoderPoster
  447. Robust-MVTON: Learning Cross-Pose Feature Alignment and Fusion for Robust Multi-View Virtual Try-OnPoster
  448. Rotation-Equivariant Self-Supervised Method in Image DenoisingPoster
  449. S2D-LFE: Sparse-to-Dense Light Field Event GenerationPoster
  450. S2Gaussian: Sparse-View Super-Resolution 3D Gaussian SplattingPoster
  451. S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual RepresentationPoster
  452. SACB-Net: Spatial-awareness Convolutions for Medical Image RegistrationHighlight
  453. SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image PretrainingPoster
  454. SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training CostPoster
  455. SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything ModelPoster
  456. SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual ScenesPoster
  457. SAM2Object: Consolidating View Consistency via SAM2 for Zero-Shot 3D Instance SegmentationPoster
  458. SAMBLE: Shape-Specific Point Cloud Sampling for an Optimal Trade-Off Between Local Detail and Global UniformityPoster
  459. SASep: Saliency-Aware Structured Separation of Geometry and Feature for Open Set Learning on Point CloudsPoster
  460. SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive TokensPoster
  461. SATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision TransformersPoster
  462. SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute PromptingPoster
  463. SCFlow2: Plug-and-Play Object Pose Refiner with Shape-Constraint Scene FlowPoster
  464. SCSA: A Plug-and-Play Semantic Continuous-Sparse Attention for Arbitrary Semantic Style TransferHighlight
  465. SDBF: Steep-Decision-Boundary Fingerprinting for Hard-Label Tampering Detection of DNN ModelsPoster
  466. SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy PredictionPoster
  467. SEAL: Semantic Attention Learning for Long Video RepresentationPoster
  468. SEEN-DA: SEmantic ENtropy guided Domain-aware Attention for Domain Adaptive Object DetectionPoster
  469. SET: Spectral Enhancement for Tiny Object DetectionPoster
  470. SFDM: Robust Decomposition of Geometry and Reflectance for Realistic Face Rendering from Sparse-view ImagesPoster
  471. SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI DetectionPoster
  472. SGCR: Spherical Gaussians for Efficient 3D Curve ReconstructionPoster
  473. SGFormer: Satellite-Ground Fusion for 3D Semantic Scene CompletionPoster
  474. SINR: Sparsity Driven Compressed Implicit Neural RepresentationsPoster
  475. SIR-DIFF: Sparse Image Sets Restoration with Multi-View Diffusion ModelPoster
  476. SKDream: Controllable Multi-view and 3D Generation with Arbitrary SkeletonsHighlight
  477. SKE-Layout: Spatial Knowledge Enhanced Layout Generation with LLMsPoster
  478. SLADE: Shielding against Dual Exploits in Large Vision-Language ModelsPoster
  479. SLVR: Super-Light Visual Reconstruction via Blueprint Controllable Convolutions and Exploring Feature Diversity RepresentationPoster
  480. SMILE: Infusing Spatial and Motion Semantics in Masked Video LearningPoster
  481. SMTPD: A New Benchmark for Temporal Prediction of Social Media PopularityPoster
  482. SOAP: Vision-Centric 3D Semantic Scene Completion with Scene-Adaptive Decoder and Occluded Region-Aware View ProjectionPoster
  483. SOGS: Second-Order Anchor for Advanced 3D Gaussian SplattingPoster
  484. SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous DrivingPoster
  485. SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic PromptsHighlight
  486. SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language ModelsPoster
  487. SPC-GS: Gaussian Splatting with Semantic-Prompt Consistency for Indoor Open-World Free-view Synthesis from Sparse InputsPoster
  488. SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual TrackingPoster
  489. SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split OptimizationHighlight
  490. STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural NetworksPoster
  491. STAR-Edge: Structure-aware Local Spherical Curve Representation for Thin-walled Edge Extraction from Unstructured Point CloudsPoster
  492. STDD: Spatio-Temporal Dual Diffusion for Video GenerationPoster
  493. STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security InspectionHighlight
  494. STOP: Integrated Spatial-Temporal Dynamic Prompting for Video UnderstandingPoster
  495. STiL: Semi-supervised Tabular-Image Learning for Comprehensive Task-Relevant Information Exploration in Multimodal ClassificationPoster
  496. SUM Parts: Benchmarking Part-Level Semantic Segmentation of Urban MeshesPoster
  497. SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation SparsityHighlight
  498. SVDC: Consistent Direct Time-of-Flight Video Depth Completion with Frequency Selective FusionPoster
  499. SVFR: A Unified Framework for Generalized Video Face RestorationPoster
  500. SVG-IR: Spatially-Varying Gaussian Splatting for Inverse RenderingPoster
  501. SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video SituationPoster
  502. S^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion PriorsPoster
  503. SaMam: Style-aware State Space Model for Arbitrary Image Style TransferHighlight
  504. Saliuitl: Ensemble Salience Guided Recovery of Adversarial Patches against CNNsPoster
  505. Samba: A Unified Mamba-based Framework for General Salient Object DetectionHighlight
  506. Sample- and Parameter-Efficient Auto-Regressive Image ModelsPoster
  507. Sampling Innovation-Based Adaptive Compressive SensingPoster
  508. Satellite Observations Guided Diffusion Model for Accurate Meteorological States at Arbitrary ResolutionHighlight
  509. Scalable Video-to-Dataset Generation for Cross-Platform Mobile AgentsPoster
  510. Scale Efficient Training for Large DatasetsPoster
  511. ScaleLSD: Scalable Deep Line Segment Detection StreamlinedPoster
  512. Scaling Down Text Encoders of Text-to-Image Diffusion ModelsPoster
  513. Scaling Inference Time Compute for Diffusion ModelsHighlight
  514. Scaling Vision Pre-Training to 4K ResolutionHighlight
  515. Scaling up Image Segmentation across Data and TasksPoster
  516. Scene-Centric Unsupervised Panoptic SegmentationHighlight
  517. Scene-agnostic Pose Regression for Visual LocalizationPoster
  518. Scene4U: Hierarchical Layered 3D Scene Reconstruction from Single Panoramic Image for Your Immerse ExplorationPoster
  519. SceneCrafter: Controllable Multi-View Driving Scene EditingPoster
  520. SceneDiffuser++: City-Scale Traffic Simulation via a Generative World ModelPoster
  521. Sea-ing in Low-lightPoster
  522. SeaLion: Semantic Part-Aware Latent Point Diffusion Models for 3D GenerationPoster
  523. Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution AnalysisPoster
  524. See Further When Clear: Curriculum Consistency ModelPoster
  525. Seeing A 3D World in A Grain of SandPoster
  526. Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal DecodingPoster
  527. Seeing More with Less: Human-like Representations in Vision ModelsHighlight
  528. Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual ScenesPoster
  529. Seeing What Matters: Empowering CLIP with Patch Generation-to-SelectionPoster
  530. Seeing is Not Believing: Adversarial Natural Object Optimization for Hard-Label 3D Scene AttacksPoster
  531. Seeing the Abstract: Translating the Abstract Language for Vision Language ModelsPoster
  532. Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment RecognitionPoster
  533. Seeking Consistent Flat Minima for Better Domain Generalization via Refining Loss LandscapesPoster
  534. SegMAN: Omni-scale Context Modeling with State Space Models and Local Attention for Semantic SegmentationPoster
  535. Segment Any Motion in VideosPoster
  536. Segment Any-Quality Images with Generative Latent Space EnhancementPoster
  537. Segment Anything, Even OccludedPoster
  538. Segment This Thing: Foveated Tokenization for Efficient Point-Prompted SegmentationPoster
  539. Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar SubjectsPoster
  540. Self-Evolving Visual Concept Library using Vision-Language CriticsPoster
  541. Self-Learning Hyperspectral and Multispectral Image Fusion via Adaptive Residual Guided Subspace Diffusion ModelPoster
  542. Self-Supervised Cross-View Correspondence with Predictive Cycle ConsistencyHighlight
  543. Self-Supervised Large Scale Point Cloud Completion for Archaeological Site RestorationPoster
  544. Self-Supervised Learning for Color Spike Camera ReconstructionPoster
  545. Self-Supervised Spatial Correspondence Across ModalitiesPoster
  546. Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolutionPoster
  547. SemAlign3D: Semantic Correspondence between RGB-Images through Aligning 3D Object-Class RepresentationsPoster
  548. SemGeoMo: Dynamic Contextual Human Motion Generation with Semantic and Geometric GuidancePoster
  549. Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic SegmentationPoster
  550. Semantic and Expressive Variations in Image Captions Across LanguagesPoster
  551. Semantic-guided Cross-Modal Prompt Learning for Skeleton-based Zero-shot Action RecognitionPoster
  552. Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video DerainingPoster
  553. SemiDAViL: Semi-supervised Domain Adaptation with Vision-Language Guidance for Semantic SegmentationPoster
  554. SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text SpottingPoster
  555. Sensitivity-Aware Efficient Fine-Tuning via Compact Dynamic-Rank AdaptationPoster
  556. Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question AnsweringPoster
  557. SeqMvRL: A Sequential Fusion Framework for Multi-view Representation LearningPoster
  558. SeriesBench: A Benchmark for Narrative-Driven Drama Series UnderstandingPoster
  559. Seurat: From Moving Points to DepthHighlight
  560. Shadow Generation Using Diffusion Model with Geometry PriorPoster
  561. Shape Abstraction via Marching Differentiable Support FunctionsHighlight
  562. Shape and Texture: What Influences Reliable Optical Flow Estimation?Poster
  563. ShapeShifter: 3D Variations Using Multiscale and Sparse Point-Voxel DiffusionPoster
  564. ShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware PromptsPoster
  565. Sharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and ManipulationPoster
  566. Shift the Lens: Environment-Aware Unsupervised Camouflaged Object DetectionPoster
  567. ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion ModelsPoster
  568. Show and Segment: Universal Medical Image Segmentation via In-Context LearningPoster
  569. Show and Tell: Visually Explainable Deep Neural Nets via Spatially-Aware Concept Bottleneck ModelsHighlight
  570. ShowMak3r: Compositional TV Show ReconstructionPoster
  571. Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head GenerationPoster
  572. SimLTD: Simple Supervised and Semi-Supervised Long-Tailed Object DetectionPoster
  573. SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity PredictionPoster
  574. Simpler Diffusion: 1.5 FID on ImageNet512 with Pixel-space DiffusionPoster
  575. Simulator HC: Regression-based Online Simulation of Starting Problem-Solution Pairs for Homotopy Continuation in Geometric VisionHighlight
  576. SinGS: Animatable Single-Image Human Gaussian Splats with Kinematic PriorsPoster
  577. Single Domain Generalization for Few-Shot Counting via Universal Representation MatchingPoster
  578. Six-CD: Benchmarking Concept Removals for Text-to-image Diffusion ModelsPoster
  579. Sketch Down the FLOPs: Towards Efficient Networks for Human SketchPoster
  580. SketchFusion: Learning Universal Sketch Features through Fusing Foundation ModelsPoster
  581. SketchVideo: Sketch-based Video Generation and EditingPoster
  582. Sketchtopia: A Dataset and Foundational Agents for Benchmarking Asynchronous Multimodal Communication with Iconic FeedbackPoster
  583. Sketchy Bounding-box Supervision for 3D Instance SegmentationPoster
  584. Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters ThemselvesPoster
  585. SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language ModelingPoster
  586. SmartCLIP: Modular Vision-language Alignment with Identification GuaranteesHighlight
  587. SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback OptimizationPoster
  588. SoMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation LearningHighlight
  589. SocialGesture: Delving into Multi-person Gesture UnderstandingPoster
  590. SocialMOIF: Multi-Order Intention Fusion for Pedestrian Trajectory PredictionPoster
  591. Soft Self-labeling and Potts Relaxations for Weakly-supervised SegmentationPoster
  592. SoftShadow: Leveraging Soft Masks for Penumbra-Aware Shadow RemovalPoster
  593. Sound Bridge: Associating Egocentric and Exocentric Videos via Audio CuesPoster
  594. SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic BindingHighlight
  595. Sparse Point Cloud Patches Rendering via Splitting 2D GaussiansPoster
  596. Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field RenderingPoster
  597. Sparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse ViewsPoster
  598. SparseAlign: a Fully Sparse Framework for Cooperative Object DetectionPoster
  599. Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image SynthesisPoster
  600. SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative LanguagePoster
  601. SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular InputPoster
  602. SpecTRe-GS: Modeling Highly Specular Surfaces with Reflected Nearby Objects by Tracing Rays in 3D Gaussian SplattingHighlight
  603. Spectral State Space Model for Rotation-Invariant Visual Representation LearningPoster
  604. SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionPoster
  605. Spherical Manifold Guided Diffusion Model for Panoramic Image GenerationPoster
  606. Spk2SRImgNet: Super-Resolve Dynamic Scene from Spike Stream via Motion Aligned Collaborative FilteringPoster
  607. SplatFlow: Self-Supervised Dynamic Gaussian Splatting in Neural Motion Flow Field for Autonomous DrivingHighlight
  608. Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous DrivingPoster
  609. Stabilizing and Accelerating Autofocus with Expert Trajectory Regularized Deep Reinforcement LearningPoster
  610. Stable-SCore: A Stable Registration-based Framework for 3D Shape CorrespondencePoster
  611. StageDesigner: Artistic Stage Generation for Scenography via Theater ScriptsPoster
  612. Star with Bilinear MappingPoster
  613. Stealthy Backdoor Attack in Self-Supervised Learning Vision Encoders for Large Vision Language ModelsPoster
  614. Steepest Descent Density Control for Compact 3D Gaussian SplattingPoster
  615. Stochastic Human Motion Prediction with Memory of Action Transition and Action CharacteristicPoster
  616. Stop Learning it all to Mitigate Visual Hallucination, Focus on the Hallucination Target.Poster
  617. Stop Walking in Circles! Bailing Out Early in Projected Gradient DescentPoster
  618. Structure from CollisionHighlight
  619. Structure-Aware Correspondence Learning for Relative Pose EstimationHighlight
  620. Structure-from-Motion with a Non-Parametric Camera ModelHighlight
  621. Style Evolving along Chain-of-Thought for Unknown-Domain Object DetectionHighlight
  622. Style Quantization for Data-Efficient GAN TrainingPoster
  623. Style-Editor: Text-driven Object-centric Style EditingHighlight
  624. StyleSSP: Sampling StartPoint Enhancement for Training-free Diffusion-based Method for Style TransferHighlight
  625. StyleStudio: Text-Driven Style Transfer with Selective Control of Style ElementsPoster
  626. Subnet-Aware Dynamic Supernet Training for Neural Architecture SearchPoster
  627. Subspace Constraint and Contribution Estimation for Heterogeneous Federated LearningPoster
  628. Supervising Sound Localization by In-the-wild EgomotionHighlight
  629. Symbolic Representation for Any-to-Any Generative TasksPoster
  630. Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D GenerationHighlight
  631. SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled SynthesisPoster
  632. SyncSDE: A Probabilistic Framework for Diffusion SynchronizationPoster
  633. SyncVP: Joint Diffusion for Synchronous Multi-Modal Video PredictionPoster
  634. Synchronized Video-to-Audio Generation via Mel Quantization-Continuum DecompositionPoster
  635. SynthLight: Portrait Relighting with Diffusion Model by Learning to Re-render Synthetic FacesPoster
  636. Synthetic Data is an Elegant GIFT for Continual Vision-Language ModelsPoster
  637. Synthetic Visual GenomePoster
  638. T-CIL: Temperature Scaling using Adversarial Perturbation for Calibration in Class-Incremental LearningPoster
  639. T-FAKE: Synthesizing Thermal Images for Facial LandmarkingPoster
  640. T2ICount: Enhancing Cross-modal Understanding for Zero-Shot CountingHighlight
  641. TADFormer: Task-Adaptive Dynamic TransFormer for Efficient Multi-Task LearningPoster
  642. TAET: Two-Stage Adversarial Equalization Training on Long-Tailed DistributionsPoster
  643. TAGA: Self-supervised Learning for Template-free Animatable Gaussian Articulated ModelPoster
  644. TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action RecognitionPoster
  645. TANGO: Training-free Embodied AI Agents for Open-world TasksPoster
  646. TAROT: Towards Essentially Domain-Invariant Robustness with Theoretical JustificationPoster
  647. TCFG: Tangential Damping Classifier-free GuidancePoster
  648. TFCustom: Customized Image Generation with Time-Aware Frequency Feature GuidanceHighlight
  649. TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time CorrectionHighlight
  650. TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion GenerationPoster
  651. TKG-DM: Training-free Chroma Key Content Generation Diffusion ModelHighlight
  652. TSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual SegmentationPoster
  653. TSP-Mamba: The Travelling Salesman Problem Meets Mamba for Image Super-resolution and BeyondPoster
  654. TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage FusionAward Candidate
  655. TailedCore: Few-Shot Sampling for Unsupervised Long-Tail Noisy Anomaly DetectionPoster
  656. Take the Bull by the Horns: Learning to Segment Hard SamplesPoster
  657. Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse InputsHighlight
  658. TaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian SplattingHighlight
  659. Targeted Forgetting of Image Subgroups in CLIP ModelsPoster
  660. Tartan IMU: A Light Foundation Model for Inertial Positioning in RoboticsPoster
  661. Task-Aware Clustering for Prompting Vision-Language ModelsPoster
  662. Task-Specific Gradient Adaptation for Few-Shot One-Class ClassificationPoster
  663. Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual GroundingPoster
  664. Taste More, Taste Better: Diverse Data and Strong Model Boost Semi-Supervised Crowd CountingPoster
  665. Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion GenerationPoster
  666. Temporal Action Detection Model Compression by Progressive Block DropPoster
  667. Temporal Alignment-Free Video Matching for Few-shot Action RecognitionPoster
  668. Temporal Score Analysis for Understanding and Correcting Diffusion ArtifactsPoster
  669. TensoFlow: Tensorial Flow-based Sampler for Inverse RenderingPoster
  670. Test-Time Domain Generalization via Universe Learning: A Multi-Graph Matching Approach for Medical Image SegmentationPoster
  671. Test-Time Fine-Tuning of Image Compression Models for Multi-Task AdaptabilityPoster
  672. Test-Time Visual In-Context TuningPoster
  673. Test-time Augmentation Improves Efficiency in Conformal PredictionPoster
  674. TexGarment: Consistent Garment UV Texture Generation via Efficient 3D Structure-Guided Diffusion TransformerPoster
  675. Text Augmented Correlation Transformer For Few-shot Classification & SegmentationPoster
  676. Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual AbductionPoster
  677. Text-guided Sparse Voxel Pruning for Efficient 3D Visual GroundingHighlight
  678. The Art of Deception: Color Visual Illusions and Diffusion ModelsPoster
  679. The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video GenerationPoster
  680. The Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion ModelsPoster
  681. The Impact Label Noise and Choice of Threshold has on Cross-Entropy and Soft-Dice in Image SegmentationPoster
  682. The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionAward Candidate
  683. The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like PhotographersPoster
  684. Theory-Inspired Deep Multi-View Multi-Label Learning with Incomplete Views and Noisy LabelsPoster
  685. Thin-Shell-SfT: Fine-Grained Monocular Non-rigid 3D Surface Tracking with Neural Deformation FieldsPoster
  686. Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot ManipulationPoster
  687. Three Cars Approaching within 100m! Enhancing Distant Geometry by Tri-Axis Voxel Scanning for Camera-based Semantic Scene CompletionPoster
  688. Three-view Focal Length Recovery From HomographiesPoster
  689. Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video GenerationPoster
  690. Tightening Robustness Verification of MaxPool-based Neural Networks via Minimizing the Over-Approximation ZonePoster
  691. Tiled DiffusionPoster
  692. Time of the Flight of the Gaussians: Optimizing Depth Indirectly in Dynamic Radiance FieldsPoster
  693. TimeTracker: Event-based Continuous Point Tracking for Video Frame Interpolation with Non-linear MotionPoster
  694. TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video GenerationPoster
  695. TopNet: Transformer-Efficient Occupancy Prediction Network for Octree-Structured Point Cloud Geometry CompressionPoster
  696. Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and ReconstructionPoster
  697. Toward Real-world BEV Perception: Depth Uncertainty Estimation via Gaussian SplattingPoster
  698. Towards All-in-One Medical Image Re-IdentificationPoster
  699. Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific ParametersPoster
  700. Towards Continual Universal SegmentationPoster
  701. Towards Cost-Effective Learning: A Synergy of Semi-Supervised and Active LearningPoster
  702. Towards Effective and Sparse Adversarial Attack on Spiking Neural Networks via Breaking Invisible Surrogate GradientsPoster
  703. Towards Efficient Foundation Model for Zero-shot Amodal SegmentationPoster
  704. Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color ConsistencyHighlight
  705. Towards Explainable and Unprecedented Accuracy in Matching Challenging Finger Crease PatternsHighlight
  706. Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse WeatherPoster
  707. Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency PartitionPoster
  708. Towards Generalizable Scene Change DetectionPoster
  709. Towards High-fidelity 3D Talking Avatar with Personalized Dynamic TexturePoster
  710. Towards Human-Understandable Multi-Dimensional Concept DiscoveryPoster
  711. Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long TextHighlight
  712. Towards In-the-wild 3D Plane Reconstruction from a Single ImageHighlight
  713. Towards Lossless Implicit Neural Representation via Bit Plane DecompositionPoster
  714. Towards Natural Language-Based Document Image Retrieval: New Dataset and BenchmarkPoster
  715. Towards Optimizing Large-Scale Multi-Graph Matching in BioimagingPoster
  716. Towards Precise Embodied Dialogue Localization via Causality Guided DiffusionPoster
  717. Towards Satellite Image Road Graph Extraction: A Global-Scale Dataset and A Novel MethodPoster
  718. Towards Scalable Human-aligned Benchmark for Text-guided Image EditingHighlight
  719. Towards Smart Point-and-Shoot PhotographyPoster
  720. Towards Source-Free Machine UnlearningPoster
  721. Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and AnticipationHighlight
  722. Towards Understanding and Quantifying Uncertainty for Text-to-Image GenerationPoster
  723. Towards Universal AI-Generated Image Detection by Variational Information Bottleneck NetworkPoster
  724. Towards Universal Dataset Distillation via Task-Driven DiffusionPoster
  725. Track Any Anomalous Object:A Granular Video Anomaly Detection PipelinePoster
  726. Tracktention: Leveraging Point Tracking to Attend Videos Faster and BetterHighlight
  727. Training Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training?Poster
  728. Training-free Dense-Aligned Diffusion Guidance for Modular Conditional Image SynthesisPoster
  729. Training-free Neural Architecture Search through Variance of Knowledge of Deep Network WeightsPoster
  730. TransPixeler: Advancing Text-to-Video Generation with TransparencyPoster
  731. Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving ScenePoster
  732. TriTex: Learning Texture from a Single Mesh via Triplane Semantic FeaturesPoster
  733. Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental LearningPoster
  734. Tuning the Frequencies: Robust Training for Sinusoidal Neural NetworksHighlight
  735. TurboFill: Adapting Few-step Text-to-image Model for Fast Image InpaintingPoster
  736. Twinner: Shining Light on Digital Twins in a Few SnapsPoster
  737. Two is Better than One: Efficient Ensemble Defense for Robust and Compact ModelsPoster
  738. U-Know-DiffPAN: An Uncertainty-aware Knowledge Distillation Diffusion Framework with Details Enhancement for PAN-SharpeningPoster
  739. UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial ReferencesPoster
  740. UCM-VeID V2: A Richer Dataset and A Pre-training Method for UAV Cross-Modality Vehicle Re-IdentificationPoster
  741. UCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label LearningHighlight
  742. UHD-processer: Unified UHD Image Restoration with Progressive Frequency Learning and Degradation-aware PromptsPoster
  743. UIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion ModelsHighlight
  744. UMFN: Unified Multi-Domain Face Normalization for Joint Cross-domain Prototype Learning and Heterogeneous Face RecognitionPoster
  745. UMotion: Uncertainty-driven Human Motion Estimation from Inertial and Ultra-wideband UnitsHighlight
  746. UNEM: UNrolled Generalized EM for Transductive Few-Shot LearningPoster
  747. UNIALIGN: Scaling Multimodal Alignment within One Unified ModelPoster
  748. UNICL-SAM: Uncertainty-Driven In-Context Segmentation with Part Prototype DiscoveryPoster
  749. UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model EvaluationPoster
  750. URWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image RestorationPoster
  751. UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video ParsingPoster
  752. UltraFusion: Ultra High Dynamic Imaging using Exposure FusionHighlight
  753. Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingPoster
  754. Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video BenchmarksPoster
  755. Unboxed: Geometrically and Temporally Consistent Video OutpaintingPoster
  756. Uncertainty Meets Diversity: A Comprehensive Active Learning Framework for Indoor 3D Object DetectionPoster
  757. Uncertainty Weighted Gradients for Model CalibrationPoster
  758. Uncertainty-Instructed Structure Injection for Generalizable HD Map ConstructionPoster
  759. Uncertainty-guided Perturbation for Image Super-Resolution Diffusion ModelPoster
  760. Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic SpacePoster
  761. Understanding Multi-Task Activities from Single-Task VideosHighlight
  762. Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream DiffusionPoster
  763. Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single VideoHighlight
  764. UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose EstimationPoster
  765. UniK3D: Universal Camera Monocular 3D EstimationPoster
  766. UniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object DetectionPoster
  767. UniNet: A Contrastive Learning-guided Unified Framework with Feature Selection for Anomaly DetectionPoster
  768. UniPhy: Learning a Unified Constitutive Model for Inverse Physics SimulationPoster
  769. UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian SplattingPoster
  770. UniSTD: Towards Unified Spatio-Temporal Learning across Diverse DisciplinesPoster
  771. Unified Dense Prediction of Video DiffusionPoster
  772. Unified Medical Lesion Segmentation via Self-referring IndicatorPoster
  773. Unified Reconstruction of Static and Dynamic Scenes from EventsHighlight
  774. Unity in Diversity: Video Editing via Gradient-Latent PurificationPoster
  775. Universal Domain Adaptation for Semantic SegmentationPoster
  776. Universal Scene Graph GenerationHighlight
  777. Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media ManipulationPoster
  778. Unlocking Generalization Power in LiDAR Point Cloud RegistrationHighlight
  779. Unlocking the Potential of Unlabeled Data in Semi-Supervised Domain GeneralizationPoster
  780. Unraveling Normal Anatomy via Fluid-Driven Anomaly RandomizationPoster
  781. Unsupervised Continual Domain Shift Learning with Multi-Prototype ModelingHighlight
  782. Unsupervised Discovery of Facial Landmarks and Head PosePoster
  783. Unveiling Differences in Generative Models: A Scalable Differential Clustering ApproachHighlight
  784. UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene SimulationPoster
  785. Using Diffusion Priors for Video Amodal SegmentationPoster
  786. Using Powerful Prior Knowledge of Diffusion Model in Deep Unfolding Networks for Image Compressive SensingPoster
  787. V-Stylist: Video Stylization via Collaboration and Reflection of MLLM AgentsPoster
  788. V2V3D: View-to-View Denoised 3D Reconstruction for Light Field MicroscopyPoster
  789. V2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object DetectionPoster
  790. VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict EntailmentPoster
  791. VEU-Bench: Towards Comprehensive Understanding of Video EditingHighlight
  792. VIRES: Video Instance Repainting via Sketch and Text Guided GenerationPoster
  793. VISTREAM: Improving Computation Efficiency of Visual Streaming Perception via Law-of-Charge-Conservation Inspired Spiking Neural NetworkPoster
  794. VI^3NR: Variance Informed Initialization for Implicit Neural RepresentationsPoster
  795. VL2Lite: Task-Specific Knowledge Distillation from Large Vision-Language Models to Lightweight NetworksPoster
  796. VLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement LearningPoster
  797. VLog: Video-Language Models by Generative Retrieval of Narration VocabularyPoster
  798. VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language ModelsPoster
  799. VODiff: Controlling Object Visibility Order in Text-to-Image GenerationPoster
  800. VSNet: Focusing on the Linguistic Characteristics of Sign LanguagePoster
  801. V^2Dial: Unification of Video and Visual Dialog via Multimodal ExpertsPoster
  802. Variance-Based Membership Inference Attacks Against Large-Scale Image Captioning ModelsPoster
  803. VasTSD: Learning 3D Vascular Tree-state Space Diffusion Model for Angiography SynthesisPoster
  804. VerbDiff: Text-Only Diffusion Models with Enhanced Interaction AwarenessPoster
  805. ViKIENet: Towards Efficient 3D Object Detection with Virtual Key Instance Enhanced NetworkPoster
  806. ViUniT: Visual Unit Tests for More Robust Visual ProgrammingPoster
  807. Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal PriorPoster
  808. Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free SimulationPoster
  809. VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic ManipulationPoster
  810. VidSeg: Training-free Video Semantic Segmentation based on Diffusion ModelsPoster
  811. Video Language Model Pretraining with Spatio-temporal MaskingPoster
  812. Video Summarization with Large Language ModelsPoster
  813. Video-Bench: Human-Aligned Video Generation BenchmarkPoster
  814. Video-ColBERT: Contextualized Late Interaction for Text-to-Video RetrievalPoster
  815. Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language ModelsPoster
  816. VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text ModelsPoster
  817. VideoDirector: Precise Video Editing via Text-to-Video ModelsPoster
  818. VideoGEM: Training-free Action Grounding in VideosPoster
  819. VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion ModelsPoster
  820. VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and EditingPoster
  821. Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation LearningPoster
  822. ViiNeuS: Volumetric Initialization for Implicit Neural Surface Reconstruction of Urban Scenes with Limited Image OverlapPoster
  823. Vision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D ScenesPoster
  824. Vision-Language Embodiment for Monocular Depth EstimationPoster
  825. Vision-Language Gradient Descent-driven All-in-One Deep Unfolding NetworksPoster
  826. Vision-Language Model IP Protection via Prompt-based LearningPoster
  827. Visual Consensus Prompting for Co-Salient Object DetectionPoster
  828. Visual Persona: Foundation Model for Full-Body Human CustomizationPoster
  829. Visual Representation Learning through Causal Intervention for Controllable Image EditingHighlight
  830. Visual and Semantic Prompt Collaboration for Generalized Zero-Shot LearningPoster
  831. Visual-Instructed Degradation Diffusion for All-in-One Image RestorationPoster
  832. VladVA: Discriminative Fine-tuning of LVLMsPoster
  833. VolFormer: Explore More Comprehensive Cube Interaction for Hyperspectral Image Restoration and BeyondPoster
  834. Volume Tells: Dual Cycle-Consistent Diffusion for 3D Fluorescence Microscopy De-noising and Super-ResolutionHighlight
  835. Volumetric Surfaces: Representing Fuzzy Geometries with Layered MeshesPoster
  836. Volumetrically Consistent 3D Gaussian RasterizationHighlight
  837. WISE: A Framework for Gigapixel Whole-Slide-Image Lossless CompressionPoster
  838. WISH: Weakly Supervised Instance Segmentation using Heterogeneous LabelsHighlight
  839. WISNet: Pseudo Label Generation on Unbalanced and Patch Annotated Waste ImagesPoster
  840. Watermarking One for All: A Robust Watermarking Scheme Against Partial Image TheftPoster
  841. Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial AnimationPoster
  842. Wavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect DetectionPoster
  843. WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and SegmentationPoster
  844. Weakly Supervised Contrastive Adversarial Training for Learning Robust Features from Semi-supervised DataPoster
  845. Weakly Supervised Semantic Segmentation via Progressive Confidence Region ExpansionPoster
  846. Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language ModelsPoster
  847. WeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba DiffusionPoster
  848. When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation LearningPoster
  849. Where the Devil Hides: Deepfake Detectors Can No Longer Be TrustedPoster
  850. Where's the Liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated ContentPoster
  851. Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional VideosHighlight
  852. Your Scale Factors are My Weapon: Targeted Bit-Flip Attacks on Vision Transformers via Scale Factor ManipulationPoster
  853. Z-Magic: Zero-shot Multiple Attributes Guided Image CreatorPoster
  854. Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video DiffusionPoster
  855. Zero-Shot 4D Lidar Panoptic SegmentationPoster
  856. Zero-Shot Blind-spot Image Denoising via Implicit Neural SamplingPoster
  857. Zero-Shot Head Swapping in Real-World ScenariosPoster
  858. Zero-Shot Image Restoration Using Few-Step Guidance of Consistency Models (and Beyond)Poster
  859. Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric DiffusionPoster
  860. Zero-Shot Styled Text Image Generation, but Make It AutoregressivePoster
  861. Zero-shot 3D Question Answering via Voxel-based Dynamic Token CompressionPoster
  862. Zero-shot RGB-D Point Cloud Registration with Pre-trained Large Vision ModelPoster
  863. ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic GraspingPoster
  864. ZeroVO: Visual Odometry with Minimal AssumptionsPoster
  865. beta-FFT: Nonlinear Interpolation and Differentiated Training Strategies for Semi-Supervised Medical Image SegmentationPoster
  866. dFLMoE: Decentralized Federated Learning via Mixture of Experts for Medical Data AnalysisPoster
  867. iG-6DoF: Model-free 6DoF Pose Estimation for Unseen Object via Iterative 3D Gaussian SplattingPoster
  868. iSegMan: Interactive Segment-and-Manipulate 3D GaussiansPoster
  869. nnWNet: Rethinking the Use of Transformers in Biomedical Image Segmentation and Calling for a Unified Evaluation BenchmarkPoster
  870. pFedMxF: Personalized Federated Class-Incremental Learning with Mixture of Frequency AggregationPoster
  871. v-CLR: View-Consistent Learning for Open-World Instance SegmentationHighlight
  872. vesselFM: A Foundation Model for Universal 3D Blood Vessel SegmentationPoster

Looking for submission deadlines instead? See the conference deadline calendar.