← All conferences

ICCV 2025 Accepted Papers

The full list of 2,620 papers accepted at ICCV 2025 (IEEE/CVF International Conference on Computer Vision). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

Poster: 2,595
  1. SDFit: 3D Object Pose and Shape by Fitting a Morphable SDF to a Single ImagePoster
  2. SDFormer: Vision-based 3D Semantic Scene Completion via SAM-assisted Dual-channel Voxel TransformerPoster
  3. SDMatte: Grafting Diffusion Models for Interactive MattingPoster
  4. SEAL: Semantic Aware Image WatermarkingPoster
  5. SEGS-SLAM: Structure-enhanced 3D Gaussian Splatting SLAM with Appearance EmbeddingPoster
  6. SEHDR: Single-Exposure HDR Novel View Synthesis via 3D Gaussian BracketingPoster
  7. SEREP: Semantic Facial Expression Representation for Robust In-the-Wild Capture and RetargetingPoster
  8. SFUOD: Source-Free Unknown Object DetectionPoster
  9. SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned DiffusionPoster
  10. SGAD: Semantic and Geometric-aware Descriptor for Local Feature MatchingPoster
  11. SHIFT: Smoothing Hallucinations by Information Flow Tuning for Multimodal Large Language ModelsPoster
  12. SHeaP: Self-Supervised Head Geometry Predictor Learned via 2D GaussiansPoster
  13. SIC: Similarity-Based Interpretable Image Classification with Neural NetworksPoster
  14. SIGMAN: Scaling 3D Human Gaussian Generation with Millions of AssetsPoster
  15. SILO: Solving Inverse Problems with Latent OperatorsPoster
  16. SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script GenerationPoster
  17. SITE: towards Spatial Intelligence Thorough EvaluationPoster
  18. SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video CreationPoster
  19. SL2A-INR: Single-Layer Learnable Activation for Implicit Neural RepresentationPoster
  20. SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
  21. SMGDiff: Soccer Motion Generation using Diffusion Probabilistic ModelsPoster
  22. SMP-Attack: Boosting the Transferability of Feature Importance-based Adversarial Attack with Semantics-aware Multi-granularity PatchoutPoster
  23. SMSTracker: Tri-path Score Mask Sigma Fusion for Multi-Modal TrackingPoster
  24. SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction TuningPoster
  25. SP2T: Sparse Proxy Attention for Dual-stream Point TransformerPoster
  26. SPA: Efficient User-Preference Alignment against Uncertainty in Medical Image SegmentationPoster
  27. SPADE: Spatial-Aware Denoising Network for Open-vocabulary Panoptic Scene Graph Generation with Long- and Local-range Context ReasoningPoster
  28. SPD: Shallow Backdoor Protecting Deep Backdoor Against Backdoor DetectionPoster
  29. SRefiner: Soft-Braid Attention for Multi-Agent Trajectory RefinementPoster
  30. SSVQ: Unleashing the Potential of Vector Quantization with Sign-SplittingPoster
  31. STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-ResolutionPoster
  32. STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic ScenePoster
  33. STDDNet: Harnessing Mamba for Video Polyp Segmentation via Spatial-aligned Temporal Modeling and Discriminative Dynamic Representation LearningPoster
  34. STEP-DETR: Advancing DETR-based Semi-Supervised Object Detection with Super Teacher and Pseudo-Label Guided Text QueriesPoster
  35. STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?Poster
  36. STIV: Scalable Text and Image Conditioned Video GenerationPoster
  37. STaR: Seamless Spatial-Temporal Aware Motion Retargeting with Penetration and Consistency ConstraintsPoster
  38. SU-RGS: Relightable 3D Gaussian Splatting from Sparse Views under Unconstrained IlluminationsPoster
  39. SUB: Benchmarking CBM Generalization via Synthetic Attribute SubstitutionsPoster
  40. SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D GenerationPoster
  41. SVG-Head: Hybrid Surface-Volumetric Gaussians for High-Fidelity Head Reconstruction and Real-Time EditingPoster
  42. SVIP: Semantically Contextualized Visual Patches for Zero-Shot LearningPoster
  43. SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionPoster
  44. SViM3D: Stable Video Material Diffusion for Single Image 3D GenerationPoster
  45. Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based AttacksPoster
  46. Saliency-Aware Quantized Imitation Learning for Efficient Robotic ControlPoster
  47. Salvaging the Overlooked: Leveraging Class-Aware Contrastive Learning for Multi-Class Anomaly DetectionPoster
  48. Sat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent DiffusionPoster
  49. Scalable Dual Fingerprinting for Hierarchical Attribution of Text-to-Image ModelsPoster
  50. Scalable Image Tokenization with Index Backpropagation QuantizationPoster
  51. Scalable Ranked Preference Optimization for Text-to-Image GenerationPoster
  52. Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention ScalingPoster
  53. Scaling 3D Compositional Models for Robust Classification and Pose EstimationPoster
  54. Scaling Action Detection: AdaTAD++ with Transformer-Enhanced Temporal-Spatial AdaptationPoster
  55. Scaling Inference-Time Search with Vision Value Model for Improved Visual ComprehensionPoster
  56. Scaling Language-Free Visual Representation LearningPoster
  57. Scaling Laws for Native Multimodal ModelsPoster
  58. Scaling Omni-modal Pretraining with Multimodal Context: Advancing Universal Representation Learning Across ModalitiesPoster
  59. Scaling Transformer-Based Novel View Synthesis with Models Token Disentanglement and Synthetic DataPoster
  60. Scaling Tumor Segmentation: Best Lessons from Real and Synthetic DataPoster
  61. Scaling and Taming Adversarial Training with Synthetic DataPoster
  62. ScanEdit: Hierarchically-Guided Functional 3D Scan EditingPoster
  63. Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
  64. Scene Coordinate Reconstruction PriorsPoster
  65. Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural RectificationPoster
  66. SceneMI: Motion In-betweening for Modeling Human-Scene InteractionPoster
  67. ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation AlignmentPoster
  68. SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language PretrainingPoster
  69. Scheduling Weight Transitions for Quantization-Aware TrainingPoster
  70. SciVid: Cross-Domain Evaluation of Video Models in Scientific ApplicationsPoster
  71. ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided DiffusionPoster
  72. Scoring, Remember, and Reference: Catching Camouflaged Objects in VideosPoster
  73. Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware OptimizationPoster
  74. SeaS: Few-shot Industrial Anomaly Image Generation with Separation and Sharing Fine-tuningPoster
  75. Seal Your Backdoor with Variational DefensePoster
  76. Seam360GS: Seamless 360deg Gaussian Splatting from Real-World Omnidirectional ImagesPoster
  77. Secure On-Device Video OOD Detection Without BackpropagationPoster
  78. Seeing Through Deepfakes: A Human-Inspired Framework for Multi-Face DetectionPoster
  79. Seeing and Seeing Through the Glass: Real and Synthetic Data for Multi-Layer Depth EstimationPoster
  80. Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual GroundingPoster
  81. Seeing the Unseen: A Semantic Alignment and Context-Aware Prompt Framework for Open-Vocabulary Camouflaged Object SegmentationPoster
  82. SegAnyPET: Universal Promptable Segmentation from Positron Emission Tomography ImagesPoster
  83. SegmentDreamer: Towards High-fidelity Text-to-3D Synthesis with Segmented Consistency Trajectory DistillationPoster
  84. Selective Contrastive Learning for Weakly Supervised Affordance GroundingPoster
  85. Self-Calibrated Variance-Stabilizing Transformations for Real-World Image DenoisingPoster
  86. Self-Calibrating Gaussian Splatting for Large Field-of-View ReconstructionPoster
  87. Self-Ensembling Gaussian Splatting for Few-Shot Novel View SynthesisPoster
  88. Self-Reinforcing Prototype Evolution with Dual-Knowledge Cooperation for Semi-Supervised Lifelong Person Re-IdentificationPoster
  89. Self-Supervised Sparse Sensor Fusion for Long Range PerceptionPoster
  90. Self-supervised Learning of Hybrid Part-aware 3D Representations of 2D Gaussians and SuperquadricsPoster
  91. SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance LearningPoster
  92. SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic EmphasisPoster
  93. Semantic Alignment and Reinforcement for Data-Free Quantization of Vision TransformersPoster
  94. Semantic Causality-Aware Vision-Based 3D Occupancy PredictionPoster
  95. Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision TokensPoster
  96. Semantic Watermarking Reinvented: Enhancing Robustness and Generation Quality with Fourier IntegrityPoster
  97. Semantic-guided Camera Ray Regression for Visual LocalizationPoster
  98. Semi-ViM: Bidirectional State Space Model for Mitigating Label Imbalance in Semi-Supervised LearningPoster
  99. Semi-supervised Concept Bottleneck ModelsPoster
  100. Semi-supervised Deep Transfer for Regression without Domain AlignmentPoster
  101. SemiVisBooster: Boosting Semi-Supervised Learning for Fine-Grained Classification through Pseudo-Label Semantic GuidancePoster
  102. SeqGrowGraph: Learning Lane Topology as a Chain of Graph ExpansionsPoster
  103. Sequential Gaussian Avatars with Hierarchical Motion ContextPoster
  104. Sequential keypoint density estimator: an overlooked baseline of skeleton-based video anomaly detectionPoster
  105. Serialization based Point Cloud OversegmentationPoster
  106. ShadowHack: Hacking Shadows via Luminance-Color Divide and ConquerPoster
  107. Shape of Motion: 4D Reconstruction from a Single VideoPoster
  108. ShortFT: Diffusion Model Alignment via Shortcut-based Fine-TuningPoster
  109. Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description GenerationPoster
  110. SiM3D: Single-instance Multiview Multimodal and Multisetup 3D Anomaly Detection BenchmarkPoster
  111. Sibai: A Few-Shot Meta-Classifier for Poisoning Detection in Federated LearningPoster
  112. SignRep: Enhancing Self-Supervised Sign RepresentationsPoster
  113. Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language GeneratorPoster
  114. Sim-DETR: Unlock DETR for Temporal Sentence GroundingPoster
  115. SimMLM: A Simple Framework for Multi-modal Learning with Missing ModalityPoster
  116. Similarity Memory Prior is All You Need for Medical Image SegmentationPoster
  117. SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language ModelsPoster
  118. Simulating Dual-Pixel Images From Ray Tracing For Depth EstimationPoster
  119. Simultaneous Motion And Noise Estimation with Event CamerasPoster
  120. Single-Scanline Relative Pose Estimation for Rolling Shutter CamerasPoster
  121. Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action SegmentationPoster
  122. SketchSplat: 3D Edge Reconstruction via Differentiable Multi-view Sketch SplattingPoster
  123. Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token SkippingPoster
  124. SkySense V2: A Unified Foundation Model for Multi-modal Remote SensingPoster
  125. SliderSpace: Decomposing the Visual Capabilities of Diffusion ModelsPoster
  126. SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversionPoster
  127. Snakes and Ladders: Two Steps Up for VideoMambaPoster
  128. Social Debiasing for Fair Multi-modal LLMsPoster
  129. Soft Local Completeness: Rethinking Completeness in XAI
  130. Soft Separation and Distillation: Toward Global Uniformity in Federated Unsupervised LearningPoster
  131. Sparfels: Fast Reconstruction from Sparse Unposed ImageryPoster
  132. Sparse Fine-Tuning of Transformers for Generative TasksPoster
  133. Sparse-Dense Side-Tuner for efficient Video Temporal GroundingPoster
  134. SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape ModelingPoster
  135. SparseLaneSTP: Leveraging Spatio-Temporal Priors with Sparse Transformers for 3D Lane DetectionPoster
  136. SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMsPoster
  137. SparseRecon: Neural Implicit Surface Reconstruction from Sparse Views with Feature and Depth ConsistenciesPoster
  138. SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
  139. Sparsity Outperforms Low-Rank Projections in Few-Shot AdaptationPoster
  140. Spatial Alignment and Temporal Matching Adapter for Video-Radar Remote Physiological MeasurementPoster
  141. Spatial Preference Rewarding for MLLMs Spatial UnderstandingPoster
  142. Spatial-Temporal Aware Visuomotor Diffusion Policy LearningPoster
  143. SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited ObservationsPoster
  144. SpatialSplat: Efficient Semantic 3D from Sparse Unposed ImagesPoster
  145. SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera MotionPoster
  146. Spatially-Varying AutofocusPoster
  147. Spatio-Spectral Pattern Illumination for Direct and Indirect Separation from a Single Hyperspectral ImagePoster
  148. SpecGuard: Spectral Projection-based Advanced Invisible WatermarkingPoster
  149. Spectral Image TokenizerPoster
  150. Spectral Sensitivity Estimation with an Uncalibrated Diffraction GratingPoster
  151. SpectralAR: Spectral Autoregressive Visual GenerationPoster
  152. Spherical Epipolar Rectification for Deep Two-View Absolute Depth EstimationPoster
  153. SpiLiFormer: Enhancing Spiking Transformers with Lateral InhibitionPoster
  154. SpikeDiff: Zero-shot High-Quality Video Reconstruction from Chromatic Spike Camera and Sub-millisecond Spike StreamsPoster
  155. SpikePack: Enhanced Information Flow in Spiking Neural Networks with High Hardware CompatibilityPoster
  156. SpinMeRound: Consistent Multi-View Identity Generation Using Diffusion ModelsPoster
  157. SplArt: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian SplattingPoster
  158. Splat-LOAM: Gaussian Splatting LiDAR Odometry and MappingPoster
  159. Splat-based 3D Scene Reconstruction with Extreme Motion-blurPoster
  160. SplatTalk: 3D VQA with Gaussian SplattingPoster
  161. Split-and-Combine: Enhancing Style Augmentation for Single Domain GeneralizationPoster
  162. St4RTrack: Simultaneous 4D Reconstruction and Tracking in the WorldPoster
  163. Stable Diffusion Models are Secretly Good at Visual In-Context LearningPoster
  164. Stable Score DistillationPoster
  165. Stable Virtual Camera: Generative View Synthesis with Diffusion ModelsPoster
  166. Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth DiffusionPoster
  167. StableCodec: Taming One-Step Diffusion for Extreme Image CompressionPoster
  168. StableDepth: Scene-Consistent and Scale-Invariant Monocular DepthPoster
  169. Staining and Locking Computer Vision Models Without RetrainingPoster
  170. Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View ImagesPoster
  171. StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided IllusionsPoster
  172. Stealthy Backdoor Attack in Federated Learning via Adaptive Layer-wise Gradient AlignmentPoster
  173. SteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric SteeringPoster
  174. Steering Guidance for Personalized Text-to-Image Diffusion ModelsPoster
  175. Stepping Out of Similar Semantic Space for Open-Vocabulary SegmentationPoster
  176. Stereo Any Video: Temporally Consistent Stereo MatchingPoster
  177. Stochastic Gradient Estimation for Higher-Order Differentiable RenderingPoster
  178. Stochastic Interpolants for Revealing Stylistic Flows across the History of Art
  179. StochasticSplats: Stochastic Rasterization for Sorting-Free 3D Gaussian SplattingPoster
  180. StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic DataPoster
  181. Straighten Viscous Rectified Flow via Noise OptimizationPoster
  182. StrandHead: Text to Hair-Disentangled 3D Head Avatars Using Human-Centric PriorsPoster
  183. StreamDiffusion: A Pipeline-level Solution for Real-Time Interactive GenerationPoster
  184. StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image StreamsPoster
  185. StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated CognitionPoster
  186. Streaming VideoLLMs for Real-Time Procedural Video UnderstandingPoster
  187. Streamlining Image Editing with Layered Diffusion BrushesPoster
  188. Street Gaussians without 3D Object TrackerPoster
  189. Stroke2Sketch: Harnessing Stroke Attributes for Training-Free Sketch GenerationPoster
  190. Stronger, Steadier & Superior: Geometric Consistency in Depth VFM Forges Domain Generalized Semantic SegmentationPoster
  191. Structure-Guided Diffusion Models for High-Fidelity Portrait Shadow RemovalPoster
  192. Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised LearningPoster
  193. Structured Policy Optimization: Enhance Large Vision-Language Model via Self-referenced DialoguePoster
  194. StyleKeeper: Prevent Content Leakage using Negative Visual Query GuidancePoster
  195. StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross FusionPoster
  196. StyleSRN: Scene Text Image Super-Resolution with Text Style EmbeddingPoster
  197. Stylized-Face: A Million-level Stylized Face Dataset for Face RecognitionPoster
  198. SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion ModelsPoster
  199. Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided DiffusionPoster
  200. SummDiff: Generative Modeling of Video Summarization with DiffusionPoster
  201. Super Resolved Imaging with Adaptive OpticsPoster
  202. SuperDec: 3D Scene Decomposition with Superquadrics PrimitivesPoster
  203. SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image EditingPoster
  204. SuperEvent: Cross-Modal Learning of Event-based Keypoint Detection for SLAMPoster
  205. SuperMat: Physically Consistent PBR Material Estimation at Interactive RatesPoster
  206. Supercharged One-step Text-to-Image Diffusion Models with Negative PromptsPoster
  207. Supercharging Floorplan Localization with Semantic RaysPoster
  208. Superpowering Open-Vocabulary Object Detectors for X-ray VisionPoster
  209. Supervised Exploratory Learning for Long-Tailed Visual RecognitionPoster
  210. SurfaceSplat: Connecting Surface Reconstruction and Gaussian SplattingPoster
  211. SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video DiscretizationPoster
  212. Switch-a-View: View Selection Learned from Unlabeled In-the-wild VideosPoster
  213. SynAD: Enhancing Real-World End-to-End Autonomous Driving Models through Synthetic Data IntegrationPoster
  214. SynCity: Training-Free Generation of 3D WorldsPoster
  215. SynFER: Towards Boosting Facial Expression Recognition with Synthetic DataPoster
  216. SynTag: Enhancing the Geometric Robustness of Inversion-based Generative Image WatermarkingPoster
  217. SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction SynthesisPoster
  218. Synchronization of Multiple VideosPoster
  219. Synchronizing Task Behavior: Aligning Multiple Tasks during Test-Time TrainingPoster
  220. Synergistic Prompting for Robust Visual Recognition with Missing ModalitiesPoster
  221. Synthesizing Near-Boundary OOD Samples for Out-of-Distribution DetectionPoster
  222. Synthetic Video Enhances Physical Fidelity in Video SynthesisPoster
  223. T2Bs: Text-to-Character Blendshapes via Video GenerationPoster
  224. T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive GenerationPoster
  225. TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language ModelsPoster
  226. TACO: Taming Diffusion for in-the-wild Video Amodal CompletionPoster
  227. TAD-E2E: A Large-scale End-to-end Autonomous Driving DatasetPoster
  228. TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion SensitivityPoster
  229. TAPNext: Tracking Any Point (TAP) as Next Token PredictionPoster
  230. TAR3D: Creating High-Quality 3D Assets via Next-Part PredictionPoster
  231. TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio SynthesisPoster
  232. TARS: Traffic-Aware Radar Scene Flow EstimationPoster
  233. TCFG: Truncated Classifier-Free Guidance for Efficient and Scalable Text-to-Image AccelerationPoster
  234. TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event CamerasPoster
  235. TF-TI2I: Training-Free Text-and-Image-to-Image Generation via Multi-Modal Implicit-Context Learning In Text-to-Image ModelsPoster
  236. TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video GenerationPoster
  237. TITAN-Guide: Taming Inference-Time Alignment for Guided Text-to-Video Diffusion ModelsPoster
  238. TITAN: Query-Token based Domain Adaptive Adversarial LearningPoster
  239. TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame InterpolationPoster
  240. TOGA: Temporally Grounded Open-Ended Video QA with Weak SupervisionPoster
  241. TOTP: Transferable Online Pedestrian Trajectory Prediction with Temporal-Adaptive Mamba Latent DiffusionPoster
  242. TPG-INR: Target Prior-Guided Implicit 3D CT Reconstruction for Enhanced Sparse-view ImagingPoster
  243. TR-PTS: Task-Relevant Parameter and Token Selection for Efficient TuningPoster
  244. TRACE: Learning 3D Gaussian Physical Dynamics from Multi-view VideosPoster
  245. TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion ModelsPoster
  246. TREAD: Token Routing for Efficient Architecture-agnostic Diffusion TrainingPoster
  247. TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-enhanced Relation-aware Knowledge TransferringPoster
  248. TRNAS: A Training-Free Robust Neural Architecture SearchPoster
  249. TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free TuningPoster
  250. Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationPoster
  251. Taming Flow Matching with Unbalanced Optimal Transport into Fast PansharpeningPoster
  252. Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the UnknownPoster
  253. Target Bias Is All You Need: Zero-Shot Debiasing of Vision-Language Models with Bias CorpusPoster
  254. Task Vector Quantization for Memory-Efficient Model MergingPoster
  255. Task-Aware Prompt Gradient Projection for Parameter-Efficient Tuning Federated Class-Incremental LearningPoster
  256. Task-Decoupled Bezier Surface Constraint for Uneven Low-Light Image EnhancementPoster
  257. Task-Oriented Human Grasp Synthesis via Context- and Task-Aware DiffusersPoster
  258. Task-Specific Zero-shot Quantization-Aware Training for Object DetectionPoster
  259. TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species GenerationPoster
  260. TeEFusion: Blending Text Embeddings to Distill Classifier-Free GuidancePoster
  261. TeRA: Rethinking Text-guided Realistic 3D Avatar GenerationPoster
  262. Teaching AI the Anatomy Behind the Scan: Addressing Anatomical Flaws in Medical Image Segmentation with Learnable PriorPoster
  263. Teaching VLMs to Localize Specific Objects from In-context ExamplesPoster
  264. Teeth Reconstruction and Performance Capture Using a Phone CameraPoster
  265. Teleportraits: Training-Free People Insertion into Any ScenePoster
  266. Temperature in Cosine-based Softmax LossPoster
  267. Temporal Overlapping Prediction: A Self-supervised Pre-training Method for LiDAR Moving Object SegmentationPoster
  268. Temporal Rate Reduction Clustering for Human Motion SegmentationPoster
  269. Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object TrackingPoster
  270. Temporal-aware Query Routing for Real-time Video Instance SegmentationPoster
  271. Tensor-aggregated LoRA in Federated Fine-tuningPoster
  272. TerraMind: Large-Scale Generative Multimodality for Earth ObservationPoster
  273. Test-Time Prompt Tuning for Zero-Shot Depth CompletionPoster
  274. Test-Time Retrieval-Augmented Adaptation for Vision-Language ModelsPoster
  275. Test-time Adaptation for Foundation Medical Segmentation Model Without Parametric UpdatesPoster
  276. Text Embedding Knows How to Quantize Text-Guided Diffusion ModelsPoster
  277. Text-IRSTD: Leveraging Semantic Text to Promote Infrared Small Target Detection in Complex ScenesPoster
  278. Text-guided Visual Prompt DINO for Generic SegmentationPoster
  279. Text-to-Any-Skeleton Motion Generation Without RetargetingPoster
  280. Text2Outfit: Controllable Outfit Generation with Multimodal Language ModelsPoster
  281. Text2VDM: Text to Vector Displacement Maps for Expressive and Interactive 3D SculptingPoster
  282. TextMaster: A Unified Framework for Realistic Text Editing via Glyph-Style Dual-ControlPoster
  283. TextSSR: Diffusion-based Data Synthesis for Scene Text RecognitionPoster
  284. Textured 3D Regenerative Morphing with 3D Diffusion PriorPoster
  285. The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video GenerationPoster
  286. The Curse of Conditions: Analyzing and Improving Optimal Transport for Conditional Flow-Based GenerationPoster
  287. The Devil is in the Spurious Correlations: Boosting Moment Retrieval with Dynamic LearningPoster
  288. The Inter-Intra Modal Measure: A Predictive Lens on Fine-Tuning Outcomes in Vision-Language ModelsPoster
  289. The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single TransformerPoster
  290. The Silent Assistant: NoiseQuery as Implicit Guidance for Goal-Driven Image GenerationPoster
  291. The Source Image is the Best Attention for Infrared and Visible Image FusionPoster
  292. Thermal Polarimetric Multi-view StereoPoster
  293. Think Twice: Test-Time Reasoning for Robust CLIP Zero-Shot ClassificationPoster
  294. TikZero: Zero-Shot Text-Guided Graphics Program SynthesisPoster
  295. Tile-wise vs. Image-wise: Random-Tile Loss and Training Paradigm for Gaussian SplattingPoster
  296. Tiling artifacts and trade-offs of feature normalization in the segmentation of large biological imagesPoster
  297. Time-Aware Auto White Balance in Mobile PhotographyPoster
  298. TimeBooth: Disentangled Facial Invariant Representation for Diverse and Personalized Face AgingPoster
  299. TimeExpert: An Expert-Guided Video LLM for Video Temporal GroundingPoster
  300. TimeFormer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust ReconstructionPoster
  301. Timestep-Aware Diffusion Model for Extreme Image RescalingPoster
  302. TinyViM: Frequency Decoupling for Tiny Hybrid Vision MambaPoster
  303. To Label or Not to Label: PALM - A Predictive Model for Evaluating Sample Efficiency in Active Learning ModelsPoster
  304. ToF-Splatting: Dense SLAM using Sparse Time-of-Flight Depth and Multi-Frame IntegrationPoster
  305. Token Activation Map to Visually Explain Multimodal LLMsPoster
  306. Token-Efficient VLM: High-Resolution Image Understanding via Dynamic Region ProposalPoster
  307. TokensGen: Harnessing Condensed Tokens for Long Video GenerationPoster
  308. ToolVQA: A Dataset for Multi-step Reasoning VQA with External ToolsPoster
  309. Top2Pano: Learning to Generate Indoor Panoramas from Top-Down ViewPoster
  310. TopicGeo: An Efficient Unified Framework for GeolocationPoster
  311. TorchAdapt: Towards Light-Agnostic Real-Time Visual PerceptionPoster
  312. Toward Better Out-painting: Improving the Image Composition with Initialization Policy ModelPoster
  313. Toward Fair and Accurate Cross-Domain Medical Image Segmentation: A VLM-Driven Active Domain Adaptation ParadigmPoster
  314. Toward Long-Tailed Online Anomaly Detection through Class-Agnostic ConceptsPoster
  315. Toward Material-Agnostic System Identification from VideosPoster
  316. Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on EdgePoster
  317. Towards Adversarial Robustness via Debiased High-Confidence Logit AlignmentPoster
  318. Towards Annotation-Free Evaluation: KPAScore for Human Keypoint DetectionPoster
  319. Towards Comprehensive Lecture Slides Understanding: Large-scale Dataset and Effective MethodPoster
  320. Towards Cross-modal Backward-compatible Representation Learning for Vision-Language ModelsPoster
  321. Towards Effective Foundation Model Adaptation for Extreme Cross-Domain Few-Shot LearningPoster
  322. Towards Efficient General Feature Prediction in Masked Skeleton ModelingPoster
  323. Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human AvatarsPoster
  324. Towards Fine-grained Interactive Segmentation in Images and VideosPoster
  325. Towards Foundational Models for Single-Chip RadarPoster
  326. Towards Higher Effective Rank in Parameter-Efficient Fine-tuning using Khatri-Rao ProductPoster
  327. Towards Human-like Virtual Beings: Simulating Human Behavior in 3D ScenesPoster
  328. Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion SynthesisPoster
  329. Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and MemoryPoster
  330. Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled ViewsPoster
  331. Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual SegmentationPoster
  332. Towards Open-World Generation of Stereo Images and Unsupervised MatchingPoster
  333. Towards Performance Consistency in Multi-Level Model CollaborationPoster
  334. Towards Real Unsupervised Anomaly Detection Via Confident Meta-LearningPoster
  335. Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based PurificationPoster
  336. Towards Robustness of Person Search against CorruptionsPoster
  337. Towards Safer and Understandable Driver Intention PredictionPoster
  338. Towards Scalable Spatial Intelligence via 2D-to-3D Data LiftingPoster
  339. Towards Stabilized and Efficient Diffusion Transformers through Long-Skip-Connections with Spectral ConstraintsPoster
  340. Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
  341. Towards Visual Localization Interoperability: Cross-Feature for Collaborative Visual Localization and MappingPoster
  342. Towards a 3D Transfer-based Black-box Attack via Critical Feature GuidancePoster
  343. Towards a Unified Copernicus Foundation Model for Earth VisionPoster
  344. Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant RepresentationPoster
  345. Towards a Universal Image Degradation Model via Content-Degradation DisentanglementPoster
  346. Trace3D: Consistent Segmentation Lifting via Gaussian Instance TracingPoster
  347. Tracing Copied Pixels and Regularizing Patch Affinity in Copy DetectionPoster
  348. TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud TrackingPoster
  349. TrackVerse: A Large-Scale Object-Centric Video Dataset for Image-Level Representation Learning
  350. Tracking Tiny Drones against Clutter: Large-Scale Infrared Benchmark with Motion-Centric Adaptive AlgorithmPoster
  351. Trade-offs in Image Generation: How Do Different Dimensions Interact?Poster
  352. TrafficLoc: Localizing Traffic Surveillance Cameras in 3D ScenesPoster
  353. Training-Free Class Purification for Open-Vocabulary Semantic SegmentationPoster
  354. Training-Free Industrial Defect Generation with Diffusion ModelsPoster
  355. Training-Free Personalization via Retrieval and Reasoning on FingerprintsPoster
  356. Training-Free Text-Guided Image Editing with Visual Autoregressive ModelPoster
  357. Training-free Generation of Temporally Consistent Rewards from VLMsPoster
  358. Training-free Geometric Image Editing on Diffusion ModelsPoster
  359. Training-free and Adaptive Sparse Attention for Efficient Long Video GenerationPoster
  360. TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion ModelsPoster
  361. Trans-Adapter: A Plug-and-Play Framework for Transparent Image InpaintingPoster
  362. Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image ModelsPoster
  363. Transformer-based Tooth Alignment Prediction with Occlusion and Collision ConstraintsPoster
  364. TransiT: Transient Transformer for Non-line-of-sight VideographyPoster
  365. Translation of Text Embedding via Delta Vector to Suppress Strongly Entangled Content in Text-to-Image Diffusion ModelsPoster
  366. Transparent Vision: A Theory of Hierarchical Invariant RepresentationsPoster
  367. Tree Skeletonization from 3D Point Clouds by Denoising DiffusionPoster
  368. Tree-NeRV: Efficient Non-Uniform Sampling for Neural Video Representation via Tree-Structured Feature GridsPoster
  369. TriDi: Trilateral Diffusion of 3D Humans, Objects, and InteractionsPoster
  370. Triad: Empowering LMM-based Anomaly Detection with Expert-guided Region-of-Interest Tokenizer and Manufacturing ProcessPoster
  371. Trial-Oriented Visual RearrangementPoster
  372. Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action RecognitionPoster
  373. Trust but Verify: Programmatic VLM Evaluation in the WildPoster
  374. TrustMark: Robust Watermarking and Watermark Removal for Arbitrary Resolution ImagesPoster
  375. TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-InterventionPoster
  376. TryOn-Refiner: Conditional Rectified-flow-based TryOn Refiner for More Accurate Detail ReconstructionPoster
  377. Tune-Your-Style: Intensity-tunable 3D Style Transfer with Gaussian SplattingPoster
  378. Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video SynthesisPoster
  379. TurboReg: TurboClique for Robust and Efficient Point Cloud RegistrationPoster
  380. TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and PredictionPoster
  381. TurboVSR: Fantastic Video Upscalers and Where to Find ThemPoster
  382. Two Losses, One Goal: Balancing Conflict Gradients for Semi-supervised Semantic SegmentationPoster
  383. U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and RegistrationPoster
  384. UAVScenes: A Multi-Modal Dataset for UAVsPoster
  385. UDC-VIT: A Real-World Video Dataset for Under-Display CamerasPoster
  386. UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital AgentsPoster
  387. UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility ConstraintPoster
  388. UIPro: Unleashing Superior Interaction Capability For GUI AgentsPoster
  389. UKBOB: One Billion MRI Labeled Masks for Generalizable 3D Medical Image SegmentationPoster
  390. ULTHO: Ultra-Lightweight yet Efficient Hyperparameter Optimization in Deep Reinforcement LearningPoster
  391. UMDATrack: Unified Multi-Domain Adaptive Tracking Under Adverse Weather ConditionsPoster
  392. UNIS: A Unified Framework for Achieving Unbiased Neural Implicit Surfaces in Volume RenderingPoster
  393. UPP: Unified Point-Level Prompting for Robust Point Cloud AnalysisPoster
  394. USP: Unified Self-Supervised Pretraining for Image Generation and UnderstandingPoster
  395. Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency AdapterPoster
  396. Ultra-Precision 6DoF Pose Estimation Using 2-D Interpolated Discrete Fourier TransformPoster
  397. UnMix-NeRF: Spectral Unmixing Meets Neural Radiance FieldsPoster
  398. UnZipLoRA: Separating Content and Style from a Single ImagePoster
  399. Unbiased Missing-modality Multimodal LearningPoster
  400. Unbiased Region-Language Alignment for Open-Vocabulary Dense PredictionPoster
  401. Uncalibrated Structure from Motion on a SpherePoster
  402. Uncertainty-Aware Diffusion-Guided Refinement of 3D ScenesPoster
  403. Uncertainty-Aware Gradient Stabilization for Small Object DetectionPoster
  404. Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language ModelsPoster
  405. Uncover Treasures in DCT: Advancing JPEG Quality Enhancement by Exploiting Latent CorrelationsPoster
  406. Understanding Co-speech Gestures in-the-wildPoster
  407. Understanding Flatness in Generative Models: Its Role and BenefitsPoster
  408. Understanding Museum Exhibits using Vision-Language ReasoningPoster
  409. Understanding Personal Concept in Open-Vocabulary Semantic SegmentationPoster
  410. Underwater Visual SLAM with Depth Uncertainty and Medium ModelingPoster
  411. Unfolding-Associative Encoder-Decoder Network with Progressive Alignment for PansharpeningPoster
  412. UniCombine: Unified Multi-Conditional Combination with Diffusion TransformerPoster
  413. UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any ScalePoster
  414. UniDxMD: Towards Unified Representation for Cross-Modal Unsupervised Domain Adaptation in 3D Semantic SegmentationPoster
  415. UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and GenerationPoster
  416. UniFuse: A Unified All-in-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and MisalignmentsPoster
  417. UniGS: Modeling Unitary 3D Gaussians for Novel View Synthesis from Sparse-view ImagesPoster
  418. UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text SynthesisPoster
  419. UniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous DrivingPoster
  420. UniOcc: A Unified Benchmark for Occupancy Forecasting and Prediction in Autonomous DrivingPoster
  421. UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character ControlPoster
  422. UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image PersonalizationPoster
  423. UniRes: Universal Image Restoration for Complex DegradationsPoster
  424. UniVG: A Generalist Diffusion Model for Unified Image Generation and EditingPoster
  425. UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field ReconstructionPoster
  426. Unified Adversarial Augmentation for Improving Palmprint RecognitionPoster
  427. Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D PrototypesPoster
  428. Unified Multi-Agent Trajectory Modeling with Masked Trajectory DiffusionPoster
  429. Unified Multimodal Understanding via Byte-Pair Visual EncodingPoster
  430. Unified Open-World Segmentation with Multi-Modal PromptsPoster
  431. Unified Video Generation via Next-Set Prediction in Continuous DomainPoster
  432. UniversalBooth: Model-Agnostic Personalized Text-to-Image GenerationPoster
  433. Unknown Text Learning for CLIP-based Few-Shot Open-set RecognitionPoster
  434. Unlearning the Noisy Correspondence Makes CLIP More RobustPoster
  435. Unleashing High-Quality Image Generation in Diffusion Sampling Using Second-Order Levenberg-Marquardt-LangevinPoster
  436. Unleashing Vecset Diffusion Model for Fast Shape GenerationPoster
  437. Unleashing the Temporal Potential of Stereo Event Cameras for Continuous-Time 3D Object DetectionPoster
  438. Unlocking Constraints: Source-Free Occlusion-Aware Seamless SegmentationPoster
  439. Unlocking the Potential of Diffusion Priors in Blind Face RestorationPoster
  440. Unraveling the Effects of Synthetic Data on End-to-End Autonomous DrivingPoster
  441. Unraveling the Smoothness Properties of Diffusion Models: A Gaussian Mixture PerspectivePoster
  442. UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AIPoster
  443. Unsupervised Histopathological Image Semantic Segmentation with Overlapping Patches Consistency ConstraintPoster
  444. Unsupervised Identification of Protein Compositions and Conformations via Implicit Content-Transformation DisentanglementPoster
  445. Unsupervised Imaging Inverse Problems with Diffusion Distribution MatchingPoster
  446. Unsupervised Joint Learning of Optical Flow and Intensity with Event CamerasPoster
  447. Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized ConstraintsPoster
  448. Unsupervised RGB-D Point Cloud Registration for Scenes with Low Overlap and Photometric InconsistencyPoster
  449. Unsupervised Visible-Infrared Person Re-identification under Unpaired SettingsPoster
  450. Unsupervised Visual Chain-of-Thought Reasoning via Preference OptimizationPoster
  451. Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURAPoster
  452. UrbanLLaVA: A Multi-modal Large Language Model for Urban IntelligencePoster
  453. V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion ModelsPoster
  454. V2M4: 4D Mesh Animation Reconstruction from a Single Monocular VideoPoster
  455. V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position EncodingPoster
  456. V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and PredictionPoster
  457. V2XScenes: A Multiple Challenging Traffic Conditions Dataset for Large-Range Vehicle-Infrastructure Collaborative PerceptionPoster
  458. VA-MoE: Variables-Adaptive Mixture of Experts for Incremental Weather ForecastingPoster
  459. VACE: All-in-One Video Creation and EditingPoster
  460. VAFlow: Video-to-Audio Generation with Cross-Modality Flow MatchingPoster
  461. VAGUE: Visual Contexts Clarify Ambiguous ExpressionsPoster
  462. VALLR: Visual ASR Language Model for Lip ReadingPoster
  463. VCA: Video Curious Agent for Long Video UnderstandingPoster
  464. VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided OptimizationPoster
  465. VGGSounder: Audio-Visual Evaluations for Foundation ModelsPoster
  466. VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual GroundingPoster
  467. VIGFace: Virtual Identity Generation for Privacy-Free Face Recognition DatasetPoster
  468. VIPerson: Flexibly Generating Virtual Identity for Person Re-IdentificationPoster
  469. VISION-XL: High Definition Video Inverse Problem Solver using Latent Image Diffusion ModelsPoster
  470. VISO: Accelerating In-orbit Object Detection with Language-Guided Mask Learning and Sparse InferencePoster
  471. VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information FlowPoster
  472. VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksPoster
  473. VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous DrivingPoster
  474. VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
  475. VLM4D: Towards Spatiotemporal Awareness in Vision Language ModelsPoster
  476. VLR-Driver: Large Vision-Language-Reasoning Models for Embodied Autonomous DrivingPoster
  477. VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsPoster
  478. VMBench: A Benchmark for Perception-Aligned Video Motion GenerationPoster
  479. VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
  480. VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation under real OcclusionsPoster
  481. VPO: Aligning Text-to-Video Generation Models with Prompt OptimizationPoster
  482. VPR-Cloak: A First Look at Privacy Cloak Against Visual Place RecognitionPoster
  483. VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative VideosPoster
  484. VRM: Knowledge Distillation via Virtual Relation MatchingPoster
  485. VSC: Visual Search Compositional Text-to-Image Diffusion ModelPoster
  486. VSP: Diagnosing the Dual Challenges of Perception and Reasoning in Spatial Planning Tasks for MLLMsPoster
  487. VSRM: A Robust Mamba-Based Framework for Video Super-ResolutionPoster
  488. VSSD: Vision Mamba with Non-Causal State Space DualityPoster
  489. VTimeCoT: Thinking by Drawing for Video Temporal Grounding and ReasoningPoster
  490. Vamba: Understanding Hour-Long Videos with Hybrid Mamba-TransformersPoster
  491. Variance-Based Pruning for Accelerating and Compressing Trained NetworksPoster
  492. Vector Contrastive Learning For Pixel-Wise Pretraining In Medical VisionPoster
  493. VehicleMAE: View-asymmetry Mutual Learning for Vehicle Re-identification Pre-training via Masked AutoEncodersPoster
  494. Verbalized Representation Learning for Interpretable Few-Shot GeneralizationPoster
  495. Versatile Transition Generation with Image-to-Video DiffusionPoster
  496. VertexRegen: Mesh Generation with Continuous Level of DetailPoster
  497. ViCTr: Vital Consistency Transfer for Pathology Aware Image SynthesisPoster
  498. ViLLa: Video Reasoning Segmentation with Large Language ModelPoster
  499. ViLU: Learning Vision-Language Uncertainties for Failure PredictionPoster
  500. ViM-VQ: Efficient Post-Training Vector Quantization for Visual MambaPoster
  501. ViSpeak: Visual Instruction Feedback in Streaming VideosPoster
  502. ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision TransformersPoster
  503. ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision ModelsPoster
  504. ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting HeadsPoster
  505. Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the WildPoster
  506. Video Color Grading via Look-Up Table GenerationPoster
  507. Video Individual Counting for Moving DronesPoster
  508. Video Motion GraphsPoster
  509. Video-T1: Test-time Scaling for Video GenerationPoster
  510. Video2BEV: Transforming Drone Videos to BEVs for Video-based Geo-localizationPoster
  511. VideoAds for Fast-Paced Video Understanding
  512. VideoAuteur: Towards Long Narrative Video GenerationPoster
  513. VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory BridgesPoster
  514. VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy OptimizationPoster
  515. VideoOrion: Tokenizing Object Dynamics in VideosPoster
  516. VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint ModelingPoster
  517. VideoSetDiff: Identifying and Reasoning Similarities and Differences in Similar VideosPoster
  518. VideoVAE+: Large Motion Video Autoencoding with Cross-modal Video VAEPoster
  519. ViewSRD: 3D Visual Grounding via Structured Multi-View DecompositionPoster
  520. VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible RegionsPoster
  521. VisNumBench: Evaluating Number Sense of Multimodal Large Language ModelsPoster
  522. VisRL: Intention-Driven Visual Perception via Reinforced ReasoningPoster
  523. Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph GenerationPoster
  524. Vision-Language Models Can't See the ObviousPoster
  525. Vision-Language Neural Graph Featurization for Extracting Retinal LesionsPoster
  526. VisionMath: Vision-Form Mathematical Problem-SolvingPoster
  527. VistaDream: Sampling multiview consistent images for single-view scene reconstructionPoster
  528. Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of ImagesPoster
  529. Visual Intention Grounding for Egocentric AssistantsPoster
  530. Visual Interestingness Decoded: How GPT-4o Mirrors Human InterestsPoster
  531. Visual Modality Prompt for Adapting Vision-Language Object DetectorsPoster
  532. Visual Relation Diffusion for Human-Object Interaction DetectionPoster
  533. Visual Surface Wave Elastography: Revealing Subsurface Physical Properties via Visible Surface WavesPoster
  534. Visual Test-time Scaling for GUI Agent GroundingPoster
  535. Visual Textualization for Image Prompted Object DetectionPoster
  536. Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language ModelsPoster
  537. VisualCloze: A Universal Image Generation Framework via Visual In-Context LearningPoster
  538. Vivid4D: Improving 4D Reconstruction from Monocular Video by Video InpaintingPoster
  539. VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language ModelsPoster
  540. VoluMe - Authentic 3D Video Calls from Live Gaussian Splat PredictionPoster
  541. VolumetricSMPL: A Neural Volumetric Body Model for Efficient Interactions, Contacts, and CollisionsPoster
  542. VoteSplat: Hough Voting Gaussian Splatting for 3D Scene UnderstandingPoster
  543. VoxelKP: A Voxel-based Network Architecture for Human Keypoint Estimation in LiDAR DataPoster
  544. Voyaging into Perpetual Dynamic Scenes from a Single ViewPoster
  545. Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video DetectionPoster
  546. WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single ImagePoster
  547. WINS: Winograd Structured Pruning for Fast Winograd ConvolutionPoster
  548. WIPES: Wavelet-based Visual PrimitivesPoster
  549. WIR3D: Visually-Informed and Geometry-Aware 3D Shape AbstractionPoster
  550. WSI-LLaVA: A Multimodal Large Language Model for Whole Slide ImagePoster
  551. WalkVLM: Aid Visually Impaired People Walking by Vision Language ModelPoster
  552. WarpHE4D: Dense 4D Head Map toward Full Head ReconstructionPoster
  553. Wasserstein Style Distribution Analysis and Transform for Stylized Image GenerationPoster
  554. Wave-MambaAD: Wavelet-driven State Space Model for Multi-class Unsupervised Anomaly DetectionPoster
  555. WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object DetectionPoster
  556. Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon TasksPoster
  557. Weakly Supervised Visible-Infrared Person Re-Identification via Heterogeneous Expert Collaborative Consistency LearningPoster
  558. Weakly-Supervised Learning of Dense Functional CorrespondencesPoster
  559. WeaveSeg: Iterative Contrast-weaving and Spectral Feature-refining for Nuclei Instance SegmentationPoster
  560. Web Artifact Attacks Disrupt Vision Language ModelsPoster
  561. What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation LearningPoster
  562. What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language ModelsPoster
  563. What If: Understanding Motion Through Sparse InteractionsPoster
  564. What Makes for Text to 360-degree Panorama Generation with Stable Diffusion?Poster
  565. What You Have is What You Track: Adaptive and Robust Multimodal TrackingPoster
  566. What to Distill? Fast Knowledge Distillation with Adaptive SamplingPoster
  567. What we need is explicit controllability: Training 3D gaze estimator using only facial imagesPoster
  568. What's Making That Sound Right Now? Video-centric Audio-Visual LocalizationPoster
  569. What's in a Latent? Leveraging Diffusion Latent Space for Domain GeneralizationPoster
  570. When Anchors Meet Cold Diffusion: A Multi-Stage Approach to Lane DetectionPoster
  571. When Confidence Fails: Revisiting Pseudo-Label Selection in Semi-supervised Semantic SegmentationPoster
  572. When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token PruningPoster
  573. When Pixel Difference Patterns Meet ViT: PiDiViT for Few-Shot Object DetectionPoster
  574. When Schrodinger Bridge Meets Real-World Image Dehazing with Unpaired TrainingPoster
  575. When and Where do Data Poisons Attack Textual Inversion?Poster
  576. Where am I? Cross-View Geo-localization with Natural Language DescriptionsPoster
  577. Where, What, Why: Towards Explainable Driver Attention PredictionPoster
  578. Who Controls the Authorization? Invertible Networks for Copyright Protection in Text-to-Image SynthesisPoster
  579. Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking HeadsPoster
  580. Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of ContextPoster
  581. Wide2Long: Learning Lens Compression and Perspective Adjustment for Wide-Angle to Telephoto TranslationPoster
  582. WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article GenerationPoster
  583. WildSAT: Learning Satellite Image Representations from Wildlife ObservationsPoster
  584. WildSeg3D: Segment Any 3D Objects in the Wild from 2D ImagesPoster
  585. WonderPlay: Dynamic 3D Scene Generation from a Single Image and ActionsPoster
  586. WonderTurbo: Generating Interactive 3D World in 0.72 SecondsPoster
  587. World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World ModelPoster
  588. WorldScore: A Unified Evaluation Benchmark for World GenerationPoster
  589. X-Capture: An Open-Source Portable Device for Multi-Sensory LearningPoster
  590. X-Dancer: Expressive Music to Human Dance Video GenerationPoster
  591. X-Fusion: Introducing New Modality to Frozen Large Language ModelsPoster
  592. X-Prompt: Generalizable Auto-Regressive Visual Learning with In-Context PromptingPoster
  593. X2-Gaussian: 4D Radiative Gaussian Splatting for Continuous-time Tomographic ReconstructionPoster
  594. X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention DistillationPoster
  595. XTrack: Multimodal Training Boosts RGB-X Video Object TrackersPoster
  596. YOLO-Count: Differentiable Object Counting for Text-to-Image GenerationPoster
  597. YOLOE: Real-Time Seeing AnythingPoster
  598. You Are Your Own Best Teacher: Achieving Centralized-level Performance in Federated Learning under Heterogeneous and Long-tailed DataPoster
  599. You Share Beliefs, I Adapt: Progressive Heterogeneous Collaborative PerceptionPoster
  600. You Think, You ACT: The New Task of Arbitrary Text to Motion GenerationPoster
  601. Your Text Encoder Can Be An Object-Level Watermarking ControllerPoster
  602. ZFusion: Efficient Deep Compositional Zero-shot Learning for Blind Image Super-Resolution with Generative Diffusion PriorPoster
  603. ZIM: Zero-Shot Image Matting for AnythingPoster
  604. ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned ModelsPoster
  605. Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech RepresentationsPoster
  606. Zero-Shot Composed Image Retrieval via Dual-Stream Instruction-Aware DistillationPoster
  607. Zero-Shot Compositional Video Learning with Coding Rate ReductionPoster
  608. Zero-Shot Depth Aware Image Editing with Diffusion ModelsPoster
  609. Zero-Shot Vision Encoder Grafting via LLM SurrogatesPoster
  610. Zero-shot Inexact CAD Model Alignment from a Single ImagePoster
  611. ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language ModelsPoster
  612. Zeroth-Order Fine-Tuning of LLMs in Random SubspacesPoster
  613. ZipVL: Accelerating Vision-Language Models through Dynamic Token SparsityPoster
  614. egoPPG: Heart Rate Estimation from Eye-Tracking Cameras in Egocentric Systems to Benefit Downstream Vision TasksPoster
  615. iManip: Skill-Incremental Learning for Robotic ManipulationPoster
  616. kh: Symmetry Understanding of 3D Shapes via Chirality DisentanglementPoster
  617. mmCooper: A Multi-agent Multi-stage Communication-efficient and Collaboration-robust Cooperative Perception FrameworkPoster
  618. monoVLN: Bridging the Observation Gap between Monocular and Panoramic Vision and Language NavigationPoster
  619. p-AVAS: Can Physics-Integrated Audio-Visual Modeling Boost Neural Acoustic Synthesis?Poster
  620. p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio DecayPoster

Looking for submission deadlines instead? See the conference deadline calendar.