← All conferences

CVPR 2025 Accepted Papers

The full list of 2,872 papers accepted at CVPR 2025 (IEEE/CVF Conference on Computer Vision and Pattern Recognition). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.

Poster: 2,468Highlight: 388Award Candidate: 15
  1. MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual RepresentationsPoster1 citations
  2. MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized DataPoster1 citations
  3. MeshArt: Generating Articulated Meshes with Structure-Guided TransformersPoster1 citations
  4. MixerMDM: Learnable Composition of Human Motion Diffusion ModelsPoster1 citations
  5. MoEE: Mixture of Emotion Experts for Audio-Driven Portrait AnimationPoster1 citations
  6. MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual EncodersPoster1 citations
  7. Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identificationHighlight1 citations
  8. Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsAward Candidate1 citations
  9. Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D SegmentationPoster1 citations
  10. Motion Modes: What Could Happen Next?Poster1 citations
  11. Multi-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain AdaptationPoster1 citations
  12. Multi-view Reconstruction via SfM-guided Monocular Depth EstimationPoster1 citations
  13. MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video RetrievalPoster1 citations
  14. Multitwine: Multi-Object Compositing with Text and Layout ControlHighlight1 citations
  15. NTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on ThermodynamicsPoster1 citations
  16. NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed ImagesPoster1 citations
  17. Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level CaptionsPoster1 citations
  18. Nested Diffusion Models Using Hierarchical Latent PriorsPoster1 citations
  19. NoPain: No-box Point Cloud Attack via Optimal Transport Singular BoundaryPoster1 citations
  20. NoT: Federated Unlearning via Weight NegationPoster1 citations
  21. Novel View Synthesis with Pixel-Space Diffusion ModelsPoster1 citations
  22. ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language ModelsPoster1 citations
  23. OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIPPoster1 citations
  24. ObjectMover: Generative Object Movement with Video PriorPoster1 citations
  25. Olympus: A Universal Task Router for Computer Vision TasksHighlight1 citations
  26. On the Generalization of Handwritten Text Recognition ModelsPoster1 citations
  27. Open-World Objectness Modeling Unifies Novel Object DetectionPoster1 citations
  28. Optical-Flow Guided Prompt Optimization for Coherent Video GenerationPoster1 citations
  29. PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language ModelsPoster1 citations
  30. PCDreamer: Point Cloud Completion Through Multi-view Diffusion PriorsPoster1 citations
  31. PEACE: Empowering Geologic Map Holistic Understanding with MLLMsPoster1 citations
  32. PERSE: Personalized 3D Generative Avatars from A Single PortraitPoster1 citations
  33. PICO: Reconstructing 3D People In Contact with ObjectsPoster1 citations
  34. PO3AD: Predicting Point Offsets toward Better 3D Point Cloud Anomaly DetectionPoster1 citations
  35. PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance PredictionPoster1 citations
  36. Panorama Generation From NFoV Image Done RightHighlight1 citations
  37. Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local PerceptionPoster1 citations
  38. Pay Attention to the Foreground in Object-Centric LearningPoster1 citations
  39. PersonaBooth: Personalized Text-to-Motion GenerationPoster1 citations
  40. Personalized Preference Fine-tuning of Diffusion ModelsPoster1 citations
  41. PhysicsGen: Can Generative Models Learn from Images to Predict Complex Physical Relations?Poster1 citations
  42. Pippo: High-Resolution Multi-View Humans from a Single ImageHighlight1 citations
  43. Point Clouds Meets Physics: Dynamic Acoustic Field Fitting Network for Point Cloud UnderstandingPoster1 citations
  44. Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among InstancesPoster1 citations
  45. Preconditioners for the Stochastic Training of Neural FieldsPoster1 citations
  46. Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge ModelsPoster1 citations
  47. ProAPO: Progressively Automatic Prompt Optimization for Visual ClassificationPoster1 citations
  48. ProReflow: Progressive Reflow with Decomposed VelocityPoster1 citations
  49. Probability Density Geodesics in Image Diffusion Latent SpacePoster1 citations
  50. PromptHash:Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing RetrievalPoster1 citations
  51. Protecting Your Video Content: Disrupting Automated Video-based LLM AnnotationsPoster1 citations
  52. ProtoDepth: Unsupervised Continual Depth Completion with PrototypesPoster1 citations
  53. Q-Bench-Video: Benchmark the Video Quality Understanding of LMMsPoster1 citations
  54. Quantization without TearsPoster1 citations
  55. RAD: Region-Aware Diffusion Models for Image InpaintingPoster1 citations
  56. RAP: Retrieval-Augmented Personalization for Multimodal Large Language ModelsPoster1 citations
  57. RNG: Relightable Neural GaussiansPoster1 citations
  58. RUBIK: A Structured Benchmark for Image Matching across Geometric ChallengesPoster1 citations
  59. RainyGS: Efficient Rain Synthesis with Physically-Based Gaussian SplattingPoster1 citations
  60. Rashomon Sets for Prototypical-Part Networks: Editing Interpretable Models in Real-TimePoster1 citations
  61. RayFlow: Instance-Aware Diffusion Acceleration via Adaptive Flow TrajectoriesPoster1 citations
  62. ReCap: Better Gaussian Relighting with Cross-Environment CapturesPoster1 citations
  63. ReNeg: Learning Negative Embedding with Reward GuidanceHighlight1 citations
  64. ReRAW: RGB-to-RAW Image Reconstruction via Stratified Sampling for Efficient Object Detection on the EdgePoster1 citations
  65. ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long VideosPoster1 citations
  66. Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image RetrievalHighlight1 citations
  67. Recovering Dynamic 3D Sketches from VideosPoster1 citations
  68. Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document RetrievalPoster1 citations
  69. Repurposing Stable Diffusion Attention for Training-Free Unsupervised Interactive SegmentationPoster1 citations
  70. Rethinking Correspondence-based Category-Level Object Pose EstimationPoster1 citations
  71. Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian SplattingPoster1 citations
  72. Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy PredictionPoster1 citations
  73. Retrieving Semantics from the Deep: an RAG Solution for Gesture SynthesisPoster1 citations
  74. Reversible Decoupling Network for Single Image Reflection RemovalPoster1 citations
  75. RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-TrainingHighlight1 citations
  76. RoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured EnvironmentsPoster1 citations
  77. Robust 3D Shape Reconstruction in Zero-Shot from a Single Image in the WildPoster1 citations
  78. SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in StructuresPoster1 citations
  79. SEC-Prompt:SEmantic Complementary Prompting for Few-Shot Class-Incremental LearningPoster1 citations
  80. SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained UnderstandingPoster1 citations
  81. SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image GenerationPoster1 citations
  82. SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D ReconstructionPoster1 citations
  83. STEPS: Sequential Probability Tensor Estimation for Text-to-Image Hard Prompt SearchPoster1 citations
  84. STINR: Deciphering Spatial Transcriptomics via Implicit Neural RepresentationPoster1 citations
  85. STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal GroundingPoster1 citations
  86. Scalable Autoregressive Monocular Depth EstimationPoster1 citations
  87. Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion ModelPoster1 citations
  88. SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World EnvironmentsPoster1 citations
  89. Science-T2I: Addressing Scientific Illusions in Image SynthesisPoster1 citations
  90. SeCap: Self-Calibrating and Adaptive Prompts for Cross-view Person Re-Identification in Aerial-Ground NetworksHighlight1 citations
  91. Search and Detect: Training-Free Long Tail Object Detection via Web-Image RetrievalPoster1 citations
  92. SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator TrajectoriesPoster1 citations
  93. Semantic and Sequential Alignment for Referring Video Object SegmentationPoster1 citations
  94. SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion ModelsPoster1 citations
  95. Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal GroundingPoster1 citations
  96. SerialGen: Personalized Image Generation by First Standardization Then PersonalizationPoster1 citations
  97. SfM-Free 3D Gaussian Splatting via Hierarchical TrainingPoster1 citations
  98. Shading Meets Motion: Self-supervised Indoor 3D Reconstruction Via Simultaneous Shape-from-Shading and Structure-from-MotionPoster1 citations
  99. Shape My Moves: Text-Driven Shape-Aware Synthesis of Human MotionsPoster1 citations
  100. ShiftwiseConv: Small Convolutional Kernel with Large Kernel EffectPoster1 citations
  101. Shining Yourself: High-Fidelity Ornaments Virtual Try-on with Diffusion ModelPoster1 citations
  102. ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual InstructionsPoster1 citations
  103. Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion ModelsPoster1 citations
  104. SimAvatar: Simulation-Ready Avatars with Layered Hair and ClothingPoster1 citations
  105. SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action AlignmentHighlight1 citations
  106. SimVS: Simulating World Inconsistencies for Robust View SynthesisPoster1 citations
  107. Similarity-Guided Layer-Adaptive Vision Transformer for UAV TrackingPoster1 citations
  108. Solving Instance Detection from an Open-World PerspectivePoster1 citations
  109. SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal ModelsHighlight1 citations
  110. Spatiotemporal Decoupling for Efficient Vision-Based Occupancy ForecastingPoster1 citations
  111. Spectral Informed Mamba for Robust Point Cloud ProcessingPoster1 citations
  112. Spiking Transformer: Introducing Accurate Addition-Only Spiking Self-Attention for TransformerPoster1 citations
  113. Split Adaptation for Pre-trained Vision TransformersPoster1 citations
  114. StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene GenerationPoster1 citations
  115. StdGEN: Semantic-Decomposed 3D Character Generation from Single ImagesPoster1 citations
  116. Steady Progress Beats Stagnation: Mutual Aid of Foundation and Conventional Models in Mixed Domain Semi-Supervised Medical Image SegmentationPoster1 citations
  117. StickMotion: Generating 3D Human Motions by Drawing a StickmanPoster1 citations
  118. Stretching Each Dollar: Diffusion Training from Scratch on a Micro-BudgetPoster1 citations
  119. SuperLightNet: Lightweight Parameter Aggregation Network for Multimodal Brain Tumor SegmentationPoster1 citations
  120. SuperPC: A Single Diffusion Model for Point Cloud Completion, Upsampling, Denoising, and ColorizationPoster1 citations
  121. SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step DiffusionPoster1 citations
  122. Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video GenerationPoster1 citations
  123. Synthetic Prior for Few-Shot Drivable Head Avatar InversionPoster1 citations
  124. Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure PriorsPoster1 citations
  125. Temporally Consistent Object-Centric Learning by Contrasting SlotsPoster1 citations
  126. Test-Time Backdoor Detection for Object Detection ModelsPoster1 citations
  127. The Change You Want To Detect: Semantic Change Detection In Earth Observation With Hybrid Data GenerationfPoster1 citations
  128. The Devil is in Low-Level Features for Cross-Domain Few-Shot SegmentationPoster1 citations
  129. Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference SystemsHighlight1 citations
  130. Token Cropr: Faster ViTs for Quite a Few TasksPoster1 citations
  131. Tokenize Image Patches: Global Context Fusion for Effective Haze Removal in Large ImagesPoster1 citations
  132. Toward Generalized Image Quality Assessment: Relaxing the Perfect Reference Quality AssumptionPoster1 citations
  133. Toward Robust Neural Reconstruction from Sparse Point SetsPoster1 citations
  134. Towards Autonomous Micromobility through Scalable Urban SimulationHighlight1 citations
  135. Towards Generalizable Trajectory Prediction using Dual-Level Representation Learning and Adaptive PromptingPoster1 citations
  136. Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual AttacksPoster1 citations
  137. Towards Realistic Example-based Modeling via 3D Gaussian StitchingPoster1 citations
  138. Towards Training-free Anomaly Detection with Vision and Language Foundation ModelsPoster1 citations
  139. Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly DetectionPoster1 citations
  140. Trajectory Mamba: Efficient Attention-Mamba Forecasting Model Based on Selective SSMPoster1 citations
  141. Traversing Distortion-Perception Tradeoff using a Single Score-Based Generative ModelPoster1 citations
  142. TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree SequencingPoster1 citations
  143. Turbo3D: Ultra-fast Text-to-3D GenerationPoster1 citations
  144. Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot ManipulationPoster1 citations
  145. Type-R: Automatically Retouching Typos for Text-to-Image GenerationHighlight1 citations
  146. UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image GenerationPoster1 citations
  147. USP-Gaussian: Unifying Spike-based Image Reconstruction, Pose Correction and Gaussian SplattingHighlight1 citations
  148. Understanding Multi-layered Transmission MatricesHighlight1 citations
  149. UniRestore: Unified Perceptual and Task-Oriented Image Restoration Model Using Diffusion PriorHighlight1 citations
  150. Unified Uncertainty-Aware Diffusion for Multi-Agent Trajectory ModelingPoster1 citations
  151. Unlearning through Knowledge Overwriting: Reversible Federated Unlearning via Selective Sparse AdapterPoster1 citations
  152. Unsupervised Foundation Model-Agnostic Slide-Level Representation LearningPoster1 citations
  153. Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering IncorrectlyPoster1 citations
  154. VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token SparsificationPoster1 citations
  155. VITED: Video Temporal Evidence DistillationPoster1 citations
  156. VTON 360: High-Fidelity Virtual Try-On from Any Viewing DirectionPoster1 citations
  157. VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's GuidePoster1 citations
  158. VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative PriorsPoster1 citations
  159. VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video UnderstandingPoster1 citations
  160. VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One StepHighlight1 citations
  161. VinaBench: Benchmark for Faithful and Consistent Visual NarrativesPoster1 citations
  162. Visual Lexicon: Rich Image Features in Language SpacePoster1 citations
  163. Visual Prompting for One-shot Controllable Video Editing without InversionPoster1 citations
  164. VoteFlow: Enforcing Local Rigidity in Self-Supervised Scene FlowPoster1 citations
  165. VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow PredictionPoster1 citations
  166. What Makes a Good Dataset for Knowledge Distillation?Poster1 citations
  167. When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven ApproachPoster1 citations
  168. Words or Vision: Do Vision-Language Models Have Blind Faith in Text?Poster1 citations
  169. Yo'Chameleon: Personalized Vision and Language GenerationPoster1 citations
  170. Zero-Shot Monocular Scene Flow Estimation in the WildAward Candidate1 citations
  171. g3D-LF: Generalizable 3D-Language Feature Fields for Embodied TasksPoster1 citations
  172. 3D Dental Model Segmentation with Geometrical Boundary PreservingPoster
  173. 3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial RepresentationsPoster
  174. 3D Gaussian Inpainting with Depth-Guided Cross-View ConsistencyPoster
  175. 3D Prior Is All You Need: Cross-Task Few-shot 2D Gaze EstimationPoster
  176. 3D Student Splatting and ScoopingAward Candidate
  177. 3D-AVS: LiDAR-based 3D Auto-Vocabulary SegmentationPoster
  178. 3D-MVP: 3D Multiview Pretraining for ManipulationPoster
  179. 3D-SLNR: A Super Lightweight Neural Representation for Large-scale 3D MappingPoster
  180. 3DEnhancer: Consistent Multi-View Diffusion for 3D EnhancementPoster
  181. 4D-Fly: Fast 4D Reconstruction from a Single Monocular VideoPoster
  182. 4DGC: Rate-Aware 4D Gaussian Compression for Efficient Streamable Free-Viewpoint VideoPoster
  183. 4DTAM: Non-Rigid Tracking and Mapping via Dynamic Surface GaussiansPoster
  184. 4Deform: Neural Surface Deformation for Robust Shape InterpolationPoster
  185. A Comprehensive Study of Decoder-Only LLMs for Text-to-Image GenerationPoster
  186. A Focused Human Body Model for Accurate Anthropometric Measurements ExtractionPoster
  187. A General Adaptive Dual-level Weighting Mechanism for Remote Sensing PansharpeningPoster
  188. A Hubness Perspective on Representation Learning for Graph-Based Multi-View ClusteringPoster
  189. A New Statistical Model of Star Speckles for Learning to Detect and Characterize Exoplanets in Direct Imaging ObservationsPoster
  190. A Physics-Informed Blur Learning Framework for Imaging SystemsPoster
  191. A Polarization-Aided Transformer for Image Deblurring via Motion Vector DecompositionHighlight
  192. A Regularization-Guided Equivariant Approach for Image RestorationPoster
  193. A Selective Re-learning Mechanism for Hyperspectral Fusion ImagingPoster
  194. A Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image SegmentationPoster
  195. A Simple yet Effective Layout Token in Large Language Models for Document UnderstandingPoster
  196. A Theory of Learning Unified Model via Knowledge Integration from Label Space Varying DomainsPoster
  197. A Unified Approach to Interpreting Self-supervised Pre-training Methods for 3D Point Clouds via InteractionsHighlight
  198. A Unified Framework for Heterogeneous Semi-supervised LearningPoster
  199. A Unified Image-Dense Annotation Generation Model for Underwater ScenesPoster
  200. A Unified Latent Schrodinger Bridge Diffusion Model for Unsupervised Anomaly Detection and LocalizationPoster
  201. A Unified, Resilient, and Explainable Adversarial Patch DetectorPoster
  202. A Universal Scale-Adaptive Deformable Transformer for Image Restoration across Diverse ArtifactsPoster
  203. A3: Few-shot Prompt Learning of Unlearnable Examples with Cross-Modal Adversarial Feature AlignmentPoster
  204. A4A: Adapter for Adapter Transfer via All-for-All Mapping for Cross-Architecture ModelsPoster
  205. ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image ObjectsPoster
  206. ABC-Former: Auxiliary Bimodal Cross-domain Transformer with Interactive Channel Attention for White BalancePoster
  207. ACAttack: Adaptive Cross Attacking RGB-T Tracker via Multi-Modal Response DecouplingPoster
  208. ACL: Activating Capability of Linear Attention for Image RestorationPoster
  209. ADD: Attribution-Driven Data Augmentation Framework for Boosting Image Super-ResolutionPoster
  210. ADU: Adaptive Detection of Unknown Categories in Black-Box Domain AdaptationPoster
  211. AIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic DataPoster
  212. AIpparel: A Multimodal Foundation Model for Digital GarmentsHighlight
  213. ALIEN: Implicit Neural Representations for Human Motion Prediction under Arbitrary LatencyHighlight
  214. AMO Sampler: Enhancing Text Rendering with OvershootingPoster
  215. AMR-Transformer: Enabling Efficient Long-range Interaction for Complex Neural Fluid SimulationPoster
  216. ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric InteractionPoster
  217. APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision TransformersPoster
  218. APT: Adaptive Personalized Training for Diffusion Models with Limited DataPoster
  219. ASHiTA: Automatic Scene-grounded HIerarchical Task AnalysisPoster
  220. ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background InpaintingPoster
  221. ATP: Adaptive Threshold Pruning for Efficient Data Encoding in Quantum Neural NetworksPoster
  222. AVF-MAE++: Scaling Affective Video Facial Masked Autoencoders via Efficient Audio-Visual Self-Supervised LearningPoster
  223. AVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual LearningPoster
  224. Acc3D: Accelerating Single Image to 3D Diffusion Models via Edge Consistency Guided Score DistillationPoster
  225. Accelerating Diffusion Transformer via Increment-Calibrated Caching with Channel-Aware Singular Value DecompositionPoster
  226. Accurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-DistillationPoster
  227. Acquire and then Adapt: Squeezing out Text-to-Image Model for Image RestorationPoster
  228. Action Detail Matters: Refining Video Recognition with Local Action QueriesPoster
  229. Activating Sparse Part Concepts for 3D Class Incremental LearningPoster
  230. Active Event-based Stereo VisionPoster
  231. Active Hyperspectral Imaging Using an Event CameraHighlight
  232. AdMiT: Adaptive Multi-Source Tuning in Dynamic EnvironmentsPoster
  233. AdaDARE-gamma: Balancing Stability and Plasticity in Multi-modal LLMs through Efficient AdaptationPoster
  234. AdaptCMVC: Robust Adaption to Incremental Views in Continual Multi-view ClusteringPoster
  235. Adapting Dense Matching for Homography Estimation with Grid-based AccelerationPoster
  236. Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal PerceptionPoster
  237. Adapting Text-to-Image Generation with Feature Difference Instruction for Generic Image RestorationPoster
  238. Adapting to Observation Length of Trajectory Prediction via Contrastive LearningPoster
  239. Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic ThresholdsPoster
  240. Adaptive Dropout: Unleashing Dropout across Layers for Generalizable Image Super-ResolutionPoster
  241. Adaptive Keyframe Sampling for Long Video UnderstandingPoster
  242. Adaptive Markup Language Generation for Contextually-Grounded Visual Document UnderstandingPoster
  243. Adaptive Non-Uniform Timestep Sampling for Accelerating Diffusion Model TrainingPoster
  244. Adaptive Parameter Selection for Tuning Vision-Language ModelsPoster
  245. Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play EnhancementPoster
  246. Adaptive Rectangular Convolution for Remote Sensing PansharpeningPoster
  247. Adaptive Unimodal Regulation for Balanced Multimodal Information AcquisitionPoster
  248. Advancing Adversarial Robustness in GNeRFs: The IL2-NeRF AttackPoster
  249. Advancing Generalizable Tumor Segmentation with Anomaly-Aware Open-Vocabulary Attention Maps and Frozen Foundation Diffusion ModelsPoster
  250. Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 DatasetPoster
  251. Advancing Multiple Instance Learning with Continual Learning for Whole Slide ImagingHighlight
  252. Adventurer: Optimizing Vision Mamba Architecture Designs for EfficiencyPoster
  253. Adversarial Domain Prompt Tuning and Generation for Single Domain GeneralizationPoster
  254. AeSPa : Attention-guided Self-supervised Parallel Imaging for MRI ReconstructionPoster
  255. AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language ModelsPoster
  256. AirRoom: Objects Matter in Room ReidentificationPoster
  257. Align-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video EditingPoster
  258. Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Large Model EnhancementPoster
  259. Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question AnsweringPoster
  260. All-Day Multi-Camera Multi-Target TrackingPoster
  261. All-Optical Nonlinear Diffractive Deep Network for Ultrafast Image DenoisingHighlight
  262. All-directional Disparity Estimation for Real-world QPD ImagesHighlight
  263. AlphaPre: Amplitude-Phase Disentanglement Model for Precipitation NowcastingPoster
  264. An Image-like Diffusion Method for Human-Object Interaction DetectionPoster
  265. Analyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose EstimationPoster
  266. Anatomical Consistency and Adaptive Prior-informed Transformation for Multi-contrast MR Image Synthesis via Diffusion ModelPoster
  267. Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2DPoster
  268. AniGrad: Anisotropic Gradient-Adaptive Sampling for 3D Reconstruction From Monocular VideoPoster
  269. AniMer: Animal Pose and Shape Estimation Using Family Aware TransformerPoster
  270. AniMo: Species-Aware Model for Text-Driven Animal Motion GenerationPoster
  271. Animate and Sound an ImagePoster
  272. Annotation Ambiguity Aware Semi-Supervised Medical Image SegmentationHighlight
  273. Anomize: Better Open Vocabulary Video Anomaly DetectionPoster
  274. Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object PerceptionPoster
  275. Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask TrackingPoster
  276. Any6D: Model-free 6D Pose Estimation of Novel ObjectsPoster
  277. AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual VideosPoster
  278. AnyMap: Learning a General Camera Model for Structure-from-Motion with Unknown Distortion in Dynamic ScenesPoster
  279. AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion ModelsPoster
  280. Anyattack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language ModelsPoster
  281. Apply Hierarchical-Chain-of-Generation to Complex Attributes Text-to-3D GenerationPoster
  282. ArcPro: Architectural Programs for Structured 3D Abstraction of Sparse PointsHighlight
  283. Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?Poster
  284. Argus: A Compact and Versatile Foundation Model for VisionPoster
  285. Argus: Vision-Centric Reasoning with Grounded Chain-of-ThoughtPoster
  286. ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image IntermediaryPoster
  287. Articulated Kinematics Distillation from Video Diffusion ModelsPoster
  288. Assessing and Learning Alignment of Unimodal Vision and Language ModelsHighlight
  289. Associative TransformerPoster
  290. Asynchronous Collaborative Graph Representation for Frames and EventsPoster
  291. Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT AccelerationPoster
  292. Attention IoU: Examining Biases in CelebA using Attention MapsPoster
  293. Attraction Diminishing and Distributing for Few-Shot Class-Incremental LearningPoster
  294. Attribute-Missing Multi-view Graph ClusteringPoster
  295. Attribute-formed Class-specific Concept Space: Endowing Language Bottleneck Model with Better Interpretability and ScalabilityPoster
  296. AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion TransformersPoster
  297. Audio-Visual Semantic Graph Network for Audio-Visual Event LocalizationPoster
  298. Augmented Deep Contexts for Spatially Embedded Video CodingHighlight
  299. Augmenting Perceptual Super-Resolution via Image Quality PredictorsPoster
  300. AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360deg Unbounded Scene InpaintingPoster
  301. Auto-Encoded Supervision for Perceptual Image Super-ResolutionPoster
  302. AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video HashingPoster
  303. Automated Proof of Polynomial Inequalities via Reinforcement LearningPoster
  304. Automatic Spectral Calibration of Hyperspectral Images: Method, Dataset and BenchmarkPoster
  305. Autoregressive Distillation of Diffusion TransformersPoster
  306. Autoregressive Sequential Pretraining for Visual TrackingPoster
  307. AvatarArtist: Open-Domain 4D AvatarizationPoster
  308. BACON: Improving Clarity of Image Captions via Bag-of-Concept GraphsPoster
  309. BADGR: Bundle Adjustment Diffusion Conditioned by Gradients for Wide-Baseline Floor Plan ReconstructionHighlight
  310. BASKET: A Large-Scale Video Dataset for Fine-Grained Skill EstimationPoster
  311. BF-STVSR: B-Splines and Fourier---Best Friends for High Fidelity Spatial-Temporal Video Super-ResolutionPoster
  312. BHViT: Binarized Hybrid Vision TransformerPoster
  313. BIGS: Bimanual Category-agnostic Interaction Reconstruction from Monocular Videos via 3D Gaussian SplattingPoster
  314. BLADE: Single-view Body Mesh Estimation through Accurate Depth EstimationPoster
  315. BOE-ViT: Boosting Orientation Estimation with Equivariance in Self-Supervised 3D Subtomogram AlignmentPoster
  316. BOLT: Boost Large Vision-Language Model Without Training for Long-form Video UnderstandingPoster
  317. BOOTPLACE: Bootstrapped Object Placement with Detection TransformersPoster
  318. Balancing Two Classifiers via A Simplex ETF Structure for Model CalibrationPoster
  319. Bayesian Prompt Flow Learning for Zero-Shot Anomaly DetectionPoster
  320. Bayesian Test-Time Adaptation for Vision-Language ModelsPoster
  321. Be More Specific: Evaluating Object-centric Realism in Synthetic ImagesPoster
  322. Believing is Seeing: Unobserved Object Detection using Generative ModelsPoster
  323. Benchmarking Object Detectors under Real-World Distribution Shifts in Satellite ImageryPoster
  324. Beyond Background Shift: Rethinking Instance Replay in Continual Semantic SegmentationPoster
  325. Beyond Clean Training Data: A Versatile and Model-Agnostic Framework for Out-of-Distribution Detection with Contaminated Training DataPoster
  326. Beyond Generation: A Diffusion-based Low-level Feature Extractor for Detecting AI-generated ImagesPoster
  327. Beyond Human Perception: Understanding Multi-Object World from Monocular ViewPoster
  328. Beyond Image Classification: A Video Benchmark and Dual-Branch Hybrid Discrimination Framework for Compositional Zero-Shot LearningPoster
  329. Beyond Local Sharpness: Communication-Efficient Global Sharpness-aware Minimization for Federated LearningPoster
  330. Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual KnowledgePoster
  331. Beyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and HarmonizationPoster
  332. Beyond Words: Augmenting Discriminative Richness via Diffusions in Unsupervised Prompt LearningPoster
  333. Bias for Action: Video Implicit Neural Representations with Bias ModulationPoster
  334. Binarized Mamba-Transformer for Lightweight Quad Bayer HybridEVS DemosaicingPoster
  335. Binarized Neural Network for Multi-spectral Image FusionPoster
  336. BioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational PathologyPoster
  337. Black Hole-Driven Identity Absorbing in Diffusion ModelsPoster
  338. Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable EventsPoster
  339. BlenderGym: Benchmarking Foundational Model Systems for Graphics EditingHighlight
  340. Blind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion ModelPoster
  341. Blood Flow Speed Estimation with Optical Coherence Tomography Angiography ImagesPoster
  342. Blurry-Edges: Photon-Limited Depth Estimation from Defocused BoundariesPoster
  343. Boltzmann Attention Sampling for Image Analysis with Small ObjectsPoster
  344. Boost the Inference with Co-training: A Depth-guided Mutual Learning Framework for Semi-supervised Medical Polyp SegmentationPoster
  345. Boosting Adversarial Transferability through Augmentation in Hypothesis SpacePoster
  346. Boosting Domain Incremental Learning: Selecting the Optimal Parameters is All You NeedPoster
  347. Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal TransportPoster
  348. Boosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty EstimationPoster
  349. Brain-Inspired Spiking Neural Networks for Energy-Efficient Object DetectionPoster
  350. Breaking the Memory Barrier of Contrastive Loss via Tile-Based StrategyHighlight
  351. BrepGiff: Lightweight Generation of Complex B-rep with 3D GAT DiffusionPoster
  352. Bridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical FlowPoster
  353. Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsPoster
  354. Bridging Viewpoint Gaps: Geometric Reasoning Boosts Semantic CorrespondencePoster
  355. Bridging the Gap between Gaussian Diffusion Models and Universal Quantization for Image CompressionPoster
  356. Bridging the Vision-Brain Gap with an Uncertainty-Aware Blur PriorPoster
  357. Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical AnalysisPoster
  358. Building Vision Models upon Heat ConductionPoster
  359. ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free WayPoster
  360. CADDreamer: CAD Object Generation from Single-view ImagesHighlight
  361. CADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature LeveragingPoster
  362. CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language ModelsPoster
  363. CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D ImageHighlight
  364. CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual InteractionHighlight
  365. CASP: Compression of Large Multimodal Models Based on Attention SparsityHighlight
  366. CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional VideoPoster
  367. CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained AlignmentPoster
  368. CCIN: Compositional Conflict Identification and Neutralization for Composed Image RetrievalHighlight
  369. CGMatch: A Different Perspective of Semi-supervised LearningPoster
  370. CH3Depth: Efficient and Flexible Depth Foundation Model with Flow MatchingHighlight
  371. CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object RepresentationPoster
  372. CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCRPoster
  373. CLIP-driven Coarse-to-fine Semantic Guidance for Fine-grained Open-set Semi-supervised LearningPoster
  374. CLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair LossPoster
  375. CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based FrameworkPoster
  376. CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AIPoster
  377. COB-GS: Clear Object Boundaries in 3DGS Segmentation Based on Boundary-Adaptive Gaussian SplittingPoster
  378. COBRA: COmBinatorial Retrieval Augmentation for Few-Shot AdaptationPoster
  379. COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time AdaptationPoster
  380. COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution ShiftsHighlight
  381. CSC-PA: Cross-image Semantic Correlation via Prototype Attentions for Single-network Semi-supervised Breast Tumor SegmentationPoster
  382. CTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D DiffusionPoster
  383. CaMuViD: Calibration-Free Multi-View DetectionPoster
  384. CamPoint: Boosting Point Cloud Segmentation with Virtual CameraPoster
  385. Camera Resection from Known Line Pencils and a Radially Distorted ScanlinePoster
  386. Camouflage Anything: Learning to Hide using Controlled Out-painting and Representation EngineeringPoster
  387. Can Generative Video Models Help Pose Estimation?Highlight
  388. Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?Poster
  389. Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and UnderstandingHighlight
  390. Can Text-to-Video Generation help Video-Language Alignment?Poster
  391. Can't Slow Me Down: Learning Robust and Hardware-Adaptive Object Detectors against Latency Attacks for Edge DevicesPoster
  392. CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo BoothPoster
  393. Category-Agnostic Neural Object RiggingPoster
  394. Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial AttacksPoster
  395. Chain of Semantics Programming in 3D Gaussian Splatting Representation for 3D Vision GroundingPoster
  396. Change3D: Revisiting Change Detection and Captioning from A Video Modeling PerspectiveHighlight
  397. Channel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image DerainingPoster
  398. Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor ScenesPoster
  399. Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMsPoster
  400. Charm: The Missing Piece in ViT Fine-Tuning for Image Aesthetic AssessmentPoster
  401. Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion ModelsPoster
  402. ChatHuman: Chatting about 3D Humans with ToolsPoster
  403. CheXWorld: Exploring Image World Modeling for Radiograph Representation LearningPoster
  404. CheXwhatsApp: A Dataset for Exploring Challenges in the Diagnosis of Chest X-rays through Mobile DevicesPoster
  405. Cheb-GR: Rethinking K-nearest Neighbor Search in Re-ranking for Person Re-identificationPoster
  406. Chebyshev Attention Depth Permutation Texture Network with Latent Texture Attribute LossPoster
  407. CheckManual: A New Challenge and Benchmark for Manual-based Appliance ManipulationHighlight
  408. Classic Video Denoising in a Machine Learning World: Robust, Fast, and ControllablePoster
  409. Classifier-guided CLIP Distillation for Unsupervised Multi-label ClassificationPoster
  410. Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual ClassifiersPoster
  411. ClimbingCap: Multi-Modal Dataset and Method for Rock Climbing in World CoordinateHighlight
  412. Closest Neighbors are Harmful for Lightweight Masked Auto-encodersPoster
  413. Co-Speech Gesture Video Generation with Implicit Motion-Audio EntanglementPoster
  414. CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused ImagesPoster
  415. CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity QuantificationPoster
  416. CoMBO: Conflict Mitigation via Branched Optimization for Class Incremental SegmentationPoster
  417. CoMapGS: Covisibility Map-based Gaussian Splatting for Sparse Novel View SynthesisPoster
  418. CoMatcher: Multi-View Collaborative Feature MatchingPoster
  419. CoSER: Towards Consistent Dense Multiview Text-to-Image Generator for 3D CreationHighlight
  420. CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language ModelsPoster
  421. CocoER: Aligning Multi-Level Feature by Competition and Coordination for Emotion RecognitionPoster
  422. ColabSfM: Collaborative Structure-from-Motion by Point Cloud RegistrationPoster
  423. Collaborative Tree Search for Enhancing Embodied Multi-Agent CollaborationPoster
  424. Color Alignment in DiffusionPoster
  425. ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle MatricesPoster
  426. Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature SpacePoster
  427. Commonsense Video Question Answering through Video-Grounded Entailment Tree ReasoningPoster
  428. Compass Control: Multi Object Orientation Control for Text-to-Image GenerationPoster
  429. Composing Parts for Expressive Object GenerationPoster
  430. Compositional Caching for Training-free Open-vocabulary Attribute DetectionHighlight
  431. Compositional Targeted Multi-Label Universal PerturbationsPoster
  432. Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision TransformersHighlight
  433. Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and HarmonizationPoster
  434. ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalPoster
  435. Concept Lancet: Image Editing with Compositional Representation TransplantPoster
  436. Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph GenerationPoster
  437. Conformal Prediction for Zero-Shot ModelsPoster
  438. Conical Visual Concentration for Efficient Large Vision-Language ModelsPoster
  439. Consistency Posterior Sampling for Diverse Image SynthesisPoster
  440. Consistency-aware Self-Training for Iterative-based Stereo MatchingPoster
  441. Consistent Normal Orientation for 3D Point Clouds via Least Squares on Delaunay GraphPoster
  442. Consistent and Controllable Image Animation with Motion Diffusion ModelsPoster
  443. Continuous Adverse Weather Removal via Degradation-Aware DistillationPoster
  444. Continuous Locomotive Crowd Behavior GenerationPoster
  445. Continuous Space-Time Video Resampling with Invertible Motion SteganographyPoster
  446. ControlFace: Harnessing Facial Parametric Control for Face RiggingPoster
  447. Controllable Human Image Generation with Personalized Multi-GarmentsPoster
  448. Convex Combination Star Shape Prior for Data-driven Image Semantic SegmentationPoster
  449. Convex Relaxation for Robust Vanishing Point Estimation in Manhattan WorldAward Candidate
  450. CorrBEV: Multi-View 3D Object Detection by Correlation Learning with Multi-modal PrototypesPoster
  451. Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt TuningPoster
  452. Crab: A Unified Audio-Visual Scene Understanding Model with Explicit CooperationPoster
  453. CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry RefinerPoster
  454. Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian SplattingHighlight
  455. CroCoDL: Cross-device Collaborative Dataset for LocalizationPoster
  456. Cross-Modal 3D Representation with Multi-View Images and Point CloudsPoster
  457. Cross-Modal Distillation for 2D/3D Multi-Object Discovery from 2D MotionPoster
  458. Cross-Modal Interactive Perception Network with Mamba for Lung Tumor Segmentation in PET-CT ImagesPoster
  459. Cross-Rejective Open-Set SAR Image RegistrationPoster
  460. CrossOver: 3D Scene Cross-Modal AlignmentHighlight
  461. CrossSDF: 3D Reconstruction of Thin Structures From Cross-SectionsPoster
  462. CryptoFace: End-to-End Encrypted Face RecognitionPoster
  463. CustAny: Customizing Anything from A Single ExamplePoster
  464. Customized Condition Controllable Generation for Video SoundtrackPoster
  465. D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective Recognition.Poster
  466. DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision TransformersPoster
  467. DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report GenerationPoster
  468. DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image FusionPoster
  469. DEAL: Data-Efficient Adversarial Learning for High-Quality Infrared ImagingPoster
  470. DFM: Differentiable Feature Matching for Anomaly DetectionPoster
  471. DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings LearningPoster
  472. DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-IDPoster
  473. DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion ModelsAward Candidate
  474. DIO: Decomposable Implicit 4D Occupancy-Flow World ModelPoster
  475. DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric VideosHighlight
  476. DKC: Differentiated Knowledge Consolidation for Cloth-Hybrid Lifelong Person Re-identificationPoster
  477. DL2G: Degradation-guided Local-to-Global Restoration for Eyeglass Reflection RemovalPoster
  478. DOF-GS: Adjustable Depth-of-Field 3D Gaussian Splatting for Post-Capture Refocusing, Defocus Rendering and Blur RemovalPoster
  479. DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-ResolutionPoster
  480. DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid FrameworkPoster
  481. DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic SegmentationPoster
  482. DSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot SegmentationPoster
  483. DTGBrepGen: A Novel B-rep Generative Model through Decoupling Topology and GeometryPoster
  484. DTOS: Dynamic Time Object Sensing with Large Multimodal ModelPoster
  485. DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D TeachersPoster
  486. DVHGNN: Multi-Scale Dilated Vision HGNN for Efficient Vision RecognitionPoster
  487. DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression ComprehensionPoster
  488. D^2iT: Dynamic Diffusion Transformer for Accurate Image GenerationPoster
  489. D^3CTTA: Domain-Dependent Decorrelation for Continual Test-Time Adaption of 3D LiDAR SegmentationPoster
  490. DaCapo: Score Distillation as Stacked Bridge for Fast and High-quality 3D EditingPoster
  491. DashGaussian: Optimizing 3D Gaussian Splatting in 200 SecondsHighlight
  492. Data Distributional Properties As Inductive Bias for Systematic GeneralizationPoster
  493. Data Synthesis with Diverse Styles for Face Recognition via 3DMM-Guided DiffusionPoster
  494. Data-Free Group-Wise Fully Quantized Winograd Convolution via Learnable ScalesPoster
  495. Data-free Universal Adversarial Perturbation with Pseudo-semantic PriorPoster
  496. DeCLIP: Decoupled Learning for Open-Vocabulary Dense PerceptionPoster
  497. DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long VideosPoster
  498. DeClotH: Decomposable 3D Cloth and Human Body Reconstruction from a Single ImagePoster
  499. DeDe: Detecting Backdoor Samples for SSL Encoders via DecodersPoster
  500. De^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze EstimationPoster
  501. Decoder Gradient Shield: Provable and High-Fidelity Prevention of Gradient-Based Box-Free Watermark RemovalPoster
  502. Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image CompressionPoster
  503. Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task LearningPoster
  504. Decoupled Motion Expression Video SegmentationPoster
  505. Decoupling Training-Free Guided Diffusion by ADMMPoster
  506. Deep Change Monitoring: A Hyperbolic Representative Learning Framework and a Dataset for Long-term Fine-grained Tree Change DetectionHighlight
  507. Deep Fair Multi-View Clustering with Attention KANHighlight
  508. DeepCompress-ViT: Rethinking Model Compression to Enhance Efficiency of Vision Transformers at the EdgePoster
  509. DeepLA-Net: Very Deep Local Aggregation Networks for Point Cloud AnalysisPoster
  510. DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual InspectionHighlight
  511. DeformCL: Learning Deformable Centerline Representation for Vessel Extraction in 3D Medical ImagePoster
  512. Degradation-Aware Feature Perturbation for All-in-One Image RestorationPoster
  513. DejaVid: Encoder-Agnostic Learned Temporal Matching for Video ClassificationPoster
  514. Dense Dispersed Structured Light for Hyperspectral 3D Imaging of Dynamic ScenesPoster
  515. Dense Match Summarization for Faster Two-view EstimationPoster
  516. Dense-SfM: Structure from Motion with Dense Consistent MatchingPoster
  517. Depth-Guided Bundle Sampling for Efficient Generalizable Neural Radiance Field ReconstructionPoster
  518. Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAIPoster
  519. Descriptor-In-Pixel : Point-Feature Tracking For Pixel Processor ArraysAward Candidate
  520. Detail-Preserving Latent Diffusion for Stable Shadow RemovalPoster
  521. Detect Any Mirrors: Boosting Learning Reliability on Large-Scale Unlabeled Data with an Iterative Data EnginePoster
  522. Detecting Open World Objects via Partial Attribute AssignmentPoster
  523. Detection-Friendly Nonuniformity Correction: A Union Framework for Infrared UAV Target DetectionHighlight
  524. Deterministic Image-to-Image Translation via Denoising Brownian Bridge Models with Dual ApproximatorsPoster
  525. Deterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion SynthesisPoster
  526. Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified AttentionPoster
  527. DexHandDiff: Interaction-aware Diffusion Planning for Adaptive Dexterous ManipulationPoster
  528. DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian SplattingPoster
  529. DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection TransformerPoster
  530. DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy LabelsPoster
  531. DiSRT-In-Bed: Diffusion-Based Sim-to-Real Transfer Framework for In-Bed Human Mesh RecoveryPoster
  532. DiSciPLE: Learning Interpretable Programs for Scientific Visual DiscoveryPoster
  533. DiTASK: Multi-Task Fine-Tuning with Diffeomorphic TransformationsPoster
  534. Diff-Palm: Realistic Palmprint Generation with Polynomial Creases and Intra-Class Variation Controllable Diffusion ModelsPoster
  535. Diff2Flow: Training Flow Matching Models via Diffusion Model AlignmentPoster
  536. DiffCAM: Data-Driven Saliency Maps by Capturing Feature DifferencesHighlight
  537. DiffLO: Semantic-Aware LiDAR Odometry with Diffusion-Based RefinementPoster
  538. DiffLocks: Generating 3D Hair from a Single Image using Diffusion ModelsPoster
  539. DiffVsgg: Diffusion-Driven Online Video Scene Graph GenerationPoster
  540. Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy GenerationPoster
  541. Differentiable Inverse Rendering with Interpretable Basis BRDFsPoster
  542. Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image CaptioningPoster
  543. Diffusion Model is Effectively Its Own TeacherPoster
  544. Diffusion-based Event Generation for High-Quality Image DeblurringPoster
  545. Diffusion-based Realistic Listening Head Generation via Hybrid Motion ModelingHighlight
  546. DirectTriGS: Triplane-based Gaussian Splatting Field Representation for 3D GenerationPoster
  547. Directional Label Diffusion Model for Learning from Noisy LabelsPoster
  548. DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text RetrievalPoster
  549. Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant LearningPoster
  550. Discrete to Continuous: Generating Smooth Transition Poses from Sign Language ObservationsPoster
  551. Disentangled Pose and Appearance Guidance for Multi-Pose GenerationPoster
  552. Disentangling Safe and Unsafe Image Corruptions via Anisotropy and LocalityPoster
  553. DiskVPS: Vanishing Point Detector via Hough Transform in a Disk RegionPoster
  554. Dissecting and Mitigating Diffusion Bias via Mechanistic InterpretabilityPoster
  555. Distilled Prompt Learning for Incomplete Multimodal Survival PredictionPoster
  556. Distilling Spatially-Heterogeneous Distortion Perception for Blind Image Quality AssessmentPoster
  557. Distinguish Then Exploit: Source-free Open Set Domain Adaptation via Weight Barcode Estimation and Sparse Label AssignmentPoster
  558. Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly DetectionPoster
  559. DiverseFlow: Sample-Efficient Diverse Mode Coverage in FlowsPoster
  560. DnLUT: Ultra-Efficient Color Image Denoising via Channel-Aware Lookup TablesPoster
  561. Do ImageNet-trained Models Learn Shortcuts? The Impact of Frequency Shortcuts on GeneralizationPoster
  562. Do Visual Imaginations Improve Vision-and-Language Navigation Agents?Poster
  563. Do Your Best and Get Enough Rest for Continual LearningPoster
  564. DoF-Gaussian: Controllable Depth-of-Field for 3D Gaussian SplattingPoster
  565. DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document UnderstandingPoster
  566. Docopilot: Improving Multimodal Models for Document-Level UnderstandingPoster
  567. Domain Adaptive Diabetic Retinopathy Grading with Model Absence and Flowing DataPoster
  568. Domain Generalization in CLIP via Learning with Diverse Text PromptsPoster
  569. Doppelgangers and Adversarial VulnerabilityHighlight
  570. Doppelgangers++: Improved Visual Disambiguation with Geometric 3D FeaturesHighlight
  571. Dragin3D: Image Editing by Dragging in 3D SpacePoster
  572. DreamTrack: Dreaming the Future for Multimodal Visual Object TrackingPoster
  573. DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous DrivingHighlight
  574. DriveScape: High-Resolution Driving Video Generation by Multi-View Feature FusionPoster
  575. DropGaussian: Structural Regularization for Sparse-view Gaussian SplattingPoster
  576. DropoutGS: Dropping Out Gaussians for Better Sparse-view RenderingPoster
  577. Dual Energy-Based Model with Open-World Uncertainty Estimation for Out-of-distribution DetectionPoster
  578. Dual Exposure Stereo for Extended Dynamic Range 3D ImagingPoster
  579. Dual Focus-Attention Transformer for Robust Point Cloud RegistrationPoster
  580. Dual Semantic Guidance for Open Vocabulary Semantic SegmentationPoster
  581. Dual-Agent Optimization framework for Cross-Domain Few-Shot SegmentationPoster
  582. Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General PansharpeningPoster
  583. DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image SegmentationPoster
  584. DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective SchedulingPoster
  585. DynPose: Largely Improving the Efficiency of Human Pose Estimation by a Simple Dynamic FrameworkPoster
  586. DynScene: Scalable Generation of Dynamic Robotic Manipulation Scenes for Embodied AIPoster
  587. DynaMoDe-NeRF: Motion-aware Deblurring Neural Radiance Field for Dynamic ScenesPoster
  588. Dynamic Camera Poses and Where to Find ThemPoster
  589. Dynamic Content Prediction with Motion-aware Priors for Blind Face Video RestorationPoster
  590. Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio SemanticsPoster
  591. Dynamic Group Normalization: Spatio-Temporal Adaptation to Evolving Data StatisticsPoster
  592. Dynamic Motion Blending for Versatile Motion EditingPoster
  593. Dynamic Neural Surfaces for Elastic 4D Shape Representation and AnalysisPoster
  594. Dynamic Pseudo Labeling via Gradient Cutting for High-Low Entropy ExplorationPoster
  595. Dynamic Stereotype Theory Induced Micro-expression Recognition with Oriented DeformationPoster
  596. Dynamic Updates for Language Adaptation in Visual-Language TrackingPoster
  597. EAP-GS: Efficient Augmentation of Pointcloud for 3D Gaussian Splatting in Few-shot Scene ReconstructionPoster
  598. EASEMVC:Efficient Dual Selection Mechanism for Deep Multi-View ClusteringPoster
  599. EBS-EKF: Accurate and High Frequency Event-based Star TrackingHighlight
  600. EDCFlow: Exploring Temporally Dense Difference Maps for Event-based Optical Flow EstimationPoster
  601. EDM: Equirectangular Projection-Oriented Dense Kernelized Feature MatchingPoster
  602. EIDT-V: Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video GenerationPoster
  603. ERUPT: Efficient Rendering with Unposed Patch TransformerPoster
  604. ESC: Erasing Space Concept for Knowledge DeletionHighlight
  605. ESCAPE: Equivariant Shape Completion via Anchor Point EncodingPoster
  606. ETAP: Event-based Tracking of Any PointHighlight
  607. EVOS: Efficient Implicit Neural Training via EVOlutionary SelectorPoster
  608. EVPGS: Enhanced View Prior Guidance for Splatting-based Extrapolated View SynthesisPoster
  609. EVolSplat: Efficient Volume-based Gaussian Splatting for Urban View SynthesisPoster
  610. Early-Bird Diffusion: Investigating and Leveraging Timestep-Aware Early-Bird Tickets in Diffusion Models for Efficient TrainingPoster
  611. Easy-editable Image Vectorization with Multi-layer Multi-scale Distributed Visual Feature EmbeddingPoster
  612. EasyCraft: A Robust and Efficient Framework for Automatic Avatar CraftingPoster
  613. EchoMatch: Partial-to-Partial Shape Matching via Correspondence ReflectionPoster
  614. EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe GuidancePoster
  615. Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional ConditioningPoster
  616. EdgeDiff: Edge-aware Diffusion Network for Building Reconstruction from Point CloudsPoster
  617. EdgeMovingNet: Edge-preserving Point Cloud Reconstruction via Joint Geometry FeaturesPoster
  618. EdgeTAM: On-Device Track Anything ModelPoster
  619. Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative EditingPoster
  620. Effective SAM Combination for Open-Vocabulary Semantic SegmentationPoster
  621. EffiDec3D: An Optimized Decoder for High-Performance and Efficient 3D Medical Image SegmentationHighlight
  622. Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise ReplacementPoster
  623. Efficient Data Driven Mixture-of-Expert Extraction from Trained NetworksPoster
  624. Efficient Decoupled Feature 3D Gaussian Splatting via Hierarchical CompressionPoster
  625. Efficient Depth Estimation for Unstable Stereo Camera Systems on AR GlassesPoster
  626. Efficient Diffusion as Low Light EnhancerPoster
  627. Efficient Dynamic Scene Editing via 4D Gaussian-based Static-Dynamic SeparationPoster
  628. Efficient Event-Based Object Detection: A Hybrid Neural Network with Spatial and Temporal AttentionPoster
  629. Efficient Motion-Aware Video MLLMHighlight
  630. Efficient Personalization of Quantized Diffusion Model without BackpropagationPoster
  631. Efficient Test-time Adaptive Object Detection via Sensitivity-Guided PruningPoster
  632. Efficient Video Super-Resolution for Real-time Rendering with Decoupled G-buffer GuidancePoster
  633. EfficientLLaVA: Generalizable Auto-Pruning for Large Vision-language ModelsPoster
  634. Effortless Active Labeling for Long-Term Test-Time AdaptationPoster
  635. Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal InputPoster
  636. EigenGS Representation: From Eigenspace to Gaussian Image SpacePoster
  637. Electromyography-Informed Facial Expression Reconstruction for Physiological-Based Synthesis and AnalysisHighlight
  638. Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context LearningPoster
  639. EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video DiffusionPoster
  640. Empowering Large Language Models with 3D Situation AwarenessPoster
  641. Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video SynthesisPoster
  642. End-to-End HOI Reconstruction Transformer with Graph-based EncodingHighlight
  643. End-to-End Implicit Neural Representations for ClassificationPoster
  644. Enduring, Efficient and Robust Trajectory Prediction Attack in Autonomous Driving via Optimization-Driven Multi-Frame Perturbation FrameworkHighlight
  645. Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal RepresentationsPoster
  646. Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute RecognitionHighlight
  647. Enhanced then Progressive Fusion with View Graph for Multi-View ClusteringPoster
  648. Enhancing 3D Gaze Estimation in the Wild using Weak Supervision with Gaze Following LabelsPoster
  649. Enhancing Adversarial Transferability with Checkpoints of a Single Model's TrainingPoster
  650. Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion ModelPoster
  651. Enhancing Facial Privacy Protection via Weakening Diffusion PurificationPoster
  652. Enhancing Few-Shot Class-Incremental Learning via Training-Free Bi-Level Modality CalibrationPoster
  653. Enhancing Online Continual Learning with Plug-and-Play State Space Model and Class-Conditional Mixture of DiscretizationPoster
  654. Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion ModelsPoster
  655. Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image SegmentationPoster
  656. Enhancing Testing-Time Robustness for Trusted Multi-View Classification in the WildPoster
  657. Enhancing Video-LLM Reasoning via Agent-of-Thoughts DistillationPoster
  658. Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise SchedulingPoster
  659. EnliveningGS: Active Locomotion of 3DGSPoster
  660. EntityErasure: Erasing Entity Cleanly via Amodal Entity Segmentation and CompletionPoster
  661. EntitySAM: Segment Everything in VideoPoster
  662. EntropyMark: Towards More Harmless Backdoor Watermark via Entropy-based Constraint for Open-source Dataset Copyright ProtectionPoster
  663. EquiPose: Exploiting Permutation Equivariance for Relative Camera Pose EstimationPoster
  664. Erase Diffusion: Empowering Object Removal Through Calibrating Diffusion PathwaysHighlight
  665. Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent SpacesPoster
  666. EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with EventsHighlight
  667. EvOcc: Accurate Semantic Occupancy for Automated Driving Using Evidence TheoryPoster
  668. Event Ellipsometer: Event-based Mueller-Matrix Video ImagingHighlight
  669. Event-Equalized Dense Video CaptioningPoster
  670. EventFly: Event Camera Perception from Ground to the SkyPoster
  671. EventPSR: Surface Normal and Reflectance Estimation from Photometric Stereo Using an Event CameraHighlight
  672. Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and BeyondPoster
  673. Explainable Saliency: Articulating Reasoning with Contextual PrioritizationPoster
  674. Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image ClassificationPoster
  675. Explaining in Diffusion: Explaining a Classifier with Diffusion SemanticsPoster
  676. Explicit Depth-Aware Blurry Video Frame Interpolation Guided by Differential CurvesPoster
  677. Exploiting Deblurring Networks for Radiance FieldsPoster
  678. Exploration-Driven Generative Interactive EnvironmentsPoster
  679. Exploring Contextual Attribute Density in Referring Expression CountingPoster
  680. Exploring Historical Information for RGBE Visual Tracking with MambaPoster
  681. Exploring Scene Affinity for Semi-Supervised LiDAR Semantic SegmentationPoster
  682. Exploring Temporally-Aware Features for Point TrackingPoster
  683. Exploring Timeline Control for Facial Motion GenerationPoster
  684. Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image SynthesisPoster
  685. Exposure-slot: Exposure-centric Representations Learning with Slot-in-Slot Attention for Region-aware Exposure CorrectionPoster
  686. FADE: Frequency-Aware Diffusion Model Factorization for Video EditingPoster
  687. FALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene UnderstandingPoster
  688. FASTer: Focal token Acquiring-and-Scaling Transformer for Long-term 3D Objection DetectionPoster
  689. FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image EditingPoster
  690. FFR: Frequency Feature Rectification for Weakly Supervised Semantic SegmentationPoster
  691. FFaceNeRF: Few-shot Face Editing in Neural Radiance FieldsPoster
  692. FG^2: Fine-Grained Cross-View Localization by Fine-Grained Feature MatchingPoster
  693. FIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze EstimationPoster
  694. FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix ApproximationHighlight
  695. FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse ViewsPoster
  696. FLAVC: Learned Video Compression with Feature Level AttentionPoster
  697. FRAME: Floor-aligned Representation for Avatar Motion from Egocentric VideoHighlight
  698. FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question AnsweringPoster
  699. FRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few ImagesHighlight
  700. FSBench: A Figure Skating Benchmark for Advancing Artistic Sports UnderstandingPoster
  701. FSHNet: Fully Sparse Hybrid Network for 3D Object DetectionPoster
  702. Face Forgery Video Detection via Temporal Forgery Cue UnravelingPoster
  703. FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMsPoster
  704. FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation ModelsPoster
  705. Fast and Accurate Gigapixel Pathological Image Classification with Hierarchical Distillation Multi-Instance LearningPoster
  706. Faster Parameter-Efficient Tuning with Token Redundancy ReductionPoster
  707. Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object DetectionPoster
  708. Feature Selection for Latent Factor ModelsPoster
  709. Feature Spectrum Learning for Remote Sensing Change DetectionPoster
  710. Feature-Preserving Mesh Decimation for Normal IntegrationPoster
  711. FedAWA: Adaptive Optimization of Aggregation Weights in Federated Learning Using Client VectorsPoster
  712. FedCALM: Conflict-aware Layer-wise Mitigation for Selective Aggregation in Deeper Personalized Federated LearningPoster
  713. FedCS: Coreset Selection for Federated LearningPoster
  714. FedSPA: Generalizable Federated Graph Learning under Homophily HeterogeneityPoster
  715. FeedEdit: Text-Based Image Editing with Dynamic Feedback RegulationPoster
  716. Ferret: An Efficient Online Continual Learning Framework under Varying Memory ConstraintsPoster
  717. Few-shot Implicit Function Generation via EquivarianceHighlight
  718. Few-shot Personalized Scanpath PredictionPoster
  719. FiRe: Fixed-points of Restoration Priors for Solving Inverse ProblemsPoster
  720. Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part SegmentationPoster
  721. FinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal GuidancePoster
  722. Finer-CAM: Spotting the Difference Reveals Finer Details for Visual ExplanationPoster
  723. Fingerprinting Denoising Diffusion Probabilistic ModelsPoster
  724. Finsler Multi-Dimensional Scaling: Manifold Learning for Asymmetric Dimensionality Reduction and EmbeddingPoster
  725. FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized SegmentationPoster
  726. Fitted Neural Lossless Image CompressionPoster
  727. Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry LocalityHighlight
  728. FlexDrive: Toward Trajectory Flexibility in Driving Scene Gaussian Splatting Reconstruction and RenderingPoster
  729. FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian SplattingPoster
  730. FlexUOD: The Answer to Real-world Unsupervised Image Outlier DetectionPoster
  731. Flexible Frame Selection for Efficient Video ReasoningPoster
  732. Flexible Group Count Enables Hassle-Free Structured PruningPoster
  733. Flow-NeRF: Joint Learning of Geometry, Poses, and Dense Flow within Unified Neural RepresentationsPoster
  734. FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic ManipulationPoster
  735. Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality EvolutionHighlight
  736. FluxSpace: Disentangled Semantic Editing in Rectified Flow ModelsPoster
  737. Focal Split: Untethered Snapshot Depth from Differential DefocusPoster
  738. Focusing on Tracks for Online Multi-Object TrackingPoster
  739. Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional FlowsPoster
  740. Font-Agent: Enhancing Font Understanding with Large Language ModelsPoster
  741. ForestLPR: LiDAR Place Recognition in Forests Attentioning Multiple BEV Density ImagesHighlight
  742. Forming Auxiliary High-confident Instance-level Loss to Promote Learning from Label ProportionsPoster
  743. Fortifying Federated Learning Towards Trustworthiness via Auditable Data Valuation and Verifiable Client ContributionPoster
  744. FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image GenerationHighlight
  745. Foveated Instance SegmentationPoster
  746. Fractal Calibration for Long-tailed Object DetectionPoster
  747. Free Lunch Enhancements for Multi-modal Crowd CountingPoster
  748. Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EMPoster
  749. Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed ViewsPoster
  750. FreeCloth: Free-form Generation Enhances Challenging Clothed Human ModelingHighlight
  751. FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian VelocityPoster
  752. FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component AnalysisHighlight
  753. FreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free PromptsPoster
  754. FreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene ReconstructionPoster
  755. FreeUV: Ground-Truth-Free Realistic Facial UV Texture Recovery via Cross-Assembly Inference StrategyPoster
  756. FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency DebiasingPoster
  757. Frequency Dynamic Convolution for Dense Image PredictionPoster
  758. Frequency-Biased Synergistic Design for Image Compression and CompensationPoster
  759. From Elements to Design: A Layered Approach for Automatic Graphic Design CompositionPoster
  760. From Head to Tail: Efficient Black-box Model Inversion Attack via Long-tailed LearningPoster
  761. From Prototypes to General Distributions: An Efficient Curriculum for Masked Image ModelingPoster
  762. From Sparse Signal to Smooth Motion: Real-Time Motion Generation with Rolling Prediction ModelsPoster
  763. From Sparse to Dense: Camera Relocalization with Scene-Specific Detector from Feature Gaussian SplattingPoster
  764. From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral PerspectivePoster
  765. FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned PriorsPoster
  766. FruitNinja: 3D Object Interior Texture Generation with Gaussian SplattingPoster
  767. Functionality Understanding and Segmentation in 3D ScenesHighlight
  768. Fuzzy Multimodal Learning for Trusted Cross-modal RetrievalPoster
  769. GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context EncodingPoster
  770. GASP: Gaussian Avatars with Synthetic PriorsPoster
  771. GBC-Splat: Generalizable Gaussian-Based Clothed Human Digitalization under Sparse RGB CamerasPoster
  772. GBlobs: Explicit Local Structure via Gaussian Blobs for Improved Cross-Domain LiDAR-based 3D Object DetectionPoster
  773. GCC: Generative Color Constancy via Diffusing a Color CheckerPoster
  774. GENIUS: A Generative Framework for Universal Multimodal SearchPoster
  775. GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following ManipulationPoster
  776. GIF: Generative Inspiration for Face Recognition at ScalePoster
  777. GIFStream: 4D Gaussian-based Immersive Video with Feature StreamPoster
  778. GIVEPose: Gradual Intra-class Variation Elimination for RGB-based Category-Level Object Pose EstimationPoster
  779. GLASS: Guided Latent Slot Diffusion for Object-Centric LearningPoster
  780. GLane3D: Detecting Lanes with Graph of 3D KeypointsPoster
  781. GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object DetectorPoster
  782. GOAL: Global-local Object Alignment LearningPoster
  783. GPAvatar: High-fidelity Head Avatars by Learning Efficient Gaussian ProjectionsPoster
  784. GPS as a Control Signal for Image GenerationPoster
  785. GPVK-VL: Geometry-Preserving Virtual Keyframes for Visual Localization under Large Viewpoint ChangesPoster
  786. GRAE-3DMOT: Geometry Relation-Aware Encoder for Online 3D Multi-Object TrackingPoster
  787. GS-2DGS: Geometrically Supervised 2DGS for Reflective Object ReconstructionPoster
  788. GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point TrackingPoster
  789. GUI-Xplore: Empowering Generalizable GUI Agents with One ExplorationPoster
  790. GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D ReconstructionPoster
  791. Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring ClassesPoster
  792. Galaxy Walker: Geometry-aware VLMs For Galaxy-scale UnderstandingHighlight
  793. GauCho: Gaussian Distributions with Cholesky Decomposition for Oriented Object DetectionPoster
  794. GaussHDR: High Dynamic Range Gaussian Splatting via Learning Unified 3D and 2D Local Tone MappingPoster
  795. Gaussian Splatting Feature Fields for (Privacy-Preserving) Visual LocalizationPoster
  796. GaussianIP: Identity-Preserving Realistic 3D Human Generation via Human-Centric Diffusion PriorPoster
  797. GazeGene: Large-scale Synthetic Gaze Dataset with 3D Eyeball AnnotationsPoster
  798. Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual ForagingPoster
  799. Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D ObjectsPoster
  800. GenAssets: Generating in-the-wild 3D Assets in Latent SpacePoster
  801. GenFusion: Closing the Loop between Reconstruction and Generation via VideosPoster
  802. GenPC: Zero-shot Point Cloud Completion via 3D Generative PriorsPoster
  803. GenVDM: Generating Vector Displacement Maps From a Single ImageHighlight
  804. Generalizable Object Keypoint Localization from Generative PriorsPoster
  805. Generalized Diffusion Detector: Mining Robust Features from Diffusion Models for Domain-Generalized DetectionPoster
  806. Generalized Gaussian Entropy Model for Point Cloud Attribute Compression with Dynamic Likelihood IntervalsPoster
  807. Generalized Zero-Shot Classification via Semantics-Free Inter-Class Feature GenerationPoster
  808. Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric VisionHighlight
  809. Generative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D ReconstructionHighlight
  810. Generative Hard Example Augmentation for Semantic Point Cloud SegmentationPoster
  811. Generative Map Priors for Collaborative BEV Semantic SegmentationPoster
  812. Generative Multimodal Pretraining with Discrete Diffusion Timestep TokensAward Candidate
  813. Generative PhotomontagePoster
  814. Generative Sparse-View Gaussian SplattingPoster
  815. Generative Zero-Shot Composed Image RetrievalPoster
  816. GeoAvatar: Geometrically-Consistent Multi-Person Avatar Reconstruction from Sparse Multi-View VideosPoster
  817. GeoDepth: From Point-to-Depth to Plane-to-Depth Modeling for Self-Supervised Monocular Depth EstimationPoster
  818. GeoMM: On Geodesic Perspective for Multi-modal LearningPoster
  819. Geometric Knowledge-Guided Localized Global Distribution Alignment for Federated LearningPoster
  820. Geometry in Style: 3D Stylization via Surface Normal DeformationPoster
  821. Geometry-guided Online 3D Video Synthesis with Multi-View Temporal ConsistencyPoster
  822. Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference UnderstandingPoster
  823. GliaNet: Adaptive Neural Network Structure Learning with Glia-DrivenPoster
  824. Glossy Object Reconstruction with Cost-effective Polarized AcquisitionHighlight
  825. GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text EditingPoster
  826. GoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View SynthesisPoster
  827. Golden Cudgel Network for Real-Time Semantic SegmentationPoster
  828. Gradient Inversion Attacks on Parameter-Efficient Fine-TuningPoster
  829. Gradient-Guided Annealing for Domain GeneralizationHighlight
  830. Graph Neural Network Combining Event Stream and Periodic Aggregation for Low-Latency Event-based VisionHighlight
  831. Graph-Embedded Structure-Aware Perceptual Hashing for Neural Network Protection and Piracy DetectionPoster
  832. GraphI2P: Image-to-Point Cloud Registration with Exploring Pattern of Correspondence via Graph LearningPoster
  833. GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy LearningPoster
  834. Gromov-Wasserstein Problem with Cyclic SymmetryPoster
  835. GroomLight: Hybrid Inverse Rendering for Relightable Human Hair Appearance ModelingPoster
  836. Ground-V: Teaching VLMs to Ground Complex Instructions in PixelsPoster
  837. Grounding 3D Object Affordance with Language Instructions, Visual Observations and InteractionsPoster
  838. GroundingFace: Fine-grained Face Understanding via Pixel Grounding Multimodal Large Language ModelHighlight
  839. GroupMamba: Efficient Group-Based Visual State Space ModelPoster
  840. GuardSplat: Efficient and Robust Watermarking for 3D Gaussian SplattingPoster
  841. H-MoRe: Learning Human-centric Motion Representation for Action AnalysisHighlight
  842. H2ST: Hierarchical Two-Sample Tests for Continual Out-of-Distribution DetectionPoster
  843. HERA: Hybrid Explicit Representation for Ultra-Realistic Head AvatarsPoster
  844. HMAR: Efficient Hierarchical Masked Auto-Regressive Image GenerationPoster
  845. HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video GenerationPoster
  846. HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture GenerationPoster
  847. HORP: Human-Object Relation Priors Guided HOI DetectionPoster
  848. HOT: Hadamard-based Optimized TrainingPoster
  849. HOTFormerLoc: Hierarchical Octree Transformer for Versatile Lidar Place Recognition Across Ground and Aerial ViewsPoster
  850. HRAvatar: High-Quality and Relightable Gaussian Head AvatarPoster
  851. HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene InteractionHighlight
  852. HSI: A Holistic Style Injector for Arbitrary Style TransferPoster
  853. HUNet: Homotopy Unfolding Network for Image Compressive SensingPoster
  854. HUSH: Holistic Panoramic 3D Scene Understanding using Spherical HarmonicsPoster
  855. HalLoc: Token-level Localization of Hallucinations for Vision Language ModelsPoster
  856. Hand-held Object Reconstruction from RGB Video with Dynamic InteractionPoster
  857. HandOS: 3D Hand Reconstruction in One StagePoster
  858. Hardware-Rasterized Ray-Based Gaussian SplattingHighlight
  859. Harnessing Frequency Spectrum Insights for Image Copyright Protection Against Diffusion ModelsPoster
  860. Harnessing Frozen Unimodal Encoders for Flexible Multimodal AlignmentPoster
  861. Harnessing Global-Local Collaborative Adversarial Perturbation for Anti-CustomizationPoster
  862. Hazy Low-Quality Satellite Video Restoration Via Learning Optimal Joint Degradation Patterns and Continuous-Scale Super-Resolution ReconstructionPoster
  863. HeMoRa: Unsupervised Heuristic Consensus Sampling for Robust Point Cloud RegistrationPoster
  864. Hearing Anywhere in Any EnvironmentPoster
  865. Hearing Hands: Generating Sounds from Physical Interactions in 3D ScenesPoster
  866. HeatFormer: A Neural Optimizer for Multiview Human Mesh RecoveryPoster
  867. Heterogeneous Skeleton-Based Action Representation LearningPoster
  868. HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face FusionPoster
  869. HiMoR: Monocular Deformable Gaussian Reconstruction with Hierarchical Motion RepresentationPoster
  870. Hiding Images in Diffusion Models by Editing Learned Score FunctionsPoster
  871. Hierarchical Adaptive Filtering Network for Text Image Specular Highlight RemovalPoster
  872. Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric LearningPoster
  873. Hierarchical Features Matter: A Deep Exploration of Progressive Parameterization Method for Dataset DistillationPoster
  874. Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D GenerationPoster
  875. Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time AdaptationPoster
  876. High Dynamic Range Video Compression: A Large-Scale Benchmark Dataset and A Learned Bit-depth Scalable Compression AlgorithmPoster
  877. High Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous FlightPoster
  878. High-Fidelity Lightweight Mesh Reconstruction from Point CloudsHighlight
  879. High-fidelity 3D Object Generation from Single Image with RGBN-Volume Gaussian Reconstruction ModelHighlight
  880. High-quality Point Cloud Oriented Normal Estimation via Hybrid Angular and Euclidean Distance EncodingPoster
  881. Higher-Order Ratio Cycles for Fast and Globally Optimal Shape MatchingPoster
  882. HistoFS: Non-IID Histopathologic Whole Slide Image Classification via Federated Style Transfer with RoI-PreservingPoster
  883. HoGS: Unified Near and Far Object Reconstruction via Homogeneous Gaussian SplattingPoster
  884. HomoGen: Enhanced Video Inpainting via Homography Propagation and DiffusionPoster
  885. Homogeneous Dynamics Space for Heterogeneous HumansPoster
  886. HotSpot: Signed Distance Function Optimization with an Asymptotically Sufficient ConditionHighlight
  887. HuMoCon: Concept Discovery for Human Motion UnderstandingPoster
  888. HuPerFlow: A Comprehensive Benchmark for Human vs. Machine Motion Estimation ComparisonHighlight
  889. Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identificationPoster
  890. HumanMM: Global Human Motion Recovery from Multi-shot VideosPoster
  891. Hybrid Concept Bottleneck ModelsPoster
  892. Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image SegmentationPoster
  893. Hybrid Reciprocal Transformer with Triplet Feature Alignment for Scene Graph GenerationPoster
  894. HybridMQA: Exploring Geometry-Texture Interactions for Colored Mesh Quality AssessmentPoster
  895. HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing ImageryPoster
  896. HyperLoRA: Parameter-Efficient Adaptive Generation for Portrait SynthesisHighlight
  897. HyperNVD: Accelerating Neural Video Decomposition via HypernetworksPoster
  898. HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight TrajectoriesPoster
  899. HyperPose: Hypernetwork-Infused Camera Pose Localization and an Extended Cambridge Landmarks DatasetPoster
  900. HyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual PerceiverPoster
  901. Hyperbolic Category DiscoveryPoster
  902. Hyperbolic Safety-Aware Vision-Language ModelsHighlight
  903. Hyperbolic Uncertainty-Aware Few-Shot Incremental Point Cloud SegmentationPoster
  904. Hypergraph Vision Transformers: Images are More than Nodes, More than EdgesPoster
  905. Hyperspectral Pansharpening via Diffusion Models with Iteratively Zero-Shot GuidancePoster
  906. I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video ModelsPoster
  907. IAAO: Interactive Affordance Learning for Articulated Objects in 3D EnvironmentsPoster
  908. ICE: Intrinsic Concept Extraction from a Single Image via Diffusion ModelsHighlight
  909. ICP: Immediate Compensation Pruning for Mid-to-high SparsityHighlight
  910. IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular VideosCPoster
  911. IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot MannerPoster
  912. IMFine: 3D Inpainting via Geometry-guided Multi-view RefinementPoster
  913. ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-OnPoster
  914. IceDiff: High Resolution and High-Quality Arctic Sea Ice Forecasting with Generative Diffusion PriorPoster
  915. Identifying and Mitigating Spurious Correlation in Multi-Task LearningPoster
  916. Identity-Clothing Similarity Modeling for Unsupervised Clothing Change Person Re-IdentificationPoster
  917. Identity-preserving Distillation Sampling by Fixed-Point IteratorPoster
  918. Illumination Spectrum Estimation for Multispectral Images via Surface Reflectance Modeling and Spatial-Spectral Feature GenerationPoster
  919. ImViD: Immersive Volumetric Videos for Enhanced VR EngagementHighlight
  920. Image Generation Diversity Issues and How to Tame ThemPoster
  921. Image Over Text: Transforming Formula Recognition Evaluation with Character Detection MatchingPoster
  922. Image Quality Assessment: From Human to Machine PreferenceHighlight
  923. Image Quality Assessment: Investigating Causal Perceptual Effects with Abductive Counterfactual InferencePoster
  924. Image Reconstruction from Readout-Multiplexed Single-Photon Detector ArraysHighlight
  925. Image Referenced Sketch Colorization Based on Animation Creation WorkflowPoster
  926. Image is All You Need to Empower Large-scale Diffusion Models for In-Domain GenerationPoster
  927. ImagineFSL: Self-Supervised Pretraining Matters on Imagined Base Set for VLM-based Few-shot LearningHighlight
  928. Implicit Bias Injection Attacks against Text-to-Image Diffusion ModelsPoster
  929. Implicit Correspondence Learning for Image-to-Point Cloud RegistrationHighlight
  930. Improve Representation for Imbalanced Regression through Geometric ConstraintsPoster
  931. Improved Monocular Depth Prediction Using Distance Transform Over Pre-semantic Contours with Self-supervised Neural NetworksPoster
  932. Improving Accuracy and Calibration via Differentiated Deep Mutual LearningPoster
  933. Improving Editability in Image Generation with Layer-wise MemoryPoster
  934. Improving Gaussian Splatting with Localized Points ManagementHighlight
  935. Improving Personalized Search with Regularized Low-Rank Parameter UpdatesHighlight
  936. Improving Semi-Supervised Semantic Segmentation with Sliced-Wasserstein Feature Alignment and UniformityPoster
  937. Improving Sound Source Localization with Joint Slot Attention on Image and AudioPoster
  938. Improving Transferable Targeted Attacks with Feature Tuning MixupPoster
  939. Improving Visual and Downstream Performance of Low-Light Enhancer with Vision Foundation Models CollaborationPoster
  940. Improving the Training of Data-Efficient GANs via Quality Aware Dynamic Discriminator Rejection SamplingPoster
  941. Improving the Transferability of Adversarial Attacks on Face Recognition with Diverse Parameters AugmentationPoster
  942. Imputation-free and Alignment-free: Incomplete Multi-view Clustering Driven by Consensus Semantic LearningPoster
  943. InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model AlignmentHighlight
  944. Incomplete Multi-View Multi-label Learning via Disentangled Representation and Label Semantic EmbeddingPoster
  945. Incomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space ModelPoster
  946. Incorporating Dense Knowledge Alignment into Unified Multimodal Representation ModelsPoster
  947. Incremental Object Keypoint LearningPoster
  948. IndoorGS: Geometric Cues Guided Gaussian Splatting for Indoor Scene ReconstructionPoster
  949. Inference-Scale Complexity in ANN-SNN Conversion for High-Performance and Low-Power ApplicationsPoster
  950. Infighting in the Dark: Multi-Label Backdoor Attack in Federated LearningPoster
  951. InsTaG: Learning Personalized 3D Talking Head from Few-Second VideoPoster
  952. Insightful Instance Features for 3D Instance SegmentationPoster
  953. Instance-wise Supervision-level Optimization in Active LearningPoster
  954. Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive LearningPoster
  955. Integral Fast Fourier Color ConstancyPoster
  956. InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance ParsingHighlight
  957. InteractionMap: Improving Online Vectorized HDMap Construction with InteractionPoster
  958. Interactive Medical Image Analysis with Concept-based Similarity ReasoningPoster
  959. Interpretable Generative Models through Post-hoc Concept BottlenecksPoster
  960. Interpretable Image Classification via Non-parametric Part Prototype LearningPoster
  961. Inversion Circle Interpolation: Diffusion-based Image Augmentation for Data-scarce ClassificationPoster
  962. Investigating the Role of Weight Decay in Enhancing Nonconvex SGDPoster
  963. Invisible Backdoor Attack against Self-supervised LearningPoster
  964. Iterative Predictor-Critic Code Decoding for Real-World Image DehazingPoster
  965. JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV SystemsPoster
  966. JiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World DataPoster
  967. Joint Optimization of Neural Radiance Fields and Continuous Camera Motion from a Monocular VideoPoster
  968. Joint Scheduling of Causal Prompts and Tasks for Multi-Task LearningPoster
  969. Joint Vision-Language Social Bias Removal for CLIPPoster
  970. Just Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly DetectionHighlight
  971. KAC: Kolmogorov-Arnold Classifier for Continual LearningHighlight
  972. KMD: Koopman Multi-modality Decomposition for Generalized Brain Tumor Segmentation under Incomplete ModalitiesPoster
  973. KVQ: Boosting Video Quality Assessment via Saliency-guided Local PerceptionPoster
  974. Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic SegmentationPoster
  975. Keyframe-Guided Creative Video InpaintingPoster
  976. Knowledge Bridger: Towards Training-Free Missing Modality CompletionPoster
  977. Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion RecognitionPoster
  978. L-SWAG: Layer-Sample Wise Activation with Gradients Information for Zero-Shot NAS on Vision TransformersPoster
  979. LAL: Enhancing 3D Human Motion Prediction with Latency-aware Auxiliary LearningPoster
  980. LATTE-MV: Learning to Anticipate Table Tennis Hits from Monocular VideosPoster
  981. LC-Mamba: Local and Continuous Mamba with Shifted Windows for Frame InterpolationPoster
  982. LEDiff: Latent Exposure Diffusion for HDR GenerationPoster
  983. LIM: Large Interpolator Model for Dynamic ReconstructionPoster
  984. LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance FieldsPoster
  985. LLM-driven Multimodal and Multi-Identity Listening Head GenerationPoster
  986. LMO: Linear Mamba Operator for MRI ReconstructionPoster
  987. LOCORE: Image Re-ranking with Long-Context Sequence ModelingPoster
  988. LOD-GS: Achieving Levels of Detail using Scalable Gaussian SoupPoster
  989. LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot LearningPoster
  990. LP-Diff: Towards Improved Restoration of Real-World Degraded License PlateHighlight
  991. LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic SegmentationPoster
  992. LSNet: See Large, Focus SmallPoster
  993. LUCAS: Layered Universal Codec AvatarsPoster
  994. Label Shift Meets Online Learning: Ensuring Consistent Adaptation with Universal Dynamic RegretHighlight
  995. Language Guided Concept Bottleneck Models for Interpretable Continual LearningPoster
  996. Language-Assisted Debiasing and Smoothing for Foundation Model-Based Semi-Supervised LearningPoster
  997. Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentPoster
  998. Language-Guided Salient Object RankingPoster
  999. Large Self-Supervised Models Bridge the Gap in Domain Adaptive Object DetectionPoster
  1000. Latent Drifting in Diffusion Models for Counterfactual Medical Image SynthesisHighlight

Looking for submission deadlines instead? See the conference deadline calendar.