CVPR 2025 Accepted Papers
The full list of 2,872 papers accepted at CVPR 2025 (IEEE/CVF Conference on Computer Vision and Pattern Recognition). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
Poster: 2,468Highlight: 388Award Candidate: 15
- Latent Space ImagingPoster
- LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand Diffusion.Poster
- Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric VideosPoster
- Learnable Infinite Taylor Gaussian for Dynamic View RenderingPoster
- Learned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus CuesPoster
- Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual SceneHighlight
- Learning Audio-guided Video Representation with Gated Attention for Video-Text RetrievalPoster
- Learning Class Prototypes for Unified Sparse-Supervised 3D Object DetectionHighlight
- Learning Compatible Multi-Prize Subnetworks for Asymmetric RetrievalPoster
- Learning Conditional Space-Time Prompt Distributions for Video Class-Incremental LearningHighlight
- Learning Dynamic Collaborative Network for Semi-supervised 3D Vessel SegmentationPoster
- Learning Endogenous Attention for Incremental Object DetectionPoster
- Learning Hazing to Dehazing: Towards Realistic Haze Generation for Real-World Image DehazingPoster
- Learning Heterogeneous Tissues with Mixture of Experts for Gigapixel Whole Slide ImagesPoster
- Learning Occlusion-Robust Vision Transformers for Real-Time UAV TrackingPoster
- Learning Partonomic 3D Reconstruction from Image CollectionsPoster
- Learning Person-Specific Animatable Face Models from In-the-Wild Images via a Shared Base ModelPoster
- Learning Phase Distortion with Selective State Space Models for Video Turbulence MitigationHighlight
- Learning Physics From Video: Unsupervised Physical Parameter Estimation for Continuous Dynamical SystemsPoster
- Learning Physics-Based Full-Body Human Reaching and Grasping from Brief Walking ReferencesPoster
- Learning Textual Prompts for Open-World Semi-Supervised LearningPoster
- Learning Visual Composition through Improved Semantic GuidancePoster
- Learning from Neighbors: Category Extrapolation for Long-Tail LearningPoster
- Learning from Streaming Video with Orthogonal GradientsPoster
- Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging ScenesPoster
- Learning on Model Weights using Tree ExpertsPoster
- Learning to Detect Objects from Multi-Agent LiDAR Scans without Manual LabelsPoster
- Learning to Filter Outlier Edges in Global SfMHighlight
- Learning to Highlight Audio by Watching MoviesPoster
- Learning to Normalize on the SPD Manifold under Bures-Wasserstein GeometryPoster
- Learning with Noisy Triplet Correspondence for Composed Image RetrievalPoster
- Learning-enabled Polynomial Lyapunov Function Synthesis via High-Accuracy Counterexample-Guided FrameworkPoster
- LesionLocator: Zero-Shot Universal Tumor Segmentation and Tracking in 3D Whole-Body ImagingPoster
- Less Attention is More: Prompt Transformer for Generalized Category DiscoveryPoster
- Less is More: Efficient Image Vectorization with Adaptive ParameterizationPoster
- Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual RecognitionHighlight
- Let Humanoids Hike! Integrative Skill Development on Complex TrailsPoster
- Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of SamplesPoster
- Let's Chorus: Partner-aware Hybrid Song-Driven 3D Head AnimationPoster
- Let's Verify and Reinforce Image Generation Step by StepPoster
- Leveraging 3D Geometric Priors in 2D Rotation Symmetry DetectionPoster
- Leveraging Global Stereo Consistency for Category-Level Shape and 6D Pose Estimation from Stereo ImagesPoster
- Leveraging Perturbation Robustness to Enhance Out-of-Distribution DetectionPoster
- Leveraging SD Map to Augment HD Map-based Trajectory PredictionPoster
- Leveraging Temporal Cues for Semi-Supervised Multi-View 3D Object DetectionPoster
- LiSu: A Dataset and Method for LiDAR Surface Normal EstimationPoster
- Libra-Merging: Importance-redundancy and Pruning-merging Trade-off for Acceleration Plug-in in Large Vision-Language ModelPoster
- LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer AttributionsPoster
- LidarGait++: Learning Local Features and Size Awareness from LiDAR Point Clouds for 3D Gait RecognitionPoster
- Lifelong Knowledge Editing for Vision Language Models with Low-Rank Mixture-of-ExpertsPoster
- Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic ManipulationPoster
- Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster InferencePoster
- Light Transport-aware Diffusion Posterior Sampling for Single-View Reconstruction of 3D VolumesHighlight
- LightLoc: Learning Outdoor LiDAR Localization at Light SpeedPoster
- Linear Attention Modeling for Learned Image CompressionPoster
- Link to the Past: Temporal Propagation for Fast 3D Human Reconstruction from Monocular VideoPoster
- Link-based Contrastive Learning for One-Shot Unsupervised Domain AdaptationPoster
- LiveCC: Learning Video LLM with Streaming Speech Transcription at ScalePoster
- LoKi: Low-dimensional KAN for Efficient Fine-tuning Image ModelsPoster
- LoRA Recycle: Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAsPoster
- LoRACLR: Contrastive Adaptation for Customization of Diffusion ModelsPoster
- LoTUS: Large-Scale Machine Unlearning with a Taste of UncertaintyPoster
- Locality-Aware Zero-Shot Human-Object Interaction DetectionPoster
- Locally Orderless Images for Optimization in Differentiable RenderingHighlight
- Logits DeConfusion with CLIP for Few-Shot LearningPoster
- LogoSP: Local-global Grouping of Superpoints for Unsupervised Semantic Segmentation of 3D Point CloudsPoster
- LongDiff: Training-Free Long Video Generation in One GoPoster
- LookingGlass: Generative Anamorphoses via Laplacian Pyramid WarpingPoster
- LotusFilter: Fast Diverse Nearest Neighbor Search via a Learned Cutoff TablePoster
- Low-Biased General Annotated Dataset GenerationPoster
- Low-Rank Adaptation in Multilinear Operator Networks for Security-Preserving Incremental LearningPoster
- Luminance-GS: Adapting 3D Gaussian Splatting to Challenging Lighting Conditions with View-Adaptive Curve AdjustmentPoster
- M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world SettingsPoster
- M3amba: Memory Mamba is All You Need for Whole Slide Image ClassificationPoster
- MAC-Ego3D: Multi-Agent Gaussian Consensus for Real-Time Collaborative Ego-Motion and Photorealistic 3D ReconstructionPoster
- MAD: Memory-Augmented Detection of 3D ObjectsPoster
- MAGE : Single Image to Material-Aware 3D via the Multi-View G-Buffer Estimation ModelPoster
- MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action AnticipationPoster
- MARBLE: Material Recomposition and Blending in CLIP-SpacePoster
- MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content CreationPoster
- MATCHA: Towards Matching AnythingHighlight
- MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse ViewsHighlight
- MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image GenerationPoster
- MDP: Multidimensional Vision Model Pruning with Latency ConstraintPoster
- MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh AttentionPoster
- MEET: Towards Memory-Efficient Temporal Sparse Deep Neural NetworksPoster
- MERGE: Multi-faceted Hierarchical Graph-based GNN for Gene Expression Prediction from Whole Slide Histopathology ImagesPoster
- MESC-3D:Mining Effective Semantic Cues for 3D Reconstruction from a Single ImagePoster
- METASCENES: Towards Automated Replica Creation for Real-world 3D ScansPoster
- MExD: An Expert-Infused Diffusion Model for Whole-Slide Image ClassificationPoster
- MFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and ForecastingPoster
- MI-DETR: An Object Detection Model with Multi-time Inquiries MechanismPoster
- MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal OutputPoster
- MIRE: Matched Implicit Neural RepresentationsPoster
- MITracker: Multi-View Integration for Visual Object TrackingHighlight
- MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio SynthesisPoster
- MNE-SLAM: Multi-Agent Neural SLAM for Mobile RobotsPoster
- MODA: Motion-Drift Augmentation for Inertial Human Motion AnalysisPoster
- MOS-Attack: A Scalable Multi-objective Adversarial Attack FrameworkPoster
- MOS: Modeling Object-Scene Associations in Generalized Category DiscoveryPoster
- MP-SfM: Monocular Surface Priors for Robust Structure-from-MotionPoster
- MTADiffusion: Mask Text Alignment Diffusion Model for Object InpaintingPoster
- MUST: The First Dataset and Unified Framework for Multispectral UAV Single Object TrackingPoster
- MV-SSM: Multi-View State Space Modeling for 3D Human Pose EstimationPoster
- MVDoppler-Pose: Multi-Modal Multi-View mmWave Sensing for Long-Distance Self-Occluded Human Walking Pose EstimationPoster
- MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion ModelPoster
- M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object SegmentationPoster
- MaDCoW: Marginal Distortion Correction for Wide-Angle Photography with Arbitrary ObjectsPoster
- MaSS13K: A Matting-level Semantic Segmentation BenchmarkPoster
- MagicArticulate: Make Your 3D Models Articulation-ReadyPoster
- Making Old Film Great Again: Degradation-aware State Space Model for Old Film RestorationPoster
- Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic SegmentationHighlight
- Mamba-Adaptor: State Space Model Adaptor for Visual RecognitionPoster
- Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space ModelsPoster
- MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training SmoothingPoster
- MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss AlpsHighlight
- MarkushGrapher: Joint Visual and Textual Recognition of Markush StructuresPoster
- Mask-Adapter: The Devil is in the Masks for Open-Vocabulary SegmentationPoster
- Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene UnderstandingPoster
- Masking meets Supervision: A Strong Learning AlliancePoster
- Matrix-Free Shared Intrinsics Bundle AdjustmentPoster
- Medusa: A Multi-Scale High-order Contrastive Dual-Diffusion Approach for Multi-View ClusteringPoster
- Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured MeshesPoster
- MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationHighlight
- Meta-Learning Hyperparameters for Parameter Efficient Fine-TuningHighlight
- MetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt TuningPoster
- MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural RepresentationHighlight
- Mimic In-Context Learning for Multimodal TasksPoster
- Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-MismatchPoster
- Mind the Gap: Detecting Black-box Adversarial Attacks in the Making through Query Update AnalysisPoster
- Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive JailbreakingHighlight
- Minding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion SegmentationPoster
- Minimal Interaction Seperated Tuning: A New Paradigm for Visual AdaptationPoster
- Minimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance SegmentationPoster
- Minority-Focused Text-to-Image Generation via Prompt OptimizationPoster
- MirrorVerse: Pushing Diffusion Models to Realistically Reflect the WorldPoster
- Mitigating Ambiguities in 3D Classification with Gaussian SplattingPoster
- Mixture of Submodules for Domain Adaptive Person SearchPoster
- MoEdit: On Learning Quantity Perception for Multi-object Image EditingPoster
- MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based DistillationPoster
- MoST: Efficient Monarch Sparse Tuning for 3D Representation LearningPoster
- MobileH2R: Learning Generalizable Human to Mobile Robot Handover Exclusively from Scalable and Diverse Synthetic DataPoster
- ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode ModelingPoster
- Model Diagnosis and Correction via Linguistic and Implicit Attribute EditingPoster
- Modeling Multiple Normal Action Representations for Error Detection in Procedural TasksPoster
- Mono2Stereo: A Benchmark and Empirical Study for Stereo ConversionPoster
- Mono3DVLT: Monocular-Video-Based 3D Visual Language TrackingPoster
- MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular DetectionPoster
- MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation ModelsPoster
- Monocular and Generalizable Gaussian Talking Head AnimationPoster
- Morpheus: Text-Driven 3D Gaussian Splat Shape and Color StylizationPoster
- Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph LearningPoster
- MotiF: Making Text Count in Image Animation with Motion Focal LossPoster
- MotionMap: Representing Multimodality in Human Pose ForecastingPoster
- MotionPRO: Exploring the Role of Pressure in Human MoCap and BeyondHighlight
- MotionPro: A Precise Motion Controller for Image-to-Video GenerationPoster
- Motions as Queries: One-Stage Multi-Person Holistic Human Motion CapturePoster
- Move-in-2D: 2D-Conditioned Human Motion GenerationPoster
- MuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image TranslationPoster
- Multi-Group Proportional Representations for Text-to-Image ModelsPoster
- Multi-Label Prototype Visual Spatial Search for Weakly Supervised Semantic SegmentationHighlight
- Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best PracticesPoster
- Multi-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEsPoster
- Multi-Modal Contrastive Masked Autoencoders: A Two-Stage Progressive Pre-training Approach for RGBD DatasetsPoster
- Multi-Modal Synergistic Implicit Image Enhancement for Efficient Optical Flow EstimationPoster
- Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual RepresentationPoster
- Multi-Scale Neighborhood Occupancy Masked Autoencoder for Self-Supervised Learning in LiDAR Point CloudsPoster
- Multi-View Pose-Agnostic Change Localization with Zero LabelsPoster
- Multi-focal Conditioned Latent Diffusion for Person Image SynthesisPoster
- Multi-modal Contrastive Learning with Negative Sampling Calibration for Phenotypic Drug DiscoveryPoster
- Multi-modal Knowledge Distillation-based Human Trajectory ForecastingPoster
- Multi-modal Medical Diagnosis via Large-small Model CollaborationPoster
- Multi-modal Topology-embedded Graph Learning for Spatially Resolved Genes Prediction from Pathology Images with Prior Gene Similarity InformationPoster
- Multi-modal Vision Pre-training for Medical Image AnalysisHighlight
- Multi-subject Open-set Personalization in Video GenerationPoster
- MultiMorph: On-demand Atlas ConstructionPoster
- MultimodalStudio: A Heterogeneous Sensor Dataset and Framework for Neural Rendering across Multiple Imaging ModalitiesPoster
- Multirate Neural Image Compression with Adaptive Lattice Vector QuantizationHighlight
- NADER: Neural Architecture Design via Multi-Agent CollaborationPoster
- NLPrompt: Noise-Label Prompt Learning for Vision-Language ModelsHighlight
- NN-Former: Rethinking Graph Structure in Neural Architecture RepresentationPoster
- NSD-Imagery: A Benchmark Dataset for Extending fMRI Vision Decoding Methods to Mental ImageryHighlight
- NTClick: Achieving Precise Interactive Segmentation With Noise-tolerant ClicksHighlight
- Navigating Image Restoration with VAR's Distribution Alignment PriorPoster
- Navigating the Unseen: Zero-shot Scene Graph Generation via Capsule-Based Equivariant FeaturesPoster
- NeISF++: Neural Incident Stokes Field for Polarized Inverse Rendering of Conductors and DielectricsPoster
- Nearly Zero-Cost Protection Against Mimicry by Personalized Diffusion ModelsPoster
- NeighborRetr: Balancing Hub Centrality in Cross-Modal RetrievalPoster
- Neural Hierarchical Decomposition for Single Image Plant ModelingPoster
- Neural Inverse Rendering from Propagating LightPoster
- Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light DiffusionPoster
- Neural Motion Simulator Pushing the Limit of World Models in Reinforcement LearningPoster
- Neural Video Compression with Context ModulationPoster
- NexusGS: Sparse View Synthesis with Epipolar Depth Priors in 3D Gaussian SplattingHighlight
- NightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene SegmentationPoster
- No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image RecognitionHighlight
- No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse WeatherPoster
- Noise Calibration and Spatial-Frequency Interactive Network for STEM Image EnhancementPoster
- Noise Diffusion for Enhancing Semantic Faithfulness in Text-to-Image SynthesisPoster
- Noise Modeling in One Hour: Minimizing Preparation Efforts for Self-supervised Low-Light RAW Image DenoisingPoster
- Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and SegmentationPoster
- Noise-Resistant Video Anomaly Detection via RGB Error-Guided Multiscale Predictive Coding and Dynamic MemoryPoster
- NoiseCtrl: A Sampling-Algorithm-Agnostic Conditional Generation Method for Diffusion ModelsPoster
- Non-Natural Image Understanding with Advancing Frequency-based Vision EncodersPoster
- Nonisotropic Gaussian Diffusion for Realistic 3D Human Motion PredictionPoster
- Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation AbilityPoster
- Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation ModelsPoster
- Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language ModelsHighlight
- Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question AnsweringPoster
- O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language ModelsHighlight
- OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation TriadPoster
- ODA-GAN: Orthogonal Decoupling Alignment GAN Assisted by Weakly-supervised Learning for Virtual Immunohistochemistry StainingPoster
- ODHSR: Online Dense 3D Reconstruction of Humans and Scenes from Monocular VideosPoster
- OFER: Occluded Face Expression ReconstructionPoster
- ONDA-Pose: Occlusion-Aware Neural Domain Adaptation for Self-Supervised 6D Object Pose EstimationPoster
- OODD: Test-time Out-of-Distribution Detection with Dynamic DictionaryPoster
- OPTICAL: Leveraging Optimal Transport for Contribution Allocation in Dataset DistillationHighlight
- ORIDa: Object-centric Real-world Image Composition DatasetPoster
- OW-OVD: Unified Open World and Open Vocabulary Object DetectionPoster
- Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic ManipulationPoster
- Object-aware Sound Source Localization via Audio-Visual Scene UnderstandingPoster
- Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object RecognitionPoster
- Odd-One-Out: Anomaly Detection by Comparing with NeighborsPoster
- OffsetOPT: Explicit Surface Reconstruction without NormalsPoster
- OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video ContextsPoster
- OmniSplat: Taming Feed-Forward 3D Gaussian Splatting for Omnidirectional Images with Editable CapabilitiesHighlight
- OmniStereo: Real-time Omnidireactional Depth Estimation with Multiview Fisheye CamerasPoster
- OmniStyle: Filtering High Quality Style Transfer Data at ScalePoster
- Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric VideosPoster
- Omnidirectional Multi-Object TrackingPoster
- On Denoising Walking Videos for Gait RecognitionPoster
- On the Out-Of-Distribution Generalization of Large Multimodal ModelsPoster
- On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free ApproachPoster
- On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only EventsPoster
- Once-Tuning-Multiple-Variants: Tuning Once and Expanded as Multiple Vision-Language Model VariantsPoster
- One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image FusionPoster
- One is Plenty: A Polymorphic Feature Interpreter for Immutable Heterogeneous Collaborative PerceptionPoster
- One-Step Event-Driven High-Speed AutofocusHighlight
- One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion ModelsPoster
- One-shot 3D Object Canonicalization based on Geometric and Semantic ConsistencyHighlight
- One2Any: One-Reference 6D Pose Estimation for Any ObjectPoster
- Online Task-Free Continual Learning via Dynamic Expansionable Memory DistributionPoster
- Online Video Understanding: OVBench and VideoChat-OnlinePoster
- OnlineAnySeg: Online Zero-Shot 3D Segmentation by Visual Foundation Model Guided 2D Mask MergingPoster
- Open Ad-hoc Categorization with Contextualized Feature LearningPoster
- Open Set Label Shift with Test Time Out-of-Distribution ReferencePoster
- Open-Canopy: Towards Very High Resolution Forest MonitoringHighlight
- OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text GenerationPoster
- OpenMIBOOD: Open Medical Imaging Benchmarks for Out-Of-Distribution DetectionPoster
- OpenSDI: Spotting Diffusion-Generated Images in the Open WorldPoster
- Opportunistic Single-Photon Time of FlightPoster
- OpticalNet: An Optical Imaging Dataset and Benchmark Beyond the Diffraction LimitHighlight
- Optimal Transport-Guided Source-Free Adaptation for Face Anti-SpoofingPoster
- Optimizing for the Shortest Path in Denoising Diffusion ModelHighlight
- Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned PolicyPoster
- OralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware MiningPoster
- OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic KernelsPoster
- Overcoming Shortcut Problem in VLM for Robust Out-of-Distribution DetectionHighlight
- PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language ModelsPoster
- PAVE: Patching and Adapting Video Large Language ModelsPoster
- PBR-NeRF: Inverse Rendering with Physics-Based Neural FieldsPoster
- PCM : Picard Consistency Model for Fast Parallel Sampling of Diffusion ModelsPoster
- PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic ManipulationPoster
- PEER Pressure: Model-to-Model Regularization for Single Source Domain GeneralizationPoster
- PGC: Physics-Based Gaussian Cloth from a Single PoseHighlight
- PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric VideosPoster
- PI-HMR: Towards Robust In-bed Temporal Human Shape Reconstruction with Contact Pressure SensingPoster
- PIAD: Pose and Illumination agnostic Anomaly DetectionPoster
- PICD: Versatile Perceptual Image Compression with Diffusion RenderingPoster
- PIDLoc: Cross-View Pose Optimization Network Inspired by PID ControllersPoster
- PIDSR: Complementary Polarized Image Demosaicing and Super-ResolutionPoster
- PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba AdapterPoster
- PMNI: Pose-free Multi-view Normal Integration for Reflective and Textureless Surface ReconstructionPoster
- POMP: Physics-consistent Motion Generative Model through Phase ManifoldsPoster
- POT: Prototypical Optimal Transport for Weakly Supervised Semantic SegmentationPoster
- POp-GS: Next Best View in 3D-Gaussian Splatting with P-OptimalityPoster
- PRaDA: Projective Radial Distortion AveragingPoster
- PS-Diffusion: Photorealistic Subject-Driven Image Editing with Disentangled Control and AttentionPoster
- PS-EIP: Robust Photometric Stereo Based on Event Interval ProfilePoster
- PSA-SSL: Pose and Size-aware Self-Supervised Learning on LiDAR Point CloudsPoster
- PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit RemeshingPoster
- PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion ModelPoster
- PURA: Parameter Update-Recovery Test-Time Adaption for RGB-T TrackingPoster
- PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language ModelsPoster
- PanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial AugmentationPoster
- PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene UnderstandingPoster
- Parallel Sequence Modeling via Generalized Spatial Propagation NetworkPoster
- Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear TransformationPoster
- Parameterized Blur Kernel Prior Learning for Local Motion DeblurringPoster
- Parametric Point Cloud Completion for Polygonal Surface ReconstructionPoster
- PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction ModelPoster
- PatchDEMUX: A Certifiably Robust Framework for Multi-label Classifiers Against Adversarial PatchesPoster
- PatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo AnomaliesPoster
- PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-ResolutionPoster
- Pattern Analogies: Learning to Perform Programmatic Image Edits by AnalogyPoster
- Percept, Memory, and Imagine: World Feature Simulating for Open-Domain Unknown Object DetectionPoster
- Perceptual Inductive Bias Is What You Need Before Contrastive LearningPoster
- Perceptual Video Compression with Neural WrappingPoster
- Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation MetricsHighlight
- Period-LLM: Extending the Periodic Capability of Multimodal Large Language ModelPoster
- Person De-reidentification: A Variation-guided Identity Shift ModelingPoster
- PersonaHOI: Effortlessly Improving Face Personalization in Human-Object Interaction GenerationPoster
- Perturb-and-Revise: Flexible 3D Editing with Generative TrajectoriesPoster
- Phoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action CorrectionPoster
- PhyS-EdiT: Physics-aware Semantic Image Editing with Text DescriptionPoster
- Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-TuningPoster
- Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot VisionPoster
- PlanarSplatting: Accurate Planar Surface Reconstruction in 3 MinutesHighlight
- Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept ControlPoster
- Plug-and-Play PPO: An Adaptive Point Prompt Optimizer Making SAM GreaterPoster
- Plug-and-Play Versatile Compressed Video EnhancementPoster
- Point Cloud Upsampling Using Conditional Diffusion Module with Adaptive Noise SuppressionPoster
- Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud AnalysisPoster
- Point-to-Region Loss for Semi-Supervised Point-Based Crowd CountingHighlight
- PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud LearningPoster
- PointSR: Self-Regularized Point Supervision for Drone-View Object DetectionPoster
- PolarFree: Polarization-based Reflection-Free ImagingPoster
- PolarNeXt: Rethink Instance Segmentation with Polar RepresentationPoster
- Polarized Color Screen MattingHighlight
- Poly-Autoregressive Prediction for Modeling InteractionsPoster
- Population Normalization for Federated LearningPoster
- Pos3R: 6D Pose Estimation for Unseen Objects Made EasyPoster
- Pose-Guided Temporal Enhancement for Robust Low-Resolution Hand ReconstructionPoster
- PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose EstimationPoster
- PoseTraj: Pose-Aware Trajectory Control in Video DiffusionPoster
- PosterO: Structuring Layout Trees to Enable Language Models in Generalized Content-Aware Layout GenerationPoster
- Potential Field Based Deep Metric LearningPoster
- Practical Solutions to the Relative Pose of Three Calibrated CamerasPoster
- Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class ImbalancePoster
- PreciseCam: Precise Camera Control for Text-to-Image GenerationPoster
- Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image EditingPoster
- Preserving Clusters in Prompt Learning for Unsupervised Domain AdaptationPoster
- Prior-free 3D Object TrackingHighlight
- ProHOC: Probabilistic Hierarchical Out-of-Distribution Classification via Multi-Depth NetworksPoster
- Probabilistic Prompt Distribution Learning for Animal Pose EstimationPoster
- ProbeSDF: Light Field Probes For Neural Surface ReconstructionPoster
- Probing the Mid-level Vision Capabilities of Self-Supervised LearningPoster
- Prof. Robot: Differentiable Robot Rendering Without Static and Self-CollisionsPoster
- Progress-Aware Video Frame CaptioningPoster
- Progressive Correspondence Regenerator for Robust 3D RegistrationPoster
- Progressive Focused Transformer for Single Image Super-ResolutionPoster
- Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D DataPoster
- ProjAttacker: A Configurable Physical Adversarial Attack for Face Recognition via ProjectorPoster
- Project-Probe-Aggregate: Efficient Fine-Tuning for Group RobustnessHighlight
- Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained AnalysisPoster
- Prompt2Perturb (P2P): Text-Guided Diffusion-Based Adversarial Attack on Breast Ultrasound ImagesPoster
- PromptHMR: Promptable Human Mesh RecoveryPoster
- Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context LearningPoster
- Proximal Algorithm Unrolling: Flexible and Efficient Reconstruction Networks for Single-Pixel ImagingPoster
- ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual GroundingPoster
- Pseudo Visible Feature Fine-Grained Fusion for Thermal Object DetectionPoster
- Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose InteractionPoster
- PyTorchGeoNodes: Enabling Differentiable Shape Programs for 3D Shape ReconstructionPoster
- Q-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction RegressionPoster
- QuCOOP: A Versatile Framework for Solving Composite and Binary-Parametrised Problems on Quantum AnnealersHighlight
- Quad-Pixel Image Defocus Deblurring: A New Benchmark and ModelPoster
- Quaffure: Real-Time Quasi-Static Neural Hair SimulationPoster
- QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the EdgePoster
- Query Efficient Black-Box Visual Prompting with Subspace LearningPoster
- Question-Aware Gaussian Experts for Audio-Visual Question AnsweringHighlight
- R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual LocalizationPoster
- R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt TuningPoster
- R2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action PlannerPoster
- RAEncoder: A Label-Free Reversible Adversarial Examples Encoder for Dataset Intellectual Property ProtectionPoster
- RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-EmbeddingsPoster
- RASP: Revisiting 3D Anamorphic Art for Shadow-Guided Packing of Irregular ObjectsPoster
- RC-AutoCalib: An End-to-End Radar-Camera Automatic Calibration NetworkPoster
- RCP-Bench: Benchmarking Robustness for Collaborative Perception Under Diverse CorruptionsPoster
- RDD: Robust Feature Detector and Descriptor using Deformable TransformerPoster
- REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity ConditioningPoster
- RGBAvatar: Reduced Gaussian Blendshapes for Online Modeling of Head AvatarsHighlight
- RICCARDO: Radar Hit Prediction and Convolution for Camera-Radar 3D Object DetectionPoster
- RL-RC-DoT: A Block-level RL agent for Task-Aware Video CompressionPoster
- ROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language ModelsPoster
- ROLL: Robust Noisy Pseudo-label Learning for Multi-View Clustering with Noisy CorrespondenceHighlight
- RORem: Training a Robust Object Remover with Human-in-the-LoopPoster
- RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera FusionPoster
- RaSS: Improving Denoising Diffusion Samplers with Reinforced Active Sampling SchedulerPoster
- Radio Frequency Ray Tracing with Neural Object Representation for Enhanced RF ModelingPoster
- Random Conditioning for Diffusion Model Compression with Distillation
- Random Conditioning with Distillation for Data-Efficient Diffusion Model CompressionPoster
- Rate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty EstimationPoster
- Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion ModelPoster
- ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence LearningPoster
- ReDiffDet: Rotation-equivariant Diffusion Model for Oriented Object DetectionPoster
- RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal ReperformancePoster
- ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data StreamsPoster
- ReWind: Understanding Long Videos with Instructed Learnable MemoryPoster
- Real-IAD D3: A Real-World 2D/Pseudo-3D/3D Dataset for Industrial Anomaly DetectionPoster
- Real-time High-fidelity Gaussian Human Avatars with Position-based Interpolation of Spatially Distributed MLPsHighlight
- Realistic Test-Time Adaptation of Vision-Language ModelsHighlight
- Reanimating Images using Neural Representations of Dynamic StimuliPoster
- ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and ReasoningPoster
- Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance GroundingPoster
- Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems ApproachHighlight
- Recognition-Synergistic Scene Text EditingPoster
- Reconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in DualPoster
- Reconstructing Animals and the WildPoster
- Reconstructing Close Human Interaction with Appearance and Proxemics ReasoningPoster
- Reconstructing Humans with a Biomechanically Accurate SkeletonPoster
- Reconstructing In-the-Wild Open-Vocabulary Human-Object InteractionsPoster
- Reconstructing People, Places, and CamerasHighlight
- Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal TransportPoster
- Rectification-specific Supervision and Constrained Estimator for Online Stereo RectificationPoster
- Recurrent Feature Mining and Keypoint Mixup Padding for Category-Agnostic Pose EstimationPoster
- Reducing Class-wise Confusion for Incremental Learning with Disentangled ManifoldsPoster
- RefPose: Leveraging Reference Geometric Correspondences for Accurate 6D Pose Estimation of Unseen ObjectsPoster
- Relation-Rich Visual Document Generator for Visual Information ExtractionPoster
- Relation3D : Enhancing Relation Modeling for Point Cloud Instance SegmentationPoster
- RelationField: Relate Anything in Radiance FieldsPoster
- Remote Photoplethysmography in Real-World and Extreme Lighting ScenariosPoster
- Reproducible Vision-Language Models Meet Concepts Out of Pre-TrainingPoster
- ResCLIP: Residual Attention for Training-free Dense Vision-language InferencePoster
- Resilient Sensor Fusion Under Adverse Sensor Failures via Multi-Modal Expert FusionPoster
- RestorGS: Depth-aware Gaussian Splatting for Efficient 3D Scene RestorationPoster
- Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher DistillationPoster
- Rethinking Decoder Design: Improving Biomarker Segmentation Using Depth-to-Space Restoration and Residual Linear AttentionPoster
- Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked AutoregressionPoster
- Rethinking Epistemic and Aleatoric Uncertainty for Active Open-Set Annotation: An Energy-Based ApproachPoster
- Rethinking Few-Shot Adaptation of Vision-Language Models in Two StagesPoster
- Rethinking Lanes and Points in Complex Scenarios for Monocular 3D Lane DetectionPoster
- Rethinking Noisy Video-Text Retrieval via Relation-aware AlignmentPoster
- Rethinking Personalized Aesthetics Assessment: Employing Physique Aesthetics Assessment as An ExemplificationHighlight
- Rethinking Query-based Transformer for Continual Image SegmentationPoster
- Rethinking Reconstruction and Denoising in the Dark: New Perspective, General Architecture and BeyondPoster
- Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking TransformersPoster
- Rethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level TasksPoster
- Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense GamePoster
- Reversing Flow for Image RestorationPoster
- Revisiting Audio-Visual Segmentation with Vision-Centric TransformerPoster
- Revisiting Backdoor Attacks against Large Vision-Language Models from Domain ShiftPoster
- Revisiting Fairness in Multitask Learning: A Performance-Driven Approach for Variance ReductionPoster
- Revisiting Generative Replay for Class Incremental Object DetectionPoster
- Revisiting Source-Free Domain Adaptation: Insights into Representativeness, Generalization, and VarietyPoster
- Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate RewardPoster
- RigGS: Rigging of 3D Gaussians for Modeling Articulated Objects in VideosPoster
- RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and SafetyPoster
- RivuletMLP: An MLP-based Architecture for Efficient Compressed Video Quality EnhancementPoster
- RoGSplat: Learning Robust Generalizable Human Gaussian Splatting from Sparse Multi-View ImagesPoster
- RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video NarrativesPoster
- RobSense: A Robust Multi-modal Foundation Model for Remote Sensing with Static, Temporal, and Incomplete Data AdaptabilityPoster
- RoboGround: Robotic Manipulation with Grounded Vision-Language PriorsPoster
- Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent AlignmentPoster
- Robust Message Embedding via Attention Flow-Based SteganographyPoster
- Robust Multi-Object 4D Generation for In-the-wild VideosPoster
- Robust Multimodal Survival Prediction with Conditional Latent Differentiation Variational AutoEncoderPoster
- Robust-MVTON: Learning Cross-Pose Feature Alignment and Fusion for Robust Multi-View Virtual Try-OnPoster
- Rotation-Equivariant Self-Supervised Method in Image DenoisingPoster
- S2D-LFE: Sparse-to-Dense Light Field Event GenerationPoster
- S2Gaussian: Sparse-View Super-Resolution 3D Gaussian SplattingPoster
- S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual RepresentationPoster
- SACB-Net: Spatial-awareness Convolutions for Medical Image RegistrationHighlight
- SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image PretrainingPoster
- SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training CostPoster
- SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything ModelPoster
- SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual ScenesPoster
- SAM2Object: Consolidating View Consistency via SAM2 for Zero-Shot 3D Instance SegmentationPoster
- SAMBLE: Shape-Specific Point Cloud Sampling for an Optimal Trade-Off Between Local Detail and Global UniformityPoster
- SASep: Saliency-Aware Structured Separation of Geometry and Feature for Open Set Learning on Point CloudsPoster
- SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive TokensPoster
- SATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision TransformersPoster
- SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute PromptingPoster
- SCFlow2: Plug-and-Play Object Pose Refiner with Shape-Constraint Scene FlowPoster
- SCSA: A Plug-and-Play Semantic Continuous-Sparse Attention for Arbitrary Semantic Style TransferHighlight
- SDBF: Steep-Decision-Boundary Fingerprinting for Hard-Label Tampering Detection of DNN ModelsPoster
- SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy PredictionPoster
- SEAL: Semantic Attention Learning for Long Video RepresentationPoster
- SEEN-DA: SEmantic ENtropy guided Domain-aware Attention for Domain Adaptive Object DetectionPoster
- SET: Spectral Enhancement for Tiny Object DetectionPoster
- SFDM: Robust Decomposition of Geometry and Reflectance for Realistic Face Rendering from Sparse-view ImagesPoster
- SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI DetectionPoster
- SGCR: Spherical Gaussians for Efficient 3D Curve ReconstructionPoster
- SGFormer: Satellite-Ground Fusion for 3D Semantic Scene CompletionPoster
- SINR: Sparsity Driven Compressed Implicit Neural RepresentationsPoster
- SIR-DIFF: Sparse Image Sets Restoration with Multi-View Diffusion ModelPoster
- SKDream: Controllable Multi-view and 3D Generation with Arbitrary SkeletonsHighlight
- SKE-Layout: Spatial Knowledge Enhanced Layout Generation with LLMsPoster
- SLADE: Shielding against Dual Exploits in Large Vision-Language ModelsPoster
- SLVR: Super-Light Visual Reconstruction via Blueprint Controllable Convolutions and Exploring Feature Diversity RepresentationPoster
- SMILE: Infusing Spatial and Motion Semantics in Masked Video LearningPoster
- SMTPD: A New Benchmark for Temporal Prediction of Social Media PopularityPoster
- SOAP: Vision-Centric 3D Semantic Scene Completion with Scene-Adaptive Decoder and Occluded Region-Aware View ProjectionPoster
- SOGS: Second-Order Anchor for Advanced 3D Gaussian SplattingPoster
- SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous DrivingPoster
- SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic PromptsHighlight
- SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language ModelsPoster
- SPC-GS: Gaussian Splatting with Semantic-Prompt Consistency for Indoor Open-World Free-view Synthesis from Sparse InputsPoster
- SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual TrackingPoster
- SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split OptimizationHighlight
- STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural NetworksPoster
- STAR-Edge: Structure-aware Local Spherical Curve Representation for Thin-walled Edge Extraction from Unstructured Point CloudsPoster
- STDD: Spatio-Temporal Dual Diffusion for Video GenerationPoster
- STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security InspectionHighlight
- STOP: Integrated Spatial-Temporal Dynamic Prompting for Video UnderstandingPoster
- STiL: Semi-supervised Tabular-Image Learning for Comprehensive Task-Relevant Information Exploration in Multimodal ClassificationPoster
- SUM Parts: Benchmarking Part-Level Semantic Segmentation of Urban MeshesPoster
- SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation SparsityHighlight
- SVDC: Consistent Direct Time-of-Flight Video Depth Completion with Frequency Selective FusionPoster
- SVFR: A Unified Framework for Generalized Video Face RestorationPoster
- SVG-IR: Spatially-Varying Gaussian Splatting for Inverse RenderingPoster
- SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video SituationPoster
- S^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion PriorsPoster
- SaMam: Style-aware State Space Model for Arbitrary Image Style TransferHighlight
- Saliuitl: Ensemble Salience Guided Recovery of Adversarial Patches against CNNsPoster
- Samba: A Unified Mamba-based Framework for General Salient Object DetectionHighlight
- Sample- and Parameter-Efficient Auto-Regressive Image ModelsPoster
- Sampling Innovation-Based Adaptive Compressive SensingPoster
- Satellite Observations Guided Diffusion Model for Accurate Meteorological States at Arbitrary ResolutionHighlight
- Scalable Video-to-Dataset Generation for Cross-Platform Mobile AgentsPoster
- Scale Efficient Training for Large DatasetsPoster
- ScaleLSD: Scalable Deep Line Segment Detection StreamlinedPoster
- Scaling Down Text Encoders of Text-to-Image Diffusion ModelsPoster
- Scaling Inference Time Compute for Diffusion ModelsHighlight
- Scaling Vision Pre-Training to 4K ResolutionHighlight
- Scaling up Image Segmentation across Data and TasksPoster
- Scene-Centric Unsupervised Panoptic SegmentationHighlight
- Scene-agnostic Pose Regression for Visual LocalizationPoster
- Scene4U: Hierarchical Layered 3D Scene Reconstruction from Single Panoramic Image for Your Immerse ExplorationPoster
- SceneCrafter: Controllable Multi-View Driving Scene EditingPoster
- SceneDiffuser++: City-Scale Traffic Simulation via a Generative World ModelPoster
- Sea-ing in Low-lightPoster
- SeaLion: Semantic Part-Aware Latent Point Diffusion Models for 3D GenerationPoster
- Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution AnalysisPoster
- See Further When Clear: Curriculum Consistency ModelPoster
- Seeing A 3D World in A Grain of SandPoster
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal DecodingPoster
- Seeing More with Less: Human-like Representations in Vision ModelsHighlight
- Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual ScenesPoster
- Seeing What Matters: Empowering CLIP with Patch Generation-to-SelectionPoster
- Seeing is Not Believing: Adversarial Natural Object Optimization for Hard-Label 3D Scene AttacksPoster
- Seeing the Abstract: Translating the Abstract Language for Vision Language ModelsPoster
- Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment RecognitionPoster
- Seeking Consistent Flat Minima for Better Domain Generalization via Refining Loss LandscapesPoster
- SegMAN: Omni-scale Context Modeling with State Space Models and Local Attention for Semantic SegmentationPoster
- Segment Any Motion in VideosPoster
- Segment Any-Quality Images with Generative Latent Space EnhancementPoster
- Segment Anything, Even OccludedPoster
- Segment This Thing: Foveated Tokenization for Efficient Point-Prompted SegmentationPoster
- Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar SubjectsPoster
- Self-Evolving Visual Concept Library using Vision-Language CriticsPoster
- Self-Learning Hyperspectral and Multispectral Image Fusion via Adaptive Residual Guided Subspace Diffusion ModelPoster
- Self-Supervised Cross-View Correspondence with Predictive Cycle ConsistencyHighlight
- Self-Supervised Large Scale Point Cloud Completion for Archaeological Site RestorationPoster
- Self-Supervised Learning for Color Spike Camera ReconstructionPoster
- Self-Supervised Spatial Correspondence Across ModalitiesPoster
- Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolutionPoster
- SemAlign3D: Semantic Correspondence between RGB-Images through Aligning 3D Object-Class RepresentationsPoster
- SemGeoMo: Dynamic Contextual Human Motion Generation with Semantic and Geometric GuidancePoster
- Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic SegmentationPoster
- Semantic and Expressive Variations in Image Captions Across LanguagesPoster
- Semantic-guided Cross-Modal Prompt Learning for Skeleton-based Zero-shot Action RecognitionPoster
- Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video DerainingPoster
- SemiDAViL: Semi-supervised Domain Adaptation with Vision-Language Guidance for Semantic SegmentationPoster
- SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text SpottingPoster
- Sensitivity-Aware Efficient Fine-Tuning via Compact Dynamic-Rank AdaptationPoster
- Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question AnsweringPoster
- SeqMvRL: A Sequential Fusion Framework for Multi-view Representation LearningPoster
- SeriesBench: A Benchmark for Narrative-Driven Drama Series UnderstandingPoster
- Seurat: From Moving Points to DepthHighlight
- Shadow Generation Using Diffusion Model with Geometry PriorPoster
- Shape Abstraction via Marching Differentiable Support FunctionsHighlight
- Shape and Texture: What Influences Reliable Optical Flow Estimation?Poster
- ShapeShifter: 3D Variations Using Multiscale and Sparse Point-Voxel DiffusionPoster
- ShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware PromptsPoster
- Sharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and ManipulationPoster
- Shift the Lens: Environment-Aware Unsupervised Camouflaged Object DetectionPoster
- ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion ModelsPoster
- Show and Segment: Universal Medical Image Segmentation via In-Context LearningPoster
- Show and Tell: Visually Explainable Deep Neural Nets via Spatially-Aware Concept Bottleneck ModelsHighlight
- ShowMak3r: Compositional TV Show ReconstructionPoster
- Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head GenerationPoster
- SimLTD: Simple Supervised and Semi-Supervised Long-Tailed Object DetectionPoster
- SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity PredictionPoster
- Simpler Diffusion: 1.5 FID on ImageNet512 with Pixel-space DiffusionPoster
- Simulator HC: Regression-based Online Simulation of Starting Problem-Solution Pairs for Homotopy Continuation in Geometric VisionHighlight
- SinGS: Animatable Single-Image Human Gaussian Splats with Kinematic PriorsPoster
- Single Domain Generalization for Few-Shot Counting via Universal Representation MatchingPoster
- Six-CD: Benchmarking Concept Removals for Text-to-image Diffusion ModelsPoster
- Sketch Down the FLOPs: Towards Efficient Networks for Human SketchPoster
- SketchFusion: Learning Universal Sketch Features through Fusing Foundation ModelsPoster
- SketchVideo: Sketch-based Video Generation and EditingPoster
- Sketchtopia: A Dataset and Foundational Agents for Benchmarking Asynchronous Multimodal Communication with Iconic FeedbackPoster
- Sketchy Bounding-box Supervision for 3D Instance SegmentationPoster
- Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters ThemselvesPoster
- SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language ModelingPoster
- SmartCLIP: Modular Vision-language Alignment with Identification GuaranteesHighlight
- SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback OptimizationPoster
- SoMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation LearningHighlight
- SocialGesture: Delving into Multi-person Gesture UnderstandingPoster
- SocialMOIF: Multi-Order Intention Fusion for Pedestrian Trajectory PredictionPoster
- Soft Self-labeling and Potts Relaxations for Weakly-supervised SegmentationPoster
- SoftShadow: Leveraging Soft Masks for Penumbra-Aware Shadow RemovalPoster
- Sound Bridge: Associating Egocentric and Exocentric Videos via Audio CuesPoster
- SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic BindingHighlight
- Sparse Point Cloud Patches Rendering via Splitting 2D GaussiansPoster
- Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field RenderingPoster
- Sparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse ViewsPoster
- SparseAlign: a Fully Sparse Framework for Cooperative Object DetectionPoster
- Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image SynthesisPoster
- SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative LanguagePoster
- SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular InputPoster
- SpecTRe-GS: Modeling Highly Specular Surfaces with Reflected Nearby Objects by Tracing Rays in 3D Gaussian SplattingHighlight
- Spectral State Space Model for Rotation-Invariant Visual Representation LearningPoster
- SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionPoster
- Spherical Manifold Guided Diffusion Model for Panoramic Image GenerationPoster
- Spk2SRImgNet: Super-Resolve Dynamic Scene from Spike Stream via Motion Aligned Collaborative FilteringPoster
- SplatFlow: Self-Supervised Dynamic Gaussian Splatting in Neural Motion Flow Field for Autonomous DrivingHighlight
- Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous DrivingPoster
- Stabilizing and Accelerating Autofocus with Expert Trajectory Regularized Deep Reinforcement LearningPoster
- Stable-SCore: A Stable Registration-based Framework for 3D Shape CorrespondencePoster
- StageDesigner: Artistic Stage Generation for Scenography via Theater ScriptsPoster
- Star with Bilinear MappingPoster
- Stealthy Backdoor Attack in Self-Supervised Learning Vision Encoders for Large Vision Language ModelsPoster
- Steepest Descent Density Control for Compact 3D Gaussian SplattingPoster
- Stochastic Human Motion Prediction with Memory of Action Transition and Action CharacteristicPoster
- Stop Learning it all to Mitigate Visual Hallucination, Focus on the Hallucination Target.Poster
- Stop Walking in Circles! Bailing Out Early in Projected Gradient DescentPoster
- Structure from CollisionHighlight
- Structure-Aware Correspondence Learning for Relative Pose EstimationHighlight
- Structure-from-Motion with a Non-Parametric Camera ModelHighlight
- Style Evolving along Chain-of-Thought for Unknown-Domain Object DetectionHighlight
- Style Quantization for Data-Efficient GAN TrainingPoster
- Style-Editor: Text-driven Object-centric Style EditingHighlight
- StyleSSP: Sampling StartPoint Enhancement for Training-free Diffusion-based Method for Style TransferHighlight
- StyleStudio: Text-Driven Style Transfer with Selective Control of Style ElementsPoster
- Subnet-Aware Dynamic Supernet Training for Neural Architecture SearchPoster
- Subspace Constraint and Contribution Estimation for Heterogeneous Federated LearningPoster
- Supervising Sound Localization by In-the-wild EgomotionHighlight
- Symbolic Representation for Any-to-Any Generative TasksPoster
- Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D GenerationHighlight
- SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled SynthesisPoster
- SyncSDE: A Probabilistic Framework for Diffusion SynchronizationPoster
- SyncVP: Joint Diffusion for Synchronous Multi-Modal Video PredictionPoster
- Synchronized Video-to-Audio Generation via Mel Quantization-Continuum DecompositionPoster
- SynthLight: Portrait Relighting with Diffusion Model by Learning to Re-render Synthetic FacesPoster
- Synthetic Data is an Elegant GIFT for Continual Vision-Language ModelsPoster
- Synthetic Visual GenomePoster
- T-CIL: Temperature Scaling using Adversarial Perturbation for Calibration in Class-Incremental LearningPoster
- T-FAKE: Synthesizing Thermal Images for Facial LandmarkingPoster
- T2ICount: Enhancing Cross-modal Understanding for Zero-Shot CountingHighlight
- TADFormer: Task-Adaptive Dynamic TransFormer for Efficient Multi-Task LearningPoster
- TAET: Two-Stage Adversarial Equalization Training on Long-Tailed DistributionsPoster
- TAGA: Self-supervised Learning for Template-free Animatable Gaussian Articulated ModelPoster
- TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action RecognitionPoster
- TANGO: Training-free Embodied AI Agents for Open-world TasksPoster
- TAROT: Towards Essentially Domain-Invariant Robustness with Theoretical JustificationPoster
- TCFG: Tangential Damping Classifier-free GuidancePoster
- TFCustom: Customized Image Generation with Time-Aware Frequency Feature GuidanceHighlight
- TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time CorrectionHighlight
- TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion GenerationPoster
- TKG-DM: Training-free Chroma Key Content Generation Diffusion ModelHighlight
- TSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual SegmentationPoster
- TSP-Mamba: The Travelling Salesman Problem Meets Mamba for Image Super-resolution and BeyondPoster
- TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage FusionAward Candidate
- TailedCore: Few-Shot Sampling for Unsupervised Long-Tail Noisy Anomaly DetectionPoster
- Take the Bull by the Horns: Learning to Segment Hard SamplesPoster
- Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse InputsHighlight
- TaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian SplattingHighlight
- Targeted Forgetting of Image Subgroups in CLIP ModelsPoster
- Tartan IMU: A Light Foundation Model for Inertial Positioning in RoboticsPoster
- Task-Aware Clustering for Prompting Vision-Language ModelsPoster
- Task-Specific Gradient Adaptation for Few-Shot One-Class ClassificationPoster
- Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual GroundingPoster
- Taste More, Taste Better: Diverse Data and Strong Model Boost Semi-Supervised Crowd CountingPoster
- Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion GenerationPoster
- Temporal Action Detection Model Compression by Progressive Block DropPoster
- Temporal Alignment-Free Video Matching for Few-shot Action RecognitionPoster
- Temporal Score Analysis for Understanding and Correcting Diffusion ArtifactsPoster
- TensoFlow: Tensorial Flow-based Sampler for Inverse RenderingPoster
- Test-Time Domain Generalization via Universe Learning: A Multi-Graph Matching Approach for Medical Image SegmentationPoster
- Test-Time Fine-Tuning of Image Compression Models for Multi-Task AdaptabilityPoster
- Test-Time Visual In-Context TuningPoster
- Test-time Augmentation Improves Efficiency in Conformal PredictionPoster
- TexGarment: Consistent Garment UV Texture Generation via Efficient 3D Structure-Guided Diffusion TransformerPoster
- Text Augmented Correlation Transformer For Few-shot Classification & SegmentationPoster
- Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual AbductionPoster
- Text-guided Sparse Voxel Pruning for Efficient 3D Visual GroundingHighlight
- The Art of Deception: Color Visual Illusions and Diffusion ModelsPoster
- The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video GenerationPoster
- The Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion ModelsPoster
- The Impact Label Noise and Choice of Threshold has on Cross-Entropy and Soft-Dice in Image SegmentationPoster
- The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionAward Candidate
- The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like PhotographersPoster
- Theory-Inspired Deep Multi-View Multi-Label Learning with Incomplete Views and Noisy LabelsPoster
- Thin-Shell-SfT: Fine-Grained Monocular Non-rigid 3D Surface Tracking with Neural Deformation FieldsPoster
- Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot ManipulationPoster
- Three Cars Approaching within 100m! Enhancing Distant Geometry by Tri-Axis Voxel Scanning for Camera-based Semantic Scene CompletionPoster
- Three-view Focal Length Recovery From HomographiesPoster
- Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video GenerationPoster
- Tightening Robustness Verification of MaxPool-based Neural Networks via Minimizing the Over-Approximation ZonePoster
- Tiled DiffusionPoster
- Time of the Flight of the Gaussians: Optimizing Depth Indirectly in Dynamic Radiance FieldsPoster
- TimeTracker: Event-based Continuous Point Tracking for Video Frame Interpolation with Non-linear MotionPoster
- TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video GenerationPoster
- TopNet: Transformer-Efficient Occupancy Prediction Network for Octree-Structured Point Cloud Geometry CompressionPoster
- Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and ReconstructionPoster
- Toward Real-world BEV Perception: Depth Uncertainty Estimation via Gaussian SplattingPoster
- Towards All-in-One Medical Image Re-IdentificationPoster
- Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific ParametersPoster
- Towards Continual Universal SegmentationPoster
- Towards Cost-Effective Learning: A Synergy of Semi-Supervised and Active LearningPoster
- Towards Effective and Sparse Adversarial Attack on Spiking Neural Networks via Breaking Invisible Surrogate GradientsPoster
- Towards Efficient Foundation Model for Zero-shot Amodal SegmentationPoster
- Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color ConsistencyHighlight
- Towards Explainable and Unprecedented Accuracy in Matching Challenging Finger Crease PatternsHighlight
- Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse WeatherPoster
- Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency PartitionPoster
- Towards Generalizable Scene Change DetectionPoster
- Towards High-fidelity 3D Talking Avatar with Personalized Dynamic TexturePoster
- Towards Human-Understandable Multi-Dimensional Concept DiscoveryPoster
- Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long TextHighlight
- Towards In-the-wild 3D Plane Reconstruction from a Single ImageHighlight
- Towards Lossless Implicit Neural Representation via Bit Plane DecompositionPoster
- Towards Natural Language-Based Document Image Retrieval: New Dataset and BenchmarkPoster
- Towards Optimizing Large-Scale Multi-Graph Matching in BioimagingPoster
- Towards Precise Embodied Dialogue Localization via Causality Guided DiffusionPoster
- Towards Satellite Image Road Graph Extraction: A Global-Scale Dataset and A Novel MethodPoster
- Towards Scalable Human-aligned Benchmark for Text-guided Image EditingHighlight
- Towards Smart Point-and-Shoot PhotographyPoster
- Towards Source-Free Machine UnlearningPoster
- Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and AnticipationHighlight
- Towards Understanding and Quantifying Uncertainty for Text-to-Image GenerationPoster
- Towards Universal AI-Generated Image Detection by Variational Information Bottleneck NetworkPoster
- Towards Universal Dataset Distillation via Task-Driven DiffusionPoster
- Track Any Anomalous Object:A Granular Video Anomaly Detection PipelinePoster
- Tracktention: Leveraging Point Tracking to Attend Videos Faster and BetterHighlight
- Training Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training?Poster
- Training-free Dense-Aligned Diffusion Guidance for Modular Conditional Image SynthesisPoster
- Training-free Neural Architecture Search through Variance of Knowledge of Deep Network WeightsPoster
- TransPixeler: Advancing Text-to-Video Generation with TransparencyPoster
- Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving ScenePoster
- TriTex: Learning Texture from a Single Mesh via Triplane Semantic FeaturesPoster
- Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental LearningPoster
- Tuning the Frequencies: Robust Training for Sinusoidal Neural NetworksHighlight
- TurboFill: Adapting Few-step Text-to-image Model for Fast Image InpaintingPoster
- Twinner: Shining Light on Digital Twins in a Few SnapsPoster
- Two is Better than One: Efficient Ensemble Defense for Robust and Compact ModelsPoster
- U-Know-DiffPAN: An Uncertainty-aware Knowledge Distillation Diffusion Framework with Details Enhancement for PAN-SharpeningPoster
- UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial ReferencesPoster
- UCM-VeID V2: A Richer Dataset and A Pre-training Method for UAV Cross-Modality Vehicle Re-IdentificationPoster
- UCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label LearningHighlight
- UHD-processer: Unified UHD Image Restoration with Progressive Frequency Learning and Degradation-aware PromptsPoster
- UIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion ModelsHighlight
- UMFN: Unified Multi-Domain Face Normalization for Joint Cross-domain Prototype Learning and Heterogeneous Face RecognitionPoster
- UMotion: Uncertainty-driven Human Motion Estimation from Inertial and Ultra-wideband UnitsHighlight
- UNEM: UNrolled Generalized EM for Transductive Few-Shot LearningPoster
- UNIALIGN: Scaling Multimodal Alignment within One Unified ModelPoster
- UNICL-SAM: Uncertainty-Driven In-Context Segmentation with Part Prototype DiscoveryPoster
- UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model EvaluationPoster
- URWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image RestorationPoster
- UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video ParsingPoster
- UltraFusion: Ultra High Dynamic Imaging using Exposure FusionHighlight
- Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingPoster
- Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video BenchmarksPoster
- Unboxed: Geometrically and Temporally Consistent Video OutpaintingPoster
- Uncertainty Meets Diversity: A Comprehensive Active Learning Framework for Indoor 3D Object DetectionPoster
- Uncertainty Weighted Gradients for Model CalibrationPoster
- Uncertainty-Instructed Structure Injection for Generalizable HD Map ConstructionPoster
- Uncertainty-guided Perturbation for Image Super-Resolution Diffusion ModelPoster
- Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic SpacePoster
- Understanding Multi-Task Activities from Single-Task VideosHighlight
- Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream DiffusionPoster
- Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single VideoHighlight
- UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose EstimationPoster
- UniK3D: Universal Camera Monocular 3D EstimationPoster
- UniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object DetectionPoster
- UniNet: A Contrastive Learning-guided Unified Framework with Feature Selection for Anomaly DetectionPoster
- UniPhy: Learning a Unified Constitutive Model for Inverse Physics SimulationPoster
- UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian SplattingPoster
- UniSTD: Towards Unified Spatio-Temporal Learning across Diverse DisciplinesPoster
- Unified Dense Prediction of Video DiffusionPoster
- Unified Medical Lesion Segmentation via Self-referring IndicatorPoster
- Unified Reconstruction of Static and Dynamic Scenes from EventsHighlight
- Unity in Diversity: Video Editing via Gradient-Latent PurificationPoster
- Universal Domain Adaptation for Semantic SegmentationPoster
- Universal Scene Graph GenerationHighlight
- Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media ManipulationPoster
- Unlocking Generalization Power in LiDAR Point Cloud RegistrationHighlight
- Unlocking the Potential of Unlabeled Data in Semi-Supervised Domain GeneralizationPoster
- Unraveling Normal Anatomy via Fluid-Driven Anomaly RandomizationPoster
- Unsupervised Continual Domain Shift Learning with Multi-Prototype ModelingHighlight
- Unsupervised Discovery of Facial Landmarks and Head PosePoster
- Unveiling Differences in Generative Models: A Scalable Differential Clustering ApproachHighlight
- UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene SimulationPoster
- Using Diffusion Priors for Video Amodal SegmentationPoster
- Using Powerful Prior Knowledge of Diffusion Model in Deep Unfolding Networks for Image Compressive SensingPoster
- V-Stylist: Video Stylization via Collaboration and Reflection of MLLM AgentsPoster
- V2V3D: View-to-View Denoised 3D Reconstruction for Light Field MicroscopyPoster
- V2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object DetectionPoster
- VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict EntailmentPoster
- VEU-Bench: Towards Comprehensive Understanding of Video EditingHighlight
- VIRES: Video Instance Repainting via Sketch and Text Guided GenerationPoster
- VISTREAM: Improving Computation Efficiency of Visual Streaming Perception via Law-of-Charge-Conservation Inspired Spiking Neural NetworkPoster
- VI^3NR: Variance Informed Initialization for Implicit Neural RepresentationsPoster
- VL2Lite: Task-Specific Knowledge Distillation from Large Vision-Language Models to Lightweight NetworksPoster
- VLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement LearningPoster
- VLog: Video-Language Models by Generative Retrieval of Narration VocabularyPoster
- VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language ModelsPoster
- VODiff: Controlling Object Visibility Order in Text-to-Image GenerationPoster
- VSNet: Focusing on the Linguistic Characteristics of Sign LanguagePoster
- V^2Dial: Unification of Video and Visual Dialog via Multimodal ExpertsPoster
- Variance-Based Membership Inference Attacks Against Large-Scale Image Captioning ModelsPoster
- VasTSD: Learning 3D Vascular Tree-state Space Diffusion Model for Angiography SynthesisPoster
- VerbDiff: Text-Only Diffusion Models with Enhanced Interaction AwarenessPoster
- ViKIENet: Towards Efficient 3D Object Detection with Virtual Key Instance Enhanced NetworkPoster
- ViUniT: Visual Unit Tests for More Robust Visual ProgrammingPoster
- Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal PriorPoster
- Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free SimulationPoster
- VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic ManipulationPoster
- VidSeg: Training-free Video Semantic Segmentation based on Diffusion ModelsPoster
- Video Language Model Pretraining with Spatio-temporal MaskingPoster
- Video Summarization with Large Language ModelsPoster
- Video-Bench: Human-Aligned Video Generation BenchmarkPoster
- Video-ColBERT: Contextualized Late Interaction for Text-to-Video RetrievalPoster
- Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language ModelsPoster
- VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text ModelsPoster
- VideoDirector: Precise Video Editing via Text-to-Video ModelsPoster
- VideoGEM: Training-free Action Grounding in VideosPoster
- VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion ModelsPoster
- VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and EditingPoster
- Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation LearningPoster
- ViiNeuS: Volumetric Initialization for Implicit Neural Surface Reconstruction of Urban Scenes with Limited Image OverlapPoster
- Vision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D ScenesPoster
- Vision-Language Embodiment for Monocular Depth EstimationPoster
- Vision-Language Gradient Descent-driven All-in-One Deep Unfolding NetworksPoster
- Vision-Language Model IP Protection via Prompt-based LearningPoster
- Visual Consensus Prompting for Co-Salient Object DetectionPoster
- Visual Persona: Foundation Model for Full-Body Human CustomizationPoster
- Visual Representation Learning through Causal Intervention for Controllable Image EditingHighlight
- Visual and Semantic Prompt Collaboration for Generalized Zero-Shot LearningPoster
- Visual-Instructed Degradation Diffusion for All-in-One Image RestorationPoster
- VladVA: Discriminative Fine-tuning of LVLMsPoster
- VolFormer: Explore More Comprehensive Cube Interaction for Hyperspectral Image Restoration and BeyondPoster
- Volume Tells: Dual Cycle-Consistent Diffusion for 3D Fluorescence Microscopy De-noising and Super-ResolutionHighlight
- Volumetric Surfaces: Representing Fuzzy Geometries with Layered MeshesPoster
- Volumetrically Consistent 3D Gaussian RasterizationHighlight
- WISE: A Framework for Gigapixel Whole-Slide-Image Lossless CompressionPoster
- WISH: Weakly Supervised Instance Segmentation using Heterogeneous LabelsHighlight
- WISNet: Pseudo Label Generation on Unbalanced and Patch Annotated Waste ImagesPoster
- Watermarking One for All: A Robust Watermarking Scheme Against Partial Image TheftPoster
- Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial AnimationPoster
- Wavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect DetectionPoster
- WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and SegmentationPoster
- Weakly Supervised Contrastive Adversarial Training for Learning Robust Features from Semi-supervised DataPoster
- Weakly Supervised Semantic Segmentation via Progressive Confidence Region ExpansionPoster
- Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language ModelsPoster
- WeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba DiffusionPoster
- When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation LearningPoster
- Where the Devil Hides: Deepfake Detectors Can No Longer Be TrustedPoster
- Where's the Liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated ContentPoster
- Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional VideosHighlight
- Your Scale Factors are My Weapon: Targeted Bit-Flip Attacks on Vision Transformers via Scale Factor ManipulationPoster
- Z-Magic: Zero-shot Multiple Attributes Guided Image CreatorPoster
- Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video DiffusionPoster
- Zero-Shot 4D Lidar Panoptic SegmentationPoster
- Zero-Shot Blind-spot Image Denoising via Implicit Neural SamplingPoster
- Zero-Shot Head Swapping in Real-World ScenariosPoster
- Zero-Shot Image Restoration Using Few-Step Guidance of Consistency Models (and Beyond)Poster
- Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric DiffusionPoster
- Zero-Shot Styled Text Image Generation, but Make It AutoregressivePoster
- Zero-shot 3D Question Answering via Voxel-based Dynamic Token CompressionPoster
- Zero-shot RGB-D Point Cloud Registration with Pre-trained Large Vision ModelPoster
- ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic GraspingPoster
- ZeroVO: Visual Odometry with Minimal AssumptionsPoster
- beta-FFT: Nonlinear Interpolation and Differentiated Training Strategies for Semi-Supervised Medical Image SegmentationPoster
- dFLMoE: Decentralized Federated Learning via Mixture of Experts for Medical Data AnalysisPoster
- iG-6DoF: Model-free 6DoF Pose Estimation for Unseen Object via Iterative 3D Gaussian SplattingPoster
- iSegMan: Interactive Segment-and-Manipulate 3D GaussiansPoster
- nnWNet: Rethinking the Use of Transformers in Biomedical Image Segmentation and Calling for a Unified Evaluation BenchmarkPoster
- pFedMxF: Personalized Federated Class-Incremental Learning with Mixture of Frequency AggregationPoster
- v-CLR: View-Consistent Learning for Open-World Instance SegmentationHighlight
- vesselFM: A Foundation Model for Universal 3D Blood Vessel SegmentationPoster
CVPR accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.