CVPR 2025 Accepted Papers
The full list of 2,872 papers accepted at CVPR 2025 (IEEE/CVF Conference on Computer Vision and Pattern Recognition). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
Poster: 2,468Highlight: 388Award Candidate: 15
- MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual RepresentationsPoster1 citations
- MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized DataPoster1 citations
- MeshArt: Generating Articulated Meshes with Structure-Guided TransformersPoster1 citations
- MixerMDM: Learnable Composition of Human Motion Diffusion ModelsPoster1 citations
- MoEE: Mixture of Emotion Experts for Audio-Driven Portrait AnimationPoster1 citations
- MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual EncodersPoster1 citations
- Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identificationHighlight1 citations
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsAward Candidate1 citations
- Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D SegmentationPoster1 citations
- Motion Modes: What Could Happen Next?Poster1 citations
- Multi-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain AdaptationPoster1 citations
- Multi-view Reconstruction via SfM-guided Monocular Depth EstimationPoster1 citations
- MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video RetrievalPoster1 citations
- Multitwine: Multi-Object Compositing with Text and Layout ControlHighlight1 citations
- NTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on ThermodynamicsPoster1 citations
- NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed ImagesPoster1 citations
- Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level CaptionsPoster1 citations
- Nested Diffusion Models Using Hierarchical Latent PriorsPoster1 citations
- NoPain: No-box Point Cloud Attack via Optimal Transport Singular BoundaryPoster1 citations
- NoT: Federated Unlearning via Weight NegationPoster1 citations
- Novel View Synthesis with Pixel-Space Diffusion ModelsPoster1 citations
- ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language ModelsPoster1 citations
- OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIPPoster1 citations
- ObjectMover: Generative Object Movement with Video PriorPoster1 citations
- Olympus: A Universal Task Router for Computer Vision TasksHighlight1 citations
- On the Generalization of Handwritten Text Recognition ModelsPoster1 citations
- Open-World Objectness Modeling Unifies Novel Object DetectionPoster1 citations
- Optical-Flow Guided Prompt Optimization for Coherent Video GenerationPoster1 citations
- PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language ModelsPoster1 citations
- PCDreamer: Point Cloud Completion Through Multi-view Diffusion PriorsPoster1 citations
- PEACE: Empowering Geologic Map Holistic Understanding with MLLMsPoster1 citations
- PERSE: Personalized 3D Generative Avatars from A Single PortraitPoster1 citations
- PICO: Reconstructing 3D People In Contact with ObjectsPoster1 citations
- PO3AD: Predicting Point Offsets toward Better 3D Point Cloud Anomaly DetectionPoster1 citations
- PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance PredictionPoster1 citations
- Panorama Generation From NFoV Image Done RightHighlight1 citations
- Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local PerceptionPoster1 citations
- Pay Attention to the Foreground in Object-Centric LearningPoster1 citations
- PersonaBooth: Personalized Text-to-Motion GenerationPoster1 citations
- Personalized Preference Fine-tuning of Diffusion ModelsPoster1 citations
- PhysicsGen: Can Generative Models Learn from Images to Predict Complex Physical Relations?Poster1 citations
- Pippo: High-Resolution Multi-View Humans from a Single ImageHighlight1 citations
- Point Clouds Meets Physics: Dynamic Acoustic Field Fitting Network for Point Cloud UnderstandingPoster1 citations
- Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among InstancesPoster1 citations
- Preconditioners for the Stochastic Training of Neural FieldsPoster1 citations
- Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge ModelsPoster1 citations
- ProAPO: Progressively Automatic Prompt Optimization for Visual ClassificationPoster1 citations
- ProReflow: Progressive Reflow with Decomposed VelocityPoster1 citations
- Probability Density Geodesics in Image Diffusion Latent SpacePoster1 citations
- PromptHash:Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing RetrievalPoster1 citations
- Protecting Your Video Content: Disrupting Automated Video-based LLM AnnotationsPoster1 citations
- ProtoDepth: Unsupervised Continual Depth Completion with PrototypesPoster1 citations
- Q-Bench-Video: Benchmark the Video Quality Understanding of LMMsPoster1 citations
- Quantization without TearsPoster1 citations
- RAD: Region-Aware Diffusion Models for Image InpaintingPoster1 citations
- RAP: Retrieval-Augmented Personalization for Multimodal Large Language ModelsPoster1 citations
- RNG: Relightable Neural GaussiansPoster1 citations
- RUBIK: A Structured Benchmark for Image Matching across Geometric ChallengesPoster1 citations
- RainyGS: Efficient Rain Synthesis with Physically-Based Gaussian SplattingPoster1 citations
- Rashomon Sets for Prototypical-Part Networks: Editing Interpretable Models in Real-TimePoster1 citations
- RayFlow: Instance-Aware Diffusion Acceleration via Adaptive Flow TrajectoriesPoster1 citations
- ReCap: Better Gaussian Relighting with Cross-Environment CapturesPoster1 citations
- ReNeg: Learning Negative Embedding with Reward GuidanceHighlight1 citations
- ReRAW: RGB-to-RAW Image Reconstruction via Stratified Sampling for Efficient Object Detection on the EdgePoster1 citations
- ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long VideosPoster1 citations
- Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image RetrievalHighlight1 citations
- Recovering Dynamic 3D Sketches from VideosPoster1 citations
- Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document RetrievalPoster1 citations
- Repurposing Stable Diffusion Attention for Training-Free Unsupervised Interactive SegmentationPoster1 citations
- Rethinking Correspondence-based Category-Level Object Pose EstimationPoster1 citations
- Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian SplattingPoster1 citations
- Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy PredictionPoster1 citations
- Retrieving Semantics from the Deep: an RAG Solution for Gesture SynthesisPoster1 citations
- Reversible Decoupling Network for Single Image Reflection RemovalPoster1 citations
- RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-TrainingHighlight1 citations
- RoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured EnvironmentsPoster1 citations
- Robust 3D Shape Reconstruction in Zero-Shot from a Single Image in the WildPoster1 citations
- SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in StructuresPoster1 citations
- SEC-Prompt:SEmantic Complementary Prompting for Few-Shot Class-Incremental LearningPoster1 citations
- SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained UnderstandingPoster1 citations
- SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image GenerationPoster1 citations
- SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D ReconstructionPoster1 citations
- STEPS: Sequential Probability Tensor Estimation for Text-to-Image Hard Prompt SearchPoster1 citations
- STINR: Deciphering Spatial Transcriptomics via Implicit Neural RepresentationPoster1 citations
- STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal GroundingPoster1 citations
- Scalable Autoregressive Monocular Depth EstimationPoster1 citations
- Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion ModelPoster1 citations
- SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World EnvironmentsPoster1 citations
- Science-T2I: Addressing Scientific Illusions in Image SynthesisPoster1 citations
- SeCap: Self-Calibrating and Adaptive Prompts for Cross-view Person Re-Identification in Aerial-Ground NetworksHighlight1 citations
- Search and Detect: Training-Free Long Tail Object Detection via Web-Image RetrievalPoster1 citations
- SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator TrajectoriesPoster1 citations
- Semantic and Sequential Alignment for Referring Video Object SegmentationPoster1 citations
- SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion ModelsPoster1 citations
- Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal GroundingPoster1 citations
- SerialGen: Personalized Image Generation by First Standardization Then PersonalizationPoster1 citations
- SfM-Free 3D Gaussian Splatting via Hierarchical TrainingPoster1 citations
- Shading Meets Motion: Self-supervised Indoor 3D Reconstruction Via Simultaneous Shape-from-Shading and Structure-from-MotionPoster1 citations
- Shape My Moves: Text-Driven Shape-Aware Synthesis of Human MotionsPoster1 citations
- ShiftwiseConv: Small Convolutional Kernel with Large Kernel EffectPoster1 citations
- Shining Yourself: High-Fidelity Ornaments Virtual Try-on with Diffusion ModelPoster1 citations
- ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual InstructionsPoster1 citations
- Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion ModelsPoster1 citations
- SimAvatar: Simulation-Ready Avatars with Layered Hair and ClothingPoster1 citations
- SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action AlignmentHighlight1 citations
- SimVS: Simulating World Inconsistencies for Robust View SynthesisPoster1 citations
- Similarity-Guided Layer-Adaptive Vision Transformer for UAV TrackingPoster1 citations
- Solving Instance Detection from an Open-World PerspectivePoster1 citations
- SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal ModelsHighlight1 citations
- Spatiotemporal Decoupling for Efficient Vision-Based Occupancy ForecastingPoster1 citations
- Spectral Informed Mamba for Robust Point Cloud ProcessingPoster1 citations
- Spiking Transformer: Introducing Accurate Addition-Only Spiking Self-Attention for TransformerPoster1 citations
- Split Adaptation for Pre-trained Vision TransformersPoster1 citations
- StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene GenerationPoster1 citations
- StdGEN: Semantic-Decomposed 3D Character Generation from Single ImagesPoster1 citations
- Steady Progress Beats Stagnation: Mutual Aid of Foundation and Conventional Models in Mixed Domain Semi-Supervised Medical Image SegmentationPoster1 citations
- StickMotion: Generating 3D Human Motions by Drawing a StickmanPoster1 citations
- Stretching Each Dollar: Diffusion Training from Scratch on a Micro-BudgetPoster1 citations
- SuperLightNet: Lightweight Parameter Aggregation Network for Multimodal Brain Tumor SegmentationPoster1 citations
- SuperPC: A Single Diffusion Model for Point Cloud Completion, Upsampling, Denoising, and ColorizationPoster1 citations
- SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step DiffusionPoster1 citations
- Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video GenerationPoster1 citations
- Synthetic Prior for Few-Shot Drivable Head Avatar InversionPoster1 citations
- Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure PriorsPoster1 citations
- Temporally Consistent Object-Centric Learning by Contrasting SlotsPoster1 citations
- Test-Time Backdoor Detection for Object Detection ModelsPoster1 citations
- The Change You Want To Detect: Semantic Change Detection In Earth Observation With Hybrid Data GenerationfPoster1 citations
- The Devil is in Low-Level Features for Cross-Domain Few-Shot SegmentationPoster1 citations
- Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference SystemsHighlight1 citations
- Token Cropr: Faster ViTs for Quite a Few TasksPoster1 citations
- Tokenize Image Patches: Global Context Fusion for Effective Haze Removal in Large ImagesPoster1 citations
- Toward Generalized Image Quality Assessment: Relaxing the Perfect Reference Quality AssumptionPoster1 citations
- Toward Robust Neural Reconstruction from Sparse Point SetsPoster1 citations
- Towards Autonomous Micromobility through Scalable Urban SimulationHighlight1 citations
- Towards Generalizable Trajectory Prediction using Dual-Level Representation Learning and Adaptive PromptingPoster1 citations
- Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual AttacksPoster1 citations
- Towards Realistic Example-based Modeling via 3D Gaussian StitchingPoster1 citations
- Towards Training-free Anomaly Detection with Vision and Language Foundation ModelsPoster1 citations
- Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly DetectionPoster1 citations
- Trajectory Mamba: Efficient Attention-Mamba Forecasting Model Based on Selective SSMPoster1 citations
- Traversing Distortion-Perception Tradeoff using a Single Score-Based Generative ModelPoster1 citations
- TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree SequencingPoster1 citations
- Turbo3D: Ultra-fast Text-to-3D GenerationPoster1 citations
- Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot ManipulationPoster1 citations
- Type-R: Automatically Retouching Typos for Text-to-Image GenerationHighlight1 citations
- UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image GenerationPoster1 citations
- USP-Gaussian: Unifying Spike-based Image Reconstruction, Pose Correction and Gaussian SplattingHighlight1 citations
- Understanding Multi-layered Transmission MatricesHighlight1 citations
- UniRestore: Unified Perceptual and Task-Oriented Image Restoration Model Using Diffusion PriorHighlight1 citations
- Unified Uncertainty-Aware Diffusion for Multi-Agent Trajectory ModelingPoster1 citations
- Unlearning through Knowledge Overwriting: Reversible Federated Unlearning via Selective Sparse AdapterPoster1 citations
- Unsupervised Foundation Model-Agnostic Slide-Level Representation LearningPoster1 citations
- Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering IncorrectlyPoster1 citations
- VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token SparsificationPoster1 citations
- VITED: Video Temporal Evidence DistillationPoster1 citations
- VTON 360: High-Fidelity Virtual Try-On from Any Viewing DirectionPoster1 citations
- VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's GuidePoster1 citations
- VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative PriorsPoster1 citations
- VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video UnderstandingPoster1 citations
- VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One StepHighlight1 citations
- VinaBench: Benchmark for Faithful and Consistent Visual NarrativesPoster1 citations
- Visual Lexicon: Rich Image Features in Language SpacePoster1 citations
- Visual Prompting for One-shot Controllable Video Editing without InversionPoster1 citations
- VoteFlow: Enforcing Local Rigidity in Self-Supervised Scene FlowPoster1 citations
- VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow PredictionPoster1 citations
- What Makes a Good Dataset for Knowledge Distillation?Poster1 citations
- When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven ApproachPoster1 citations
- Words or Vision: Do Vision-Language Models Have Blind Faith in Text?Poster1 citations
- Yo'Chameleon: Personalized Vision and Language GenerationPoster1 citations
- Zero-Shot Monocular Scene Flow Estimation in the WildAward Candidate1 citations
- g3D-LF: Generalizable 3D-Language Feature Fields for Embodied TasksPoster1 citations
- 3D Dental Model Segmentation with Geometrical Boundary PreservingPoster
- 3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial RepresentationsPoster
- 3D Gaussian Inpainting with Depth-Guided Cross-View ConsistencyPoster
- 3D Prior Is All You Need: Cross-Task Few-shot 2D Gaze EstimationPoster
- 3D Student Splatting and ScoopingAward Candidate
- 3D-AVS: LiDAR-based 3D Auto-Vocabulary SegmentationPoster
- 3D-MVP: 3D Multiview Pretraining for ManipulationPoster
- 3D-SLNR: A Super Lightweight Neural Representation for Large-scale 3D MappingPoster
- 3DEnhancer: Consistent Multi-View Diffusion for 3D EnhancementPoster
- 4D-Fly: Fast 4D Reconstruction from a Single Monocular VideoPoster
- 4DGC: Rate-Aware 4D Gaussian Compression for Efficient Streamable Free-Viewpoint VideoPoster
- 4DTAM: Non-Rigid Tracking and Mapping via Dynamic Surface GaussiansPoster
- 4Deform: Neural Surface Deformation for Robust Shape InterpolationPoster
- A Comprehensive Study of Decoder-Only LLMs for Text-to-Image GenerationPoster
- A Focused Human Body Model for Accurate Anthropometric Measurements ExtractionPoster
- A General Adaptive Dual-level Weighting Mechanism for Remote Sensing PansharpeningPoster
- A Hubness Perspective on Representation Learning for Graph-Based Multi-View ClusteringPoster
- A New Statistical Model of Star Speckles for Learning to Detect and Characterize Exoplanets in Direct Imaging ObservationsPoster
- A Physics-Informed Blur Learning Framework for Imaging SystemsPoster
- A Polarization-Aided Transformer for Image Deblurring via Motion Vector DecompositionHighlight
- A Regularization-Guided Equivariant Approach for Image RestorationPoster
- A Selective Re-learning Mechanism for Hyperspectral Fusion ImagingPoster
- A Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image SegmentationPoster
- A Simple yet Effective Layout Token in Large Language Models for Document UnderstandingPoster
- A Theory of Learning Unified Model via Knowledge Integration from Label Space Varying DomainsPoster
- A Unified Approach to Interpreting Self-supervised Pre-training Methods for 3D Point Clouds via InteractionsHighlight
- A Unified Framework for Heterogeneous Semi-supervised LearningPoster
- A Unified Image-Dense Annotation Generation Model for Underwater ScenesPoster
- A Unified Latent Schrodinger Bridge Diffusion Model for Unsupervised Anomaly Detection and LocalizationPoster
- A Unified, Resilient, and Explainable Adversarial Patch DetectorPoster
- A Universal Scale-Adaptive Deformable Transformer for Image Restoration across Diverse ArtifactsPoster
- A3: Few-shot Prompt Learning of Unlearnable Examples with Cross-Modal Adversarial Feature AlignmentPoster
- A4A: Adapter for Adapter Transfer via All-for-All Mapping for Cross-Architecture ModelsPoster
- ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image ObjectsPoster
- ABC-Former: Auxiliary Bimodal Cross-domain Transformer with Interactive Channel Attention for White BalancePoster
- ACAttack: Adaptive Cross Attacking RGB-T Tracker via Multi-Modal Response DecouplingPoster
- ACL: Activating Capability of Linear Attention for Image RestorationPoster
- ADD: Attribution-Driven Data Augmentation Framework for Boosting Image Super-ResolutionPoster
- ADU: Adaptive Detection of Unknown Categories in Black-Box Domain AdaptationPoster
- AIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic DataPoster
- AIpparel: A Multimodal Foundation Model for Digital GarmentsHighlight
- ALIEN: Implicit Neural Representations for Human Motion Prediction under Arbitrary LatencyHighlight
- AMO Sampler: Enhancing Text Rendering with OvershootingPoster
- AMR-Transformer: Enabling Efficient Long-range Interaction for Complex Neural Fluid SimulationPoster
- ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric InteractionPoster
- APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision TransformersPoster
- APT: Adaptive Personalized Training for Diffusion Models with Limited DataPoster
- ASHiTA: Automatic Scene-grounded HIerarchical Task AnalysisPoster
- ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background InpaintingPoster
- ATP: Adaptive Threshold Pruning for Efficient Data Encoding in Quantum Neural NetworksPoster
- AVF-MAE++: Scaling Affective Video Facial Masked Autoencoders via Efficient Audio-Visual Self-Supervised LearningPoster
- AVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual LearningPoster
- Acc3D: Accelerating Single Image to 3D Diffusion Models via Edge Consistency Guided Score DistillationPoster
- Accelerating Diffusion Transformer via Increment-Calibrated Caching with Channel-Aware Singular Value DecompositionPoster
- Accurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-DistillationPoster
- Acquire and then Adapt: Squeezing out Text-to-Image Model for Image RestorationPoster
- Action Detail Matters: Refining Video Recognition with Local Action QueriesPoster
- Activating Sparse Part Concepts for 3D Class Incremental LearningPoster
- Active Event-based Stereo VisionPoster
- Active Hyperspectral Imaging Using an Event CameraHighlight
- AdMiT: Adaptive Multi-Source Tuning in Dynamic EnvironmentsPoster
- AdaDARE-gamma: Balancing Stability and Plasticity in Multi-modal LLMs through Efficient AdaptationPoster
- AdaptCMVC: Robust Adaption to Incremental Views in Continual Multi-view ClusteringPoster
- Adapting Dense Matching for Homography Estimation with Grid-based AccelerationPoster
- Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal PerceptionPoster
- Adapting Text-to-Image Generation with Feature Difference Instruction for Generic Image RestorationPoster
- Adapting to Observation Length of Trajectory Prediction via Contrastive LearningPoster
- Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic ThresholdsPoster
- Adaptive Dropout: Unleashing Dropout across Layers for Generalizable Image Super-ResolutionPoster
- Adaptive Keyframe Sampling for Long Video UnderstandingPoster
- Adaptive Markup Language Generation for Contextually-Grounded Visual Document UnderstandingPoster
- Adaptive Non-Uniform Timestep Sampling for Accelerating Diffusion Model TrainingPoster
- Adaptive Parameter Selection for Tuning Vision-Language ModelsPoster
- Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play EnhancementPoster
- Adaptive Rectangular Convolution for Remote Sensing PansharpeningPoster
- Adaptive Unimodal Regulation for Balanced Multimodal Information AcquisitionPoster
- Advancing Adversarial Robustness in GNeRFs: The IL2-NeRF AttackPoster
- Advancing Generalizable Tumor Segmentation with Anomaly-Aware Open-Vocabulary Attention Maps and Frozen Foundation Diffusion ModelsPoster
- Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 DatasetPoster
- Advancing Multiple Instance Learning with Continual Learning for Whole Slide ImagingHighlight
- Adventurer: Optimizing Vision Mamba Architecture Designs for EfficiencyPoster
- Adversarial Domain Prompt Tuning and Generation for Single Domain GeneralizationPoster
- AeSPa : Attention-guided Self-supervised Parallel Imaging for MRI ReconstructionPoster
- AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language ModelsPoster
- AirRoom: Objects Matter in Room ReidentificationPoster
- Align-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video EditingPoster
- Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Large Model EnhancementPoster
- Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question AnsweringPoster
- All-Day Multi-Camera Multi-Target TrackingPoster
- All-Optical Nonlinear Diffractive Deep Network for Ultrafast Image DenoisingHighlight
- All-directional Disparity Estimation for Real-world QPD ImagesHighlight
- AlphaPre: Amplitude-Phase Disentanglement Model for Precipitation NowcastingPoster
- An Image-like Diffusion Method for Human-Object Interaction DetectionPoster
- Analyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose EstimationPoster
- Anatomical Consistency and Adaptive Prior-informed Transformation for Multi-contrast MR Image Synthesis via Diffusion ModelPoster
- Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2DPoster
- AniGrad: Anisotropic Gradient-Adaptive Sampling for 3D Reconstruction From Monocular VideoPoster
- AniMer: Animal Pose and Shape Estimation Using Family Aware TransformerPoster
- AniMo: Species-Aware Model for Text-Driven Animal Motion GenerationPoster
- Animate and Sound an ImagePoster
- Annotation Ambiguity Aware Semi-Supervised Medical Image SegmentationHighlight
- Anomize: Better Open Vocabulary Video Anomaly DetectionPoster
- Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object PerceptionPoster
- Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask TrackingPoster
- Any6D: Model-free 6D Pose Estimation of Novel ObjectsPoster
- AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual VideosPoster
- AnyMap: Learning a General Camera Model for Structure-from-Motion with Unknown Distortion in Dynamic ScenesPoster
- AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion ModelsPoster
- Anyattack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language ModelsPoster
- Apply Hierarchical-Chain-of-Generation to Complex Attributes Text-to-3D GenerationPoster
- ArcPro: Architectural Programs for Structured 3D Abstraction of Sparse PointsHighlight
- Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?Poster
- Argus: A Compact and Versatile Foundation Model for VisionPoster
- Argus: Vision-Centric Reasoning with Grounded Chain-of-ThoughtPoster
- ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image IntermediaryPoster
- Articulated Kinematics Distillation from Video Diffusion ModelsPoster
- Assessing and Learning Alignment of Unimodal Vision and Language ModelsHighlight
- Associative TransformerPoster
- Asynchronous Collaborative Graph Representation for Frames and EventsPoster
- Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT AccelerationPoster
- Attention IoU: Examining Biases in CelebA using Attention MapsPoster
- Attraction Diminishing and Distributing for Few-Shot Class-Incremental LearningPoster
- Attribute-Missing Multi-view Graph ClusteringPoster
- Attribute-formed Class-specific Concept Space: Endowing Language Bottleneck Model with Better Interpretability and ScalabilityPoster
- AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion TransformersPoster
- Audio-Visual Semantic Graph Network for Audio-Visual Event LocalizationPoster
- Augmented Deep Contexts for Spatially Embedded Video CodingHighlight
- Augmenting Perceptual Super-Resolution via Image Quality PredictorsPoster
- AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360deg Unbounded Scene InpaintingPoster
- Auto-Encoded Supervision for Perceptual Image Super-ResolutionPoster
- AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video HashingPoster
- Automated Proof of Polynomial Inequalities via Reinforcement LearningPoster
- Automatic Spectral Calibration of Hyperspectral Images: Method, Dataset and BenchmarkPoster
- Autoregressive Distillation of Diffusion TransformersPoster
- Autoregressive Sequential Pretraining for Visual TrackingPoster
- AvatarArtist: Open-Domain 4D AvatarizationPoster
- BACON: Improving Clarity of Image Captions via Bag-of-Concept GraphsPoster
- BADGR: Bundle Adjustment Diffusion Conditioned by Gradients for Wide-Baseline Floor Plan ReconstructionHighlight
- BASKET: A Large-Scale Video Dataset for Fine-Grained Skill EstimationPoster
- BF-STVSR: B-Splines and Fourier---Best Friends for High Fidelity Spatial-Temporal Video Super-ResolutionPoster
- BHViT: Binarized Hybrid Vision TransformerPoster
- BIGS: Bimanual Category-agnostic Interaction Reconstruction from Monocular Videos via 3D Gaussian SplattingPoster
- BLADE: Single-view Body Mesh Estimation through Accurate Depth EstimationPoster
- BOE-ViT: Boosting Orientation Estimation with Equivariance in Self-Supervised 3D Subtomogram AlignmentPoster
- BOLT: Boost Large Vision-Language Model Without Training for Long-form Video UnderstandingPoster
- BOOTPLACE: Bootstrapped Object Placement with Detection TransformersPoster
- Balancing Two Classifiers via A Simplex ETF Structure for Model CalibrationPoster
- Bayesian Prompt Flow Learning for Zero-Shot Anomaly DetectionPoster
- Bayesian Test-Time Adaptation for Vision-Language ModelsPoster
- Be More Specific: Evaluating Object-centric Realism in Synthetic ImagesPoster
- Believing is Seeing: Unobserved Object Detection using Generative ModelsPoster
- Benchmarking Object Detectors under Real-World Distribution Shifts in Satellite ImageryPoster
- Beyond Background Shift: Rethinking Instance Replay in Continual Semantic SegmentationPoster
- Beyond Clean Training Data: A Versatile and Model-Agnostic Framework for Out-of-Distribution Detection with Contaminated Training DataPoster
- Beyond Generation: A Diffusion-based Low-level Feature Extractor for Detecting AI-generated ImagesPoster
- Beyond Human Perception: Understanding Multi-Object World from Monocular ViewPoster
- Beyond Image Classification: A Video Benchmark and Dual-Branch Hybrid Discrimination Framework for Compositional Zero-Shot LearningPoster
- Beyond Local Sharpness: Communication-Efficient Global Sharpness-aware Minimization for Federated LearningPoster
- Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual KnowledgePoster
- Beyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and HarmonizationPoster
- Beyond Words: Augmenting Discriminative Richness via Diffusions in Unsupervised Prompt LearningPoster
- Bias for Action: Video Implicit Neural Representations with Bias ModulationPoster
- Binarized Mamba-Transformer for Lightweight Quad Bayer HybridEVS DemosaicingPoster
- Binarized Neural Network for Multi-spectral Image FusionPoster
- BioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational PathologyPoster
- Black Hole-Driven Identity Absorbing in Diffusion ModelsPoster
- Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable EventsPoster
- BlenderGym: Benchmarking Foundational Model Systems for Graphics EditingHighlight
- Blind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion ModelPoster
- Blood Flow Speed Estimation with Optical Coherence Tomography Angiography ImagesPoster
- Blurry-Edges: Photon-Limited Depth Estimation from Defocused BoundariesPoster
- Boltzmann Attention Sampling for Image Analysis with Small ObjectsPoster
- Boost the Inference with Co-training: A Depth-guided Mutual Learning Framework for Semi-supervised Medical Polyp SegmentationPoster
- Boosting Adversarial Transferability through Augmentation in Hypothesis SpacePoster
- Boosting Domain Incremental Learning: Selecting the Optimal Parameters is All You NeedPoster
- Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal TransportPoster
- Boosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty EstimationPoster
- Brain-Inspired Spiking Neural Networks for Energy-Efficient Object DetectionPoster
- Breaking the Memory Barrier of Contrastive Loss via Tile-Based StrategyHighlight
- BrepGiff: Lightweight Generation of Complex B-rep with 3D GAT DiffusionPoster
- Bridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical FlowPoster
- Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsPoster
- Bridging Viewpoint Gaps: Geometric Reasoning Boosts Semantic CorrespondencePoster
- Bridging the Gap between Gaussian Diffusion Models and Universal Quantization for Image CompressionPoster
- Bridging the Vision-Brain Gap with an Uncertainty-Aware Blur PriorPoster
- Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical AnalysisPoster
- Building Vision Models upon Heat ConductionPoster
- ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free WayPoster
- CADDreamer: CAD Object Generation from Single-view ImagesHighlight
- CADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature LeveragingPoster
- CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language ModelsPoster
- CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D ImageHighlight
- CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual InteractionHighlight
- CASP: Compression of Large Multimodal Models Based on Attention SparsityHighlight
- CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional VideoPoster
- CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained AlignmentPoster
- CCIN: Compositional Conflict Identification and Neutralization for Composed Image RetrievalHighlight
- CGMatch: A Different Perspective of Semi-supervised LearningPoster
- CH3Depth: Efficient and Flexible Depth Foundation Model with Flow MatchingHighlight
- CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object RepresentationPoster
- CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCRPoster
- CLIP-driven Coarse-to-fine Semantic Guidance for Fine-grained Open-set Semi-supervised LearningPoster
- CLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair LossPoster
- CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based FrameworkPoster
- CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AIPoster
- COB-GS: Clear Object Boundaries in 3DGS Segmentation Based on Boundary-Adaptive Gaussian SplittingPoster
- COBRA: COmBinatorial Retrieval Augmentation for Few-Shot AdaptationPoster
- COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time AdaptationPoster
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution ShiftsHighlight
- CSC-PA: Cross-image Semantic Correlation via Prototype Attentions for Single-network Semi-supervised Breast Tumor SegmentationPoster
- CTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D DiffusionPoster
- CaMuViD: Calibration-Free Multi-View DetectionPoster
- CamPoint: Boosting Point Cloud Segmentation with Virtual CameraPoster
- Camera Resection from Known Line Pencils and a Radially Distorted ScanlinePoster
- Camouflage Anything: Learning to Hide using Controlled Out-painting and Representation EngineeringPoster
- Can Generative Video Models Help Pose Estimation?Highlight
- Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?Poster
- Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and UnderstandingHighlight
- Can Text-to-Video Generation help Video-Language Alignment?Poster
- Can't Slow Me Down: Learning Robust and Hardware-Adaptive Object Detectors against Latency Attacks for Edge DevicesPoster
- CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo BoothPoster
- Category-Agnostic Neural Object RiggingPoster
- Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial AttacksPoster
- Chain of Semantics Programming in 3D Gaussian Splatting Representation for 3D Vision GroundingPoster
- Change3D: Revisiting Change Detection and Captioning from A Video Modeling PerspectiveHighlight
- Channel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image DerainingPoster
- Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor ScenesPoster
- Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMsPoster
- Charm: The Missing Piece in ViT Fine-Tuning for Image Aesthetic AssessmentPoster
- Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion ModelsPoster
- ChatHuman: Chatting about 3D Humans with ToolsPoster
- CheXWorld: Exploring Image World Modeling for Radiograph Representation LearningPoster
- CheXwhatsApp: A Dataset for Exploring Challenges in the Diagnosis of Chest X-rays through Mobile DevicesPoster
- Cheb-GR: Rethinking K-nearest Neighbor Search in Re-ranking for Person Re-identificationPoster
- Chebyshev Attention Depth Permutation Texture Network with Latent Texture Attribute LossPoster
- CheckManual: A New Challenge and Benchmark for Manual-based Appliance ManipulationHighlight
- Classic Video Denoising in a Machine Learning World: Robust, Fast, and ControllablePoster
- Classifier-guided CLIP Distillation for Unsupervised Multi-label ClassificationPoster
- Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual ClassifiersPoster
- ClimbingCap: Multi-Modal Dataset and Method for Rock Climbing in World CoordinateHighlight
- Closest Neighbors are Harmful for Lightweight Masked Auto-encodersPoster
- Co-Speech Gesture Video Generation with Implicit Motion-Audio EntanglementPoster
- CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused ImagesPoster
- CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity QuantificationPoster
- CoMBO: Conflict Mitigation via Branched Optimization for Class Incremental SegmentationPoster
- CoMapGS: Covisibility Map-based Gaussian Splatting for Sparse Novel View SynthesisPoster
- CoMatcher: Multi-View Collaborative Feature MatchingPoster
- CoSER: Towards Consistent Dense Multiview Text-to-Image Generator for 3D CreationHighlight
- CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language ModelsPoster
- CocoER: Aligning Multi-Level Feature by Competition and Coordination for Emotion RecognitionPoster
- ColabSfM: Collaborative Structure-from-Motion by Point Cloud RegistrationPoster
- Collaborative Tree Search for Enhancing Embodied Multi-Agent CollaborationPoster
- Color Alignment in DiffusionPoster
- ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle MatricesPoster
- Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature SpacePoster
- Commonsense Video Question Answering through Video-Grounded Entailment Tree ReasoningPoster
- Compass Control: Multi Object Orientation Control for Text-to-Image GenerationPoster
- Composing Parts for Expressive Object GenerationPoster
- Compositional Caching for Training-free Open-vocabulary Attribute DetectionHighlight
- Compositional Targeted Multi-Label Universal PerturbationsPoster
- Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision TransformersHighlight
- Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and HarmonizationPoster
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalPoster
- Concept Lancet: Image Editing with Compositional Representation TransplantPoster
- Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph GenerationPoster
- Conformal Prediction for Zero-Shot ModelsPoster
- Conical Visual Concentration for Efficient Large Vision-Language ModelsPoster
- Consistency Posterior Sampling for Diverse Image SynthesisPoster
- Consistency-aware Self-Training for Iterative-based Stereo MatchingPoster
- Consistent Normal Orientation for 3D Point Clouds via Least Squares on Delaunay GraphPoster
- Consistent and Controllable Image Animation with Motion Diffusion ModelsPoster
- Continuous Adverse Weather Removal via Degradation-Aware DistillationPoster
- Continuous Locomotive Crowd Behavior GenerationPoster
- Continuous Space-Time Video Resampling with Invertible Motion SteganographyPoster
- ControlFace: Harnessing Facial Parametric Control for Face RiggingPoster
- Controllable Human Image Generation with Personalized Multi-GarmentsPoster
- Convex Combination Star Shape Prior for Data-driven Image Semantic SegmentationPoster
- Convex Relaxation for Robust Vanishing Point Estimation in Manhattan WorldAward Candidate
- CorrBEV: Multi-View 3D Object Detection by Correlation Learning with Multi-modal PrototypesPoster
- Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt TuningPoster
- Crab: A Unified Audio-Visual Scene Understanding Model with Explicit CooperationPoster
- CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry RefinerPoster
- Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian SplattingHighlight
- CroCoDL: Cross-device Collaborative Dataset for LocalizationPoster
- Cross-Modal 3D Representation with Multi-View Images and Point CloudsPoster
- Cross-Modal Distillation for 2D/3D Multi-Object Discovery from 2D MotionPoster
- Cross-Modal Interactive Perception Network with Mamba for Lung Tumor Segmentation in PET-CT ImagesPoster
- Cross-Rejective Open-Set SAR Image RegistrationPoster
- CrossOver: 3D Scene Cross-Modal AlignmentHighlight
- CrossSDF: 3D Reconstruction of Thin Structures From Cross-SectionsPoster
- CryptoFace: End-to-End Encrypted Face RecognitionPoster
- CustAny: Customizing Anything from A Single ExamplePoster
- Customized Condition Controllable Generation for Video SoundtrackPoster
- D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective Recognition.Poster
- DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision TransformersPoster
- DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report GenerationPoster
- DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image FusionPoster
- DEAL: Data-Efficient Adversarial Learning for High-Quality Infrared ImagingPoster
- DFM: Differentiable Feature Matching for Anomaly DetectionPoster
- DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings LearningPoster
- DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-IDPoster
- DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion ModelsAward Candidate
- DIO: Decomposable Implicit 4D Occupancy-Flow World ModelPoster
- DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric VideosHighlight
- DKC: Differentiated Knowledge Consolidation for Cloth-Hybrid Lifelong Person Re-identificationPoster
- DL2G: Degradation-guided Local-to-Global Restoration for Eyeglass Reflection RemovalPoster
- DOF-GS: Adjustable Depth-of-Field 3D Gaussian Splatting for Post-Capture Refocusing, Defocus Rendering and Blur RemovalPoster
- DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-ResolutionPoster
- DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid FrameworkPoster
- DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic SegmentationPoster
- DSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot SegmentationPoster
- DTGBrepGen: A Novel B-rep Generative Model through Decoupling Topology and GeometryPoster
- DTOS: Dynamic Time Object Sensing with Large Multimodal ModelPoster
- DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D TeachersPoster
- DVHGNN: Multi-Scale Dilated Vision HGNN for Efficient Vision RecognitionPoster
- DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression ComprehensionPoster
- D^2iT: Dynamic Diffusion Transformer for Accurate Image GenerationPoster
- D^3CTTA: Domain-Dependent Decorrelation for Continual Test-Time Adaption of 3D LiDAR SegmentationPoster
- DaCapo: Score Distillation as Stacked Bridge for Fast and High-quality 3D EditingPoster
- DashGaussian: Optimizing 3D Gaussian Splatting in 200 SecondsHighlight
- Data Distributional Properties As Inductive Bias for Systematic GeneralizationPoster
- Data Synthesis with Diverse Styles for Face Recognition via 3DMM-Guided DiffusionPoster
- Data-Free Group-Wise Fully Quantized Winograd Convolution via Learnable ScalesPoster
- Data-free Universal Adversarial Perturbation with Pseudo-semantic PriorPoster
- DeCLIP: Decoupled Learning for Open-Vocabulary Dense PerceptionPoster
- DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long VideosPoster
- DeClotH: Decomposable 3D Cloth and Human Body Reconstruction from a Single ImagePoster
- DeDe: Detecting Backdoor Samples for SSL Encoders via DecodersPoster
- De^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze EstimationPoster
- Decoder Gradient Shield: Provable and High-Fidelity Prevention of Gradient-Based Box-Free Watermark RemovalPoster
- Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image CompressionPoster
- Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task LearningPoster
- Decoupled Motion Expression Video SegmentationPoster
- Decoupling Training-Free Guided Diffusion by ADMMPoster
- Deep Change Monitoring: A Hyperbolic Representative Learning Framework and a Dataset for Long-term Fine-grained Tree Change DetectionHighlight
- Deep Fair Multi-View Clustering with Attention KANHighlight
- DeepCompress-ViT: Rethinking Model Compression to Enhance Efficiency of Vision Transformers at the EdgePoster
- DeepLA-Net: Very Deep Local Aggregation Networks for Point Cloud AnalysisPoster
- DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual InspectionHighlight
- DeformCL: Learning Deformable Centerline Representation for Vessel Extraction in 3D Medical ImagePoster
- Degradation-Aware Feature Perturbation for All-in-One Image RestorationPoster
- DejaVid: Encoder-Agnostic Learned Temporal Matching for Video ClassificationPoster
- Dense Dispersed Structured Light for Hyperspectral 3D Imaging of Dynamic ScenesPoster
- Dense Match Summarization for Faster Two-view EstimationPoster
- Dense-SfM: Structure from Motion with Dense Consistent MatchingPoster
- Depth-Guided Bundle Sampling for Efficient Generalizable Neural Radiance Field ReconstructionPoster
- Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAIPoster
- Descriptor-In-Pixel : Point-Feature Tracking For Pixel Processor ArraysAward Candidate
- Detail-Preserving Latent Diffusion for Stable Shadow RemovalPoster
- Detect Any Mirrors: Boosting Learning Reliability on Large-Scale Unlabeled Data with an Iterative Data EnginePoster
- Detecting Open World Objects via Partial Attribute AssignmentPoster
- Detection-Friendly Nonuniformity Correction: A Union Framework for Infrared UAV Target DetectionHighlight
- Deterministic Image-to-Image Translation via Denoising Brownian Bridge Models with Dual ApproximatorsPoster
- Deterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion SynthesisPoster
- Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified AttentionPoster
- DexHandDiff: Interaction-aware Diffusion Planning for Adaptive Dexterous ManipulationPoster
- DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian SplattingPoster
- DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection TransformerPoster
- DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy LabelsPoster
- DiSRT-In-Bed: Diffusion-Based Sim-to-Real Transfer Framework for In-Bed Human Mesh RecoveryPoster
- DiSciPLE: Learning Interpretable Programs for Scientific Visual DiscoveryPoster
- DiTASK: Multi-Task Fine-Tuning with Diffeomorphic TransformationsPoster
- Diff-Palm: Realistic Palmprint Generation with Polynomial Creases and Intra-Class Variation Controllable Diffusion ModelsPoster
- Diff2Flow: Training Flow Matching Models via Diffusion Model AlignmentPoster
- DiffCAM: Data-Driven Saliency Maps by Capturing Feature DifferencesHighlight
- DiffLO: Semantic-Aware LiDAR Odometry with Diffusion-Based RefinementPoster
- DiffLocks: Generating 3D Hair from a Single Image using Diffusion ModelsPoster
- DiffVsgg: Diffusion-Driven Online Video Scene Graph GenerationPoster
- Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy GenerationPoster
- Differentiable Inverse Rendering with Interpretable Basis BRDFsPoster
- Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image CaptioningPoster
- Diffusion Model is Effectively Its Own TeacherPoster
- Diffusion-based Event Generation for High-Quality Image DeblurringPoster
- Diffusion-based Realistic Listening Head Generation via Hybrid Motion ModelingHighlight
- DirectTriGS: Triplane-based Gaussian Splatting Field Representation for 3D GenerationPoster
- Directional Label Diffusion Model for Learning from Noisy LabelsPoster
- DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text RetrievalPoster
- Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant LearningPoster
- Discrete to Continuous: Generating Smooth Transition Poses from Sign Language ObservationsPoster
- Disentangled Pose and Appearance Guidance for Multi-Pose GenerationPoster
- Disentangling Safe and Unsafe Image Corruptions via Anisotropy and LocalityPoster
- DiskVPS: Vanishing Point Detector via Hough Transform in a Disk RegionPoster
- Dissecting and Mitigating Diffusion Bias via Mechanistic InterpretabilityPoster
- Distilled Prompt Learning for Incomplete Multimodal Survival PredictionPoster
- Distilling Spatially-Heterogeneous Distortion Perception for Blind Image Quality AssessmentPoster
- Distinguish Then Exploit: Source-free Open Set Domain Adaptation via Weight Barcode Estimation and Sparse Label AssignmentPoster
- Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly DetectionPoster
- DiverseFlow: Sample-Efficient Diverse Mode Coverage in FlowsPoster
- DnLUT: Ultra-Efficient Color Image Denoising via Channel-Aware Lookup TablesPoster
- Do ImageNet-trained Models Learn Shortcuts? The Impact of Frequency Shortcuts on GeneralizationPoster
- Do Visual Imaginations Improve Vision-and-Language Navigation Agents?Poster
- Do Your Best and Get Enough Rest for Continual LearningPoster
- DoF-Gaussian: Controllable Depth-of-Field for 3D Gaussian SplattingPoster
- DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document UnderstandingPoster
- Docopilot: Improving Multimodal Models for Document-Level UnderstandingPoster
- Domain Adaptive Diabetic Retinopathy Grading with Model Absence and Flowing DataPoster
- Domain Generalization in CLIP via Learning with Diverse Text PromptsPoster
- Doppelgangers and Adversarial VulnerabilityHighlight
- Doppelgangers++: Improved Visual Disambiguation with Geometric 3D FeaturesHighlight
- Dragin3D: Image Editing by Dragging in 3D SpacePoster
- DreamTrack: Dreaming the Future for Multimodal Visual Object TrackingPoster
- DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous DrivingHighlight
- DriveScape: High-Resolution Driving Video Generation by Multi-View Feature FusionPoster
- DropGaussian: Structural Regularization for Sparse-view Gaussian SplattingPoster
- DropoutGS: Dropping Out Gaussians for Better Sparse-view RenderingPoster
- Dual Energy-Based Model with Open-World Uncertainty Estimation for Out-of-distribution DetectionPoster
- Dual Exposure Stereo for Extended Dynamic Range 3D ImagingPoster
- Dual Focus-Attention Transformer for Robust Point Cloud RegistrationPoster
- Dual Semantic Guidance for Open Vocabulary Semantic SegmentationPoster
- Dual-Agent Optimization framework for Cross-Domain Few-Shot SegmentationPoster
- Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General PansharpeningPoster
- DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image SegmentationPoster
- DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective SchedulingPoster
- DynPose: Largely Improving the Efficiency of Human Pose Estimation by a Simple Dynamic FrameworkPoster
- DynScene: Scalable Generation of Dynamic Robotic Manipulation Scenes for Embodied AIPoster
- DynaMoDe-NeRF: Motion-aware Deblurring Neural Radiance Field for Dynamic ScenesPoster
- Dynamic Camera Poses and Where to Find ThemPoster
- Dynamic Content Prediction with Motion-aware Priors for Blind Face Video RestorationPoster
- Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio SemanticsPoster
- Dynamic Group Normalization: Spatio-Temporal Adaptation to Evolving Data StatisticsPoster
- Dynamic Motion Blending for Versatile Motion EditingPoster
- Dynamic Neural Surfaces for Elastic 4D Shape Representation and AnalysisPoster
- Dynamic Pseudo Labeling via Gradient Cutting for High-Low Entropy ExplorationPoster
- Dynamic Stereotype Theory Induced Micro-expression Recognition with Oriented DeformationPoster
- Dynamic Updates for Language Adaptation in Visual-Language TrackingPoster
- EAP-GS: Efficient Augmentation of Pointcloud for 3D Gaussian Splatting in Few-shot Scene ReconstructionPoster
- EASEMVC:Efficient Dual Selection Mechanism for Deep Multi-View ClusteringPoster
- EBS-EKF: Accurate and High Frequency Event-based Star TrackingHighlight
- EDCFlow: Exploring Temporally Dense Difference Maps for Event-based Optical Flow EstimationPoster
- EDM: Equirectangular Projection-Oriented Dense Kernelized Feature MatchingPoster
- EIDT-V: Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video GenerationPoster
- ERUPT: Efficient Rendering with Unposed Patch TransformerPoster
- ESC: Erasing Space Concept for Knowledge DeletionHighlight
- ESCAPE: Equivariant Shape Completion via Anchor Point EncodingPoster
- ETAP: Event-based Tracking of Any PointHighlight
- EVOS: Efficient Implicit Neural Training via EVOlutionary SelectorPoster
- EVPGS: Enhanced View Prior Guidance for Splatting-based Extrapolated View SynthesisPoster
- EVolSplat: Efficient Volume-based Gaussian Splatting for Urban View SynthesisPoster
- Early-Bird Diffusion: Investigating and Leveraging Timestep-Aware Early-Bird Tickets in Diffusion Models for Efficient TrainingPoster
- Easy-editable Image Vectorization with Multi-layer Multi-scale Distributed Visual Feature EmbeddingPoster
- EasyCraft: A Robust and Efficient Framework for Automatic Avatar CraftingPoster
- EchoMatch: Partial-to-Partial Shape Matching via Correspondence ReflectionPoster
- EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe GuidancePoster
- Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional ConditioningPoster
- EdgeDiff: Edge-aware Diffusion Network for Building Reconstruction from Point CloudsPoster
- EdgeMovingNet: Edge-preserving Point Cloud Reconstruction via Joint Geometry FeaturesPoster
- EdgeTAM: On-Device Track Anything ModelPoster
- Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative EditingPoster
- Effective SAM Combination for Open-Vocabulary Semantic SegmentationPoster
- EffiDec3D: An Optimized Decoder for High-Performance and Efficient 3D Medical Image SegmentationHighlight
- Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise ReplacementPoster
- Efficient Data Driven Mixture-of-Expert Extraction from Trained NetworksPoster
- Efficient Decoupled Feature 3D Gaussian Splatting via Hierarchical CompressionPoster
- Efficient Depth Estimation for Unstable Stereo Camera Systems on AR GlassesPoster
- Efficient Diffusion as Low Light EnhancerPoster
- Efficient Dynamic Scene Editing via 4D Gaussian-based Static-Dynamic SeparationPoster
- Efficient Event-Based Object Detection: A Hybrid Neural Network with Spatial and Temporal AttentionPoster
- Efficient Motion-Aware Video MLLMHighlight
- Efficient Personalization of Quantized Diffusion Model without BackpropagationPoster
- Efficient Test-time Adaptive Object Detection via Sensitivity-Guided PruningPoster
- Efficient Video Super-Resolution for Real-time Rendering with Decoupled G-buffer GuidancePoster
- EfficientLLaVA: Generalizable Auto-Pruning for Large Vision-language ModelsPoster
- Effortless Active Labeling for Long-Term Test-Time AdaptationPoster
- Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal InputPoster
- EigenGS Representation: From Eigenspace to Gaussian Image SpacePoster
- Electromyography-Informed Facial Expression Reconstruction for Physiological-Based Synthesis and AnalysisHighlight
- Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context LearningPoster
- EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video DiffusionPoster
- Empowering Large Language Models with 3D Situation AwarenessPoster
- Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video SynthesisPoster
- End-to-End HOI Reconstruction Transformer with Graph-based EncodingHighlight
- End-to-End Implicit Neural Representations for ClassificationPoster
- Enduring, Efficient and Robust Trajectory Prediction Attack in Autonomous Driving via Optimization-Driven Multi-Frame Perturbation FrameworkHighlight
- Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal RepresentationsPoster
- Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute RecognitionHighlight
- Enhanced then Progressive Fusion with View Graph for Multi-View ClusteringPoster
- Enhancing 3D Gaze Estimation in the Wild using Weak Supervision with Gaze Following LabelsPoster
- Enhancing Adversarial Transferability with Checkpoints of a Single Model's TrainingPoster
- Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion ModelPoster
- Enhancing Facial Privacy Protection via Weakening Diffusion PurificationPoster
- Enhancing Few-Shot Class-Incremental Learning via Training-Free Bi-Level Modality CalibrationPoster
- Enhancing Online Continual Learning with Plug-and-Play State Space Model and Class-Conditional Mixture of DiscretizationPoster
- Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion ModelsPoster
- Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image SegmentationPoster
- Enhancing Testing-Time Robustness for Trusted Multi-View Classification in the WildPoster
- Enhancing Video-LLM Reasoning via Agent-of-Thoughts DistillationPoster
- Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise SchedulingPoster
- EnliveningGS: Active Locomotion of 3DGSPoster
- EntityErasure: Erasing Entity Cleanly via Amodal Entity Segmentation and CompletionPoster
- EntitySAM: Segment Everything in VideoPoster
- EntropyMark: Towards More Harmless Backdoor Watermark via Entropy-based Constraint for Open-source Dataset Copyright ProtectionPoster
- EquiPose: Exploiting Permutation Equivariance for Relative Camera Pose EstimationPoster
- Erase Diffusion: Empowering Object Removal Through Calibrating Diffusion PathwaysHighlight
- Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent SpacesPoster
- EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with EventsHighlight
- EvOcc: Accurate Semantic Occupancy for Automated Driving Using Evidence TheoryPoster
- Event Ellipsometer: Event-based Mueller-Matrix Video ImagingHighlight
- Event-Equalized Dense Video CaptioningPoster
- EventFly: Event Camera Perception from Ground to the SkyPoster
- EventPSR: Surface Normal and Reflectance Estimation from Photometric Stereo Using an Event CameraHighlight
- Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and BeyondPoster
- Explainable Saliency: Articulating Reasoning with Contextual PrioritizationPoster
- Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image ClassificationPoster
- Explaining in Diffusion: Explaining a Classifier with Diffusion SemanticsPoster
- Explicit Depth-Aware Blurry Video Frame Interpolation Guided by Differential CurvesPoster
- Exploiting Deblurring Networks for Radiance FieldsPoster
- Exploration-Driven Generative Interactive EnvironmentsPoster
- Exploring Contextual Attribute Density in Referring Expression CountingPoster
- Exploring Historical Information for RGBE Visual Tracking with MambaPoster
- Exploring Scene Affinity for Semi-Supervised LiDAR Semantic SegmentationPoster
- Exploring Temporally-Aware Features for Point TrackingPoster
- Exploring Timeline Control for Facial Motion GenerationPoster
- Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image SynthesisPoster
- Exposure-slot: Exposure-centric Representations Learning with Slot-in-Slot Attention for Region-aware Exposure CorrectionPoster
- FADE: Frequency-Aware Diffusion Model Factorization for Video EditingPoster
- FALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene UnderstandingPoster
- FASTer: Focal token Acquiring-and-Scaling Transformer for Long-term 3D Objection DetectionPoster
- FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image EditingPoster
- FFR: Frequency Feature Rectification for Weakly Supervised Semantic SegmentationPoster
- FFaceNeRF: Few-shot Face Editing in Neural Radiance FieldsPoster
- FG^2: Fine-Grained Cross-View Localization by Fine-Grained Feature MatchingPoster
- FIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze EstimationPoster
- FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix ApproximationHighlight
- FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse ViewsPoster
- FLAVC: Learned Video Compression with Feature Level AttentionPoster
- FRAME: Floor-aligned Representation for Avatar Motion from Egocentric VideoHighlight
- FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question AnsweringPoster
- FRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few ImagesHighlight
- FSBench: A Figure Skating Benchmark for Advancing Artistic Sports UnderstandingPoster
- FSHNet: Fully Sparse Hybrid Network for 3D Object DetectionPoster
- Face Forgery Video Detection via Temporal Forgery Cue UnravelingPoster
- FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMsPoster
- FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation ModelsPoster
- Fast and Accurate Gigapixel Pathological Image Classification with Hierarchical Distillation Multi-Instance LearningPoster
- Faster Parameter-Efficient Tuning with Token Redundancy ReductionPoster
- Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object DetectionPoster
- Feature Selection for Latent Factor ModelsPoster
- Feature Spectrum Learning for Remote Sensing Change DetectionPoster
- Feature-Preserving Mesh Decimation for Normal IntegrationPoster
- FedAWA: Adaptive Optimization of Aggregation Weights in Federated Learning Using Client VectorsPoster
- FedCALM: Conflict-aware Layer-wise Mitigation for Selective Aggregation in Deeper Personalized Federated LearningPoster
- FedCS: Coreset Selection for Federated LearningPoster
- FedSPA: Generalizable Federated Graph Learning under Homophily HeterogeneityPoster
- FeedEdit: Text-Based Image Editing with Dynamic Feedback RegulationPoster
- Ferret: An Efficient Online Continual Learning Framework under Varying Memory ConstraintsPoster
- Few-shot Implicit Function Generation via EquivarianceHighlight
- Few-shot Personalized Scanpath PredictionPoster
- FiRe: Fixed-points of Restoration Priors for Solving Inverse ProblemsPoster
- Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part SegmentationPoster
- FinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal GuidancePoster
- Finer-CAM: Spotting the Difference Reveals Finer Details for Visual ExplanationPoster
- Fingerprinting Denoising Diffusion Probabilistic ModelsPoster
- Finsler Multi-Dimensional Scaling: Manifold Learning for Asymmetric Dimensionality Reduction and EmbeddingPoster
- FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized SegmentationPoster
- Fitted Neural Lossless Image CompressionPoster
- Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry LocalityHighlight
- FlexDrive: Toward Trajectory Flexibility in Driving Scene Gaussian Splatting Reconstruction and RenderingPoster
- FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian SplattingPoster
- FlexUOD: The Answer to Real-world Unsupervised Image Outlier DetectionPoster
- Flexible Frame Selection for Efficient Video ReasoningPoster
- Flexible Group Count Enables Hassle-Free Structured PruningPoster
- Flow-NeRF: Joint Learning of Geometry, Poses, and Dense Flow within Unified Neural RepresentationsPoster
- FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic ManipulationPoster
- Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality EvolutionHighlight
- FluxSpace: Disentangled Semantic Editing in Rectified Flow ModelsPoster
- Focal Split: Untethered Snapshot Depth from Differential DefocusPoster
- Focusing on Tracks for Online Multi-Object TrackingPoster
- Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional FlowsPoster
- Font-Agent: Enhancing Font Understanding with Large Language ModelsPoster
- ForestLPR: LiDAR Place Recognition in Forests Attentioning Multiple BEV Density ImagesHighlight
- Forming Auxiliary High-confident Instance-level Loss to Promote Learning from Label ProportionsPoster
- Fortifying Federated Learning Towards Trustworthiness via Auditable Data Valuation and Verifiable Client ContributionPoster
- FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image GenerationHighlight
- Foveated Instance SegmentationPoster
- Fractal Calibration for Long-tailed Object DetectionPoster
- Free Lunch Enhancements for Multi-modal Crowd CountingPoster
- Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EMPoster
- Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed ViewsPoster
- FreeCloth: Free-form Generation Enhances Challenging Clothed Human ModelingHighlight
- FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian VelocityPoster
- FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component AnalysisHighlight
- FreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free PromptsPoster
- FreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene ReconstructionPoster
- FreeUV: Ground-Truth-Free Realistic Facial UV Texture Recovery via Cross-Assembly Inference StrategyPoster
- FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency DebiasingPoster
- Frequency Dynamic Convolution for Dense Image PredictionPoster
- Frequency-Biased Synergistic Design for Image Compression and CompensationPoster
- From Elements to Design: A Layered Approach for Automatic Graphic Design CompositionPoster
- From Head to Tail: Efficient Black-box Model Inversion Attack via Long-tailed LearningPoster
- From Prototypes to General Distributions: An Efficient Curriculum for Masked Image ModelingPoster
- From Sparse Signal to Smooth Motion: Real-Time Motion Generation with Rolling Prediction ModelsPoster
- From Sparse to Dense: Camera Relocalization with Scene-Specific Detector from Feature Gaussian SplattingPoster
- From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral PerspectivePoster
- FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned PriorsPoster
- FruitNinja: 3D Object Interior Texture Generation with Gaussian SplattingPoster
- Functionality Understanding and Segmentation in 3D ScenesHighlight
- Fuzzy Multimodal Learning for Trusted Cross-modal RetrievalPoster
- GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context EncodingPoster
- GASP: Gaussian Avatars with Synthetic PriorsPoster
- GBC-Splat: Generalizable Gaussian-Based Clothed Human Digitalization under Sparse RGB CamerasPoster
- GBlobs: Explicit Local Structure via Gaussian Blobs for Improved Cross-Domain LiDAR-based 3D Object DetectionPoster
- GCC: Generative Color Constancy via Diffusing a Color CheckerPoster
- GENIUS: A Generative Framework for Universal Multimodal SearchPoster
- GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following ManipulationPoster
- GIF: Generative Inspiration for Face Recognition at ScalePoster
- GIFStream: 4D Gaussian-based Immersive Video with Feature StreamPoster
- GIVEPose: Gradual Intra-class Variation Elimination for RGB-based Category-Level Object Pose EstimationPoster
- GLASS: Guided Latent Slot Diffusion for Object-Centric LearningPoster
- GLane3D: Detecting Lanes with Graph of 3D KeypointsPoster
- GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object DetectorPoster
- GOAL: Global-local Object Alignment LearningPoster
- GPAvatar: High-fidelity Head Avatars by Learning Efficient Gaussian ProjectionsPoster
- GPS as a Control Signal for Image GenerationPoster
- GPVK-VL: Geometry-Preserving Virtual Keyframes for Visual Localization under Large Viewpoint ChangesPoster
- GRAE-3DMOT: Geometry Relation-Aware Encoder for Online 3D Multi-Object TrackingPoster
- GS-2DGS: Geometrically Supervised 2DGS for Reflective Object ReconstructionPoster
- GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point TrackingPoster
- GUI-Xplore: Empowering Generalizable GUI Agents with One ExplorationPoster
- GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D ReconstructionPoster
- Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring ClassesPoster
- Galaxy Walker: Geometry-aware VLMs For Galaxy-scale UnderstandingHighlight
- GauCho: Gaussian Distributions with Cholesky Decomposition for Oriented Object DetectionPoster
- GaussHDR: High Dynamic Range Gaussian Splatting via Learning Unified 3D and 2D Local Tone MappingPoster
- Gaussian Splatting Feature Fields for (Privacy-Preserving) Visual LocalizationPoster
- GaussianIP: Identity-Preserving Realistic 3D Human Generation via Human-Centric Diffusion PriorPoster
- GazeGene: Large-scale Synthetic Gaze Dataset with 3D Eyeball AnnotationsPoster
- Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual ForagingPoster
- Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D ObjectsPoster
- GenAssets: Generating in-the-wild 3D Assets in Latent SpacePoster
- GenFusion: Closing the Loop between Reconstruction and Generation via VideosPoster
- GenPC: Zero-shot Point Cloud Completion via 3D Generative PriorsPoster
- GenVDM: Generating Vector Displacement Maps From a Single ImageHighlight
- Generalizable Object Keypoint Localization from Generative PriorsPoster
- Generalized Diffusion Detector: Mining Robust Features from Diffusion Models for Domain-Generalized DetectionPoster
- Generalized Gaussian Entropy Model for Point Cloud Attribute Compression with Dynamic Likelihood IntervalsPoster
- Generalized Zero-Shot Classification via Semantics-Free Inter-Class Feature GenerationPoster
- Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric VisionHighlight
- Generative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D ReconstructionHighlight
- Generative Hard Example Augmentation for Semantic Point Cloud SegmentationPoster
- Generative Map Priors for Collaborative BEV Semantic SegmentationPoster
- Generative Multimodal Pretraining with Discrete Diffusion Timestep TokensAward Candidate
- Generative PhotomontagePoster
- Generative Sparse-View Gaussian SplattingPoster
- Generative Zero-Shot Composed Image RetrievalPoster
- GeoAvatar: Geometrically-Consistent Multi-Person Avatar Reconstruction from Sparse Multi-View VideosPoster
- GeoDepth: From Point-to-Depth to Plane-to-Depth Modeling for Self-Supervised Monocular Depth EstimationPoster
- GeoMM: On Geodesic Perspective for Multi-modal LearningPoster
- Geometric Knowledge-Guided Localized Global Distribution Alignment for Federated LearningPoster
- Geometry in Style: 3D Stylization via Surface Normal DeformationPoster
- Geometry-guided Online 3D Video Synthesis with Multi-View Temporal ConsistencyPoster
- Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference UnderstandingPoster
- GliaNet: Adaptive Neural Network Structure Learning with Glia-DrivenPoster
- Glossy Object Reconstruction with Cost-effective Polarized AcquisitionHighlight
- GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text EditingPoster
- GoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View SynthesisPoster
- Golden Cudgel Network for Real-Time Semantic SegmentationPoster
- Gradient Inversion Attacks on Parameter-Efficient Fine-TuningPoster
- Gradient-Guided Annealing for Domain GeneralizationHighlight
- Graph Neural Network Combining Event Stream and Periodic Aggregation for Low-Latency Event-based VisionHighlight
- Graph-Embedded Structure-Aware Perceptual Hashing for Neural Network Protection and Piracy DetectionPoster
- GraphI2P: Image-to-Point Cloud Registration with Exploring Pattern of Correspondence via Graph LearningPoster
- GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy LearningPoster
- Gromov-Wasserstein Problem with Cyclic SymmetryPoster
- GroomLight: Hybrid Inverse Rendering for Relightable Human Hair Appearance ModelingPoster
- Ground-V: Teaching VLMs to Ground Complex Instructions in PixelsPoster
- Grounding 3D Object Affordance with Language Instructions, Visual Observations and InteractionsPoster
- GroundingFace: Fine-grained Face Understanding via Pixel Grounding Multimodal Large Language ModelHighlight
- GroupMamba: Efficient Group-Based Visual State Space ModelPoster
- GuardSplat: Efficient and Robust Watermarking for 3D Gaussian SplattingPoster
- H-MoRe: Learning Human-centric Motion Representation for Action AnalysisHighlight
- H2ST: Hierarchical Two-Sample Tests for Continual Out-of-Distribution DetectionPoster
- HERA: Hybrid Explicit Representation for Ultra-Realistic Head AvatarsPoster
- HMAR: Efficient Hierarchical Masked Auto-Regressive Image GenerationPoster
- HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video GenerationPoster
- HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture GenerationPoster
- HORP: Human-Object Relation Priors Guided HOI DetectionPoster
- HOT: Hadamard-based Optimized TrainingPoster
- HOTFormerLoc: Hierarchical Octree Transformer for Versatile Lidar Place Recognition Across Ground and Aerial ViewsPoster
- HRAvatar: High-Quality and Relightable Gaussian Head AvatarPoster
- HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene InteractionHighlight
- HSI: A Holistic Style Injector for Arbitrary Style TransferPoster
- HUNet: Homotopy Unfolding Network for Image Compressive SensingPoster
- HUSH: Holistic Panoramic 3D Scene Understanding using Spherical HarmonicsPoster
- HalLoc: Token-level Localization of Hallucinations for Vision Language ModelsPoster
- Hand-held Object Reconstruction from RGB Video with Dynamic InteractionPoster
- HandOS: 3D Hand Reconstruction in One StagePoster
- Hardware-Rasterized Ray-Based Gaussian SplattingHighlight
- Harnessing Frequency Spectrum Insights for Image Copyright Protection Against Diffusion ModelsPoster
- Harnessing Frozen Unimodal Encoders for Flexible Multimodal AlignmentPoster
- Harnessing Global-Local Collaborative Adversarial Perturbation for Anti-CustomizationPoster
- Hazy Low-Quality Satellite Video Restoration Via Learning Optimal Joint Degradation Patterns and Continuous-Scale Super-Resolution ReconstructionPoster
- HeMoRa: Unsupervised Heuristic Consensus Sampling for Robust Point Cloud RegistrationPoster
- Hearing Anywhere in Any EnvironmentPoster
- Hearing Hands: Generating Sounds from Physical Interactions in 3D ScenesPoster
- HeatFormer: A Neural Optimizer for Multiview Human Mesh RecoveryPoster
- Heterogeneous Skeleton-Based Action Representation LearningPoster
- HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face FusionPoster
- HiMoR: Monocular Deformable Gaussian Reconstruction with Hierarchical Motion RepresentationPoster
- Hiding Images in Diffusion Models by Editing Learned Score FunctionsPoster
- Hierarchical Adaptive Filtering Network for Text Image Specular Highlight RemovalPoster
- Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric LearningPoster
- Hierarchical Features Matter: A Deep Exploration of Progressive Parameterization Method for Dataset DistillationPoster
- Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D GenerationPoster
- Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time AdaptationPoster
- High Dynamic Range Video Compression: A Large-Scale Benchmark Dataset and A Learned Bit-depth Scalable Compression AlgorithmPoster
- High Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous FlightPoster
- High-Fidelity Lightweight Mesh Reconstruction from Point CloudsHighlight
- High-fidelity 3D Object Generation from Single Image with RGBN-Volume Gaussian Reconstruction ModelHighlight
- High-quality Point Cloud Oriented Normal Estimation via Hybrid Angular and Euclidean Distance EncodingPoster
- Higher-Order Ratio Cycles for Fast and Globally Optimal Shape MatchingPoster
- HistoFS: Non-IID Histopathologic Whole Slide Image Classification via Federated Style Transfer with RoI-PreservingPoster
- HoGS: Unified Near and Far Object Reconstruction via Homogeneous Gaussian SplattingPoster
- HomoGen: Enhanced Video Inpainting via Homography Propagation and DiffusionPoster
- Homogeneous Dynamics Space for Heterogeneous HumansPoster
- HotSpot: Signed Distance Function Optimization with an Asymptotically Sufficient ConditionHighlight
- HuMoCon: Concept Discovery for Human Motion UnderstandingPoster
- HuPerFlow: A Comprehensive Benchmark for Human vs. Machine Motion Estimation ComparisonHighlight
- Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identificationPoster
- HumanMM: Global Human Motion Recovery from Multi-shot VideosPoster
- Hybrid Concept Bottleneck ModelsPoster
- Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image SegmentationPoster
- Hybrid Reciprocal Transformer with Triplet Feature Alignment for Scene Graph GenerationPoster
- HybridMQA: Exploring Geometry-Texture Interactions for Colored Mesh Quality AssessmentPoster
- HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing ImageryPoster
- HyperLoRA: Parameter-Efficient Adaptive Generation for Portrait SynthesisHighlight
- HyperNVD: Accelerating Neural Video Decomposition via HypernetworksPoster
- HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight TrajectoriesPoster
- HyperPose: Hypernetwork-Infused Camera Pose Localization and an Extended Cambridge Landmarks DatasetPoster
- HyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual PerceiverPoster
- Hyperbolic Category DiscoveryPoster
- Hyperbolic Safety-Aware Vision-Language ModelsHighlight
- Hyperbolic Uncertainty-Aware Few-Shot Incremental Point Cloud SegmentationPoster
- Hypergraph Vision Transformers: Images are More than Nodes, More than EdgesPoster
- Hyperspectral Pansharpening via Diffusion Models with Iteratively Zero-Shot GuidancePoster
- I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video ModelsPoster
- IAAO: Interactive Affordance Learning for Articulated Objects in 3D EnvironmentsPoster
- ICE: Intrinsic Concept Extraction from a Single Image via Diffusion ModelsHighlight
- ICP: Immediate Compensation Pruning for Mid-to-high SparsityHighlight
- IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular VideosCPoster
- IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot MannerPoster
- IMFine: 3D Inpainting via Geometry-guided Multi-view RefinementPoster
- ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-OnPoster
- IceDiff: High Resolution and High-Quality Arctic Sea Ice Forecasting with Generative Diffusion PriorPoster
- Identifying and Mitigating Spurious Correlation in Multi-Task LearningPoster
- Identity-Clothing Similarity Modeling for Unsupervised Clothing Change Person Re-IdentificationPoster
- Identity-preserving Distillation Sampling by Fixed-Point IteratorPoster
- Illumination Spectrum Estimation for Multispectral Images via Surface Reflectance Modeling and Spatial-Spectral Feature GenerationPoster
- ImViD: Immersive Volumetric Videos for Enhanced VR EngagementHighlight
- Image Generation Diversity Issues and How to Tame ThemPoster
- Image Over Text: Transforming Formula Recognition Evaluation with Character Detection MatchingPoster
- Image Quality Assessment: From Human to Machine PreferenceHighlight
- Image Quality Assessment: Investigating Causal Perceptual Effects with Abductive Counterfactual InferencePoster
- Image Reconstruction from Readout-Multiplexed Single-Photon Detector ArraysHighlight
- Image Referenced Sketch Colorization Based on Animation Creation WorkflowPoster
- Image is All You Need to Empower Large-scale Diffusion Models for In-Domain GenerationPoster
- ImagineFSL: Self-Supervised Pretraining Matters on Imagined Base Set for VLM-based Few-shot LearningHighlight
- Implicit Bias Injection Attacks against Text-to-Image Diffusion ModelsPoster
- Implicit Correspondence Learning for Image-to-Point Cloud RegistrationHighlight
- Improve Representation for Imbalanced Regression through Geometric ConstraintsPoster
- Improved Monocular Depth Prediction Using Distance Transform Over Pre-semantic Contours with Self-supervised Neural NetworksPoster
- Improving Accuracy and Calibration via Differentiated Deep Mutual LearningPoster
- Improving Editability in Image Generation with Layer-wise MemoryPoster
- Improving Gaussian Splatting with Localized Points ManagementHighlight
- Improving Personalized Search with Regularized Low-Rank Parameter UpdatesHighlight
- Improving Semi-Supervised Semantic Segmentation with Sliced-Wasserstein Feature Alignment and UniformityPoster
- Improving Sound Source Localization with Joint Slot Attention on Image and AudioPoster
- Improving Transferable Targeted Attacks with Feature Tuning MixupPoster
- Improving Visual and Downstream Performance of Low-Light Enhancer with Vision Foundation Models CollaborationPoster
- Improving the Training of Data-Efficient GANs via Quality Aware Dynamic Discriminator Rejection SamplingPoster
- Improving the Transferability of Adversarial Attacks on Face Recognition with Diverse Parameters AugmentationPoster
- Imputation-free and Alignment-free: Incomplete Multi-view Clustering Driven by Consensus Semantic LearningPoster
- InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model AlignmentHighlight
- Incomplete Multi-View Multi-label Learning via Disentangled Representation and Label Semantic EmbeddingPoster
- Incomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space ModelPoster
- Incorporating Dense Knowledge Alignment into Unified Multimodal Representation ModelsPoster
- Incremental Object Keypoint LearningPoster
- IndoorGS: Geometric Cues Guided Gaussian Splatting for Indoor Scene ReconstructionPoster
- Inference-Scale Complexity in ANN-SNN Conversion for High-Performance and Low-Power ApplicationsPoster
- Infighting in the Dark: Multi-Label Backdoor Attack in Federated LearningPoster
- InsTaG: Learning Personalized 3D Talking Head from Few-Second VideoPoster
- Insightful Instance Features for 3D Instance SegmentationPoster
- Instance-wise Supervision-level Optimization in Active LearningPoster
- Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive LearningPoster
- Integral Fast Fourier Color ConstancyPoster
- InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance ParsingHighlight
- InteractionMap: Improving Online Vectorized HDMap Construction with InteractionPoster
- Interactive Medical Image Analysis with Concept-based Similarity ReasoningPoster
- Interpretable Generative Models through Post-hoc Concept BottlenecksPoster
- Interpretable Image Classification via Non-parametric Part Prototype LearningPoster
- Inversion Circle Interpolation: Diffusion-based Image Augmentation for Data-scarce ClassificationPoster
- Investigating the Role of Weight Decay in Enhancing Nonconvex SGDPoster
- Invisible Backdoor Attack against Self-supervised LearningPoster
- Iterative Predictor-Critic Code Decoding for Real-World Image DehazingPoster
- JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV SystemsPoster
- JiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World DataPoster
- Joint Optimization of Neural Radiance Fields and Continuous Camera Motion from a Monocular VideoPoster
- Joint Scheduling of Causal Prompts and Tasks for Multi-Task LearningPoster
- Joint Vision-Language Social Bias Removal for CLIPPoster
- Just Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly DetectionHighlight
- KAC: Kolmogorov-Arnold Classifier for Continual LearningHighlight
- KMD: Koopman Multi-modality Decomposition for Generalized Brain Tumor Segmentation under Incomplete ModalitiesPoster
- KVQ: Boosting Video Quality Assessment via Saliency-guided Local PerceptionPoster
- Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic SegmentationPoster
- Keyframe-Guided Creative Video InpaintingPoster
- Knowledge Bridger: Towards Training-Free Missing Modality CompletionPoster
- Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion RecognitionPoster
- L-SWAG: Layer-Sample Wise Activation with Gradients Information for Zero-Shot NAS on Vision TransformersPoster
- LAL: Enhancing 3D Human Motion Prediction with Latency-aware Auxiliary LearningPoster
- LATTE-MV: Learning to Anticipate Table Tennis Hits from Monocular VideosPoster
- LC-Mamba: Local and Continuous Mamba with Shifted Windows for Frame InterpolationPoster
- LEDiff: Latent Exposure Diffusion for HDR GenerationPoster
- LIM: Large Interpolator Model for Dynamic ReconstructionPoster
- LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance FieldsPoster
- LLM-driven Multimodal and Multi-Identity Listening Head GenerationPoster
- LMO: Linear Mamba Operator for MRI ReconstructionPoster
- LOCORE: Image Re-ranking with Long-Context Sequence ModelingPoster
- LOD-GS: Achieving Levels of Detail using Scalable Gaussian SoupPoster
- LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot LearningPoster
- LP-Diff: Towards Improved Restoration of Real-World Degraded License PlateHighlight
- LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic SegmentationPoster
- LSNet: See Large, Focus SmallPoster
- LUCAS: Layered Universal Codec AvatarsPoster
- Label Shift Meets Online Learning: Ensuring Consistent Adaptation with Universal Dynamic RegretHighlight
- Language Guided Concept Bottleneck Models for Interpretable Continual LearningPoster
- Language-Assisted Debiasing and Smoothing for Foundation Model-Based Semi-Supervised LearningPoster
- Language-Guided Audio-Visual Learning for Long-Term Sports AssessmentPoster
- Language-Guided Salient Object RankingPoster
- Large Self-Supervised Models Bridge the Gap in Domain Adaptive Object DetectionPoster
- Latent Drifting in Diffusion Models for Counterfactual Medical Image SynthesisHighlight
CVPR accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.