ICCV 2025 Accepted Papers
The full list of 2,620 papers accepted at ICCV 2025 (IEEE/CVF International Conference on Computer Vision). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
Poster: 2,595
- SDFit: 3D Object Pose and Shape by Fitting a Morphable SDF to a Single ImagePoster
- SDFormer: Vision-based 3D Semantic Scene Completion via SAM-assisted Dual-channel Voxel TransformerPoster
- SDMatte: Grafting Diffusion Models for Interactive MattingPoster
- SEAL: Semantic Aware Image WatermarkingPoster
- SEGS-SLAM: Structure-enhanced 3D Gaussian Splatting SLAM with Appearance EmbeddingPoster
- SEHDR: Single-Exposure HDR Novel View Synthesis via 3D Gaussian BracketingPoster
- SEREP: Semantic Facial Expression Representation for Robust In-the-Wild Capture and RetargetingPoster
- SFUOD: Source-Free Unknown Object DetectionPoster
- SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned DiffusionPoster
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature MatchingPoster
- SHIFT: Smoothing Hallucinations by Information Flow Tuning for Multimodal Large Language ModelsPoster
- SHeaP: Self-Supervised Head Geometry Predictor Learned via 2D GaussiansPoster
- SIC: Similarity-Based Interpretable Image Classification with Neural NetworksPoster
- SIGMAN: Scaling 3D Human Gaussian Generation with Millions of AssetsPoster
- SILO: Solving Inverse Problems with Latent OperatorsPoster
- SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script GenerationPoster
- SITE: towards Spatial Intelligence Thorough EvaluationPoster
- SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video CreationPoster
- SL2A-INR: Single-Layer Learnable Activation for Implicit Neural RepresentationPoster
- SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
- SMGDiff: Soccer Motion Generation using Diffusion Probabilistic ModelsPoster
- SMP-Attack: Boosting the Transferability of Feature Importance-based Adversarial Attack with Semantics-aware Multi-granularity PatchoutPoster
- SMSTracker: Tri-path Score Mask Sigma Fusion for Multi-Modal TrackingPoster
- SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction TuningPoster
- SP2T: Sparse Proxy Attention for Dual-stream Point TransformerPoster
- SPA: Efficient User-Preference Alignment against Uncertainty in Medical Image SegmentationPoster
- SPADE: Spatial-Aware Denoising Network for Open-vocabulary Panoptic Scene Graph Generation with Long- and Local-range Context ReasoningPoster
- SPD: Shallow Backdoor Protecting Deep Backdoor Against Backdoor DetectionPoster
- SRefiner: Soft-Braid Attention for Multi-Agent Trajectory RefinementPoster
- SSVQ: Unleashing the Potential of Vector Quantization with Sign-SplittingPoster
- STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-ResolutionPoster
- STD-GS: Exploring Frame-Event Interaction for SpatioTemporal-Disentangled Gaussian Splatting to Reconstruct High-Dynamic ScenePoster
- STDDNet: Harnessing Mamba for Video Polyp Segmentation via Spatial-aligned Temporal Modeling and Discriminative Dynamic Representation LearningPoster
- STEP-DETR: Advancing DETR-based Semi-Supervised Object Detection with Super Teacher and Pseudo-Label Guided Text QueriesPoster
- STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?Poster
- STIV: Scalable Text and Image Conditioned Video GenerationPoster
- STaR: Seamless Spatial-Temporal Aware Motion Retargeting with Penetration and Consistency ConstraintsPoster
- SU-RGS: Relightable 3D Gaussian Splatting from Sparse Views under Unconstrained IlluminationsPoster
- SUB: Benchmarking CBM Generalization via Synthetic Attribute SubstitutionsPoster
- SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D GenerationPoster
- SVG-Head: Hybrid Surface-Volumetric Gaussians for High-Fidelity Head Reconstruction and Real-Time EditingPoster
- SVIP: Semantically Contextualized Visual Patches for Zero-Shot LearningPoster
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionPoster
- SViM3D: Stable Video Material Diffusion for Single Image 3D GenerationPoster
- Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based AttacksPoster
- Saliency-Aware Quantized Imitation Learning for Efficient Robotic ControlPoster
- Salvaging the Overlooked: Leveraging Class-Aware Contrastive Learning for Multi-Class Anomaly DetectionPoster
- Sat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent DiffusionPoster
- Scalable Dual Fingerprinting for Hierarchical Attribution of Text-to-Image ModelsPoster
- Scalable Image Tokenization with Index Backpropagation QuantizationPoster
- Scalable Ranked Preference Optimization for Text-to-Image GenerationPoster
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention ScalingPoster
- Scaling 3D Compositional Models for Robust Classification and Pose EstimationPoster
- Scaling Action Detection: AdaTAD++ with Transformer-Enhanced Temporal-Spatial AdaptationPoster
- Scaling Inference-Time Search with Vision Value Model for Improved Visual ComprehensionPoster
- Scaling Language-Free Visual Representation LearningPoster
- Scaling Laws for Native Multimodal ModelsPoster
- Scaling Omni-modal Pretraining with Multimodal Context: Advancing Universal Representation Learning Across ModalitiesPoster
- Scaling Transformer-Based Novel View Synthesis with Models Token Disentanglement and Synthetic DataPoster
- Scaling Tumor Segmentation: Best Lessons from Real and Synthetic DataPoster
- Scaling and Taming Adversarial Training with Synthetic DataPoster
- ScanEdit: Hierarchically-Guided Functional 3D Scan EditingPoster
- Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
- Scene Coordinate Reconstruction PriorsPoster
- Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural RectificationPoster
- SceneMI: Motion In-betweening for Modeling Human-Scene InteractionPoster
- ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation AlignmentPoster
- SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language PretrainingPoster
- Scheduling Weight Transitions for Quantization-Aware TrainingPoster
- SciVid: Cross-Domain Evaluation of Video Models in Scientific ApplicationsPoster
- ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided DiffusionPoster
- Scoring, Remember, and Reference: Catching Camouflaged Objects in VideosPoster
- Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware OptimizationPoster
- SeaS: Few-shot Industrial Anomaly Image Generation with Separation and Sharing Fine-tuningPoster
- Seal Your Backdoor with Variational DefensePoster
- Seam360GS: Seamless 360deg Gaussian Splatting from Real-World Omnidirectional ImagesPoster
- Secure On-Device Video OOD Detection Without BackpropagationPoster
- Seeing Through Deepfakes: A Human-Inspired Framework for Multi-Face DetectionPoster
- Seeing and Seeing Through the Glass: Real and Synthetic Data for Multi-Layer Depth EstimationPoster
- Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual GroundingPoster
- Seeing the Unseen: A Semantic Alignment and Context-Aware Prompt Framework for Open-Vocabulary Camouflaged Object SegmentationPoster
- SegAnyPET: Universal Promptable Segmentation from Positron Emission Tomography ImagesPoster
- SegmentDreamer: Towards High-fidelity Text-to-3D Synthesis with Segmented Consistency Trajectory DistillationPoster
- Selective Contrastive Learning for Weakly Supervised Affordance GroundingPoster
- Self-Calibrated Variance-Stabilizing Transformations for Real-World Image DenoisingPoster
- Self-Calibrating Gaussian Splatting for Large Field-of-View ReconstructionPoster
- Self-Ensembling Gaussian Splatting for Few-Shot Novel View SynthesisPoster
- Self-Reinforcing Prototype Evolution with Dual-Knowledge Cooperation for Semi-Supervised Lifelong Person Re-IdentificationPoster
- Self-Supervised Sparse Sensor Fusion for Long Range PerceptionPoster
- Self-supervised Learning of Hybrid Part-aware 3D Representations of 2D Gaussians and SuperquadricsPoster
- SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance LearningPoster
- SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic EmphasisPoster
- Semantic Alignment and Reinforcement for Data-Free Quantization of Vision TransformersPoster
- Semantic Causality-Aware Vision-Based 3D Occupancy PredictionPoster
- Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision TokensPoster
- Semantic Watermarking Reinvented: Enhancing Robustness and Generation Quality with Fourier IntegrityPoster
- Semantic-guided Camera Ray Regression for Visual LocalizationPoster
- Semi-ViM: Bidirectional State Space Model for Mitigating Label Imbalance in Semi-Supervised LearningPoster
- Semi-supervised Concept Bottleneck ModelsPoster
- Semi-supervised Deep Transfer for Regression without Domain AlignmentPoster
- SemiVisBooster: Boosting Semi-Supervised Learning for Fine-Grained Classification through Pseudo-Label Semantic GuidancePoster
- SeqGrowGraph: Learning Lane Topology as a Chain of Graph ExpansionsPoster
- Sequential Gaussian Avatars with Hierarchical Motion ContextPoster
- Sequential keypoint density estimator: an overlooked baseline of skeleton-based video anomaly detectionPoster
- Serialization based Point Cloud OversegmentationPoster
- ShadowHack: Hacking Shadows via Luminance-Color Divide and ConquerPoster
- Shape of Motion: 4D Reconstruction from a Single VideoPoster
- ShortFT: Diffusion Model Alignment via Shortcut-based Fine-TuningPoster
- Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description GenerationPoster
- SiM3D: Single-instance Multiview Multimodal and Multisetup 3D Anomaly Detection BenchmarkPoster
- Sibai: A Few-Shot Meta-Classifier for Poisoning Detection in Federated LearningPoster
- SignRep: Enhancing Self-Supervised Sign RepresentationsPoster
- Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language GeneratorPoster
- Sim-DETR: Unlock DETR for Temporal Sentence GroundingPoster
- SimMLM: A Simple Framework for Multi-modal Learning with Missing ModalityPoster
- Similarity Memory Prior is All You Need for Medical Image SegmentationPoster
- SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language ModelsPoster
- Simulating Dual-Pixel Images From Ray Tracing For Depth EstimationPoster
- Simultaneous Motion And Noise Estimation with Event CamerasPoster
- Single-Scanline Relative Pose Estimation for Rolling Shutter CamerasPoster
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action SegmentationPoster
- SketchSplat: 3D Edge Reconstruction via Differentiable Multi-view Sketch SplattingPoster
- Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token SkippingPoster
- SkySense V2: A Unified Foundation Model for Multi-modal Remote SensingPoster
- SliderSpace: Decomposing the Visual Capabilities of Diffusion ModelsPoster
- SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversionPoster
- Snakes and Ladders: Two Steps Up for VideoMambaPoster
- Social Debiasing for Fair Multi-modal LLMsPoster
- Soft Local Completeness: Rethinking Completeness in XAI
- Soft Separation and Distillation: Toward Global Uniformity in Federated Unsupervised LearningPoster
- Sparfels: Fast Reconstruction from Sparse Unposed ImageryPoster
- Sparse Fine-Tuning of Transformers for Generative TasksPoster
- Sparse-Dense Side-Tuner for efficient Video Temporal GroundingPoster
- SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape ModelingPoster
- SparseLaneSTP: Leveraging Spatio-Temporal Priors with Sparse Transformers for 3D Lane DetectionPoster
- SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMsPoster
- SparseRecon: Neural Implicit Surface Reconstruction from Sparse Views with Feature and Depth ConsistenciesPoster
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
- Sparsity Outperforms Low-Rank Projections in Few-Shot AdaptationPoster
- Spatial Alignment and Temporal Matching Adapter for Video-Radar Remote Physiological MeasurementPoster
- Spatial Preference Rewarding for MLLMs Spatial UnderstandingPoster
- Spatial-Temporal Aware Visuomotor Diffusion Policy LearningPoster
- SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited ObservationsPoster
- SpatialSplat: Efficient Semantic 3D from Sparse Unposed ImagesPoster
- SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera MotionPoster
- Spatially-Varying AutofocusPoster
- Spatio-Spectral Pattern Illumination for Direct and Indirect Separation from a Single Hyperspectral ImagePoster
- SpecGuard: Spectral Projection-based Advanced Invisible WatermarkingPoster
- Spectral Image TokenizerPoster
- Spectral Sensitivity Estimation with an Uncalibrated Diffraction GratingPoster
- SpectralAR: Spectral Autoregressive Visual GenerationPoster
- Spherical Epipolar Rectification for Deep Two-View Absolute Depth EstimationPoster
- SpiLiFormer: Enhancing Spiking Transformers with Lateral InhibitionPoster
- SpikeDiff: Zero-shot High-Quality Video Reconstruction from Chromatic Spike Camera and Sub-millisecond Spike StreamsPoster
- SpikePack: Enhanced Information Flow in Spiking Neural Networks with High Hardware CompatibilityPoster
- SpinMeRound: Consistent Multi-View Identity Generation Using Diffusion ModelsPoster
- SplArt: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian SplattingPoster
- Splat-LOAM: Gaussian Splatting LiDAR Odometry and MappingPoster
- Splat-based 3D Scene Reconstruction with Extreme Motion-blurPoster
- SplatTalk: 3D VQA with Gaussian SplattingPoster
- Split-and-Combine: Enhancing Style Augmentation for Single Domain GeneralizationPoster
- St4RTrack: Simultaneous 4D Reconstruction and Tracking in the WorldPoster
- Stable Diffusion Models are Secretly Good at Visual In-Context LearningPoster
- Stable Score DistillationPoster
- Stable Virtual Camera: Generative View Synthesis with Diffusion ModelsPoster
- Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth DiffusionPoster
- StableCodec: Taming One-Step Diffusion for Extreme Image CompressionPoster
- StableDepth: Scene-Consistent and Scale-Invariant Monocular DepthPoster
- Staining and Locking Computer Vision Models Without RetrainingPoster
- Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View ImagesPoster
- StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided IllusionsPoster
- Stealthy Backdoor Attack in Federated Learning via Adaptive Layer-wise Gradient AlignmentPoster
- SteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric SteeringPoster
- Steering Guidance for Personalized Text-to-Image Diffusion ModelsPoster
- Stepping Out of Similar Semantic Space for Open-Vocabulary SegmentationPoster
- Stereo Any Video: Temporally Consistent Stereo MatchingPoster
- Stochastic Gradient Estimation for Higher-Order Differentiable RenderingPoster
- Stochastic Interpolants for Revealing Stylistic Flows across the History of Art
- StochasticSplats: Stochastic Rasterization for Sorting-Free 3D Gaussian SplattingPoster
- StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic DataPoster
- Straighten Viscous Rectified Flow via Noise OptimizationPoster
- StrandHead: Text to Hair-Disentangled 3D Head Avatars Using Human-Centric PriorsPoster
- StreamDiffusion: A Pipeline-level Solution for Real-Time Interactive GenerationPoster
- StreamGS: Online Generalizable Gaussian Splatting Reconstruction for Unposed Image StreamsPoster
- StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated CognitionPoster
- Streaming VideoLLMs for Real-Time Procedural Video UnderstandingPoster
- Streamlining Image Editing with Layered Diffusion BrushesPoster
- Street Gaussians without 3D Object TrackerPoster
- Stroke2Sketch: Harnessing Stroke Attributes for Training-Free Sketch GenerationPoster
- Stronger, Steadier & Superior: Geometric Consistency in Depth VFM Forges Domain Generalized Semantic SegmentationPoster
- Structure-Guided Diffusion Models for High-Fidelity Portrait Shadow RemovalPoster
- Structure-aware Semantic Discrepancy and Consistency for 3D Medical Image Self-supervised LearningPoster
- Structured Policy Optimization: Enhance Large Vision-Language Model via Self-referenced DialoguePoster
- StyleKeeper: Prevent Content Leakage using Negative Visual Query GuidancePoster
- StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross FusionPoster
- StyleSRN: Scene Text Image Super-Resolution with Text Style EmbeddingPoster
- Stylized-Face: A Million-level Stylized Face Dataset for Face RecognitionPoster
- SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion ModelsPoster
- Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided DiffusionPoster
- SummDiff: Generative Modeling of Video Summarization with DiffusionPoster
- Super Resolved Imaging with Adaptive OpticsPoster
- SuperDec: 3D Scene Decomposition with Superquadrics PrimitivesPoster
- SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image EditingPoster
- SuperEvent: Cross-Modal Learning of Event-based Keypoint Detection for SLAMPoster
- SuperMat: Physically Consistent PBR Material Estimation at Interactive RatesPoster
- Supercharged One-step Text-to-Image Diffusion Models with Negative PromptsPoster
- Supercharging Floorplan Localization with Semantic RaysPoster
- Superpowering Open-Vocabulary Object Detectors for X-ray VisionPoster
- Supervised Exploratory Learning for Long-Tailed Visual RecognitionPoster
- SurfaceSplat: Connecting Surface Reconstruction and Gaussian SplattingPoster
- SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video DiscretizationPoster
- Switch-a-View: View Selection Learned from Unlabeled In-the-wild VideosPoster
- SynAD: Enhancing Real-World End-to-End Autonomous Driving Models through Synthetic Data IntegrationPoster
- SynCity: Training-Free Generation of 3D WorldsPoster
- SynFER: Towards Boosting Facial Expression Recognition with Synthetic DataPoster
- SynTag: Enhancing the Geometric Robustness of Inversion-based Generative Image WatermarkingPoster
- SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction SynthesisPoster
- Synchronization of Multiple VideosPoster
- Synchronizing Task Behavior: Aligning Multiple Tasks during Test-Time TrainingPoster
- Synergistic Prompting for Robust Visual Recognition with Missing ModalitiesPoster
- Synthesizing Near-Boundary OOD Samples for Out-of-Distribution DetectionPoster
- Synthetic Video Enhances Physical Fidelity in Video SynthesisPoster
- T2Bs: Text-to-Character Blendshapes via Video GenerationPoster
- T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive GenerationPoster
- TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language ModelsPoster
- TACO: Taming Diffusion for in-the-wild Video Amodal CompletionPoster
- TAD-E2E: A Large-scale End-to-end Autonomous Driving DatasetPoster
- TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion SensitivityPoster
- TAPNext: Tracking Any Point (TAP) as Next Token PredictionPoster
- TAR3D: Creating High-Quality 3D Assets via Next-Part PredictionPoster
- TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio SynthesisPoster
- TARS: Traffic-Aware Radar Scene Flow EstimationPoster
- TCFG: Truncated Classifier-Free Guidance for Efficient and Scalable Text-to-Image AccelerationPoster
- TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event CamerasPoster
- TF-TI2I: Training-Free Text-and-Image-to-Image Generation via Multi-Modal Implicit-Context Learning In Text-to-Image ModelsPoster
- TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video GenerationPoster
- TITAN-Guide: Taming Inference-Time Alignment for Guided Text-to-Video Diffusion ModelsPoster
- TITAN: Query-Token based Domain Adaptive Adversarial LearningPoster
- TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame InterpolationPoster
- TOGA: Temporally Grounded Open-Ended Video QA with Weak SupervisionPoster
- TOTP: Transferable Online Pedestrian Trajectory Prediction with Temporal-Adaptive Mamba Latent DiffusionPoster
- TPG-INR: Target Prior-Guided Implicit 3D CT Reconstruction for Enhanced Sparse-view ImagingPoster
- TR-PTS: Task-Relevant Parameter and Token Selection for Efficient TuningPoster
- TRACE: Learning 3D Gaussian Physical Dynamics from Multi-view VideosPoster
- TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion ModelsPoster
- TREAD: Token Routing for Efficient Architecture-agnostic Diffusion TrainingPoster
- TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-enhanced Relation-aware Knowledge TransferringPoster
- TRNAS: A Training-Free Robust Neural Architecture SearchPoster
- TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free TuningPoster
- Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationPoster
- Taming Flow Matching with Unbalanced Optimal Transport into Fast PansharpeningPoster
- Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the UnknownPoster
- Target Bias Is All You Need: Zero-Shot Debiasing of Vision-Language Models with Bias CorpusPoster
- Task Vector Quantization for Memory-Efficient Model MergingPoster
- Task-Aware Prompt Gradient Projection for Parameter-Efficient Tuning Federated Class-Incremental LearningPoster
- Task-Decoupled Bezier Surface Constraint for Uneven Low-Light Image EnhancementPoster
- Task-Oriented Human Grasp Synthesis via Context- and Task-Aware DiffusersPoster
- Task-Specific Zero-shot Quantization-Aware Training for Object DetectionPoster
- TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species GenerationPoster
- TeEFusion: Blending Text Embeddings to Distill Classifier-Free GuidancePoster
- TeRA: Rethinking Text-guided Realistic 3D Avatar GenerationPoster
- Teaching AI the Anatomy Behind the Scan: Addressing Anatomical Flaws in Medical Image Segmentation with Learnable PriorPoster
- Teaching VLMs to Localize Specific Objects from In-context ExamplesPoster
- Teeth Reconstruction and Performance Capture Using a Phone CameraPoster
- Teleportraits: Training-Free People Insertion into Any ScenePoster
- Temperature in Cosine-based Softmax LossPoster
- Temporal Overlapping Prediction: A Self-supervised Pre-training Method for LiDAR Moving Object SegmentationPoster
- Temporal Rate Reduction Clustering for Human Motion SegmentationPoster
- Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object TrackingPoster
- Temporal-aware Query Routing for Real-time Video Instance SegmentationPoster
- Tensor-aggregated LoRA in Federated Fine-tuningPoster
- TerraMind: Large-Scale Generative Multimodality for Earth ObservationPoster
- Test-Time Prompt Tuning for Zero-Shot Depth CompletionPoster
- Test-Time Retrieval-Augmented Adaptation for Vision-Language ModelsPoster
- Test-time Adaptation for Foundation Medical Segmentation Model Without Parametric UpdatesPoster
- Text Embedding Knows How to Quantize Text-Guided Diffusion ModelsPoster
- Text-IRSTD: Leveraging Semantic Text to Promote Infrared Small Target Detection in Complex ScenesPoster
- Text-guided Visual Prompt DINO for Generic SegmentationPoster
- Text-to-Any-Skeleton Motion Generation Without RetargetingPoster
- Text2Outfit: Controllable Outfit Generation with Multimodal Language ModelsPoster
- Text2VDM: Text to Vector Displacement Maps for Expressive and Interactive 3D SculptingPoster
- TextMaster: A Unified Framework for Realistic Text Editing via Glyph-Style Dual-ControlPoster
- TextSSR: Diffusion-based Data Synthesis for Scene Text RecognitionPoster
- Textured 3D Regenerative Morphing with 3D Diffusion PriorPoster
- The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video GenerationPoster
- The Curse of Conditions: Analyzing and Improving Optimal Transport for Conditional Flow-Based GenerationPoster
- The Devil is in the Spurious Correlations: Boosting Moment Retrieval with Dynamic LearningPoster
- The Inter-Intra Modal Measure: A Predictive Lens on Fine-Tuning Outcomes in Vision-Language ModelsPoster
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single TransformerPoster
- The Silent Assistant: NoiseQuery as Implicit Guidance for Goal-Driven Image GenerationPoster
- The Source Image is the Best Attention for Infrared and Visible Image FusionPoster
- Thermal Polarimetric Multi-view StereoPoster
- Think Twice: Test-Time Reasoning for Robust CLIP Zero-Shot ClassificationPoster
- TikZero: Zero-Shot Text-Guided Graphics Program SynthesisPoster
- Tile-wise vs. Image-wise: Random-Tile Loss and Training Paradigm for Gaussian SplattingPoster
- Tiling artifacts and trade-offs of feature normalization in the segmentation of large biological imagesPoster
- Time-Aware Auto White Balance in Mobile PhotographyPoster
- TimeBooth: Disentangled Facial Invariant Representation for Diverse and Personalized Face AgingPoster
- TimeExpert: An Expert-Guided Video LLM for Video Temporal GroundingPoster
- TimeFormer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust ReconstructionPoster
- Timestep-Aware Diffusion Model for Extreme Image RescalingPoster
- TinyViM: Frequency Decoupling for Tiny Hybrid Vision MambaPoster
- To Label or Not to Label: PALM - A Predictive Model for Evaluating Sample Efficiency in Active Learning ModelsPoster
- ToF-Splatting: Dense SLAM using Sparse Time-of-Flight Depth and Multi-Frame IntegrationPoster
- Token Activation Map to Visually Explain Multimodal LLMsPoster
- Token-Efficient VLM: High-Resolution Image Understanding via Dynamic Region ProposalPoster
- TokensGen: Harnessing Condensed Tokens for Long Video GenerationPoster
- ToolVQA: A Dataset for Multi-step Reasoning VQA with External ToolsPoster
- Top2Pano: Learning to Generate Indoor Panoramas from Top-Down ViewPoster
- TopicGeo: An Efficient Unified Framework for GeolocationPoster
- TorchAdapt: Towards Light-Agnostic Real-Time Visual PerceptionPoster
- Toward Better Out-painting: Improving the Image Composition with Initialization Policy ModelPoster
- Toward Fair and Accurate Cross-Domain Medical Image Segmentation: A VLM-Driven Active Domain Adaptation ParadigmPoster
- Toward Long-Tailed Online Anomaly Detection through Class-Agnostic ConceptsPoster
- Toward Material-Agnostic System Identification from VideosPoster
- Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on EdgePoster
- Towards Adversarial Robustness via Debiased High-Confidence Logit AlignmentPoster
- Towards Annotation-Free Evaluation: KPAScore for Human Keypoint DetectionPoster
- Towards Comprehensive Lecture Slides Understanding: Large-scale Dataset and Effective MethodPoster
- Towards Cross-modal Backward-compatible Representation Learning for Vision-Language ModelsPoster
- Towards Effective Foundation Model Adaptation for Extreme Cross-Domain Few-Shot LearningPoster
- Towards Efficient General Feature Prediction in Masked Skeleton ModelingPoster
- Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human AvatarsPoster
- Towards Fine-grained Interactive Segmentation in Images and VideosPoster
- Towards Foundational Models for Single-Chip RadarPoster
- Towards Higher Effective Rank in Parameter-Efficient Fine-tuning using Khatri-Rao ProductPoster
- Towards Human-like Virtual Beings: Simulating Human Behavior in 3D ScenesPoster
- Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion SynthesisPoster
- Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and MemoryPoster
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled ViewsPoster
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual SegmentationPoster
- Towards Open-World Generation of Stereo Images and Unsupervised MatchingPoster
- Towards Performance Consistency in Multi-Level Model CollaborationPoster
- Towards Real Unsupervised Anomaly Detection Via Confident Meta-LearningPoster
- Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based PurificationPoster
- Towards Robustness of Person Search against CorruptionsPoster
- Towards Safer and Understandable Driver Intention PredictionPoster
- Towards Scalable Spatial Intelligence via 2D-to-3D Data LiftingPoster
- Towards Stabilized and Efficient Diffusion Transformers through Long-Skip-Connections with Spectral ConstraintsPoster
- Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
- Towards Visual Localization Interoperability: Cross-Feature for Collaborative Visual Localization and MappingPoster
- Towards a 3D Transfer-based Black-box Attack via Critical Feature GuidancePoster
- Towards a Unified Copernicus Foundation Model for Earth VisionPoster
- Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant RepresentationPoster
- Towards a Universal Image Degradation Model via Content-Degradation DisentanglementPoster
- Trace3D: Consistent Segmentation Lifting via Gaussian Instance TracingPoster
- Tracing Copied Pixels and Regularizing Patch Affinity in Copy DetectionPoster
- TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud TrackingPoster
- TrackVerse: A Large-Scale Object-Centric Video Dataset for Image-Level Representation Learning
- Tracking Tiny Drones against Clutter: Large-Scale Infrared Benchmark with Motion-Centric Adaptive AlgorithmPoster
- Trade-offs in Image Generation: How Do Different Dimensions Interact?Poster
- TrafficLoc: Localizing Traffic Surveillance Cameras in 3D ScenesPoster
- Training-Free Class Purification for Open-Vocabulary Semantic SegmentationPoster
- Training-Free Industrial Defect Generation with Diffusion ModelsPoster
- Training-Free Personalization via Retrieval and Reasoning on FingerprintsPoster
- Training-Free Text-Guided Image Editing with Visual Autoregressive ModelPoster
- Training-free Generation of Temporally Consistent Rewards from VLMsPoster
- Training-free Geometric Image Editing on Diffusion ModelsPoster
- Training-free and Adaptive Sparse Attention for Efficient Long Video GenerationPoster
- TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion ModelsPoster
- Trans-Adapter: A Plug-and-Play Framework for Transparent Image InpaintingPoster
- Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image ModelsPoster
- Transformer-based Tooth Alignment Prediction with Occlusion and Collision ConstraintsPoster
- TransiT: Transient Transformer for Non-line-of-sight VideographyPoster
- Translation of Text Embedding via Delta Vector to Suppress Strongly Entangled Content in Text-to-Image Diffusion ModelsPoster
- Transparent Vision: A Theory of Hierarchical Invariant RepresentationsPoster
- Tree Skeletonization from 3D Point Clouds by Denoising DiffusionPoster
- Tree-NeRV: Efficient Non-Uniform Sampling for Neural Video Representation via Tree-Structured Feature GridsPoster
- TriDi: Trilateral Diffusion of 3D Humans, Objects, and InteractionsPoster
- Triad: Empowering LMM-based Anomaly Detection with Expert-guided Region-of-Interest Tokenizer and Manufacturing ProcessPoster
- Trial-Oriented Visual RearrangementPoster
- Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action RecognitionPoster
- Trust but Verify: Programmatic VLM Evaluation in the WildPoster
- TrustMark: Robust Watermarking and Watermark Removal for Arbitrary Resolution ImagesPoster
- TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-InterventionPoster
- TryOn-Refiner: Conditional Rectified-flow-based TryOn Refiner for More Accurate Detail ReconstructionPoster
- Tune-Your-Style: Intensity-tunable 3D Style Transfer with Gaussian SplattingPoster
- Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video SynthesisPoster
- TurboReg: TurboClique for Robust and Efficient Point Cloud RegistrationPoster
- TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and PredictionPoster
- TurboVSR: Fantastic Video Upscalers and Where to Find ThemPoster
- Two Losses, One Goal: Balancing Conflict Gradients for Semi-supervised Semantic SegmentationPoster
- U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and RegistrationPoster
- UAVScenes: A Multi-Modal Dataset for UAVsPoster
- UDC-VIT: A Real-World Video Dataset for Under-Display CamerasPoster
- UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital AgentsPoster
- UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility ConstraintPoster
- UIPro: Unleashing Superior Interaction Capability For GUI AgentsPoster
- UKBOB: One Billion MRI Labeled Masks for Generalizable 3D Medical Image SegmentationPoster
- ULTHO: Ultra-Lightweight yet Efficient Hyperparameter Optimization in Deep Reinforcement LearningPoster
- UMDATrack: Unified Multi-Domain Adaptive Tracking Under Adverse Weather ConditionsPoster
- UNIS: A Unified Framework for Achieving Unbiased Neural Implicit Surfaces in Volume RenderingPoster
- UPP: Unified Point-Level Prompting for Robust Point Cloud AnalysisPoster
- USP: Unified Self-Supervised Pretraining for Image Generation and UnderstandingPoster
- Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency AdapterPoster
- Ultra-Precision 6DoF Pose Estimation Using 2-D Interpolated Discrete Fourier TransformPoster
- UnMix-NeRF: Spectral Unmixing Meets Neural Radiance FieldsPoster
- UnZipLoRA: Separating Content and Style from a Single ImagePoster
- Unbiased Missing-modality Multimodal LearningPoster
- Unbiased Region-Language Alignment for Open-Vocabulary Dense PredictionPoster
- Uncalibrated Structure from Motion on a SpherePoster
- Uncertainty-Aware Diffusion-Guided Refinement of 3D ScenesPoster
- Uncertainty-Aware Gradient Stabilization for Small Object DetectionPoster
- Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language ModelsPoster
- Uncover Treasures in DCT: Advancing JPEG Quality Enhancement by Exploiting Latent CorrelationsPoster
- Understanding Co-speech Gestures in-the-wildPoster
- Understanding Flatness in Generative Models: Its Role and BenefitsPoster
- Understanding Museum Exhibits using Vision-Language ReasoningPoster
- Understanding Personal Concept in Open-Vocabulary Semantic SegmentationPoster
- Underwater Visual SLAM with Depth Uncertainty and Medium ModelingPoster
- Unfolding-Associative Encoder-Decoder Network with Progressive Alignment for PansharpeningPoster
- UniCombine: Unified Multi-Conditional Combination with Diffusion TransformerPoster
- UniConvNet: Expanding Effective Receptive Field while Maintaining Asymptotically Gaussian Distribution for ConvNets of Any ScalePoster
- UniDxMD: Towards Unified Representation for Cross-Modal Unsupervised Domain Adaptation in 3D Semantic SegmentationPoster
- UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and GenerationPoster
- UniFuse: A Unified All-in-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and MisalignmentsPoster
- UniGS: Modeling Unitary 3D Gaussians for Novel View Synthesis from Sparse-view ImagesPoster
- UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text SynthesisPoster
- UniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous DrivingPoster
- UniOcc: A Unified Benchmark for Occupancy Forecasting and Prediction in Autonomous DrivingPoster
- UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character ControlPoster
- UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image PersonalizationPoster
- UniRes: Universal Image Restoration for Complex DegradationsPoster
- UniVG: A Generalist Diffusion Model for Unified Image Generation and EditingPoster
- UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field ReconstructionPoster
- Unified Adversarial Augmentation for Improving Palmprint RecognitionPoster
- Unified Category-Level Object Detection and Pose Estimation from RGB Images using 3D PrototypesPoster
- Unified Multi-Agent Trajectory Modeling with Masked Trajectory DiffusionPoster
- Unified Multimodal Understanding via Byte-Pair Visual EncodingPoster
- Unified Open-World Segmentation with Multi-Modal PromptsPoster
- Unified Video Generation via Next-Set Prediction in Continuous DomainPoster
- UniversalBooth: Model-Agnostic Personalized Text-to-Image GenerationPoster
- Unknown Text Learning for CLIP-based Few-Shot Open-set RecognitionPoster
- Unlearning the Noisy Correspondence Makes CLIP More RobustPoster
- Unleashing High-Quality Image Generation in Diffusion Sampling Using Second-Order Levenberg-Marquardt-LangevinPoster
- Unleashing Vecset Diffusion Model for Fast Shape GenerationPoster
- Unleashing the Temporal Potential of Stereo Event Cameras for Continuous-Time 3D Object DetectionPoster
- Unlocking Constraints: Source-Free Occlusion-Aware Seamless SegmentationPoster
- Unlocking the Potential of Diffusion Priors in Blind Face RestorationPoster
- Unraveling the Effects of Synthetic Data on End-to-End Autonomous DrivingPoster
- Unraveling the Smoothness Properties of Diffusion Models: A Gaussian Mixture PerspectivePoster
- UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AIPoster
- Unsupervised Histopathological Image Semantic Segmentation with Overlapping Patches Consistency ConstraintPoster
- Unsupervised Identification of Protein Compositions and Conformations via Implicit Content-Transformation DisentanglementPoster
- Unsupervised Imaging Inverse Problems with Diffusion Distribution MatchingPoster
- Unsupervised Joint Learning of Optical Flow and Intensity with Event CamerasPoster
- Unsupervised Part Discovery via Descriptor-Based Masked Image Restoration with Optimized ConstraintsPoster
- Unsupervised RGB-D Point Cloud Registration for Scenes with Low Overlap and Photometric InconsistencyPoster
- Unsupervised Visible-Infrared Person Re-identification under Unpaired SettingsPoster
- Unsupervised Visual Chain-of-Thought Reasoning via Preference OptimizationPoster
- Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURAPoster
- UrbanLLaVA: A Multi-modal Large Language Model for Urban IntelligencePoster
- V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion ModelsPoster
- V2M4: 4D Mesh Animation Reconstruction from a Single Monocular VideoPoster
- V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position EncodingPoster
- V2XPnP: Vehicle-to-Everything Spatio-Temporal Fusion for Multi-Agent Perception and PredictionPoster
- V2XScenes: A Multiple Challenging Traffic Conditions Dataset for Large-Range Vehicle-Infrastructure Collaborative PerceptionPoster
- VA-MoE: Variables-Adaptive Mixture of Experts for Incremental Weather ForecastingPoster
- VACE: All-in-One Video Creation and EditingPoster
- VAFlow: Video-to-Audio Generation with Cross-Modality Flow MatchingPoster
- VAGUE: Visual Contexts Clarify Ambiguous ExpressionsPoster
- VALLR: Visual ASR Language Model for Lip ReadingPoster
- VCA: Video Curious Agent for Long Video UnderstandingPoster
- VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided OptimizationPoster
- VGGSounder: Audio-Visual Evaluations for Foundation ModelsPoster
- VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual GroundingPoster
- VIGFace: Virtual Identity Generation for Privacy-Free Face Recognition DatasetPoster
- VIPerson: Flexibly Generating Virtual Identity for Person Re-IdentificationPoster
- VISION-XL: High Definition Video Inverse Problem Solver using Latent Image Diffusion ModelsPoster
- VISO: Accelerating In-orbit Object Detection with Language-Guided Mask Learning and Sparse InferencePoster
- VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information FlowPoster
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksPoster
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous DrivingPoster
- VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
- VLM4D: Towards Spatiotemporal Awareness in Vision Language ModelsPoster
- VLR-Driver: Large Vision-Language-Reasoning Models for Embodied Autonomous DrivingPoster
- VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsPoster
- VMBench: A Benchmark for Perception-Aligned Video Motion GenerationPoster
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
- VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation under real OcclusionsPoster
- VPO: Aligning Text-to-Video Generation Models with Prompt OptimizationPoster
- VPR-Cloak: A First Look at Privacy Cloak Against Visual Place RecognitionPoster
- VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative VideosPoster
- VRM: Knowledge Distillation via Virtual Relation MatchingPoster
- VSC: Visual Search Compositional Text-to-Image Diffusion ModelPoster
- VSP: Diagnosing the Dual Challenges of Perception and Reasoning in Spatial Planning Tasks for MLLMsPoster
- VSRM: A Robust Mamba-Based Framework for Video Super-ResolutionPoster
- VSSD: Vision Mamba with Non-Causal State Space DualityPoster
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and ReasoningPoster
- Vamba: Understanding Hour-Long Videos with Hybrid Mamba-TransformersPoster
- Variance-Based Pruning for Accelerating and Compressing Trained NetworksPoster
- Vector Contrastive Learning For Pixel-Wise Pretraining In Medical VisionPoster
- VehicleMAE: View-asymmetry Mutual Learning for Vehicle Re-identification Pre-training via Masked AutoEncodersPoster
- Verbalized Representation Learning for Interpretable Few-Shot GeneralizationPoster
- Versatile Transition Generation with Image-to-Video DiffusionPoster
- VertexRegen: Mesh Generation with Continuous Level of DetailPoster
- ViCTr: Vital Consistency Transfer for Pathology Aware Image SynthesisPoster
- ViLLa: Video Reasoning Segmentation with Large Language ModelPoster
- ViLU: Learning Vision-Language Uncertainties for Failure PredictionPoster
- ViM-VQ: Efficient Post-Training Vector Quantization for Visual MambaPoster
- ViSpeak: Visual Instruction Feedback in Streaming VideosPoster
- ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision TransformersPoster
- ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision ModelsPoster
- ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting HeadsPoster
- Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the WildPoster
- Video Color Grading via Look-Up Table GenerationPoster
- Video Individual Counting for Moving DronesPoster
- Video Motion GraphsPoster
- Video-T1: Test-time Scaling for Video GenerationPoster
- Video2BEV: Transforming Drone Videos to BEVs for Video-based Geo-localizationPoster
- VideoAds for Fast-Paced Video Understanding
- VideoAuteur: Towards Long Narrative Video GenerationPoster
- VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory BridgesPoster
- VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy OptimizationPoster
- VideoOrion: Tokenizing Object Dynamics in VideosPoster
- VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint ModelingPoster
- VideoSetDiff: Identifying and Reasoning Similarities and Differences in Similar VideosPoster
- VideoVAE+: Large Motion Video Autoencoding with Cross-modal Video VAEPoster
- ViewSRD: 3D Visual Grounding via Structured Multi-View DecompositionPoster
- VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible RegionsPoster
- VisNumBench: Evaluating Number Sense of Multimodal Large Language ModelsPoster
- VisRL: Intention-Driven Visual Perception via Reinforced ReasoningPoster
- Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph GenerationPoster
- Vision-Language Models Can't See the ObviousPoster
- Vision-Language Neural Graph Featurization for Extracting Retinal LesionsPoster
- VisionMath: Vision-Form Mathematical Problem-SolvingPoster
- VistaDream: Sampling multiview consistent images for single-view scene reconstructionPoster
- Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of ImagesPoster
- Visual Intention Grounding for Egocentric AssistantsPoster
- Visual Interestingness Decoded: How GPT-4o Mirrors Human InterestsPoster
- Visual Modality Prompt for Adapting Vision-Language Object DetectorsPoster
- Visual Relation Diffusion for Human-Object Interaction DetectionPoster
- Visual Surface Wave Elastography: Revealing Subsurface Physical Properties via Visible Surface WavesPoster
- Visual Test-time Scaling for GUI Agent GroundingPoster
- Visual Textualization for Image Prompted Object DetectionPoster
- Visual-Oriented Fine-Grained Knowledge Editing for MultiModal Large Language ModelsPoster
- VisualCloze: A Universal Image Generation Framework via Visual In-Context LearningPoster
- Vivid4D: Improving 4D Reconstruction from Monocular Video by Video InpaintingPoster
- VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language ModelsPoster
- VoluMe - Authentic 3D Video Calls from Live Gaussian Splat PredictionPoster
- VolumetricSMPL: A Neural Volumetric Body Model for Efficient Interactions, Contacts, and CollisionsPoster
- VoteSplat: Hough Voting Gaussian Splatting for 3D Scene UnderstandingPoster
- VoxelKP: A Voxel-based Network Architecture for Human Keypoint Estimation in LiDAR DataPoster
- Voyaging into Perpetual Dynamic Scenes from a Single ViewPoster
- Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video DetectionPoster
- WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single ImagePoster
- WINS: Winograd Structured Pruning for Fast Winograd ConvolutionPoster
- WIPES: Wavelet-based Visual PrimitivesPoster
- WIR3D: Visually-Informed and Geometry-Aware 3D Shape AbstractionPoster
- WSI-LLaVA: A Multimodal Large Language Model for Whole Slide ImagePoster
- WalkVLM: Aid Visually Impaired People Walking by Vision Language ModelPoster
- WarpHE4D: Dense 4D Head Map toward Full Head ReconstructionPoster
- Wasserstein Style Distribution Analysis and Transform for Stylized Image GenerationPoster
- Wave-MambaAD: Wavelet-driven State Space Model for Multi-class Unsupervised Anomaly DetectionPoster
- WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object DetectionPoster
- Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon TasksPoster
- Weakly Supervised Visible-Infrared Person Re-Identification via Heterogeneous Expert Collaborative Consistency LearningPoster
- Weakly-Supervised Learning of Dense Functional CorrespondencesPoster
- WeaveSeg: Iterative Contrast-weaving and Spectral Feature-refining for Nuclei Instance SegmentationPoster
- Web Artifact Attacks Disrupt Vision Language ModelsPoster
- What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation LearningPoster
- What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language ModelsPoster
- What If: Understanding Motion Through Sparse InteractionsPoster
- What Makes for Text to 360-degree Panorama Generation with Stable Diffusion?Poster
- What You Have is What You Track: Adaptive and Robust Multimodal TrackingPoster
- What to Distill? Fast Knowledge Distillation with Adaptive SamplingPoster
- What we need is explicit controllability: Training 3D gaze estimator using only facial imagesPoster
- What's Making That Sound Right Now? Video-centric Audio-Visual LocalizationPoster
- What's in a Latent? Leveraging Diffusion Latent Space for Domain GeneralizationPoster
- When Anchors Meet Cold Diffusion: A Multi-Stage Approach to Lane DetectionPoster
- When Confidence Fails: Revisiting Pseudo-Label Selection in Semi-supervised Semantic SegmentationPoster
- When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token PruningPoster
- When Pixel Difference Patterns Meet ViT: PiDiViT for Few-Shot Object DetectionPoster
- When Schrodinger Bridge Meets Real-World Image Dehazing with Unpaired TrainingPoster
- When and Where do Data Poisons Attack Textual Inversion?Poster
- Where am I? Cross-View Geo-localization with Natural Language DescriptionsPoster
- Where, What, Why: Towards Explainable Driver Attention PredictionPoster
- Who Controls the Authorization? Invertible Networks for Copyright Protection in Text-to-Image SynthesisPoster
- Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking HeadsPoster
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of ContextPoster
- Wide2Long: Learning Lens Compression and Perspective Adjustment for Wide-Angle to Telephoto TranslationPoster
- WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article GenerationPoster
- WildSAT: Learning Satellite Image Representations from Wildlife ObservationsPoster
- WildSeg3D: Segment Any 3D Objects in the Wild from 2D ImagesPoster
- WonderPlay: Dynamic 3D Scene Generation from a Single Image and ActionsPoster
- WonderTurbo: Generating Interactive 3D World in 0.72 SecondsPoster
- World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World ModelPoster
- WorldScore: A Unified Evaluation Benchmark for World GenerationPoster
- X-Capture: An Open-Source Portable Device for Multi-Sensory LearningPoster
- X-Dancer: Expressive Music to Human Dance Video GenerationPoster
- X-Fusion: Introducing New Modality to Frozen Large Language ModelsPoster
- X-Prompt: Generalizable Auto-Regressive Visual Learning with In-Context PromptingPoster
- X2-Gaussian: 4D Radiative Gaussian Splatting for Continuous-time Tomographic ReconstructionPoster
- X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention DistillationPoster
- XTrack: Multimodal Training Boosts RGB-X Video Object TrackersPoster
- YOLO-Count: Differentiable Object Counting for Text-to-Image GenerationPoster
- YOLOE: Real-Time Seeing AnythingPoster
- You Are Your Own Best Teacher: Achieving Centralized-level Performance in Federated Learning under Heterogeneous and Long-tailed DataPoster
- You Share Beliefs, I Adapt: Progressive Heterogeneous Collaborative PerceptionPoster
- You Think, You ACT: The New Task of Arbitrary Text to Motion GenerationPoster
- Your Text Encoder Can Be An Object-Level Watermarking ControllerPoster
- ZFusion: Efficient Deep Compositional Zero-shot Learning for Blind Image Super-Resolution with Generative Diffusion PriorPoster
- ZIM: Zero-Shot Image Matting for AnythingPoster
- ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned ModelsPoster
- Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech RepresentationsPoster
- Zero-Shot Composed Image Retrieval via Dual-Stream Instruction-Aware DistillationPoster
- Zero-Shot Compositional Video Learning with Coding Rate ReductionPoster
- Zero-Shot Depth Aware Image Editing with Diffusion ModelsPoster
- Zero-Shot Vision Encoder Grafting via LLM SurrogatesPoster
- Zero-shot Inexact CAD Model Alignment from a Single ImagePoster
- ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language ModelsPoster
- Zeroth-Order Fine-Tuning of LLMs in Random SubspacesPoster
- ZipVL: Accelerating Vision-Language Models through Dynamic Token SparsityPoster
- egoPPG: Heart Rate Estimation from Eye-Tracking Cameras in Egocentric Systems to Benefit Downstream Vision TasksPoster
- iManip: Skill-Incremental Learning for Robotic ManipulationPoster
- kh: Symmetry Understanding of 3D Shapes via Chirality DisentanglementPoster
- mmCooper: A Multi-agent Multi-stage Communication-efficient and Collaboration-robust Cooperative Perception FrameworkPoster
- monoVLN: Bridging the Observation Gap between Monocular and Panoramic Vision and Language NavigationPoster
- p-AVAS: Can Physics-Integrated Audio-Visual Modeling Boost Neural Acoustic Synthesis?Poster
- p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio DecayPoster
ICCV accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.