CVPR 2026 Accepted Papers
The full list of 4,068 papers accepted at CVPR 2026 (IEEE/CVF Conference on Computer Vision and Pattern Recognition). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
accepted: 4,068
- SAM2Text: Towards Prompt-Free and Multi-Resolution Video Scene Text Segmentationaccepted
- SAME: Sparse and Anchored Model Editing for Heterogeneous Incremental Learning under Limited Dataaccepted
- SAMIX: Reinforcing SAM2 with Semantic Adapter and Reference Selecting Policy for Mix-Supervised Segmentationaccepted
- SAMTok: Representing Any Mask with Two Wordsaccepted
- SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment Anythingaccepted
- SANER: Switchable Adapter with Non-parametric Enhanced Routing for Person De-Reidentificationaccepted
- SAQN: Semantic-based Adaptive Query Network for 3D Referring Expression Segmentationaccepted
- SAR2Net: Learning Spatially Anchored Representations for Retrieval-Guided Cross-Stain Alignmentaccepted
- SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video Groundingaccepted
- SARMAE: Masked Autoencoder for SAR Representation Learningaccepted
- SASNet: Spatially-Adaptive Sinusoidal Networks for INRsaccepted
- SAT-RRG: LLM-Guided Self-Adaptive Training for Radiology Report Generation with Token-Level Push-Pull Optimizationaccepted
- SATTC: Structure-Aware Label-Free Test-Time Calibration for Cross-Subject EEG-to-Image Retrievalaccepted
- SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusionaccepted
- SAVE: Speech-Aware Video Representation Learning for Video-Text Retrievalaccepted
- SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observationaccepted
- SCE-Depth: A Spherical Compound Eye Framework for Wide FOV Depth Estimationaccepted
- SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate Embeddingsaccepted
- SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal Modelsaccepted
- SCoRe: Salience-Coverage Reduction for Vision Token Pruning in Vision-Language Modelsaccepted
- SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentationaccepted
- SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detectionaccepted
- SDGS: Spatial Difference Guided Gaussian Splatting for Simultaneous Localization and 3D Reconstructionaccepted
- SDTrack: A Baseline for Event-based Tracking via Spiking Neural Networksaccepted
- SDUIE: Semi-Supervised Diffusion for Underwater Image Enhancement with Quant-Text Dual Controlaccepted
- SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose Estimationaccepted
- SEA-Flow3D: Simplified, Efficient, and Accurate Scene Flow via Spatial Vector Sampling and Multi-scale Refinementaccepted
- SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asiaaccepted
- SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answeringaccepted
- SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decodingaccepted
- SEATrack: Simple, Efficient, and Adaptive Multimodal Trackeraccepted
- SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learningaccepted
- SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learningaccepted
- SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detectionaccepted
- SG-LoRA: Semantic-guided LoRA Parameters Generationaccepted
- SGAD-SLAM: Splatting Gaussians at Adjusted Depth for Better Radiance Fields in RGBD SLAMaccepted
- SGDE: Self-supervised Geometry Degradation Estimation Framework for Coded Aperture Compressive Spectral Imagingaccepted
- SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Drivingaccepted
- SGI: Structured 2D Gaussians for Efficient and Compact Large Image Representationaccepted
- SGS-Intrinsic: Semantic-Invariant Gaussian Splatting for Sparse-View Indoor Inverse Renderingaccepted
- SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signalsaccepted
- SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentationaccepted
- SHARP: Short-Window Streaming for Accurate and Robust Prediction in Motion Forecastingaccepted
- SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wildaccepted
- SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Trainingaccepted
- SIF: Semantically In-Distribution Fingerprints for Large Vision-Language Modelsaccepted
- SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Imagesaccepted
- SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokensaccepted
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language Modelsaccepted
- SIMPLEPOSTER: A SIMPLE BASELINE FOR PRODUCT POSTER GENERATIONaccepted
- SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarkingaccepted
- SIR: Structured Image Representations for Explainable Robot Learningaccepted
- SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuationaccepted
- SLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic Scenesaccepted
- SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generationaccepted
- SMAP: Semantic Route Planning with Map-Grounded Multimodal Alignmentaccepted
- SMRABooth: Subject and Motion Representation Alignment for Customized Video Generationaccepted
- SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognitionaccepted
- SMVRT: Implicit Human 3D Modeling Using Sparse Multi-View Volumetric Reconstruction with Transformer Fusionaccepted
- SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth Estimationaccepted
- SO-Bench: A Structural Output Evaluation of Multimodal LLMaccepted
- SODA: Sensitivity-Oriented Dynamic Acceleration for Diffusion Transformeraccepted
- SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation Modelsaccepted
- SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt Contextsaccepted
- SPAN: Spatial-Projection Alignment for Monocular 3D Object Detectionaccepted
- SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentationaccepted
- SPARK: Sim-ready Part-level Articulated Reconstruction with VLM Knowledgeaccepted
- SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMsaccepted
- SPDMark: Selective Parameter Displacement for Robust Video Watermarkingaccepted
- SPE-MVS: Spatial Position Encoding Enhanced Multi-View Stereo with Monocular Depth Priorsaccepted
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understandingaccepted
- SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentationaccepted
- SPOT: Spatiotemporal Prompt Optimization for Motion-Stabilized MLLM-Guided Video Segmentationaccepted
- SPREAD: Spatial-Physical REasoning via geometry Aware Diffusionaccepted
- SR3R: Rethinking Super-Resolution 3D Reconstruction With Feed-Forward Gaussian Splattingaccepted
- SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion Trainingaccepted
- SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Namesaccepted
- SRGCD: Stability-Driven Region Growth Framework for 3D Change Detectionaccepted
- SRPO: Self-Referential Policy Optimization for Vision-Language-Action Modelsaccepted
- SSM-Aware Token-Efficient VMamba via Adaptive Patch Pruning and Merging for Person Re-Identificationaccepted
- ST4R-Splat: Spatio-Temporal Referring Segmentation in 4D Gaussian Splattingaccepted
- STAC: Plug-and-Play Spatio-Temporal Aware Cache Compression for Streaming 3D Reconstructionaccepted
- STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrativeaccepted
- STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMsaccepted
- STAR: Test-Time Adaptation Can Enhance Universal Prompt Learning for Vision-Language Modelsaccepted
- STARFlow-V: End-to-End Video Generative Modeling with Autoregressive Normalizing Flowsaccepted
- STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars Reconstructionaccepted
- STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolutionaccepted
- STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecastingaccepted
- STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregationaccepted
- STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detectionaccepted
- STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrievalaccepted
- SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detectionaccepted
- SURF: Signature-Retained Fast Video Generationaccepted
- SV-GS: Sparse View 4D Reconstruction with Skeleton-Driven Gaussian Splattingaccepted
- SVAgent: Storyline-guided Long Video Understanding via Cross-Modal Multi-Agent Collaborationaccepted
- SVBench: Evaluation of Video Generation Models on Social Reasoningaccepted
- SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Modelsaccepted
- SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attributionaccepted
- SaPaVe: Towards Active Perception and Manipulation in Vision-Language Action Models for Roboticsaccepted
- SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse Worldaccepted
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimizationaccepted
- SafeLogo: Turning Your Logos into Jailbreak Shields via Micro-Regional Adversarial Trainingaccepted
- SafeRoPE: Risk-specific Head-wise Embedding Rotation for Safe Generation in Rectified Flow Transformersaccepted
- Saliency-Driven Token Merging for Vision Transformersaccepted
- Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learningaccepted
- Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Rewardaccepted
- Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucinationaccepted
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMsaccepted
- Same or Not? Enhancing Visual Perception in Vision-Language Modelsaccepted
- Sampling-Aware Quantization for Diffusion Modelsaccepted
- Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Editsaccepted
- Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstructionaccepted
- Scalable Feature Matching via State Space Modeling and Sparse Correlationaccepted
- Scalable Multi-View Subspace Clustering with Tensorized Anchor Guidanceaccepted
- Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Modelsaccepted
- Scalable Trajectory Generation for Whole-Body Mobile Manipulationaccepted
- Scale Space Diffusionaccepted
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMsaccepted
- Scaling Dense Event-Stream Pretraining from Visual Foundation Modelsaccepted
- Scaling Instruction-Based Video Editing with a High-Quality Synthetic Datasetaccepted
- Scaling Multi-Identity Consistency for Image Customization via Multi-to-Multi Matching Paradigmaccepted
- Scaling Parallel Sequence Models to Vision Foundation Modelsaccepted
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformersaccepted
- Scaling Spatial Intelligence with Multimodal Foundation Modelsaccepted
- Scaling Test-Time Robustness of Vision-Language Models via Self-Critical Inference Frameworkaccepted
- Scaling Up AI-Generated Image Detection with Generator-Aware Prototypesaccepted
- Scaling View Synthesis Transformersaccepted
- Scaling Zero-Shot Reference-to-Video Generationaccepted
- Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanismaccepted
- Scaling-Aware Data Selection for End-to-End Autonomous Driving Systemsaccepted
- Scaling4D: Pushing the Frontier of Video Novel View Synthesis through Large-Scale Monocular Videosaccepted
- Scan Clusters, Not Pixels: A Cluster-Centric Paradigm for Efficient Ultra-high-definition Image Restorationaccepted
- SceMoS: Scene-Aware 3D Human Motion Synthesis by Planning with Geometry-Grounded Tokensaccepted
- ScenDi: 3D-to-2D Scene Diffusion Cascades for Urban Generationaccepted
- Scene Grounding in the Wildaccepted
- Scene Reconstruction as Mapping Priors for 3D Detectionaccepted
- Scene-Centric Unsupervised Video Panoptic Segmentationaccepted
- Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Modelsaccepted
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Modelaccepted
- SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotationsaccepted
- SceneTok: A Compressed, Diffusable Token Space for 3D Scenesaccepted
- Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understandingaccepted
- SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent Systemaccepted
- Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modelingaccepted
- Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoringaccepted
- Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformersaccepted
- SeD-UD: An Influence-Driven and Hierarchically-Decoupled Information Bottleneck for Multimodal Intent Recognitionaccepted
- SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion Modelsaccepted
- SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Drivingaccepted
- See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflectionaccepted
- See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMsaccepted
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoningaccepted
- See Through the Noise: Improving Domain Generalization in Gaze Estimationaccepted
- See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understandingaccepted
- See What We Cannot See: A Geo-guided Reasoning Benchmark for Object Counting under Adverse Earth Observation Conditionsaccepted
- See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesisaccepted
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Togglesaccepted
- SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Groupingaccepted
- SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generationaccepted
- SeeU: Seeing the Unseen World via 4D Dynamics-aware Generationaccepted
- Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videosaccepted
- Seeing Beyond: Extrapolative Domain Adaptive Panoramic Segmentationaccepted
- Seeing Both Sides: Towards Bidirectional Semantic Alignment for Open-Vocabulary Camouflaged Object Segmentationaccepted
- Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindnessaccepted
- Seeing Conversations: Communication Context Identification in Egocentric Videoaccepted
- Seeing Depth Through Frequency and Motion: A Progressive Training Paradigm for Monocular Depth Estimationaccepted
- Seeing Motion Through Polarity for Event-based Action Recognitionaccepted
- Seeing Through Blur: Tackling Defocus in Spike-Based Imagingaccepted
- Seeing Through Touch: Tactile-Driven Visual Localization of Material Regionsaccepted
- Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspectiveaccepted
- Seeing Through the Shift: Causality-Inspired Robust Generalized Category Discoveryaccepted
- Seeing What Matters: A Training-Free Self-Guided Framework for Multimodal Detail Perception and Reasoningaccepted
- Seeing What Matters: Visual Preference Policy Optimization for Visual Generationaccepted
- Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understandingaccepted
- Seeing is Improving: Visual Feedback for Iterative Text Layout Refinementaccepted
- Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmarkaccepted
- Seeing through Light and Darkness: Sensor-Physics Grounded Deblurring HDR NeRF from Single-Exposure Images and Eventsaccepted
- Seeing through boxes: Non-Line-of-Sight 3D Reconstruction from Radar Signalsaccepted
- Seeing without Pixels: Perception from Camera Trajectoriesaccepted
- Seele: A Unified Acceleration Framework for Real-Time Gaussian Splatting on Mobile Devicesaccepted
- SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentationaccepted
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Imagesaccepted
- SegGBC: Justifiable Coarse-to-Fine Granular-Ball Computing for Enhancing Clustering Image Segmentationaccepted
- SegMo: Co-Designing Content-Aware Sparsity and Locally-Cohesive Segment Parallelism for Efficient VLM Inferenceaccepted
- SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentationaccepted
- SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Modelsaccepted
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMsaccepted
- Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoningaccepted
- Select, Hypothesize and Verify: Towards Verified Neuron Concept Interpretationaccepted
- Selection-as-Nonlinearity: Bridging Attention and Activation via a Joint Game-Decision Lens for Interpretable, Discriminative Visual Representationsaccepted
- Selective Amnesia using Contrastive Subnet Erasure for Class Level Unlearning in Vision Modelsaccepted
- Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentationaccepted
- Selectively Extracting and Injecting Visual Attributes into Text-to-Image Modelsaccepted
- Self-Attention Driven Tensor Representation for High-Order Data Recoveryaccepted
- Self-Consistency for LLM-Based Motion Trajectory Generation and Verificationaccepted
- Self-Corrected Image Generation with Explainable Latent Rewardsaccepted
- Self-Critical Distillation Network for Video-based Commonsense Captioningaccepted
- Self-Diffusion Driven Blind Imagingaccepted
- Self-Evaluation Unlocks Any-Step Text-to-Image Generationaccepted
- Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generationaccepted
- Self-guided Semantic Inspection for Zero-Shot Composed Image Retrievalaccepted
- Self-supervised Dynamic Heterogeneous Degradation Modeling for Unified Zero-Shot Image Restorationaccepted
- SelfHVD: Self-Supervised Handheld Video Deblurringaccepted
- Selfi: Self-improving Reconstruction Engine via 3D Geometric Feature Alignmentaccepted
- SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detectionaccepted
- SemLayer: Semantic-aware Generative Segmentation and Layer Construction for Abstract Iconsaccepted
- SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidanceaccepted
- Semantic Alignment for Pose-Invariant Identity Preserving Diffusionaccepted
- Semantic Audio-Visual Navigation in Continuous Environmentsaccepted
- Semantic Context Matters: Improving Conditioning for Autoregressive Modelsaccepted
- Semantic Derivative Flow: Graph-Guided Diffusion for Controllable Instance Interactionsaccepted
- Semantic Foam: Unifying Spatial and Semantic Scene Decompositionaccepted
- Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learningaccepted
- Semantic Scale Space: A Framework for Controllable Image Abstractionaccepted
- Semantic-Adaptive Diffusion for Dynamic Spatiotemporal Fusionaccepted
- Semantic-Guided Global-Local Collaborative Prompt Learning for Few-Shot Class Incremental Learningaccepted
- SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action Planningaccepted
- Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusionaccepted
- Semi-Supervised Conformal Prediction With Unlabeled Nonconformity Scoreaccepted
- Semi-supervised Echocardiography Video Segmentation via Anchor Semantic Awareness and Continuous Pseudo-label Reforgingaccepted
- SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentationaccepted
- SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Cachingaccepted
- SenseSearch: Empowering Vision-Language Models with High-Resolution Agentic Search-Reasoning via Reinforcement Learningaccepted
- Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Drivingaccepted
- ShadowDraw: From Any Object to Shadow-Drawing Compositional Artaccepted
- Shape-of-You: Fused Gromov-Wasserstein Optimal Transport for Semantic Correspondence in-the-Wildaccepted
- ShapeAR: Generating Editable Shape Layers via Autoregressive Diffusionaccepted
- ShapeR: Robust Conditional 3D Shape Generation from Casual Capturesaccepted
- SharpTimeGS: Sharp and Stable Dynamic Gaussian Splatting via Lifespan Modulationaccepted
- Shedding Light on VLN Robustness: A Black-box Framework for Indoor Lighting-based Adversarial Attackaccepted
- ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimationaccepted
- ShiftLUT: Spatial Shift Enhanced Look-Up Tables for Efficient Image Restorationaccepted
- Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimationaccepted
- ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitionsaccepted
- ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinementaccepted
- ShowUI-p: Flow-based Generative Models as GUI Dexterous Handsaccepted
- ShreddingNet: Coarse-to-Fine Restoration for Multi-Source Shredded Manuscriptsaccepted
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Modelsaccepted
- SignPR: A Progressive Vector-Quantized Diffusion Framework for Sign Language Productionaccepted
- SimLBR: Learning to Detect Fake Images by Learning to Detect Real Imagesaccepted
- SimRecon: SimReady Compositional Scene Reconstruction from Real Videosaccepted
- SimScale: Learning to Drive via Real-World Simulation at Scaleaccepted
- Similarity-Consistent Likelihood Diffusion enables Hidden Person Detection from Wall Reflectionsaccepted
- Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learningaccepted
- Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimizationaccepted
- Simple but Effective Triplet-Based Compression Strategies for Compact Visual Localizationaccepted
- Simple-ViLMedSAM: Simple Text Prompts Meet Vision-Language Models for Medical Image Segmentationaccepted
- SinGeo: Unlock Single Model's Potential for Robust Cross-View Geo-Localizationaccepted
- SineProject: Machine Unlearning for Stable Vision-Language Alignmentaccepted
- Single-Round Scalable Analytic Federated Learningaccepted
- Single-step Diffusion-based Video Coding with Semantic-Temporal Guidanceaccepted
- SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognitionaccepted
- Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume Generationaccepted
- Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow Distillationaccepted
- SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawingaccepted
- SketchDeco: Training-Free Latent Composition for Precise Sketch Colourisationaccepted
- SketchFaceGS: Real-Time Sketch-Driven Face Editing and Generation with Gaussian Splattingaccepted
- SketchRevive: Fine-Grained Pixel-to-Vector Sketch Completion with Diffusion-Prior-Guided Multimodal LLMsaccepted
- SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and Moreaccepted
- SkillSight: Efficient First-Person Skill Assessment with Gazeaccepted
- Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal Predictionaccepted
- Sky2Ground: A Benchmark for Site Modeling under Varying Altitudeaccepted
- SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Designaccepted
- SkySense-VITA: Towards Universal In-context Segmentation of Multi-modal Remote Sensing Imageryaccepted
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoningaccepted
- SliderEdit: Continuous Image Editing with Fine-Grained Instruction Controlaccepted
- Small Object, Great Challenge: A Benchmark for Small Object Visual Groundingaccepted
- Smart Replay: Adaptive Scheduling of Memory Rehearsal for Computational Resource-Aware Incremental Learningaccepted
- SmokeSVD: Smoke Reconstruction from A Single View via Progressive Novel View Synthesis and Refinement with Diffusion Modelsaccepted
- Smoothing the Score Function to Enhance Generalization in Diffusion Modelsaccepted
- SoC: Semantic Orthogonal Calibration for Test-Time Prompt Tuningaccepted
- SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsaccepted
- SoccerMaster: A Vision Foundation Model for Soccer Understandingaccepted
- SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigationaccepted
- Socratic-Geo: Synthetic Data Generation and Cross-Modal Geometric Reasoning via Multi-Agent Interactionaccepted
- Soft Modality-Guided Expert Specialization in MoE-VLMsaccepted
- SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Modelsaccepted
- Solvability of the Viewing Graph Under the Affine Camera Modelaccepted
- Solving Minimal Problems Without Matrix Inversion Using FFT-Based Interpolationaccepted
- Solving a Nonlinear Blind Inverse Problem for Tagged MRI with Physics and Deep Generative Priorsaccepted
- SonoWorld: From One Image to a 3D Audio-Visual Sceneaccepted
- Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animationaccepted
- SounDiT: Geo-Contextual Soundscape-to-Landscape Generationaccepted
- Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Groupingaccepted
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Drivingaccepted
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Modelsaccepted
- SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Timeaccepted
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RLaccepted
- SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free Accelerationaccepted
- Sparse Spectral LoRA: Routed Experts for Medical VLMsaccepted
- Sparse Task Vector Mixup with Hypernetworks for Efficient Knowledge Transfer in Whole-Slide Image Prognosisaccepted
- Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Modelsaccepted
- Sparse-View Localization via Online Neural 3D Regressionaccepted
- SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse Camerasaccepted
- SparseOIT: Improving Order-Independent Transparency 3DGS via Active Set Methodaccepted
- SparseSplat: Towards Applicable Feed-Forward 3D Gaussian Splatting with Pixel-Unaligned Predictionaccepted
- SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Modelaccepted
- Sparsely Timing the Change: A Spiking Temporal Framework for Remote Sensing Interpretationaccepted
- Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detectionaccepted
- Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completionaccepted
- Spatia: Video Generation with Updatable Spatial Memoryaccepted
- SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Modelsaccepted
- Spatial Matters: Position-Guided 3D Referring Expression Segmentationaccepted
- Spatial Retrieval Augmented Autonomous Drivingaccepted
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videosaccepted
- Spatial-Frequency Collaborative Learning for Occluded Visible-Infrared Person Re-Identificationaccepted
- Spatial-SAM: Spatially Consistent 3D Electron Microscopy Segmentation with SDF Memory and Semi-Supervised Learningaccepted
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learningaccepted
- Spatial-Spectral Residuals Informed Diffusion Neural Operator for Pan-sharpeningaccepted
- SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modelingaccepted
- SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generationaccepted
- SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligenceaccepted
- SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoningaccepted
- SpatialTree: How Spatial Intelligence Branches Out in MLLMsaccepted
- SpatialVID: A Large-Scale Video Dataset with Spatial Annotationsaccepted
- Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Trackingaccepted
- Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensoraccepted
- Spatiotemporal Pyramid Flow Matching for Climate Emulationaccepted
- Spe-BEVHead: Rethinking the Detection Head Design for Bird's-Eye-View Object Detectionaccepted
- Specificity-aware reinforcement learning for fine-grained open-world classificationaccepted
- Spectral Conformal Risk Control: Distribution-Free Tail Guarantees via Bayesian Quadratureaccepted
- Spectral Mixture-of-Experts for Continual Learningaccepted
- Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action Segmentationaccepted
- Spectral Super-Resolution via Adversarial Unfolding and Data-Driven Spectrum Regularization: From Multispectral Satellite Data to NASA Hyperspectral Imageaccepted
- Spectral-Geometric Neural Fields for Pose-Free LiDAR View Synthesisaccepted
- Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imageryaccepted
- Spectrum from Defocus: Fast Spectral Imaging with Chromatic Focal Stackaccepted
- SpeeDe3DGS: Speedy Deformable 3D Gaussian Splatting with Temporal Pruning and Motion Groupingaccepted
- SpeeDiff: Scalable Pixel-Anchored End-to-End Latent Diffusion Modelaccepted
- Speeding Up the Learning of 3D Gaussians with Much Shorter Gaussian Listsaccepted
- Spherical Leech Quantization for Visual Tokenization and Generationaccepted
- Spherical Voronoi: Directional Appearance as a Differentiable Partition of the Sphereaccepted
- SpiderCam: Low-Power Snapshot Depth from Differential Defocusaccepted
- Spike-driven Discrete Aggregation for Event-based Object Detectionaccepted
- SpikeTrack: A Spike-driven Framework for Efficient Visual Trackingaccepted
- SpikeTrack: High-performance and Energy-efficient Event-Based Object Tracking with Spiking Neural Networkaccepted
- SpiralDiff: Spiral Diffusion with LoRA for RGB-to-RAW Conversion Across Camerasaccepted
- Spk2VidNet: A Hierarchical Recurrent Architecture for High-Fidelity Video Reconstruction from Long Spike-Camera Streamsaccepted
- Splat-Based Metal Artifact Reduction in Cone-Beam CT via Compact Attenuation Modelingaccepted
- SplatSuRe: Selective Super-Resolution for Multi-view Consistent 3D Gaussian Splattingaccepted
- Splatent: Splatting Diffusion Latents for Novel View Synthesisaccepted
- SplitFlux: Learning to Decouple Content and Style from a Single Imageaccepted
- Spot The Ball: A Benchmark for Visual Social Inferenceaccepted
- SpotEdit: Selective Region Editing in Diffusion Transformersaccepted
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representationaccepted
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answeringaccepted
- Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulationaccepted
- Stabilizing Feature Geometry in Noisy Pretrained Models for Robust Downstream Tasksaccepted
- Stabilizing Streaming Video Geometry via Dynamic Feature Normalizationaccepted
- Stable Mean Flow: Lyapunov-Inspired One-Step Flow Matchingaccepted
- Stable Spike: Dual Consistency Optimization via Bitwise AND Operations for Spiking Neural Networksaccepted
- Stable and Efficient Single-Rollout RL for Multimodal Reasoningaccepted
- StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasetsaccepted
- StableMaterials: Enhancing Diversity in Material Generation via Semi-Supervised Learningaccepted
- Stake the Points: Structure-Faithful Instance Unlearningaccepted
- Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generationaccepted
- Statistical Characteristic-Guided Denoising for Rapid High-Resolution Transmission Electron Microscopy Imagingaccepted
- Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioningaccepted
- Stealing Split Learning Bottom Models by Recovering Embedding Geometryaccepted
- Steering Where to Diffuse: Generative Modeling of Phenotypic Response Simulation with Steered Diffusion Bridgeaccepted
- Stepwise Credit Assignment for GRPO on Flow-Matching Modelsaccepted
- Stereo World Model: Camera-Guided Stereo Video Generationaccepted
- StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generationaccepted
- Stitch-a-Demo: Creating Video Demonstrations from Multistep Descriptionsaccepted
- Stochastic Ray Tracing for the Reconstruction of 3D Gaussian Splattingaccepted
- StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narrativesaccepted
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatarsaccepted
- StreamDiT: Real-Time Streaming Text-to-Video Generationaccepted
- StreamRAG: Enhancing Real-Time Video Understanding with Retrieval Augmentationaccepted
- StreamReady: Learning What to Answer and When in Long Streaming Videosaccepted
- StreamVLO: Streaming Visual-LiDAR Odometry with Cumulative Drift Compensationaccepted
- Streaming Diffusion Model for Fast Infrared and Visible Video Fusionaccepted
- Streaming Video Crime Anticipation with Spatio-Temporal Causal Reasoningaccepted
- Streaming Video Instruction Tuningaccepted
- StreamingTOM: Streaming Token Compression for Efficient Video Understandingaccepted
- Streamlined Knowledge Distillationaccepted
- Streamlined Open-Vocabulary Human-Object Interaction Detectionaccepted
- Stronger Normalization-Free Transformersaccepted
- StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cuesaccepted
- Structural Action Transformer for 3D Dexterous Manipulationaccepted
- Structural Graph Probing of Vision-Language Modelsaccepted
- Structure-Aware Representation Distillation for Tiny-Dense Object Segmentationaccepted
- Structure-to-Intensity Diffusion for Adverse-Weather LiDAR Generationaccepted
- Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward Modelingaccepted
- StyleDoctor: Towards Specialist Reward Model for Style-centric Generation Tasksaccepted
- StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image Referencesaccepted
- StyleTextGen: Style-Conditioned Multilingual Scene Text Generationaccepted
- SuP: Sub-cloud Driven Point Cloud Registrationaccepted
- Submodel Extraction for Efficient and Personalized Federated Learning via Optimal Transportaccepted
- Subspace Alignment for CLIP-based Continual Learning via Canonical Correlation Analysisaccepted
- SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modelingaccepted
- SunFaded: Illumination-Aware Gaussian Splatting for Dark Scenes with Camera-Mounted Active Lightingaccepted
- Superman: Unifying Skeleton and Vision for Human Motion Perception and Generationaccepted
- Suppressing Non-Semantic Noise in Masked Image Modeling Representationsaccepted
- SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmarkaccepted
- SwiftTailor: Efficient 3D Garment Generation with Geometry Image Representationaccepted
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overheadaccepted
- SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controlsaccepted
- SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generationaccepted
- Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understandingaccepted
- SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perceptionaccepted
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generationaccepted
- SyncDreamer: Controllable and Expressive Avatar Generation Beyond the Talking Headaccepted
- SyncMos: Scalable Motion Synchronisation for Multi-Agent Scene Interactionaccepted
- Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videosaccepted
- SynthRGB-T: Language-Vision Guided Image Translation for Diversity Synthesisaccepted
- Synthesizing Visual Concepts as Vision-Language Programsaccepted
- Synthetic Curriculum Reinforces Compositional Text-to-Image Generationaccepted
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Groundingaccepted
- T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Groundingaccepted
- TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detectionaccepted
- TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Expertsaccepted
- TALO: Pushing 3D Vision Foundation Models Towards Globally Consistent Online Reconstructionaccepted
- TALON: Test-time Adaptive Learning for On-the-Fly Category Discoveryaccepted
- TAMER: A Tri-Modal Contrastive Alignment and Multi-Scale Embedding Refinement Framework for Zero-Shot ECG Diagnosisaccepted
- TANGO: Learning Distribution-wise Foundation Prior Consistency and Instance-wise Style Calibration for Medical Image Generalizationaccepted
- TANGO: Text-Anchored Guided Optimization for Robust Fine-tuning Vision-Language Models under Label Noiseaccepted
- TAP: A Token-Adaptive Predictor Framework for Training-Free Diffusion Accelerationaccepted
- TAPE: Task-Adaptive Prototype Evolution in Audio-Language Models for Fully Few-shot Class-incremental Audio Classificationaccepted
- TAR: Token-Aware Refinement for Fine-grained Generalized Category Discoveryaccepted
- TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Expertsaccepted
- TAlignDiff: Automatic Tooth Alignment assisted by Diffusion-based Transformation Learningaccepted
- TC-Pade: Trajectory-Consistent Pade Approximation for Diffusion Accelerationaccepted
- TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignmentaccepted
- TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Modelsaccepted
- TESO: Online Tracking of Essential Matrix by Stochastic Optimizationaccepted
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysisaccepted
- TEXTRIX: Latent Attribute Grid for Native Texture Generation and Beyondaccepted
- TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detectionaccepted
- TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detectionaccepted
- TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completionaccepted
- TGT: Text-Grounded Trajectories for Locally Controlled Video Generationaccepted
- TGTrack: Temporal Generative Learning for Unified Single Object Trackingaccepted
- THE MORE, THE MERRIER: CONTRASTIVE FUSION FOR HIGHER-ORDER MULTIMODAL ALIGNMENTaccepted
- TIACam: Text-Anchored Invariant Feature Learning with Auto-Augmentation for Camera-Robust Zero-Watermarkingaccepted
- TIGER: A Unified Framework for Time, Images and Geo-location Retrievalaccepted
- TIM: Temporal Decoupling with Iterative Mutual-Refinement Model for Longitudinal Radiology Report Generationaccepted
- TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Modelsaccepted
- TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignmentaccepted
- TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detectionaccepted
- TM-BSN: Triangular-Masked Blind-Spot Network for Real-World Self-Supervised Image Denoisingaccepted
- TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Dual-Level Scale-Oriented Contrastaccepted
- TRANSPORTER: Transferring Visual Semantics from VLM Manifoldsaccepted
- TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognitionaccepted
- TRIDENT: A Trimodal Cascade Generative Framework for Drug and RNA-Conditioned Cellular Morphology Synthesisaccepted
- TRM-VLA: Temporal-Aware Chain-of-Thought Reasoning and Memorization for Vision-Language-Action Modelsaccepted
- TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videosaccepted
- TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognitionaccepted
- TSTM: Temporal Segmentation for Task-relevant Mask in Visual Reinforcement Learning Generalizationaccepted
- TTAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Eventsaccepted
- TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Modelsaccepted
- TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Modelsaccepted
- TTRV: Test-Time Reinforcement Learning for Vision Language Modelsaccepted
- TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolutionaccepted
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Modelsaccepted
- TV2TV: A Unified Framework for Interleaved Language and Video Generationaccepted
- TVHighlights: LLM-Guided Human-Free Collaborative Training for Video Highlight Detection in Movies and TV Dramasaccepted
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummiesaccepted
- TWINGS: Thin Plate Splines Warp-aligned Initialization for Sparse-View Gaussian Splattingaccepted
- TableMix: Enhancing Multimodal Table Reasoning in MLLMs from a Data-Centric Perspectiveaccepted
- TacSIm: A Dataset and Benchmark for Football Tactical Style Imitationaccepted
- Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute Miningaccepted
- Tackling Model Bias via Game-theoretic Multi-agent Collaboration Framework for Hateful Meme Classificationaccepted
- TagSplat: Topology-Aware Gaussian Splatting for Dynamic Mesh Modeling and Trackingaccepted
- Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenesaccepted
- Talking Together: Synthesizing Co-Located 3D Conversations from Audioaccepted
- Taming Generative Diffusion Model for Task-Oriented Infrared Imagingaccepted
- Taming Noise-Induced Prototype Degradation for Privacy-Preserving Personalized Federated Fine-Tuningaccepted
- Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learningaccepted
- Taming Sampling Perturbations with Variance Expansion Loss for Latent Diffusion Modelsaccepted
- Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Controlaccepted
- Taming the Long Tail: Rebalancing Adversarial Training via Adaptive Perturbationaccepted
- Target-Aware Invertible Encoder with Reconstruction Guidance for Infrared Small Target Detectionaccepted
- Task-Aware Image Signal Processor for Advanced Visual Perceptionaccepted
- Task-Driven Implicit Representations for Automated Design of LiDAR Systemsaccepted
- Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentationaccepted
- TaskForce: Cooperative Multi-agent Reinforcement Learning for Multi-task Optimizationaccepted
- TaskIT: Memory-Efficient Fine-Tuning of Multi-LoRA LLMs via Cross-Task Importance Transferaccepted
- Tavatar: Topology-Aware Gaussian Attribute Derivation for Animatable Human Avatarsaccepted
- Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Modelsaccepted
- TeFlow: Enabling Multi-frame Supervision for Self-Supervised Feed-forward Scene Flow Estimationaccepted
- TeHOR: Text-Guided 3D Human and Object Reconstruction with Texturesaccepted
- Tea-Adapter: Teacher Adapter for Efficient Conditional Generationaccepted
- Teacher-Guided Routing for Sparse Vision Mixture-of-Expertsaccepted
- Teaching DINOv3 About Partial 3D Geometry: A Self-Supervised Geometry-Aware Approachaccepted
- TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Sizeaccepted
- Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attentionaccepted
- Tell2Adapt: A Unified Framework for Source Free Unsupervised Domain Adaptation via Vision Foundation Modelaccepted
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learningaccepted
- TempoControl: Temporal Attention Guidance for Text-to-Video Modelsaccepted
- TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Predictionaccepted
- Temporal Equilibrium MeanFlow: Bridging the Scale Gap for One-Step Generationaccepted
- Temporal Imbalance of Positive and Negative Supervision in Class-Incremental Learningaccepted
- Temporal Interaction in Spiking Transformers with Multi-Delay Mixeraccepted
- Temporal Inversion for Learning Interval Change in Chest X-Raysaccepted
- Temporal Representation Enhancement (TRE): Learning to Forget Dominant Patterns for Enhanced Temporal Spiking Featuresaccepted
- TerraScope: Pixel-Grounded Visual Reasoning for Earth Observationaccepted
- TerraSeg: Self-Supervised Ground Segmentation for Any LiDARaccepted
- Test-Time 3D Occupancy Predictionaccepted
- Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisationaccepted
- Test-Time Attention Purification for Backdoored Large Vision Language Modelsaccepted
- Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative Modelingaccepted
- Test-Time Multi-Prompt Adaptation for Open-Vocabulary Remote Sensing Image Segmentationaccepted
- Test-Time Perturbation Tuning with Delayed Feedback for Vision-Language-Action Modelsaccepted
- Test-Time Training for LiDAR Semantic Segmentation under Corruption via Geometric Inlier Discriminationaccepted
- Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistencyaccepted
- Test-time Sparsity for Extreme Fast Action Diffusionaccepted
- Text-Driven 3D Hand Motion Generation from Sign Language Dataaccepted
- Text-Image Conditioned 3D Generationaccepted
- Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image Retrievalaccepted
- Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Modelsaccepted
- Text-guided Feature Disentanglement for Cross-modal Gait Recognitionaccepted
- TextFM: Robust Semi-dense Feature Matching with Language Guidanceaccepted
- TextOVSR: Text-Guided Real-World Opera Video Super-Resolutionaccepted
- TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Renderingaccepted
- Texvent: Asynchronous Event Data Simulation via Text Promptaccepted
- The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Modelsaccepted
- The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contextsaccepted
- The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignmentaccepted
- The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category Discoveryaccepted
- The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergyaccepted
- The Drift Kernel: Why Diffusion Models Change Even When Told Not Toaccepted
- The Geometry of Robustness: Optimizing Loss Landscape Curvature and Feature Manifold Alignment for Robust Finetuning of Vision-Language Modelsaccepted
- The Golden Subspace: Where Efficiency Meets Generalization in Continual Test-Time Adaptationaccepted
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generationaccepted
- The Invisible Gorilla Effect in Out-of-distribution Detectionaccepted
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognitionaccepted
- The Midas Touch for Metric Depthaccepted
- The Missing GAP: From Solving Square Jigsaw Puzzles to Handling Real World Archaeological Fragmentsaccepted
- The Missing Point in Vision Transformers for Universal Image Segmentationaccepted
- The Power of Decaying Steps: Enhancing Attack Stability and Transferability for Sign-based Optimizersaccepted
- The Power of Prior: Training-Free Open-Vocabulary Semantic Segmentation with LLaVAaccepted
- The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detectionaccepted
- The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identificationaccepted
- The Surprising Effectiveness of Noise Pretraining for Implicit Neural Representationsaccepted
- The Universal Normal Embeddingaccepted
- The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injectionaccepted
- TherA: Thermal-Aware Visual-Language Prompting for Controllable RGB-to-Thermal Infrared Translationaccepted
- Thermal Diffusion Matters: Infrared Spatial-Temporal Video Super-Resolution through Heat Conduction Priorsaccepted
- Thermal is Always Wild: Characterizing and Addressing Challenges in Thermal-Only Novel View Synthesisaccepted
- Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detectionaccepted
- Thermally Activated Dual-Modal Adversarial Clothing against AI Surveillance Systemsaccepted
- Think 360deg: Beyond Depth: Evaluating the Width-centric Reasoning Capability of MLLMsaccepted
- Think Before You Drive: World Model-Inspired Multimodal Groundingaccepted
- Think Visually, Reason Textually: Vision-Language Synergy in Abstract Reasoningaccepted
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Viewsaccepted
- Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understandingaccepted
- Think-Then-Generate: Structural Chain-of-Thought Reasoning for Consistent 3D Generationaccepted
- Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Modelsaccepted
- ThinkGen: Generalized Thinking for Visual Generationaccepted
- Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMsaccepted
- Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Modelsaccepted
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoningaccepted
- Thinking in 360deg: Humanoid Visual Search in the Wildaccepted
- Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D Worldaccepted
- Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decodingaccepted
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understandingaccepted
- Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Modelaccepted
- Thinking with Programming Vision: Towards a Unified View for Thinking with Imagesaccepted
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigmaccepted
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generationaccepted
- ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inferenceaccepted
- Through the Frequency Lens: Cross-Domain Generalisable Gaze Estimation with Adaptive Modulationaccepted
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generationaccepted
- Time Blindness: Why Video-Language Models Can't See What Humans Can?accepted
- Time Without Time: Pseudo-Temporal Representation for Space-Time Super-Resolutionaccepted
- Time-Aware One Step Diffusion Network for Real-World Image Super-Resolutionaccepted
- Time-Specialized Event-Image Alignment for Blur-to-Video Decompositionaccepted
- TimeBridge: Self-Supervised Video Representation Learning via Start-End Joint Embedding and In-Between Frame Predictionaccepted
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMsaccepted
- TimeRipples: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Spaceaccepted
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understandingaccepted
- Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Modelsaccepted
- Token Warping Helps MLLMs Look from Nearby Viewpointsaccepted
- TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokensaccepted
- TokenHand: Discrete Token Representation for Efficient Hand Mesh Reconstructionaccepted
- TokenLight: Precise Lighting Control in Images using Attribute Tokensaccepted
- TokenSplat: Token-aligned 3D Gaussian Splatting for Feed-forward Pose-free Reconstructionaccepted
- TokenTrace: Multi-Concept Attribution through Watermarked Token Recoveryaccepted
- Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor Plansaccepted
- Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelityaccepted
- TopoCL: Topological Contrastive Learning for Medical Imagingaccepted
- TopoHR: Hierarchical Centerline Representation for Cyclic Topology Reasoning in Driving Scenes with Point-to-Instance Relationsaccepted
- TopoMA: Topology-Guided Multi-Agent Dense RGB 3D Reconstruction via Distributed Inferenceaccepted
- TopoMesh: High-Fidelity Mesh Autoencoding via Topological Unificationaccepted
- TopoSlide: Topologically-Informed Histopathology Whole Slide Image Representation Learningaccepted
- Topology-aware Feature Propagation for Unsupervised Non-rigid Point Cloud Correspondenceaccepted
- TouchDream: 3D Object Completion through Imagined Touchaccepted
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspectiveaccepted
- Toward Early Quality Assessment of Text-to-Image Diffusion Modelsaccepted
- Toward Generalizable Whole Brain Representations with High-Resolution Light-Sheet Dataaccepted
- Toward Low-Cost yet Effective Temporal Learning for UAV Trackingaccepted
- Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Datasetaccepted
- Towards Balanced Multi-Modal Learning in 3D Human Pose Estimationaccepted
- Towards Calibrating Prompt Tuning of Vision- Language Modelsaccepted
- Towards Cross-Modal Preservation, Consistency and Alignment for Privacy-Preserving Visible-Infrared Person Re-Identificationaccepted
- Towards Decompositional Human Motion Generation with Energy-Based Diffusion Modelsaccepted
- Towards Dynamic Modality Alignment in Multimodal Continual Learningaccepted
- Towards Efficient Medical Reasoning with Minimal Fine-Tuning Dataaccepted
- Towards Fine-Grained Attribution: Instance-Aware Preference Optimization for Aligning Diffusion Modelsaccepted
- Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Cloudsaccepted
- Towards GUI Agents: Vision-Language Diffusion Models for GUI Groundingaccepted
- Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learningaccepted
- Towards Generalized Multimodal Homography Estimationaccepted
- Towards Generalized Representations for Low-Light Understanding: When Signal Constancy Meets Semantic Enrichmentaccepted
- Towards High-Quality Image Segmentation: Improving Topology Accuracy by Penalizing Neighbor Pixelsaccepted
- Towards High-resolution and Disentangled Reference-based Sketch Colorizationaccepted
- Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interactionaccepted
- Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimizationaccepted
- Towards Holistic Modeling for Video Frame Interpolation with Auto-regressive Diffusion Transformersaccepted
- Towards Human-Imperceptible Backdoor Attacks on Text-to-Image Diffusion Modelsaccepted
- Towards Human-Like Robot Handwriting via Contour-Aware Generationaccepted
- Towards Intrinsic-Aware Monocular 3D Object Detectionaccepted
- Towards Knowledge-augmented Bayesian Deep Learning For Computer Visionaccepted
- Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid Robotsaccepted
- Towards Multimodal Domain Generalization with Few Labelsaccepted
- Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoningaccepted
- Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Datasetaccepted
- Towards Persistence: Learning Topological Constraints for Event-based Small Object Detectionaccepted
- Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Frameworkaccepted
- Towards Policy-Adaptive Image Guardrail: Benchmark and Methodaccepted
- Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Trainingaccepted
- Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priorsaccepted
- Towards Reasoning-Preserving Unlearning in Multimodal Large Language Modelsaccepted
- Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networksaccepted
- Towards Robust Multi-Modal Semantic Segmentation with Teacher-Student Framework and Hybrid Prototype Distillationaccepted
- Towards Robust Multimodal Large Language Models Against Jailbreak Attacksaccepted
- Towards Robust Sequential Decomposition for Complex Image Editingaccepted
- Towards Robust Vision Transformers: Path Dependency Analysis and a Simple Two-Stage Adversarial Trainingaccepted
- Towards Sparse Video Understanding and Reasoningaccepted
- Towards Stable Federated Continual Test-Time Adaptation in Wild Worldaccepted
- Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Videoaccepted
- Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approachaccepted
- Towards Storytelling Animations: Joint Synthesis of Human and Camera Motionsaccepted
- Towards Streaming Referring Video Segmentation via Large Language Modelaccepted
- Towards Training-free Scene Text Editingaccepted
- Towards Uncertainty-aware Unsupervised Domain Adaptation for Videos and Time-Series with Causal Optimal Transportaccepted
- Towards Unified Human Perception and Machine Understanding: Token Flow Guided Compression Frameworkaccepted
- Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysisaccepted
- Towards Visual Query Localization in the 3D Worldaccepted
- Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspectiveaccepted
- TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videosaccepted
- TrackMAE: Video Representation Learning via Track Mask and Predictaccepted
- Tracking by Predicting 3-D Gaussians Over Timeaccepted
- Tracking through Severe Occlusion via Event-Derived Transient Cuesaccepted
- Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animationaccepted
- TrafficAlign: Aligning Large Language Models for Traffic Scenario Generationaccepted
- Trainable Log-linear Sparse Attention for Efficient Diffusion Transformersaccepted
- Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automationaccepted
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learningaccepted
- Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid Promptaccepted
- Training-Only Heterogeneous Image-Patch-Text Graph Supervision for Advancing Few-Shot Learning Adaptersaccepted
- Training-free Detection of Generated Videos via Spatial-Temporal Likelihoodsaccepted
- Training-free Mixed-Resolution Latent Upsampling for Spatially Accelerated Diffusion Transformersaccepted
- Training-free Motion Factorization for Compositional Video Generationaccepted
- Training-free, Perceptually Consistent Low-Resolution Previews with High-Resolution Image for Efficient Workflows of Diffusion Modelsaccepted
- TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigationaccepted
- TrajTok: Learning Trajectory Tokens Enhances Video Understandingaccepted
- TransPrune: Token Transition Pruning for Efficient Large Vision-Language Modelaccepted
- Transform to Transfer: Boosting Adversarial Attack Transferability on Vision-Language Pre-training Modelsaccepted
- Transition Matching Distillation for Fast Video Generationaccepted
- Transition Models: Rethinking the Generative Learning Objectiveaccepted
- Translating Signals to Languages for sEMG-Based Activity Recognitionaccepted
- TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Explorationaccepted
- Tri-Modal Fusion Transformers for UAV-based Object Detectionaccepted
- Tri-Subspaces Disentanglement for Multimodal Sentiment Analysisaccepted
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detectionaccepted
- TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglementaccepted
- TriSim: Tri-Dimensional Similarity Modeling with Extreme Value Theory for False-Negative Mitigation in Remote Sensing Image-Text Retrievalaccepted
- TruckDrive: Long-Range Autonomous Highway Driving Datasetaccepted
- Trust-calibrated Collaborative Learning for Long-Tailed Visual Recognitionaccepted
- Tunable Soft Equivariance with Guaranteesaccepted
- Turbo-GS: Accelerating 3D Gaussian Fitting for High-Resolution Radiance Fieldsaccepted
- Turning Pre-Trained Vision Transformers into End-to-End Histopathology Whole Slide Image Models for Survival Predictionaccepted
- Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detectionaccepted
- Twin-T & TwintVQA: A Reliable Structure-Detail Separating VLM and a Comprehensive Benchmark for Chart and Table Tasksaccepted
- U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generationaccepted
- U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequencesaccepted
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancementaccepted
- UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVsaccepted
- UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detectionaccepted
- UAVLight: A Benchmark for Illumination-Robust 3D Reconstruction in Unmanned Aerial Vehicle (UAV) Scenesaccepted
- UCAN: Unified Convolutional Attention Network for Expansive Receptive Fields in Lightweight Super-Resolutionaccepted
- UCMNet: Uncertainty-Aware Context Memory Network for Under-Display Camera Image Restorationaccepted
- UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimationaccepted
- UETrack: A Unified and Efficient Framework for Single Object Trackingaccepted
- UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modelingaccepted
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Modelsaccepted
- UI-Lens: Assessing General MLLMs' Potential to Automate UI Display Quality Assuranceaccepted
- UIKA: Fast Universal Head Avatar from Pose-Free Imagesaccepted
- ULF-Loc: Unbiased Landmark Feature for Robust Visual Localization with 3D Gaussian Splattingaccepted
- UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language Modelingaccepted
- UNICBench: UNIfied Counting Benchmark for MLLMaccepted
- UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attendersaccepted
- URICA: A Uniformity Region Affine Identifier Capture Algorithm for Arbitrary Region Retrieval in Pathology Imagesaccepted
- URScenes: A Multi-scenario Dataset for Unstructured Road Environmentsaccepted
- UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimationaccepted
- UTPTrack: Towards Simple and Unified Token Pruning for Visual Trackingaccepted
- UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigmaccepted
- UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language Conditionsaccepted
- U^2Flow: Uncertainty-Aware Unsupervised Optical Flow Estimationaccepted
- Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking from Sparse Inertial Sensors and Ranging-based Between-sensor Distancesaccepted
- Ultra-Fast Neural Video Compressionaccepted
- Ultra-Low Bitrate Perceptual Image Compression with Shallow Encoderaccepted
- UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratiosaccepted
- Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understandingaccepted
- UnReflectAnything: RGB-Only Highlight Removal by Rendering Synthetic Specular Supervisionaccepted
- Unblur-SLAM: Dense Neural SLAM for Blurry Inputsaccepted
- Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Modelsaccepted
- Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Modelsaccepted
- Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object Detectionaccepted
- Uncertainty-driven 3D Gaussian Splatting Active Mapping via Anisotropic Visibility Fieldaccepted
- Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Modelsaccepted
- Underground Plant Exploration: Non-Destructive 3D Root Assessment with GPR Based on Point Graph Neural Networkaccepted
- Understanding Counting Mechanisms in Large Language and Vision-Language Modelsaccepted
- Understanding Task Transfer in Vision-Language Modelsaccepted
- Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminabilityaccepted
- Understanding and Enforcing Weight Disentanglement in Task Arithmeticaccepted
- Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Modelsaccepted
- Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Modelsaccepted
- Understanding, Accelerating, and Improving MeanFlow Trainingaccepted
- Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generationaccepted
- Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing Modalitiesaccepted
- Uni-Hema: Unified Model for Digital Hematopathologyaccepted
- Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Imagesaccepted
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactionsaccepted
- UniChange: Unifying Change Detection with Multimodal Large Language Modelaccepted
- UniComp: Rethinking Video Compression Through Informational Uniquenessaccepted
- UniCompress: Token Compression for Unified Vision-Language Understanding and Generationaccepted
- UniCorrn: Unified Correspondence Transformer Across 2D and 3Daccepted
- UniDAC: Universal Metric Depth Estimation for Any Cameraaccepted
- UniDef: Universal Defense Against Unauthorized Image Manipulationaccepted
- UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videosaccepted
- UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifyingaccepted
- UniFusion: A Unified Image Fusion Framework with Robust Representation and Source-Aware Preservationaccepted
- UniGame: Turning a Unified Multimodal Model Into Its Own Adversaryaccepted
- UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in RLaccepted
- UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detectionaccepted
- UniGeoRS: A Unified Benchmark for Tri-view Geo-Localizationaccepted
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenesaccepted
- UniLDiff: Unlocking the Power of Diffusion Priors for All-in-One Image Restorationaccepted
- UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speakingaccepted
- UniLight: A Unified Representation for Lightingaccepted
- UniM: A Unified Any-to-Any Interleaved Multimodal Benchmarkaccepted
- UniMERNet: A Universal Network for Real-World Mathematical Expression Recognitionaccepted
- UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompressionaccepted
- UniPR: Unified Object-level Real-to-Sim Perception and Reconstruction from a Single Stereo Pairaccepted
- UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latentsaccepted
- UniPercept: A Unified Diffusion Model for Generalizable Visual Perceptionaccepted
- UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matchingaccepted
- UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimizationaccepted
- UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Registeraccepted
- UniSER: A Foundation Model for Unified Soft Effects Removalaccepted
- UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward Passaccepted
- UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual Promptingaccepted
- UniT: Unified Multimodal Chain-of-Thought Test-time Scalingaccepted
- UniTEX: Universal High Fidelity Generative Texturing for 3D Shapesaccepted
- UniVBench: Towards Unified Evaluation for Video Foundation Modelsaccepted
- UniVerse: A Unified Modulation Framework for Segmentation-Free, Disentangled Multi-Concept Personalizationaccepted
- UniVerse: Empower Unified Generation with Reasoning and Knowledgeaccepted
- UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Editsaccepted
- Unified Camera Positional Encoding for Controlled Video Generationaccepted
- Unified Customized Generation by Disentangled Reward Modelingaccepted
- Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimizationaccepted
- Unified Latent Space for Understanding and Generation via Semantic Auto-encoderaccepted
- Unified Multimodal Models as Auto-Encodersaccepted
- Unified Number-Free Text-to-Motion Generation Via Flow Matchingaccepted
- Unified Personalized Understanding, Generating and Editingaccepted
- Unified Primitive Proxies for Structured Shape Completionaccepted
- Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retentionaccepted
- Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Cameraaccepted
- Unified Vector Floorplan Generation via Markup Representationaccepted
- Unifying Language-Action Understanding and Generation for Autonomous Drivingaccepted
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generationaccepted
- Unifying Precise Keyframes and Semantic Control via Multi-level Diffusionaccepted
- Unique Lives, Shared World: Learning from Single-Life Videosaccepted
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generationaccepted
- Universal 3D Shape Matching via Coarse-to-Fine Language Guidanceaccepted
- Universal Guideline-Driven Image Clustering via a Hybrid LLM Agentaccepted
- Universal-to-Specific: Dynamic Knowledge-Guided Multiple Instance Learning for Few-Shot Whole Slide Image Classificationaccepted
- Unlearning without Forgetting: Securely Removing Targeted Concepts from Large-Scale Vision-Language Open-Vocabulary Detectorsaccepted
- Unleashing Stealthy Backdoor Pandemic by Infecting a Single Diffusion Modelaccepted
- Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from Failuresaccepted
- Unleashing Vision-Language Semantics for Deepfake Video Detectionaccepted
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Modelsaccepted
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detectionaccepted
- Unlocking 3D Affordance Segmentation with 2D Semantic Knowledgeaccepted
- Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait Recognitionaccepted
- Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Frameworkaccepted
- Unlocking Pre-trained Weights: Parameter Inheritance for Zero-Shot Initializationaccepted
- Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methodsaccepted
- Unlocking Token Rewards via Training-Free Reward Attributionaccepted
- Unlocking the Power of Critical Factors for 3D Visual Geometry Estimationaccepted
- Unpaired Image Deraining Using Reward-Guided Self-Reinforcement Strategyaccepted
- Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Imagesaccepted
- Unsafe2Safe: Controllable Image Anonymization for Downstream Utilityaccepted
- Unstitching the Chimera: Frame-Level Risk and Train-Free Mitigation for Video Hallucinationaccepted
- Unsupervised 3d Motion Estimation Using Event Cameraaccepted
- Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priorsaccepted
- Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Modelaccepted
- Unsupervised Multi-agent and Single-agent Perception from Cooperative Viewsaccepted
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsamplingaccepted
- Urban-GS: A Unified 3D Gaussian Splatting Framework for Compact and High-Fidelity Aerial-to-Street Reconstructionaccepted
- V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMsaccepted
- V-DPM: 4D Video Reconstruction with Dynamic Point Mapsaccepted
- V-RGBX: Video Editing with Accurate Controls over Intrinsic Propertiesaccepted
- V2U4Real: A Real-world Large-scale Dataset for Vehicle-to-UAV Cooperative Perceptionaccepted
- VA-p: Variational Policy Alignment for Pixel-Aware Autoregressive Generationaccepted
- VABench: A Comprehensive Benchmark for Audio-Video Generationaccepted
- VAD-GS: Visibility-Aware Densification for 3D Gaussian Splatting in Dynamic Urban Scenesaccepted
- VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generationaccepted
- VAST: Video Ability-Stratified Taxonomy for Data-Efficient Video Reasoningaccepted
- VCP-Attack: Visual-Contrastive Projection for Transferable Black-Box Targeted Attacks on Large Vision-Language Modelsaccepted
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridgingaccepted
- VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimationaccepted
- VDFE: Difference-Aware 3D Scene Editing with Non-Intrusive Video Diffusion Priors for Multi-View Consistency and Efficiencyaccepted
- VDOT: Efficient Unified Video Creation via Optimal Transport Distillationaccepted
- VEMamba: Efficient Isotropic Reconstruction of Volume Electron Microscopy with Axial-Lateral Consistent Mambaaccepted
- VENI: Variational Encoder for Natural Illuminationaccepted
- VES-RFT: Rewarding Visual Evidence Sensitivity to Mitigate Hallucinations in Large Vision-Language Modelsaccepted
- VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluationaccepted
- VGA: Empowering Aerial-Ground Localization by Visual Geometry Alignmentaccepted
- VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scaleaccepted
- VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Drivingaccepted
- VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimationaccepted
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detectionaccepted
- VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentationaccepted
- VGGT-Ωaccepted
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Predictionaccepted
- VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinementaccepted
- VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Networkaccepted
- VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Datasetaccepted
- VIRAL: Visual Sim-to-Real at Scale for Humanoid Loco-Manipulationaccepted
- VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimationaccepted
- VIRO: Robust and Efficient Neuro-Symbolic Reasoning with Verification for Referring Expression Comprehensionaccepted
- VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentationaccepted
- VISTA: A Test-Time Self-Improving Video Generation Agentaccepted
- VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactionsaccepted
- VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessmentaccepted
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimizationaccepted
- VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Modelsaccepted
- VL-Eraser: Vacuum Distillation for Machine Unlearning in Vision-Language Modelsaccepted
- VL-RouterBench: A Benchmark for Vision-Language Model Routingaccepted
- VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modelingaccepted
- VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compressionaccepted
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstructionaccepted
- VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recoveryaccepted
- VLM-Loc: Localization in Point Cloud Maps via Vision-Language Modelsaccepted
- VLM-PTQ: Efficient Post-Training Quantization for Large Vision-Language Modelsaccepted
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigmaccepted
- VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detectionaccepted
- VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformationaccepted
- VMonarch: Efficient Video Diffusion Transformers with Structured Attentionaccepted
- VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillationaccepted
- VOSR: A Vision-Only Generative Model for Image Super-Resolutionaccepted
- VQ-VA World: Towards High-Quality Visual Question-Visual Answeringaccepted
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstructionaccepted
- VRCLIP: Multimodal Canonical Correlation Alignment for CLIP-Driven Vision-Radio Person Re-Identificationaccepted
- VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cuesaccepted
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environmentsaccepted
- VSRELL: A Simple Baseline for Video Super-Resolution and Enhancement in Low-Light Environmentaccepted
- VT-Intrinsic: Physics-Based Decomposition of Reflectance and Shading using a Single Visible-Thermal Image Pairaccepted
- VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangementaccepted
- VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skippingaccepted
- V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondenceaccepted
- Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervisionaccepted
- VarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAMaccepted
- Variation-aware Vision Token Dropping for Faster Large Vision-Language Modelsaccepted
- Variational Graph-based Normal Integrationaccepted
- VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inferenceaccepted
- VecGlypher: Unified Vector Glyph Generation with Language Modelsaccepted
- Vector Prism: Animating Vector Graphics by Stratifying Semantic Structureaccepted
- VectorArk: Learning Practical Image Vectorization with Rounded Polygon Representationaccepted
- Velox: Learning Representations of 4D Geometry and Appearanceaccepted
- Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Croppingaccepted
- Verifying Neural Network Robustness with Dual Perturbationsaccepted
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Controlaccepted
- VesMamba: 3D Pulmonary Vessel Segmentation from CT images via Mamba with Structural Perception and Scale-aware Filteringaccepted
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Bodyaccepted
- ViHOI: Human-Object Interaction Synthesis with Visual Priorsaccepted
- ViKey: Enhancing Temporal Understanding in Videos via Visual Promptingaccepted
- ViLearn: Accelerating Training Convergence of Image-to-3D Generation via Visibility Learningaccepted
- ViLoMem: Agentic Learner with Grow-and-Refine Multimodal Semantic Memoryaccepted
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunkingaccepted
- ViStoryBench: Comprehensive Benchmark Suite for Story Visualizationaccepted
- ViT$^3$: Unlocking Test-Time Training in Visionaccepted
- ViTPrompt: Training-Free Prompt Refinement with Visual Tokens for Open-Vocabulary Detectionaccepted
- Vibe Spaces for Creatively Connecting and Expressing Visual Conceptsaccepted
- VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generationsaccepted
- VidEoMT: Your ViT is Secretly Also a Video Segmentation Modelaccepted
- VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transferaccepted
- VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scaleaccepted
- Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learneraccepted
- Video Panels for Long Video Understandingaccepted
- Video-CoE: Reinforcing Video Event Prediction via Chain of Eventsaccepted
- Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Modelsaccepted
- Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPOaccepted
- Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot Learningaccepted
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understandingaccepted
- VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twiceaccepted
- VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learningaccepted
- VideoCoF: Unified Video Editing with Temporal Reasoneraccepted
- VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusionaccepted
- VideoITG: Multimodal Video Understanding with Instructed Temporal Groundingaccepted
- VideoMaMa: Mask-Guided Video Matting via Generative Prioraccepted
- VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognitionaccepted
- VideoRealBench: A Chain-of-Thought Realism Evaluation Benchmark for Generated Human-Centric Videosaccepted
- VideoSSR: Video Self-Supervised Reinforcement Learningaccepted
- VideoSeek: Long-Horizon Video Agent with Tool-Guided Seekingaccepted
- VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agentsaccepted
- VideoWorld 2: Learning Transferable Knowledge from Real-world Videosaccepted
- View-Aware Semantic Alignment for Aerial-Ground Person Re-Identificationaccepted
- VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QAaccepted
- Vinedresser3D: Towards Agentic Text-guided 3D Editingaccepted
- Virtual Full-stack Scanning of Brain MRI via Imputing Any Quantised Codeaccepted
- Virtual Immunohistochemistry Staining with Dual-Aligned Multi-Task Feature Guidanceaccepted
- Virtual Nodes Guided Dynamic Graph Neural Network for Brain Tumor Segmentation with Missing Modalitiesaccepted
- VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understandingaccepted
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Modelsaccepted
- VisPlay: Self-Evolving Vision-Language Modelsaccepted
- VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Modelsaccepted
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsaccepted
- VisiLock: Authorizing Instruction-based Image editing with Dual Score Distillationaccepted
- Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Modelsaccepted
- Vision Transformers Need More Than Registersaccepted
- Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identificationaccepted
- Vision-Language Model Guided Source-Free Domain Adaptation via Optimal Transportaccepted
- Vision-Oriented Lightweight Neural Architecture Search with Budget-Adaptive Evaluationaccepted
- Vision-Speech Models: Teaching Speech Models to Converse about Imagesaccepted
- VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesisaccepted
- VisionLeaf: Entropy-Guided Leaf-First Reasoning for Efficient and Accurate Think-with-Imageaccepted
- Vista4D: Video Reshooting with 4D Point Cloudsaccepted
- Visual Diffusion Models are Geometric Solversaccepted
- Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scalingaccepted
- Visual Grounding for Object Questionsaccepted
- Visual Personalization Turing Testaccepted
- Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detectionaccepted
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Modelsaccepted
- Visual-RRT: Finding Paths toward Visual-Goals via Differentiable Renderingaccepted
- VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformeraccepted
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenesaccepted
- ViterbiPlanNet: Injecting Procedural Knowledge via Differentiable Viterbi for Planning in Instructional Videosaccepted
- VoDaSuRe: A Large-Scale Dataset Revealing Domain Shift in Volumetric Super-Resolutionaccepted
- Vocabulary Scaling Law: Tuning Open-vocabulary Predictors for Their Opennessaccepted
- Volumetric Functional Mapsaccepted
- VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentationaccepted
- Voxify3D: Pixel Art Meets Volumetric Renderingaccepted
- W2W: Language-Model-Based Trajectory Prediction with Reinforcement Learningaccepted
- WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Drivingaccepted
- WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generationaccepted
- WHU-MARS: A Multispectral Aerial-Ground Benchmark Towards Any-Scenario Person Re-Identificationaccepted
- WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrievalaccepted
- WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenariosaccepted
- WPT: World-to-Policy Transfer via Online World Model Distillationaccepted
- WRIVINDER: Towards Spatial Intelligence for Geo-locating Ground Images onto Satellite Imageryaccepted
- WaDi: Weight Direction-aware Distillation for One-step Image Synthesisaccepted
- WaTeRFlow: Watermark Temporal Robustness via Flow Consistencyaccepted
- WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigationaccepted
- Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Trainingaccepted
- Wanderland: Geometrically Grounded Simulation for Open-World Embodied AIaccepted
- Watch and Learn: Learning to Use Computers from Online Videosaccepted
- Wave-Former: Through-Occlusion 3D Reconstruction via Wireless Shape Completionaccepted
- Wavelet-Driven 3D Anomaly Detection under Pose-Agnostic and Sparse-Viewaccepted
- Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understandingaccepted
- WeDetect: Fast Open-Vocabulary Object Detection as Retrievalaccepted
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokensaccepted
- Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoningaccepted
- WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformationaccepted
- WeaveTime: Streaming from Earlier Frames into Emergent Memory in VideoLLMsaccepted
- WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Tracesaccepted
- WebGym: Scaling Training Environments for Long-Horizon Visual Web Agents with Realistic Tasksaccepted
- Weight Space Representation Learning via Neural Field Adaptationaccepted
- What Are You Doing? A Closer Look at Controllable Human Video Generationaccepted
- What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Modelsaccepted
- What Is It Like to Be a Noise? An Entropy-based Gaussian Noise Regularization for Diffusion Modelsaccepted
- What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is Rarely F1accepted
- What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?accepted
- What Matters in Practical Learned Image Compressionaccepted
- What Your Features Reveal: Data-Efficient Black-Box Feature Inversion Attack for Split DNNsaccepted
- What's Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolutionaccepted
- When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapseaccepted
- When Anonymity Breaks: Identifying Models Behind Text-to-Image Leaderboardsaccepted
- When CLIP Sees More, It Fights Back Harder: Multi-View Guided Adaptive Counterattacks for Test-Time Adversarial Robustnessaccepted
- When Do Models Actually Decide? Mapping the Layer-Wise Decision Timeline in Pretrained Neural Networksaccepted
- When Lines Meet Textures: Spatial-Frequency Aligned Diffusion Features for Cross-Sparsity Correspondenceaccepted
- When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adaptersaccepted
- When Local Rules Create Global Order: Self-Organized Representation Learning for Latent Diffusion Modelsaccepted
- When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Modelsaccepted
- When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generatorsaccepted
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Modelsaccepted
- When Robots Should Say ''I Don't Know'': Benchmarking Abstention in Embodied Question Answeringaccepted
- When Safety Collides: Resolving Multi-Category Harmful Conflicts in Text-to-Image Diffusion via Adaptive Safety Guidanceaccepted
- When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMsaccepted
- When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detectionaccepted
- When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigmaccepted
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thoughtaccepted
- When to Think and When to Look: Uncertainty-Guided Lookbackaccepted
- Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generationaccepted
- Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attentionaccepted
CVPR accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.