ICCV 2025 Accepted Papers
The full list of 2,620 papers accepted at ICCV 2025 (IEEE/CVF International Conference on Computer Vision). Click any title for details, similar papers, and links to the original source. You can also search these papers by meaning, not just keywords.
Poster: 2,595
- Generalized Few-Shot Point Cloud Segmentation via LLM-Assisted Hyper-Relation MatchingPoster
- Generalized Tensor-based Parameter-Efficient Fine-Tuning via Lie Group TransformationsPoster
- Generalized and Efficient 2D Gaussian Splatting for Arbitrary-scale Super-ResolutionPoster
- Generate, Refine, and Encode: Leveraging Synthesized Novel Samples for On-the-Fly Fine-Grained Category DiscoveryPoster
- Generate, Transduct, Adapt: Iterative Transduction with VLMsPoster
- Generating Multi-Image Synthetic Data for Text-to-Image CustomizationPoster
- Generating Physically Stable and Buildable Brick Structures from TextPoster
- Generating, Fast and Slow: Scalable Parallel Video Generation with Video Interface NetworksPoster
- Generative Active Learning for Long-tail Trajectory Prediction via Controllable Diffusion ModelPoster
- Generative Adversarial DiffusionPoster
- Generative Gaussian Splatting: Generating 3D Scenes with Video Diffusion PriorsPoster
- Generative Modeling of Shape-Dependent Self-Contact Human PosesPoster
- Generative Video Bi-flowPoster
- Generative ZooPoster
- Generic Event Boundary Detection via Denoising DiffusionPoster
- GenieBlue: Integrating both Linguistic and Multimodal Capabilities for Large Language Models on Mobile DevicesPoster
- Geo4D: Leveraging Video Generators for Geometric 4D Scene ReconstructionPoster
- GeoAvatar: Adaptive Geometrical Gaussian Splatting for 3D Head AvatarPoster
- GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image GenerationPoster
- GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View LocalizationPoster
- GeoExplorer: Active Geo-localization with Curiosity-Driven ExplorationPoster
- GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-based TransformerPoster
- GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video DiffusionPoster
- GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language FieldsPoster
- GeoSplatting: Towards Geometry Guided Gaussian Splatting for Physically-based Inverse RenderingPoster
- Geometric Alignment and Prior Modulation for View-Guided Point Cloud Completion on Unseen CategoriesPoster
- Geometry DistributionsPoster
- GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion PriorsPoster
- GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented GenerationPoster
- GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal ModelingPoster
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image GenerationPoster
- GlassWizard: Harvesting Diffusion Priors for Glass Surface DetectionPoster
- GloPER: Unsupervised Animal Pattern Extraction from Local ReconstructionPoster
- Global Motion Corresponder for 3D Point-Based Scene Interpolation under Large MotionPoster
- Global Regulation and Excitation via Attention Tuning for Stereo MatchingPoster
- Global and Local Entailment Learning for Natural World ImageryPoster
- Global-Aware Monocular Semantic Scene Completion with State Space ModelsPoster
- Go to Zero: Towards Zero-shot Motion Generation with Million-scale DataPoster
- Golden Noise for Diffusion Models: A Learning FrameworkPoster
- Gradient Decomposition and Alignment for Incremental Object DetectionPoster
- Gradient Extrapolation for Debiased Representation LearningPoster
- Gradient Short-Circuit: Efficient Out-of-Distribution Detection via Feature InterventionPoster
- Gradient-Reweighted Adversarial Camouflage for Physical Object Detection EvasionPoster
- Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept RepresentationsPoster
- Graph Domain Adaptation with Dual-branch Encoder and Two-level Alignment for Whole Slide Image-based Survival PredictionPoster
- GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language InstructionsPoster
- Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-ReferringPoster
- GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential GroundingPoster
- GroundingSuite: Measuring Complex Multi-Granular Pixel GroundingPoster
- Group Inertial Poser: Multi-Person Pose and Global Translation from Sparse Inertial Sensors and Ultra-Wideband Ranging
- Group-wise Scaling and Orthogonal Decomposition for Domain-Invariant Feature Extraction in Face Anti-SpoofingPoster
- Grouped Speculative Decoding for Autoregressive Image GenerationPoster
- Growing a Twig to Accelerate Large Vision-Language ModelsPoster
- Guiding Diffusion Models with Adaptive Negative Sampling Without External ResourcesPoster
- Guiding Diffusion-Based Articulated Object Generation by Partial Point Cloud Alignment and Physical Plausibility ConstraintsPoster
- Guiding Noisy Label Conditional Diffusion Models with Score-based Discriminator CorrectionPoster
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D ReconstructionPoster
- HADES: Human Avatar with Dynamic Explicit Hair StrandsPoster
- HAMSt3R: Human-Aware Multi-view Stereo 3D ReconstructionPoster
- HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReIDPoster
- HDR Image Generation via Gain Map Decomposed DiffusionPoster
- HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and GenerationPoster
- HERMES: temporal-coHERent long-forM understanding with Episodes and SemanticsPoster
- HERO: Human Reaction Generation from VideosPoster
- HFD-Teacher: High-Frequency Depth Distillation from Depth Foundation Models for Enhanced Depth CompletionPoster
- HIS-GPT: Towards 3D Human-In-Scene Multimodal UnderstandingPoster
- HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature AdaptationPoster
- HOMO-Feature: Cross-Arbitrary-Modal Image Matching with Homomorphism of Organized Major OrientationPoster
- HORT: Monocular Hand-held Objects Reconstruction with TransformersPoster
- HPSv3: Towards Wide-Spectrum Human Preference ScorePoster
- HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP ModelsPoster
- HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?Poster
- HUG: Hierarchical Urban Gaussian Splatting with Block-Based Reconstruction for Large-Scale Aerial ScenesPoster
- HUMOTO: A 4D Dataset of Mocap Human Object InteractionsPoster
- HUST: High-Fidelity Unbiased Skin Tone Estimation via Texture QuantizationPoster
- HVPUNet: Hybrid-Voxel Point-cloud Upsampling NetworkPoster
- HairCUP: Hair Compositional Universal Prior for 3D Gaussian AvatarsPoster
- Hallucinatory Image Tokens: A Training-free EAZY Approach to Detecting and Mitigating Object Hallucinations in LVLMsPoster
- Harmonizing Visual Representations for Unified Multimodal Understanding and GenerationPoster
- HarmonySeg: Tubular Structure Segmentation with Deep-Shallow Feature Fusion and Growth-Suppression Balanced LossPoster
- Harnessing Input-Adaptive Inference for Efficient VLNPoster
- Harnessing Massive Satellite Imagery with Efficient Masked Image ModelingPoster
- Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised LearningPoster
- Harnessing Uncertainty-aware Bounding Boxes for Unsupervised 3D Object DetectionPoster
- Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary SegmentationPoster
- Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful IllusionsPoster
- HazeFlow: Revisit Haze Physical Model as ODE and Non-Homogeneous Haze Generation for Real-World DehazingPoster
- HccePose(BF): Predicting Front & Back Surfaces to Construct Ultra-Dense 2D-3D Correspondences for Pose EstimationPoster
- Head2Body: Body Pose Generation from Multi-sensory Head-mounted InputsPoster
- Heatmap Regression without Soft-Argmax for Facial Landmark DetectionPoster
- Heavy Labels Out! Dataset Distillation with Label Space LighteningPoster
- Height-Fidelity Dense Global Fusion for Multi-modal 3D Object DetectionPoster
- Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language ModelsPoster
- Hi-Gaussian: Hierarchical Gaussians under Normalized Spherical Projection for Single-View 3D ReconstructionPoster
- Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal BridgingPoster
- HiERO: Understanding the Hierarchy of Human Behavior Enhances Reasoning on Egocentric VideosPoster
- HiGarment: Cross-modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment ImagePoster
- HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal ModelPoster
- HiNeuS: High-fidelity Neural Surface Mitigating Low-texture and Reflective AmbiguityPoster
- HiP-AD: Hierarchical and Multi-Granularity Planning with Deformable Attention for Autonomous Driving in a Single DecoderPoster
- Hierarchical 3D Scene Graphs Construction OutdoorsPoster
- Hierarchical Cross-modal Prompt Learning for Vision-Language ModelsPoster
- Hierarchical Divide-and-Conquer Grouping for Classification Adaptation of Pre-Trained ModelsPoster
- Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal GroundingPoster
- Hierarchical Material Recognition from Local AppearancePoster
- Hierarchical Variational Test-Time Prompt Generation for Zero-Shot GeneralizationPoster
- Hierarchical Visual Prompt Learning for Continual Video Instance SegmentationPoster
- Hierarchical-aware Orthogonal Disentanglement Framework for Fine-grained Skeleton-based Action RecognitionPoster
- Hierarchy UGP: Hierarchy Unified Gaussian Primitive for Large-Scale Dynamic Scene ReconstructionPoster
- Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image RetrievalPoster
- High-Precision 3D Measurement of Complex Textured Surfaces Using Multiple Filtering ApproachPoster
- High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose EstimationPoster
- Highlight What You Want: Weakly-Supervised Instance-Level Controllable Infrared-Visible Image FusionPoster
- Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous DrivingPoster
- Hipandas: Hyperspectral Image Joint Denoising and Super-Resolution by Image Fusion with the Panchromatic ImagePoster
- HoliTracer: Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing ImageryPoster
- Holistic Tokenizer for Autoregressive Image GenerationPoster
- Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model UnlearningPoster
- HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion ModelsPoster
- HouseTour: A Virtual Real Estate A(I)gentPoster
- How Can Objects Help Video-Language Understanding?Poster
- How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in An Extensible Escape GamePoster
- How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?Poster
- How Far are AI-generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation ApproachPoster
- How To Make Your Cell Tracker Say "I dunno!"Poster
- How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor ScenesPoster
- Human-Object Interaction from Human-Level InstructionsPoster
- Human-in-the-Loop Local Corrections of 3D Scene Layouts via InfillingPoster
- HumanOLAT: A Large-Scale Dataset for Full-Body Human Relighting and Novel-View SynthesisPoster
- HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion AnomalyPoster
- Humans as Checkerboards: Calibrating Camera Motion Scale for World-Coordinate Human Mesh RecoveryPoster
- Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated VideosPoster
- HumorDB: Can AI understand graphical humor?Poster
- HyPiDecoder: Hybrid Pixel Decoder for Efficient Segmentation and DetectionPoster
- HyTIP: Hybrid Temporal Information Propagation for Masked Conditional Residual Video CodingPoster
- Hybrid Layout Control for Diffusion Transformer: Fewer Annotations, Superior AestheticsPoster
- Hybrid-TTA: Continual Test-time Adaptation via Dynamic Domain Shift DetectionPoster
- Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth EstimationPoster
- Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop TrainingPoster
- HypDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-shot Image GenerationPoster
- Hyper-Depth: Hypergraph-based Multi-Scale Representation Fusion for Monocular Depth EstimationPoster
- HyperGCT: A Dynamic Hyper-GNN-Learned Geometric Constraint for 3D RegistrationPoster
- Hypergraph Clustering Network with Partial Attribute ImputationPoster
- I Am Big, You Are Little; I Am Right, You Are WrongPoster
- I2-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene ForecastingPoster
- I2V3D: Controllable Image-to-video Generation with 3D GuidancePoster
- I2VControl: Disentangled and Unified Video Motion Synthesis ControlPoster
- IAP: Invisible Adversarial Patch Attack through Perceptibility-Aware Localization and Perturbation OptimizationPoster
- ICE-Bench: A Unified and Comprehensive Benchmark for Image Creating and EditingPoster
- IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using ThemselvesPoster
- IDF: Iterative Dynamic Filtering Networks for Generalizable Image DenoisingPoster
- IDFace: Face Template Protection for Efficient and Secure IdentificationPoster
- IFAdapter: Instance Feature Control for Grounded Text-to-Image GenerationPoster
- IGD: Instructional Graphic Design with Multimodal Layer GenerationPoster
- IGL-Nav: Incremental 3D Gaussian Localization for Image-goal NavigationPoster
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-EnhancePoster
- IM-LUT: Interpolation Mixing Look-Up Tables for Image Super-ResolutionPoster
- IM360: Large-scale Indoor Mapping with 360 CamerasPoster
- IMG: Calibrating Diffusion Models via Implicit Multimodal GuidancePoster
- IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&APoster
- INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in InsurancePoster
- INSTINCT: Instance-Level Interaction Architecture for Query-Based Collaborative PerceptionPoster
- INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance SamplingPoster
- IQA-Adapter: Exploring Knowledge Transfer from Image Quality Assessment to Diffusion-based Generative ModelsPoster
- IRASim: A Fine-Grained World Model for Robot ManipulationPoster
- IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale BenchmarkPoster
- ISP2HRNet: Learning to Reconstruct High Resolution Image from Irregularly Sampled Pixels via Hierarchical Gradient LearningPoster
- Identity Preserving 3D Head Stylization with Multiview Score DistillationPoster
- Identity-aware Language Gaussian Splatting for Open-vocabulary 3D Semantic SegmentationPoster
- Im2Haircut: Single-view Strand-based Hair Reconstruction for Human AvatarsPoster
- ImHead: A Large-scale Implicit Morphable Model for Localized Head ModelingPoster
- Image Intrinsic Scale Assessment: Bridging the Gap Between Quality and ResolutionPoster
- Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred ImagePoster
- Image-Guided Shape-from-Template Using Mesh Inextensibility ConstraintsPoster
- ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model PersonalizationPoster
- ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought ReasoningPoster
- Images as Noisy Labels: Unleashing the Potential of the Diffusion Model for Open-Vocabulary Semantic SegmentationPoster
- Imbalance in Balance: Online Concept Balancing in Generation Models
- Implicit Counterfactual Learning for Audio-Visual SegmentationPoster
- Importance-Based Token Merging for Efficient Image and Video GenerationPoster
- Improved Noise Schedule for Diffusion TrainingPoster
- Improving Large Vision and Language Models by Learning from a Panel of PeersPoster
- Improving Multimodal Learning via Imbalanced LearningPoster
- Improving Noise Efficiency in Privacy-preserving Dataset DistillationPoster
- Improving Rectified Flow with Boundary ConditionsPoster
- Improving SAM for Camouflaged Object Detection via Dual Stream AdaptersPoster
- Incremental Few-Shot Semantic Segmentation via Multi-Level Switchable Visual PromptsPoster
- InfGen: A Resolution-Agnostic Paradigm for Scalable Image SynthesisPoster
- Inference-Time Diffusion Model DistillationPoster
- InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video ModelsPoster
- InfiniDreamer: Arbitrarily Long Human Motion Generation via Segment Score DistillationPoster
- InfiniteYou: Flexible Photo Recrafting While Preserving Your IdentityPoster
- InfoBridge: Balanced Multimodal Integration through Conditional Dependency ModelingPoster
- Information Density Principle for MLLM BenchmarksPoster
- Information-Bottleneck Driven Binary Neural Network for Change DetectionPoster
- Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional WarpingPoster
- InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset ConstructionPoster
- InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video GenerationPoster
- InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered ScenesPoster
- Instance-Level Video Depth in Groups Beyond OcclusionsPoster
- Instant GaussianImage: A Generalizable and Self-Adaptive Image Representation via 2D Gaussian SplattingPoster
- InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified FlowPoster
- InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language ModelsPoster
- Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language ModelsPoster
- Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMsPoster
- Instruction-based Image Editing with Planning, Reasoning, and GenerationPoster
- Integrating Biological Knowledge for Robust Microscopy Image Profiling on De Novo Cell LinesPoster
- Integrating Task-Specific and Universal Adapters for Pre-Trained Model-based Class-Incremental LearningPoster
- Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem SolvingPoster
- Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive SegmentationPoster
- InterGSEdit: Interactive 3D Gaussian Splatting Editing with 3D Geometry-Consistent Attention PriorPoster
- InteractAvatar: Modeling Hand-Face Interaction in Photorealistic Avatars with Deformable GaussiansPoster
- Interaction-Merged Motion Planning: Effectively Leveraging Diverse Motion Datasets for Robust PlanningPoster
- Intermediate Connectors and Geometric Priors for Language-Guided Affordance Segmentation on Unseen Object CategoriesPoster
- Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language ModelPoster
- Interpretable point cloud classification using multiple instance learningPoster
- Intervening in Black Box: Concept Bottleneck Model for Enhancing Human Neural Network Mutual UnderstandingPoster
- Intra-modal and Cross-modal Synchronization for Audio-visual Deepfake Detection and Temporal LocalizationPoster
- Intra-view and Inter-view Correlation Guided Multi-view Novel Class DiscoveryPoster
- IntrinsicControlNet: Cross-distribution Image Generation with Real and UnrealPoster
- IntroStyle: Training-Free Introspective Style Attribution using Diffusion FeaturesPoster
- InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance ModelingPoster
- Inverse 3D Microscopy Rendering for Cell Shape Inference with Active MeshPoster
- Inverse Image-Based Rendering for Light Field Generation from Single ImagesPoster
- Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking DesignPoster
- Iris: Breaking GUI Complexity with Adaptive Focus and Self-RefiningPoster
- Is CLIP ideal? No. Can we fix it? Yes!Poster
- Is Less More? Exploring Token Condensation as Training-free Test-time AdaptationPoster
- Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited EntropyPoster
- Is Tracking Really More Challenging in First Person Egocentric Vision?Poster
- Is Visual in-Context Learning for Compositional Medical Tasks within Reach?Poster
- JPEG Processing Neural Operator for Backward-Compatible CodingPoster
- JailbreakDiffBench: A Comprehensive Benchmark for Jailbreaking Diffusion ModelsPoster
- Jailbreaking Multimodal Large Language Models via Shuffle InconsistencyPoster
- Jigsaw++: Imagining Complete Shape Priors for Object ReassemblyPoster
- Joint Asymmetric Loss for Learning with Noisy LabelsPoster
- Joint Diffusion Models in Continual LearningPoster
- Joint Learning of Pose Regression and Denoising Diffusion with Score Scaling Sampling for Category-level 6D Pose EstimationPoster
- Joint Self-Supervised Video Alignment and Action SegmentationPoster
- Joint Semantic and Rendering Enhancements in 3D Gaussian Modeling with Anisotropic Local EncodingPoster
- JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion TransformersPoster
- KOEnsAttack: Towards Efficient Data-Free Black-Box Adversarial Attacks via Knowledge-Orthogonalized Substitute EnsemblesPoster
- KV-Edit: Training-Free Image Editing for Precise Background PreservationPoster
- Kaleidoscopic Background Attack: Disrupting Pose Estimation with Multi-Fold Radial Symmetry TexturesPoster
- Kaputt: A Large-Scale Dataset for Visual Defect DetectionPoster
- Keep Your Friends Close, and Your Enemies Farther: Distance-aware Voxel-wise Contrastive Learning for Semi-supervised Multi-organ SegmentationPoster
- Kestrel: 3D Multimodal LLM for Part-Aware Grounded DescriptionPoster
- Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video ProcessingPoster
- KinMo: Kinematic-aware Human Motion Understanding and GenerationPoster
- Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIPPoster
- Know Your Attention Maps: Class-specific Token Masking for Weakly Supervised Semantic SegmentationPoster
- Knowledge Distillation for Learned Image CompressionPoster
- Knowledge Distillation with Refined LogitsPoster
- Knowledge Transfer from Interaction LearningPoster
- LA-MOTR: End-to-End Multi-Object Tracking by Learnable AssociationPoster
- LACONIC: A 3D Layout Adapter for Controllable Image CreationPoster
- LANGTRAJ: Diffusion Model and Dataset for Language-Conditioned Trajectory SimulationPoster
- LATINO-PRO: LAtent consisTency INverse sOlver with PRompt OptimizationPoster
- LBM: Latent Bridge Matching for Fast Image-to-Image TranslationPoster
- LDIP: Long Distance Information Propagation for Video Super-ResolutionPoster
- LDPose: Towards Inclusive Human Pose Estimation for Limb-Deficient Individuals in the WildPoster
- LEGION: Learning to Ground and Explain for Synthetic Image DetectionPoster
- LEGO-Maker: A Semantic-Driven Algorithm for Text-to-3D GenerationPoster
- LGA-Net: Learning Local and Global Affinities for Sparse Scribble based Image ColorizationPoster
- LHM: Large Animatable Human Reconstruction Model for Single Image to 3D in SecondsPoster
- LIFT: Latent Implicit Functions for Task- and Data-Agnostic EncodingPoster
- LINR-PCGC: Lossless Implicit Neural Representations for Point Cloud Geometry CompressionPoster
- LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region AssistancePoster
- LIRA: Reasoning Reconstruction via Multimodal Large Language ModelsPoster
- LLM Thought Divergence and Convergence for Dialogue-Based Image Generation ControlPoster
- LLM-Assisted Semantic Guidance for Sparsely Annotated Remote Sensing Object DetectionPoster
- LLM-assisted Entropy-based Adaptive Distillation for Unsupervised Fine-grained Visual Representation LearningPoster
- LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text MatchingPoster
- LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMsPoster
- LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D CapabilitiesPoster
- LLaVA-CoT: Let Vision Language Models Reason Step-by-StepPoster
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsPoster
- LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsPoster
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMsPoster
- LMM-Det: Make Large Multimodal Models Excel in Object DetectionPoster
- LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image EditingPoster
- LOMM: Latest Object Memory Management for Temporally Consistent Video Instance SegmentationPoster
- LONG3R: Long Sequence Streaming 3D ReconstructionPoster
- LOTA: Bit-Planes Guided AI-Generated Image DetectionPoster
- LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text PairingPoster
- LUDVIG: Learning-Free Uplifting of 2D Visual Features to Gaussian Splatting ScenesPoster
- LUSD: Localized Update Score Distillation for Text-Guided Image EditingPoster
- LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up TablesPoster
- LV-MAE: Learning Long Video Representations through Masked-Embedding AutoencodersPoster
- LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM AgentsPoster
- LVBench: An Extreme Long Video Understanding BenchmarkPoster
- LVFace: Progressive Cluster Optimization for Large Vision Models in Face RecognitionPoster
- LaCoOT: Layer Collapse through Optimal TransportPoster
- LaRender: Training-Free Occlusion Control in Image Generation via Latent RenderingPoster
- Laboring on less labors: RPCA Paradigm for Pan-sharpeningPoster
- LaneDiffusion: Improving Centerline Graph Learning via Prior Injected BEV Feature GenerationPoster
- LangBridge: Interpreting Image as a Combination of Language EmbeddingsPoster
- LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video DiffusionPoster
- Language Decoupling with Fine-grained Knowledge Guidance for Referring Multi-object TrackingPoster
- Language Driven Occupancy PredictionPoster
- Language-Driven Multi-Label Zero-Shot Learning with Semantic GranularityPoster
- Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and CompressibilityPoster
- Large Multi-modal Models Can Interpret Features in Large Multi-modal ModelsPoster
- Large Scene Generation with Cube-Absorb Discrete DiffusionPoster
- Large-scale Pre-training for Grounded Video Caption GenerationPoster
- Lark: Low-Rank Updates After Knowledge Localization for Few-shot Class-Incremental LearningPoster
- Latent Diffusion Models with Masked AutoEncodersPoster
- Latent Expression Generation for Referring Image Segmentation and GroundingPoster
- Latent Swap Joint Diffusion for 2D Long-Form Latent GenerationPoster
- Latent-Reframe: Enabling Camera Control for Video Diffusion Models without TrainingPoster
- Latte: Collaborative Test-Time Adaptation of Vision-Language Models in Federated LearningPoster
- LawDIS: Language-Window-based Controllable Dichotomous Image SegmentationPoster
- Lay-Your-Scene: Natural Scene Layout Generation with Diffusion TransformersPoster
- Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story GenerationPoster
- LayerAnimate: Layer-level Control for AnimationPoster
- LayerD: Decomposing Raster Graphic Designs into LayersPoster
- LayerLock: Non-collapsing Representation Learning with Progressive FreezingPoster
- LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion TransformerPoster
- LazyMAR: Accelerating Masked Autoregressive Models via Feature CachingPoster
- LeGrad: An Explainability Method for Vision Transformers via Feature Formation SensitivityPoster
- LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion ModelsPoster
- Leaps and Bounds: An Improved Point Cloud Winding Number Formulation for Fast Normal Estimation and Surface ReconstructionPoster
- Learn2Synth: Learning Optimal Data Synthesis Using Hypergradients for Brain Image SegmentationPoster
- Learnable Feature Patches and Vectors for Boosting Low-light Image Enhancement without External KnowledgePoster
- Learnable Fractional Reaction-Diffusion Dynamics for Under-Display ToF Imaging and BeyondPoster
- Learnable Logit Adjustment for Imbalanced Semi-Supervised Learning under Class Distribution MismatchPoster
- Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report GenerationPoster
- Learned Image Compression with Hierarchical Progressive Context ModelingPoster
- Learning 3D Object Spatial Relationships from Pre-trained 2D Diffusion ModelsPoster
- Learning 3D Scene Analogies with Neural Contextual Scene MapsPoster
- Learning 4D Embodied World ModelsPoster
- Learning A Unified Template for Gait RecognitionPoster
- Learning Beyond Still Frames: Scaling Vision-Language Models with VideoPoster
- Learning Deblurring Texture Prior from Unpaired Data with Diffusion ModelPoster
- Learning Dense Feature Matching via Lifting Single 2D Image to 3D SpacePoster
- Learning Efficient and Generalizable Human Representation with Human Gaussian ModelPoster
- Learning Few-Step Diffusion Models by Trajectory Distribution MatchingPoster
- Learning Hierarchical Line Buffer for Image ProcessingPoster
- Learning Implicit Features with Flow-Infused Transformations for Realistic Virtual Try-OnPoster
- Learning Interpretable Queries for Explainable Image Classification with Information PursuitPoster
- Learning Large Motion Estimation from Intermediate Representations with a High-Resolution Optical Flow Dataset Featuring Long-Range Dynamic MotionPoster
- Learning Neural Scene Representation from iToF ImagingPoster
- Learning Normal Flow Directly From EventsPoster
- Learning Normals of Noisy Points by Local Gradient-Aware Surface FilteringPoster
- Learning Null Geodesics for Gravitational Lensing Rendering in General RelativityPoster
- Learning Pixel-adaptive Multi-layer Perceptrons for Real-time Image EnhancementPoster
- Learning Precise Affordances from Egocentric Videos for Robotic ManipulationPoster
- Learning Robust Image Watermarking with Lossless Cover RecoveryPoster
- Learning Robust Stereo Matching in the Wild with Selective Mixture-of-ExpertsPoster
- Learning Separable Fine-Grained Representation via Dendrogram Construction from Coarse Labels for Fine-grained Visual RecognitionPoster
- Learning Streaming Video Representation via Multitask TrainingPoster
- Learning Visual Hierarchies in Hyperbolic Space for Image RetrievalPoster
- Learning Visual Proxy for Compositional Zero-Shot LearningPoster
- Learning Yourself: Class-Incremental Semantic Segmentation with Language-Inspired Bootstrapped DisentanglementPoster
- Learning an Implicit Physics Model for Image-based Fluid SimulationPoster
- Learning on the Go: A Meta-learning Object Navigation ModelPoster
- Learning to Generalize without Bias for Open-Vocabulary Action RecognitionPoster
- Learning to Inference Adaptively for Multimodal Large Language ModelsPoster
- Learning to See Inside Opaque Liquid Containers using Speckle VibrometryPoster
- Learning to See in the Extremely DarkPoster
- Learning to Unlearn while Retaining: Combating Gradient Conflicts in Machine UnlearningPoster
- Less Static, More Private: Towards Transferable Privacy-Preserving Action Recognition by Generative Decoupled LearningPoster
- Less is More: Empowering GUI Agent with Context-Aware SimplificationPoster
- Less is More: Improving Motion Diffusion Models with Sparse KeyframesPoster
- Less-to-More Generalization: Unlocking More Controllability by In-Context GenerationPoster
- Leveraging 2D Priors and SDF Guidance for Urban Scene RenderingPoster
- Leveraging BEV Paradigm for Ground-to-Aerial Image SynthesisPoster
- Leveraging Debiased Cross-modal Attention Maps and Code-based Reasoning for Zero-shot Referring Expression ComprehensionPoster
- Leveraging Local Patch Alignment to Seam-cutting for Large Parallax Image StitchingPoster
- Leveraging Panoptic Scene Graph for Evaluating Fine-Grained Text-to-Image GenerationPoster
- Leveraging Prior Knowledge of Diffusion Model for Person SearchPoster
- Leveraging Spatial Invariance to Boost Adversarial TransferabilityPoster
- Leveraging the Power of MLLMs for Gloss-Free Sign Language TranslationPoster
- LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video DiffusionPoster
- LiT: Delving into a Simple Linear Diffusion Transformer for Image GenerationPoster
- Liberated-GS: 3D Gaussian Splatting Independent from SfM Point CloudsPoster
- Lidar Waveforms are Worth 40x128x33 WordsPoster
- Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary ModelingPoster
- LightBSR: Towards Lightweight Blind Super-Resolution via Discriminative Implicit Degradation Representation LearningPoster
- LightCity: An Urban Dataset for Outdoor Inverse Rendering and Reconstruction under Multi-illumination ConditionsPoster
- LightSwitch: Multi-view Relighting with Material-guided DiffusionPoster
- LightsOut: Diffusion-based Outpainting for Enhanced Lens Flare RemovalPoster
- Lightweight Gradient-Aware Upscaling of 3D Gaussian Splatting ImagesPoster
- Lightweight and Fast Real-time Image Enhancement via Decomposition of the Spatial-aware Lookup TablesPoster
- LoD-Loc v2: Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette AlignmentPoster
- LoRA-FAIR: Federated LoRA Fine-Tuning with Aggregation and Initialization RefinementPoster
- LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image GenerationPoster
- LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion ModelsPoster
- Local Scale Equivariance with Latent Deep Equilibrium CanonicalizerPoster
- LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature DecouplingPoster
- LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation ModelsPoster
- Long Context Tuning for Video GenerationPoster
- Long-Context State-Space Video World ModelsPoster
- Long-LRM: Long-sequence Large Reconstruction Model for Wide-coverage Gaussian SplatsPoster
- Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario GenerationPoster
- LongAnimation: Long Animation Generation with Dynamic Global-Local MemoryPoster
- LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long VideosPoster
- LookOut: Real-World Humanoid Egocentric NavigationPoster
- Looking in the Mirror: A Faithful Counterfactual Explanation Method for Interpreting Deep Image Classification ModelsPoster
- Loss Functions for Predictor-based Neural Architecture SearchPoster
- Low-Light Image Enhancement Using Event-Based Illumination EstimationPoster
- Lumina-Image 2.0: A Unified and Efficient Image Generative FrameworkPoster
- Lyra: An Efficient and Speech-Centric Framework for Omni-CognitionPoster
- M-Net: MRI Brain Tumor Sequential Segmentation Network via Mesh-CastPoster
- M-SpecGene: Generalized Foundation Model for RGBT Multispectral VisionPoster
- M2EIT: Multi-Domain Mixture of Experts for Robust Neural Inertial TrackingPoster
- M2SFormer: Multi-Spectral and Multi-Scale Attention with Edge-Aware Difficulty Guidance for Image Forgery LocalizationPoster
- MA-CIR: A Multimodal Arithmetic Benchmark for Composed Image RetrievalPoster
- MAESTRO: Task-Relevant Optimization via Adaptive Feature Enhancement and Suppression for Multi-task 3D PerceptionPoster
- MATE: Motion-Augmented Temporal Consistency for Event-based Point TrackingPoster
- MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual TranslationPoster
- MAVias: Mitigate any Visual BiasPoster
- MBTI: Masked Blending Transformers with Implicit Positional Encoding for Frame-rate Agnostic Motion EstimationPoster
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsPoster
- MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video UnderstandingPoster
- MCID: Multi-aspect Copyright Infringement Detection for Generated ImagesPoster
- MDD: A Dataset for Text-and-Music Conditioned Duet Dance GenerationPoster
- MDP-Omni: Parameter-free Multimodal Depth Prior-based Sampling for Omnidirectional Stereo MatchingPoster
- MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMsPoster
- MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic ScenesPoster
- MEH: A Multi-Style Dataset and Toolkit for Advancing Egyptian Hieroglyph RecognitionPoster
- MEMFOF: High-Resolution Training for Memory-Efficient Multi-Frame Optical Flow EstimationPoster
- METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language ModelsPoster
- MGSR: 2D/3D Mutual-boosted Gaussian Splatting for High-fidelity Surface Reconstruction under Various Light ConditionsPoster
- MGSfM: Multi-Camera Geometry Driven Global Structure-from-MotionPoster
- MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video CodingPoster
- MIEB: Massive Image Embedding BenchmarkPoster
- MINERVA: Evaluating Complex Video ReasoningPoster
- MIORe & VAR-MIORe: Benchmarks to Push the Boundaries of RestorationPoster
- MM-IFEngine: Towards Multimodal Instruction FollowingPoster
- MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMsPoster
- MMAD: Multi-label Micro-Action Detection in VideosPoster
- MMAIF: Multi-task and Multi-degradation All-in-One for Image Fusion with Language GuidancePoster
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent TuningPoster
- MMCR: Benchmarking Cross-Source Reasoning in Scientific PapersPoster
- MMGeo: Multimodal Compositional Geo-Localization for UAVsPoster
- MMOne: Representing Multiple Modalities in One ScenePoster
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGIPoster
- MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder PruningPoster
- MOERL: When Mixture-of-Experts Meet Reinforcement Learning for Adverse Weather Image RestorationPoster
- MOSAIC: Generating Consistent, Privacy-Preserving Scenes from Multiple Depth Views in Multi-Room EnvironmentsPoster
- MOSCATO: Predicting Multiple Object State Change Through ActionsPoster
- MOVE: Motion-Guided Few-Shot Video Object SegmentationPoster
- MP-HSIR: A Multi-Prompt Framework for Universal Hyperspectral Image RestorationPoster
- MPBR: Multimodal Progressive Bidirectional Reasoning for Open-Set Fine-Grained RecognitionPoster
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationPoster
- MR-FIQA: Face Image Quality Assessment with Multi-Reference Representations from Synthetic Data GenerationPoster
- MRGen: Segmentation Data Engine For Underrepresented MRI ModalitiesPoster
- MS3D: High-Quality 3D Generation via Multi-Scale Representation ModelingPoster
- MSA2: Multi-task Framework with Structure-aware and Style-adaptive Character Representation for Open-set Chinese Text RecognitionPoster
- MSQ: Memory-Efficient Bit Sparsification QuantizationPoster
- MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video ParsingPoster
- MUNBa: Machine Unlearning via Nash BargainingPoster
- MUSE-VL: Modeling Unified VLM through Semantic Discrete EncodingPoster
- MUSE: Multi-Subject Unified Synthesis via Explicit Layout Semantic ExpansionPoster
- MV-Adapter: Multi-View Consistent Image Generation Made EasyPoster
- MVGBench: a Comprehensive Benchmark for Multi-view Generation ModelsPoster
- MVQA: Mamba with Unified Sampling for Efficient Video Quality AssessmentPoster
- MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance CostPoster
- MaGS: Reconstructing and Simulating Dynamic 3D Objects with Mesh-adsorbed Gaussian SplattingPoster
- MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language ModelingPoster
- MagShield: Towards Better Robustness in Sparse Inertial Motion Capture Under Magnetic DisturbancesPoster
- Magic Insert: Style-Aware Drag-and-DropPoster
- MagicCity: Geometry-Aware 3D City Generation from Satellite Imagery with Multi-View ConsistencyPoster
- MagicColor: Multi-Instance Sketch ColorizationPoster
- MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive ControlPoster
- MagicHOI: Leveraging 3D Priors for Accurate Hand-object Reconstruction from Short Monocular Video ClipsPoster
- MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video CustomizationPoster
- MagicMirror: ID-Preserved Video Generation in Video Diffusion TransformersPoster
- MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory GuidancePoster
- Make Me Happier: Evoking Emotions Through Image Diffusion ModelsPoster
- Make Your Training Flexible: Towards Deployment-Efficient Video ModelsPoster
- MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric SequencePoster
- MamV2XCalib: V2X-based Target-less Infrastructure Camera Calibration with State Space ModelPoster
- Mamba-3VL: Taming State Space Model for 3D Vision Language LearningPoster
- MambaML: Exploring State Space Models for Multi-Label Image ClassificationPoster
- Manual-PA: Learning 3D Part Assembly from Instruction DiagramsPoster
- Marigold-DC: Zero-Shot Monocular Depth Completion with Guided DiffusionPoster
- MaskControl: Spatio-Temporal Control for Masked Motion SynthesisPoster
- MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the WildPoster
- MaskSAM: Auto-prompt SAM with Mask Classification for Volumetric Medical Image SegmentationPoster
- Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and RepresentativenessPoster
- MatchDiffusion: Training-free Generation of Match-CutsPoster
- MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR DiffusionPoster
- MeasureXpert: Automatic Anthropometric Measurement Extraction from Two Unregistered, Partial, Posed, and Dressed Body ScansPoster
- Measuring the Impact of Rotation Equivariance on Aerial Object DetectionPoster
- MedSegFactory: Text-Guided Generation of Medical Image-Mask PairsPoster
- MedVSR: Medical Video Super-Resolution with Cross State-Space PropagationPoster
- Medical World ModelPoster
- MemDistill: Distilling LiDAR Knowledge into Memory for Camera-Only 3D Object DetectionPoster
- Membership Inference Attacks with False Discovery Rate ControlPoster
- Memory-Efficient 4-bit Preconditioned Stochastic OptimizationPoster
- Memory-Efficient Generative Models via Product QuantizationPoster
- MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided StylizationPoster
- MergeOcc: Bridge the Domain Gap between Different LiDARs for Robust Occupancy PredictionPoster
- MeshAnything V2: Artist-Created Mesh Generation with Adjacent Mesh TokenizationPoster
- MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D MeshPoster
- MeshMamba: State Space Models for Articulated 3D Mesh Generation and ReconstructionPoster
- MeshPad: Interactive Sketch-Conditioned Artist-Reminiscent Mesh Generation and EditingPoster
- Met2Net: A Decoupled Two-Stage Spatio-Temporal Forecasting Model for Complex Meteorological SystemsPoster
- Meta-Learning Dynamic Center Distance: Hard Sample Mining for Learning with Noisy LabelsPoster
- Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned ConceptsPoster
- MetaMorph: Multimodal Understanding and Generation via Instruction TuningPoster
- MetaScope: Optics-Driven Neural Network for Ultra-Micro Metalens EndoscopyPoster
- Metric Convolutions: A Unifying Theory to Adaptive Image ConvolutionsPoster
- MiDSummer: Multi-Guidance Diffusion for Controllable Zero-Shot Immersive Gaussian Splatting Scene GenerationPoster
- MikuDance: Animating Character Art with Mixed Motion DynamicsPoster
- MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnPPoster
- Mind the Cost of Scaffold! Benign Clients May Even Become Accomplices of Backdoor AttackPoster
- Mind the Gap: Aligning Vision Foundation Models to Image Feature MatchingPoster
- Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual LearningPoster
- MissRAG: Addressing the Missing Modality Challenge in Multimodal Large Language ModelsPoster
- MistSense: Versatile Online Detection of Procedural and Execution MistakesPoster
- Mitigating Catastrophic Overfitting in Fast Adversarial Training via Label Information EliminationPoster
- Mitigating Geometric Degradation in Fast DownSampling via FastAdapter for Point Cloud SegmentationPoster
- Mitigating Object Hallucinations via Sentence-Level Early InterventionPoster
- MixA-Q: Revisiting Activation Sparsity for Vision Transformers from a Mixed-Precision Quantization PerspectivePoster
- MixA: A Mixed Attention approach with Stable Lightweight Linear Attention to enhance Efficiency of Vision Transformers at the EdgePoster
- MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action AnticipationPoster
- MixRI: Mixing Features of Reference Images for Novel Object Pose EstimationPoster
- Mixed Signals: A Diverse Point Cloud Dataset for Heterogeneous LiDAR V2X CollaborationPoster
- Mixture of Experts Guided by Gaussian Splatters Matters: A new Approach to Weakly-Supervised Video Anomaly DetectionPoster
- Mixture-of-Scores: Robust Image-Text Data Valuation via Three Lines of CodePoster
- MoFRR: Mixture of Diffusion Models for Face Retouching RestorationPoster
- MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar ReconstructionPoster
- MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile ManipulationPoster
- MoMaps: Semantics-Aware Scene Motion Generation with Motion MapsPoster
- MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised LearningPoster
- Mobile Video DiffusionPoster
- MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile DevicesPoster
- MobileViCLIP: An Efficient Video-Text Model for Mobile DevicesPoster
- ModSkill: Physical Character Skill ModularizationPoster
- ModalTune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-task Learning in Digital PathologyPoster
- Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion ModelsPoster
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath PredictionPoster
- Modeling Saliency Dataset BiasPoster
- Moderating the Generalization of Score-based Generative ModelPoster
- MolParser: End-to-end Visual Recognition of Molecule Structures in the WildPoster
- Moment Quantization for Video Temporal GroundingPoster
- Momentum-GS: Momentum Gaussian Self-Distillation for High-Quality Large Scene ReconstructionPoster
- MonSTeR: a Unified Model for Motion, Scene, Text RetrievalPoster
- MonoFusion: Sparse-View 4D Reconstruction via Monocular FusionPoster
- MonoMVSNet: Monocular Priors Guided Multi-View Stereo NetworkPoster
- MonoMobility: Zero-Shot 3D Mobility Analysis from Monocular VideosPoster
- MonoSOWA: Scalable Monocular 3D Object Detector Without Human AnnotationsPoster
- Monocular Facial Appearance Capture in the WildPoster
- Monocular Semantic Scene Completion via Masked Recurrent NetworksPoster
- More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label LearningPoster
- MorphoGen: Efficient Unconditional Generation of Long-Range Projection Neuronal Morphology via a Global-to-Local FrameworkPoster
- MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining DynamicsPoster
- Motal: Unsupervised 3D Object Detection by Modality and Task-specific Knowledge TransferPoster
- Motion Synthesis with Sparse and Flexible Keyjoint ControlPoster
- Motion-2-to-3: Leveraging 2D Motion Data for 3D Motion GenerationsPoster
- MotionAgent: Fine-grained Controllable Video Generation via Motion Field AgentPoster
- MotionCtrl: A Real-time Controllable Vision-Language-Motion ModelPoster
- MotionDiff: Training-free Zero-shot Interactive Motion Editing via Flow-assisted Multi-view DiffusionPoster
- MotionFollower: Editing Video Motion via Score-Guided DiffusionPoster
- MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion ParadigmPoster
- MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video GenerationPoster
- MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent SpacePoster
- Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from VideosPoster
- Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied NavigationPoster
- MuGS: Multi-Baseline Generalizable Gaussian Splatting ReconstructionPoster
- Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language ModelsPoster
- Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMsPoster
- Multi-Modal Few-Shot Temporal Action SegmentationPoster
- Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM): A Task-Adaptive Representation Learning FrameworkPoster
- Multi-Object Sketch Animation by Scene Decomposition and Motion PlanningPoster
- Multi-Schema Proximity Network for Composed Image RetrievalPoster
- Multi-View 3D Point TrackingPoster
- Multi-View Slot Attention Using Paraphrased Texts for Face Anti-SpoofingPoster
- Multi-identity Human Image Animation with Structural Video DiffusionPoster
- Multi-modal Identity ExtractionPoster
- Multi-modal Multi-platform Person Re-Identification: Benchmark and MethodPoster
- Multi-modal Segment Anything Model for Camouflaged Scene SegmentationPoster
- Multi-scenario Overlapping Text Segmentation with Depth AwarenessPoster
- Multi-turn Consistent Image EditingPoster
- Multi-view Gaze Target EstimationPoster
- MultiADS: Defect-aware Supervision for Multi-type Anomaly Detection and Segmentation in Zero-Shot LearningPoster
- MultiModal Action Conditioned Video Simulation
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language ModelsPoster
- Multidimensional Byte Pair Encoding: Shortened Sequences for Improved Visual Data GenerationPoster
- Multimodal LLM Guided Exploration and Active Mapping using Fisher InformationPoster
- Multimodal LLMs as Customized Reward Models for Text-to-Image GenerationPoster
- Multimodal Large Language Model-Guided ISP Hyperparameter Optimization with Dynamic Preference LearningPoster
- Multimodal Latent Diffusion Model for Complex Sewing Pattern GenerationPoster
- Multimodal Prompt Alignment for Facial Expression RecognitionPoster
- Multispectral Demosaicing via Dual CamerasPoster
- MultiverSeg: Scalable Interactive Segmentation of Biomedical Imaging Datasets with In-Context GuidancePoster
- Music Grounding by Short VideoPoster
- Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion ModelingPoster
- NAPPure: Adversarial Purification for Robust Image Classification under Non-Additive PerturbationsPoster
- NATRA: Noise-Agnostic Framework for Trajectory Prediction with Noisy ObservationsPoster
- NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic ReasoningPoster
- NETracer: A Topology-Aware Iterative Tracing Approach for Tubular Structure ExtractionPoster
- NGD: Neural Gradient Based Deformation for Monocular Garment ReconstructionPoster
- Nautilus: Locality-aware Autoencoder for Scalable Mesh GenerationPoster
- NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous EnvironmentsPoster
- NavQ: Learning a Q-Model for Foresighted Vision-and-Language NavigationPoster
- NeRF Is a Valuable Assistant for 3D Gaussian SplattingPoster
- NegRefine: Refining Negative Label-Based Zero-Shot OOD DetectionPoster
- NeuFrameQ: Neural Frame Fields for Scalable and Generalizable Anisotropic QuadrangulationPoster
- NeurOp-Diff: Continuous Remote Sensing Image Super-Resolution via Neural Operator DiffusionPoster
- Neural Architecture Search Driven by Locally Guided Diffusion for Personalized Federated LearningPoster
- Neural Compression for 3D Geometry SetsPoster
- Neural Inverse Rendering for High-Accuracy 3D Measurement of Moving Objects with Fewer Phase-Shifting PatternsPoster
- Neural Multi-View Self-Calibrated Photometric Stereo without Photometric Stereo CuesPoster
- Neural Shell Texture Splatting: More Details and Fewer PrimitivesPoster
- Neural Solver of Dichromatic Reflection Model for Specular Highlight RemovalPoster
- NeuralSVG: An Implicit Representation for Text-to-Vector GenerationPoster
- Neuromanifold-Regularized KANs for Shape-fair Feature RepresentationsPoster
- Neurons: Emulating the Human Visual Cortex Improves Fidelity and Interpretability in fMRI-to-Video ReconstructionPoster
- Neuroverse3D: Developing In-Context Learning Universal Model for Neuroimaging in 3DPoster
- No More Sibling Rivalry: Debiasing Human-Object Interaction DetectionPoster
- No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse ViewsPoster
- Noise-Modeled Diffusion Models for Low-Light Spike Image RestorationPoster
- Noise2Score3D: Tweedie's Approach for Unsupervised Point Cloud DenoisingPoster
- NoiseController: Towards Consistent Multi-view Video Generation via Noise Decomposition and CollaborationPoster
- Normal and Abnormal Pathology Knowledge-Augmented Vision-Language Model for Anomaly Detection in Pathology ImagesPoster
- NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion PriorsPoster
- NormalLoc: Visual Localization on Textureless 3D Models using Surface NormalsPoster
- Not All Degradations Are Equal: A Targeted Feature Denoising Framework for Generalizable Image Super-ResolutionPoster
- Not All Frame Features Are Equal: Video-to-4D Generation via Decoupling Dynamic-Static FeaturesPoster
- Not Only Vision: Evolve Visual Speech Recognition via Peripheral InformationPoster
- Not all Views are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation ModelsPoster
- NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language ModelsPoster
- NuiScene: Exploring Efficient Generation of Unbounded Outdoor ScenesPoster
- NullSwap: Proactive Identity Cloaking Against Deepfake Face SwappingPoster
- O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric ViewsPoster
- OCK: Unsupervised Dynamic Video Prediction with Object-Centric KinematicsPoster
- OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented GenerationPoster
- OCSplats: Observation Completeness Quantification and Label Noise Separation in 3DGSPoster
- OD-RASE: Ontology-Driven Risk Assessment and Safety Enhancement for Autonomous DrivingPoster
- ODDR: Outlier Detection & Dimension Reduction Based Defense Against Adversarial PatchesPoster
- ODP-Bench: Benchmarking Out-of-Distribution Performance PredictionPoster
- OMNI-DC: Highly Robust Depth Completion with Multiresolution Depth IntegrationPoster
- ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language ModelsPoster
- ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action GenerationPoster
- OURO: A Self-Bootstrapped Framework for Enhancing Multimodal Scene UnderstandingPoster
- OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object DetectionPoster
- OV3D-CG: Open-vocabulary 3D Instance Segmentation with Contextual GuidancePoster
- OVA-Fields: Weakly Supervised Open-Vocabulary Affordance Fields for Robot Operational Part DetectionPoster
- OVG-HQ: Online Video Grounding with Hybrid-modal QueriesPoster
- Oasis: One Image is All You Need for Multimodal Instruction Data SynthesisPoster
- Object-centric Video Question Answering with Visual Grounding and ReferringPoster
- Object-level Correlation for Few-Shot SegmentationPoster
- ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian SplattingPoster
- ObjectMate: A Recurrence Prior for Object Insertion and Subject-Driven GenerationPoster
- ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric PerspectivesPoster
- OcRFDet: Object-Centric Radiance Fields for Multi-View 3D Object Detection in Autonomous DrivingPoster
- OccluGaussian: Occlusion-Aware Gaussian Splatting for Large Scene Reconstruction and RenderingPoster
- Occlusion-robust Stylization for Drawing-based 3D AnimationPoster
- Occupancy Learning with Spatiotemporal MemoryPoster
- Omegance: A Single Parameter for Various Granularities in Diffusion-Based SynthesisPoster
- OminiControl: Minimal and Universal Control for Diffusion TransformerPoster
- Omni-scene Perception-oriented Point Cloud Geometry Enhancement for Coordinate QuantizationPoster
- OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer ModelsPoster
- OmniDiff: A Comprehensive Benchmark for Fine-grained Image Difference CaptioningPoster
- OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation ModelsPoster
- OmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal InpaintingPoster
- OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic SegmentationPoster
- OmniVTON: Training-Free Universal Virtual Try-OnPoster
- On Large Multimodal Models as Open-World Image ClassifiersPoster
- On the Complexity-Faithfulness Trade-off of Gradient-Based ExplanationsPoster
- On the Generalization of Representation Uncertainty in Earth ObservationPoster
- On the Provable Importance of Gradients for Autonomous Language-Assisted Image ClusteringPoster
- On the Recovery of Cameras from Fundamental MatricesPoster
- On the Robustness Tradeoff in Fine-TuningPoster
- On-Device Diffusion Transformer Policy for Efficient Robot ManipulationPoster
- One Encoder to Rule them All: Representation Learning for Model-free Visual Reinforcement Learning using Fourier Neural OperatorsPoster
- One Look is Enough: Seamless Patchwise Refinement for Zero-Shot Monocular Depth Estimation on High-Resolution ImagesPoster
- One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language ModelsPoster
- One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training ModelsPoster
- One Polyp Identifies All: One-Shot Polyp Segmentation with SAM via Cascaded Priors and Iterative Prompt EvolutionPoster
- One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object TrajectoryPoster
- One-Shot Knowledge Transfer for Scalable Person Re-IdentificationPoster
- One-Step Specular Highlight Removal with Adapted Diffusion ModelsPoster
- OneGT: One-Shot Geometry-Texture Neural Rendering for Head AvatarsPoster
- Online Dense Point Tracking with Streaming MemoryPoster
- Online Generic Event Boundary DetectionPoster
- Online Language SplattingPoster
- Online Reasoning Video Segmentation with Just-in-Time Digital TwinsPoster
- Open-Unfairness Adversarial Mitigation for Generalized Deepfake DetectionPoster
- Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept CalibrationPoster
- Open-World Skill Discovery from Unsegmented Demonstration VideosPoster
- Open-ended Hierarchical Streaming Video Understanding with Vision Language ModelsPoster
- OpenAnimals: Revisiting Person Re-Identification for Animals Towards Better GeneralizationPoster
- OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human AnnotationsPoster
- OpenRSD: Towards Open-prompts for Object Detection in Remote Sensing ImagesPoster
- OpenSubstance: A High-quality Measured Dataset of Multi-View and -Lighting Images and ShapesPoster
- OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal LearningPoster
- OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingPoster
- Optical Model-Driven Sharpness Mapping for Autofocus in Small Depth-of-Field and Severe Defocus ScenariosPoster
- Optimal Transport for Brain-Image Alignment: Unveiling Redundancy and Synergy in Neural Information ProcessingPoster
- OracleFusion: Assisting the Decipherment of Oracle Bone Script with Structurally Constrained Semantic TypographyPoster
- Orchid: Image Latent Diffusion for Joint Appearance and Geometry GenerationPoster
- OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLMPoster
- OuroMamba: A Data-Free Quantization Framework for Vision MambaPoster
- Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse RenderingPoster
- Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian PointmapsPoster
- Outlier-Aware Post-Training Quantization for Image Super-ResolutionPoster
- Overcoming Dual Drift for Continual Long-Tailed Visual Question AnsweringPoster
- PAN-Crafter: Learning Modality-Consistent Alignment for PAN-SharpeningPoster
- PARTE: Part-Guided Texturing for 3D Human Reconstruction from a Single ImagePoster
- PASD: A Pixel-Adaptive Swarm Dynamics Approach for Unsupervised Low-Light Image EnhancementPoster
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic ManipulationPoster
- PASTA: Part-Aware Sketch-to-3D Shape Generation with Text-Aligned PriorPoster
- PBCAT: Patch-Based Composite Adversarial Training against Physically Realizable Attacks on Object DetectionPoster
- PBFG: A New Physically-Based Dataset and Removal of Lens Flares and GlaresPoster
- PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-RegularizationsPoster
- PEFTDiff: Diffusion-Guided Transferability Estimation for Parameter-Efficient Fine-TuningPoster
- PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single ImagePoster
- PHATNet: A Physics-guided Haze Transfer Network for Domain-adaptive Real-world Image DehazingPoster
- PHD: Personalized 3D Human Body Fitting with Point DiffusionPoster
- PINO: Person-Interaction Noise Optimization for Long-Duration and Customizable Motion Generation of Arbitrary-Sized GroupsPoster
- PLA: Prompt Learning Attack against Text-to-Image Generative ModelsPoster
- PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging SparsityPoster
- PLAN: Proactive Low-Rank Allocation for Continual LearningPoster
- PLMP - Point-Line Minimal Problems for Projective SfMPoster
- POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D ReconstructionPoster
- PRE-Mamba: A 4D State Space Model for Ultra-High-Frequent Event Camera DerainingPoster
- PRIMAL: Physically Reactive and Interactive Motor Model for Avatar LearningPoster
- PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding ProjectionPoster
- PRM: Photometric Stereo based Large Reconstruction ModelPoster
- PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt Relocation
- PROGRESSOR: A Perceptually Guided Reward Estimator with Self-Supervised Online RefinementPoster
- PROL : Rehearsal Free Continual Learning in Streaming Data via Prompt Online LearningPoster
- PS-Mamba: Spatial-Temporal Graph Mamba for Pose Sequence RefinementPoster
- PS3: A Multimodal Transformer Integrating Pathology Reports with Histology Images and Biological Pathways for Cancer Survival PredictionPoster
- PUMA: Empowering Unified MLLM with Multi-granular Visual GenerationPoster
- PUMPS: Skeleton-Agnostic Point-based Universal Motion Pre-Training for Synthesis in Human Motion TasksPoster
- PVChat: Personalized Video Chat with One-Shot LearningPoster
- PVMamba: Parallelizing Vision Mamba via Dynamic State AggregationPoster
- PacGDC: Label-Efficient Generalizable Depth Completion with Projection Ambiguity and ConsistencyPoster
- PanSt3R: Multi-view Consistent Panoptic SegmentationPoster
- PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMsPoster
- PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama ReconstructionPoster
- Parameter-Efficient Adaptation of Geospatial Foundation Models through Embedding DeflectionPoster
- Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion ModelsPoster
- PartField: Learning 3D Feature Fields for Part Segmentation and BeyondPoster
- Partial Forward Blocking: A Novel Data Pruning Paradigm for Lossless Training AccelerationPoster
- Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud LocalizationPoster
- Passing the Driving Knowledge TestPoster
- PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-ResolutionPoster
- PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask ConditionsPoster
- PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to HistopathologyPoster
- PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion ModelPoster
- Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative ModelsPoster
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsPoster
- Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion RepresentationPoster
- Performing Defocus Deblurring by Modeling its Formation ProcessPoster
- PersPose: 3D Human Pose Estimation with Perspective Encoding and Perspective RotationPoster
- PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion
- PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic DegradationPoster
- Personalized Federated Learning under Local SupervisionPoster
- Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery SimulationPoster
- Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous DistillationPoster
- Perspective-Invariant 3D Object DetectionPoster
- Perspective-aware 3D Gaussian Inpainting with Multi-view ConsistencyPoster
- Ph-GAN: Physics-Inspired GAN for Generating SAR Images Under Limited DataPoster
- Phantom: Subject-Consistent Video Generation via Cross-Modal AlignmentPoster
- Photolithography Overlay Map Generation with Implicit Knowledge Distillation Diffusion TransformerPoster
- PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object ModelingPoster
- PhysSplat: Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian SplattingPoster
- PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from VideosPoster
- Physical Degradation Model-Guided Interferometric Hyperspectral Reconstruction with Unfolding TransformerPoster
- Physics Context Builders: A Modular Framework for Physical Reasoning in Vision-Language ModelsPoster
- Pi-GPS: Enhancing Geometry Problem Solving by Unleashing the Power of Diagrammatic InformationPoster
- Pinco: Position-induced Consistent Adapter for Diffusion Transformer in Foreground-conditioned InpaintingPoster
- PixTalk: Controlling Photorealistic Image Processing and Editing with LanguagePoster
- PixelStitch: Structure-Preserving Pixel-Wise Bidirectional Warps for Unsupervised Image StitchingPoster
- PlaceIt3D: Language-Guided Object Placement in Real 3D ScenesPoster
- PlanGen: Towards Unified Layout Planning and Image Generation in Auto-Regressive Vision Language ModelsPoster
- Planar Affine Rectification from Local Change of Scale and OrientationPoster
- PlaneRAS: Learning Planar Primitives for 3D Plane RecoveryPoster
- Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video CaptioningPoster
- Plug-in Feedback Self-adaptive Attention in CLIP for Training-free Open-Vocabulary SegmentationPoster
- PlugMark: A Plug-in Zero-Watermarking Framework for Diffusion ModelsPoster
- Point Cloud Self-supervised Learning via 3D to Multi-view Masked LearnerPoster
- PointGAC: Geometric-Aware Codebook for Masked Point ModelingPoster
- PolGS: Polarimetric Gaussian Splatting for Fast Reflective Surface ReconstructionPoster
- PolarAnything: Diffusion-based Polarimetric Image SynthesisPoster
- Polarimetric Neural Field via Unified Complex-Valued Wave RepresentationPoster
- Ponimator: Unfolding Interactive Pose for Versatile Human-human Interaction AnimationPoster
- PoseAnchor: Robust Root Position Estimation for 3D Human Pose EstimationPoster
- PoseSyn: Synthesizing Diverse 3D Pose Data from In-the-Wild 2D DataPoster
- PossLoss: A Reliable and Sensitive Facial Landmark Detection Loss FunctionPoster
- Power of Cooperative Supervision: Multiple Teachers Framework for Advanced 3D Semi-Supervised Object DetectionPoster
- Preacher: Paper-to-Video Agentic SystemPoster
- Precise Action-to-Video Generation Through Visual Action PromptsPoster
- Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object UnderstandingPoster
- Preserve Anything: Controllable Image Synthesis with Object PreservationPoster
- Pretend Benign: A Stealthy Adversarial Attack by Exploiting Vulnerabilities in Cooperative PerceptionPoster
- Pretrained Reversible Generation as Unsupervised Visual Representation LearningPoster
- PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal ViewPoster
- PrimHOI: Compositional Human-Object Interaction via Reusable Primitives
- Princeton365: A Diverse Dataset with Accurate Camera PosePoster
- Principles of Visual Tokens for Efficient Video UnderstandingPoster
- Prior-aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose EstimationPoster
- Prior2Former - Evidential Modeling of Mask Transformers for Assumption-Free Open-World Panoptic SegmentationPoster
- Privacy-centric Deep Motion Retargeting for Anonymization of Skeleton-Based Motion VisualizationPoster
- ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis UsersPoster
- ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process JudgesPoster
- ProSAM: Enhancing the Robustness of SAM-based Visual Reference Segmentation with Probabilistic PromptsPoster
- Proactive Scene Decomposition and ReconstructionPoster
- ProbMED: A Probabilistic Framework for Medical Multimodal BindingPoster
- ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity RecognitionPoster
- Probabilistic Inertial Poser (ProbIP): Uncertainty-aware Human Motion Modeling from Sparse Inertial SensorsPoster
- Probabilistic Prototype Calibration of Vision-language Models for Generalized Few-shot Semantic SegmentationPoster
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?Poster
- Progressive Artwork Outpainting via Latent Diffusion ModelsPoster
- Progressive Distribution Bridging: Unsupervised Adaptation for Large-scale Pre-trained Models via Adaptive Auxiliary DataPoster
- Progressive Growing of Video Tokenizers for Temporally Compact Latent SpacesPoster
- Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental LearningPoster
- Progressive Test Time Energy Adaptation for Medical Image SegmentationPoster
- Prompt Guidance and Human Proximal Perception for HOT Prediction with Regional Joint LossPoster
- Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLMPoster
- Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual InversionPoster
- PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware MaskPoster
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity DiscriminationPoster
- Prototype Guided Backdoor Defense via Activation Space Manipulation
- Prototype-based Contrastive Learning with Stage-wise Progressive Augmentation for Self-Supervised Fine-Grained LearningPoster
- Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video RetrievalPoster
- Proxy-Bridged Game Transformer for Interactive Extreme Motion PredictionPoster
- Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language ModelsPoster
- Pseudo-SD: Pseudo Controlled Stable Diffusion for Semi-Supervised and Cross-Domain Semantic SegmentationPoster
- PseudoMapTrainer: Learning Online Mapping without HD MapsPoster
- Punching Bag vs. Punching Person: Motion Transferability in VideosPoster
- Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level DynamicsPoster
- Purge-Gate: Backpropagation-Free Test-Time Adaptation for Point Clouds Classification via Token purgingPoster
- Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene ReconstructionsPoster
- Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMsPoster
- Q-Norm: Robust Representation Learning via Quality-Adaptive NormalizationPoster
- QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingPoster
- QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized GenerationPoster
- QuEST: Low-bit Diffusion Model Quantization via Efficient Selective FinetuningPoster
- Quadratic Gaussian Splatting: High Quality Surface Reconstruction with Second-order Geometric PrimitivesPoster
- Quanta Neural Networks: From Photons to PerceptionPoster
- Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval via Uncertainty MinimizationPoster
- QuickSplat: Fast 3D Surface Reconstruction via Learned Gaussian InitializationPoster
- R-LiViT: A LiDAR-Visual-Thermal Dataset Enabling Vulnerable Road User Focused Roadside PerceptionPoster
- R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal FormalizationPoster
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy OptimizationPoster
- RA-BUSSeg: Relation-aware Semi-supervised Breast Ultrasound Image Segmentation via Adjacent Propagation and Cross-layer AlignmentPoster
- RAGD: Regional-Aware Diffusion Model for Text-to-Image GenerationPoster
- RAGDiffusion: Faithful Cloth Generation via External Knowledge AssimilationPoster
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General GraspingPoster
- RALoc: Enhancing Outdoor LiDAR Localization via Rotation AwarenessPoster
- RANKCLIP: Ranking-Consistent Language-Image PretrainingPoster
- RARE: Refine Any Registration of Pairwise Point Clouds via Zero-Shot LearningPoster
- RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal FusionPoster
- REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion LatentsPoster
- REGEN: Learning Compact Video Embedding with (Re-)Generative DecoderPoster
- REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion TransformersPoster
- REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout AlignmentPoster
- RESCUE: Crowd Evacuation Simulation via Controlling SDM-United CharactersPoster
- RGE-GS: Reward-Guided Expansive Driving Scene Reconstruction via Diffusion PriorsPoster
- RI3D: Few-Shot Gaussian Splatting With Repair and Inpainting Diffusion PriorsPoster
- RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy PredictionPoster
- RIPE: Reinforcement Learning on Unlabeled Image Pairs for Robust Keypoint ExtractionPoster
- RMultiplex200K: Toward Reliable Multimodal Process Supervision for Visual Language Models on TelecommunicationsPoster
- ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work ZonesPoster
- ROAR: Reducing Inversion Error in Generative Image WatermarkingPoster
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image GenerationPoster
- RS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation ModelPoster
- RTMap: Real-Time Recursive Mapping with Change Detection and LocalizationPoster
- RadGPT: Constructing 3D Image-Text Tumor DatasetsPoster
- RadarSplat: Radar Gaussian Splatting for High-Fidelity Data Synthesis and 3D Reconstruction of Autonomous Driving ScenesPoster
- Radiant Foam: Real-Time Differentiable Ray TracingPoster
- RainbowPrompt: Diversity-Enhanced Prompt-Evolving for Continual LearningPoster
- Randomized Autoregressive Visual GenerationPoster
- RapVerse: Coherent Vocals and Whole-Body Motion Generation from TextPoster
- RareCLIP: Rarity-aware Online Zero-shot Industrial Anomaly DetectionPoster
- RayGaussX: Accelerating Gaussian-Based Ray Marching for Real-Time and High-Quality Novel View SynthesisPoster
- RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose EstimationPoster
- RayZer: A Self-supervised Large View Synthesis ModelPoster
- RayletDF: Raylet Distance Fields for Generalizable 3D Surface Reconstruction from Point Clouds or GaussiansPoster
- ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous DrivingPoster
- ReCamMaster: Camera-Controlled Generative Rendering from A Single VideoPoster
- ReCoT: Reflective Self-Correction Training for Mitigating Confirmation Bias in Large Vision-Language ModelsPoster
- ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention AdaptationPoster
- ReME: A Data-Centric Framework for Training-Free Open-Vocabulary SegmentationPoster
- ReMP-AD: Retrieval-enhanced Multi-modal Prompt Fusion for Few-Shot Industrial Visual Anomaly DetectionPoster
- RePoseD: Efficient Relative Pose Estimation With Known Depth InformationPoster
- ReTracker: Exploring Image Matching for Robust Online Any Point TrackingPoster
- Real3D: Towards Scaling Large Reconstruction Models with Real ImagesPoster
- RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera ControlPoster
- RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video ModelsPoster
- Reangle-A-Video: 4D Video Generation as Video-to-Video TranslationPoster
- ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question AnsweringPoster
- ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco ReconstructionPoster
- Recognizing Actions from Robotic View for Natural Human-Robot InteractionPoster
- ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene RepresentationPoster
- Recover Biological Structure from Sparse-View Diffraction Images with Neural Volumetric PriorPoster
- Recovering Parametric Scenes from Very Few Time-of-Flight PixelsPoster
- Rectifying Magnitude Neglect in Linear AttentionPoster
- Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy RegularizationPoster
- RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring ExpressionsPoster
- Refer to Any Segmentation Mask Group With Vision-Language PromptsPoster
- ReferDINO: Referring Video Object Segmentation with Visual Grounding FoundationsPoster
- ReferEverything: Towards Segmenting Everything We Can Speak of in VideosPoster
- Reference-based Super-Resolution via Image-based Retrieval-Augmented Generation DiffusionPoster
- Referring Expression Comprehension for Small ObjectsPoster
- Referring to Any PersonPoster
- Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context ReflectionPoster
- RegGS: Unposed Sparse Views Gaussian Splatting with 3DGS RegistrationPoster
- Region-Level Data Attribution for Text-to-Image Generative ModelsPoster
- Region-aware Anchoring Mechanism for Efficient Referring Visual GroundingPoster
- Region-based Cluster Discrimination for Visual Representation LearningPoster
- Registration beyond Points: General Affine Subspace Alignment via Geodesic Distance on Grassmann ManifoldPoster
- Reinforcement Learning-Guided Data Selection via Redundancy AssessmentPoster
- Relative Illumination Fields: Learning Medium and Light Independent Underwater ScenesPoster
- Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for PrivacyPoster
- Removing Cost Volumes from Optical Flow EstimatorsPoster
- Removing Out-of-Focus Reflective Flares via Color AlignmentPoster
- Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task LearningPoster
- Representation Shift: Unifying Token Compression with FlashAttentionPoster
- Representing 3D Shapes with 64 Latent Vectors for 3D Diffusion ModelsPoster
- Repurposing 2D Diffusion Models with Gaussian Atlas for 3D GenerationPoster
- ResGS: Residual Densification of 3D Gaussian for Efficient Detail RecoveryPoster
- ResQ: A Novel Framework to Implement Residual Neural Networks on Analog Rydberg Atom Quantum ComputersPoster
- ResidualViT for Efficient Temporally Dense Video EncodingPoster
- Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task LearningPoster
- Resonance: Learning to Predict Social-Aware Pedestrian Trajectories as Co-VibrationsPoster
- Rethink Sparse Signals for Pose-guided Text-to-image GenerationPoster
- Rethinking Bimanual Robotic Manipulation: Learning with Decoupled Interaction FrameworkPoster
- Rethinking Cross-Modal Interaction in Multimodal Diffusion TransformersPoster
- Rethinking DPO-style Diffusion Aligning FrameworksPoster
- Rethinking Detecting Salient and Camouflaged Objects in Unconstrained ScenesPoster
- Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image SynthesisPoster
- Rethinking Few Shot CLIP Benchmarks: A Critical Analysis in the Inductive SettingPoster
- Rethinking Key-frame-based Micro-expression Recognition: A Robust and Accurate Framework Against Key-frame ErrorsPoster
- Rethinking Layered Graphic Design Generation with a Top-Down ApproachPoster
- Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual DisparitiesPoster
- Retinex-MEF: Retinex-based Glare Effects Aware Unsupervised Multi-Exposure Image FusionPoster
- RetinexMCNet: A Memory Controller Dominated Network for Low-Light Video Enhancement Based on RetinexPoster
- Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image SetsPoster
- Revelio: Interpreting and leveraging semantic information in diffusion modelsPoster
- Reverse Convolution and Its Applications to Image RestorationPoster
- Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New InsightsPoster
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature AlignmentPoster
- Revisiting Image Fusion for Multi-Illuminant White-Balance CorrectionPoster
- Revisiting Point Cloud Completion: Are We Ready For The Real-World?Poster
- Revisiting Pool-based Prompt Learning for Few-shot Class-incremental LearningPoster
- RhythmGuassian: Repurposing Generalizable Gaussian Model For Remote Physiological MeasurementPoster
- Riemannian-Geometric Fingerprints of Generative ModelsPoster
- RnGCam: High-speed video from rolling & global shutter measurementsPoster
- RoCo-Sim: Enhancing Roadside Collaborative Perception through Foreground SimulationPoster
- RoMo: Robust Motion Segmentation Improves Structure from MotionPoster
- RobAVA: A Large-scale Dataset and Baseline Towards Video based Robotic Arm Action UnderstandingPoster
- Robin3D: Improving 3D Large Language Model via Robust Instruction TuningPoster
- RoboAnnotatorX: A Comprehensive and Universal Annotation Framework for Accurate Understanding of Long-horizon Robot DemonstrationPoster
- RoboPearls: Editable Video Simulation for Robot ManipulationPoster
- RoboTrom-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and PredictionPoster
- RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous DrivingPoster
- RoboTron-Mani: All-in-One Multimodal Large Model for Robotic ManipulationPoster
- RoboTron-Sim: Improving Real-World Driving via Simulated Hard-CasePoster
- Robust 3D Object Detection using Probabilistic Point Clouds from Single-Photon LiDARs
- Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation SamplingPoster
- Robust Adverse Weather Removal via Spectral-based Spatial GroupingPoster
- Robust Dataset Condensation using Supervised Contrastive LearningPoster
- Robust Low-light Scene Restoration via Illumination TransitionPoster
- Robust Machine Unlearning for Quantized Neural Networks via Adaptive Gradient Reweighting with Similar LabelsPoster
- Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated PerturbationPoster
- Robust Test-Time Adaptation for Single Image Denoising Using Deep Gaussian PriorPoster
- Robust Unfolding Network for HDR Imaging with Modulo CamerasPoster
- Robust and Efficient 3D Gaussian Splatting for Urban Scene ReconstructionPoster
- RobustSplat: Decoupling Densification and Dynamics for Transient-Free 3DGSPoster
- Robustifying Zero-Shot Vision Language Models by Subspaces AlignmentPoster
- RogSplat: Robust Gaussian Splatting via Generative PriorsPoster
- RomanTex: Decoupling 3D-aware Rotary Positional Embedded Multi-Attention Network for Texture SynthesisPoster
- Ross3D: Reconstructive Visual Instruction Tuning with 3D-AwarenessPoster
- S2M2: Scalable Stereo Matching Model for Reliable Depth EstimationPoster
- S3E: Self-Supervised State Estimation for Radar-Inertial SystemPoster
- S3R-GS: Streamlining the Pipeline for Large-Scale Street Scene ReconstructionPoster
- S4M: Boosting Semi-Supervised Instance Segmentation with SAMPoster
- SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style TransferPoster
- SA-Occ: Satellite-Assisted 3D Occupancy Prediction in Real WorldPoster
- SAC-GNC: SAmple Consensus for adaptive Graduated Non-ConvexityPoster
- SAFER: Sharpness Aware layer-selective Finetuning for Enhanced Robustness in vision transformersPoster
- SAFT: Shape and Appearance of Fabrics from Template via Differentiable Physical Simulations from Monocular VideoPoster
- SAGI: Semantically Aligned and Uncertainty Guided AI Image InpaintingPoster
- SALAD -- Semantics-Aware Logical Anomaly DetectionPoster
- SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model FailuresPoster
- SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory TreePoster
- SAM4D: Segment Anything in Camera and LiDAR StreamsPoster
- SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of ExpertsPoster
- SAMO: A Lightweight Sharpness-Aware Approach for Multi-Task Optimization with Joint Global-Local PerturbationPoster
- SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action RecognitionPoster
- SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical ImagesPoster
- SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency DistillationPoster
- SAS: Segment Any 3D Scene with Integrated 2D PriorsPoster
- SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse AutoencodersPoster
- SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement LearningPoster
- SC-Lane: Slope-aware and Consistent Road Height Estimation Framework for 3D Lane DetectionPoster
- SCAN: Bootstrapping Contrastive Pre-training for Data EfficiencyPoster
- SCFlow: Implicitly Learning Style and Content Disentanglement with Flow ModelsPoster
- SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance SegmentationPoster
- SD2Actor: Continuous State Decomposition via Diffusion Embeddings for Robotic ManipulationPoster
ICCV accepted papers in other years
Looking for submission deadlines instead? See the conference deadline calendar.