← Search

Yaowei Wang

106 accepted papers

2026

AffIn-Space: Learning Affine-Invariant Representations for 3D Spatial Understanding with MLLMs

ICML 2026poster

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general visual understanding, they suffer from a fundamental geometric fragility: standard visual representations often degrade rapidly under changes in viewpoint and viewing distance. Our analysis identifies that ex…

Cited by 0SourceScholar
2026

Beyond Heuristics: Learnable Density Control for 3D Gaussian Splatting

ICML 2026poster

While 3D Gaussian Splatting (3DGS) has demonstrated impressive real-time rendering performance, its efficacy remains constrained by a reliance on heuristic density control. Despite numerous refinements to these handcrafted rules, such methods inherently lack the flexibility to adapt to diverse scene…

Cited by 0SourceScholar
2026

CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception

ICML 2026poster

High-resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods struggle with the trade-off between coverage and efficiency. Visual expert-assisted search is efficient but prone to blind s…

Cited by 0SourceScholar
2026

Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining

CVPR 2026

Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, they still suffer from two fundamental limitations: severe visual information lo

Cited by 0SourcecodeScholar
2026

CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation

ICLR 2026poster

Existing works on reasoning segmentation either connect hidden features from a language model directly to a mask decoder or represent positions in text, which limits interpretability and semantic detail. To solve this, we present CoPRS, a Multi-modal Chain-of-Thought (MCoT)–based positional percepti…

Cited by 0SourceScholar
2026

Compositional Transformation Reasoning for Composed Video Retrieval

CVPR 2026

Composed Video Retrieval aims to retrieve a target video given a reference video and a textual modification describing the desired change. The core challenge lies in modeling compositional multimodal transformations, i.e., how entities, actions, and scenes evolve across video and language modalities

Cited by 0SourcecodeScholar
2026

Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style Bridging

CVPR 2026

Continual Test-Time Adaptation (CTTA) aims to empower perception systems to handle dynamic distribution shifts encountered after deployment. Existing methods predominantly follow a backward-alignment paradigm, which rigidly aligns incoming data with supervisory surrogates derived from the source dom

Cited by 0SourcecodeScholar
2026

DuetMerging: Synergizing Dynamic and Static Strategies for Mitigating Task Interference in Model Merging

CVPR 2026

Model merging offers a promising paradigm for consolidating multiple expert models into a single multitask architecture. However, its effectiveness is often hindered by task interference, where conflicting parameter updates from different tasks degrade performance. While dynamic, Mixture-of-Experts

Cited by 0SourceScholar
2026

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

AAAI 2026technical

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric) vision, overlooking the unique challenges of first-person (egoce

Cited by 0SourcePDFScholar
2026

HIGH QUALITY UNDERWATER IMAGE COMPRESSION WITH ADAPTIVE COLOR CORRECTION

ICASSP 2026oral

With the increasing exploration and exploitation of the underwater world, underwater images have become a critical medium for human interaction with marine environments, driving extensive research into their efficient transmission and storage. However, contemporary underwater image compression algor…

Cited by 0SourcePDFScholar
2026

Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval

CVPR 2026

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos based on text queries that describe only partial events. Existing methods suffer from incomplete global contextual perception, struggling with query ambiguity and local noise induced by spurious responses. To address these i

Cited by 0SourcecodeScholar
2026

Interactive Tracking: A Human-in-the-Loop Paradigm with Memory-Augmented Adaptation

CVPR 2026

Existing visual trackers mainly operate in a non-interactive, fire-and-forget manner, making them impractical for real-world scenarios that require human-in-the-loop adaptation. To overcome this limitation, we introduce Interactive Tracking, a new paradigm that allows users to guide the tracker at a

Cited by 0SourcecodeScholar
2026

Joint Learning of General and Diverse Patterns with Mixture of Memory Experts for Weakly-Supervised Video Anomaly Detection

CVPR 2026

Weakly-supervised Video Anomaly Detection (wVAD) aims to detect abnormal events using only binary labels, making it challenging to capture both the diversity of anomalies and their shared semantic cues. Existing methods either focus on a generic anomaly pattern, achieving strong generalization but w

Cited by 0SourceScholar
2026

Norm$\times$Direction: Restoring the Missing Query Norm in Vision Linear Attention

ICML 2026poster

Linear attention mitigates the quadratic complexity of softmax attention but suffers from a critical loss of expressiveness. We identify two primary causes: (1) The normalization operation cancels the query norm, which breaks the correlation between a query's norm and the spikiness (entropy) of the …

Cited by 0SourceScholar
2026

Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation

ICLR 2026poster

Dataset distillation aims to generate compact synthetic datasets that enable models trained on them to achieve performance comparable to those trained on full real datasets, while substantially reducing storage and computational costs. Early bi-level optimization methods (e.g., MTT) have shown promi…

Cited by 0SourcecodeScholar
2026

Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval

ICML 2026poster

Partially relevant video retrieval aims to retrieve untrimmed videos using text queries that describe only partial content. However, the inherent asymmetry between brief queries and rich video content inevitably introduces uncertainty into the retrieval process. In this setting, vague queries often …

Cited by 0SourceScholar
2026

SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation

CVPR 2026

While large language models provide strong compositional reasoning, existing reasoning segmentation pipelines fail to transparently connect this reasoning to visual perception. Current methods, such as latent query alignment, are end-to-end yet opaque "black boxes". Conversely, textual localization

Cited by 0SourcecodeScholar
2026

TGV-KV: Text-Grounded KV Eviction for Vision-Language Models

ICML 2026poster

Vision-Language Models (VLMs) inherit the auto-regressive generation paradigm and cache the keys and values (KV) of all previous tokens to accelerate inference, resulting in memory consumption that scales linearly with context length. This issue is particularly pronounced in VLMs due to substantial …

Cited by 0SourceScholar
2026

VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved impressive progress in vision–language alignment, yet they remain limited in visual–spatial reasoning. We first identify that this limitation arises from the attention mechanism: visual tokens are overshadowed by language tokens, preventing the…

Cited by 0SourceScholar
2025

A Survey on the Feedback Mechanism of LLM-based AI Agents

IJCAI 2025

Large language models (LLMs) are increasingly being adopted to develop general-purpose AI agents. However, it remains challenging for these LLM-based AI agents to efficiently learn from feedback and iteratively optimize their strategies. To address this challenge, tremendous efforts have been dedica

2025

A Unified Agentic Framework for Evaluating Conditional Image Generation

ACL 2025long

Conditional image generation has gained significant attention for its ability to personalize content. However, the field faces challenges in developing task-agnostic, reliable, and explainable evaluation metrics. This paper introduces CIGEval, a unified agentic framework for comprehensive evaluation…

2025

An Exploration with Entropy Constrained 3D Gaussians for 2D Video Compression

ICLR 2025poster

3D Gaussian Splatting (3DGS) has witnessed its rapid development in novel view synthesis, which attains high quality reconstruction and real-time rendering. At the same time, there is still a gap before implicit neural representation (INR) can become a practical compressor due to the lack of stream…

2025

AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing

CVPR 2025poster

Self-Supervised Video Hashing (SSVH) compresses videos into hash codes for efficient indexing and retrieval using unlabeled training videos. Existing approaches rely on random frame sampling to learn video features and treat all frames equally. This results in suboptimal hash codes, as it ignores fr…

2025

Building Vision Models upon Heat Conduction

CVPR 2025poster

Visual representation models leveraging attention mechanisms are challenged by significant computational overhead, particularly when pursuing large receptive fields. In this study, we aim to mitigate this challenge by introducing the Heat Conduction Operator (HCO) built upon the physical heat conduc…

2025

C2AD: Dual Consistency Learning for Zero-Shot Anomaly Detection

ICASSP 2025accepted

Zero-shot anomaly detection (ZSAD) is dedicated to detecting anomalies without having any seen normal or abnormal samples for the target set. Existing approaches utilize the pre-trained CLIP to assess normality/abnormality by exploiting the similarity between images and text with the frozen visual e…

Cited by 0SourceScholar
2025

Cassic: Towards Content-Adaptive State-Space Models for Learned Image Compression

ICCV 2025poster

Learned image compression (LIC) demonstrates superior rate-distortion (RD) performance compared to traditional methods. Recent method MambaVC attempts to introduce Mamba, a variant of state space models, into this field aim to establish a new paradigm beyond convolutional neural networks and transfo…

Cited by 0SourcePDFScholar
2025

Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic Scenarios

ICCV 2025poster

In practice, environments constantly change over time and space, posing significant challenges for object detectors trained based on a closed-set assumption, i.e., training and test data share the same distribution. To this end, continual test-time adaptation has attracted much attention, aiming to…

Cited by 0SourcePDFScholar
2025

Core Context Aware Transformers for Long Context Language Modeling

ICML 2025poster

Transformer-based Large Language Models (LLMs) have exhibited remarkable success in extensive tasks primarily attributed to self-attention mechanism, which requires a token to consider all preceding tokens as its context to compute attention. However, when the context length L becomes very large (e.…

Cited by 12SourcePDFScholar
2025

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

CVPR 2025poster

3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. However, existing methods usually rely on global scene perception from pure 3D point c…

2025

DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic Refinement

ICLR 2025poster

Reconstructing high-quality images under low bitrates conditions presents a challenge, and previous methods have made this task feasible by leveraging the priors of diffusion models. However, the effective exploration of pre-trained latent diffusion models and semantic information integration in im…

Cited by 0SourcePDFScholar
2025

EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment

ICLR 2025poster

Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language models (MLLM) are insufficient in extracting visual features, leading to imbalanc…

2025

Efficient Safety Alignment of Large Language Models via Preference Re-ranking and Representation-based Reward Modeling

ACL 2025long

Reinforcement Learning (RL) algorithms for safety alignment of Large Language Models (LLMs), such as Direct Preference Optimization (DPO), encounter the challenge of distribution shift. Current approaches typically address this issue through online sampling from the target policy, which requires sig…

Cited by 0SourcePDFScholar
2025

Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning

CVPR 2025poster

Visual In-Context Learning (VICL) enables adaptively solving vision tasks by leveraging pixel demonstrations, mimicking human-like task completion through analogy. Prompt selection is critical in VICL, but current methods assume the existence of a single "ideal" prompt in a pool of candidates, which…

2025

Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

ICCV 2025poster

Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos…

2025

FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation

NeurIPS 2025poster

Recent Large Vision Language Models (LVLMs) demonstrate promising capabilities in unifying visual understanding and generative modeling, enabling both accurate content understanding and flexible editing. However, current approaches treat \textbf{\textit{"what to see"}} and \textbf{\textit{"how to ed…

Cited by 0SourceScholar
2025

LaneDiffusion: Improving Centerline Graph Learning via Prior Injected BEV Feature Generation

ICCV 2025poster

Centerline graphs, crucial for path planning in autonomous driving, are traditionally learned using deterministic methods. However, these methods often lack spatial reasoning and struggle with occluded or invisible centerlines. Generative approaches, despite their potential, remain underexplored in…

2025

Learning Compatible Multi-Prize Subnetworks for Asymmetric Retrieval

CVPR 2025poster

Asymmetric retrieval is a typical scenario in real-world retrieval systems, where compatible models of varying capacities are deployed on platforms with different resource configurations. Existing methods generally train pre-defined networks or subnetworks with capacities specifically designed for p…

2025

Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval

ICLR 2025poster

With the explosive growth of video data, finding videos that meet detailed requirements in large datasets has become a challenge. To address this, the composed video retrieval task has been introduced, enabling users to retrieve videos using complex queries that involve both visual and textual infor…

2025

Learning Spatial-Semantic Features for Robust Video Object Segmentation

ICLR 2025poster

Tracking and segmenting multiple similar objects with distinct or complex parts in long-term videos is particularly challenging due to the ambiguity in identifying target components and the confusion caused by occlusion, background clutter, and changes in appearance or environment over time. In this…

Cited by 2SourcePDFScholar
2025

LoRATv2: Enabling Low-Cost Temporal Modeling in One-Stream Trackers

NeurIPS 2025spotlight

Transformer-based algorithms, such as LoRAT, have significantly enhanced object-tracking performance. However, these approaches rely on a standard attention mechanism, which incurs quadratic token complexity, making real-time inference computationally expensive. In this paper, we introduce LoRATv2,…

Cited by 0SourcecodeScholar
2025

Modeling Uncertainty in Composed Image Retrieval via Probabilistic Embeddings

ACL 2025long

Composed Image Retrieval (CIR) enables users to search for images using multimodal queries that combine text and reference images. While metric learning methods have shown promise, they rely on deterministic point embeddings that fail to capture the inherent uncertainty in the input data, in which u…

2025

NN-Former: Rethinking Graph Structure in Neural Architecture Representation

CVPR 2025poster

The growing use of deep learning necessitates efficient network design and deployment, making neural predictors vital for estimating attributes such as accuracy and latency. Recently, Graph Neural Networks (GNNs) and transformers have shown promising performance in representing neural architectures.…

2025

PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models

ACL 2025long

Large Language Models (LLMs) suffer severe performance degradation when facing extremely low-bit (sub 2-bit) quantization. Several existing sub 2-bit post-training quantization (PTQ) methods utilize a mix-precision scheme by leveraging an unstructured fine-grained mask to explicitly distinguish sali…

2025

Perceptually Constrained Precipitation Nowcasting Model

ICML 2025poster

Most current precipitation nowcasting methods aim to capture the underlying spatiotemporal dynamics of precipitation systems by minimizing the mean square error (MSE). However, these methods often neglect effective constraints on the data distribution, leading to unsatisfactory prediction accuracy a…

Cited by 0SourcePDFScholar
2025

Pilot: Building the Federated Multimodal Instruction Tuning Framework

AAAI 2025technical

In this paper, we explore a novel federated multimodal instruction tuning task(FedMIT), which is significant for collaboratively fine-tuning MLLMs on different types of multimodal instruction data on distributed devices. To solve the new task, we propose a federated multimodal instruction tuning fra…

Cited by 1SourcePDFScholar
2025

RS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation Model

ICCV 2025poster

Remote sensing foundation models largely break away from the traditional paradigm of designing task-specific models, offering greater scalability across multiple tasks. However, they face challenges such as low computational efficiency and limited interpretability, especially when dealing with large…

Cited by 0SourcePDFScholar
2025

STE-Mamba: Automated Multimodal Depression Detection through Emotional Analysis and Spatio-Temporal Information Ensemble

ICASSP 2025accepted

Automatic Depression Detection (ADD) garners widespread attention due to its convenience and objectivity. While existing research makes significant progress, challenges remain. First, most current ADD methods struggle to balance computational overhead and prediction accuracy. Second, these methods p…

Cited by 0SourceScholar
2025

Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues

CVPR 2025poster

Understanding human behavior and the environmental information in the egocentric video is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-person view. Leveraging the corresponding exocentric video to provide global context has…

2025

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data

NeurIPS 2025spotlight

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial uncertainty and data scarcity, limiting the 3D spatial reasoning capab…

Cited by 0SourceScholar
2025

Unsupervised Degradation Representation Aware Transform for Real-World Blind Image Super-Resolution

AAAI 2025technical

Blind image super-resolution (blind SR) aims to restore a high-resolution (HR) image from a low-resolution (LR) image with unknown degradation. Many existing methods explicitly estimate degradation information from various LR images. However, in most cases, image degradations are independent of imag…

2025

Video Language Model Pretraining with Spatio-temporal Masking

CVPR 2025poster

The development of self-supervised video-language models based on mask learning has significantly advanced downstream video tasks. These models leverage masked reconstruction to facilitate joint learning of visual and linguistic information. However, recent study reveals that reconstructing image fe…

Cited by 0SourcePDFScholar
2025

VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension

ACL 2025long

Assessing the video comprehension capabilities of multimodal AI systems can effectively measure their understanding and reasoning abilities. Most video evaluation benchmarks are limited to a single language, typically English, and predominantly feature videos rooted in Western cultural contexts. In…

Cited by 0SourcePDFScholar
2025

Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors

EMNLP 2025

The misuse of large language models (LLMs), such as academic plagiarism, has driven the development of detectors to identify LLM-generated texts. To bypass these detectors, paraphrase attacks have emerged to purposely rewrite these texts to evade detection. Despite the success, existing methods requ

2024

Clip-Based Synergistic Knowledge Transfer for text-based Person Retrieval

ICASSP 2024accepted

Text-based Person Retrieval (TPR) aims to retrieve the target person images given a textual query. The primary challenge lies in bridging the substantial gap between vision and language modalities, especially when dealing with limited large-scale datasets. In this paper, we introduce a CLIP-based Sy…

Cited by 0SourceScholar
2024

CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition

CVPR 2024poster

Over the past decade most methods in visual place recognition (VPR) have used neural networks to produce feature representations. These networks typically produce a global representation of a place image using only this image itself and neglect the cross-image variations (e.g. viewpoint and illumina…

2024

Feature Distribution Matching by Optimal Transport for Effective and Robust Coreset Selection

AAAI 2024technical

Training neural networks with good generalization requires large computational costs in many deep learning methods due to large-scale datasets and over-parameterized models. Despite the emergence of a number of coreset selection methods to reduce the computational costs, the problem of coreset distr…

2024

HARDVS: Revisiting Human Activity Recognition with Dynamic Vision Sensors

AAAI 2024technical

The main streams of human activity recognition (HAR) algorithms are developed based on RGB cameras which usually suffer from illumination, fast motion, privacy preservation, and large energy consumption. Meanwhile, the biologically inspired event cameras attracted great interest due to their unique…

2024

LG-VQ: Language-Guided Codebook Learning

NeurIPS 2024poster

Vector quantization (VQ) is a key technique in high-resolution and high-fidelity image synthesis, which aims to learn a codebook to encode an image with a sequence of discrete codes and then generate an image in an auto-regression manner. Although existing methods have shown superior performance,…

Cited by 3SourcePDFScholar
2024

M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

NeurIPS 2024poster

This paper presents M$^3$GPT, an advanced $\textbf{M}$ultimodal, $\textbf{M}$ultitask framework for $\textbf{M}$otion comprehension and generation. M$^3$GPT operates on three fundamental principles. The first focuses on creating a unified representation space for various motion-relevant modalities…

2024

MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection

IJCAI 2024poster

Popular transformer-based detectors detect objects in a one-to-one manner, where both the bounding box and category of each object are predicted only by the single query, leading to the box-sensitive category predictions. Additionally, the initialization of positional queries solely based on the pre…

2024

Modality-Collaborative Test-Time Adaptation for Action Recognition

CVPR 2024poster

Video-based Unsupervised Domain Adaptation (VUDA) method improves the generalization of the video model enabling it to be applied to action recognition tasks in different environments. However these methods require continuous access to source data during the adaptation process which are impractical…

Cited by 6SourcePDFScholar
2024

Multi-Factor Adaptive Vision Selection for Egocentric Video Question Answering

ICML 2024poster

The challenge of interpreting the world from a human perspective in Artificial Intelligence (AI) is particularly evident in egocentric video question answering, which grapples with issues like small object recognition, noise suppression, and spatial-temporal reasoning. To address these challenges, w…

2024

OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling

NeurIPS 2024poster

Constrained by the separate encoding of vision and language, existing grounding and referring segmentation works heavily rely on bulky Transformer-based fusion en-/decoders and a variety of early-stage interaction technologies. Simultaneously, the current mask visual language modeling (MVLM) fails t…

2024

Prompt-Driven Dynamic Object-Centric Learning for Single Domain Generalization

CVPR 2024poster

Single-domain generalization aims to learn a model from single source domain data attaining generalized performance on other unseen target domains. Existing works primarily focus on improving the generalization ability of static networks. However static networks are unable to dynamically adapt to th…

Cited by 12SourcePDFScholar
2024

RTracker: Recoverable Tracking via PN Tree Structured Memory

CVPR 2024poster

Existing tracking methods mainly focus on learning better target representation or developing more robust prediction models to improve tracking performance. While tracking performance has significantly improved the target loss issue occurs frequently due to tracking failures complete occlusion or ou…

2024

Regressor-Segmenter Mutual Prompt Learning for Crowd Counting

CVPR 2024poster

Crowd counting has achieved significant progress by training regressors to predict instance positions. In heavily crowded scenarios however regressors are challenged by uncontrollable annotation variance which causes density map bias and context information inaccuracy. In this study we propose mutua…

2024

Reshaping the Online Data Buffering and Organizing Mechanism for Continual Test-Time Adaptation

ECCV 2024poster

"Continual Test-Time Adaptation (CTTA) involves adapting a pre-trained source model to continually changing unsupervised target domains. In this paper, we systematically analyze the challenges of this task: online environment, unsupervised nature, and the risks of error accumulation and catastrophic…

2024

StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion

ECCV 2024poster

"Story visualization aims to generate a series of realistic and coherent images based on a storyline. Current models adopt a frame-by-frame architecture by transforming the pre-trained text-to-image model into an auto-regressive manner. Although these models have shown notable progress, there are st…

2024

Towards Robust and Efficient Cloud-Edge Elastic Model Adaptation via Selective Entropy Distillation

ICLR 2024poster

The conventional deep learning paradigm often involves training a deep model on a server and then deploying the model or its distilled ones to resource-limited edge devices. Usually, the models shall remain fixed once deployed (at least for some period) due to the potential high cost of model adapta…

2024

Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition

ICLR 2024poster

Recent studies show that vision models pre-trained in generic visual learning tasks with large-scale data can provide useful feature representations for a wide range of visual perception problems. However, few attempts have been made to exploit pre-trained foundation models in visual place recogniti…

2024

VMamba: Visual State Space Model

NeurIPS 2024spotlight

Designing computationally efficient network architectures remains an ongoing necessity in computer vision. In this paper, we adapt Mamba, a state-space language model, into VMamba, a vision backbone with linear time complexity. At the core of VMamba is a stack of Visual State-Space (VSS) blocks with…

2023

AsyFOD: An Asymmetric Adaptation Paradigm for Few-Shot Domain Adaptive Object Detection

CVPR 2023poster

In this work, we study few-shot domain adaptive object detection (FSDAOD), where only a few target labeled images are available for training in addition to sufficient source labeled images. Critically, in FSDAOD, the data-scarcity in the target domain leads to an extreme data imbalance between the s…

2023

CIGAR: Cross-Modality Graph Reasoning for Domain Adaptive Object Detection

CVPR 2023poster

Unsupervised domain adaptive object detection (UDA-OD) aims to learn a detector by generalizing knowledge from a labeled source domain to an unlabeled target domain. Though the existing graph-based methods for UDA-OD perform well in some cases, they cannot learn a proper node set for the graph. In a…

Cited by 34SourcePDFScholar
2023

CiteTracker: Correlating Image and Text for Visual Tracking

ICCV 2023poster

Existing visual tracking methods typically take an image patch as the reference of the target to perform tracking. However, a single image patch cannot provide a complete and precise concept of the target object as images are limited in their ability to abstract and can be ambiguous, which makes it…

Cited by 51PDFcodeScholar
2023

Digging out Discrimination Information from Generated Samples for Robust Visual Question Answering

ACL 2023findings

Visual Question Answering (VQA) aims to answer a textual question based on a given image. Nevertheless, recent studies have shown that VQA models tend to capture the biases to answer the question, instead of using the reasoning ability, resulting in poor generalisation ability. To alleviate the issu…

Cited by 9SourcePDFScholar
2023

Integrally Pre-Trained Transformer Pyramid Networks

CVPR 2023poster

In this paper, we present an integral pre-training framework based on masked image modeling (MIM). We advocate for pre-training the backbone and neck jointly so that the transfer gap between MIM and downstream recognition tasks is minimal. We make two technical contributions. First, we unify the rec…

2023

Isolation and Impartial Aggregation: A Paradigm of Incremental Learning without Interference

AAAI 2023technical

This paper focuses on the prevalent stage interference and stage performance imbalance of incremental learning. To avoid obvious stage learning bottlenecks, we propose a new incremental learning framework, which leverages a series of stage-isolated classifiers to perform the learning task at each st…

2023

KERM: Knowledge Enhanced Reasoning for Vision-and-Language Navigation

CVPR 2023poster

Vision-and-language navigation (VLN) is the task to enable an embodied agent to navigate to a remote location following the natural language instruction in real scenes. Most of the previous approaches utilize the entire features or object-centric features to represent navigable candidates. However,…

2023

Learned Distributed Image Compression with Multi-Scale Patch Matching in Feature Domain

AAAI 2023technical

Beyond achieving higher compression efficiency over classical image compression codecs, deep image compression is expected to be improved with additional side information, e.g., another image from a different perspective of the same scene. To better utilize the side information under the distributed…

Cited by 13SourcePDFScholar
2023

Learning Mask-aware CLIP Representations for Zero-Shot Segmentation

NeurIPS 2023poster

Recently, pre-trained vision-language models have been increasingly used to tackle the challenging zero-shot segmentation task. Typical solutions follow the paradigm of first generating mask proposals and then adopting CLIP to classify them. To maintain the CLIP's zero-shot transferability, previous…

2023

Manifold-Aware Self-Training for Unsupervised Domain Adaptation on Regressing 6D Object Pose

IJCAI 2023poster

Domain gap between synthetic and real data in visual regression (e.g., 6D pose estimation) is bridged in this paper via global feature alignment and local refinement on the coarse classification of discretized anchor classes in target space, which imposes a piece-wise target manifold regularization…

2023

MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning

ACL 2023long

Supervised visual captioning models typically require a large scale of images or videos paired with descriptions in a specific language (i.e., the vision-caption pairs) for training. However, collecting and labeling large-scale datasets is time-consuming and expensive for many scenarios and language…

2023

Recurrent Fine-Grained Self-Attention Network for Video Crowd Counting

ICASSP 2023accepted

Striking a balance between exploring the spatio-temporal correlation and controlling model complexity is vital for video-based crowd counting methods. In this paper, we propose a Recurrent Fine-Grained Self-Attention Network (RFSNet) to achieve efficient and accurate counting in video scenes via the…

Cited by 0SourceScholar
2023

Spikformer: When Spiking Neural Network Meets Transformer

ICLR 2023poster

We consider two biologically plausible structures, the Spiking Neural Network (SNN) and the self-attention mechanism. The former offers an energy-efficient and event-driven paradigm for deep learning, while the latter has the ability to capture feature dependencies, enabling Transformer to achieve g…

2023

Strip-MLP: Efficient Token Interaction for Vision MLP

ICCV 2023poster

Token interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token interaction on the spatial dimension is highly dependent on the spatial resolution of the feature maps, which limits the m…

Cited by 13PDFcodeScholar
2023

Unlearnable Clusters: Towards Label-Agnostic Unlearnable Examples

CVPR 2023poster

There is a growing interest in developing unlearnable examples (UEs) against visual privacy leaks on the Internet. UEs are training samples added with invisible but unlearnable noise, which have been found can prevent unauthorized training of machine learning models. UEs typically are generated via…

2022

DAS: Densely-Anchored Sampling for Deep Metric Learning

ECCV 2022poster

"Deep Metric Learning (DML) serves to learn an embedding function to project semantically similar data into nearby embedding space and plays a vital role in many applications, such as image retrieval and face recognition. However, the performance of DML methods often highly depends on sampling metho…

2022

Fine-Grained Object Classification via Self-Supervised Pose Alignment

CVPR 2022poster

Semantic patterns of fine-grained objects are determined by subtle appearance difference of local parts, which thus inspires a number of part-based methods. However, due to uncontrollable object poses in images, distinctive details carried by local regions can be spatially distributed or even self-o…

Cited by 78PDFcodeScholar
2022

M5Product: Self-Harmonized Contrastive Learning for E-Commercial Multi-Modal Pretraining

CVPR 2022poster

Despite the potential of multi-modal pre-training to learn highly discriminative feature representations from complementary data modalities, current progress is being slowed by the lack of large-scale modality-diverse datasets. By leveraging the natural suitability of E-commerce, where different mod…

Cited by 44PDFcodeScholar
2022

Mixed-Precision Neural Network Quantization via Learned Layer-Wise Importance

ECCV 2022poster

"The exponentially large discrete search space in mixed-precision quantization (MPQ) makes it hard to determine the optimal bit-width for each layer. Previous works usually resort to iterative search methods on the training set, which consume hundreds or even thousands of GPU-hours. In this study, w…

2022

Towards End-to-End Image Compression and Analysis with Transformers

AAAI 2022technical

We propose an end-to-end image compression and analysis model with Transformers, targeting to the cloud-based image classification application. Instead of placing an existing Transformer-based image classification model directly after an image codec, we aim to redesign the Vision Transformer (ViT) m…

2021

Conformer: Local Features Coupling Global Representations for Visual Recognition

ICCV 2021poster

Within Convolutional Neural Network (CNN), the convolution operations are good at extracting local features but experience difficulty to capture global representations. Within visual transformer, the cascaded self-attention modules can capture long-distance feature dependencies but unfortunately det…

Cited by 890PDFcodeScholar
2021

Contrastive Neural Architecture Search With Neural Architecture Comparators

CVPR 2021poster

One of the key steps in Neural Architecture Search (NAS) is to estimate the performance of candidate architectures. Existing methods either directly use the validation performance or learn a predictor to estimate the performance. However, these methods can be either computationally expensive or very…

Cited by 87PDFcodeScholar
2021

Direct Measure Matching for Crowd Counting

IJCAI 2021poster

Traditional crowd counting approaches usually use Gaussian assumption to generate pseudo density ground truth, which suffers from problems like inaccurate estimation of the Gaussian kernel sizes. In this paper, we propose a new measure-based counting approach to regress the predicted density maps to…

Cited by 48SourcePDFScholar
2021

Hierarchically and Cooperatively Learning Traffic Signal Control

AAAI 2021technical

Deep reinforcement learning (RL) has been applied to traffic signal control recently and demonstrated superior performance to conventional control methods. However, there are still several challenges we have to address before fully applying deep RL to traffic signal control. Firstly, the objective o…

Cited by 78SourcePDFScholar
2021

Learning Scalable lY=-Constrained Near-Lossless Image Compression via Joint Lossy Image and Residual Compression

CVPR 2021poster

We propose a novel joint lossy image and residual compression framework for learning l_infinity-constrained near-lossless image compression. Specifically, we obtain a lossy reconstruction of the raw image through lossy image compression and uniformly quantize the corresponding residual to satisfy a…

Cited by 33PDFScholar
2021

Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and Benchmark

CVPR 2021poster

Tracking by natural language specification is a new rising research topic that aims at locating the target object in the video sequence based on its language description. Compared with traditional bounding box (BBox) based tracking, this setting guides object tracking with high-level semantic inform…

Cited by 220PDFScholar
2020

An Asymmetric Modeling for Action Assessment

ECCV 2020poster

Action assessment is a task of assessing the performance of an action. It is widely applicable to many real-world scenarios such as medical treatment and sporting events. However, existing methods for action assessment are mostly limited to individual actions, especially lacking modeling of the asym…

Cited by 58SourcePDFScholar
2020

Large Batch Optimization for Object Detection: Training COCO in 12 Minutes

ECCV 2020poster

Most of existing object detectors usually adopt a small training batch size ( ~16), which severely hinders the whole community from exploring large-scale datasets due to the extremely long training procedure. In this paper, we propose a versatile large batch optimization framework for object detecti…

2019

Transductive Episodic-Wise Adaptive Metric for Few-Shot Learning

ICCV 2019poster

Few-shot learning, which aims at extracting new concepts rapidly from extremely few examples of novel classes, has been featured into the meta-learning paradigm recently. Yet, the key challenge of how to learn a generalizable classifier with the capability of adapting to specific tasks with severely…

Cited by 248PDFScholar
2017

Exploiting Multi-Grain Ranking Constraints for Precisely Searching Visually-Similar Vehicles

ICCV 2017poster

Precise search of visually-similar vehicles poses a great challenge in computer vision, which needs to find exactly the same vehicle among a massive vehicles with visually similar appearances for a given query image. In this paper, we model the relationship of vehicle images as multiple grains. Foll…

Cited by 130PDFScholar
2017

Learning Long-Term Dependencies for Action Recognition With a Biologically-Inspired Deep Network

ICCV 2017poster

Despite a lot of research efforts devoted in recent years, how to efficiently learn long-term dependencies from sequences still remains a pretty challenging task. As one of the key models for sequence learning, recurrent neural network (RNN) and its variants such as long short term memory (LSTM) and…

Cited by 85PDFcodeScholar
2016

Unsupervised Cross-Dataset Transfer Learning for Person Re-Identification

CVPR 2016poster

Most existing person re-identification (Re-ID) approaches follow a supervised learning framework, in which a large number of labelled matching pairs are required for training. This severely limits their scalability in real-world applications. To overcome this limitation, we develop a novel cross-dat…

Cited by 457PDFScholar