← Search

Wei Feng

109 accepted papers

2026

AutoPP: Towards Automated Product Poster Generation and Optimization

AAAI 2026technical

Product posters blend striking visuals with informative text to highlight the product and capture customer attention. However, crafting appealing posters and manually optimizing them based on online performance is laborious and resource-consuming. To address this, we introduce AutoPP, an automated p

Cited by 0SourcePDFScholar
2026

Beyond Visual Reconstruction Quality: Object Perception-aware 3D Gaussian Splatting for Autonomous Driving

ICLR 2026poster

Reconstruction techniques, such as 3D Gaussian Splatting (3DGS), are increasingly used for generating scenarios in autonomous driving system (ADS) research. Existing 3DGS-based works for autonomous driving scenario generation have, through various optimizations, achieved high visual similarity in re…

Cited by 0SourcecodeScholar
2026

Debiased and Denoised Projection Learning for Incomplete Multi-view Clustering

ICLR 2026poster

Multi-view clustering achieves outstanding performance but relies on the assumption of complete multi-view samples. However, certain views may be partially unavailable due to failures during acquisition or storage, resulting in distribution shifts across views. Although some incomplete multi-view cl…

Cited by 0SourceScholar
2026

Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models

CVPR 2026

Generating realistic and user-preferred advertisements is a key challenge in e-commerce. Existing approaches utilize multiple independent models driven by click-through-rate (CTR) to controllably create attractive image or text advertisements. However, their pipelines lack cross-modal perception and

Cited by 0SourcecodeScholar
2026

Discriminative Graph Embedding Framework via Label-Free Marginal Fisher Analysis

AAAI 2026technical

Marginal Fisher Analysis (MFA) is a classical dimensionality reduction (DR) method that leverages dual graphs to capture intra-class compactness and inter-class separability. However, MFA’s reliance on high-quality labels limits its practical application. For another, existing unsupervised DR method

Cited by 0SourcePDFScholar
2026

E-Logic Prompt: Unified Energy-Logic Framework for Continual Visual Question Answering

AAAI 2026technical

Prompt tuning has shown promise for continual visual question answering (CVQA), facilitating modular and transferable knowledge across tasks. However, existing approaches often overlook the guiding role of prompts in the model’s implicit reasoning process. This oversight can lead to inconsistent re

Cited by 0SourcePDFScholar
2026

EgoRoC: Towards Egocentric Robotic Control via Task-Agnostic Visual Alignment

CVPR 2026

Recent Vision-Language-Action (VLA) models map visual-textual inputs to robotic actions via end-to-end architectures, yet this approach entangles visual understanding with task-specific actions. This leads to an exhaustive collection of full operational sequences and parameter redundancy across task

Cited by 0SourceScholar
2026

Exposing Mixture and Annotating Confusion for Active Universal Test-Time Adaptation

ICLR 2026poster

Universal Test-Time Adaptation (UTTA) tackles the challenge of handling both class and domain shifts in unsupervised settings with stream testing data. Currently, most UTTA methods can only deal with minor shifts and heavily rely on heuristic approaches. To advance UTTA under dual shifts, we propose…

Cited by 0SourceScholar
2026

Federated Incomplete Multi-View Clustering with Tensorized Low-Rank Constraint

AAAI 2026technical

Federated Multi-View Clustering has gained increasing attention for its ability to discover complementary clustering structures of distributed multi-view data while preserving data privacy. However, real-world clients often only have access to partial views, and the view incompleteness poses great c

Cited by 0SourcePDFScholar
2026

HOPS: Hierarchical Open-vocabulary Part Segmentation with Attention-Aware Filtering and Affinity-Guided Enhancement

CVPR 2026

Open-vocabulary part segmentation (OVPS) aims to segment objects into fine-grained parts while generalizing to unseen categories. Existing VLM-based methods face two challenges: (1) object over-segmentation, caused by overly broad semantic activations, and (2) part under-segmentation, resulting from

Cited by 0SourcecodeScholar
2026

InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation

CVPR 2026

E-commerce product poster generation aims to automatically synthesize a single image that effectively conveys product information by presenting a subject, text, and a designed style. Recent diffusion models with fine-grained and efficient controllability have advanced product poster synthesis, yet t

Cited by 0SourceScholar
2026

KD-CVG: A KNOWLEDGE-DRIVEN APPROACH FOR CREATIVE VIDEO GENERATION

ICASSP 2026poster

Creative Generation (CG) leverages generative models to automatically produce advertising content that highlights product features, and it has been a significant focus of recent research. However, while CG has advanced considerably, most efforts have concentrated on generating advertising text and i…

Cited by 0SourcePDFScholar
2026

MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation

AAAI 2026technical

Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural g

Cited by 0SourcePDFScholar
2026

Multi-modal Frequency Decomposition Network for Semantic Scene Completion

CVPR 2026

Based on an RGB-D image pair, semantic scene completion (SSC) provides a description for 3D scene understanding by predicting 3D semantic occupancy map. Recent methods extract RGB-D multi-modal features and fuse them in spatial domain, which disregards the misalignment caused by the imperfect raw mu

Cited by 0SourceScholar
2026

NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection

CVPR 2026

Despite the remarkable progress in open-vocabulary object detection (OVD), a significant gap remains between the training and testing phases. During training, the RPN and RoI heads often misclassify unlabeled novel-category objects as background, causing some proposals to be prematurely filtered out

Cited by 0SourceScholar
2026

OBJVanish: Prompt-Driven Generation of Physically Realizable 3D LiDAR-Invisible Objects

ICML 2026poster

LiDAR-based 3D object detectors are fundamental to autonomous driving, where missed detections pose severe safety risks. While adversarial attacks are crucial for evaluating the robustness of these detectors, existing point-level perturbation methods rarely cause complete object disappearance and pr…

Cited by 0SourceScholar
2026

PRISM: Progressive Robust Learning for Open-World Continual Category Discovery

ICLR 2026poster

Continual Category Discovery (CCD) aims to leverage models trained on known categories to automatically discover novel category concepts from continuously arriving streams of unlabeled data, while retaining the ability to recognize previously known classes. Despite recent progress, existing methods…

Cited by 0SourceScholar
2026

RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving Scenes

CVPR 2026

High-quality video editing and processing are crucial in domains such as filmmaking and autonomous driving, where accurate visual refinement and data preparation are essential. However, it is challenging to achieve precise control over dynamic objects while maintaining spatiotemporal consistency. Cu

Cited by 0SourcecodeScholar
2026

RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers

AAAI 2026technical

The Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled diffusion transformer methods incur significant parameter and computational overheads and suffer from inefficient resource

Cited by 0SourcePDFScholar
2026

Rethinking Cross-Modal Anchor Alignment for Mitigating Error Accumulation

CVPR 2026

Mitigating noisy correspondence in cross-modal matching poses a serious challenge due to the problem of error accumulation. Existing methods primarily attribute this accumulation to errors caused by noisy sample pairs. However, a novel source of error from clean sample pairs (also termed anchor pair

Cited by 0SourceScholar
2026

Seeing Through the Shift: Causality-Inspired Robust Generalized Category Discovery

CVPR 2026

Generalized Category Discovery (GCD) aims to transfer knowledge from known categories to automatically discover new, unseen ones while preserving recognition of the known classes. Despite recent progress, existing GCD approaches typically assume that all data are drawn from the same distribution, wh

Cited by 0SourceScholar
2026

Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding

CVPR 2026

Recent advancements in multimodal large reasoning models (MLRMs) have significantly improved performance in visual question answering. However, we observe that transition words (e.g., because, however, and wait) are closely associated with hallucinations and tend to exhibit high-entropy states. We a

Cited by 0SourcecodeScholar
2026

UV-RGS: Relightable 3D Gaussian Splatting from Unposed Views Under Varied Illuminations

AAAI 2026technical

The latest advancements in scene relighting have been predominantly driven by inverse rendering with 3D Gaussian Splatting (3DGS). However, existing methods remain overly reliant on precise camera parameters under static illumination conditions, which is prohibitively expensive and even impractical

Cited by 0SourcePDFScholar
2026

VSRELL: A Simple Baseline for Video Super-Resolution and Enhancement in Low-Light Environment

CVPR 2026

We propose an integrated learning scheme of Video Super-Resolution and Enhancement in Low-Light environment, named VSRELL, which aims to recover Well-Illuminated High-Resolution (WIHR) sequence from Low-Light Low-Resolution (LLLR) counterparts. Due to the complex coupling of multiple degradations, t

Cited by 0SourcecodeScholar
2026

Write Where It Matters: Policy-Guided Watermarks for 3D Gaussian Splatting

CVPR 2026

Recent advances in 3D Gaussian Splatting (3DGS) enable photorealistic real-time rendering but also increase the risks of unauthorized copying and redistribution. Existing 3DGS watermarking methods typically rely on handcrafted thresholds or globally fixed hyperparameters to balance invisibility and

Cited by 0SourceScholar
2026

iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models

ICLR 2026poster

Recent methods have made notable progress in accelerating Large Vision-Language Models (LVLMs) by exploiting the inherent redundancy in visual inputs. Most existing approaches, however, focus narrowly on reducing image tokens before or within the Large Language Model (LLM) stage to lower computation…

Cited by 0SourcecodeScholar
2025

A Block Term Decomposition Model Based Algorithm for Tensor Completion of Multidimensional Harmonic Signals

ICASSP 2025accepted

We consider tensor data completion of an incomplete observation of multidimensional harmonic (MH) signals. Unlike existing tensor-based techniques for MH retrieval (MHR), which mostly adopt the canonical polyadic decomposition (CPD) to model the simple "one-to-one" correspondence among harmonics acr…

Cited by 0SourceScholar
2025

A Parametric Non-Negative Coupled Canonical Polyadic Decomposition Algorithm for Hyperspectral Super-Resolution

ICASSP 2025accepted

Recently, coupled tensor decomposition has been widely used in data fusion of a hyperspectral image (HSI) and a multispectral image (MSI) for hyperspectral super-resolution (HSR). However, exsiting works often ignore the inherent non-negative (NN) property of the image data, or impose the NN constra…

Cited by 0SourceScholar
2025

Beyond Background Shift: Rethinking Instance Replay in Continual Semantic Segmentation

CVPR 2025poster

In this work, we focus on continual semantic segmentation (CSS), where segmentation networks are required to continuously learn new classes without erasing knowledge of previously learned ones. Although storing images of old classes and directly incorporating them into the training of new models has…

2025

Contrastive Multi-view Subspace Clustering via Tensor Transformers Autoencoder

AAAI 2025technical

Multi-view clustering aims to identify consistent and complementary information across multiple views to partition data into clusters, emerging as a popular unsupervised method for multi-view data analysis. However, existing methods often design view-specific encoders to extract distinct features fr…

Cited by 0SourcePDFScholar
2025

Deep Multi-modal Graph Clustering via Graph Transformer Network

AAAI 2025technical

Current deep multi-modal graph clustering methods primarily rely on Graph Neural Network (GNN) to fully exploit attribute features and graph structures, including message propagation and low-dimensional feature embedding. However, these methods lack further exploration of graph structural informatio…

Cited by 0SourcePDFScholar
2025

Dual Semantic Guidance for Open Vocabulary Semantic Segmentation

CVPR 2025poster

Open-vocabulary semantic segmentation aims to enable models to segment arbitrary categories. Currently, though pre-trained Vision-Language Models (VLMs) like CLIP have established a robust foundation for this task by learning to match text and image representations from large-scale data, their lack…

Cited by 0SourcePDFScholar
2025

Enhanced Unsupervised Discriminant Dimensionality Reduction for Nonlinear Data

IJCAI 2025

Linear Discriminant Analysis (LDA) is a classical supervised dimensionality reduction algorithm. However, LDA focuses more on global structure and overly depends on reliable data labels. For data with outliers and nonlinear structures, LDA cannot effectively capture the true structure of the data. M

Cited by 0SourcePDFScholar
2025

Enhancing Interpretable Image Classification Through LLM Agents and Conditional Concept Bottleneck Models

ACL 2025long

Concept Bottleneck Models (CBMs) decompose image classification into a process governed by interpretable, human-readable concepts. Recent advances in CBMs have used Large Language Models (LLMs) to generate candidate concepts. However, a critical question remains: What is the optimal number of concep…

Cited by 0SourcePDFScholar
2025

Fair Incomplete Multi-View Clustering via Distribution Alignment

IJCAI 2025

Incomplete multi-view clustering (IMVC) extracts consistent and complementary information from multi-source/modality data with missing views, aiming to partition the data into different clusters. It can effectively address the problem of unsupervised multi-source data analysis in complex environment

Cited by 0SourcePDFScholar
2025

FedAGC: Federated Continual Learning with Asymmetric Gradient Correction

ICCV 2025poster

Federated Continual Learning (FCL) has emerged as a prominent distributed learning paradigm and aims at addressing model learning challenges in both federated and continual learning settings. Efficient personalization in FCL remains a major challenge, as it must handle not only conflicts between old…

Cited by 0SourcePDFScholar
2025

GReg: Geometry-Aware Region Refinement for Sign Language Video Generation

ICCV 2025poster

Sign Language Video Generation (SLVG) aims to transform sign language sequences into natural and fluent sign language videos. Existing SLVG methods lack geometric modeling of human anatomical structures, leading to anatomically implausible and temporally inconsistent generation. To address these cha…

Cited by 0SourcePDFScholar
2025

Generate E-commerce Product Background by Integrating Category Commonality and Personalized Style

ICASSP 2025accepted

The state-of-the-art methods for e-commerce product background generation suffer from the inefficiency of designing product-wise prompts when scaling up the production, as well as the ineffectiveness of describing fine-grained styles when customizing personalized backgrounds for some specific brands…

Cited by 0SourceScholar
2025

Generative Hard Example Augmentation for Semantic Point Cloud Segmentation

CVPR 2025poster

The recent progress in semantic point cloud segmentation is attributed to deep networks, which require a large amount of point cloud data for training. However, how to collect substantial point-wise annotations of the point clouds at affordable cost for the end-to-end network training still needs to…

Cited by 0SourcePDFScholar
2025

GeoFlow-SLAM: A Robust Tightly-Coupled RGBD-Inertial and Legged Odometry Fusion SLAM for Dynamic Legged Robotics

IROS 2025

This paper presents GeoFlow-SLAM, a robust and effective Tightly-Coupled RGBD-Inertial and Legged Odometry Fusion SLAM for legged robotics undergoing aggressive and high-frequency motions. By integrating geometric consistency, legged odometry constraints, and dual-stream optical flow (GeoFlow), our

Cited by 1SourcecodeScholar
2025

Hypergraph Clustering Network with Partial Attribute Imputation

ICCV 2025poster

Existing hypergraph clustering methods typically assume that node attributes are fully available. However, in real-world scenarios, missing node attributes are common for the sake of privacy or due to data noise. While some approaches attempt to handle missing attributes in traditional graphs, they…

Cited by 0SourcePDFScholar
2025

Improving Generalization in Federated Learning with Highly Heterogeneous Data via Momentum-Based Stochastic Controlled Weight Averaging

ICML 2025poster

For federated learning (FL) algorithms such as FedSAM, their generalization capability is crucial for real-word applications. In this paper, we revisit the generalization problem in FL and investigate the impact of data heterogeneity on FL generalization. We find that FedSAM usually performs worse t…

Cited by 0SourcePDFScholar
2025

Investigating Numerical Translation with Large Language Models

ICASSP 2025accepted

The inaccurate translation of numbers can lead to significant security issues, ranging from financial setbacks to medical inaccuracies. While large language models (LLMs) have made significant advancements in machine translation, their capacity for translating numbers has not been thoroughly explore…

Cited by 0SourceScholar
2025

QBasicVSR: Temporal Awareness Adaptation Quantization for Video Super-Resolution

NeurIPS 2025poster

While model quantization has become pivotal for deploying super-resolution (SR) networks on mobile devices, existing works focus on quantization methods only for image super-resolution. Different from image super-resolution, the temporal error propagation, shared temporal parameterization, and tempo…

Cited by 0SourceScholar
2025

SU-RGS: Relightable 3D Gaussian Splatting from Sparse Views under Unconstrained Illuminations

ICCV 2025poster

The latest advancements in scene relighting have been predominantly driven by inverse rendering with 3D Gaussian Splatting (3DGS). However, existing methods remain overly reliant on densely sampled images under static illumination conditions, which is prohibitively expensive and even impractical in…

Cited by 0SourcePDFScholar
2025

Scalable Federated One-Step Multi-View Clustering with Tensorized Regularization

AAAI 2025technical

Multi-view clustering (MVC) methods have garnered considerable attention within centralized data frameworks. However, real-world multi-view data are often collected and stored by different organizations, complicating the practical deployment of MVC and motivating the emergence of federated multi-vie…

Cited by 0SourcePDFScholar
2025

Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding

CVPR 2025poster

Recent advancements in multimodal large language models (MLLMs) have significantly improved performance in visual question answering. However, they often suffer from hallucinations. In this work, hallucinations are categorized into two main types: initial hallucinations and snowball hallucinations.…

Cited by 0SourcePDFScholar
2025

TorchTitan: One-stop PyTorch native solution for production ready LLM pretraining

ICLR 2025poster

The development of large language models (LLMs) has been instrumental in advancing state-of-the-art natural language processing applications. Training LLMs with billions of parameters and trillions of tokens requires sophisticated distributed systems that enable composing and comparing several state…

2024

Deep Correlated Prompting for Visual Recognition with Missing Modalities

NeurIPS 2024poster

Large-scale multimodal models have shown excellent performance over a series of tasks powered by the large corpus of paired multimodal training data. Generally, they are always assumed to receive modality-complete inputs. However, this simple assumption may not always hold in the real world due to p…

2024

Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language Recognition

COLING 2024main

Skeleton-aware sign language recognition (SLR) has gained popularity due to its ability to remain unaffected by background information and its lower computational requirements. Current methods utilize spatial graph modules and temporal modules to capture spatial and temporal features, respectively.…

2024

Efficient Federated Multi-View Clustering with Integrated Matrix Factorization and K-Means

IJCAI 2024poster

Multi-view clustering is a popular unsupervised multi-view learning method. Real-world multi-view data are often distributed across multiple entities, presenting a challenge for performing multi-view clustering. Federated learning provides a solution by enabling multiple entities to collaboratively…

Cited by 1SourcePDFScholar
2024

Federated Multi-View Clustering via Tensor Factorization

IJCAI 2024poster

Multi-view clustering is an effective method to process massive unlabeled multi-view data. Since data of different views may be collected and held by different parties, it becomes impractical to train a multi-view clustering model in a centralized way, for the sake of privacy. However, federated mul…

Cited by 1SourcePDFScholar
2024

Hunting Attributes: Context Prototype-Aware Learning for Weakly Supervised Semantic Segmentation

CVPR 2024poster

Recent weakly supervised semantic segmentation (WSSS) methods strive to incorporate contextual knowledge to improve the completeness of class activation maps (CAM). In this work we argue that the knowledge bias between instances and contexts affects the capability of the prototype to sufficiently un…

2024

LRR: Language-Driven Resamplable Continuous Representation against Adversarial Tracking Attacks

ICLR 2024poster

Visual object tracking plays a critical role in visual-based autonomous systems, as it aims to estimate the position and size of the object of interest within a live video. Despite significant progress made in this field, state-of-the-art (SOTA) trackers often fail when faced with adversarial pertur…

2024

Long-Tailed Learning as Multi-Objective Optimization

AAAI 2024technical

Real-world data is extremely imbalanced and presents a long-tailed distribution, resulting in models biased towards classes with sufficient samples and performing poorly on rare classes. Recent methods propose to rebalance classes but they undertake the seesaw dilemma (what is increasing performance…

2024

Partial Multi-View Clustering via Self-Supervised Network

AAAI 2024technical

Partial multi-view clustering is a challenging and practical research problem for data analysis in real-world applications, due to the potential data missing issue in different views. However, most existing methods have not fully explored the correlation information among various incomplete views. I…

Cited by 7SourcePDFScholar
2024

Reconstruction Weighting Principal Component Analysis with Fusion Contrastive Learning

IJCAI 2024poster

Principal component analysis (PCA) is a popular unsupervised dimensionality reduction method to extract the principal components of data. However, there are two problems with the existing PCA: (1) Traditional PCA methods treat each sample equally and ignore sample differences. (2) They fail to extra…

2024

Robust Multi-Robot Global Localization with Unknown Initial Pose based on Neighbor Constraints

ICRA 2024poster

Multi-robot global localization (MR-GL) with unknown initial positions in a large scale environment is a challenging task. The key point is the data association between different robots’ viewpoints. It also makes traditional Appearance-based localization methods unusable. Recently, researchers have…

Cited by 1SourceScholar
2024

SAVSR: Arbitrary-Scale Video Super-Resolution via a Learned Scale-Adaptive Network

AAAI 2024technical

Deep learning-based video super-resolution (VSR) networks have gained significant performance improvements in recent years. However, existing VSR networks can only support a fixed integer scale super-resolution task, and when we want to perform VSR at multiple scales, we need to train several models…

2024

Towards Reliable Advertising Image Generation Using Human Feedback

ECCV 2024poster

"In the e-commerce realm, compelling advertising images are pivotal for attracting customer attention. While generative models automate image generation, they often produce substandard images that may mislead customers and require significant labor costs to inspect. This paper delves into increasing…

2024

VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding

NeurIPS 2024poster

Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding that hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine-grained VCMR benchmark requiring methods to localize the best-matched moment from…

2023

Continuous Sign Language Recognition With Correlation Network

CVPR 2023poster

Human body trajectories are a salient cue to identify actions in video. Such body trajectories are mainly conveyed by hands and face across consecutive frames in sign language. However, current methods in continuous sign language recognition(CSLR) usually process frames independently to capture fram…

2023

Leveraging Inpainting for Single-Image Shadow Removal

ICCV 2023poster

Fully-supervised shadow removal methods achieve the best restoration qualities on public datasets but still generate some shadow remnants. One of the reasons is the lack of large-scale shadow & shadow-free image pairs. Unsupervised methods can alleviate the issue but their restoration qualities are…

Cited by 28PDFcodeScholar
2023

Measuring Asymmetric Gradient Discrepancy in Parallel Continual Learning

ICCV 2023poster

In Parallel Continual Learning (PCL), the parallel multiple tasks start and end training unpredictably, thus suffering from training conflict and catastrophic forgetting issues. The two issues are raised because the gradients from parallel tasks differ in directions and magnitudes. Thus, in this pap…

Cited by 15PDFcodeScholar
2023

NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding

NeurIPS 2023poster

The application of deep learning to nursing procedure activity understanding has the potential to greatly enhance the quality and safety of nurse-patient interactions. By utilizing the technique, we can facilitate training and education, improve quality control, and enable operational compliance mon…

2023

OFCOURSE: A Multi-Agent Reinforcement Learning Environment for Order Fulfillment

NeurIPS 2023poster

The dramatic growth of global e-commerce has led to a surge in demand for efficient and cost-effective order fulfillment which can increase customers' service levels and sellers' competitiveness. However, managing order fulfillment is challenging due to a series of interdependent online sequential d…

2023

Open Compound Domain Adaptation with Object Style Compensation for Semantic Segmentation

NeurIPS 2023poster

Many methods of semantic image segmentation have borrowed the success of open compound domain adaptation. They minimize the style gap between the images of source and target domains, more easily predicting the accurate pseudo annotations for target domain's images that train segmentation network. Th…

Cited by 8SourcePDFScholar
2023

Self-Emphasizing Network for Continuous Sign Language Recognition

AAAI 2023technical

Hand and face play an important role in expressing sign language. Their features are usually especially leveraged to improve system performance. However, to effectively extract visual representations and capture trajectories for hands and face, previous methods always come at high computations with…

2023

Social Relation Reasoning Based on Triangular Constraints

AAAI 2023technical

Social networks are essentially in a graph structure where persons act as nodes and the edges connecting nodes denote social relations. The prediction of social relations, therefore, relies on the context in graphs to model the higher-order constraints among relations, which has not been exploited s…

Cited by 9SourcePDFScholar
2023

Unsupervised Domain Adaptation for Medical Image Segmentation by Selective Entropy Constraints and Adaptive Semantic Alignment

AAAI 2023technical

Generalizing a deep learning model to new domains is crucial for computer-aided medical diagnosis systems. Most existing unsupervised domain adaptation methods have made significant progress in reducing the domain distribution gap through adversarial training. However, these methods may still produc…

2022

Can You Spot the Chameleon? Adversarially Camouflaging Images From Co-Salient Object Detection

CVPR 2022poster

Co-salient object detection (CoSOD) has recently achieved significant progress and played a key role in retrieval-related tasks. However, it inevitably poses an entirely new safety and security issue, i.e., highly personal and sensitive content can potentially be extracting by powerful CoSOD methods…

Cited by 25PDFcodeScholar
2022

Connecting the Complementary-View Videos: Joint Camera Identification and Subject Association

CVPR 2022poster

We attempt to connect the data from complementary views, i.e., top view from drone-mounted cameras in the air, and side view from wearable cameras on the ground. Collaborative analysis of such complementary-view data can facilitate to build the air-ground cooperative visual system for various kinds…

Cited by 13PDFcodeScholar
2022

Exploring Example Influence in Continual Learning

NeurIPS 2022accept

Continual Learning (CL) sequentially learns new tasks like human beings, with the goal to achieve better Stability (S, remembering past tasks) and Plasticity (P, adapting to new tasks). Due to the fact that past training data is not available, it is valuable to explore the influence difference on S…

2022

MISF: Multi-Level Interactive Siamese Filtering for High-Fidelity Image Inpainting

CVPR 2022poster

Although achieving significant progress, existing deep generative inpainting methods still show low generalization across different scenes. As a result, the generated images usually contain artifacts or the filled pixels differ greatly from the ground truth, making them far from real-world applicati…

Cited by 111PDFcodeScholar
2022

Panoramic Human Activity Recognition

ECCV 2022poster

"To obtain a more comprehensive activity understanding for a crowded scene, in this paper, we propose a new problem of panoramic human activity recognition (PAR), which aims to simultaneously achieve the the recognition of individual actions, social group activities, and global activities. This is a…

2022

Regularized Modal Regression on Markov-Dependent Observations: A Theoretical Assessment

AAAI 2022technical

Modal regression, a widely used regression protocol, has been extensively investigated in statistical and machine learning communities due to its robustness to outlier and heavy-tailed noises. Understanding modal regression's theoretical behavior can be fundamental in learning theory. Despite signif…

Cited by 1SourcePDFScholar
2022

Rethinking Video Rain Streak Removal: A New Synthesis Model and a Deraining Network with Video Rain Prior

ECCV 2022poster

"Existing video synthetic models and deraining methods are mostly built on a simplified video rain model assuming that rain streak layers of different video frames are uncorrelated, thereby producing degraded performance on real-world rainy videos. To address this problem, we devise a new video rain…

2022

Self-Supervised Social Relation Representation for Human Group Detection

ECCV 2022poster

"Human group detection, which splits crowd of people into groups, is an important step for video-based human social activity analysis. The core of human group detection is the human social relation representation and division. In this paper, we propose a new two-stage multi-head framework for human…

2022

Temporal Lift Pooling for Continuous Sign Language Recognition

ECCV 2022poster

"Pooling methods are necessities for modern neural networks for increasing receptive fields and lowering down computational costs. However, commonly used hand-crafted pooling approaches, e.g. max pooling and average pooling, may not well preserve discriminative features. While many researchers have…

2021

Auto-Exposure Fusion for Single-Image Shadow Removal

CVPR 2021poster

Shadow removal is still a challenging task due to its inherent background-dependent and spatial-variant properties, leading to unknown and diverse shadow patterns. Even powerful deep neural networks could hardly recover traceless shadow-removed background. This paper proposes a new solution for this…

Cited by 174PDFcodeScholar
2021

EfficientDeRain: Learning Pixel-wise Dilation Filtering for High-Efficiency Single-Image Deraining

AAAI 2021technical

Single-image deraining is rather challenging due to the unknown rain model. Existing methods often make specific assumptions of the rain model, which can hardly cover many diverse circumstances in the real world, compelling them to employ complex optimization or progressive refinement. This, however…

2021

Multi-Domain Multi-Task Rehearsal for Lifelong Learning

AAAI 2021technical

Rehearsal, seeking to remind the model by storing old knowledge in lifelong learning, is one of the most effective ways to mitigate catastrophic forgetting, i.e., biased forgetting of previous knowledge when moving to new tasks. However, the old tasks of the most previous rehearsal-based methods suf…

Cited by 31SourcePDFScholar
2021

VIL-100: A New Dataset and a Baseline Model for Video Instance Lane Detection

ICCV 2021poster

Lane detection plays a key role in autonomous driving. While car cameras always take streaming videos on the way, current lane detection works mainly focus on individual images (frames) by ignoring dynamics along the video. In this work, we collect a new video instance lane detection (VIL-100) datas…

Cited by 63PDFcodeScholar
2020

A Multi-Task Mean Teacher for Semi-Supervised Shadow Detection

CVPR 2020poster

Existing shadow detection methods suffer from an intrinsic limitation in relying on limited labeled datasets, and they may produce poor results in some complicated situations. To boost the shadow detection performance, this paper presents a multi-task mean teacher model for semi-supervised shadow de…

Cited by 190PDFcodeScholar
2020

Dynamically Pruned Message Passing Networks for Large-scale Knowledge Graph Reasoning

ICLR 2020poster

We propose Dynamically Pruned Message Passing Networks (DPMPN) for large-scale knowledge graph reasoning. In contrast to existing models, embedding-based or path-based, we learn an input-dependent subgraph to explicitly model a sequential reasoning process. Each subgraph is dynamically constructed,…

Cited by 86SourcecodeScholar
2020

Key Action and Joint CTC-Attention based Sign Language Recognition

ICASSP 2020accepted

Sign Language Recognition (SLR) translates sign language video into natural language. In practice, sign language video, owning a large number of redundant frames, is necessary to be selected the essential. However, unlike common video that describes actions, sign language video is characterized as c…

Cited by 0SourceScholar
2020

SPARK: Spatial-aware Online Incremental Attack Against Visual Tracking

ECCV 2020poster

Adversarial attacks of deep neural networks have been intensively studied on image, audio, natural language, patch, and pixel classification tasks. Nevertheless, as a typical, while important real-world application, the adversarial attacks of online video object tracking that traces an object's movi…

Cited by 112SourcePDFScholar
2020

Watch out! Motion is Blurring the Vision of Your Deep Neural Networks

NeurIPS 2020poster

The state-of-the-art deep neural networks (DNNs) are vulnerable against adversarial examples with additive random-like noise perturbations. While such examples are hardly found in the physical world, the image blurring effect caused by object motion, on the other hand, commonly occurs in practice, m…

2019

Learned Map Prediction for Enhanced Mobile Robot Exploration

ICRA 2019poster

We demonstrate an autonomous ground robot capable of exploring unknown indoor environments for reconstructing their 2D maps. This problem has been traditionally tackled by geometric heuristics and information theory. More recently, deep learning and reinforcement learning based approaches have been…

Cited by 123SourceScholar
2019

TextDragon: An End-to-End Framework for Arbitrary Shaped Text Spotting

ICCV 2019poster

Most existing text spotting methods either focus on horizontal/oriented texts or perform arbitrary shaped text spotting with character-level annotations. In this paper, we propose a novel text spotting framework to detect and recognize text of arbitrary shapes in an end-to-end manner, using only wor…

Cited by 255PDFScholar
2017

Frequency-tuned ACM for biomedical image segmentation

ICASSP 2017accepted

Biomedical images are usually corrupted by strong noise and intensity inhomogeneity simultaneously. Existing region-based active contour models (RACMs) easily fail when segmenting such images. In the frequency domain, we propose a generalized RACM that presents a new way to understand the essence of…

Cited by 0SourceScholar
2017

Learning Dynamic Siamese Network for Visual Object Tracking

ICCV 2017poster

How to effectively learn temporal variation of target appearance, to exclude the interference of cluttered background, while maintaining real-time response, is an essential problem of visual object tracking. Recently, Siamese networks have shown great potentials of matching based trackers in achievi…

Cited by 1046PDFScholar
2015

Fine-Grained Change Detection of Misaligned Scenes With Varied Illuminations

ICCV 2015poster

Detecting fine-grained subtle changes among a scene is critically important in practice. Previous change detection methods, focusing on detecting large-scale significant changes, cannot do this well. This paper proposes a feasible end-to-end approach to this challenging problem. We start from active…

Cited by 43PDFScholar