← Search

Lin Wang

101 accepted papers

2026

Beyond Single-Point Perturbation: A Hierarchical, Manifold-Aware Approach to Diffusion Attacks

AAAI 2026technical

Latent Diffusion Models have become a powerful tool for generating high-fidelity unrestricted adversarial examples. However, the existing methods typically perturb only the initial latent or rely on prompt engineering, which is ill-suited to the iterative nature of the diffusion process, plus optimi

Cited by 0SourcePDFScholar
2026

Beyond the Trade-off: Unifying Fairness and Performance in Federated Learning

ICML 2026poster

Federated Learning (FL) often suffers from a trade-off between global model performance and client-level fairness due to data heterogeneity, which often leads to inconsistent performance of the globally trained models, resulting in unfair outcomes among users. Existing fair FL algorithms face a trad…

Cited by 0SourceScholar
2026

CLEX: Complementary Label Exchange Learning for Noisy Facial Expression Recognition

CVPR 2026

Facial expression recognition (FER) in the wild is severely hampered by label noise and annotation ambiguity. Existing methods, including sample selection, label ensembling, and consistency regularization, primarily rely on ordinary label supervision and offer limited control over non-target predict

Cited by 0SourceScholar
2026

CT2BSE: 3D BSE Microstructural Image Cross-Device Generation from µCT for Cement Hydration via Voxel Swin Transformer

IJCAI 2026

Acquiring three-dimensional(3D) microstructural images of cement hydration reveals critical microscale features essential for understanding hydration mechanisms and advancing material development. Micro-computed tomography(µCT) is widely used to capture these images due to its non-destructive, repea

Cited by 0Scholar
2026

Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection

CVPR 2026

Open-Vocabulary Temporal Action Detection (OV-TAD) aims to classify and localize action segments in untrimmed videos for unseen categories. Previous methods rely solely on global alignment between label-level semantics and visual features, which is insufficient to transfer temporal consistent visual

Cited by 0SourceScholar
2026

Emergent Co-Adaptive Strategies in Heterogeneous Multi-Robot Systems Via Meta-Learning

ICRA 2026poster

Abstract— As teamed robots increasingly share public spaces with humans, the ability to co-adapt—to mutually adjust behavior in response to one another—becomes essential for safe, efficient, and socially acceptable operation. This paper introduces a socially co-adaptive framework for heterogeneous m…

Cited by 0Scholar
2026

Ev-iCRF: Self-supervised Event-guided iCRF Estimation for HDR Image Reconstruction

AAAI 2026technical

In this paper, we present Ev-iCRF, a novel self-supervised pipeline for high dynamic range (HDR) image reconstruction from a single-exposure low dynamic range (LDR) image, guided by asynchronous event streams generated by a bio-inspired event camera. The highlight of Ev-iCRF lies in its formulation

Cited by 0SourcePDFScholar
2026

EvDiff3D: Event-Aware Diffusion Repair for High-Fidelity Event-Based 3D Reconstruction

AAAI 2026technical

Event cameras are bio-inspired sensors that capture visual information through asynchronous brightness changes, offering distinct advantages including high temporal resolution and wide dynamic range. While prior research has investigated event-based 3D reconstruction for extreme scenarios, existing

Cited by 0SourcePDFScholar
2026

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs) typically assume a uniform spatial fidelity across the entire field of view of visual inputs, dedicating equal precision to even the uninformative regions. By contrast, human vision is neither uniform nor static. It is adaptive, selective, and resource-efficient. In lig

Cited by 0SourceScholar
2026

MOC: Multi-Order Communication in LLM-based Multi-Agent Systems

ICML 2026poster

Despite the remarkable progress of Large Language Model (LLM) based Multi-Agent Systems, most research focuses on optimizing coordination topology while largely underexploring the equally critical problem: how to transmit and optimize messages among agents effectively? Current communication schemes …

Cited by 0SourceScholar
2026

OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera

CVPR 2026

Robust 3D semantic occupancy is essential for legged and humanoid robots, yet most Semantic Scene Completion (SSC) systems are built for wheeled platforms with forward-facing sensors. We present OneOcc, a vision-only panoramic SSC framework tailored to severe body jitter and 360deg continuity. OneOc

Cited by 0SourcecodeScholar
2026

PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving Systems

AAAI 2026technical

Multimodal Large Language Models (MLLMs) are becoming integral to autonomous driving (AD) systems due to their strong vision-language reasoning capabilities. However, MLLMs are vulnerable to adversarial attacks—particularly adversarial patch attacks—which can pose serious threats in real-world scen

Cited by 0SourcePDFScholar
2026

STEDiff: Revealing the Spatial and Temporal Redundancy of Backdoor Attacks in Text-to-Image Diffusion Models

ICLR 2026poster

Recently, diffusion models have been recognized as state-of-the-art models for image generation due to their ability to produce high-quality images. However, recent studies have shown that diffusion models are susceptible to backdoor attacks, where an attacker can activate hidden biases using a spec…

Cited by 0SourcecodeScholar
2026

Temporal and Spatial Representation Learning for Multimodal Low-Beam 3D Object Detection

AAAI 2026technical

To facilitate the large-scale deployment of autonomous driving in real-world scenarios, developing low-cost and high-performance 3D object detection systems has become a critical technical challenge. Although high-beam LiDARs provide denser point cloud data, their prohibitive hardware cost and high

Cited by 0SourcePDFScholar
2025

A Method for Removing Reflections from Water Surface Images Based on Pre-trained Image Restoration

ICASSP 2025accepted

Reflections on the water surface hinder the extraction of valuable information from water surface images. To remove reflections from water surface images, we construct a synthetic dataset and propose a multi-task network for water surface reflection detection and removal. Specifically, we first use…

Cited by 0SourceScholar
2025

AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts

ICCV 2025poster

Despite rapid advancements in text-to-image (T2I) models, their safety mechanisms are vulnerable to adversarial prompts, which maliciously generate unsafe images. Current red-teaming methods for proactively assessing such vulnerabilities usually require white-box access to T2I models, and rely on in…

Cited by 0SourcePDFScholar
2025

BiCo-Fusion: Bidirectional Complementary LiDAR-Camera Fusion for Semantic- and Spatial-Aware 3D Object Detection

RA-L 2025

3D object detection is an important task that has been widely applied in autonomous driving. To perform this task, a new trend is to fuse multi-modal inputs, i.e., LiDAR and camera. Under such a trend, recent methods fuse these two modalities by unifying them in the same 3D space. However, during di

Cited by 16SourceScholar
2025

CUBE360: Learning Cubic Field Representation for Monocular Panoramic Depth Estimation

RA-L 2025

Panoramic depth estimation presents significant challenges due to the severe distortion caused by equirectangular projection (ERP) and the limited availability of panoramic RGB-D datasets. Inspired by the recent success of neural rendering, we propose a self-supervised method, named CUBE360, that le

Cited by 0SourceScholar
2025

Client2Vec: Improving Federated Learning by Distribution Shifts Aware Client Indexing

ICCV 2025poster

Federated Learning (FL) is a privacy-preserving distributed machine learning paradigm. Nonetheless, the substantial distribution shifts among clients pose a considerable challenge to the performance of current FL algorithms. To mitigate this challenge, various methods have been proposed to enhance t…

2025

DAP-LED: Learning Degradation-Aware Priors with Clip for Joint Low-Light Enhancement and Deblurring

ICRA 2025

Autonomous vehicles and robots often struggle with reliable visual perception at night due to the low illumination and motion blur caused by the long exposure time of RGB cameras. Existing methods address this challenge by sequentially connecting the off-the-shelf pretrained lowlight enhancement and

Cited by 5SourcecodeScholar
2025

Fair Text-Attributed Graph Representation Learning

EMNLP 2025

Text-Attributed Graphs (TAGs), which integrate text and graph structures, have recently gained traction, especially in web applications. However, as a graph structure, TAG representation learning (TAGRL) naturally inherits issues from Graph Neural Networks (GNNs), such as fairness. Moreover, previou

Cited by 0SourcePDFScholar
2025

Foresee and Act Ahead: Task Prediction and Pre-Scheduling Enabled Efficient Robotic Warehousing

ICRA 2025

In warehousing systems, to enhance efficiency amid surging demand volumes, much attention has been placed on how to reasonably allocate tasks of delivery to robots. However, the labor of robots is still inevitably wasted to some extent. In this paper, we propose a pre-scheduling enhanced warehousing

Cited by 1SourceScholar
2025

Leader360V: A Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment

NeurIPS 2025poster

360 video captures the complete surrounding scenes with the ultra-large field of view of 360x180. This makes 360 scene understanding tasks, *e.g.*, segmentation and tracking, crucial for appications, such as autonomous driving, robotics. With the recent emergence of foundation models, the community…

Cited by 0SourcecodeScholar
2025

PanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial Augmentation

CVPR 2025poster

Recently, Depth Anything Models (DAMs) - a type of depth foundation models - have demonstrated impressive zero-shot capabilities across diverse perspective images. Despite its success, it remains an open question regarding DAMs' performance on panorama images that enjoy a large field-of-view (180x36…

Cited by 0SourcePDFScholar
2025

PolypSense3D: A Multi-Source Benchmark Dataset for Depth-Aware Polyp Size Measurement in Endoscopy

NeurIPS 2025poster

Accurate polyp sizing during endoscopy is crucial for cancer risk assessment but is hindered by subjective methods and inadequate datasets lacking integrated 2D appearance, 3D structure, and real-world size information. We introduce PolypSense3D, the first multi-source benchmark dataset specifically…

Cited by 0SourcecodeScholar
2025

ProteinConformers: Benchmark Dataset for Simulating Protein Conformational Landscape Diversity and Plausibility

NeurIPS 2025poster

Understanding the conformational landscape of proteins is essential for elucidating protein function and facilitating drug design. However, existing protein conformation benchmarks fail to capture the full energy landscape, limiting their ability to evaluate the diversity and physical plausibility o…

Cited by 0SourcecodeScholar
2025

Robo-GS: A Physics Consistent Spatial-Temporal Model for Robotic Arm with Hybrid Representation

ICRA 2025

The Real2Sim2Real (R2S2R) paradigm is critical for advancing robotic learning. Existing methods lack a comprehensive solution to accurately reconstruct real-world objects with both spatial representations and their associated physics attributes in the Real2Sim stage. We propose a Real2Sim pipeline t

Cited by 73SourceScholar
2025

SDAFE: A Dual-filter Stable Diffusion Data Augmentation Method for Facial Expression Recognition

ICASSP 2025accepted

Facial expressions are a powerful medium for conveying emotions. In facial expression recognition (FER) field, the difficulty of collecting specific expressions often leads to class imbalance in mainstream datasets, significantly reducing the classification accuracy of deep neural networks. To addre…

Cited by 0SourceScholar
2025

Single Pump-Valve Pneumatic Actuation With Continuous Flow Rate Control for Soft Robots

RA-L 2025

Pneumatic actuated soft robots attract increasing interest of the researchers due to the availability and simplicity in actuation. The soft robots driven by soft pneumatic actuators (SPAs) of various active volumes demand pneumatic systems with various range of flow rate. However, the usually bulky

Cited by 4SourceScholar
2025

TimeTracker: Event-based Continuous Point Tracking for Video Frame Interpolation with Non-linear Motion

CVPR 2025poster

Video frame interpolation (VFI) that leverages the bio-inspired event cameras as guidance has recently shown better performance and memory efficiency than the frame-based methods, thanks to the event cameras' advantages, such as high temporal resolution. A hurdle for event-based VFI is how to effect…

Cited by 0SourcePDFScholar
2024

"UniINR: Event-guided Unified Rolling Shutter Correction, Deblurring, and Interpolation"

ECCV 2024poster

"Video frames captured by rolling shutter (RS) cameras during fast camera movement frequently exhibit RS distortion and blur simultaneously. Naturally, recovering high-frame-rate global shutter (GS) sharp frames from an RS blur frame must simultaneously consider RS correction, deblur, and frame inte…

2024

Benchmarking Implicit Neural Representation and Geometric Rendering in Real-Time RGB-D SLAM

CVPR 2024poster

Implicit neural representation (INR) in combination with geometric rendering has recently been employed in real-time dense RGB-D SLAM. Despite active research endeavors being made there lacks a unified protocol for fair evaluation impeding the evolution of this area. In this work we establish to our…

2024

Chasing Day and Night: Towards Robust and Efficient All-Day Object Detection Guided by an Event Camera

ICRA 2024poster

The ability to detect objects in all lighting (i.e., normal-, over-, and under-exposed) conditions is crucial for real-world applications, such as self-driving. Traditional RGB-based detectors often fail under such varying lighting conditions. Therefore, recent works utilize novel event cameras to s…

Cited by 18SourcecodeScholar
2024

Co-Occ: Coupling Explicit Feature Fusion With Volume Rendering Regularization for Multi-Modal 3D Semantic Occupancy Prediction

RA-L 2024

3D semantic occupancy prediction is a pivotal task in the field of autonomous driving. Recent approaches have made great advances in 3D semantic occupancy predictions on a single modality. However, multi-modal semantic occupancy prediction approaches have encountered difficulties in dealing with the

Cited by 58SourceScholar
2024

Cooking-Clip: Context-Aware Language-Image Pretraining for Zero-Shot Recipe Generation

ICASSP 2024accepted

Cooking is one of the oldest and the most common human activities in everyone’s daily life. Instructional cooking videos have also become one of the most common data sources for multimodal visual understanding researches. Compared to other domains, multimodal cooking videos: 1. not only have signifi…

Cited by 0SourceScholar
2024

DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling

ECCV 2024poster

"Text-to-3D scene generation holds immense potential for the gaming, film, and architecture sectors. Despite significant progress, existing methods struggle with maintaining high quality, consistency, and editing flexibility. In this paper, we propose , a 3D Gaussian-based novel text-to-3D scene gen…

2024

Elite360D: Towards Efficient 360 Depth Estimation via Semantic- and Distance-Aware Bi-Projection Fusion

CVPR 2024poster

360 depth estimation has recently received great attention for 3D reconstruction owing to its omnidirectional field of view (FoV). Recent approaches are predominantly focused on cross-projection fusion with geometry-based re-projection: they fuse 360 images with equirectangular projection (ERP) and…

Cited by 9SourcePDFScholar
2024

EventBind: Learning a Unified Representation to Bind Them All for Event-based Open-world Understanding

ECCV 2024poster

"In this paper, we propose EventBind, a novel and effective framework that unleashes the potential of vision-language models (VLMs) for event-based recognition to compensate for the lack of large-scale event-based datasets. In particular, due to the distinct modality gap with the image-text data and…

Cited by 9SourcePDFScholar
2024

EventDance: Unsupervised Source-free Cross-modal Adaptation for Event-based Object Recognition

CVPR 2024poster

In this paper we make the first attempt at achieving the cross-modal (i.e. image-to-events) adaptation for event-based object recognition without accessing any labeled source image data owning to privacy and commercial issues. Tackling this novel problem is non-trivial due to the novelty of event ca…

Cited by 11SourcePDFScholar
2024

ExACT: Language-guided Conceptual Reasoning and Uncertainty Estimation for Event-based Action Recognition and More

CVPR 2024highlight

Event cameras have recently been shown beneficial for practical vision tasks such as action recognition thanks to their high temporal resolution power efficiency and reduced privacy concerns. However current research is hindered by 1) the difficulty in processing events because of their prolonged du…

Cited by 23SourcePDFScholar
2024

Exploring Targeted Universal Adversarial Attack for Deep Hashing

ICASSP 2024accepted

Although image-dependent adversarial attacks have been studied, the more challenging image-agnostic adversarial attack for deep hashing remains an unexplored territory. In this paper, we take the first attempt on the more efficient and malicious targeted universal adversarial attack (TUAA) for deep…

Cited by 0SourceScholar
2024

GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-aware Panoramic Semantic Segmentation

CVPR 2024poster

This paper tackles a novel yet challenging problem: how to transfer knowledge from the emerging Segment Anything Model (SAM) -- which reveals impressive zero-shot instance segmentation capacity -- to learn a compact panoramic semantic segmentation model i.e. student without requiring any labeled dat…

Cited by 7SourcePDFScholar
2024

LaSe-E2V: Towards Language-guided Semantic-aware Event-to-Video Reconstruction

NeurIPS 2024poster

Event cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video (E2V) reconstruction to bridge event-based and standard comput…

Cited by 2SourcePDFScholar
2024

LinNet: Linear Network for Efficient Point Cloud Representation Learning

NeurIPS 2024poster

Point-based methods have made significant progress, but improving their scalability in large-scale 3D scenes is still a challenging problem. In this paper, we delve into the point-based method and develop a simpler, faster, stronger variant model, dubbed as LinNet. In particular, we first propose th…

Cited by 1SourcePDFScholar
2024

OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding

ECCV 2024poster

"Surgical scene perception via videos is critical for advancing robotic surgery, telesurgery, and AI-assisted surgery, particularly in ophthalmology. However, the scarcity of diverse and richly annotated video datasets has hindered the development of intelligent systems for surgical workflow analysi…

2024

SRFNet: Monocular Depth Estimation with Fine-grained Structure via Spatial Reliability-oriented Fusion of Frames and Events

ICRA 2024poster

Monocular depth estimation is a crucial task to measure distance relative to a camera, which is important for applications, such as robot navigation and self-driving. Traditional frame-based methods suffer from performance drops due to the limited dynamic range and motion blur. Therefore, recent wor…

Cited by 8SourcecodeScholar
2024

Semantics Distortion and Style Matter: Towards Source-free UDA for Panoramic Segmentation

CVPR 2024poster

This paper addresses an interesting yet challenging problem-- source-free unsupervised domain adaptation (SFUDA) for pinhole-to-panoramic semantic segmentation--given only a pinhole image-trained model (i.e. source) and unlabeled panoramic images (i.e. target). Tackling this problem is nontrivial du…

Cited by 14SourcePDFScholar
2024

Towards Dynamic and Small Objects Refinement for Unsupervised Domain Adaptative Nighttime Semantic Segmentation

IROS 2024poster

Nighttime semantic segmentation plays a crucial role in practical applications, such as autonomous driving, where it frequently encounters difficulties caused by inadequate illumination conditions and the absence of well-annotated datasets. Moreover, semantic segmentation models trained on daytime d…

Cited by 2SourcecodeScholar
2024

Towards Robust Event-guided Low-Light Image Enhancement: A Large-Scale Real-World Event-Image Dataset and Novel Approach

CVPR 2024poster

Event camera has recently received much attention for low-light image enhancement (LIE) thanks to their distinct advantages such as high dynamic range. However current research is prohibitively restricted by the lack of large-scale real-world and spatial-temporally aligned event-image datasets. To t…

2024

Transformer-CNN Cohort: Semi-supervised Semantic Segmentation by the Best of Both Students

ICRA 2024poster

The popular methods for semi-supervised semantic segmentation mostly adopt a unitary network model using convolutional neural networks (CNNs) and enforce consistency of the model’s predictions over perturbations applied to the inputs or model. However, such a learning paradigm suffers from two criti…

Cited by 19SourcecodeScholar
2024

UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All

CVPR 2024poster

We present UniBind a flexible and efficient approach that learns a unified representation space for seven diverse modalities-- images text audio point cloud thermal video and event data. Existing works eg. ImageBind treat the image as the central modality and build an image-centered representation s…

Cited by 13SourcePDFScholar
2024

When Cohesion Lies in the Embedding Space: Embedding-Based Reference-Free Metrics for Topic Segmentation

COLING 2024main

In this paper we propose a new framework and new methods for the reference-free evaluation of topic segmentation systems directly in the embedding space. Specifically, we define a common framework for reference-free, embedding-based topic segmentation metrics, and show how this applies to an existin…

Cited by 2SourcePDFScholar
2023

A Good Student is Cooperative and Reliable: CNN-Transformer Collaborative Learning for Semantic Segmentation

ICCV 2023poster

In this paper, we strive to answer the question 'how to collaboratively learn convolutional neural network (CNN)-based and vision transformer (ViT)-based models by selecting and exchanging the reliable knowledge between them for semantic segmentation?' Accordingly, we propose an online knowledge dis…

Cited by 36PDFScholar
2023

Benchmarking and Analyzing Robust Point Cloud Recognition: Bag of Tricks for Defending Adversarial Examples

ICCV 2023poster

Deep Neural Networks (DNNs) for 3D point cloud recognition are vulnerable to adversarial examples, threatening their practical deployment. Despite the many research endeavors have been made to tackle this issue in recent years, the diversity of adversarial examples on 3D point clouds makes them more…

Cited by 5PDFcodeScholar
2023

Both Style and Distortion Matter: Dual-Path Unsupervised Domain Adaptation for Panoramic Semantic Segmentation

CVPR 2023poster

The ability of scene understanding has sparked active research for panoramic image semantic segmentation. However, the performance is hampered by distortion of the equirectangular projection (ERP) and a lack of pixel-wise annotations. For this reason, some works treat the ERP and pinhole images equa…

Cited by 36SourcePDFScholar
2023

DELTA: Diverse Client Sampling for Fasting Federated Learning

NeurIPS 2023poster

Partial client participation has been widely adopted in Federated Learning (FL) to reduce the communication burden efficiently. However, an inadequate client sampling scheme can lead to the selection of unrepresentative subsets, resulting in significant variance in model updates and slowed convergen…

2023

HRDFuse: Monocular 360deg Depth Estimation by Collaboratively Learning Holistic-With-Regional Depth Distributions

CVPR 2023poster

Depth estimation from a monocular 360 image is a burgeoning problem owing to its holistic sensing of a scene. Recently, some methods, e.g., OmniFusion, have applied the tangent projection (TP) to represent a 360 image and predicted depth values via patch-wise regressions, which are merged to get a d…

Cited by 32SourcePDFScholar
2023

Improved Event-Based Dense Depth Estimation via Optical Flow Compensation

ICRA 2023poster

Event cameras have the potential to overcome the limitations of classical computer vision in real-world applications. Depth estimation is a crucial step for high-level robotics tasks and has attracted much attention from the community. In this paper, we propose an event-based dense depth estimation…

Cited by 7SourceScholar
2023

Learning Spatial-Temporal Implicit Neural Representations for Event-Guided Video Super-Resolution

CVPR 2023poster

Event cameras sense the intensity changes asynchronously and produce event streams with high dynamic range and low latency. This has inspired research endeavors utilizing events to guide the challenging video super-resolution (VSR) task. In this paper, we make the first at tempt to address a novel p…

2023

Look at the Neighbor: Distortion-aware Unsupervised Domain Adaptation for Panoramic Semantic Segmentation

ICCV 2023poster

Endeavors have been recently made to transfer knowledge from the labeled pinhole image domain to the unlabeled panoramic image domain via Unsupervised Domain Adaptation (UDA). The aim is to tackle the domain gaps caused by the style disparities and distortion problem of the non-uniformly distributed…

Cited by 25PDFScholar
2023

NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding

NeurIPS 2023poster

The application of deep learning to nursing procedure activity understanding has the potential to greatly enhance the quality and safety of nurse-patient interactions. By utilizing the technique, we can facilitate training and education, improve quality control, and enable operational compliance mon…

2023

OPT-GAN: A Broad-Spectrum Global Optimizer for Black-Box Problems by Learning Distribution

AAAI 2023technical

Black-box optimization (BBO) algorithms are concerned with finding the best solutions for problems with missing analytical details. Most classical methods for such problems are based on strong and fixed a priori assumptions, such as Gaussianity. However, the complex real-world problems, especially w…

2023

OmniZoomer: Learning to Move and Zoom in on Sphere at High-Resolution

ICCV 2023poster

Omnidirectional images (ODIs) have become increasingly popular, as their large field-of-view (FoV) can offer viewers the chance to freely choose the view directions in immersive environments such as virtual reality. The Mobius transformation is typically employed to further provide the opportunity f…

Cited by 9PDFcodeScholar
2023

Patch-Mix Transformer for Unsupervised Domain Adaptation: A Game Perspective

CVPR 2023highlight

Endeavors have been recently made to leverage the vision transformer (ViT) for the challenging unsupervised domain adaptation (UDA) task. They typically adopt the cross-attention in ViT for direct domain alignment. However, as the performance of cross-attention highly relies on the quality of pseudo…

2023

Pixel Is All You Need: Adversarial Trajectory-Ensemble Active Learning for Salient Object Detection

AAAI 2023technical

Although weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by…

Cited by 10SourcePDFScholar
2023

SEPT: Towards Scalable and Efficient Visual Pre-training

AAAI 2023technical

Recently, the self-supervised pre-training paradigm has shown great potential in leveraging large-scale unlabeled data to improve downstream task performance. However, increasing the scale of unlabeled pre-training data in real-world scenarios requires prohibitive computational costs and faces the c…

Cited by 1SourcePDFScholar
2023

STS-GAN: Can We Synthesize Solid Texture with High Fidelity from Arbitrary 2D Exemplar?

IJCAI 2023poster

Solid texture synthesis (STS), an effective way to extend a 2D exemplar to a 3D solid volume, exhibits advantages in computational photography. However, existing methods generally fail to accurately learn arbitrary textures, which may result in the failure to synthesize solid textures with high fide…

Cited by 12SourcePDFScholar
2023

Unsupervised Domain Adaptation for Medical Image Segmentation by Selective Entropy Constraints and Adaptive Semantic Alignment

AAAI 2023technical

Generalizing a deep learning model to new domains is crucial for computer-aided medical diagnosis systems. Most existing unsupervised domain adaptation methods have made significant progress in reducing the domain distribution gap through adversarial training. However, these methods may still produc…

2022

BIPS: Bi-modal Indoor Panorama Synthesis via Residual Depth-Aided Adversarial Learning

ECCV 2022poster

"Providing omnidirectional depth along with RGB information is important for numerous applications. However, as omnidirectional RGB-D data is not always available, synthesizing RGB-D panorama data from limited information of a scene can be useful. Therefore, some prior works tried to synthesize RGB…

2022

CS-GResNet: A Simple and Highly Efficient Network for Facial Expression Recognition

ICASSP 2022accepted

Facial expression recognition (FER) has recently attracted attention in computer vision. However, existing methods mostly focus on the explicit performance and overlook their computational resources and memory consumption. Hence, achieving promising performance while maintaining the efficiency of mo…

Cited by 0SourceScholar
2022

Deconvolutional Density Network: Modeling Free-Form Conditional Distributions

AAAI 2022technical

Conditional density estimation (CDE) is the task of estimating the probability of an event conditioned on some inputs. A neural network (NN) can also be used to compute the output distribution for continuous-domain, which can be viewed as an extension of regression task. Nevertheless, it is difficul…

2022

Efficient Video Deblurring Guided by Motion Magnitude

ECCV 2022poster

"Video deblurring is a highly under-constrained problem due to the spatially and temporally varying blur. An intuitive approach for video deblurring includes two steps: a) detecting the blurry region in the current frame; b) utilizing the information from clear regions in adjacent frames for current…

2022

Prototype-Based Inter-Camera Learning for Person Re-Identification

ICASSP 2022accepted

Person re-identification (ReID) aims at retrieving images of the same person across non-overlapping camera views. The prior works focus on either fully supervised or unsupervised ReID settings, and achieve remarkable performances. In real scenarios, however, the major annotation cost comes from matc…

Cited by 0SourceScholar
2022

SphereSR: 360deg Image Super-Resolution With Arbitrary Projection via Continuous Spherical Image Representation

CVPR 2022oral

The 360deg imaging has recently gained much attention; however, its angular resolution is relatively lower than that of a narrow field-of-view (FOV) perspective image as it is captured using a fisheye lens with the same sensor size. Therefore, it is beneficial to super-resolve a 360deg image. Severa…

Cited by 58PDFScholar
2022

Unbiased Manifold Augmentation for Coarse Class Subdivision

ECCV 2022poster

"Class Subdivision (CCS) is important for many practical applications, where the training set originally annotated for a coarse class (e.g. bird) needs to further support its sub-classes recognition (e.g. swan, crow) with only very few fine-grained labeled samples. From the perspective of causal rep…

2021

CDNet: Centripetal Direction Network for Nuclear Instance Segmentation

ICCV 2021poster

Nuclear instance segmentation is a challenging task due to a large number of touching and overlapping nuclei in pathological images. Existing methods cannot effectively recognize the accurate boundary owing to neglecting the relationship between pixels (e.g., direction information). In this paper, w…

Cited by 60PDFcodeScholar
2021

Dual Transfer Learning for Event-Based End-Task Prediction via Pluggable Event to Image Translation

ICCV 2021poster

Event cameras are novel sensors that perceive the per-pixel intensity changes and output asynchronous event streams with high dynamic range and less motion blur. It has been shown that events alone can be used for end-task learning, e.g., semantic segmentation, based on encoder-decoder-like networks…

Cited by 43PDFcodeScholar
2021

EvDistill: Asynchronous Events To End-Task Learning via Bidirectional Reconstruction-Guided Cross-Modal Knowledge Distillation

CVPR 2021poster

Event cameras sense per-pixel intensity changes and produce asynchronous event streams with high dynamic range and less motion blur, showing advantages over the conventional cameras. A hurdle of training event-based models is the lack of large qualitative labeled data. Prior works learning end-tasks…

Cited by 85PDFcodeScholar
2020

Deceiving Image-to-Image Translation Networks for Autonomous Driving With Adversarial Perturbations

RA-L 2020

Deep neural networks (DNNs) have achieved impressive performance on handling computer vision problems. However, it has been found that DNNs are vulnerable to adversarial examples. For such reason, adversarial perturbations have been recently studied in several respects. However, most previous works

Cited by 29SourceScholar
2020

EventSR: From Asynchronous Events to Image Reconstruction, Restoration, and Super-Resolution via End-to-End Adversarial Learning

CVPR 2020poster

Event cameras sense intensity changes and have many advantages over conventional cameras. To take advantage of event cameras, some methods have been proposed to reconstruct intensity images from event streams. However, the outputs are still in low resolution (LR), noisy, and unrealistic. The low-qua…

Cited by 123PDFcodeScholar
2020

The Compressed Nested Array for Underdetermined DOA Estimation by Fourth-order Difference Coarrays

ICASSP 2020accepted

In this paper, a new sparse array structure, which further improves the degrees of freedom (DOFs) and enhanced the DOA estimation performance, for the fourth-order cumulant based direction of arrival (DOA) estimation is proposed. The new-formed array is hole-free and can achieve a large consecutive…

Cited by 0SourceScholar
2019

Audio-visual sensing from a quadcopter: dataset and baselines for source localization and sound enhancement

IROS 2019poster

We present an audio-visual dataset recorded outdoors from a quadcopter and discuss baseline results for multiple applications. The dataset includes a scenario for source localization and sound enhancement with up to two static sources, and a scenario for source localization and tracking with a movin…

Cited by 30SourceScholar
2019

Event-Based High Dynamic Range Image and Very High Frame Rate Video Generation Using Conditional Generative Adversarial Networks

CVPR 2019poster

Event cameras have a lot of advantages over traditional cameras, such as low latency, high temporal resolution, and high dynamic range. However, since the outputs of event cameras are the sequences of asynchronous events over time rather than actual intensity images, existing algorithms could not be…

Cited by 241PDFScholar