← Search

Bin Chen

102 accepted papers

2026

$A_2$DEPT: Large Language Model–Driven Automated Algorithm Design via Evolutionary Program Trees

ICML 2026poster

Designing heuristics for combinatorial optimization problems (COPs) is a fundamental yet challenging task that traditionally requires extensive domain expertise. Recently, Large Language Model (LLM)-based Automated Heuristic Design (AHD) has shown promise in autonomously generating heuristic compone…

Cited by 0SourceScholar
2026

ASTPKEFormer: Adaptive Spatiotemporal Prior Knowledge Embedding-Induced Transformers for Traffic Data Forecasting

IJCAI 2026

Traffic forecasting is fundamentally challenging due to the complex and dynamic spatiotemporal dependencies inherent in road networks. Although existing prediction models are able to achieve certain results on this task, existing Transformer-based models usually rely on simple embedding strategies a

Cited by 0Scholar
2026

Autoregressive-based Progressive Coding for Ultra-Low Bitrate Image Compression

ICLR 2026poster

Generative models have demonstrated significant results in ultra-low bitrate image compression, owing to their powerful capabilities for content generation and texture completion. Existing works primarily based on diffusion models still face challenges such as limited bitrate adaptability and high c…

Cited by 0SourceScholar
2026

CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception

ICML 2026poster

High-resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods struggle with the trade-off between coverage and efficiency. Visual expert-assisted search is efficient but prone to blind s…

Cited by 0SourceScholar
2026

Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models

ICLR 2026poster

The rapid progress of visual autoregressive (VAR) models has brought new opportunities for text-to-image generation, but also heightened safety concerns. Existing concept erasure techniques, primarily designed for diffusion models, fail to generalize to VARs due to their next-scale token prediction…

Cited by 0SourcecodeScholar
2026

CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation

ICLR 2026poster

Existing works on reasoning segmentation either connect hidden features from a language model directly to a mask decoder or represent positions in text, which limits interpretability and semantic detail. To solve this, we present CoPRS, a Multi-modal Chain-of-Thought (MCoT)–based positional percepti…

Cited by 0SourceScholar
2026

DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles

CVPR 2026

Prompt learning is a dominant paradigm for adapting pre-trained Vision-Language Models (VLMs) to downstream tasks. However, existing methods often rely on a simplistic, layer-centric view, assuming shallow layers capture general features while deep layers handle task-specific knowledge. This assumpt

Cited by 0SourceScholar
2026

FreqSIC: Frequency-aware Stereo Image Compression with Bi-directional Checkerboard Context Model

CVPR 2026

Stereo image compression is essential for a wide range of 3D vision. Recent methods have demonstrated strong capabilities in eliminating inter-view redundancy and enabling compact entropy coding via spatial-domain stereo transformation and advanced autoregressive entropy models. However, these appro

Cited by 0SourceScholar
2026

HIGH QUALITY UNDERWATER IMAGE COMPRESSION WITH ADAPTIVE COLOR CORRECTION

ICASSP 2026oral

With the increasing exploration and exploitation of the underwater world, underwater images have become a critical medium for human interaction with marine environments, driving extensive research into their efficient transmission and storage. However, contemporary underwater image compression algor…

Cited by 0SourcePDFScholar
2026

Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval

CVPR 2026

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos based on text queries that describe only partial events. Existing methods suffer from incomplete global contextual perception, struggling with query ambiguity and local noise induced by spurious responses. To address these i

Cited by 0SourcecodeScholar
2026

Improved Adversarial Diffusion Compression for Real-World Video Super-Resolution

ICLR 2026poster

While many diffusion models have achieved impressive results in real-world video super-resolution (Real-VSR) by generating rich and realistic details, their reliance on multi-step sampling leads to slow inference. One-step networks like SeedVR2, DOVE, and DLoRAL alleviate this through condensing gen…

Cited by 0SourceScholar
2026

Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning

CVPR 2026

Visual in-context learning (VICL) enables visual foundation models to handle multiple tasks by steering them with demonstrative prompts. The choice of such prompts largely influences VICL performance, standing out as a key challenge. Prior work has made substantial progress on prompt retrieval and r

Cited by 0SourcecodeScholar
2026

MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy Model

CVPR 2026

Stereo image compression (SIC) has become increasingly vital with its applications surging in fields such as 3D reconstruction and autonomous navigation. Previous methods leverage cross-attention to model inter-view redundancy and employ autoregressive entropy models to predict probability distribut

Cited by 0SourceScholar
2026

PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment

ICLR 2026poster

Visual In-Context Learning (VICL) aims to complete vision tasks by imitating pixel demonstrations. Recent work Condenser pioneered prompt fusion that combines the advantages of various demonstrations, which shows a promising way to extend VICL. Unfortunately, the patch-wise fusion framework and mode…

Cited by 0SourceScholar
2026

Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation

ICLR 2026poster

Dataset distillation aims to generate compact synthetic datasets that enable models trained on them to achieve performance comparable to those trained on full real datasets, while substantially reducing storage and computational costs. Early bi-level optimization methods (e.g., MTT) have shown promi…

Cited by 0SourcecodeScholar
2026

SDiD:Shared diffusion prior for efficient distributed stereo image compression

ICML 2026poster

Stereo vision is widely utilized in automotive imagery and 3D reconstruction, creating a demand for compressing stereo images. Existing methods for stereo image compression often employ VAE-like architectures based on distortion optimization, leading to subpar perceptual quality at low bitrates. Whi…

Cited by 0SourceScholar
2026

SyncMos: Scalable Motion Synchronisation for Multi-Agent Scene Interaction

CVPR 2026

Text-guided motion generation in 3D scenes has advanced the synthesis of human-scene interactions, contributing to embodied AI, scene understanding, and virtual agent simulation. While recent studies have begun exploring multi-agent scenarios, achieving temporally synchronised interactions among mul

Cited by 0SourceScholar
2026

Towards Efficient Low-rate Image Compression with Frequency-aware Diffusion Prior Refinement

AAAI 2026technical

Recent advancements in diffusion-based generative priors have enabled visually plausible image compression at extremely low bit rates. However, existing approaches suffer from slow sampling processes and suboptimal bit allocation due to fragmented training paradigms. In this work, we propose Acceler

Cited by 0SourcePDFScholar
2026

UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement

CVPR 2026

Image quality assessment (IQA) and image restoration are fundamental problems in low-level vision. Although IQA and restoration are closely connected conceptually, most existing work treats them in isolation. Recent advances in unified multimodal understanding-generation models demonstrate promising

Cited by 0SourcecodeScholar
2025

3D-LMVIC: Learning-based Multi-View Image Compression with 3D Gaussian Geometric Priors

ICML 2025poster

Existing multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more complex disparities encountered in wide-baseline multi-camer…

Cited by 0SourcePDFScholar
2025

Adversarial Diffusion Compression for Real-World Image Super-Resolution

CVPR 2025poster

Real-world image super-resolution (Real-ISR) aims to reconstruct high-resolution images from low-resolution inputs degraded by complex, unknown processes. While many Stable Diffusion (SD)-based Real-ISR methods have achieved remarkable success, their slow, multi-step inference hinders practical depl…

2025

An Exploration with Entropy Constrained 3D Gaussians for 2D Video Compression

ICLR 2025poster

3D Gaussian Splatting (3DGS) has witnessed its rapid development in novel view synthesis, which attains high quality reconstruction and real-time rendering. At the same time, there is still a gap before implicit neural representation (INR) can become a practical compressor due to the lack of stream…

2025

AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing

CVPR 2025poster

Self-Supervised Video Hashing (SSVH) compresses videos into hash codes for efficient indexing and retrieval using unlabeled training videos. Existing approaches rely on random frame sampling to learn video features and treat all frames equally. This results in suboptimal hash codes, as it ignores fr…

2025

Cassic: Towards Content-Adaptive State-Space Models for Learned Image Compression

ICCV 2025poster

Learned image compression (LIC) demonstrates superior rate-distortion (RD) performance compared to traditional methods. Recent method MambaVC attempts to introduce Mamba, a variant of state space models, into this field aim to establish a new paradigm beyond convolutional neural networks and transfo…

Cited by 0SourcePDFScholar
2025

Clients Collaborate: Flexible Differentially Private Federated Learning with Guaranteed Improvement of Utility-Privacy Trade-off

ICML 2025poster

To defend against privacy leakage of user data, differential privacy is widely used in federated learning, but it is not free. The addition of noise randomly disrupts the semantic integrity of the model and this disturbance accumulates with increased communication rounds. In this paper, we introduce…

2025

Continuously Learning Video-level Object Tokens for Robust UAV tracking

ICASSP 2025accepted

Due to the dynamic changes in flight motion and viewpoint, the objects in unmanned aerial vehicle (UAV) tracking scenarios often suffer from drastic appearance variations. Existing UAV trackers often leverage a frame-level matching mechanism, which measures the appearance similarity between the obje…

Cited by 0SourceScholar
2025

DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception

CVPR 2025poster

Dense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have shown promise in open-vocabulary tasks, their direct applicatio…

2025

DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic Refinement

ICLR 2025poster

Reconstructing high-quality images under low bitrates conditions presents a challenge, and previous methods have made this task feasible by leveraging the priors of diffusion models. However, the effective exploration of pre-trained latent diffusion models and semantic information integration in im…

Cited by 0SourcePDFScholar
2025

Efficient Self-Supervised Video Hashing with Selective State Spaces

AAAI 2025technical

Self-supervised video hashing (SSVH) is a practical task in video indexing and retrieval. Although Transformers are predominant in SSVH for their impressive temporal modeling capabilities, they often suffer from computational and memory inefficiencies. Drawing inspiration from Mamba, an advanced sta…

2025

Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning

CVPR 2025poster

Visual In-Context Learning (VICL) enables adaptively solving vision tasks by leveraging pixel demonstrations, mimicking human-like task completion through analogy. Prompt selection is critical in VICL, but current methods assume the existence of a single "ideal" prompt in a pool of candidates, which…

2025

Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

ICCV 2025poster

Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos…

2025

Expert-Enhanced Masked Point Modeling for Point Cloud Self-Supervised Learning

ICRA 2025

Recently, learning-based point cloud analysis has played a crucial role in robotic perception. Masked Point Modeling (MPM), owing to its powerful representational capabilities, has become the mainstream point cloud self-supervised learning method. However, existing MPM-based methods often suffer fro

Cited by 0SourcecodeScholar
2025

GaussianSR: High Fidelity 2D Gaussian Splatting for Arbitrary-Scale Image Super-Resolution

AAAI 2025technical

Implicit neural representations (INRs) have revolutionized arbitrary-scale super-resolution (ASSR) by modeling images as continuous functions. Most existing INR-based ASSR networks first extract features from the given low-resolution image using an encoder, and then render the super-resolved result…

Cited by 3SourcePDFScholar
2025

Going Beyond Feature Similarity: Effective Dataset distillation based on Class-aware Conditional Mutual Information

ICLR 2025poster

Dataset distillation (DD) aims to minimize the time and memory consumption needed for training deep neural networks on large datasets, by creating a smaller synthetic dataset that has similar performance to that of the full real dataset. However, current dataset distillation methods often result in…

2025

Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs

NeurIPS 2025poster

Large Vision-Language Models (LVLMs) are susceptible to hallucinations, where generated responses seem semantically plausible yet exhibit little or no relevance to the input image. Previous studies reveal that this issue primarily stems from LVLMs' over-reliance on language priors while disregarding…

Cited by 0SourceScholar
2025

Hierarchical Features Matter: A Deep Exploration of Progressive Parameterization Method for Dataset Distillation

CVPR 2025poster

Dataset distillation is an emerging dataset reduction method, which condenses large-scale datasets while maintaining task accuracy. Current parameterization methods achieve enhanced performance under extremely high compression ratio by optimizing determined synthetic dataset in informative feature d…

2025

KOEnsAttack: Towards Efficient Data-Free Black-Box Adversarial Attacks via Knowledge-Orthogonalized Substitute Ensembles

ICCV 2025poster

Data-free black-box attacks aim to attack a model without access to either the model parameters or training data. Existing methods use a generator to synthesize training samples and then train a substitute model to imitate the victim model. The adversarial examples (AEs) are finally generated using…

Cited by 0SourcePDFScholar
2025

LNeRV: Learnable Hierarchical Encoding Improve Neural Representation Video Codec

ICASSP 2025accepted

Existing Implicit Neural Representation (INR) video compression techniques have opened up new avenues in the field of video compression. NeRV maps the temporal coordinates to high-resolution images using neural networks, providing a more flexible and efficient encoding method for video data. However…

Cited by 1SourceScholar
2025

Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior

NeurIPS 2025poster

Recent advances in Video Large Language Models (VLLMs) have achieved remarkable video understanding capabilities, yet face critical efficiency bottlenecks due to quadratic computational growth with lengthy visual token sequences of long videos. While existing keyframe sampling methods can improve te…

Cited by 0SourceScholar
2025

MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models

ICASSP 2025accepted

Question-answering (QA) is a natural approach for humans to understand a piece of music audio. However, for machines, accessing a large-scale dataset covering diverse aspects of music is crucial, yet challenging, due to the scarcity of publicly available music data of this type. This paper introduce…

Cited by 0SourceScholar
2025

MST-HA: Multi-Modal Signal Fusion with Bayesian Optimization for Robust Industrial Robot Joint Health Assessment

ICASSP 2025accepted

This paper presents a novel multi-modal deep learning framework for industrial robot joint health assessment and prediction, leveraging non-invasive signal fusion and Bayesian optimization. The proposed method addresses the challenges of comprehensive joint state monitoring in complex industrial env…

Cited by 0SourceScholar
2025

Meta-Conscious Driven Domain-Aware Federated Learning

ICASSP 2025accepted

Cross-domain collaboration can drive comprehensive knowledge innovation and foster synergistic advancements. Federated learning (FL) enables such collaboration while ensuring data security. However, cross-domain FL often faces challenges due to knowledge interference between domains, which can resul…

Cited by 0SourceScholar
2025

MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds

EMNLP 2025

The rapid advancement of large language models has intensified public concerns about the potential misuse. Therefore, it is important to build trustworthy AI-generated text detection systems. Existing methods neglect stylistic modeling and mostly rely on static thresholds, which greatly limits the d

2025

Modeling Uncertainty in Composed Image Retrieval via Probabilistic Embeddings

ACL 2025long

Composed Image Retrieval (CIR) enables users to search for images using multimodal queries that combine text and reference images. While metric learning methods have shown promise, they rely on deterministic point embeddings that fail to capture the inherent uncertainty in the input data, in which u…

2025

OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision

AAAI 2025technical

Open-vocabulary detection aims to detect objects from novel categories beyond the base categories on which the detector is trained. However, existing open-vocabulary detectors trained on base category data tend to assign higher confidence to trained categories and confuse novel categories with the b…

2025

OmniGuard: Hybrid Manipulation Localization via Augmented Versatile Deep Image Watermarking

CVPR 2025poster

With the rapid growth of generative AI and its widespread application in image editing, new risks have emerged regarding the authenticity and integrity of digital content. Existing versatile watermarking approaches suffer from trade-offs between tamper localization precision and visual quality. Cons…

Cited by 4SourcePDFScholar
2025

One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models

ICCV 2025poster

Vision-Language Pre-training (VLP) models have exhibited unprecedented capability in many applications by taking full advantage of the learned multimodal alignment. However, previous studies have shown they are vulnerable to maliciously crafted adversarial samples. Despite recent success, these atta…

2025

PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba Adapter

CVPR 2025poster

Applying pre-trained models to assist point cloud understanding has recently become a mainstream paradigm in 3D perception. However, existing application strategies are straightforward, utilizing only the final output of the pre-trained model for various task heads. It neglects the rich complementar…

2025

Point Cloud Mixture-of-Domain-Experts Model for 3D Self-supervised Learning

IJCAI 2025

Point clouds, as a primary representation of 3D data, can be categorized into scene domain point clouds and object domain point clouds. Point cloud self-supervised learning (SSL) has become a mainstream paradigm for learning 3D representations. However, existing point cloud SSL primarily focuses on

Cited by 0SourcePDFScholar
2025

RobNAS: Robust Neural Architecture Search for Point Cloud Adversarial Defense

ICASSP 2025accepted

As point clouds gain widespread application in fields such as autonomous driving and scene modeling, an increasing number of point cloud learning networks have emerged. As a result, research on 3D adversarial attacks and defenses has rapidly advanced. To the best of our knowledge, existing 3D defens…

Cited by 0SourceScholar
2025

Stealthy Shield Defense: A Conditional Mutual Information-Based Approach against Black-Box Model Inversion Attacks

ICLR 2025poster

Model inversion attacks (MIAs) aim to reconstruct the private training data by accessing the public model, raising concerns about privacy leakage. Black-box MIAs, where attackers can only query the model and obtain outputs, are closer to real-world scenarios. The latest black-box attacks have outper…

2025

Text-guided Multimodal Fusion for the Multimodal Emotion and Intent Joint Understanding

ICASSP 2025accepted

Emotion and Intent Joint Understanding in Multi-modal Conversation is a challenging task in the field of affective computing, aiming to decode the semantic information manifested in the multimodal conversational while simultaneously inferring the emotions and intents of the utterance. To address thi…

Cited by 0SourceScholar
2025

Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors

EMNLP 2025

The misuse of large language models (LLMs), such as academic plagiarism, has driven the development of detectors to identify LLM-generated texts. To bypass these detectors, paraphrase attacks have emerged to purposely rewrite these texts to evade detection. Despite the success, existing methods requ

2024

BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional Bootstrapping

NeurIPS 2024poster

Adaptation of pretrained vision-language models such as CLIP to various downstream tasks have raised great interest in recent researches. Previous works have proposed a variety of test-time adaptation (TTA) methods to achieve strong generalization without any knowledge of the target domain. Howev…

2024

COSMIC: Compress Satellite Image Efficiently via Diffusion Compensation

NeurIPS 2024poster

With the rapidly increasing number of satellites in space and their enhanced capabilities, the amount of earth observation images collected by satellites is exceeding the transmission limits of satellite-to-ground links. Although existing learned image compression solutions achieve remarkable perfor…

Cited by 1SourcePDFScholar
2024

Comparing a BERT Classifier and a GPT classifier for Detecting Connective Language Across Multiple Social Media

EMNLP 2024main

This study presents an approach for detecting connective language—defined as language that facilitates engagement, understanding, and conversation—from social media discussions. We developed and evaluated two types of classifiers: BERT and GPT-3.5 turbo. Our results demonstrate that the BERT classif…

Cited by 1SourcePDFScholar
2024

Decoding at the Speed of Thought: Harnessing Parallel Decoding of Lexical Units for LLMs

COLING 2024main

Large language models have demonstrated exceptional capability in natural language understanding and generation. However, their generation speed is limited by the inherently sequential nature of their decoding process, posing challenges for real-time applications. This paper introduces Lexical Unit…

2024

GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval

AAAI 2024technical

Given a text query, partially relevant video retrieval (PRVR) seeks to find untrimmed videos containing pertinent moments in a database. For PRVR, clip modeling is essential to capture the partial relationship between texts and videos. Current PRVR methods adopt scanning-based clip construction to a…

2024

GladCoder: Stylized QR Code Generation with Grayscale-Aware Denoising Process

IJCAI 2024poster

Traditional QR codes consist of a grid of black-and-white square modules, which lack aesthetic appeal and meaning for human perception. This has motivated recent research to beautify the visual appearance of QR codes. However, there exists a trade-off between the visual quality and scanning-robustne…

Cited by 0SourcePDFScholar
2024

Interpreting Temporal Knowledge Graph Reasoning (Student Abstract)

AAAI 2024technical

Temporal knowledge graph reasoning is an essential task that holds immense value in diverse real-world applications. Existing studies mainly focus on leveraging structural and sequential dependencies, excelling in tasks like entity and link prediction. However, they confront a notable interpretabili…

Cited by 2SourcePDFScholar
2024

LCM: Locally Constrained Compact Point Cloud Model for Masked Point Modeling

NeurIPS 2024poster

The pre-trained point cloud model based on Masked Point Modeling (MPM) has exhibited substantial improvements across various tasks. However, these models heavily rely on the Transformer, leading to quadratic complexity and limited decoder, hindering their practice application. To address this limita…

2024

MEAT: Median-Ensemble Adversarial Training for Improving Robustness and Generalization

ICASSP 2024accepted

Self-ensemble adversarial training methods improve model robustness by ensembling models at different training epochs, such as model weight averaging (WA). However, previous research has shown that self-ensemble defense methods in adversarial training (AT) still suffer from robust overfitting, which…

Cited by 0SourceScholar
2024

Mitigating Linguistic Artifacts in Emotion Recognition for Conversations from TV Scripts to Daily Conversations

COLING 2024main

Emotion Recognition in Conversations (ERC) is a well-studied task with numerous potential real-world applications. However, existing ERC models trained on the MELD dataset derived from TV series, struggle when applied to daily conversation datasets. A closer examination of the datasets unveils the p…

Cited by 0SourcePDFScholar
2024

MuseChat: A Conversational Music Recommendation System for Videos

CVPR 2024highlight

Music recommendation for videos attracts growing interest in multi-modal research. However existing systems focus primarily on content compatibility often ignoring the users' preferences. Their inability to interact with users for further refinements or to provide explanations leads to a less satisf…

2024

Natural Evolution-based Dual-Level Aggregation for Temporal Knowledge Graph Reasoning

EMNLP 2024finding

Temporal knowledge graph (TKG) reasoning aims to predict missing facts based on a given history. Most of the existing methods unifiedly model the evolution process of different events and ignore their inherent asynchronous characteristics, resulting in suboptimal performance. To tackle this challeng…

Cited by 0SourcePDFScholar
2024

Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-Experts

NeurIPS 2024poster

Designing single-task image restoration models for specific degradation has seen great success in recent years. To achieve generalized image restoration, all-in-one methods have recently been proposed and shown potential for multiple restoration tasks using one single model. Despite the promising re…

2024

ReFIR: Grounding Large Restoration Models with Retrieval Augmentation

NeurIPS 2024poster

Recent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from the hallucination dilemma, i.e., producing incorrect contents…

2024

Shallow Diffusion for Fast Speech Enhancement (Student Abstract)

AAAI 2024technical

Recently, the field of Speech Enhancement has witnessed the success of diffusion-based generative models. However, these diffusion-based methods used to take multiple iterations to generate high-quality samples, leading to high computational costs and inefficiency. In this paper, we propose SDFEN (S…

Cited by 0SourcePDFScholar
2024

Towards Compact 3D Representations via Point Feature Enhancement Masked Autoencoders

AAAI 2024technical

Learning 3D representation plays a critical role in masked autoencoder (MAE) based pre-training methods for point cloud, including single-modal and cross-modal based MAE. Specifically, although cross-modal MAE methods learn strong 3D representations via the auxiliary of other modal knowledge, they…

2024

Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization

ICLR 2024poster

Recently, the remarkable advance of the Large Language Model (LLM) has inspired researchers to transfer its extraordinary reasoning capability to both vision and language data. However, the prevailing approaches primarily regard the visual input as a prompt and focus exclusively on optimizing the te…

2024

Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding

AAAI 2024technical

In recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. However, when extended to point cloud data, existing works mainly focus on building ta…

2023

An Adaptive Model Ensemble Adversarial Attack for Boosting Adversarial Transferability

ICCV 2023poster

While the transferability property of adversarial examples allows the adversary to perform black-box attacks i.e., the attacker has no knowledge about the target model), the transfer-based adversarial attacks have gained great attention. Previous works mostly study gradient variation or image transf…

Cited by 47PDFcodeScholar
2023

An Exploratory Study on Model Compression for Text-to-SQL

ACL 2023findings

Text-to-SQL translates user queries into SQL statements that can retrieve relevant answers from relational databases. Recent approaches to Text-to-SQL rely on pre-trained language models that are computationally expensive and technically challenging to deploy in real-world applications that require…

2023

Battle of the Large Language Models: Dolly vs LLaMA vs Vicuna vs Guanaco vs Bard vs ChatGPT - A Text-to-SQL Parsing Comparison

EMNLP 2023long findings

The success of ChatGPT has ignited an AI race, with researchers striving to develop new large language models (LLMs) that can match or surpass the language understanding and generation abilities of commercial ones. In recent times, a number of models have emerged, claiming performance near that of…

Cited by 0SourceScholar
2023

Contrastive Masked Autoencoders for Self-Supervised Video Hashing

AAAI 2023technical

Self-Supervised Video Hashing (SSVH) models learn to generate short binary representations for videos without ground-truth supervision, facilitating large-scale video retrieval efficiency and attracting increasing research attention. The success of SSVH lies in the understanding of video content and…

2023

Difficulty-Aware Data Augmentor for Scene Text Recognition

ICASSP 2023accepted

Deep neural network (DNN) based scene text recognition (STR) methods usually require a large amount of annotated data for training, which is time-consuming and cost-expensive in practice. To address this issue, many data augmentation methods have been developed to train recognizers by improving the…

Cited by 0SourceScholar
2023

FSR: A General Frequency-Oriented Framework to Accelerate Image Super-resolution Networks

AAAI 2023technical

Deep neural networks (DNNs) have witnessed remarkable achievement in image super-resolution (SR), and plenty of DNN-based SR models with elaborated network designs have recently been proposed. However, existing methods usually require substantial computations by operating in spatial domain. To addre…

2023

GIFD: A Generative Gradient Inversion Method with Feature Domain Optimization

ICCV 2023poster

Federated Learning (FL) has recently emerged as a promising distributed machine learning framework to preserve clients' privacy, by allowing multiple clients to upload the gradients calculated from their local data to a central server. Recent studies find that the exchanged gradients also take the r…

Cited by 40PDFcodeScholar
2023

GlowGAN: Unsupervised Learning of HDR Images from LDR Images in the Wild

ICCV 2023poster

Most in-the-wild images are stored in Low Dynamic Range (LDR) form, serving as a partial observation of the High Dynamic Range (HDR) visual world. Despite limited dynamic range, these LDR images are often captured with different exposures, implicitly containing information about the underlying HDR i…

Cited by 13PDFScholar
2023

Instance-aware Dynamic Prompt Tuning for Pre-trained Point Cloud Models

ICCV 2023poster

Pre-trained point cloud models have found extensive applications in 3D understanding tasks like object classification and part segmentation. However, the prevailing strategy of full fine-tuning in downstream tasks leads to large per-task storage overhead for model parameters, which limits the effici…

Cited by 48PDFcodeScholar
2023

Learned Distributed Image Compression with Multi-Scale Patch Matching in Feature Domain

AAAI 2023technical

Beyond achieving higher compression efficiency over classical image compression codecs, deep image compression is expected to be improved with additional side information, e.g., another image from a different perspective of the same scene. To better utilize the side information under the distributed…

Cited by 13SourcePDFScholar
2023

Learning Transferable Spatiotemporal Representations From Natural Script Knowledge

CVPR 2023poster

Pre-training on large-scale video data has become a common recipe for learning transferable spatiotemporal representations in recent years. Despite some progress, existing methods are mostly limited to highly curated datasets (e.g., K400) and exhibit unsatisfactory out-of-the-box representations. We…

2023

Sph2Pob: Boosting Object Detection on Spherical Images with Planar Oriented Boxes Methods

IJCAI 2023poster

Object detection on panoramic/spherical images has been developed rapidly in the past few years, where IoU-calculator is a fundamental part of various detector components, i.e. Label Assignment, Loss and NMS. Due to the low efficiency and non-differentiability of spherical Unbiased IoU, spherical ap…

2023

Subspace Modeling Enabled High-Sensitivity X-Ray Chemical Imaging

ICASSP 2023accepted

Resolving morphological chemical phase transformations at the nanoscale is of vital importance to many scientific and industrial applications across various disciplines. The TXM-XANES imaging technique, by combining full-field transmission X-ray microscopy (TXM) and X-ray absorption near edge struct…

Cited by 5SourceScholar
2022

APD: Learning Diverse Behaviors for Reinforcement Learning Through Unsupervised Active Pre-Training

RA-L 2022

Unsupervised pre-training in reinforcement learning enables the agent to gain prior environmental knowledge, which is then fine-tuned in the supervised stage to quickly adapt to various downstream tasks. In the absence of task-related rewards, pre-training aims to acquire policies (i.e., behaviors)

Cited by 5SourceScholar
2022

Contrastive Quantization with Code Memory for Unsupervised Image Retrieval

AAAI 2022technical

The high efficiency in computation and storage makes hashing (including binary hashing and quantization) a common strategy in large-scale retrieval systems. To alleviate the reliance on expensive annotations, unsupervised deep hashing becomes an important research problem. This paper provides a nove…

2022

Is Discourse Role Important for Emotion Recognition in Conversation?

AAAI 2022technical

A conversation is a sequence of utterances, where each utterance plays a specific discourse role while expressing a particular emotion. This paper proposes a novel method to exploit latent discourse role information of an utterance to determine the emotion it conveys in a conversation. Specifically,…

Cited by 30SourcePDFScholar
2022

Learning to Deblur Using Light Field Generated and Real Defocus Images

CVPR 2022oral

Defocus deblurring is a challenging task due to the spatially varying nature of defocus blur. While deep learning approach shows great promise in solving image restoration problems, defocus deblurring demands accurate training data that consists of all-in-focus and defocus image pairs, which is diff…

Cited by 84PDFcodeScholar
2022

Unbiased IoU for Spherical Image Object Detection

AAAI 2022technical

As one of the fundamental components of object detection, intersection-over-union (IoU) calculations between two bounding boxes play an important role in samples selection, NMS operation and evaluation of object detection algorithms. This procedure is well-defined and solved for planar images, while…

Cited by 12SourcePDFScholar
2021

Efficient Face Manipulation Via Deep Feature Disentanglement And Reintegration Net

ICASSP 2021accepted

Deep neural networks (DNNs) have been widely used in facial manipulation. Existing methods focus on training deeper networks in indirect supervision ways (e.g., feature constraint), or in unsupervised ways (e.g., cycle-consistency loss) due to the lack of ground-truth face images for manipulated out…

Cited by 1SourceScholar
2021

HOCA: Higher-Order Channel Attention for Single Image Super-Resolution

ICASSP 2021accepted

Convolutional neural networks (CNNs) have obtained great success in single image super-resolution (SR). More recent works (e.g., RCAN and SAN) have obtained remarkable performance with channel attention based on first- or second-order statistics of features. However, these methods neglect the rich f…

Cited by 5SourceScholar
2021

Weakly Supervised Deep Hyperspherical Quantization for Image Retrieval

AAAI 2021technical

Deep quantization methods have shown high efficiency on large-scale image retrieval. However, current models heavily rely on ground-truth information, hindering the application of quantization in label-hungry scenarios. A more realistic demand is to learn from inexhaustible uploaded images that are…

2020

Targeted Attack for Deep Hashing based Retrieval

ECCV 2020poster

The deep hashing based retrieval method is widely adopted in large-scale image and video retrieval. However, there is little investigation on its security. In this paper, we propose a novel method, dubbed deep hashing targeted attack (DHTA), to study the targeted attack on such retrieval. Specifical…