← Search

Jun Chen

93 accepted papers

2026

G-Merging: Graph Models Merging for Parameter-Efficient Multi-Task Knowledge Consolidation

ICLR 2026poster

The pretrain-finetuning paradigm has achieved notable success in graph learning. Moreover, merging models fine-tuned on different tasks to enable a parameter-efficient model with multi-task capabilities is gaining increasing attention for its practicality. However, existing model merging methods, su…

Cited by 0SourcecodeScholar
2026

HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation

IJCAI 2026

Large language models (LLMs) are increasingly used for hardware and firmware code generation, but existing studies primarily evaluate functional correctness while largely overlooking security. However, LLM-generated code that appears functionally sound may embed security flaws which could induce cat

Cited by 0Scholar
2026

KRAMABENCH: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes

ICLR 2026poster

Discovering insights from a real-world data lake potentially containing unclean, semi-structured, and unstructured data requires a variety of data processing tasks, ranging from extraction and cleaning to integration, analysis, and modeling. This process often also demands domain knowledge and proje…

Cited by 0SourcecodeScholar
2026

Learn to Merge: Meta-Learning for Adaptive Multi-Task Model Merging

ICML 2026poster

Model merging in the pretrain-finetune paradigm has proven effective by combining multiple finetuned models into one with multi-task capabilities. However, existing methods rely on fix or manually tuned merging coefficients, making the unified model sensitive to the initial merging strategy and subo…

Cited by 0SourceScholar
2026

Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition

CVPR 2026

High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory and copyright constraints. This scarcity hampers model development--ironically, in settings where generative models are most needed to compensate for

Cited by 0SourceScholar
2026

StormInsight: Hierarchical Environmental Forcing and Vertical Coupling for Weather System Evolution

ICML 2026poster

Nowcasting forms the first line of defense against rapidly evolving weather hazards, where even minutes of delay can lead to severe societal impacts. However, existing systems predominantly extrapolate 2D radar reflectivity, which struggles under rapid intensification regimes. We introduce \N, a mul…

Cited by 0SourceScholar
2026

VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice

CVPR 2026

Chain-of-thought (CoT) reasoning has emerged as a powerful tool for multimodal large language models on video understanding tasks. However, its necessity and advantages over direct answering remain underexplored. In this paper, we first demonstrate that for RL-trained video models, direct answering

Cited by 0SourceScholar
2026

WHU-MARS: A Multispectral Aerial-Ground Benchmark Towards Any-Scenario Person Re-Identification

CVPR 2026

Recent person re-identification (ReID) leverages heterogeneous sensing with multiple modalities and viewpoints to improve robustness across diverse conditions. However, most approaches target predefined scenario pairs (e.g., visible-infrared or aerial-ground) and train separate task-specific models.

Cited by 0SourcecodeScholar
2026

Zero-Reference Joint Low-Light Enhancement and Deblurring via Visual Autoregressive Modeling with VLM-Derived Modulation

AAAI 2026technical

Real-world dark images commonly exhibit not only low visibility and contrast but also complex noise and blur, posing significant restoration challenges. Existing methods often rely on paired data or fail to model dynamic illumination and blur characteristics, leading to poor generalization. To tackl

Cited by 0SourcePDFScholar
2026

dTRPO : Trajectory Reduction in Policy Optimization of Diffusion Large Language Models

ICML 2026poster

Diffusion Large Language Models (dLLMs) introduce a new paradigm for language generation and thus induce new challenges in aligning dLLMs for human preference. In this work, aim to optimize the dLLM generation process by developing a theoretical formulation and an efficient and effective quantificat…

Cited by 0SourceScholar
2025

4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding

ICCV 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated impressive 2D image/video understanding capabilities.However, there are no publicly standardized benchmarks to assess the abilities of MLLMs in understanding the 4D objects.In this paper, we introduce 4D-Bench, the first benchmark to evaluat…

2025

Balancing Privacy and Performance: A Many-in-One Approach for Image Anonymization

AAAI 2025technical

The effective utilization of data through Deep Neural Networks (DNNs) has profoundly influenced various aspects of society. The growing demand for high-quality, particularly personalized, data has spurred research efforts to prevent data leakage and protect privacy in recent years. Early privacy-pre…

Cited by 0SourcePDFScholar
2025

Diffusion-Based Imaginative Coordination for Bimanual Manipulation

ICCV 2025poster

Bimanual manipulation is crucial in robotics, enabling complex tasks in industrial automation and household services. However, it poses significant challenges due to the high-dimensional action space and intricate coordination requirements. While video prediction has been recently studied for repres…

2025

Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents

CVPR 2025poster

Large multimodal models (LMMs) have achieved impressive progress in vision-language understanding, yet they face limitations in real-world applications requiring complex reasoning over a large number of images. Existing benchmarks for multi-image question-answering are limited in scope, each questio…

2025

Error Analysis Affected by Heavy-Tailed Gradients for Non-Convex Pairwise Stochastic Gradient Descent

AAAI 2025technical

In recent years, there have been a growing number of works studying the generalization properties of stochastic gradient descent (SGD) from the perspective of algorithmic stability. However, few of them devote to simultaneously studying the generalization and optimization for the non-convex setting,…

Cited by 0SourcePDFScholar
2025

GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models

ICLR 2025poster

In this paper, we introduce GoodDrag, a novel approach to improve the stability and image quality of drag editing. Unlike existing methods that struggle with accumulated perturbations and often result in distortions, GoodDrag introduces an AlDD framework that alternates between drag and denoising op…

2025

How does Labeling Error Impact Contrastive Learning? A Perspective from Data Dimensionality Reduction

ICML 2025poster

In recent years, contrastive learning has achieved state-of-the-art performance in the territory of self-supervised representation learning. Many previous works have attempted to provide the theoretical understanding underlying the success of contrastive learning. Almost all of them rely on a defau…

Cited by 0SourcePDFScholar
2025

Improving Retrieval Augmented Language Model with Self-Reasoning

AAAI 2025technical

The Retrieval-Augmented Language Model (RALM) has demonstrated remarkable performance on knowledge-intensive tasks by integrating external knowledge during inference, which mitigates the factual hallucinations inherited in large language models (LLMs). Despite these advancements, challenges persist…

Cited by 7SourcePDFScholar
2025

LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement

ACL 2025long

Recent advancements in language models (LMs) have demonstrated strong capabilities in semantic understanding and contextual modeling, which have flourished in generative speech enhancement (SE). However, many LM-based SE approaches primarily focus on semantic information, often neglecting the critic…

2025

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

ICML 2025poster

Multimodal Large Language Models (MLLMs) have shown promising progress in understanding and analyzing video content. However, processing long videos remains a significant challenge constrained by LLM's context size. To address this limitation, we propose \textbf{LongVU}, a spatiotemporal adaptive co…

2025

Manifold Constraint Reduces Exposure Bias in Accelerated Diffusion Sampling

ICLR 2025poster

Diffusion models have demonstrated significant potential for generating high-quality images, audio, and videos. However, their iterative inference process entails substantial computational costs, limiting practical applications. Recently, researchers have introduced accelerated sampling methods that…

Cited by 0SourcePDFScholar
2025

Unbiased Prototype Consistency Learning for Multi-Modal and Multi-Task Object Re-Identification

NeurIPS 2025spotlight

In object re-identification (ReID) task, both cross-modal and multi-modal retrieval methods have achieved notable progress. However, existing approaches are designed for specific modality and category (person or vehicle) retrieval task, lacking generalizability to others. Acquiring multiple task-spe…

Cited by 0SourcecodeScholar
2025

Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding

NeurIPS 2025spotlight

Understanding and reasoning over long videos pose significant challenges for large video language models (LVLMs) due to the difficulty in processing intensive video tokens beyond context window and retaining long-term sequential information. Retrieval-Augmented Generation (RAG) has demonstrated effe…

Cited by 0SourcecodeScholar
2025

WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation

ICCV 2025poster

Knowledge discovery and collection are intelligence-intensive tasks that traditionally require significant human effort to ensure high-quality outputs. Recent research has explored multi-agent frameworks for automating Wikipedia-style article generation by retrieving and synthesizing information fro…

2024

A Multimodal, Multi-Task Adapting Framework for Video Action Recognition

AAAI 2024technical

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance a…

Cited by 17SourcePDFScholar
2024

Cross-Scale Domain Adaptation with Comprehensive Information for Pansharpening

IJCAI 2024poster

Deep learning-based pansharpening methods typically use simulated data at the reduced-resolution scale for training. It limits their performance when generalizing the trained model to the full-resolution scale due to incomprehensive information utilization of panchromatic (PAN) images at the full-re…

2024

Decentralized Riemannian Conjugate Gradient Method on the Stiefel Manifold

ICLR 2024poster

The conjugate gradient method is a crucial first-order optimization method that generally converges faster than the steepest descent method, and its computational cost is much lower than that of second-order methods. However, while various types of conjugate gradient methods have been studied in Euc…

Cited by 11SourcePDFScholar
2024

ECMamba: Consolidating Selective State Space Model with Retinex Guidance for Efficient Multiple Exposure Correction

NeurIPS 2024poster

Exposure Correction (EC) aims to recover proper exposure conditions for images captured under over-exposure or under-exposure scenarios. While existing deep learning models have shown promising results, few have fully embedded Retinex theory into their architecture, highlighting a gap in current met…

2024

Explore 3D Dance Generation via Reward Model from Automatically-Ranked Demonstrations

AAAI 2024technical

This paper presents an Exploratory 3D Dance generation framework, E3D2, designed to address the exploration capability deficiency in existing music-conditioned 3D dance generation models. Current models often generate monotonous and simplistic dance sequences that misalign with human preferences bec…

Cited by 4SourcePDFScholar
2024

Generating Stereophonic Music with Single-Stage Language Models

ICASSP 2024accepted

The recent success of audio language models (LMs) has revolutionized the field of neural music generation. Among all audio LM approaches, MusicGen has demonstrated the success of a single-stage LMs based music generation framework, without needing to train multiple LMs. Despite its promising perform…

Cited by 0SourceScholar
2024

How Does Black-Box Impact the Learning Guarantee of Stochastic Compositional Optimization?

NeurIPS 2024poster

Stochastic compositional optimization (SCO) problem constitutes a class of optimization problems characterized by the objective function with a compositional form, including the tasks with known derivatives, such as AUC maximization, and the derivative-free tasks exemplified by black-box vertical fe…

Cited by 0SourcePDFScholar
2024

Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time

ECCV 2024poster

"Leveraging Large Language Models’ remarkable proficiency in text-based tasks, recent works on Multi-modal LLMs (MLLMs) extend them to other modalities like vision and audio. However, the progress in these directions has been mostly focused on tasks that only require a coarse-grained understanding o…

2024

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

ICLR 2024poster

The recent GPT-4 has demonstrated extraordinary multi-modal abilities, such as directly generating websites from handwritten text and identifying humorous elements within images. These features are rarely observed in previous vision-language models. However, the technical details behind GPT-4 contin…

2024

Multi-View Midivae: Fusing Track- and Bar-View Representations for Long Multi-Track Symbolic Music Generation

ICASSP 2024accepted

Variational Autoencoders (VAEs) constitute a crucial component of neural symbolic music generation, among which some works have yielded outstanding results and attracted considerable attention. Nevertheless, previous VAEs still encounter issues with overly long feature sequences and generated result…

Cited by 0SourceScholar
2024

SCNet: Sparse Compression Network for Music Source Separation

ICASSP 2024accepted

Deep learning-based methods have made significant achievements in music source separation. However, obtaining good results while maintaining a low model complexity remains challenging in super wide-band music source separation. Previous works either overlook the differences in subbands or inadequate…

Cited by 0SourceScholar
2024

Shallow-Deep Collaborative Learning for Unsupervised Visible-Infrared Person Re-Identification

CVPR 2024poster

Unsupervised visible-infrared person re-identification (US-VI-ReID) centers on learning a cross-modality retrieval model without labels reducing the reliance on expensive cross-modality manual annotation. Previous US-VI-ReID works gravitate toward learning cross-modality information with the deep fe…

2024

SimCalib: Graph Neural Network Calibration Based on Similarity between Nodes

AAAI 2024technical

Graph neural networks (GNNs) have exhibited impressive performance in modeling graph data as exemplified in various applications. Recently, the GNN calibration problem has attracted increasing attention, especially in cost-sensitive scenarios. Previous work has gained empirical insights on the issue…

Cited by 6SourcePDFScholar
2024

Structured Optimal Brain Pruning for Large Language Models

EMNLP 2024main

The massive parameters and computational demands hinder the widespread application of Large Language Models (LLMs). Network pruning provides a practical solution to this problem. However, existing pruning works for LLMs mainly focus on unstructured pruning or necessitate post-pruning fine-tuning. Th…

Cited by 1SourcePDFScholar
2024

SumCSE: Summary as a transformation for Contrastive Learning

NAACL 2024findings

Sentence embedding models are typically trained using contrastive learning (CL), either using human annotations directly or by repurposing other annotated datasets. In this work, we explore the recently introduced paradigm of generating CL data using generative language models (LM). In CL for comput…

2024

UNICORN: A Unified Causal Video-Oriented Language-Modeling Framework for Temporal Video-Language Tasks

EMNLP 2024main

The great success of large language models has encouraged the development of large multimodal models, with a focus on image-language interaction. Despite promising results in various image-language downstream tasks, it is still challenging and unclear how to extend the capabilities of these models t…

2023

A Synthetic Corpus Generation Method for Neural Vocoder Training

ICASSP 2023accepted

Nowadays, neural vocoders are preferred for their ability to synthesize high-fidelity audio. However, training a neural vocoder requires a massive corpus of high-quality real audio, and the audio recording process is often labor-intensive. In this work, we propose a synthetic corpus generation metho…

Cited by 0SourceScholar
2023

Av-Sepformer: Cross-Attention Sepformer for Audio-Visual Target Speaker Extraction

ICASSP 2023accepted

Visual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross- and self-attention to fuse and model features from audio an…

Cited by 0SourceScholar
2023

Contrastive Semi-Supervised Learning for Underwater Image Restoration via Reliable Bank

CVPR 2023poster

Despite the remarkable achievement of recent underwater image restoration techniques, the lack of labeled data has become a major hurdle for further progress. In this work, we propose a mean-teacher based Semi-supervised Underwater Image Restoration (Semi-UIR) framework to incorporate the unlabeled…

2023

Development of a Four-Wheel Steering Scale Vehicle for Research and Education on Autonomous Vehicle Motion Control

RA-L 2023

Autonomous vehicle motion control development requires testing and evaluation at all stages of the process. The development phase involving the instrumentation and operation of a full-size vehicle can be especially costly. Scale vehicles have been developed in the literature to serve as a cost-effec

Cited by 18SourceScholar
2023

Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation Only

ICCV 2023poster

Semantic segmentation is a crucial task in computer vision that involves segmenting images into semantically meaningful regions at the pixel level. However, existing approaches often rely on expensive human annotations as supervision for model training, limiting their scalability to large, unlabeled…

Cited by 33PDFcodeScholar
2023

Fine-Grained Theoretical Analysis of Federated Zeroth-Order Optimization

NeurIPS 2023poster

Federated zeroth-order optimization (FedZO) algorithm enjoys the advantages of both zeroth-order optimization and federated learning, and has shown exceptional performance on black-box attack and softmax regression tasks. However, there is no generalization analysis for FedZO, and its analysis on co…

Cited by 10SourcePDFScholar
2023

Gesper: A Unified Framework for General Speech Restoration

ICASSP 2023accepted

This paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by spee…

Cited by 0SourceScholar
2023

Inter-Subnet: Speech Enhancement with Subband Interaction

ICASSP 2023accepted

Subband-based approaches process subbands in parallel through the model with shared parameters to learn the commonality of local spectrums for noise reduction. In this way, they have achieved remarkable results with fewer parameters. However, in some complex environments, the lack of global spectral…

Cited by 0SourceScholar
2023

Learning Global-aware Kernel for Image Harmonization

ICCV 2023poster

Image harmonization aims to solve the visual inconsistency problem in composited images by adaptively adjusting the foreground pixels with the background as references. Existing methods employ local color transformation or region matching between foreground and background, which neglects powerful pr…

Cited by 9PDFScholar
2023

MammalNet: A Large-Scale Video Benchmark for Mammal Recognition and Behavior Understanding

CVPR 2023poster

Monitoring animal behavior can facilitate conservation efforts by providing key insights into wildlife health, population status, and ecosystem function. Automatic recognition of animals and their behaviors is critical for capitalizing on the large unlabeled datasets generated by modern video device…

2023

On the Stability and Generalization of Triplet Learning

AAAI 2023technical

Triplet learning, i.e. learning from triplet data, has attracted much attention in computer vision tasks with an extremely large number of categories, e.g., face recognition and person re-identification. Albeit with rapid progress in designing and applying triplet learning algorithms, there is a lac…

Cited by 5SourcePDFScholar
2023

On the choice of Perception Loss Function for Learned Video Compression

NeurIPS 2023poster

We study causal, low-latency, sequential video compression when the output is subjected to both a mean squared-error (MSE) distortion loss as well as a perception loss to target realism. Motivated by prior approaches, we consider two different perception loss functions (PLFs). The first, PLF-JD, co…

Cited by 15SourcePDFScholar
2023

SUBP: Soft Uniform Block Pruning for 1$\times$N Sparse CNNs Multithreading Acceleration

NeurIPS 2023poster

The study of sparsity in Convolutional Neural Networks (CNNs) has become widespread to compress and accelerate models in environments with limited resources. By constraining N consecutive weights along the output channel to be group-wise non-zero, the recent network with 1$\times$N sparsity has rece…

2023

Speech Enhancement with Intelligent Neural Homomorphic Synthesis

ICASSP 2023accepted

Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement. Specifically, we use homomorphic signal processing and cepstral ana…

Cited by 0SourceScholar
2023

Stability-Based Generalization Analysis for Mixtures of Pointwise and Pairwise Learning

AAAI 2023technical

Recently, some mixture algorithms of pointwise and pairwise learning (PPL) have been formulated by employing the hybrid error metric of “pointwise loss + pairwise loss” and have shown empirical effectiveness on feature selection, ranking and recommendation tasks. However, to the best of our knowledg…

Cited by 3SourcePDFScholar
2023

TEA-PSE 3.0: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System For ICASSP 2023 Dns-Challenge

ICASSP 2023accepted

This paper introduces the Unbeatable Team’s submission to the ICASSP 2023 Deep Noise Suppression (DNS) Challenge. We expand our previous work, TEA-PSE, to its upgraded version – TEA-PSE 3.0. Specifically, TEA-PSE 3.0 incorporates a residual LSTM after squeezed temporal convolution network (S-TCN) to…

Cited by 0SourceScholar
2023

TFCnet: Time-Frequency Domain Corrector for Speech Separation

ICASSP 2023accepted

Deep learning-based methods have made significant achievements in speech separation. Especially the time-domain separation methods have achieved the best performance in recent years. However, time-domain methods are unstable for waveform transformation, which is prone to amplitude and phase errors.…

Cited by 0SourceScholar
2023

Top-K Visual Tokens Transformer: Selecting Tokens for Visible-Infrared Person Re-Identification

ICASSP 2023accepted

Visible modality and infrared modality person re-identification (VI-ReID) is an extremely important and challenging task. Existing works mainly focus on reducing the modality gap with Convolutional Neural Networks (CNN). However, the features extracted by CNN may contain useless identity-irrelevant…

Cited by 0SourceScholar
2023

Towards Grand Unified Representation Learning for Unsupervised Visible-Infrared Person Re-Identification

ICCV 2023poster

Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) is an extremely important and challenging task, which can alleviate the issue of expensive cross-modality annotations. Existing works focus on handling the cross-modality discrepancy under unsupervised conditions. However,…

Cited by 52PDFcodeScholar
2023

Unified Data-Free Compression: Pruning and Quantization without Fine-Tuning

ICCV 2023poster

Structured pruning and quantization are promising approaches for reducing the inference time and memory footprint of neural networks. However, most existing methods require the original training dataset to fine-tune the model. This not only brings heavy resource consumption but also is not possible…

Cited by 21PDFScholar
2022

Few-Shot Class-Incremental Learning via Entropy-Regularized Data-Free Replay

ECCV 2022poster

"Few-shot class-incremental learning (FSCIL) has been proposed aiming to enable a deep learning system to incrementally learn new classes with limited data. Recently, a pioneer claims that the commonly used replay-based method in class-incremental learning (CIL) is ineffective and thus not preferred…

2022

FullSubNet+: Channel Attention Fullsubnet with Complex Spectrograms for Speech Enhancement

ICASSP 2022accepted

Previously proposed FullSubNet has achieved outstanding performance in Deep Noise Suppression (DNS) Challenge and attracted much attention. However, it still encounters issues such as input-output mismatch and coarse processing for frequency bands. In this paper, we propose an extended single-channe…

Cited by 0SourceScholar
2022

LOSSY COMPRESSION WITH DISTRIBUTION SHIFT AS ENTROPY CONSTRAINED OPTIMAL TRANSPORT

ICLR 2022poster

We study an extension of lossy compression where the reconstruction distribution is different from the source distribution in order to account for distributional shift due to processing. We formulate this as a generalization of optimal transport with an entropy bottleneck to account for the rate con…

Cited by 15SourcePDFScholar
2022

Learning to Train a Point Cloud Reconstruction Network without Matching

ECCV 2022poster

"Reconstruction networks for well-ordered data such as 2D images and 1D continuous signals are easy to optimize through element-wised squared errors, while permutation-arbitrary point clouds cannot be constrained directly because their points permutations are not fixed. Though existing works design…

2022

REMOTE: Reinforced Motion Transformation Network for Semi-supervised 2D Pose Estimation in Videos

AAAI 2022technical

Existing approaches for 2D pose estimation in videos often require a large number of dense annotations, which are costly and labor intensive to acquire. In this paper, we propose a semi-supervised REinforced MOtion Transformation nEtwork (REMOTE) to leverage a few labeled frames and temporal pose va…

Cited by 13SourcePDFScholar
2022

RelTransformer: A Transformer-Based Long-Tail Visual Relationship Recognition

CVPR 2022poster

The visual relationship recognition (VRR) task aims at understanding the pairwise visual relationships between interacting objects in an image. These relationships typically have a long-tail distribution due to their compositional nature. This problem gets more severe when the vocabulary becomes lar…

Cited by 21PDFcodeScholar
2022

Resolution-Free Point Cloud Sampling Network with Data Distillation

ECCV 2022poster

"Down-sampling algorithms are adopted to simplify the point clouds and save the computation cost on subsequent tasks. Existing learning-based sampling methods often need to train a big sampling network to support sampling under different resolutions, which must generate sampled points with the costl…

2022

SuperLine3D: Self-Supervised Line Segmentation and Description for LiDAR Point Cloud

ECCV 2022poster

"Poles and building edges are frequently observable objects on urban roads, conveying reliable hints for various computer vision tasks. To repetitively extract them as features and perform association between discrete LiDAR frames for registration, we propose the first learning-based feature segment…

2022

Towards Multi-Domain Single Image Dehazing via Test-Time Training

CVPR 2022poster

Recent years have witnessed significant progress in the area of single image dehazing, thanks to the employment of deep neural networks and diverse datasets. Most of the existing methods perform well when the training and testing are conducted on a single dataset. However, they are not able to handl…

Cited by 64PDFScholar
2022

VisualGPT: Data-Efficient Adaptation of Pretrained Language Models for Image Captioning

CVPR 2022poster

The limited availability of annotated data often hinders real-world applications of machine learning. To efficiently learn from small quantities of multimodal data, we leverage the linguistic knowledge from a large pre-trained language model (PLM) and quickly adapt it to new domains of image caption…

Cited by 277PDFcodeScholar
2021

A Novel Sequence-to-Subgraph Framework for Diagnosis Classification

IJCAI 2021poster

Text-based diagnosis classification is a critical problem in AI-enabled healthcare studies, which assists clinicians in making correct decision and lowering the rate of diagnostic errors. Previous studies follow the routine of sequence based deep learning models in NLP literature to deal with clinic…

2021

Distributed Multi-Target Tracking for Heterogeneous Mobile Sensing Networks with Limited Field of Views

ICRA 2021poster

This paper introduces the normalized unused sensing capacity to measure the amount of information that a sensor is currently gathering relative to its theoretical maximum. This quantity can be computed using entirely local information and works for arbitrary sensor models, unlike previous literature…

Cited by 31SourceScholar
2021

Exploring Long Tail Visual Relationship Recognition With Large Vocabulary

ICCV 2021poster

Several approaches have been proposed in recent literature to alleviate the long-tail problem, mainly in object classification tasks. In this paper, we make the first large-scale study concerning the task of Long-Tail Visual Relationship Recognition (LTVRR). LTVRR aims at improving the learning of s…

Cited by 22PDFcodeScholar
2021

Universal Rate-Distortion-Perception Representations for Lossy Compression

NeurIPS 2021poster

In the context of lossy compression, Blau \& Michaeli (2019) adopt a mathematical notion of perceptual quality and define the information rate-distortion-perception function, generalizing the classical rate-distortion tradeoff. We consider the notion of universal representations in which one may fix…

Cited by 70SourcePDFScholar
2020

Collision-Free Distributed Multi-Target Tracking Using Teams of Mobile Robots with Localization Uncertainty

IROS 2020poster

Accurately tracking dynamic targets relies on robots accounting for uncertainties in their own states to share information and maintain safety. The problem becomes even more challenging when there is an unknown and time-varying number of targets in the environment. In this paper we address this prob…

Cited by 36SourceScholar
2020

Image Super-Resolution Using Residual Global Context Network

ICASSP 2020accepted

Recent studies have showed that convolutional neural networks (CNN) can effectively improve the performance of single image super-resolution (SR). However, previous methods rarely considered long-range dependencies between pixels and channel-wise interdependencies at the same time. They ignores the…

Cited by 0SourceScholar
2020

Temporal Positive-unlabeled Learning for Biomedical Hypothesis Generation via Risk Estimation

NeurIPS 2020poster

Understanding the relationships between biomedical terms like viruses, drugs, and symptoms is essential in the fight against diseases. Many attempts have been made to introduce the use of machine learning to the scientific process of hypothesis generation (HG), which refers to the discovery of meani…

Cited by 15SourcePDFScholar
2020

The Graph-based Mutual Attentive Network for Automatic Diagnosis

IJCAI 2020poster

The automatic diagnosis has been suffering from the problem of inadequate reliable corpus to train a trustworthy predictive model. Besides, most of the previous deep learning based diagnosis models adopt the sequence learning techniques (CNN or RNN), which is difficult to extract the complex structu…

2020

When Pedestrian Detection Meets Nighttime Surveillance: A New Benchmark

IJCAI 2020poster

Pedestrian detection at nighttime is a crucial and frontier problem in surveillance, but has not been well explored by the computer vision and artificial intelligence communities. Most of existing methods detect pedestrians under favorable lighting conditions (e.g. daytime) and achieve promising per…

2019

Cross-view Identical Part Area Alignment for Person Re-identification

ICASSP 2019accepted

Person re-identification aims to associate images captured by non-overlapping cameras. It is a challenging task because images are often in different conditions such as background clutter, illumination variation, viewpoint changes and different camera settings. Viewpoint changes and pose variations…

Cited by 0SourceScholar
2019

Fast Free-viewpoint Video Synthesis Algorithm for Sports Scenes

IROS 2019poster

In this paper, we report on a parallel free-viewpoint video synthesis algorithm that can efficiently reconstruct a high-quality 3D scene representation of sports scenes. The proposed method focuses on a scene that is captured by multiple synchronized cameras featuring wide-baselines. The following s…

Cited by 21SourceScholar
2019

GridDehazeNet: Attention-Based Multi-Scale Network for Image Dehazing

ICCV 2019poster

We propose an end-to-end trainable Convolutional Neural Network (CNN), named GridDehazeNet, for single image dehazing. The GridDehazeNet consists of three modules: pre-processing, backbone, and post-processing. The trainable pre-processing module can generate learned inputs with better diversity and…

Cited by 1096PDFcodeScholar
2019

Rain Streak Removal via Multi-scale Mixture Exponential Power Model

ICASSP 2019accepted

Rain streaks severely hamper the visible performance of the outdoor surveillance videos, which becomes an attractive issue in recent computer vision research. Existing methods usually encode rain streaks into Gaussian Mixture Model (GM-M). However, the limited number of Gaussian components in the GM…

Cited by 0SourceScholar
2017

Transferring clothing parsing from fashion dataset to surveillance

ICASSP 2017accepted

In this paper we address the problem of automatic clothing parsing in surveillance video with the information from user-generated tags such as “jeans” and “T-shirt”. Although clothing parsing has achieved great success in fashion clothing, it is quite challenging to parse clothing in practical surve…

Cited by 0SourceScholar
2016

Multiple instance discriminative dictionary learning for action recognition

ICASSP 2016accepted

Action recognition from video is a prominent research area in computer vision, with far-reaching applications. Current state-of-the-art action recognition methods is Fisher Vector (FV) coding model based on spatio-temporal local features. Though high dimensional local features have more representati…

Cited by 0SourceScholar