← Search

Jie Chen

214 accepted papers

2026

Anchor Frame Bridging for Coherent First-Last Frame Video Generation

ICLR 2026poster

First-last frame video generation has recently gained significant attention. It enables coherent motion generation between specified first and last frames. However, this approach suffers from semantic degradation in intermediate frames, causing scene distortion and subject deformation that undermine…

Cited by 0SourceScholar
2026

AnoMamba: Aligning Reconstruction with Time Series Anomaly Detection via Selective Global Dependency Modeling

IJCAI 2026

Reconstruction-based frameworks are widely adopted in Time Series Anomaly Detection (TSAD), assuming that models reconstruct normal behavior well but yield larger errors on anomalies. However, in unsupervised TSAD, minimizing reconstruction loss alone often breaks this assumption. Models tend to ove

Cited by 0Scholar
2026

ApET: Approximation-Error Guided Token Compression for Efficient VLMs

CVPR 2026

Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead and degrade inference efficiency. Prior studies typically relies on [CLS] attention or text-vision cross-attention to iden

Cited by 0SourcecodeScholar
2026

BiHiTo: Biomolecular Hierarchy-inspired Tokenization

AAAI 2026technical

Three-dimensional atomic arrangements of biomolecules are key to demystifying biological functions. The rapid expansion of accessible structural data, driven by advances in AI for science, highlights the critical challenge of efficiently modeling large-scale biomolecular structures, which are high-d

Cited by 0SourcePDFScholar
2026

BioDynaSpec: Harmonic-Guided Spatio-Spectral Autoregressive Diffusion for Protein Dynamics Generation

ICML 2026poster

Generating long-horizon, all-atom molecular dynamics (MD) is difficult due to error accumulation in time-domain autoregressive models (causing drift) and fixed step-size constraints on temporal resolution. We propose **BioDynaSpec**, which reformulates protein dynamics as spatio-spectral generation:…

Cited by 0SourceScholar
2026

CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

CVPR 2026

With the rapid advancement of artificial intelligence generated content (AIGC) technologies, video forgery has become increasingly prevalent, posing new challenges to public discourse and societal security. Despite remarkable progress in existing deepfake detection methods, AIGC forgery detection re

Cited by 0SourcecodeScholar
2026

CoGenSAM: Codebook-Interactive Generative Labeling for Adapting SAM to Crack Segmentation

AAAI 2026technical

The goal of this work is to adapt Segment Anything Models (SAM) into crack segmentation tasks via automatic label generation, thus eliminating manual annotation cost. In this regard, an intuitive approach is to extract edges of crack samples and generate labels via the dilation and erosion processes

Cited by 0SourcePDFScholar
2026

Comp-Attn: Present-and-Align Attention for Compositional Video Genneration

ICML 2026poster

In the domain of text-to-video (T2V) generation, reliably synthesizing compositional content involving multiple subjects with intricate relations is still underexplored. The main challenges are twofold: 1) Subject presence, where not all subjects can be presented in the video; 2) Inter-subject relat…

Cited by 0SourceScholar
2026

Condition-Aware Graph Flow Matching for Modeling the Distributions of Complex Physical Systems

ICML 2026poster

Accurately modeling the full distributions of possible states is crucial for understanding statistical properties and enabling reliable predictions in complex, unsteady physical systems. Recently, diffusion models and flow matching have shown promise in these tasks. However, they remain limited in u…

Cited by 0SourceScholar
2026

Deep Inverse Shading: Consistent Albedo and Surface Detail Recovery via Generative Refinement

AAAI 2026technical

Reconstructing human avatars using generative priors is essential for achieving versatile and realistic avatar models. Traditional approaches often rely on volumetric representations guided by generative models, but these methods require extensive volumetric rendering queries, leading to slow traini

Cited by 0SourcePDFScholar
2026

Graph Diffusion Transformers are In-Context Molecular Designers

ICLR 2026poster

In-context learning lets large models adapt to new tasks from a few demonstrations, but it has shown limited success in molecular design, where labeled data are scarce and properties span millions of biological assays and material measurements. We introduce demonstration-conditioned diffusion models…

Cited by 0SourcecodeScholar
2026

Graph-of-Agents: A Graph-based Framework for Multi-Agent LLM Collaboration

ICLR 2026poster

With an ever-growing zoo of LLMs and benchmarks, the need to orchestrate multiple models for improved task performance has never been more pressing. While frameworks like Mixture-of-Agents (MoA) attempt to coordinate LLMs, they often fall short in terms of (1) selecting relevant agents, (2) facilita…

Cited by 0SourceScholar
2026

Mixture-of-World Models: Scaling Multi-Task Reinforcement Learning with Modular Latent Dynamics

ICLR 2026poster

A fundamental challenge in multi-task reinforcement learning (MTRL) is achieving sample efficiency in visual domains where tasks exhibit significant heterogeneity in both observations and dynamics. Model-based RL (MBRL) offers a promising path to sample efficiency through world models, but standard…

Cited by 0SourceScholar
2026

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

AAAI 2026technical

Although Vision Language Models (VLMs) have shown generalization in medical imaging, pathology presents unique challenges due to ultra-high resolution, complex tissue structures, and nuanced semantics. These factors make pathology VLMs prone to hallucinations, i.e., generating outputs inconsistent w

Cited by 0SourcePDFScholar
2026

Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner

AAAI 2026technical

Recent advances in vision-language models (VLMs) have enabled broad progress in the general medical field. However, pathology still remains a more challenging sub-domain, with current pathology-specific VLMs exhibiting limitations in both diagnostic accuracy and reasoning plausibility. Such shortcom

Cited by 0SourcePDFScholar
2026

ProAR: Probabilistic Autoregressive Modeling for Molecular Dynamics

AAAI 2026technical

Understanding the structural dynamics of biomolecules is crucial for uncovering biological functions. As molecular dynamics (MD) simulation data becomes more available, deep generative models have been developed to synthesize realistic MD trajectories. However, existing methods produce fixed-length

Cited by 0SourcePDFScholar
2026

Reliable Policy Transfer for Safety-Aware End-to-End Driving with Deep Reinforcement Learning

CVPR 2026

End-to-End (E2E) Reinforcement Learning (RL) for autonomous driving still struggles with safety and generalization under distribution shift, as perception-heavy encoders, sparse rewards, and ad hoc uncertainty handling yield brittle closed-loop behavior. This work introduces a unified Deep RL (DRL)

Cited by 0SourcecodeScholar
2026

SOSControl: Enhancing Human Motion Generation Through Saliency-Aware Symbolic Orientation and Timing Control

AAAI 2026technical

Traditional text-to-motion frameworks often lack precise control, and existing approaches based on joint keyframe locations provide only positional guidance, making it challenging and unintuitive to specify body part orientations and motion timing. To address these limitations, we introduce the Sali

Cited by 0SourcePDFScholar
2026

The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path

ICML 2026poster

Large Language Models (LLMs) fundamentally suffer from representation collapse, a bottleneck that severely degrades performance in long contexts. We identify that existing approaches risk drifting into one of two pathological extremes: Homogenization Collapse (e.g., attention sinks causing rank defi…

Cited by 0SourceScholar
2026

UniAPO: Unified Multimodal Automated Prompt Optimization

AAAI 2026technical

Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, demonstrating effectiveness primarily in text-only input scenarios. However, extending existing APO methods to multimodal t

Cited by 0SourcePDFScholar
2026

WaveFormer: Frequency-Time Decoupled Vision Modeling with Wave Equation

AAAI 2026technical

Vision modeling has advanced rapidly with Transformers, whose attention mechanisms capture visual dependencies but lack a principled account of how semantic information propagates spatially. We revisit this problem from a wave-based perspective: feature maps are treated as spatial signals whose evol

Cited by 0SourcePDFScholar
2025

Achieving Lift-to-Weight Ratio >3.5 in Piezoelectric Direct-Driven Insect-Scale Flapping-Wing MAVs

IROS 2025

Insect-scale flapping-wing micro aerial vehicles (FWMAVs) employing piezoelectric direct-drive configurations eliminate traditional kinematic chains through direct coupling of the wing and actuator. While this design approach significantly reduces structural complexity and manufacturing costs compar

Cited by 0SourceScholar
2025

Adversarial Diffusion Compression for Real-World Image Super-Resolution

CVPR 2025poster

Real-world image super-resolution (Real-ISR) aims to reconstruct high-resolution images from low-resolution inputs degraded by complex, unknown processes. While many Stable Diffusion (SD)-based Real-ISR methods have achieved remarkable success, their slow, multi-step inference hinders practical depl…

2025

Adversarial Learning Under Hybrid Perturbations for Robust Acute Lymphoblastic Leukemia Classification

AAAI 2025technical

Acute lymphoblastic leukemia is a childhood cancer prevalent worldwide, which can prove fatal within weeks or months. However, current diagnosis models based on machine learning and deep learning methods fail to consider device noise (pixel-level perturbations) and rotation/translation (spatial-tran…

Cited by 0SourcePDFScholar
2025

Aligning Instance Brownian Bridge with Texts for Open-Vocabulary Video Instance Segmentation

AAAI 2025technical

Temporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video data, previous methods leverage the image-text pretraining model for recognizing object instances by separately aligning…

Cited by 0SourcePDFScholar
2025

Attack-inspired Calibration Loss for Calibrating Crack Recognition

AAAI 2025technical

Deep neural networks (DNNs) have substantially achieved high predictive accuracy in many vision tasks. However, we find that they are poorly calibrated for crack recognition tasks, as these DNNs tend to produce both under-confident and over-confident predictions in such safety-critical applications,…

2025

Binary Representation Learning for Discriminative Acoustic Unit Discovery

ICASSP 2025accepted

Acoustic Unit Discovery (AUD) aims to obtain phoneme-like units that preserve linguistically significant information while removing paralinguistic details. Although Contrastive Predictive Coding (CPC) has emerged as a leading self-supervised representation learning method for this task, CPC-based me…

Cited by 0SourceScholar
2025

CLEP: A Novel Contrastive Learning Method for Evolutionary Reentrancy Vulnerability Detection

AAAI 2025technical

Reentrancy vulnerabilities in smart contracts have been exploited to steal enormous amounts of money, thus detecting reentrancy vulnerabilities is a hotspot issue in security research. However, a new attack is emerging in which attackers continuously release new reentrancy patterns to exploit fresh…

Cited by 0SourcePDFScholar
2025

Causality Meets the Table: Debiasing LLMs for Faithful TableQA via Front-Door Intervention

NeurIPS 2025poster

Table Question Answering (TableQA) combines natural language understanding and structured data reasoning, posing challenges in semantic interpretation and logical inference. Recent advances in Large Language Models (LLMs) have improved TableQA performance through Direct Prompting and Agent paradigms…

Cited by 0SourceScholar
2025

Cross-View Graph Consistency Learning for Invariant Graph Representations

AAAI 2025technical

Graph representation learning is fundamental for analyzing graph-structured data. Exploring invariant graph representations remains a challenge for most existing graph representation learning methods. In this paper, we propose a cross-view graph consistency learning (CGCL) method that learns invaria…

2025

DASH: 4D Hash Encoding with Self-Supervised Decomposition for Real-Time Dynamic Scene Rendering

ICCV 2025poster

Dynamic scene reconstruction is a long-term challenge in 3D vision. Existing plane-based methods in dynamic Gaussian splatting suffer from an unsuitable low-rank assumption, causing feature overlap and poor rendering quality. Although 4D hash encoding provides an explicit representation without low-…

2025

Deep Compositional Phase Diffusion for Long Motion Sequence Generation

NeurIPS 2025oral

Recent research on motion generation has shown significant progress in generating semantically aligned motion with singular semantics. However, when employing these models to create composite sequences containing multiple semantically generated motion clips, they often struggle to preserve the conti…

Cited by 0SourcecodeScholar
2025

Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy

AAAI 2025technical

Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. It is challenging to timely prevent such model-stealing attacks to achieve strong protection and maintain utility. In this paper, we propose a novel non-parametric detector…

2025

Derailer-Rerailer: Adaptive Verification for Efficient and Reliable Language Model Reasoning

ACL 2025finding

Large Language Models (LLMs) have shown impressive reasoning capabilities, yet existing prompting methods face a critical trade-off: simple approaches often struggle with complex tasks and reasoning stability, while more sophisticated methods require multiple inferences and substantial computational…

2025

DigitalLLaVA: Incorporating Digital Cognition Capability for Physical World Comprehension in Multimodal LLMs

AAAI 2025technical

Multimodal Large Language Models (MLLMs) have shown remarkable cognitive capabilities in various cross-modal tasks.However, existing MLLMs struggle with tasks that require physical digital cognition, such as accurately reading an electric meter or pressure gauge. This limitation significantly reduce…

Cited by 0SourcePDFScholar
2025

Directed Graph Grammars for Sequence-based Learning

ICML 2025poster

Directed acyclic graphs (DAGs) are a class of graphs commonly used in practice, with examples that include electronic circuits, Bayesian networks, and neural architectures. While many effective encoders exist for DAGs, it remains challenging to decode them in a principled manner, because the nodes o…

2025

Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection

NeurIPS 2025poster

Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel objects with only a handful of labeled samples from previously unseen domains. While data augmentation and generative methods have shown promise in few-shot learning, their effectiveness for CD-FSOD remains unclear due to the need…

Cited by 0SourcecodeScholar
2025

DyMoDreamer: World Modeling with Dynamic Modulation

NeurIPS 2025poster

A critical bottleneck in deep reinforcement learning (DRL) is sample inefficiency, as training high-performance agents often demands extensive environmental interactions. Model-based reinforcement learning (MBRL) mitigates this by building world models that simulate environmental dynamics and genera…

Cited by 0SourceScholar
2025

Efficient Spiking Point Mamba for Point Cloud Analysis

ICCV 2025poster

Bio-inspired Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. However, existing 3D SNNs have struggled with long-range dependencies until the recent emergence of Mamba, which offers superior computational efficiency and sequence modeling capabili…

2025

Enhancing Language Model Hypernetworks with Restart: A Study on Optimization

NAACL 2025long

Hypernetworks are a class of meta-networks that generate weights for main neural networks. Their unique parameter spaces necessitate exploring suitable optimization strategies to enhance performance, especially for language models. However, a comprehensive investigation into optimization strategies…

2025

Exploring Simple Siamese Network for High-Resolution Video Quality Assessment

ICASSP 2025accepted

In the research of video quality assessment (VQA), two-branch network [1] has emerged as a promising solution. It decouples VQA with separate technical and aesthetic branches to measure the perception of low-level distortions and high-level semantics respectively. However, we argue that while techni…

Cited by 0SourceScholar
2025

Exploring Spectral Signatures of Chinese liquor using Machine Learning and SHapley Additive exPlanations

ICASSP 2025accepted

Chinese liquor holds great cultural and economic significance globally. The accurate classification of aroma types and alcohol content is crucial for quality control in Chinese liquor production. To address limitations such as subjectivity and sensor drift in current methods, this study introduces a…

Cited by 0SourceScholar
2025

Forget the Unneeded: Backdooring Large Language Models via Contrastive-enhanced Machine Unlearning

EMNLP 2025

Prompt tuning for Large Language Models (LLMs) is vulnerable to backdoor attacks. Existing methods find backdoor attacks to be a significant threat in data-rich scenarios. However, in data-limited scenarios, these methods have difficulty capturing precise backdoor patterns, leading to weakened backd

2025

Foundation Molecular Grammar: Multi-Modal Foundation Models Induce Interpretable Molecular Graph Languages

ICML 2025poster

Recent data-efficient molecular generation approaches exploit graph grammars to introduce interpretability into the generative models. However, grammar learning therein relies on expert annotation or unreliable heuristics for algorithmic inference. We propose Foundation Molecular Grammar (FMG), whic…

2025

GPEN: Global Position Encoding Network for Enhanced Subgraph Representation Learning

ICML 2025poster

Subgraph representation learning has attracted growing interest due to its wide applications in various domains. However, existing methods primarily focus on local neighborhood structures while overlooking the significant impact of global structural information, in particular the influence of multi-…

Cited by 0SourcePDFScholar
2025

Lessons Learned: A Multi-Agent Framework for Code LLMs to Learn and Improve

NeurIPS 2025poster

Recent studies show that LLMs possess different skills and specialize in different tasks. In fact, we observe that their varied performance occur in several levels of granularity. For example, in the code optimization task, code LLMs excel at different optimization categories and no one dominates ot…

Cited by 0SourceScholar
2025

Lightweight Self-Supervised Monocular Depth Estimation for All-Day Scenes Using Generative Adversarial Network

ICASSP 2025accepted

Self-supervised monocular depth estimation (MDE) has achieved performance levels comparable to supervised methods in well-lit environments. However, current methods struggle particularly with challenging nighttime scenes. Existing all-day self-supervised MDE methods often rely on specialized nightti…

Cited by 0SourceScholar
2025

MTPNet: Multi-Grained Target Perception for Unified Activity Cliff Prediction

IJCAI 2025

Activity cliff prediction is a critical task in drug discovery and material design. Existing computational methods are limited to handling single binding targets, which restricts the applicability of these prediction models. In this paper, we present the Multi-Grained Target Perception network (MTPN

2025

Multimodal Large Language Models for Inverse Molecular Design with Retrosynthetic Planning

ICLR 2025poster

While large language models (LLMs) have integrated images, adapting them to graphs remains challenging, limiting their applications in materials and drug design. This difficulty stems from the need for coherent autoregressive generation across texts and graphs. To address this, we introduce Llamole,…

2025

NCDI-Diffusion: Neural Contextual and Directional Inversion for Novel View Synthesis through Diffusion Models

ICASSP 2025accepted

Novel view synthesis typically requires a comprehensive set of multi-view images for either image-based rendering or scene representation-based optimization. However, achieving high-fidelity novel view rendering often demands a large number of images. To address this limitation, we propose NCDI-Diff…

Cited by 0SourceScholar
2025

PhySpec: Physically Consistent Spectral Reconstruction via Orthogonal Subspace Decomposition and Self-Supervised Meta-Auxiliary Learning

ICML 2025spotlight

This paper presents a novel approach to hyperspectral image (HSI) reconstruction from RGB images, addressing fundamental limitations in existing learning-based methods from a physical perspective. We discuss and aim to address the ``colorimetric dilemma": failure to consistently reproduce ground-tru…

Cited by 0SourcePDFScholar
2025

Procedural Synthesis of Synthesizable Molecules

ICLR 2025poster

Designing synthetically accessible molecules and recommending analogs to unsynthesizable molecules are important problems for accelerating molecular discovery. We reconceptualize both problems using ideas from program synthesis. Drawing inspiration from syntax-guided synthesis approaches, we decoupl…

2025

RA-NeRF: Robust Neural Radiance Field Reconstruction with Accurate Camera Pose Estimation under Complex Trajectories

IROS 2025

Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have emerged as powerful tools for 3D reconstruction and SLAM tasks. However, their performance depends heavily on accurate camera pose priors. Existing approaches attempt to address this issue by introducing external constraints but fal

Cited by 1SourceScholar
2025

Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling

NAACL 2025long

Self-consistency mitigates hallucinations in Large Language Models (LLMs) by sampling multiple reasoning paths, but it lacks a systematic approach to determine the optimal number of samples or select the most faithful rationale. To address this limitation, we introduce Reasoning-Aware Self-Consisten…

Cited by 0SourcePDFScholar
2025

RetinaStereo: Dynamic-Volume Stereo Matching Network

ICASSP 2025accepted

Existing stereo matching techniques often struggle with detailing subtle objects on depth edges. To alleviate this problem, we introduced the Dynamic-Range Disparity Initialization module, which integrates three complementary branches: the dynamic dense volume for localized disparity sampling, the s…

Cited by 0SourceScholar
2025

Robust Offline Imitation Learning Through State-level Trajectory Stitching

IROS 2025

Imitation learning (IL) has proven effective for enabling robots to acquire visuomotor skills through expert demonstrations. However, traditional IL methods are limited by their reliance on high-quality, often scarce, expert data, and suffer from covariate shift. To address these challenges, recent

Cited by 0SourcecodeScholar
2025

SE-GUI: Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

NeurIPS 2025poster

Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to precise interface elements remains challenging—especially in complex, high-resolution, professional environments. Tradit…

Cited by 0SourceScholar
2025

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework

EMNLP 2025

Large reasoning models (LRMs) have exhibited strong performance on complex reasoning tasks, with further gains achievable through increased computational budgets at inference. However, current test-time scaling methods predominantly rely on redundant sampling, ignoring the historical experience util

2025

Teaching Language Models to Critique via Reinforcement Learning

ICML 2025poster

Teaching large language models (LLMs) to critique and refine their outputs is crucial for building systems that can iteratively improve, yet it is fundamentally limited by the ability to provide *accurate judgments* and *actionable suggestions*. In this work, we study LLM critics for code generation…

Cited by 2SourcePDFScholar
2025

Temporal-aware Query Routing for Real-time Video Instance Segmentation

ICCV 2025poster

With the rise of applications such as embodied intelligence, developing high real-time online video instance segmentation (VIS) has become increasingly important. However, through time profiling of the components in advanced online VIS architecture (i.e., transformer-based architecture), we find tha…

Cited by 0SourcePDFScholar
2025

Three-DOF controlled flight in palm-scale micro robotic blimp driven by flapping wings

IROS 2025

Micro blimps exhibit significant potential for applications in environmental monitoring and disaster rescue. Nonetheless, traditional propulsion methods for micro blimps encounter challenges such as complex mechanical structures, intricate attitude control, and large volumes. This paper present a no

Cited by 0SourceScholar
2025

Towards Effective and Efficient Continual Pre-training of Large Language Models

ACL 2025long

Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. In this paper, we comprehensively study its key designs to balance the new abilities while retaining the original abilities, and present an effective CPT method that can greatly imp…

2025

Triples as the Key: Structuring Makes Decomposition and Verification Easier in LLM-based TableQA

ICLR 2025poster

As the mainstream approach, LLMs have been widely applied and researched in TableQA tasks. Currently, the core of LLM-based TableQA methods typically include three phases: question decomposition, sub-question TableQA reasoning, and answer verification. However, several challenges remain in this proc…

Cited by 0SourcePDFScholar
2025

Tune-Your-Style: Intensity-tunable 3D Style Transfer with Gaussian Splatting

ICCV 2025poster

3D style transfer refers to the artistic stylization of 3D assets based on reference style images. Recently, 3DGS-based stylization methods have drawn considerable attention, primarily due to their markedly enhanced training and rendering speeds. However, a vital challenge for 3D style transfer is t…

2025

Unified Arbitrary-Time Video Frame Interpolation and Prediction

ICASSP 2025accepted

Video frame interpolation and prediction aim to synthesize frames in-between and subsequent to existing frames, respectively. Despite being closely-related, these two tasks are traditionally studied with different model architectures, or same architecture but individually trained weights. Furthermor…

Cited by 0SourceScholar
2025

Unraveling Metameric Dilemma for Spectral Reconstruction: A High-Fidelity Approach via Semi-Supervised Learning

NeurIPS 2025poster

Spectral reconstruction from RGB images often suffers from a metameric dilemma, where distinct spectral distributions map to nearly identical RGB values, making them indistinguishable to current models and leading to unreliable reconstructions. In this paper, we present Diff-Spectra that integrates…

Cited by 0SourceScholar
2025

YuLan-Mini: Pushing the Limits of Open Data-efficient Language Model

ACL 2025long

Due to the immense resource demands and the involved complex techniques, it is still challenging for successfully pre-training a large language models (LLMs) with state-of-the-art performance. In this paper, we explore the key bottlenecks and designs during pre-training, and make the following contr…

Cited by 0SourcePDFScholar
2025

iSegMan: Interactive Segment-and-Manipulate 3D Gaussians

CVPR 2025poster

The efficient rendering and explicit nature of 3DGS promote the advancement of 3D scene manipulation.However, existing methods typically encounter challenges in controlling the manipulation region and are unable to furnish the user with interactive feedback, which inevitably leads to unexpected resu…

Cited by 0SourcePDFScholar
2024

A Point-Line Features Fusion Method for Fast and Robust Monocular Visual-Inertial Initialization

IROS 2024poster

Fast and robust initialization is essential for highly accurate monocular visual-inertial odometer (VIO), but at present majority of initialization methods rely only on point features, unstable in low texture and blurring situations. Therefore, we propose a novel point-line features fusion method fo…

Cited by 0SourceScholar
2024

Automated Label Unification for Multi-Dataset Semantic Segmentation with GNNs

NeurIPS 2024poster

Deep supervised models possess significant capability to assimilate extensive training data, thereby presenting an opportunity to enhance model performance through training on multiple datasets. However, conflicts arising from different label spaces among datasets may adversely affect model performa…

2024

Boundary Exploration for Bayesian Optimization With Unknown Physical Constraints

ICML 2024poster

Bayesian optimization has been successfully applied to optimize black-box functions where the number of evaluations is severely limited. However, in many real-world applications, it is hard or impossible to know in advance which designs are feasible due to some physical or system limitations. These…

2024

CF-NeRF: Camera Parameter Free Neural Radiance Fields with Incremental Learning

AAAI 2024technical

Neural Radiance Fields have demonstrated impressive performance in novel view synthesis. However, NeRF and most of its variants still rely on traditional complex pipelines to provide extrinsic and intrinsic camera parameters, such as COLMAP. Recent works, like NeRFmm, BARF, and L2G-NeRF, directly tr…

Cited by 10SourcePDFScholar
2024

DETRs Beat YOLOs on Real-time Object Detection

CVPR 2024poster

The YOLO series has become the most popular framework for real-time object detection due to its reasonable trade-off between speed and accuracy. However we observe that the speed and accuracy of YOLOs are negatively affected by the NMS. Recently end-to-end Transformer-based detectors (DETRs) have pr…

2024

Dynamic Video Frame Interpolation with Integrated Difficulty Pre-Assessment

ICASSP 2024accepted

Video frame interpolation (VFI) has witnessed great progress in recent years. However, existing VFI models still struggle to achieve a good trade-off between accuracy and efficiency. Accurate VFI models typically rely on heavy compute to process all samples, ignoring the fact that easy samples with…

Cited by 0SourceScholar
2024

Exploring Latent Cross-Channel Embedding for Accurate 3d Human Pose Reconstruction in a Diffusion Framework

ICASSP 2024accepted

Monocular 3D human pose estimation poses significant challenges due to the inherent depth ambiguities that arise during the reprojection process from 2D to 3D. Conventional approaches that rely on estimating an over-fit projection matrix struggle to effectively address these challenges and often res…

Cited by 0SourceScholar
2024

Exploring Mathematical Extrapolation of Large Language Models with Synthetic Data

ACL 2024findings

While large language models (LLMs) have shown excellent capabilities in language understanding, text generation and many other tasks, they still struggle in complex multi-step reasoning problems such as mathematical reasoning. In this paper, through a newly proposed arithmetical puzzle problem, we s…

Cited by 2SourcePDFScholar
2024

FaceChain-SuDe: Building Derived Class to Inherit Category Attributes for One-shot Subject-Driven Generation

CVPR 2024poster

Recently subject-driven generation has garnered significant interest due to its ability to personalize text-to-image generation. Typical works focus on learning the new subject's private attributes. However an important fact has not been taken seriously that a subject is not an isolated new concept…

2024

GraCo: Granularity-Controllable Interactive Segmentation

CVPR 2024highlight

Interactive Segmentation (IS) segments specific objects or parts in the image according to user input. Current IS pipelines fall into two categories: single-granularity output and multi-granularity output. The latter aims to alleviate the spatial ambiguity present in the former. However the multi-gr…

2024

Graph Neural Flows for Unveiling Systemic Interactions Among Irregularly Sampled Time Series

NeurIPS 2024poster

Interacting systems are prevalent in nature. It is challenging to accurately predict the dynamics of the system if its constituent components are analyzed independently. We develop a graph-based model that unveils the systemic interactions of time series observed at irregular time points, by using a…

2024

HiCoM: Hierarchical Coherent Motion for Dynamic Streamable Scenes with 3D Gaussian Splatting

NeurIPS 2024poster

The online reconstruction of dynamic scenes from multi-view streaming videos faces significant challenges in training, rendering and storage efficiency. Harnessing superior learning speed and real-time rendering capabilities, 3D Gaussian Splatting (3DGS) has recently demonstrated considerable potent…

Cited by 3SourcePDFScholar
2024

Hyperspectral Image Reconstruction via Combinatorial Embedding of Cross-Channel Spatio-Spectral Clues

AAAI 2024technical

Existing learning-based hyperspectral reconstruction methods show limitations in fully exploiting the information among the hyperspectral bands. As such, we propose to investigate the chromatic inter-dependencies in their respective hyperspectral embedding space. These embedded features can be fully…

2024

LLMBox: A Comprehensive Library for Large Language Models

ACL 2024system demonstrations

To facilitate the research on large language models (LLMs), this paper presents a comprehensive and unified library, LLMBox, to ease the development, use, and evaluation of LLMs. This library is featured with three main merits: (1) a unified data interface that supports the flexible implementation o…

2024

Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation

ECCV 2024poster

"Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily focus on the direct synthesis of global motions while negle…

2024

MA-Stereo: Real-Time Stereo Matching via Multi-Scale Attention Fusion and Spatial Error-Aware Refinement

RA-L 2024

Stereo matching is a fundamental task in computer vision. Real-time stereo matching has recently shown great potential in robotics and autonomous driving applications. However, the existing cost aggregation in real-time stereo matching suffers from accuracy limitations in ill-posed regions. Furtherm

Cited by 4SourceScholar
2024

Make a Strong Teacher with Label Assistance: A Novel Knowledge Distillation Approach for Semantic Segmentation

ECCV 2024poster

"In this paper, we introduce a novel knowledge distillation approach for the semantic segmentation task. Unlike previous methods that rely on power-trained teachers or other modalities to provide additional knowledge, our approach does not require complex teacher models or information from extra sen…

2024

Mind Marginal Non-Crack Regions: Clustering-Inspired Representation Learning for Crack Segmentation

CVPR 2024poster

Crack segmentation datasets make great efforts to obtain the ground truth crack or non-crack labels as clearly as possible. However it can be observed that ambiguities are still inevitable when considering the marginal non-crack region due to low contrast and heterogeneous texture. To solve this pro…

Cited by 13SourcePDFScholar
2024

ParCo: Part-Coordinating Text-to-Motion Synthesis

ECCV 2024poster

"We study a challenging task: text-to-motion synthesis, aiming to generate motions that align with textual descriptions and exhibit coordinated movements. Currently, the part-based methods introduce part partition into the motion synthesis process to achieve finer-grained generation. However, these…

2024

Parallel Vertex Diffusion for Unified Visual Grounding

AAAI 2024technical

Unified visual grounding (UVG) capitalizes on a wealth of task-related knowledge across various grounding tasks via one-shot training, which curtails retraining costs and task-specific architecture design efforts. Vertex generation-based UVG methods achieve this versatility by unified modeling objec…

Cited by 27SourcePDFScholar
2024

Parameterized Approximation Schemes for Fair-Range Clustering

NeurIPS 2024poster

Fair-range clustering extends classical clustering formulations by associating each data point with one or more demographic labels. It imposes lower and upper bound constraints on the number of facilities opened for each label, ensuring fair representation of all demographic groups by the selected f…

Cited by 0SourcePDFScholar
2024

Practical Privacy-Preserving MLaaS: When Compressive Sensing Meets Generative Networks

AAAI 2024technical

The Machine-Learning-as-a-Service (MLaaS) framework allows one to grab low-hanging fruit of machine learning techniques and data science, without either much expertise for this sophisticated sphere or provision of specific infrastructures. However, the requirement of revealing all training data to t…

Cited by 1SourcePDFScholar
2024

Protein-Ligand Interaction Prior for Binding-aware 3D Molecule Diffusion Models

ICLR 2024poster

Generating 3D ligand molecules that bind to specific protein targets via diffusion models has shown great promise for structure-based drug design. The key idea is to disrupt molecules into noise through a fixed forward process and learn its reverse process to generate molecules from noise in a denoi…

2024

Representing Molecules as Random Walks Over Interpretable Grammars

ICML 2024spotlight

Recent research in molecular discovery has primarily been devoted to small, drug-like molecules, leaving many similarly important applications in material design without adequate technology. These applications often rely on more complex molecular structures with fewer examples that are carefully des…

Cited by 3SourcePDFScholar
2024

STL-SLAM: A Structured-Constrained RGB-D SLAM Approach to Texture-Limited Environments

IROS 2024poster

Most RGB-D-based SLAM methods assume texture-rich environments, making them susceptible to significant tracking errors or complete failures in the absence of texture features. Moreover, many existing methods encounter substantial rotation estimation errors, leading to long-term drift in tracking. Th…

Cited by 1SourceScholar
2024

Secure Distributed Sparse Gaussian Process Models Using Multi-Key Homomorphic Encryption

AAAI 2024technical

Distributed sparse Gaussian process (dGP) models provide an ability to achieve accurate predictive performance using data from multiple devices in a time efficient and scalable manner. The distributed computation of model, however, risks exposure of privately owned data to public manipulation. In th…

Cited by 1SourcePDFScholar
2024

Textual Grounding for Open-vocabulary Visual Information Extraction in Layout-diversified Documents

ECCV 2024poster

"Current methodologies have achieved notable success in the closed-set visual information extraction (VIE) task, while the exploration into open-vocabulary settings is comparatively underdeveloped, which is practical for individual users in terms of inferring information across documents of diverse…

Cited by 1SourcePDFScholar
2024

The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language Models

ACL 2024long

In the era of large language models (LLMs), hallucination (the tendency to generate factually incorrect content) poses great challenges to trustworthy and reliable deployment of LLMs in real-world applications. To tackle the hallucination, three key questions should be well studied: how to detect ha…

2024

Unveiling the Flaws: Exploring Imperfections in Synthetic Data and Mitigation Strategies for Large Language Models

EMNLP 2024finding

Synthetic data has been proposed as a solution to address the issue of high-quality data scarcity in the training of large language models (LLMs). Studies have shown that synthetic data can effectively improve the performance of LLMs on downstream benchmarks. However, despite its potential benefits,…

Cited by 8SourcePDFScholar
2023

A Gromov--Wasserstein Geometric View of Spectrum-Preserving Graph Coarsening

ICML 2023poster

Graph coarsening is a technique for solving large-scale graph problems by working on a smaller version of the original graph, and possibly interpolating the results back to the original graph. It has a long history in scientific computing and has recently gained popularity in machine learning, parti…

2023

A Unified Pyramid Recurrent Network for Video Frame Interpolation

CVPR 2023poster

Flow-guided synthesis provides a common framework for frame interpolation, where optical flow is estimated to guide the synthesis of intermediate frames between consecutive inputs. In this paper, we present UPR-Net, a novel Unified Pyramid Recurrent Network for frame interpolation. Cast in a flexibl…

2023

ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation

CVPR 2023highlight

Recently, self-supervised large-scale visual pre-training models have shown great promise in representing pixel-level semantic relationships, significantly promoting the development of unsupervised dense prediction tasks, e.g., unsupervised semantic segmentation (USS). The extracted relationship amo…

Cited by 50SourcePDFScholar
2023

Change Point Detection with Neural Online Density-Ratio Estimator

ICASSP 2023accepted

Detecting change points in streaming time series data is a long standing problem in signal processing. A plethora of methods have been proposed to address it, depending on the hypotheses at hand. Non-parametric approaches are particularly interesting as they do not make any assumption on the distrib…

Cited by 0SourceScholar
2023

Compressed Decentralized Proximal Stochastic Gradient Method for Nonconvex Composite Problems with Heterogeneous Data

ICML 2023poster

We first propose a decentralized proximal stochastic gradient tracking method (DProxSGT) for nonconvex stochastic composite problems, with data heterogeneously distributed on multiple workers in a decentralized connected network. To save communication cost, we then extend DProxSGT to a compressed me…

Cited by 12SourcePDFScholar
2023

DiffusionRet: Generative Text-Video Retrieval with Diffusion Model

ICCV 2023poster

Existing text-video retrieval solutions are, in essence, discriminant models focused on maximizing the conditional likelihood, i.e., p(candidates|query). While straightforward, this de facto paradigm overlooks the underlying data distribution p(query), which makes it challenging to identify out-of-d…

Cited by 72PDFcodeScholar
2023

Discover and Align Taxonomic Context Priors for Open-world Semi-Supervised Learning

NeurIPS 2023poster

Open-world Semi-Supervised Learning (OSSL) is a realistic and challenging task, aiming to classify unlabeled samples from both seen and novel classes using partially labeled samples from the seen classes. Previous works typically explore the relationship of samples as priors on the pre-defined sing…

2023

Edge Devices Friendly Self-Supervised Monocular Depth Estimation via Knowledge Distillation

RA-L 2023

Self-supervised monocular depth estimation (MDE) has great potential for deployment in a wide range of applications, including virtual reality, autonomous driving, and robotics. Nevertheless, most previous studies focused on complex architectures to pursue better performance in MDE. In this letter,

Cited by 11SourceScholar
2023

Efficient and Robust Time-Optimal Trajectory Planning and Control for Agile Quadrotor Flight

RA-L 2023

Agile quadrotor flight relies on rapidly planning and accurately tracking time-optimal trajectories, a technology critical to their application in the wild. However, the computational burden of computing time-optimal trajectories based on the full quadrotor dynamics (typically on the order of minute

Cited by 30SourcecodeScholar
2023

Explain-then-translate: an analysis on improving program translation with self-generated explanations

EMNLP 2023long findings

This work explores the use of self-generated natural language explanations as an intermediate step for code-to-code translation with language models. Across three types of explanations and 19 programming languages constructed from the MultiPL-E dataset, we find the explanations to be particularly e…

Cited by 0SourcecodeScholar
2023

FISS+: Efficient and Focused Trajectory Generation and Refinement Using Fast Iterative Search and Sampling Strategy

IROS 2023poster

Trajectory planning plays a crucial role in autonomous driving systems, as it is tasked to generate feasible trajectories under highly dynamic scenarios within the time constraint. This paper proposes a novel two-stage coarse-to-fine framework for efficient sampling-based trajectory planning. The pr…

Cited by 7SourceScholar
2023

Federated learning of models pre-trained on different features with consensus graphs

UAI 2023poster

Learning an effective global model on private and decentralized datasets has become an increasingly important challenge of machine learning when applied in practice. Existing distributed learning paradigms, such as Federated Learning, enable this via model aggregation which enforces a strong form of…

2023

From Node Interaction To Hop Interaction: New Effective and Scalable Graph Learning Paradigm

CVPR 2023poster

Existing Graph Neural Networks (GNNs) follow the message-passing mechanism that conducts information interaction among nodes iteratively. While considerable progress has been made, such node interaction paradigms still have the following limitation. First, the scalability limitation precludes the br…

2023

Fuzzy Positive Learning for Semi-Supervised Semantic Segmentation

CVPR 2023poster

Semi-supervised learning (SSL) essentially pursues class boundary exploration with less dependence on human annotations. Although typical attempts focus on ameliorating the inevitable error-prone pseudo-labeling, we think differently and resort to exhausting informative semantics from multiple proba…

2023

GC-Flow: A Graph-Based Flow Network for Effective Clustering

ICML 2023poster

Graph convolutional networks (GCNs) are *discriminative models* that directly model the class posterior $p(y|\mathbf{x})$ for semi-supervised classification of graph data. While being effective, as a representation learning approach, the node representations extracted from a GCN often miss useful in…

2023

Graph Neural Network-Inspired Kernels for Gaussian Processes in Semi-Supervised Learning

ICLR 2023poster

Gaussian processes (GPs) are an attractive class of machine learning models because of their simplicity and flexibility as building blocks of more complex Bayesian models. Meanwhile, graph neural networks (GNNs) emerged recently as a promising class of models for graph-structured data in semi-superv…

2023

Hierarchical Grammar-Induced Geometry for Data-Efficient Molecular Property Prediction

ICML 2023poster

The prediction of molecular properties is a crucial task in the field of material and drug discovery. The potential benefits of using deep learning techniques are reflected in the wealth of recent literature. Still, these techniques are faced with a common challenge in practice: Labeled data are lim…

2023

High-Frequency Stereo Matching Network

CVPR 2023highlight

In the field of binocular stereo matching, remarkable progress has been made by iterative methods like RAFT-Stereo and CREStereo. However, most of these methods lose information during the iterative process, making it difficult to generate more detailed difference maps that take full advantage of hi…

Cited by 86SourcePDFScholar
2023

LaPE: Layer-adaptive Position Embedding for Vision Transformers with Independent Layer Normalization

ICCV 2023poster

Position information is critical for Vision Transformers (VTs) due to the permutation-invariance of self-attention operations. A typical way to introduce position information is adding the absolute Position Embedding (PE) to patch embedding before entering VTs. However, this approach operates the sa…

Cited by 10PDFcodeScholar
2023

Learnable Blur Kernel for Single-Image Defocus Deblurring in the Wild

AAAI 2023technical

Recent research showed that the dual-pixel sensor has made great progress in defocus map estimation and image defocus deblurring. However, extracting real-time dual-pixel views is troublesome and complex in algorithm deployment. Moreover, the deblurred image generated by the defocus deblurring netwo…

Cited by 6SourcePDFScholar
2023

Learning to Distill Global Representation for Sparse-View CT

ICCV 2023poster

Sparse-view computed tomography (CT)---using a small number of projections for tomographic reconstruction---enables much lower radiation dose to patients and accelerated data acquisition. The reconstructed images, however, suffer from strong artifacts, greatly limiting their diagnostic value. Curren…

Cited by 12PDFcodeScholar
2023

LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech

ICASSP 2023accepted

Recent advances in neural text-to-speech (TTS) models bring thousands of TTS applications into daily life, where models are deployed in cloud to provide services for customs. Among these models are diffusion probabilistic models (DPMs), which can be stably trained and are more parameter-efficient co…

Cited by 0SourceScholar
2023

Multi-granularity Interaction Simulation for Unsupervised Interactive Segmentation

ICCV 2023poster

Interactive segmentation enables users to segment as needed by providing cues of objects, which introduces human-computer interaction for many fields, such as image editing and medical image analysis. Typically, massive and expansive pixel-level annotations are spent to train deep models by object-o…

Cited by 10PDFScholar
2023

Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation

CVPR 2023poster

Weakly supervised semantic segmentation is typically inspired by class activation maps, which serve as pseudo masks with class-discriminative regions highlighted. Although tremendous efforts have been made to recall precise and complete locations for each class, existing methods still commonly suffe…

2023

Out-of-Distributed Semantic Pruning for Robust Semi-Supervised Learning

CVPR 2023poster

Recent advances in robust semi-supervised learning (SSL) typical filters out-of-distribution (OOD) information at the sample level. We argue that an overlooked problem of robust SSL is its corrupted information on semantic level, practically limiting the development of the field. In this paper, we t…

2023

Proximal Stochastic Recursive Momentum Methods for Nonconvex Composite Decentralized Optimization

AAAI 2023technical

Consider a network of N decentralized computing agents collaboratively solving a nonconvex stochastic composite problem. In this work, we propose a single-loop algorithm, called DEEPSTORM, that achieves optimal sample complexity for this setting. Unlike double-loop algorithms that require a large ba…

2023

Recurrent Fine-Grained Self-Attention Network for Video Crowd Counting

ICASSP 2023accepted

Striking a balance between exploring the spatio-temporal correlation and controlling model complexity is vital for video-based crowd counting methods. In this paper, we propose a Recurrent Fine-Grained Self-Attention Network (RFSNet) to achieve efficient and accurate counting in video scenes via the…

Cited by 0SourceScholar
2023

TG-VQA: Ternary Game of Video Question Answering

IJCAI 2023poster

Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions, i.e., annotations or priors, current contrastive learning-based VideoQA methods remains challenging to perform fine-gr…

Cited by 17SourcePDFScholar
2023

TJ-FlyingFish: Design and Implementation of an Aerial-Aquatic Quadrotor with Tiltable Propulsion Units

ICRA 2023poster

Aerial-aquatic vehicles are capable to move in the two most dominant fluids, making them more promising for a wide range of applications. We propose a prototype with special designs for propulsion and thruster configuration to cope with the vast differences in the fluid properties of water and air.…

Cited by 34SourceScholar
2023

Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment

IJCAI 2023poster

Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local details or are computationally expensive. What's worse, they fail to leverage the heterogeneous concepts in data. In this…

2023

The Devil is in the Crack Orientation: A New Perspective for Crack Detection

ICCV 2023poster

Cracks are usually curve-like structures that are the focus of many computer-vision applications (e.g., road safety inspection and surface inspection of industrial facilities). The existing pixel-based crack segmentation methods rely on time-consuming and costly pixel-level annotations. And the obje…

Cited by 24PDFScholar
2023

TopoSeg: Topology-Aware Nuclear Instance Segmentation

ICCV 2023poster

Nuclear instance segmentation has been critical for pathology image analysis in medical science, e.g., cancer diagnosis. Current methods typically adopt pixel-wise optimization for nuclei boundary exploration, where rich structural information could be lost for subsequent quantitative morphology ass…

Cited by 25PDFcodeScholar
2023

Towards Real-World Burst Image Super-Resolution: Benchmark and Method

ICCV 2023poster

Despite substantial advances, single-image super-resolution (SISR) is always in a dilemma to reconstruct high-quality images with limited information from one input image, especially in realistic scenarios. In this paper, we establish a large-scale real-world burst super-resolution dataset, i.e., Re…

Cited by 16PDFcodeScholar
2023

Video-Text As Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning

CVPR 2023highlight

Contrastive learning-based video-language representation learning approaches, e.g., CLIP, have achieved outstanding performance, which pursue semantic interaction upon pre-defined video-text pairs. To clarify this coarse-grained global interaction and move a step further, we have to encounter challe…

2023

WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image Segmentation

IJCAI 2023poster

The top-down and bottom-up methods are two mainstreams of referring segmentation, while both methods have their own intrinsic weaknesses. Top-down methods are chiefly disturbed by Polar Negative (PN) errors owing to the lack of fine-grained cross-modal alignment. Bottom-up methods are mainly perturb…

Cited by 4SourcePDFScholar
2022

Data-Efficient Graph Grammar Learning for Molecular Generation

ICLR 2022oral

The problem of molecular generation has received significant attention recently. Existing methods are typically based on deep neural networks and require training on large datasets with tens of thousands of samples. In practice, however, the size of class-specific chemical datasets is usually limite…

2022

Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations

NeurIPS 2022accept

Most video-and-language representation learning approaches employ contrastive learning, e.g., CLIP, to project the video and text features into a common latent space according to the semantic similarities of text-video pairs. However, such learned shared latent spaces are not often optimal, and the…

2022

Forest: A Lightweight Semantic Image Descriptor for Robust Visual Place Recognition

RA-L 2022

Visual place recognition (VPR) is the process of identifying previously visited places using visual information. It is crucial for a robot to achieve fast and accurate VPR under appearance and viewpoint changes. Inspired by human perception intuition, this letter proposes a lightweight semantic imag

Cited by 5SourceScholar
2022

Geometry-Aware Guided Loss for Deep Crack Recognition

CVPR 2022poster

Despite the substantial progress of deep models for crack recognition, due to the inconsistent cracks in varying sizes, shapes, and noisy background textures, there still lacks the discriminative power of the deeply learned features when supervised by the cross-entropy loss. In this paper, we propos…

Cited by 35PDFScholar
2022

Hyperspectral Image Super-Resolution with Deep Priors and Degradation Model Inversion

ICASSP 2022accepted

To overcome inherent hardware limitations of hyperspectral imaging systems with respect to their spatial resolution, fusion-based hyper-spectral image (HSI) super-resolution is attracting increasing attention. This technique aims to fuse a low-resolution (LR) HSI and a conventional high-resolution (…

Cited by 0SourceScholar
2022

Locality Guidance for Improving Vision Transformers on Tiny Datasets

ECCV 2022poster

"While the Vision Transformer (VT) architecture is becoming trendy in computer vision, pure VT models perform poorly on tiny datasets. To address this issue, this paper proposes the locality guidance for improving the performance of VTs on tiny datasets. We first analyze that the local information,…

2022

Memory-Based Message Passing: Decoupling the Message for Propagation from Discrimination

ICASSP 2022accepted

Message passing is a fundamental procedure for graph neural networks in the field of graph representation learning. Based on the homophily assumption, the current message passing always aggregates features of connected nodes, such as the graph Laplacian smoothing process. However, real-world graphs…

Cited by 0SourceScholar
2022

Temporal-MPI: Enabling Multi-Plane Images for Dynamic Scene Modelling via Temporal Basis Learning

ECCV 2022poster

"Novel view synthesis of static scenes has achieved remarkable advancements in producing photo-realistic results. However, key challenges remain for immersive rendering of dynamic scenes. One of the seminal image-based rendering method, the multi-plane image (MPI), produces high novel-view synthesis…

Cited by 15SourcePDFScholar
2022

Transient Analysis of Clustered Multitask Diffusion RLS Algorithm

ICASSP 2022accepted

In this paper, we propose a novel clustered multitask diffusion RLS (MT-DRLS) algorithm over network to further improve the performance of its counterpart, the multitask diffusion LMS (MT-DLMS) algorithm. Its transient behavior is investigated, in the mean and mean-square error sense. Simulation res…

Cited by 0SourceScholar
2022

ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

CVPR 2022poster

Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of s…

Cited by 79PDFScholar
2022

When Active Learning Meets Implicit Semantic Data Augmentation

ECCV 2022poster

"Active learning (AL) is a label-efficient technique for training deep models when only a limited labeled set is available and the manual annotation is expensive. Implicit semantic data augmentation (ISDA) effectively extends the limited amount of labeled samples and increases the diversity of label…

Cited by 18SourcePDFScholar
2021

A Multi-Stage Progressive Learning Strategy for Covid-19 Diagnosis Using Chest Computed Tomography with Imbalanced Data

ICASSP 2021accepted

In this paper, a multi-stage progressive learning strategy is investigated to train classifiers for COVID-19 Diagnosis using imbalanced Chest Computed Tomography Data acquired from patients infected with COVID-19 Pneumonia, Community Acquired Pneumonia (CAP) and from normal healthy subjects. In the…

Cited by 0SourceScholar
2021

CDNet: Centripetal Direction Network for Nuclear Instance Segmentation

ICCV 2021poster

Nuclear instance segmentation is a challenging task due to a large number of touching and overlapping nuclei in pathological images. Existing methods cannot effectively recognize the accurate boundary owing to neglecting the relationship between pixels (e.g., direction information). In this paper, w…

Cited by 60PDFcodeScholar
2021

CentripetalText: An Efficient Text Instance Representation for Scene Text Detection

NeurIPS 2021poster

Scene text detection remains a grand challenge due to the variation in text curvatures, orientations, and aspect ratios. One of the hardest problems in this task is how to represent text instances of arbitrary shapes. Although many methods have been proposed to model irregular texts in a flexible ma…

2021

CoLA: Weakly-Supervised Temporal Action Localization With Snippet Contrastive Learning

CVPR 2021poster

Weakly-supervised temporal action localization (WS-TAL) aims to localize actions in untrimmed videos with only video-level labels. Most existing models follow the "localization by classification" procedure: locate temporal regions contributing most to the video-level classification. Generally, they…

Cited by 184PDFcodeScholar
2021

CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks

NeurIPS 2021poster

Over the last several decades, software has been woven into the fabric of every aspect of our society. As software development surges and code infrastructure of enterprise applications ages, it is now more critical than ever to increase software development productivity and modernize legacy applicat…

Cited by 327SourcecodeScholar
2021

Discover Cross-Modality Nuances for Visible-Infrared Person Re-Identification

CVPR 2021poster

Visible-infrared person re-identification (Re-ID) aims to match the pedestrian images of the same identity from different modalities. Existing works mainly focus on alleviating the modality discrepancy by aligning the distributions of features from different modalities. However, nuanced but discrimi…

Cited by 286PDFcodeScholar
2021

Dynamic visualization for L1 fusion convex clustering in near-linear time

UAI 2021poster

Convex clustering has drawn recent attention because of its competitive performance and nice property to guarantee global optimality. However, convex clustering is infeasible due to its high computational cost for large-scale data sets. We propose a novel method to solve the L1 fusion convex cluster…

Cited by 3SourcePDFScholar
2021

Graph Universal Adversarial Attacks: A Few Bad Actors Ruin Graph Learning Models

IJCAI 2021poster

Deep neural networks, while generalize well, are known to be sensitive to small adversarial perturbations. This phenomenon poses severe security threat and calls for in-depth investigation of the robustness of deep learning models. With the emergence of neural networks for graph structured data, sim…

2021

RR-Net: Injecting Interactive Semantics in Human-Object Interaction Detection

IJCAI 2021poster

Human-Object Interaction (HOI) detection devotes to learn how humans interact with surrounding objects. Latest end-to-end HOI detectors are short of relation reasoning, which leads to inability to learn HOI-specific interactive semantics for predictions. In this paper, we therefore propose novel rel…

Cited by 4SourcePDFScholar
2021

ReCU: Reviving the Dead Weights in Binary Neural Networks

ICCV 2021poster

Binary neural networks (BNNs) have received increasing attention due to their superior reductions of computation and memory. Most existing works focus on either lessening the quantization error by minimizing the gap between the full-precision weights and their binarization or designing a gradient ap…

Cited by 114PDFcodeScholar
2021

Unsupervised Learning of Graph Hierarchical Abstractions with Differentiable Coarsening and Optimal Transport

AAAI 2021technical

Hierarchical abstractions are a methodology for solving large-scale graph problems in various disciplines. Coarsening is one such approach: it generates a pyramid of graphs whereby the one in the next level is a structural summary of the prior one. With a long history in scientific computing, many c…

2021

Variational Autoencoders for Hyperspectral Unmixing with Endmember Variability

ICASSP 2021accepted

Spectral signatures are usually affected by variations in environmental conditions. The spectral variability is thus one of the most important and challenging problems to be addressed in hyperspectral unmixing. Generally, it is a non-trivial task to model the endmember variability, and existing spec…

Cited by 0SourceScholar
2020

AD-Cluster: Augmented Discriminative Clustering for Domain Adaptive Person Re-Identification

CVPR 2020poster

Domain adaptive person re-identification (re-ID) is a challenging task, especially when person identities in target domains are unknown. Existing methods attempt to address this challenge by transferring image styles or aligning feature distributions across domains, whereas the rich unlabeled sample…

Cited by 383PDFScholar
2020

Deep Spatial-angular Regularization for Compressive Light Field Reconstruction over Coded Apertures

ECCV 2020poster

Coded aperture is a promising approach for capturing the 4-D light field (LF), in which the 4-D data are compressively modulated into 2-D coded measurements that are further decoded by reconstruction algorithms. The bottleneck lies in the reconstruction algorithms, resulting in rather limited recons…

2020

Dfsmn-San with Persistent Memory Model for Automatic Speech Recognition

ICASSP 2020accepted

Self-attention networks (SAN) have been introduced into automatic speech recognition (ASR) and achieved state-of-the-art performance owing to its superior ability in capturing long term dependency. One of the key ingredients is the self-attention mechanism which can be effectively performed on the w…

Cited by 0SourceScholar
2020

Exploring Entity-Level Spatial Relationships for Image-Text Matching

ICASSP 2020accepted

Exploring the entity-level (i.e., objects in an image, words in a text) spatial relationship contributes to understanding multimedia content precisely. The ignorance of spatial information in previous works probably leads to misunderstandings of image contents. For instance, sentences `Boats are on…

Cited by 0SourceScholar
2020

Integration of Multi-Look Beamformers for Multi-Channel Keyword Spotting

ICASSP 2020accepted

Keyword spotting (KWS) is in great demand in smart devices in the era of Internet of Things. Albeit recent progresses, the performance of KWS, measured in false alarms and false rejects, may still degrade significantly under the far field and noisy conditions. In this paper, we propose integrating m…

Cited by 0SourceScholar
2020

Learning Spectral-Spatial Prior Via 3DDNCNN for Hyperspectral Image Deconvolution

ICASSP 2020accepted

Hyperspectral image (HSI) deconvolution is an ill-posed problem aiming at recovering sharp images with tens or hundreds of spectral channels from blurred and noisy observations. In order to successfully conduct the deconvolution, proper priors are required to regularize the optimization problem. How…

Cited by 0SourceScholar
2020

Learning a Weakly-Supervised Video Actor-Action Segmentation Model With a Wise Selection

CVPR 2020oral

We address weakly-supervised video actor-action segmentation (VAAS), which extends general video object segmentation (VOS) to additionally consider action labels of the actors. The most successful methods on VOS synthesize a pool of pseudo-annotations (PAs) and then refine them iteratively. However,…

Cited by 19PDFScholar
2020

Light Field Spatial Super-Resolution via Deep Combinatorial Geometry Embedding and Structural Consistency Regularization

CVPR 2020poster

Light field (LF) images acquired by hand-held devices usually suffer from low spatial resolution as the limited sampling resources have to be shared with the angular dimension. LF spatial super-resolution (SR) thus becomes an indispensable part of the LF camera processing pipeline. The high-dimensio…

Cited by 188PDFcodeScholar
2020

Low-Rank Approximation of Matrices Via A Rank-Revealing Factorization with Randomization

ICASSP 2020accepted

Given a matrix A with numerical rank k, the two-sided orthogonal decomposition (TSOD) computes a factorization A = UDV <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">T</sup> , where U and V are unitary, and D is (upper/lower) triangular. TSOD is rank-r…

Cited by 0SourceScholar
2020

Online Convex Optimization Over Erdos-Renyi Random Networks

NeurIPS 2020poster

The work studies how node-to-node communications over an Erd\H{o}s-R\'enyi random network influence distributed online convex optimization, which is vital in solving large-scale machine learning in antagonistic or changing environments. At per step, each node (computing unit) makes a local decisi…

Cited by 20SourcePDFScholar
2020

Pixel-Wise Linear/Nonlinear Nonnegative Matrix Factorization for Unmixing of Hyperspectral Data

ICASSP 2020accepted

Nonlinear spectral unmixing is a challenging and important task in hyperspectral image analysis. The kernel-based bi-objective non-negative matrix factorization (Bi-NMF) has shown its usefulness in nonlinear unmixing; However, it suffers several issues that prohibit its practical application. In thi…

Cited by 0SourceScholar
2020

Proximal Multitask Learning Over Distributed Networks with Jointly Sparse Structure

ICASSP 2020accepted

Modeling relations between local optimum parameter vectors in multitask networks has attracted much attention over the last years. This work considers a distributed optimization problem for parameter vectors with a jointly sparse structure among nodes, that is, the parameter vectors share the same s…

Cited by 0SourceScholar
2019

Adaptively Aligned Image Captioning via Adaptive Attention Time

NeurIPS 2019poster

Recent neural models for image captioning usually employ an encoder-decoder framework with an attention mechanism. However, the attention mechanism in such a framework aligns one single (attended) image feature vector to one caption word, assuming one-to-one mapping from source image regions and tar…

2019

Learning Discriminative Features in Sequence Training without Requiring Framewise Labelled Data

ICASSP 2019accepted

In this work, we try to answer two questions: Can deeply learned features with discriminative power benefit an ASR system’s robustness to acoustic variability? And how to learn them without requiring framewise labelled sequence training data? As existing methods usually require knowing where the lab…

Cited by 0SourceScholar
2019

Robust Sparse Multichannel Active Noise Control

ICASSP 2019accepted

Multichannel active noise control (MC-ANC) aims to cancel low-frequency noise in an enclosure. If noise sources are distributed sparsely in space, adding an ℓ <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</inf> -norm constraint to the standard MC-AN…

Cited by 0SourceScholar
2018

Constrained Generation of Semantically Valid Graphs via Regularizing Variational Autoencoders

NeurIPS 2018poster

Deep generative models have achieved remarkable success in various data domains, including images, time series, and natural languages. There remain, however, substantial challenges for combinatorial structures, including graphs. One of the key challenges lies in the difficulty of ensuring semantic v…

2018

Fast Light Field Reconstruction With Deep Coarse-To-Fine Modeling of Spatial-Angular Clues

ECCV 2018poster

Densely-sampled light fields (LFs) are beneficial to many applications such as depth inference and post-capture refocusing. However, it is costly and challenging to capture them. In this paper, we propose a learning based algorithm to reconstruct a densely-sampled LF fast and accurately from a spars…

2018

FastGCN: Fast Learning with Graph Convolutional Networks via Importance Sampling

ICLR 2018poster

The graph convolutional networks (GCN) recently proposed by Kipf and Welling are an effective graph model for semi-supervised learning. Such a model, however, is transductive in nature because parameters are learned through convolutions with both training and test data. Moreover, the recursive neigh…

2018

Robust Video Content Alignment and Compensation for Rain Removal in a CNN Framework

CVPR 2018poster

Rain removal is important for improving the robustness of outdoor vision based systems. Current rain removal methods show limitations either for complex dynamic scenes shot from fast moving cameras, or under torrential rain fall with opaque occlusions. We propose a novel derain algorithm, which appl…

Cited by 209SourcePDFScholar
2018

Super Wide Regression Network for Unsupervised Cross-Database Facial Expression Recognition

ICASSP 2018accepted

Unsupervised cross-database facial expression recognition (FER) is a challenging problem, in which the training and testing samples belong to different facial expression databases. For this reason, the training (source) and testing (target) facial expression samples would have different feature dist…

Cited by 0SourceScholar