← Search

Siyuan Li

103 accepted papers

2026

An HMDP-MPC Decision-Making Framework with Adaptive Safety Margins and Hysteresis for Autonomous Driving

ICRA 2026poster

This paper presents a unified decision-making framework that integrates Hybrid Markov Decision Processes (HMDPs) with Model Predictive Control (MPC), augmented by velocity-dependent safety margins and a prediction-aware hysteresis mechanism. Both the ego and surrounding vehicles are modeled as HMDPs…

2026

Boosting Zero-Shot VLN Via Abstract Obstacle Map-Based Waypoint Prediction with TopoGraph-And-VisitInfo-Aware Prompting

ICRA 2026poster

With the rapid progress of foundation models and robotics, vision-language navigation (VLN) has emerged as a key task for embodied agents with broad practical applications. We address VLN in continuous environments, a particularly challenging setting where an agent must jointly interpret natural lan…

2026

CDBridge: A Cross-omics Post-training Bridge Strategy for Context-aware Biological Modeling

ICLR 2026poster

Linking genomic DNA to quantitative, context-specific expression remains a central challenge in computational biology. Current foundation models capture either tissue context or sequence features, but not both. Cross-omics systems, in turn, often overlook critical mechanisms such as alternative spli…

Cited by 0SourceScholar
2026

Doloris: Dual Conditional Diffusion Implicit Bridges with Sparsity Masking Strategy for Unpaired Single-Cell Perturbation Estimation

ICLR 2026poster

Estimating single-cell responses across various perturbations facilitates the identification of key genes and enhances drug screening, significantly boosting experimental efficiency. However, single-cell sequencing is a destructive process, making it impossible to capture the same cell's phenotype b…

Cited by 0SourcecodeScholar
2026

GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models

CVPR 2026

Unified Multimodal Models (UMMs) are redefining the landscape of artificial intelligence by coupling perception and generation across language, vision, and structured reasoning. Yet, despite their growing sophistication, a critical gap persists in evaluation: existing benchmarks largely measure disc

Cited by 0SourceScholar
2026

Geometric Flow Grounding: A Unified Manifold Decoupling Framework for Dynamics Discovery and Verification

ICML 2026oral

Modeling complex dynamics from observational data is fundamental to scientific discovery and artificial intelligence. However, existing approaches ranging from Neural ODEs to diffusion models are often plagued by the entanglement of static state representations and instantaneous motion, leading to a…

Cited by 0SourceScholar
2026

GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent System

ICLR 2026poster

In the realm of multi-agent systems, the challenge of partial observability is a critical barrier to effective coordination and decision-making. Existing approaches, such as belief state estimation and inter-agent communication, often fall short. Belief-based methods are limited by their focus on pa…

Cited by 0SourceScholar
2026

How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning

ICML 2026spotlight

Solving complex geometric problems inherently requires \textit{interleaved reasoning}: a tight alternation between constructing diagrams and performing logical deductions. Although recent Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities in visual generation and plotting…

Cited by 0SourceScholar
2026

Interpreting Fedspeak with Confidence: A LLM-Based Uncertainty-Aware Framework Guided by Monetary Policy Transmission Paths

AAAI 2026technical

"Fedspeak", the stylized and often nuanced language used by the U.S. Federal Reserve, encodes implicit policy signals and strategic stances. The Federal Open Market Committee strategically employs Fedspeak as a communication tool to shape market expectations and influence both domestic and global e

Cited by 0SourcePDFScholar
2026

LacTokGen: Latent Consistency Tokenizer for 1024-pixel Image Generation by 256 Tokens

CVPR 2026

Image tokenization has significantly advanced visual generation and multimodal modeling, particularly when paired with autoregressive models. However, current methods face challenges in balancing efficiency and quality: high-resolution image generation either requires an excessive number of tokens o

Cited by 0SourcecodeScholar
2026

MergeDNA: Context-Aware Genome Modeling with Dynamic Tokenization Through Token Merging

AAAI 2026technical

Modeling genomic sequences faces two unsolved challenges: the information density varies widely across different regions, while there is no clearly defined minimum vocabulary unit. Relying on either four primitive bases or independently designed DNA tokenizers, existing approaches with naive masked

Cited by 0SourcePDFScholar
2026

MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding

ICLR 2026poster

Vision-language alignment in multi-modal large language models (MLLMs) relies on supervised fine-tuning (SFT) or reinforcement learning (RL). To align multi-modal large language models (MLLMs) in the post-training stage, supervised fine-tuning (SFT) is a stable choice but requires human annotations…

Cited by 0SourcecodeScholar
2026

Model-Agnostic Sentiment Distribution Stability Analysis for Robust LLM-Generated Texts Detection

AAAI 2026technical

The rapid advancement of large language models (LLMs) has resulted in increasingly sophisticated AI-generated content, posing significant challenges in distinguishing LLM-generated text from human-written language. Existing detection methods, primarily based on lexical heuristics or fine-tuned class

Cited by 0SourcePDFScholar
2026

MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios

AAAI 2026technical

Model-based reinforcement learning (MBRL) is a crucial approach to enhance the generalization capabilities and improve the sample efficiency of RL algorithms. However, current MBRL methods focus primarily on building world models for single tasks and rarely address generalization across different s

Cited by 0SourcePDFScholar
2026

SAM 3: Segment Anything with Concepts

ICLR 2026poster

We present Segment Anything Model (SAM) 3, a unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short noun phrases (e.g., “yellow school bus”), image exemplars, or a combination of both. Promptable Concept Segmentation (P…

Cited by 687SourcecodeScholar
2026

SteinsGate: Adding Causality to Diffusions for Long Video Generation via Path Integral

ICLR 2026poster

Video generation has advanced rapidly, but current models remain limited to short clips, far from the length and complexity of real-world narratives. Long video generation is thus both important and challenging. Existing approaches either attempt to extend the modeling length of video diffusion mode…

Cited by 0SourceScholar
2026

TrinityDNA: A Bio-Inspired Foundational Model for Efficient Long-Sequence DNA Modeling

AAAI 2026technical

The modeling of genomic sequences presents unique challenges due to their long length and structural complexity. Traditional sequence models struggle to capture long-range dependencies and biological features inherent in DNA. In this work, we propose TrinityDNA, a novel DNA foundational model design

Cited by 0SourcePDFScholar
2026

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

ICML 2026poster

Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelmed by global context. Recent "Thinking-with-Images" methods alleviate this by iteratively zooming into regions of interes…

Cited by 0SourceScholar
2025

3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection

ICCV 2025poster

Monocular 3D object detection is valuable for various applications such as robotics and AR/VR. Existing methods are confined to closed-set settings, where the training and testing sets consist of the same scenes and/or object categories. However, real-world applications often introduce new environme…

2025

AlphaFold Database Debiasing for Robust Inverse Folding

NeurIPS 2025poster

The AlphaFold Protein Structure Database (AFDB) offers unparalleled structural coverage at near-experimental accuracy, positioning it as a valuable resource for data-driven protein design. However, its direct use in training deep models that are sensitive to fine-grained atomic geometry—such as inve…

Cited by 0SourceScholar
2025

Basis Function Learning for Variable-Length and Continuous-Indexed Signals

ICASSP 2025accepted

Representing variable-length and continuous-indexed signals through a linear combination of basis functions poses a fundamental challenge in science and engineering. Current approaches resort to preprocessing steps, such as interpolation and extrapolation, to handle irregular and off-grid measuremen…

Cited by 0SourceScholar
2025

DaCapo: Score Distillation as Stacked Bridge for Fast and High-quality 3D Editing

CVPR 2025poster

Score Distillation Sampling (SDS) has been successfully extended to text-driven 3D scene editing with 2D pretrained diffusion models. However, SDS-based editing methods suffer from lengthy optimization processes with slow inference and low quality. We attribute the issue of lengthy optimization to t…

Cited by 0SourcePDFScholar
2025

Dual-branch Graph Feature Learning for NLOS Imaging

AAAI 2025technical

The domain of non-line-of-sight (NLOS) imaging is advancing rapidly, offering the capability to reveal occluded scenes that are not directly visible. However, contemporary NLOS systems face several significant challenges: (1) The computational and storage requirements are profound due to the inheren…

Cited by 0SourcePDFScholar
2025

EVA: Geometric Inverse Design for Fast Protein Motif-Scaffolding with Coupled Flow

ICLR 2025poster

Motif-scaffolding is a fundamental component of protein design, which aims to construct the scaffold structure that stabilizes motifs conferring desired functions. Recent advances in generative models are promising for designing scaffolds, with two main approaches: training-based and sampling-based…

Cited by 0SourcePDFScholar
2025

Enhancing Image Generation Fidelity via Progressive Prompts

ICASSP 2025accepted

Diffusion transformer (DiT) architecture catches much attention in image generation, which achieves better fidelity, performance, and diversity. However, most existing DiT-based image generation methods are global-aware synthesis and regional prompt control is less explored. In this paper, we propos…

Cited by 0SourceScholar
2025

Enhancing Rumor Detection Methods with Propagation Structure Infused Language Model

COLING 2025main

Pretrained Language Models (PLMs) have excelled in various Natural Language Processing tasks, benefiting from large-scale pretraining and self-attention mechanism’s ability to capture long-range dependencies. However, their performance on social media application tasks like rumor detection remains s…

2025

From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing

CVPR 2025highlight

We introduce the task of text-to-diagram generation, which focuses on creating structured visual representations directly from textual descriptions. Existing approaches in text-to-image and text-to-code generation lack the logical organization and flexibility needed to produce accurate, editable dia…

Cited by 2SourcePDFScholar
2025

Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation

CVPR 2025poster

Recent advances in zero-shot referring image segmentation (RIS), driven by models such as the Segment Anything Model (SAM) and CLIP, have made substantial progress in aligning visual and textual information. Despite these successes, the extraction of precise and high-quality mask region representati…

2025

MeToken: Uniform Micro-environment Token Boosts Post-Translational Modification Prediction

ICLR 2025poster

Post-translational modifications (PTMs) profoundly expand the complexity and functionality of the proteome, regulating protein attributes and interactions that are crucial for biological processes. Accurately predicting PTM sites and their specific types is therefore essential for elucidating protei…

2025

MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization

CVPR 2025poster

Masked Image Modeling (MIM) with Vector Quantization (VQ) has achieved great success in both self-supervised pre-training and image generation. However, most existing methods struggle to address the trade-off in the shared latent space for generation quality vs. representation learning and efficienc…

2025

Multi-View 3D Point Tracking

ICCV 2025poster

We introduce the first data-driven multi-view 3D point tracker, designed to track arbitrary points in dynamic scenes using multiple camera views. Unlike existing monocular trackers, which struggle with depth ambiguities and occlusion, or prior multi-camera methods that require over 20 cameras and te…

2025

One2Any: One-Reference 6D Pose Estimation for Any Object

CVPR 2025poster

6D object pose estimation remains challenging for many applications due to dependencies on complete 3D models, multi-view images, or training limited to specific object categories. These requirements make generalization to novel objects difficult for which neither 3D models nor multi-view images may…

2025

Prior-guided Hierarchical Harmonization Network for Efficient Image Dehazing

AAAI 2025technical

Image dehazing is a crucial task that involves the enhancement of degraded images to recover their sharpness and textures. While vision Transformers have exhibited impressive results in diverse dehazing tasks, their quadratic complexity and lack of dehazing priors pose significant drawbacks for real…

Cited by 0SourcePDFScholar
2025

ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular Videos

NeurIPS 2025poster

Achieving truly practical dynamic 3D reconstruction requires online operation, global pose and map consistency, detailed appearance modeling, and the flexibility to handle both RGB and RGB-D inputs. However, existing SLAM methods typically merely remove the dynamic parts or require RGB-D input, whil…

Cited by 0SourceScholar
2025

Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning

ICCV 2025poster

Despite the promise of Multi-Task Learning (MTL) in leveraging complementary knowledge across tasks, existing multi-task optimization (MTO) techniques remain fixated on resolving conflicts through optimizer-centric loss scaling and gradient manipulation, yet fail to deliver consistent gains. In this…

Cited by 0SourcePDFScholar
2025

SOLAR: Serendipity Optimized Language Model Aligned for Recommendation

EMNLP 2025

Recently, Large Language Models (LLMs) have shown strong potential in recommendation tasks due to their broad world knowledge and reasoning capabilities. However, applying them to serendipity-oriented recommendation remains challenging, mainly due to a domain gap of LLMs in modeling personalized use

2025

Safe Planner: Empowering Safety Awareness in Large Pre-Trained Models for Robot Task Planning

AAAI 2025technical

Robot task planning is an important problem for autonomous robots in long-horizon challenging tasks. As large pre-trained models have demonstrated superior planning ability, recent research investigates utilizing large models to achieve autonomous planning for robots in diverse tasks. However, sinc…

Cited by 3SourcePDFScholar
2025

Samba: Synchronized Set-of-Sequences Modeling for Multiple Object Tracking

ICLR 2025spotlight

Multiple object tracking in complex scenarios - such as coordinated dance performances, team sports, or dynamic animal groups - presents unique challenges. In these settings, objects frequently move in coordinated patterns, occlude each other, and exhibit long-term dependencies in their trajectories…

Cited by 2SourcePDFScholar
2025

SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches

IJCAI 2025

Hand-drawn sketches are a natural and efficient medium for capturing and conveying ideas. Despite significant advancements in controllable natural image generation, translating freehand sketches into structured, machine-readable diagrams remains a labor-intensive and predominantly manual task. The p

Cited by 0SourcePDFScholar
2025

SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control Tasks

AAAI 2025technical

Deep reinforcement learning (DRL) has achieved remarkable success in various domains, yet its reliance on neural networks results in a lack of transparency, which limits its practical applications in safety-critical and human-agent interaction domains. Decision trees, known for their notable explain…

2025

Taming LLMs with Gradient Grouping

ACL 2025long

Training large language models (LLMs) poses challenges due to their massive scale and heterogeneous architectures. While adaptive optimizers like AdamW help address gradient variations, they still struggle with efficient and effective parameter-wise learning rate estimation, resulting in training in…

2025

Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial Recordings

ICLR 2025poster

Recent advancements in brain-computer interfaces (BCIs) and deep learning have made decoding lexical tones from intracranial recordings possible, providing the potential to restore the communication ability of speech-impaired tonal language speakers. However, data heterogeneity induced by both physi…

Cited by 0SourcePDFScholar
2025

UniK3D: Universal Camera Monocular 3D Estimation

CVPR 2025poster

Monocular 3D estimation is crucial for visual perception. However, current methods fall short by relying on oversimplified assumptions, such as pinhole camera models or rectified images. These limitations severely restrict their general applicability, causing poor performance in real-world scenarios…

2025

VN-GT: Optimizing Virtual Network Deployment via Game Theory

ICASSP 2025accepted

The static and homogeneous nature of traditional networks presents a significant challenge for our defense efforts. These characteristics enable an experienced attacker to quickly determine our network topology and gather detailed information about the internal hosts through systematic scanning tech…

Cited by 0SourceScholar
2025

Video-Bench: Human-Aligned Video Generation Benchmark

CVPR 2025poster

Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories: traditional benchmarks, which use metrics and embeddings to evaluate…

2024

Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-Consistency Training

ECCV 2024poster

"Multimodal reasoning is a challenging task that requires models to reason across multiple modalities to answer questions. Existing approaches have made progress by incorporating language and visual modalities into a two-stage reasoning framework, separating rationale generation from answer inferenc…

2024

FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled Audio

CVPR 2024poster

In this paper we abstract the process of people hearing speech extracting meaningful cues and creating various dynamically audio-consistent talking faces termed Listening and Imagining into the task of high-fidelity diverse talking faces generation from a single audio. Specifically it involves two c…

2024

IMM: An Imitative Reinforcement Learning Approach with Predictive Representation Learning for Automatic Market Making

IJCAI 2024poster

Market making (MM) via Reinforcement Learning (RL) has attracted significant attention in financial trading. Most existing RL-based MM methods focus on optimizing single-price level strategies which fail at frequent order cancellations and loss of queue priority. By comparison, strategies involving…

Cited by 2SourcePDFScholar
2024

Instructor-inspired Machine Learning for Robust Molecular Property Prediction

NeurIPS 2024poster

Machine learning catalyzes a revolution in chemical and biological science. However, its efficacy is heavily dependent on the availability of labeled data, and annotating biochemical data is extremely laborious. To surmount this data sparsity challenge, we present an instructive learning algorithm n…

Cited by 1SourcePDFScholar
2024

KW-Design: Pushing the Limit of Protein Design via Knowledge Refinement

ICLR 2024poster

Recent studies have shown competitive performance in protein inverse folding, while most of them disregard the importance of predictive confidence, fail to cover the vast protein space, and do not incorporate common protein knowledge. Given the great success of pretrained models on diverse protein-r…

2024

Learning to Predict Mutational Effects of Protein-Protein Interactions by Microenvironment-aware Hierarchical Prompt Learning

ICML 2024poster

Protein-protein bindings play a key role in a variety of fundamental biological processes, and thus predicting the effects of amino acid mutations on protein-protein binding is crucial. To tackle the scarcity of annotated mutation data, pre-training with massive unlabeled data has emerged as a promi…

Cited by 16SourcePDFScholar
2024

LongVQ: Long Sequence Modeling with Vector Quantization on Structured Memory

IJCAI 2024poster

Transformer models have been successful in various sequence processing tasks, but the self-attention mechanism's computational cost limits its practicality for long sequences. Although there are existing attention variants that improve computational efficiency, they have a limited ability to abstrac…

Cited by 2SourcePDFScholar
2024

MAPE-PPI: Towards Effective and Efficient Protein-Protein Interaction Prediction via Microenvironment-Aware Protein Embedding

ICLR 2024spotlight

Protein-Protein Interactions (PPIs) are fundamental in various biological processes and play a key role in life activities. The growing demand and cost of experimental PPI assays require computational methods for efficient PPI prediction. While existing methods rely heavily on protein sequence for P…

2024

Matching Anything by Segmenting Anything

CVPR 2024highlight

The robust association of the same objects across video frames in complex scenes is crucial for many applications especially object tracking. Current methods predominantly rely on labeled domain-specific video datasets which limits cross-domain generalization of learned similarity embeddings. We pro…

2024

MogaNet: Multi-order Gated Aggregation Network

ICLR 2024poster

By contextualizing the kernel as global as possible, Modern ConvNets have shown great potential in computer vision tasks. However, recent progress on \textit{multi-order game-theoretic interaction} within deep neural networks (DNNs) reveals the representation bottleneck of modern ConvNets, where the…

2024

PhyloGen: Language Model-Enhanced Phylogenetic Inference via Graph Structure Generation

NeurIPS 2024poster

Phylogenetic trees elucidate evolutionary relationships among species, but phylogenetic inference remains challenging due to the complexity of combining continuous (branch lengths) and discrete parameters (tree topology). Traditional Markov Chain Monte Carlo methods face slow convergence and co…

Cited by 2SourcePDFScholar
2024

Protein 3D Graph Structure Learning for Robust Structure-Based Protein Property Prediction

AAAI 2024technical

Protein structure-based property prediction has emerged as a promising approach for various biological tasks, such as protein function prediction and sub-cellular location estimation. The existing methods highly rely on experimental protein structure data and fail in scenarios where these data are u…

Cited by 12SourcePDFScholar
2024

RDesign: Hierarchical Data-efficient Representation Learning for Tertiary Structure-based RNA Design

ICLR 2024poster

While artificial intelligence has made remarkable strides in revealing the relationship between biological macromolecules' primary sequence and tertiary structure, designing RNA sequences based on specified tertiary structures remains challenging. Though existing approaches in protein design have th…

2024

Re-Dock: Towards Flexible and Realistic Molecular Docking with Diffusion Bridge

ICML 2024spotlight

Accurate prediction of protein-ligand binding structures, a task known as molecular docking is crucial for drug design but remains challenging. While deep learning has shown promise, existing methods often depend on holo-protein structures (docked, and not accessible in realistic tasks) or neglect p…

Cited by 10SourcePDFScholar
2024

Rethinking Memory and Communication Costs for Efficient Data Parallel Training of Large Language Models

NeurIPS 2024poster

Recently, various strategies for distributed training of large language models (LLMs) have been proposed. By categorizing them into basic strategies and composite strategies, we have discovered that existing basic strategies provide limited options in specific scenarios, leaving considerable room fo…

Cited by 0SourcePDFScholar
2024

Robust Visual Imitation Learning with Inverse Dynamics Representations

AAAI 2024technical

Imitation learning (IL) has achieved considerable success in solving complex sequential decision-making problems. However, current IL methods mainly assume that the environment for learning policies is the same as the environment for collecting expert datasets. Therefore, these methods may fail to w…

Cited by 2SourcePDFScholar
2024

SemiReward: A General Reward Model for Semi-supervised Learning

ICLR 2024poster

Semi-supervised learning (SSL) has witnessed great progress with various improvements in the self-training framework with pseudo labeling. The main challenge is how to distinguish high-quality pseudo labels against the confirmation bias. However, existing pseudo-label selection strategies are limite…

2024

Short-Long Convolutions Help Hardware-Efficient Linear Attention to Focus on Long Sequences

ICML 2024poster

To mitigate the computational complexity in the self-attention mechanism on long sequences, linear attention utilizes computation tricks to achieve linear complexity, while state space models (SSMs) popularize a favourable practice of using non-data-dependent memory pattern, *i.e.,* emphasize the ne…

Cited by 6SourcePDFScholar
2024

TopoFR: A Closer Look at Topology Alignment on Face Recognition

NeurIPS 2024poster

The field of face recognition (FR) has undergone significant advancements with the rise of deep learning. Recently, the success of unsupervised learning and graph neural networks has demonstrated the effectiveness of data structure information. Considering that the FR task can leverage large-scale…

2024

UniDepth: Universal Monocular Metric Depth Estimation

CVPR 2024highlight

Accurate monocular metric depth estimation (MMDE) is crucial to solving downstream tasks in 3D perception and modeling. However the remarkable accuracy of recent MMDE methods is confined to their training domains. These methods fail to generalize to unseen domains even in the presence of moderate do…

2024

UniIF: Unified Molecule Inverse Folding

NeurIPS 2024poster

Molecule inverse folding has been a long-standing challenge in chemistry and biology, with the potential to revolutionize drug discovery and material science. Despite specified models have been proposed for different small- or macro-molecules, few have attempted to unify the learning process, result…

Cited by 16SourcePDFScholar
2024

VQDNA: Unleashing the Power of Vector Quantization for Multi-Species Genomic Sequence Modeling

ICML 2024poster

Similar to natural language models, pre-trained genome language models are proposed to capture the underlying intricacies within genomes with unsupervised sequence modeling. They have become essential tools for researchers and practitioners in biology. However, the hand-crafted tokenization policies…

Cited by 9SourcePDFScholar
2024

Walker: Self-supervised Multiple Object Tracking by Walking on Temporal Object Appearance Graphs

ECCV 2024poster

"The supervision of state-of-the-art multiple object tracking (MOT) methods requires enormous annotation efforts to provide bounding boxes for all frames of all videos, and instance IDs to associate them through time. To this end, we introduce Walker, the first self-supervised tracker that learns fr…

2024

Wavelet-Driven Spatiotemporal Predictive Learning: Bridging Frequency and Time Variations

AAAI 2024technical

Spatiotemporal predictive learning is a paradigm that empowers models to learn spatial and temporal patterns by predicting future frames from past frames in an unsupervised manner. This method typically uses recurrent units to capture long-term dependencies, but these units often come with high comp…

2023

Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN

ICML 2023poster

Masked image modeling, an emerging self-supervised pre-training method, has shown impressive success across numerous downstream vision tasks with Vision transformers. Its underlying idea is simple: a portion of the input image is masked out and then reconstructed via a pre-text task. However, the wo…

Cited by 47SourcePDFScholar
2023

Behavior Contrastive Learning for Unsupervised Skill Discovery

ICML 2023poster

In reinforcement learning, unsupervised skill discovery aims to learn diverse skills without extrinsic rewards. Previous methods discover skills by maximizing the mutual information (MI) between states and skills. However, such an MI objective tends to learn simple and static skills and may hinder e…

2023

CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text Labels

AAAI 2023technical

Pre-trained vision-language models like CLIP have recently shown superior performances on various downstream tasks, including image classification and segmentation. However, in fine-grained image re-identification (ReID), the labels are indexes, lacking concrete text descriptions. Therefore, it rema…

2023

CVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition With Variational Alignment

CVPR 2023highlight

Sign language recognition (SLR) is a weakly supervised task that annotates sign videos as textual glosses. Recent studies show that insufficient training caused by the lack of large-scale available sign datasets becomes the main bottleneck for SLR. Most SLR works thereby adopt pretrained visual modu…

2023

Cascade-DETR: Delving into High-Quality Universal Object Detection

ICCV 2023poster

Object localization in general environments is a fundamental part of vision systems. While dominating on the COCO benchmark, recent Transformer-based detection methods are not competitive in diverse domains. Moreover, these methods still struggle to very accurately estimate the object bounding boxes…

Cited by 39PDFcodeScholar
2023

Flow to Control: Offline Reinforcement Learning with Lossless Primitive Discovery

AAAI 2023technical

Offline reinforcement learning (RL) enables the agent to effectively learn from logged data, which significantly extends the applicability of RL algorithms in real-world scenarios where exploration can be expensive or unsafe. Previous works have shown that extracting primitive skills from the recurr…

Cited by 18SourcePDFScholar
2023

Functional-Group-Based Diffusion for Pocket-Specific Molecule Generation and Elaboration

NeurIPS 2023poster

In recent years, AI-assisted drug design methods have been proposed to generate molecules given the pockets' structures of target proteins. Most of them are {\em atom-level-based} methods, which consider atoms as basic components and generate atom positions and types. In this way, however, it is ha…

Cited by 25SourcePDFScholar
2023

Harnessing Hard Mixed Samples with Decoupled Regularizer

NeurIPS 2023poster

Mixup is an efficient data augmentation approach that improves the generalization of neural networks by smoothing the decision boundary with mixed data. Recently, dynamic mixup methods have improved previous \textit{static} policies effectively (e.g., linear interpolation) by maximizing target-relat…

2023

Learning to Solve Tasks with Exploring Prior Behaviours

IROS 2023poster

Demonstrations are widely used in Deep Reinforcement Learning (DRL) for facilitating solving tasks with sparse rewards. However, the tasks in real-world scenarios can often have varied initial conditions from the demonstration, which would require additional prior behaviours. For example, consider w…

Cited by 3SourcecodeScholar
2023

Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules

ICLR 2023poster

Recent years have witnessed the prosperity of pre-training graph neural networks (GNNs) for molecules. Typically, atom types as node attributes are randomly masked, and GNNs are then trained to predict masked types as in AttrMask \citep{hu2020strategies}, following the Masked Language Modeling (MLM)…

2023

OVTrack: Open-Vocabulary Multiple Object Tracking

CVPR 2023poster

The ability to recognize, localize and track dynamic objects in a scene is fundamental to many real-world applications, such as self-driving and robotic systems. Yet, traditional multiple object tracking (MOT) benchmarks rely only on a few object categories that hardly represent the multitude of pos…

Cited by 66SourcePDFScholar
2023

OpenSTL: A Comprehensive Benchmark of Spatio-Temporal Predictive Learning

NeurIPS 2023poster

Spatio-temporal predictive learning is a learning paradigm that enables models to learn spatial and temporal patterns by predicting future frames from given past frames in an unsupervised manner. Despite remarkable progress in recent years, a lack of systematic understanding persists due to the dive…

2023

Rethinking Explaining Graph Neural Networks via Non-parametric Subgraph Matching

ICML 2023poster

The success of graph neural networks (GNNs) provokes the question about explainability: ``Which fraction of the input graph is the most determinant of the prediction?'' Particularly, parametric explainers prevail in existing approaches because of their more robust capability to decipher the black-bo…

2023

Temporal Attention Unit: Towards Efficient Spatiotemporal Predictive Learning

CVPR 2023poster

Spatiotemporal predictive learning aims to generate future frames by learning from historical frames. In this paper, we investigate existing methods and present a general framework of spatiotemporal predictive learning, in which the spatial encoder and decoder capture intra-frame features and the mi…

2023

Understanding the Limitations of Deep Models for Molecular property prediction: Insights and Solutions

NeurIPS 2023poster

Molecular Property Prediction (MPP) is a crucial task in the AI-driven Drug Discovery (AIDD) pipeline, which has recently gained considerable attention thanks to advancements in deep learning. However, recent research has revealed that deep models struggle to beat traditional non-deep ones on MPP. I…

Cited by 37SourcePDFScholar
2022

Active Hierarchical Exploration with Stable Subgoal Representation Learning

ICLR 2022poster

Goal-conditioned hierarchical reinforcement learning (GCHRL) provides a promising approach to solving long-horizon tasks. Recently, its success has been extended to more general settings by concurrently learning hierarchical policies and subgoal representations. Although GCHRL possesses superior exp…

2022

AutoMix: Unveiling the Power of Mixup for Stronger Classifiers

ECCV 2022poster

"Data mixing augmentation have proved to be effective for improving the generalization ability of deep neural networks. While early methods mix samples by hand-crafted policies (\textit{e.g.}, linear interpolation), recent methods utilize saliency information to match the mixed samples and labels vi…

2022

DLME: Deep Local-Flatness Manifold Embedding

ECCV 2022poster

"Manifold learning (ML) aims to seek low-dimensional embedding from high-dimensional data. The problem is challenging on real-world datasets, especially with under-sampling data, and we find that previous methods perform poorly in this case. Generally, ML methods first transform input data into a lo…

2022

Divide and Contrast: Source-free Domain Adaptation via Adaptive Contrastive Learning

NeurIPS 2022accept

We investigate a practical domain adaptation task, called source-free domain adaptation (SFUDA), where the source pretrained model is adapted to the target domain without access to the source data. Existing techniques mainly leverage self-supervised pseudo-labeling to achieve class-wise global align…

2022

Style Transformer for Image Inversion and Editing

CVPR 2022poster

Existing GAN inversion methods fail to provide codes for reliable reconstruction and flexible editing simultaneously. This paper presents a transformer-based image inversion and editing model for pretrained StyleGAN which is not only with less distortions, but also of high quality and flexibility fo…

Cited by 69PDFcodeScholar
2022

UMT: Unified Multi-Modal Transformers for Joint Video Moment Retrieval and Highlight Detection

CVPR 2022poster

Finding relevant moments and highlights in videos according to natural language queries is a natural and highly valuable common need in the current video content explosion era. Nevertheless, jointly conducting moment retrieval and highlight detection is an emerging research topic, even though its co…

Cited by 191PDFcodeScholar
2021

Offline Reinforcement Learning with Reverse Model-based Imagination

NeurIPS 2021poster

In offline reinforcement learning (offline RL), one of the main challenges is to deal with the distributional shift between the learning policy and the given dataset. To address this problem, recent offline RL methods attempt to introduce conservatism bias to encourage learning in high-confidence a…

Cited by 71SourcePDFScholar
2020

Deformation-Aware Unpaired Image Translation for Pose Estimation on Laboratory Animals

CVPR 2020poster

Our goal is to capture the pose of real animals using synthetic training examples, without using any manual supervision. Our focus is on neuroscience model organisms, to be able to study how neural circuits orchestrate behaviour. Human pose estimation attains remarkable accuracy when trained on real…

Cited by 51PDFScholar
2020

TLPG-Tracker: Joint Learning of Target Localization and Proposal Generation for Visual Tracking

IJCAI 2020poster

Target localization and proposal generation are two essential subtasks in generic visual tracking, and it is a challenge to address both the two efficiently. In this paper, we propose an efficient two-stage architecture which makes full use of the complementarity of two subtasks to achieve robust lo…

Cited by 0SourcePDFScholar
2019

Hierarchical Reinforcement Learning with Advantage-Based Auxiliary Rewards

NeurIPS 2019poster

Hierarchical Reinforcement Learning (HRL) is a promising approach to solving long-horizon problems with sparse and delayed rewards. Many existing HRL algorithms either use pre-trained low-level skills that are unadaptable, or require domain-specific information to define low-level rewards. In this p…

Cited by 102SourcePDFScholar
2019

Single Image Deraining: A Comprehensive Benchmark Analysis

CVPR 2019poster

We present a comprehensive study and evaluation of existing single image deraining algorithms, using a new large-scale benchmark consisting of both synthetic and real-world rainy images.This dataset highlights diverse data sources and image contents, and is divided into three subsets (rain streak, r…

Cited by 368PDFcodeScholar