← Search

Bin Li

195 accepted papers

2026

ACCFormer: Predicting Analog Circuit Performance Metrics via Topology-Aware Transformers

IJCAI 2026

Reusing and migrating analog circuit intellectual property (IP) across process nodes poses a significant challenge in modern chip design. Efficient and generalizable circuit performance prediction methods for analog circuits are crucial to achieving this goal. Current data-driven approaches typicall

Cited by 0Scholar
2026

Adaptive Morph-Patch Transformer for Aortic Vessel Segmentation

AAAI 2026technical

Accurate segmentation of aortic vascular structures is critical for diagnosing and treating cardiovascular diseases. Traditional Transformer-based models have shown promise in this domain by capturing long-range dependencies between vascular features. However, their reliance on fixed-size rectangul

Cited by 1SourcePDFScholar
2026

Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision-Language Models

ICML 2026poster

Large Vision-Language Models (LVLMs) typically process visual inputs as a prefix to the language decoder. As the model autoregressively generates text, this initial visual information inevitably undergoes ``dilution'', leading the model to over-rely on language priors and hallucinate objects. Existi…

Cited by 0SourceScholar
2026

Autoregressive End-To-End Planning with Time-Invariant Spatial Alignment and Multi-Objective Policy Refinement

ICRA 2026poster

The inherent sequential modeling capabilities of autoregressive models make them a formidable baseline for end-to-end planning in autonomous driving. Nevertheless, their performance is constrained by a spatio-temporal misalignment, as the planner must condition future actions on past sensory data. T…

2026

Autoregressive Meta-Actions for Unified Controllable Trajectory Generation in Autonomous Driving

RA-L 2026

Generating trajectories from high-level commands is critical for autonomous driving, but prevailing methods suffer from a flaw we term semantic misalignment. By associating long trajectories with a single, static meta-action (e.g., “lane change”), these methods corrupt training data during maneuver

Cited by 0SourcecodeScholar
2026

CACR: Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning

ICML 2026poster

The task of temporal answer grounding in instructional videos (TAGV), which aims to locate precise video segments that respond to natural language queries, is increasingly important for direct video answer retrieval. This task remains challenging due to the need to comprehend semantically complex qu…

Cited by 0SourceScholar
2026

CME-CAD: Heterogeneous Collaborative Multi-Expert Reinforcement Learning for CAD Code Generation

CVPR 2026

Computer-Aided Design (CAD) is essential in industrial design, but the complexity of traditional CAD modeling and workflows presents significant challenges for automating the generation of high-precision, editable CAD models. Existing methods, such as 3D reconstruction from sketches, often produce n

Cited by 0SourceScholar
2026

CoD: A Diffusion Foundation Model for Image Compression

CVPR 2026

Existing diffusion codecs typically build on text-to-image diffusion foundation models like Stable Diffusion.However, text conditioning is suboptimal from a compression perspective, hindering the potential of downstream diffusion codecs, particularly at ultra-low bitrates.To address it, we introduce

Cited by 0SourcecodeScholar
2026

Decomposition of Concept-Level Rules in Visual Scenes

ICLR 2026poster

Human cognition is compositional, and one can parse a visual scene into independent concepts and the corresponding concept-changing rules. By contrast, many vision-language systems process images holistically, with limited support for explicit decomposition. And previous methods of decomposing conce…

Cited by 0SourceScholar
2026

Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models

ICML 2026poster

Multimodal large language models (MLLMs) have been increasingly adopted in forensics for their robust semantic understanding. As AI-generated images become realistic, semantic-level inconsistencies alone are often insufficient for reliable detection. This motivates a critical question: *whether MLLM…

Cited by 0SourceScholar
2026

Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding

CVPR 2026

The application of Large Multimodal Models (LMMs) to long-form video understanding is constrained by limited context lengths and the computationally prohibitive cost of processing dense video tokens. Consequently, recent research has focused on query-aware frame selection, methods that often incur s

Cited by 0SourceScholar
2026

DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving

CVPR 2026

Dynamic scene reconstruction in autonomous driving remains a fundamental challenge due to significant temporal variations, moving objects, and complex scene dynamics. Existing feed-forward 3D models have demonstrated strong performance in static reconstruction but still struggle to capture dynamic m

Cited by 0SourceScholar
2026

E-MaT:Event-oriented Mamba for Egocentric Point Tracking

AAAI 2026technical

Egocentric point tracking aims to localize points on object surfaces from a first-person perspective and serves as a critical step toward embodied intelligence. Recent methods rely on video input, tracking query points through feature matching across consecutive frames. However, these methods strug

Cited by 0SourcePDFScholar
2026

Enabling Your Forensic Detector Know How Well It Performs on Distorted Samples

ICLR 2026poster

Generative AI has substantially facilitated realistic image synthesizing, posing great challenges for reliable forensics. When image forensic detectors are deployed in the wild, the inputs usually undergone various distortions including compression, rescaling, and lossy transmission. Such distortion…

Cited by 0SourceScholar
2026

Envision, Attend, Then Respond: Counterfactual Hallucination Mitigation in Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) often hallucinate when visual evidence conflicts with world knowledge, i.e., in counterfactual scenarios. We propose Envision-Attend-Respond (EnAR), a training-free framework that leverages visual priors to steer the model's attention toward counterfactual elemen

Cited by 0SourcecodeScholar
2026

From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code Generation

AAAI 2026technical

Computer-Aided Design (CAD) plays a vital role in engineering and manufacturing, yet current CAD workflows require extensive domain expertise and manual modeling effort. Recent advances in large language models (LLMs) have made it possible to generate code from natural language, opening new opportun

Cited by 0SourcePDFScholar
2026

Generative Video Compression with One-Dimensional Latent Representation

CVPR 2026

Recent advancements in generative video codec (GVC) typically encode video into a 2D latent grid and employ high-capacity generative decoders for reconstruction. However, this paradigm still leaves two key challenges in fully exploiting spatial-temporal redundancy: Spatially, the 2D latent grid inev

Cited by 0SourceScholar
2026

InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs

CVPR 2026

Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overlooking the physically plausible interplay essential for multi-agent interactions. To bridge this gap, we propose InterAg

Cited by 0SourcecodeScholar
2026

KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls

ICML 2026poster

Large Multimodal Models encode extensive factual knowledge in their pre-trained weights. However, its knowledge remains static and limited, unable to keep pace with real-world developments, which hinders continuous knowledge acquisition. Effective knowledge injection thus becomes critical, involving…

Cited by 0SourceScholar
2026

KiGRAS: Kinematic-Driven Generative Model for Realistic Agent Simulation

ICRA 2026poster

Trajectory generation is a pivotal task in autonomous driving. Recent studies have introduced the autoregressive paradigm, leveraging the state transition model to approximate future trajectory distributions. This paradigm closely mirrors the real-world trajectory generation process and has achieved…

2026

MEDFACT-R1: TOWARDS FACTUAL MEDICAL REASONING VIA PSEUDO-LABEL AUGMENTATION

ICASSP 2026poster

Ensuring factual consistency and reliable reasoning remains a critical challenge for medical vision-language models. We introduce MEDFACT-R1, a two-stage framework that integrates external knowledge grounding with reinforcement learning to improve the factual medical reasoning. The first stage uses…

Cited by 0SourcePDFScholar
2026

OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding

AAAI 2026technical

LVLMs have been shown to perform excellently in image-level tasks such as VQA and caption. However, in many instance-level tasks, such as visual grounding and object detection, LVLMs still show performance gaps compared to previous expert models. Meanwhile, although pedestrian tracking is a classica

Cited by 0SourcePDFScholar
2026

Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation

CVPR 2026

Multimodal autoregressive (AR) models, based on next-token prediction and transformer architecture, have demonstrated remarkable capabilities in various multimodal tasks including text-to-image (T2I) generation. Despite their strong performance in general T2I tasks, our research reveals that these m

Cited by 0SourceScholar
2026

REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment

ICML 2026poster

Knowledge-intensive Visual Question Answering (KI-VQA) frequently suffers from severe knowledge conflicts caused by the inherent limitations of open-domain retrieval. However, existing paradigms face critical limitations, including the lack of generalizable conflict detection and intra-model constra…

Cited by 0SourceScholar
2026

Real-Time and Lightweight Diffusion Image Compression

ICML 2026poster

Recent advanced diffusion methods typically derive strong generative priors by scaling diffusion transformers. However, scaling fails to generalize when adapted for real-time compression scenarios that demand lightweight models. In this paper, we explore the design of real-time and lightweight diffu…

Cited by 0SourceScholar
2026

Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images

ICLR 2026poster

The rapid advancement of AI-generated content (AIGC) has enabled the synthesis of visually convincing images; however, many such outputs exhibit subtle \textbf{semantic anomalies}, including unrealistic object configurations, violations of physical laws, or commonsense inconsistencies, which comprom…

Cited by 0SourceScholar
2026

Single-step Diffusion-based Video Coding with Semantic-Temporal Guidance

CVPR 2026

While traditional and neural video codecs (NVCs) have achieved remarkable rate-distortion performance, improving perceptual quality at low bitrates remains challenging. Some NVCs incorporate perceptual or adversarial objectives but still suffer from artifacts due to limited generation capacity, wher

Cited by 0SourcecodeScholar
2026

Tackling Heavy-Tailed Q-Value Bias in Offline-to-Online Reinforcement Learning with Laplace-Robust Modeling

ICLR 2026poster

Offline-to-online reinforcement learning (O2O RL) aims to improve the performance of offline pretrained agents through online fine-tuning. Existing O2O RL methods have achieved advances in mitigating the overestimation of Q-value biases (i.e., biases of cumulative rewards), improving the performance…

Cited by 0SourceScholar
2026

Unbiased Gradient Estimation for Event Binning via Functional Backpropagation

ICLR 2026poster

Event-based vision encodes dynamic scenes as asynchronous spatio-temporal spikes called events. To leverage conventional image processing pipelines, events are typically binned into frames. However, binning functions are discontinuous, which truncates gradients at the frame level and forces most eve…

Cited by 0SourcecodeScholar
2026

VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network

CVPR 2026

The rapid advances in deep learning have significantly enhanced the accuracy of multimodal 3D human pose estimation (HPE). However, the state-of-the-art (SOTA) HPE pipelines still rely on Transformers, whose quadratic complexity makes real-time processing for long sequences impractical. Mamba addres

Cited by 0SourcecodeScholar
2026

Vision in One Vector: Implicit Visual Compression with Diffusion Foundation Models

ICML 2026poster

Modern visual generative models acquire rich visual knowledge through large-scale training, yet existing visual representations (such as pixels, latents, or tokens) remain external to the model and cannot directly exploit this knowledge for compact storage or reuse. In this work, we introduce a new …

Cited by 0SourceScholar
2026

Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge Intelligence

ICLR 2026poster

Large Vision Transformers (ViTs) must often be compressed before they can be deployed on resource-constrained edge devices. However, many edge devices require only part of the *all-classes* knowledge of a pre-trained ViT in their corresponding application scenarios. This is overlooked by existing c…

Cited by 0SourcecodeScholar
2026

When Large Multimodal Models Confront Evolving Knowledge: Challenges and Explorations

ICLR 2026poster

Large Multimodal Models (LMMs) store vast amounts of pretrained knowledge but struggle to remain aligned with real-world updates, making it difficult to avoid capability degradation when acquiring evolving knowledge. Furthermore, most current work focuses on exploring static textual knowledge inject…

Cited by 0SourceScholar
2026

When MLLMs Meets Compression Distortion: A Coding Paradigm Tailored to MLLMs

ICLR 2026poster

The increasing deployment of powerful Multimodal Large Language Models (MLLMs), typically hosted on cloud platforms, urgently requires effective compression techniques to efficiently transmit signal inputs (e.g., images, videos) from edge devices with minimal bandwidth usage. However, conventional i…

Cited by 0SourcecodeScholar
2026

Why Attention Patterns Exist: A Unifying Temporal Perspective Analysis

ICLR 2026poster

Attention patterns play a crucial role in both training and inference of large language models (LLMs). Prior works have identified individual patterns—such as retrieval heads, sink heads, and diagonal traces—but these observations remain fragmented and lack a unifying explanation. To bridge this gap…

Cited by 0SourcecodeScholar
2025

5%>100%: Breaking Performance Shackles of Full Fine-Tuning on Visual Recognition Tasks

CVPR 2025poster

Pre-training & fine-tuning can enhance the transferring efficiency and performance in visual tasks. Recent delta-tuning methods provide more options for visual classification tasks. Despite their success, existing visual delta-tuning art fails to exceed the upper limit of full fine-tuning on challen…

2025

ALARM: Safe Reinforcement Learning With Reliable Mimicry for Robust Legged Locomotion

RA-L 2025

Legged robots are supposed to traverse complicated environments, which makes it challenging to design a model-based controller due to their functional complexity. Currently, using deep reinforcement learning to improve the adaptability of robots in complex scenarios has been a major research trend.

Cited by 3SourceScholar
2025

Accurate and Scalable Graph Neural Networks via Message Invariance

ICLR 2025poster

Message passing-based graph neural networks (GNNs) have achieved great success in many real-world applications. For a sampled mini-batch of target nodes, the message passing process is divided into two parts: message passing between nodes within the batch (MP-IB) and message passing from nodes outsi…

2025

AttentionPredictor: Temporal Patterns Matter for KV Cache Compression

NeurIPS 2025poster

With the development of large language models (LLMs), efficient inference through Key-Value (KV) cache compression has attracted considerable attention, especially for long-context generation. To compress the KV cache, recent methods identify critical KV tokens through static modeling of attention s…

Cited by 0SourcecodeScholar
2025

BeSimulator: A Large Language Model Powered Text-based Behavior Simulator

EMNLP 2025

Traditional robot simulators focus on physical process modeling and realistic rendering, often suffering from high computational costs, inefficiencies, and limited adaptability. To handle this issue, we concentrate on behavior simulation in robotics to analyze and validate the logic behind robot beh

2025

Benchmarking End-To-End Performance of AI-Based Chip Placement Algorithms

NeurIPS 2025poster

Chip placement is a critical step in the Electronic Design Automation (EDA) workflow, which aims to arrange chip modules on the canvas to optimize the performance, power, and area (PPA) metrics of final designs. Recent advances show great potential of AI-based algorithms in chip placement. However,…

Cited by 0SourceScholar
2025

Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning

ICML 2025poster

Abstract visual reasoning (AVR) enables humans to quickly discover and generalize abstract rules to new scenarios. Designing intelligent systems with human-like AVR abilities has been a long-standing topic in the artificial intelligence community. Deep AVR solvers have recently achieved remarkable s…

2025

Breaking Latent Prior Bias in Detectors for Generalizable AIGC Image Detection

NeurIPS 2025poster

Current AIGC detectors often achieve near-perfect accuracy on images produced by the same generator used for training but struggle to generalize to outputs from unseen generators. We trace this failure in part to latent prior bias: detectors learn shortcuts tied to patterns stemming from the initial…

Cited by 0SourceScholar
2025

CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning

NeurIPS 2025poster

Computer-Aided Design (CAD) is pivotal in industrial manufacturing, with orthographic projection reasoning foundational to its entire workflow—encompassing design, manufacturing, and simulation. However, prevailing deep-learning approaches employ standard 3D reconstruction pipelines as an alternativ…

Cited by 0SourcecodeScholar
2025

ChatReID: Open-ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language Models

ICCV 2025poster

Person re-identification (Re-ID) is a crucial task in computer vision, aiming to recognize individuals across non-overlapping camera views. While recent advanced vision-language models (VLMs) excel in logical reasoning and multi-task generalization, their applications in Re-ID tasks remain limited.…

Cited by 0SourcePDFScholar
2025

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models

ICASSP 2025accepted

The impressive performance of Large Language Model (LLM) has prompted researchers to develop Multi-modal LLM (MLLM), which has shown great potential for various multi-modal tasks. However, current MLLM often struggles to effectively address fine-grained multi-modal challenges. We argue that this lim…

Cited by 0SourceScholar
2025

Code-BT: A Code-Driven Approach to Behavior Tree Generation for Robot Tasks Planning with Large Language Models

IJCAI 2025

Behavior trees(BTs) provide a systematic and structured control architecture extensively employed in game AI and robotic behavior control, owing to their modularity, reactivity, and reusability. Nonetheless, manual BTs design requires significant expertise and becomes inefficient as task complexity

Cited by 0SourcePDFScholar
2025

Content and Salient Semantics Collaboration for Cloth-Changing Person Re-Identification

ICASSP 2025accepted

Cloth-changing person re-identification aims at recognizing the same person with clothing changes across non-overlapping cameras. Advanced methods either resort to identity-related auxiliary modalities (e.g., sketches, silhouettes, and keypoints) or clothing labels to mitigate the impact of clothes.…

Cited by 0SourceScholar
2025

DLF: Extreme Image Compression with Dual-generative Latent Fusion

ICCV 2025poster

Recent studies in extreme image compression have achieved remarkable performance by compressing the tokens from generative tokenizers. However, these methods often prioritize clustering common semantics within the dataset, while overlooking the diverse details of individual objects. Consequently, th…

2025

Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding

NeurIPS 2025poster

Long-form video understanding presents significant challenges due to extensive temporal-spatial complexity and the difficulty of question answering under such extended contexts. While Large Language Models (LLMs) have demonstrated considerable advancements in video analysis capabilities and long co…

Cited by 0SourceScholar
2025

Exploiting Robust Model Watermarking Against the Model Fine-Tuning Attack via Flat Minima Aware Optimizers

ICASSP 2025accepted

With the rapid advancement of deep neural networks (DNNs), model watermarking has emerged as a widely adopted technique for safeguarding model copyrights. A prevalent method involves utilizing a watermark decoder to retrieve watermark bits from generated outputs, but such methods are often vulnerabl…

Cited by 0SourceScholar
2025

Foundation Model Driven Appearance Extraction for Robust Multiple Object Tracking

AAAI 2025technical

Multiple Object Tracking (MOT) is a fundamental task in computer vision. Existing methods utilize motion information or appearance information to perform object tracking. However, these algorithms still struggle with special circumstances, such as occlusion and blurring in complex scenes. Inspired b…

2025

Global Tropical Cyclone Intensity Forecasting with Multi-modal Multi-scale Causal Autoregressive Model

ICASSP 2025accepted

Accurate forecasting of tropical cyclone (TC) intensity is crucial for formulating disaster risk reduction strategies. Current methods predominantly rely on limited spatiotemporal information from ERA5 data and neglect the causal relationships between these physical variables, failing to fully captu…

Cited by 0SourceScholar
2025

Guard Me If You Know Me: Protecting Specific Face-Identity from Deepfakes

NeurIPS 2025poster

Securing personal identity against deepfake attacks is increasingly critical in the digital age, especially for celebrities and political figures whose faces are easily accessible and frequently targeted. Most existing deepfake detection methods focus on general-purpose scenarios and often ignore th…

Cited by 0SourcecodeScholar
2025

KiGRAS: Kinematic-Driven Generative Model for Realistic Agent Simulation

RA-L 2025

Trajectory generation is a pivotal task in autonomous driving. Recent studies have introduced the autoregressive paradigm, leveraging the state transition model to approximate future trajectory distributions. This paradigm closely mirrors the real-world trajectory generation process and has achieved

Cited by 23SourceScholar
2025

Learning Robust Representations with Long-Term Information for Generalization in Visual Reinforcement Learning

ICLR 2025poster

Generalization in visual reinforcement learning (VRL) aims to learn agents that can adapt to test environments with unseen visual distractions. Despite advances in robust representations learning, many methods do not take into account the essential downstream task of sequential decision-making. This…

Cited by 0SourcePDFScholar
2025

LoRA-PAR: A Flexible Dual-System LoRA Partitioning Approach to Efficient LLM Fine-Tuning

EMNLP 2025

Large-scale generative models like DeepSeek-R1 and OpenAI-O1 benefit substantially from chain-of-thought (CoT) reasoning, yet pushing their performance typically requires vast data, large model sizes, and full-parameter fine-tuning. While parameter-efficient fine-tuning (PEFT) helps reduce cost, mos

2025

LogicTree: Improving Complex Reasoning of LLMs via Instantiated Multi-step Synthetic Logical Data

NeurIPS 2025spotlight

Despite their remarkable performance on various tasks, Large Language Models (LLMs) still struggle with logical reasoning, particularly in complex and multi-step reasoning processes. Among various efforts to enhance LLMs' reasoning capabilities, synthesizing large-scale, high-quality logical reason…

Cited by 0SourceScholar
2025

MATE: Motion-Augmented Temporal Consistency for Event-based Point Tracking

ICCV 2025poster

Tracking Any Point (TAP) plays a crucial role in motion analysis. Video-based approaches rely on iterative local matching for tracking, but they assume linear motion during the blind time between frames, which leads to point loss under large displacements or nonlinear motion. The high temporal resol…

Cited by 0SourcePDFScholar
2025

Multi-modal Multi-platform Person Re-Identification: Benchmark and Method

ICCV 2025poster

Conventional person re-identification (ReID) research is often limited to single-modality sensor data from static cameras, which fails to address the complexities of real-world scenarios where multi-modal signals are increasingly prevalent. For instance, consider an urban ReID system integrating sta…

Cited by 0SourcePDFScholar
2025

One-Shot Heterogeneous Federated Learning with Local Model-Guided Diffusion Models

ICML 2025poster

In recent years, One-shot Federated Learning (OSFL) methods based on Diffusion Models (DMs) have garnered increasing attention due to their remarkable performance. However, most of these methods require the deployment of foundation models on client devices, which significantly raises the computation…

Cited by 0SourcePDFScholar
2025

One-Step Diffusion-Based Image Compression with Semantic Distillation

NeurIPS 2025poster

While recent diffusion-based generative image codecs have shown impressive performance, their iterative sampling process introduces unpleasant latency. In this work, we revisit the design of a diffusion-based codec and argue that multi-step sampling is not necessary for generative compression. Based…

Cited by 0SourcecodeScholar
2025

PICD: Versatile Perceptual Image Compression with Diffusion Rendering

CVPR 2025poster

Recently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perc…

Cited by 0SourcePDFScholar
2025

Points, Images and Texts: Boosting Point Cloud Completion with Multi-Modal Features

ICRA 2025

Point cloud completion is crucial for reconstructing accurate shapes in many 3D visual applications. Recent approaches incorporate images into the completion pipeline, introducing geometric clues and global constraints. However, their fusion processes often fail to reconstruct detailed parts and mai

Cited by 0SourceScholar
2025

Query-efficient Attack for Black-box Image Inpainting Forensics via Reinforcement Learning

AAAI 2025technical

Recently, image inpainting has become a common tool for manipulating nature images in a malicious manner, which has led to the rapid advancement of inpainting forensics. Although current forensics methods have shown precise location of inpainting regions and reliable robustness against image post-pr…

Cited by 0SourcePDFScholar
2025

SPAC: Sparse Partitioning and Adaptive Core Tensor Pruning Model for Knowledge Graph Completion

AAAI 2025technical

Tensor decomposition (TD) models are promising solutions for knowledge graph completion due to their simple structures but powerful representation capacities. The TD models typically adopt Tucker decomposition with a structured core tensor. Some models with a sparse core tensor, such as DistMult and…

Cited by 0SourcePDFScholar
2025

Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection

AAAI 2025technical

The proliferation of deepfake faces poses huge potential negative impacts on our daily lives. Despite substantial advancements in deepfake detection over these years, the generalizability of existing methods against forgeries from unseen datasets or created by emerging generative models remains cons…

Cited by 2SourcePDFScholar
2025

Strategyproofness and Monotone Allocation of Auction in Social Networks

IJCAI 2025

Strategyproofness in network auctions requires that bidders not only report their valuations truthfully, but also do their best to invite neighbours from the social network. In contrast to canonical auctions, where the value-monotone allocation in Myerson's Lemma is a cornerstone, a general principl

Cited by 0SourcePDFScholar
2025

Towards Practical Real-Time Neural Video Compression

CVPR 2025poster

We introduce a practical real-time neural video codec (NVC) designed to deliver high compression ratio, low latency and broad versatility. In practice, the coding speed of NVCs depends on 1) computational costs, and 2) non-computational operational costs, such as memory I/O and the number of functio…

2025

VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language Models

NeurIPS 2025poster

Faces synthesized by diffusion models (DMs) with high-quality and controllable attributes pose a significant challenge for Deepfake detection. Most state-of-the-art detectors only yield a binary decision, incapable of forgery localization, attribution of forgery methods, and providing analysis on th…

Cited by 0SourceScholar
2025

VQLTI: Long-Term Tropical Cyclone Intensity Forecasting with Physical Constraints

AAAI 2025technical

Tropical cyclone (TC) intensity forecasting is crucial for early disaster warning and emergency decision-making. Numerous researchers have explored deep-learning methods to address computational and post-processing issues in operational forecasting. Regrettably, they exhibit subpar long-term forecas…

2024

"Idling Neurons, Appropriately Lenient Workload During Fine-tuning Leads to Better Generalization"

ECCV 2024poster

"Pre-training on large-scale datasets has become a fundamental method for training deep neural networks. Pre-training provides a better set of parameters than random initialization, which reduces the training cost of deep neural networks on the target task. In addition, pre-training also provides a…

Cited by 0SourcePDFScholar
2024

A Keyless Extraction Framework Targeting at Deep Learning Based Image-Within-Image Models

ICASSP 2024accepted

Image-within-image technique aims to establish covert communication by concealing a secret image within a cover image. Compared with traditional steganography algorithms, the security of image-within-image technique has not been rigorously evaluated by steganalysis. Existing attack methods just brut…

Cited by 0SourceScholar
2024

Accelerating Data Generation for Neural Operators via Krylov Subspace Recycling

ICLR 2024spotlight

Learning neural operators for solving partial differential equations (PDEs) has attracted great attention due to its high inference efficiency. However, training such operators requires generating a substantial amount of labeled data, i.e., PDE problems together with their solutions. The data genera…

2024

CMA: A Chromaticity Map Adapter for Robust Detection of Screen-Recapture Document Images

CVPR 2024poster

The rebroadcasting of screen-recaptured document images introduces a significant risk to the confidential documents processed in government departments and commercial companies. However detecting recaptured document images subjected to distortions from online social networks (OSNs) is challenging si…

Cited by 2SourcePDFScholar
2024

Coarse-to-Fine Highlighting: Reducing Knowledge Hallucination in Large Language Models

ICML 2024poster

Generation of plausible but incorrect factual information, often termed hallucination, has attracted significant research interest. Retrieval-augmented language model (RALM)---which enhances models with up-to-date knowledge---emerges as a promising method to reduce hallucination. However, existing R…

Cited by 10SourcePDFScholar
2024

DiffForensics: Leveraging Diffusion Prior to Image Forgery Detection and Localization

CVPR 2024poster

As manipulating images may lead to misinterpretation of the visual content addressing the image forgery detection and localization (IFDL) problem has drawn serious public concerns. In this work we propose a simple assumption that the effective forensic method should focus on the mesoscopic propertie…

Cited by 20SourcePDFScholar
2024

Discriminative Forests Improve Generative Diversity for Generative Adversarial Networks

AAAI 2024technical

Improving the diversity of Artificial Intelligence Generated Content (AIGC) is one of the fundamental problems in the theory of generative models such as generative adversarial networks (GANs). Previous studies have demonstrated that the discriminator in GANs should have high capacity and robustness…

2024

Efficient Backdoor Attacks for Deep Neural Networks in Real-world Scenarios

ICLR 2024poster

Recent deep neural networks (DNNs) have came to rely on vast amounts of training data, providing an opportunity for malicious attackers to exploit and contaminate the data to carry out backdoor attacks. However, existing backdoor attack methods make unrealistic assumptions, assuming that all trainin…

2024

Enhanced Fine-Grained Motion Diffusion for Text-Driven Human Motion Synthesis

AAAI 2024technical

The emergence of text-driven motion synthesis technique provides animators with great potential to create efficiently. However, in most cases, textual expressions only contain general and qualitative motion descriptions, while lack fine depiction and sufficient intensity, leading to the synthesized…

Cited by 6SourcePDFScholar
2024

Exploring One-Shot Semi-supervised Federated Learning with Pre-trained Diffusion Models

AAAI 2024technical

Recently, semi-supervised federated learning (semi-FL) has been proposed to handle the commonly seen real-world scenarios with labeled data on the server and unlabeled data on the clients. However, existing methods face several challenges such as communication costs, data heterogeneity, and training…

Cited by 25SourcePDFScholar
2024

Fast Adaptation for Human Pose Estimation via Meta-Optimization

CVPR 2024poster

Domain shift is a challenge for supervised human pose estimation where the source data and target data come from different distributions. This is why pose estimation methods generally perform worse on the test set than on the training set. Recently test-time adaptation has proven to be an effective…

Cited by 8SourcePDFScholar
2024

Federated Adaptive Prompt Tuning for Multi-Domain Collaborative Learning

AAAI 2024technical

Federated learning (FL) enables multiple clients to collaboratively train a global model without disclosing their data. Previous researches often require training the complete model parameters. However, the emergence of powerful pre-trained models makes it possible to achieve higher performance with…

2024

Few-Shot Semantic Dependency Parsing via Graph Contrastive Learning

COLING 2024main

Graph neural networks (GNNs) have achieved promising performance on semantic dependency parsing (SDP), owing to their powerful graph representation learning ability. However, training a high-performing GNN-based model requires a large amount of labeled data and it is prone to over-fitting in the abs…

2024

FreqBlender: Enhancing DeepFake Detection by Blending Frequency Knowledge

NeurIPS 2024poster

Generating synthetic fake faces, known as pseudo-fake faces, is an effective way to improve the generalization of DeepFake detection. Existing methods typically generate these faces by blending real or fake faces in spatial domain. While these methods have shown promise, they overlook the simulation…

Cited by 8SourcePDFScholar
2024

GMM-Based Heuristic Decision Framework for Safe Automated Laparoscope Control

RA-L 2024

Automated laparoscope field of view (FoV) control in minimal invasive surgery (MIS) poses challenges, as existing solutions failed to address dynamic surgical FoV requirements across different phases and they neglected the misorientation effect or potential obstacles during the control process which

Cited by 11SourceScholar
2024

Generative Latent Coding for Ultra-Low Bitrate Image Compression

CVPR 2024poster

Most existing image compression approaches perform transform coding in the pixel space to reduce its spatial redundancy. However they encounter difficulties in achieving both high-realism and high-fidelity at low bitrate as the pixel-space distortion may not align with human perception. To address t…

Cited by 13SourcePDFScholar
2024

Human Motion Forecasting in Dynamic Domain Shifts: A Homeostatic Continual Test-time Adaptation Framework

ECCV 2024poster

"Existing motion forecasting models, while making progress, struggle to bridge the gap between the source and target domains. Recent solutions often rely on an unrealistic assumption that the target domain remains stationary. Due to the ever-changing environment, however, the real-world test distrib…

Cited by 1SourcePDFScholar
2024

HySense: Hybrid Event Occurrence Detection Method for IoT Devices

ICASSP 2024accepted

Integrating IoT devices into physical environments through automation can result in device state inconsistency (DSI) anomalies. These anomalies can be caused by malicious attacks or device malfunctions, making it crucial to detect them for the security and reliability of IoT systems. Physical finger…

Cited by 0SourceScholar
2024

Improving Viewpoint-Independent Object-Centric Representations through Active Viewpoint Selection

NeurIPS 2024poster

Given the complexities inherent in visual scenes, such as object occlusion, a comprehensive understanding often requires observation from multiple viewpoints. Existing multi-viewpoint object-centric learning methods typically employ random or sequential viewpoint selection strategies. While applicab…

Cited by 0SourcePDFScholar
2024

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

CVPR 2024poster

The exponential growth of large language models (LLMs) has opened up numerous possibilities for multi-modal AGI systems. However the progress in vision and vision-language foundation models which are also critical elements of multi-modal AGI has not kept pace with LLMs. In this work we design a larg…

2024

Long-term Temporal Context Gathering for Neural Video Compression

ECCV 2024poster

"Most existing neural video codecs (NVCs) only extract short-term temporal context by optical flow-based motion compensation. However, such short-term temporal context suffers from error propagation and lacks awareness of long-term relevant information. This limits their performance, particularly in…

2024

MILP-StuDio: MILP Instance Generation via Block Structure Decomposition

NeurIPS 2024poster

Mixed-integer linear programming (MILP) is one of the most popular mathematical formulations with numerous applications. In practice, improving the performance of MILP solvers often requires a large amount of high-quality data, which can be challenging to collect. Researchers thus turn to generation…

Cited by 11SourcePDFScholar
2024

Mastering Symbolic Operations: Augmenting Language Models with Compiled Neural Networks

ICLR 2024poster

Language models' (LMs) proficiency in handling deterministic symbolic reasoning and rule-based tasks remains limited due to their dependency implicit learning on textual data. To endow LMs with genuine rule comprehension abilities, we propose "Neural Comprehension" - a framework that synergistically…

2024

MoML: Online Meta Adaptation for 3D Human Motion Prediction

CVPR 2024poster

In the academic field the research on human motion prediction tasks mainly focuses on exploiting the observed information to forecast human movements accurately in the near future horizon. However a significant gap appears when it comes to the application field as current models are all trained offl…

Cited by 2SourcePDFScholar
2024

NeRM: Learning Neural Representations for High-Framerate Human Motion Synthesis

ICLR 2024poster

Generating realistic human motions with high framerate is an underexplored task, due to the varied framerates of training data, huge memory burden brought by high framerates and slow sampling speed of generative models. Recent advances make a compromise for training by downsampling high-framerate de…

Cited by 6SourcePDFScholar
2024

Rethinking Branching on Exact Combinatorial Optimization Solver: The First Deep Symbolic Discovery Framework

ICLR 2024poster

Machine learning (ML) has been shown to successfully accelerate solving NP-hard combinatorial optimization (CO) problems under the branch and bound framework. However, the high training and inference cost and limited interpretability of ML approaches severely limit their wide application to modern…

Cited by 11SourcePDFScholar
2024

Simultaneous Estimation of Shape and Force along Highly Deformable Surgical Manipulators Using Sparse FBG Measurement

ICRA 2024poster

Recently, fiber optic sensors such as fiber Bragg gratings (FBGs) have been widely investigated for shape reconstruction and force estimation of flexible surgical robots. However, most existing approaches need precise model parameters of FBGs inside the fiber and their alignments with the flexible r…

Cited by 2SourceScholar
2024

Towards Generative Abstract Reasoning: Completing Raven’s Progressive Matrix via Rule Abstraction and Selection

ICLR 2024poster

Endowing machines with abstract reasoning ability has been a long-term research topic in artificial intelligence. Raven's Progressive Matrix (RPM) is widely used to probe abstract visual reasoning in machine intelligence, where models will analyze the underlying rules and select one image from candi…

2024

Towards Next-Generation Logic Synthesis: A Scalable Neural Circuit Generation Framework

NeurIPS 2024poster

Logic Synthesis (LS) aims to generate an optimized logic circuit satisfying a given functionality, which generally consists of circuit translation and optimization. It is a challenging and fundamental combinatorial optimization problem in integrated circuit design. Traditional LS approaches rely on…

Cited by 5SourcePDFScholar
2024

Trajectory-prediction-based Dynamic Tracking of a UGV to a Moving Target under Multi-disturbed Conditions

ICRA 2024poster

Tracking dynamic targets poses a significant challenge for Unmanned Ground Vehicles (UGVs). Existing methods often lack research on multi-disturbed conditions. To address this issue, we propose a trajectory-prediction-based dynamic tracking scheme, which includes target localization, trajectory pred…

Cited by 0SourceScholar
2024

Uncertainty-based Offline Variational Bayesian Reinforcement Learning for Robustness under Diverse Data Corruptions

NeurIPS 2024poster

Real-world offline datasets are often subject to data corruptions (such as noise or adversarial attacks) due to sensor failures or malicious attacks. Despite advances in robust offline reinforcement learning (RL), existing methods struggle to learn robust agents under high uncertainty caused by the…

2024

Unsupervised Cross-Domain Image Retrieval via Prototypical Optimal Transport

AAAI 2024technical

Unsupervised cross-domain image retrieval (UCIR) aims to retrieve images sharing the same category across diverse domains without relying on labeled data. Prior approaches have typically decomposed the UCIR problem into two distinct tasks: intra-domain representation learning and cross-domain featur…

2024

Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning

NeurIPS 2024poster

Recently, vision model pre-training has evolved from relying on manually annotated datasets to leveraging large-scale, web-crawled image-text data. Despite these advances, there is no pre-training method that effectively exploits the interleaved image-text data, which is very prevalent on the Intern…

2023

Autonomous Intelligent Navigation for Flexible Endoscopy Using Monocular Depth Guidance and 3-D Shape Planning

ICRA 2023poster

Recent advancements toward perception and decision-making of flexible endoscopes have shown great potential in computer-aided surgical interventions. However, owing to modeling uncertainty and inter-patient anatomical variation in flexible endoscopy, the challenge remains for efficient and safe navi…

Cited by 12SourceScholar
2023

Chinese Text Recognition with A Pre-Trained CLIP-Like Model Through Image-IDS Aligning

ICCV 2023oral

Scene text recognition has been studied for decades due to its broad applications. However, despite Chinese characters possessing different characteristics from Latin characters, such as complex inner structures and large categories, few methods have been proposed for Chinese Text Recognition (CTR).…

Cited by 30PDFcodeScholar
2023

DeFeeNet: Consecutive 3D Human Motion Prediction With Deviation Feedback

CVPR 2023poster

Let us rethink the real-world scenarios that require human motion prediction techniques, such as human-robot collaboration. Current works simplify the task of predicting human motions into a one-off process of forecasting a short future sequence (usually no longer than 1 second) based on a historica…

Cited by 14SourcePDFScholar
2023

Demonstration-Guided Reinforcement Learning with Efficient Exploration for Task Automation of Surgical Robot

ICRA 2023poster

Task automation of surgical robot has the potentials to improve surgical efficiency. Recent reinforcement learning (RL) based approaches provide scalable solutions to surgical automation, but typically require extensive data collection to solve a task if no prior knowledge is given. This issue is kn…

Cited by 29SourcecodeScholar
2023

Domain Re-Modulation for Few-Shot Generative Domain Adaptation

NeurIPS 2023poster

In this study, we delve into the task of few-shot Generative Domain Adaptation (GDA), which involves transferring a pre-trained generator from one domain to a new domain using only a few reference images. Inspired by the way human brains acquire knowledge in new domains, we present an innovative gen…

2023

End-to-End Learning of Deep Visuomotor Policy for Needle Picking

IROS 2023poster

Needle picking is a challenging manipulation task in robot-assisted surgery due to the characteristics of small slender shapes of needles, needles' variations in shapes and sizes, and demands for millimeter-level control. Prior works, heavily relying on the prior of needles (e.g., geometric models),…

Cited by 6SourceScholar
2023

Enhancing Robustness and Imperceptibility of Blind Watermarking with Improved Message Processor

ICASSP 2023accepted

The current state-of-the-art(SOTA) blind watermark embedding method MBRS based on deep learning is less robust to Crop, and additional diffusion layers need to be added for optimization. However, the diffusion layer will make the model less robust to noise other than Crop. Therefore, MBRS which need…

Cited by 0SourceScholar
2023

Human Joint Kinematics Diffusion-Refinement for Stochastic Motion Prediction

AAAI 2023technical

Stochastic human motion prediction aims to forecast multiple plausible future motions given a single pose sequence from the past. Most previous works focus on designing elaborate losses to improve the accuracy, while the diversity is typically characterized by randomly sampling a set of latent varia…

2023

Integrating Syntactic and Semantic Knowledge in AMR Parsing with Heterogeneous Graph Attention Network

ICASSP 2023accepted

Abstract Meaning Representation (AMR) parsing is the task of translating a sentence to an AMR semantic graph which captures the basic meaning of the sentence, and is empowered by pre-trained Transformer models recently. These models encode the syntactic and semantic knowledge implicitly through self…

Cited by 0SourceScholar
2023

Large Language Models are Better Reasoners with Self-Verification

EMNLP 2023long findings

Recently, with the chain of thought (CoT) prompting, large language models (LLMs), e.g., GPT-3, have shown strong reasoning ability in several natural language processing tasks such as arithmetic, commonsense, and logical reasoning. However, LLMs with CoT require multi-step prompting and multi-token…

Cited by 0SourcecodeScholar
2023

Learning Accurate 3D Shape Based on Stereo Polarimetric Imaging

CVPR 2023poster

Shape from Polarization (SfP) aims to recover surface normal using the polarization cues of light. The accuracy of existing SfP methods is affected by two main problems. First, the ambiguity of polarization cues partially results in false normal estimation. Second, the widely-used assumption about o…

Cited by 12SourcePDFScholar
2023

Learning to Locate the Text Forgery in Smartphone Screenshots

ICASSP 2023accepted

In this paper, we present the Screenshot Text Forgery Dataset (STFD), which is the first public dataset for the smartphone screenshot text forgery localization task. To address such a task, we propose a novel Screenshot Text Forgery Localization Network (STFL-Net). Specifically, we introduce the OCR…

Cited by 0SourceScholar
2023

Meta-Auxiliary Learning for Adaptive Human Pose Prediction

AAAI 2023technical

Predicting high-fidelity future human poses, from a historically observed sequence, is crucial for intelligent robots to interact with humans. Deep end-to-end learning approaches, which typically train a generic pre-trained model on external datasets and then directly apply it to all test samples, e…

Cited by 5SourcePDFScholar
2023

Orientation-Independent Chinese Text Recognition in Scene Images

IJCAI 2023poster

Scene text recognition (STR) has attracted much attention due to its broad applications. The previous works pay more attention to dealing with the recognition of Latin text images with complex backgrounds by introducing language models or other auxiliary networks. Different from Latin texts, many ve…

2023

Siamese Image Modeling for Self-Supervised Vision Representation Learning

CVPR 2023poster

Self-supervised learning (SSL) has delivered superior performance on a variety of downstream vision tasks. Two main-stream SSL frameworks have been proposed, i.e., Instance Discrimination (ID) and Masked Image Modeling (MIM). ID pulls together representations from different views of the same image,…

2023

Smoothing Point Adjustment-Based Evaluation of Time Series Anomaly Detection

ICASSP 2023accepted

Anomalies in time series appear consecutively, forming anomaly segments. Applying the classical point-based evaluation metrics to evaluate the detection performance of segments leads to considerable underestimation, so most related studies resort to point adjustment. This operation treats all points…

Cited by 0SourceScholar
2023

TODE-Trans: Transparent Object Depth Estimation with Transformer

ICRA 2023poster

Transparent objects are widely used in industrial automation and daily life. However, robust visual recognition and perception of transparent objects have always been a major challenge. Currently, most commercial-grade depth cameras are still not good at sensing the surfaces of transparent objects d…

Cited by 24SourcecodeScholar
2023

Test-time Personalizable Forecasting of 3D Human Poses

ICCV 2023poster

Current motion forecasting approaches typically train a deep end-to-end model from the source domain data, and then apply it directly to target subjects. Despite promising results, they remain non-optimal, due to privacy considerations, the test person and his/her natural properties (e.g., stature,…

Cited by 7PDFScholar
2023

Time-Conditioned Generative Modeling of Object-Centric Representations for Video Decomposition and Prediction

UAI 2023poster

When perceiving the world from multiple viewpoints, humans have the ability to reason about the complete objects in a compositional manner even when an object is completely occluded from certain viewpoints. Meanwhile, humans are able to imagine novel views after observing multiple viewpoints. Recent…

2023

Towards Accurate Video Text Spotting with Text-wise Semantic Reasoning

IJCAI 2023poster

Video text spotting (VTS) aims at extracting texts from videos, where text detection, tracking and recognition are conducted simultaneously. There have been some works that can tackle VTS; however, they may ignore the underlying semantic relationships among texts within a frame. We observe that the…

2023

Towards All-in-One Pre-Training via Maximizing Multi-Modal Mutual Information

CVPR 2023poster

To effectively exploit the potential of large-scale models, various pre-training strategies supported by massive data from different sources are proposed, including supervised pre-training, weakly-supervised pre-training, and self-supervised pre-training. It has been proved that combining multiple p…

2023

Training-free Diffusion Model Adaptation for Variable-Sized Text-to-Image Synthesis

NeurIPS 2023poster

Diffusion models (DMs) have recently gained attention with state-of-the-art performance in text-to-image synthesis. Abiding by the tradition in deep learning, DMs are trained and evaluated on the images with fixed sizes. However, users are demanding for various images with specific sizes and various…

Cited by 31SourcePDFScholar
2022

3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic Surgery

ICRA 2022poster

Automatic laparoscope motion control is fundamentally important for surgeons to efficiently perform operations. However, its traditional control methods based on tool tracking without considering information hidden in surgical scenes are not intelligent enough, while the latest supervised imitation…

Cited by 15SourceScholar
2022

Actor-Critic Policy Optimization in a Large-Scale Imperfect-Information Game

ICLR 2022poster

The deep policy gradient method has demonstrated promising results in many large-scale games, where the agent learns purely from its own experience. Yet, policy gradient methods with self-play suffer convergence problems to a Nash Equilibrium (NE) in multi-agent situations. Counterfactual regret min…

Cited by 33SourcePDFScholar
2022

Cross-Modal Knowledge Distillation for Depth Privileged Monocular Visual Odometry

RA-L 2022

Most self-supervised monocular visual odometry (VO) suffer from the scale ambiguity problem. A promising way to address this problem is to introduce additional information for training. In this work, we propose a new depth privileged framework to learn a monocular VO. It assumes that sparse depth is

Cited by 9SourceScholar
2022

Distilled Visual and Robot Kinematics Embeddings for Metric Depth Estimation in Monocular Scene Reconstruction

IROS 2022poster

Estimating precise metric depth and scene reconstruction from monocular endoscopy is a fundamental task for surgical navigation in robotic surgery. However, traditional stereo matching adopts binocular images to perceive the depth information, which is difficult to transfer to the soft robotics-base…

Cited by 11SourceScholar
2022

DynGL-SDP: Dynamic Graph Learning for Semantic Dependency Parsing

COLING 2022main

A recent success in semantic dependency parsing shows that graph neural networks can make significant accuracy improvements, owing to its powerful ability in learning expressive graph representations. However, this work learns graph representations based on a static graph constructed by an existing…

2022

Equivalence Analysis between Counterfactual Regret Minimization and Online Mirror Descent

ICML 2022spotlight

Follow-the-Regularized-Leader (FTRL) and Online Mirror Descent (OMD) are regret minimization algorithms for Online Convex Optimization (OCO), they are mathematically elegant but less practical in solving Extensive-Form Games (EFGs). Counterfactual Regret Minimization (CFR) is a technique for approxi…

2022

FakeCLR: Exploring Contrastive Learning for Solving Latent Discontinuity in Data-Efficient GANs

ECCV 2022poster

"Data-Efficient GANs (DE-GANs), which aim to learn generative models with a limited amount of training data, encounter several challenges for generating high-quality samples. Since data augmentation strategies have largely alleviated the training instability, how to further improve the generative pe…

2022

Learning General Gaussian Mixture Model with Integral Cosine Similarity

IJCAI 2022poster

Gaussian mixture model (GMM) is a powerful statistical tool in data modeling, especially for unsupervised learning tasks. Traditional learning methods for GMM such as expectation maximization (EM) require the covariance of the Gaussian components to be non-singular, a condition that is often not sat…

2022

Learning Laparoscope Actions via Video Features for Proactive Robotic Field-of-View Control

RA-L 2022

Smart laparoscope motion control for adjusting surgical field-of-view is an increasingly hot topic in robot-assisted surgery. Previous off-the-shelf methods have been conducted in reactive ways which heavily rely on human input signals, e.g., gaze or voice, thus cannot avoid cognitive burdens to sur

Cited by 18SourceScholar
2022

Learning Robust Policy against Disturbance in Transition Dynamics via State-Conservative Policy Optimization

AAAI 2022technical

Deep reinforcement learning algorithms can perform poorly in real-world tasks due to the discrepancy between source and target environments. This discrepancy is commonly viewed as the disturbance in transition dynamics. Many existing algorithms learn robust policies by modeling the disturbance and a…

Cited by 24SourcePDFScholar
2022

Multi-Dimensional Proprioception and Stiffness Tuning for Soft Robotic Joints

ICRA 2022poster

Proprioception and variable stiffness are two trending topics in soft robotics research. The former could endow soft robots with the ability to perceive the environment as well as their internal states without the need of dedicated sensors, while the latter could strengthen the otherwise excessive c…

Cited by 5SourceScholar
2022

Overlooked Poses Actually Make Sense: Distilling Privileged Knowledge for Human Motion Prediction

ECCV 2022poster

"Previous works on human motion prediction follow the pattern of building a mapping relation between the sequence observed and the one to be predicted. However, due to the inherent complexity of multivariate time series data, it still remains a challenge to find the extrapolation relation between mo…

Cited by 8SourcePDFScholar
2022

Roadblocks for Temporarily Disabling Shortcuts and Learning New Knowledge

NeurIPS 2022accept

Deep learning models have been found with a tendency of relying on shortcuts, i.e., decision rules that perform well on standard benchmarks but fail when transferred to more challenging testing conditions. Such reliance may hinder deep learning models from learning other task-related features and se…

Cited by 7SourcePDFScholar
2022

Sample-Efficient Reinforcement Learning via Conservative Model-Based Actor-Critic

AAAI 2022technical

Model-based reinforcement learning algorithms, which aim to learn a model of the environment to make decisions, are more sample efficient than their model-free counterparts. The sample efficiency of model-based approaches relies on whether the model can well approximate the environment. However, lea…

Cited by 45SourcePDFScholar
2022

Text Gestalt: Stroke-Aware Scene Text Image Super-resolution

AAAI 2022technical

In the last decade, the blossom of deep learning has witnessed the rapid development of scene text recognition. However, the recognition of low-resolution scene text images remains a challenge. Even though some super-resolution methods have been proposed to tackle this problem, they usually treat te…

2022

Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified Viewpoints

AAAI 2022technical

Visual scenes are extremely rich in diversity, not only because there are infinite combinations of objects and background, but also because the observations of the same scene may vary greatly with the change of viewpoints. When observing a visual scene that contains multiple objects from multiple vi…

2021

An Efficient Pessimistic-Optimistic Algorithm for Stochastic Linear Bandits with General Constraints

NeurIPS 2021poster

This paper considers stochastic linear bandits with general nonlinear constraints. The objective is to maximize the expected cumulative reward over horizon $T$ subject to a set of constraints in each round $\tau\leq T$. We propose a pessimistic-optimistic algorithm for this problem, which is efficie…

Cited by 53SourcePDFScholar
2021

Continuous-time edge modelling using non-parametric point processes

NeurIPS 2021poster

The mutually-exciting Hawkes process (ME-HP) is a natural choice to model reciprocity, which is an important attribute of continuous-time edge (dyadic) data. However, existing ways of implementing the ME-HP for such data are either inflexible, as the exogenous (background) rate functions are typical…

Cited by 8SourcePDFScholar
2021

Data-driven Holistic Framework for Automated Laparoscope Optimal View Control with Learning-based Depth Perception

ICRA 2021poster

Laparoscopic Field of View (FOV) control is one of the most fundamental and important components in Minimally Invasive Surgery (MIS), nevertheless the traditional manual holding paradigm may easily bring fatigue to surgical assistants, and misunderstanding between surgeons also hinders assistants to…

Cited by 28SourceScholar
2021

Deformable DETR: Deformable Transformers for End-to-End Object Detection

ICLR 2021oral

DETR has been recently proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance. However, it suffers from slow convergence and limited feature spatial resolution, due to the limitation of Transformer attention modules in processing ima…

2021

Dual-Stream Multiple Instance Learning Network for Whole Slide Image Classification With Self-Supervised Contrastive Learning

CVPR 2021poster

We address the challenging problem of whole slide image (WSI) classification. WSIs have very high resolutions and usually lack localized annotations. WSI classification can be cast as a multiple instance learning (MIL) problem when only slide-level labels are available. We propose a MIL-based method…

Cited by 1004PDFcodeScholar
2021

Fden: Mining Effective Information of Features in Detecting Network Anomalies

ICASSP 2021accepted

Network anomaly detection is important for detecting and reacting to the presence of network attacks. In this paper, we propose a novel method to effectively leverage the features in detecting network anomalies, named FDEn, consisting of flow-based Feature Derivation (FD) and prior knowledge incorpo…

Cited by 0SourceScholar
2021

Image Steganography Based on Iterative Adversarial Perturbations Onto a Synchronized-Directions Sub-Image

ICASSP 2021accepted

Nowadays a steganography has to face challenges to both feature-based staganalysis and convolutional neural network (CNN) based steganalysis. In this paper, we present a novel steganographic scheme to incorporate synchronizing modification directions and iterative adversarial perturbations to enhanc…

Cited by 0SourceScholar
2021

PENet: Towards Precise and Efficient Image Guided Depth Completion

ICRA 2021poster

Image guided depth completion is the task of generating a dense depth map from a sparse depth map and a high quality image. In this task, how to fuse the color and depth modalities plays an important role in achieving good performance. This paper proposes a two-branch backbone that consists of a col…

Cited by 369SourcecodeScholar
2021

Poisson-Randomised DirBN: Large Mutation is Needed in Dirichlet Belief Networks

ICML 2021spotlight

The Dirichlet Belief Network (DirBN) was recently proposed as a promising deep generative model to learn interpretable deep latent distributions for objects. However, its current representation capability is limited since its latent distributions across different layers is prone to form similar patt…

2021

Robust Three-Dimensional Shape Sensing for Flexible Endoscopic Surgery Using Multi-Core FBG Sensors

RA-L 2021

In this letter, we propose a novel 3D shape sensing algorithm for flexible endoscopic surgery using multi-core fiber Bragg grating (FBG) sensors. Considering the signal noises and environmental perturbations, the direct use of FBG measurements for shape sensing and position estimation is regarded as

Cited by 55SourceScholar
2021

SurRoL: An Open-source Reinforcement Learning Centered and dVRK Compatible Platform for Surgical Robot Learning

IROS 2021poster

Autonomous surgical execution relieves tedious routines and surgeon’s fatigue. Recent learning-based methods, especially reinforcement learning (RL) based methods, achieve promising performance for dexterous manipulation, which usually requires the simulation to collect data efficiently and reduce t…

Cited by 95SourcecodeScholar
2021

Zero-Shot Chinese Character Recognition with Stroke-Level Decomposition

IJCAI 2021poster

Chinese character recognition has attracted much research interest due to its wide applications. Although it has been studied for many years, some issues in this field have not been completely resolved yet, \textit{e.g.} the zero-shot problem. Previous character-based and radical-based methods have…

2020

Distributed Verification of Belief Precisions Convergence in Gaussian Belief Propagation

ICASSP 2020accepted

Gaussian belief propagation (BP) finds extensive applications in signal processing but it is not guaranteed to converge in loopy graphs. In order to determine whether Gaussian BP would converge, one could directly use the classical convergence conditions of Gaussian BP, such as diagonal dominance, w…

Cited by 0SourceScholar
2020

Recurrent Dirichlet Belief Networks for interpretable Dynamic Relational Data Modelling

IJCAI 2020poster

The Dirichlet Belief Network~(DirBN) has been recently proposed as a promising approach in learning interpretable deep latent representations for objects. In this work, we leverage its interpretable modelling architecture and propose a deep dynamic probabilistic framework -- the Recurrent Dirichle…

Cited by 0SourcePDFScholar
2020

VL-BERT: Pre-training of Generic Visual-Linguistic Representations

ICLR 2020poster

We introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short). VL-BERT adopts the simple yet powerful Transformer model as the backbone, and extends it to take both visual and linguistic embedded features as input. In it, each…

Cited by 2015SourcecodeScholar
2019

A New Spatial Steganographic Scheme by Modeling Image Residuals with Multivariate Gaussian Model

ICASSP 2019accepted

Embedding costs used in content-adaptive image steganographic schemes can be defined in a heuristic way or with a statistical model. Inspired by previous steganographic methods, i.e., MG (multivariate Gaussian model) and MiPOD (minimizing the power of optimal detector), we propose a model-driven sch…

Cited by 0SourceScholar
2019

Adaptive Neural Admittance Control for Collision Avoidance in Human-Robot Collaborative Tasks

IROS 2019poster

This paper proposed an adaptive neural admittance control strategy for collision avoidance in human-robot collaborative tasks. In order to ensure that the robot end-effector can avoid collisions with surroundings, robot should be operated compliantly by human within a constrained task space. An impe…

Cited by 7SourceScholar
2019

Generative Modeling of Infinite Occluded Objects for Compositional Scene Representation

ICML 2019oral

We present a deep generative model which explicitly models object occlusions for compositional scene representation. Latent representations of objects are disentangled into location, size, shape, and appearance, and the visual scene can be generated compositionally by integrating these representatio…

Cited by 23SourcePDFScholar
2019

Scalable Deep Generative Relational Model with High-Order Node Dependence

NeurIPS 2019poster

In this work, we propose a probabilistic framework for relational data modelling and latent structure exploring. Given the possible feature information for the nodes in a network, our model builds up a deep architecture that can approximate to the possible nonlinear mappings between the nodes' featu…

2018

Affinity Derivation and Graph Merge for Instance Segmentation

ECCV 2018poster

We present an instance segmentation scheme based on pixel affinity information, which is the relationship of two pixels belonging to a same instance. In our scheme, we use two neural networks with similar structure. One is to predict pixel level semantic score and the other is designed to derive pix…

2015

A low complexity optimization algorithm for zero-forcing precoding under per-antenna power constraints

ICASSP 2015accepted

Zero-forcing beamforming (ZFBF) is a popular pre-coding scheme for MIMO systems. Most of the studies in the literature are under total power constraints. However, the per-antenna power constraints (PAPC) are more realistic. The state-of-the-art method is interior point method which is expensive to r…

Cited by 0SourceScholar
2015

Low-latency list decoding of polar codes with double thresholding

ICASSP 2015accepted

For polar codes with short-to-medium code length, list successive cancellation decoding is used to achieve a good error-correcting performance. However, list pruning in the current list decoding is based on the sorting strategy and its timing complexity is high. This results in a long decoding laten…

Cited by 0SourceScholar