← Search

Zheng WANG

203 accepted papers

2026

Any2Any: Unified Arbitrary Modality Translation for Remote Sensing

ICML 2026poster

Multi-modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross-modal translation methods treat each modality pair as an independent task, resulting in quadratic complexity and limited ge…

Cited by 0SourceScholar
2026

BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving

ICLR 2026poster

Diffusion-based planners have shown great promise for autonomous driving due to their ability to capture multi-modal driving behaviors. However, guiding these models effectively in reactive, closed-loop environments remains a significant challenge. Simple conditioning often fails to provide sufficie…

Cited by 0SourcecodeScholar
2026

Characteristic Root Analysis and Regularization for Linear Time Series Forecasting

ICLR 2026poster

Time series forecasting remains a critical challenge across numerous domains, yet the effectiveness of complex models often varies unpredictably across datasets. Recent studies highlight the surprising competitiveness of simple linear models, suggesting that their robustness and interpretability wa…

Cited by 0SourcecodeScholar
2026

Cross-Tactile Sensor Representation Learning

ICML 2026poster

Visuo-tactile sensors have been widely adopted in robotic manipulation. However, inherent heterogeneity in sensor designs hinders the learning of unified tactile representations in cross-sensor scenarios. Existing methods that focus on reconstruction or task-specific supervision often fail to captur…

Cited by 0SourceScholar
2026

D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies

AAAI 2026technical

Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial General Intelligence. While most existing datasets and benchmarks for training and evaluating GUI agents are static and id

Cited by 5SourcePDFScholar
2026

DrugTrail: Explainable Drug Discovery via Structured Reasoning and Druggability‑Tailored Preference Optimization

ICLR 2026poster

Machine learning promises to revolutionize drug discovery, but its "black-box" nature and narrow focus limit adoption by experts. While Large Language Models (LLMs) offer a path forward with their broad knowledge and interactivity, existing methods remain data-intensive and lack transparent reasonin…

Cited by 0SourceScholar
2026

Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics

ICRA 2026poster

近年来,视觉-语言-行动(VLA)模型通过无缝整合视觉感知、语言理解和动作生成,在端到端的学习框架中彻底革新了机器人作。然而,由于这些模型设计为直接与物理世界和人类交互,其安全性至关重要,即使是小漏洞也可能导致灾难性故障。在本研究中,我们提出了通用对抗对象,这是一种表面纹理优化的球体,当置于机器人视野内时,任务成功率会显著降低。具体来说,我们的方法引入了一个多层次攻击框架,能够共同干扰轨迹规划、任务执行和动作控制。我们在模拟和现实机器人环境中验证了我们的方法。实验结果表明,对抗对象在两种代表性VLA模型(Pi0和RDT&#

Cited by 0Scholar
2026

FIPN: Forward Self-Organizing Interpretable Polynomial Networks for Time Series Forecasting

ICML 2026poster

Most existing time series forecasting models are trained with backpropagation, which often brings high computational cost and limited transparency, so it can be hard to understand why a model makes a given prediction. This paper presents FIPN, a forward self-organizing interpretable polynomial netwo…

Cited by 0SourceScholar
2026

FW-VTON: FLATTENING-AND-WARPING FOR PERSON-TO-PERSON VIRTUAL TRY-ON

ICASSP 2026poster

Traditional virtual try-on methods primarily focus on the garment-to-person try-on task, which requires flat garment representations. In contrast, this paper introduces a novel approach to the person-to-person try-on task. Unlike the garment-to-person try-on task, the person-to-person task only invo…

Cited by 0SourcePDFScholar
2026

FedARA: Resource-adaptive Low-rank Personalized Federated Learning via Anchor-driven Representation Alignment on Heterogeneous Edge Devices

CVPR 2026

Personalized Federated Learning (PFL) has gained significant attention for enabling participating clients to train customized personalized models on non-IID local data. However, current PFL methods mainly suffer from two limitations: 1) Only the personalized part supports heterogeneous design, while

Cited by 0SourceScholar
2026

FedRAC: Rolling Submodel Allocation for Collaborative Fairness in Federated Learning

CVPR 2026

Collaborative fairness in federated learning ensures that clients are rewarded according to their contributions, thereby fostering long-term participation among clients. However, existing methods often under-reward low-contributing clients in the early training stage and neglect critical issues (con

Cited by 0SourcecodeScholar
2026

From Collapse to Control: Understanding and Extending Context Length in Emerging Hybrid Models via Universal Position Interpolation

ICLR 2026poster

Hybrid Mamba-Transformer models have emerged as promising alternatives to pure Transformers, offering efficiency and competitive performance. However, they struggle to generalize beyond their training context windows, collapsing on long-context tasks. We provide the first systematic analysis of this…

Cited by 0SourcecodeScholar
2026

Heuristic Self-Paced Learning for Domain Adaptive Semantic Segmentation under Adverse Conditions

CVPR 2026

The learning order of semantic classes significantly impacts unsupervised domain adaptation for semantic segmentation, especially under adverse weather conditions. Most existing curricula rely on handcrafted heuristics (e.g., fixed uncertainty metrics) and follow a static schedule, which fails to ad

Cited by 0SourceScholar
2026

Hyper-Opinion Vagueness Quantification for Robust Multimodal Learning

AAAI 2026technical

Robust Multimodal Learning (RML) aims to address the issues of unreliable predictions of multimodal models. Nevertheless, previous RML works often struggle to distinguish between different categories that rely on identical intra-modal cues, making ambiguous predictions. We defined this degree of ``u

Cited by 0SourcePDFScholar
2026

InfoDet: A Dataset for Infographic Element Detection

ICLR 2026poster

Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of existing VLMs lies in their inaccurate visual grounding of infographic elements,…

Cited by 0SourcecodeScholar
2026

Language-Grounded Decoupled Action Representation for Robotic Manipulation

CVPR 2026

The heterogeneity between high-level vision-language understanding and low-level action control remains a fundamental challenge in robotic manipulation. Although recent methods have advanced task-specific action alignment, they often struggle to generate robust and accurate actions for novel or sema

Cited by 0SourceScholar
2026

Non-Parametric Probabilistic Robustness: A Conservative Risk Estimator under Unknown Perturbation Distributions

ICML 2026poster

Deep learning (DL) models, despite their remarkable success, remain vulnerable to small input perturbations that can cause erroneous outputs, motivating probabilistic robustness (PR) as a complementary notion to adversarial robustness (AR) for stochastic reliability assessment. However, existing PR …

Cited by 0SourceScholar
2026

ORBIT: A Prognostic World Model for Ocular Reasoning Based on Imagined Trajectories

ICML 2026poster

The longitudinal management of blinding fundus diseases constitutes a Partially Observable Markov Decision Process (POMDP) necessitating a critical precision-risk trade-off between intervention and over-treatment, as true pathology is often obscured in static observations. However, existing paradigm…

Cited by 0SourceScholar
2026

PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image–Text Retrieval

AAAI 2026technical

Remote sensing (RS) image–text retrieval faces significant challenges in real-world datasets due to the presence of Pseudo-Matched Pairs (PMPs), semantically mismatched or weakly aligned image–text pairs, which hinder the learning of reliable cross-modal alignments. To address this issue, we propose

Cited by 0SourcePDFScholar
2026

Position: The Systemic Lack of Agency in Visual Reasoning

ICML 2026poster

This paper argues that a systemic lack of Agency constrains the implicit reasoning capabilities of current Vision-Language Models (VLMs). Implicit reasoning refers to the ability to autonomously discover and utilize hidden visual evidence to bridge information gaps, rather than merely relying on exp…

Cited by 0SourceScholar
2026

Probing RLVR Training Instability through the Lens of Objective-Level Hacking

ICML 2026poster

Prolonged reinforcement learning with verifiable rewards (RLVR) has been shown to drive continuous improvements in the reasoning capabilities of large language models, but the training is often prone to instabilities, especially in Mixture-of-Experts (MoE) architectures. Training instability severel…

Cited by 0SourceScholar
2026

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference

ICML 2026poster

Mixture-of-Experts (MoE) have shown strong potential in scaling language models efficiently by activating only a small subset of experts per input. However, their deployment remains limited due to the high memory overhead associated with storing all expert parameters, particularly as the number of e…

Cited by 0SourceScholar
2026

QuantWear: Quantum-scale Wear Particle Detection for Jet Engine Diagnosis

ICML 2026poster

The quantity and 3-D shape of wear particles are essential indicators for assessing the health of jet engines, enabling early detection of potential damage and preventing accidents caused by catastrophic failures. However, capturing wear particles is difficult due to their minute sizes and ultra hig…

Cited by 0SourceScholar
2026

Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition

CVPR 2026

High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory and copyright constraints. This scarcity hampers model development--ironically, in settings where generative models are most needed to compensate for

Cited by 0SourceScholar
2026

Reliable-View 2D-3D Key-Part Aligned Transformer with Reinforced Masking for 3D Point Cloud Understanding

AAAI 2026technical

Self-supervised 3D point cloud understanding is crucial for scene understanding, where Masked Autoencoders (MAE) have achieved excellent performance in point cloud representation learning. However, existing MAE-style methods fail to consider spatial-semantic variations in masking strategies, and joi

Cited by 0SourcePDFScholar
2026

S2-Boost: Synergistic Semantic Boosting for Coarse-to-Fine Ensemble Learning

AAAI 2026technical

Neuroscientific evidence reveals that human visual recognition is not an instantaneous event but a hierarchical process, where the brain constructs a holistic perception by progressively integrating simple features like edges or texture into complex scenes. Ensemble learning successfully utilizes th

Cited by 0SourcePDFScholar
2026

ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management

ICML 2026poster

LLM-based multi-agent simulations are increasingly adopted across application domains, but remain difficult to scale due to GPU memory pressure. Each agent maintains private GPU-resident states, including models, prefix caches, and adapters, which quickly exhaust device memory as the agent count gro…

Cited by 0SourceScholar
2026

SimROD: A Simple Baseline for Raw Object Detection with Global and Local Enhancements

AAAI 2026technical

Most visual models are designed for sRGB images, yet RAW data offers significant advantages for object detection by preserving sensor information before ISP processing. This enables improved detection accuracy and more efficient hardware designs by bypassing the ISP. However, RAW object detection is

Cited by 0SourcePDFScholar
2026

Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning

ICLR 2026poster

Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs). Yet on-policy algorithms such as Group Relative Policy Optimization (GRPO) often suffer in early training: noisy gradients from low-quality rollouts lead to unstable updates and inefficient explora…

Cited by 0SourcecodeScholar
2026

StyleDistillation: A New Insight of Image Style Enables Personalized Aesthetic Manipulation

ICML 2026poster

Text-guided stylized image generation has yielded promising advances by leveraging the powerful capabilities of text-to-image diffusion models. However, the inherent coupling of style and content information within the reference image presents a significant challenge. To address this, we propose Sty…

Cited by 0SourceScholar
2026

The Labyrinth and the Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models

ICML 2026poster

Sequential editing of structured knowledge in large language models allows targeted factual updates without retraining, yet existing methods often rely on complex regularization or constraint mechanisms whose necessity remains unclear. In this work, we systematically investigate the mechanisms under…

Cited by 0SourceScholar
2026

Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding

CVPR 2026

Long video understanding is challenging due to dense visual redundancy, long-range temporal dependencies, and the tendency of chain-of-thought and retrieval-based agents to accumulate semantic drift and correlation-driven errors. We argue that long-video reasoning should begin not with reactive retr

Cited by 0SourcecodeScholar
2026

Towards Federated Clustering: A Client-wise Private Graph Aggregation Framework

AAAI 2026technical

Federated clustering addresses the critical challenge of extracting patterns from decentralized, unlabeled data. However, it is hampered by the flaw that current approaches are forced to accept a compromise between performance and privacy: transmitting embedding representations risks sensitive data

Cited by 0SourcePDFScholar
2025

A New Unsupervised Infrared and Visible Image Fusion Method Based on Salient Object Segmentation under Poor Illumination

IROS 2025

This work aims to propose a new unsupervised infrared and visible image fusion method based on salient object segmentation, which can obtain a fused image with more information on salient object and realize the salient object segmentation under poor illumination.The new method can be divided into fo

Cited by 0SourceScholar
2025

A Survey of LLM-based Agents in Medicine: How far are we from Baymax?

ACL 2025finding

Large Language Models (LLMs) are transforming healthcare through LLM-based agents that can understand and assist with medical tasks. This survey examines the architectures, applications, and challenges of LLM-based agents in medicine. We analyze key components including system profiles, clinical pla…

Cited by 0SourcePDFScholar
2025

A Unified Spatiotemporal Frequency Graph Neural Network for fMRI-based Brain Functional Connectivity Analysis

ICASSP 2025accepted

Analyzing functional connectivity patterns from resting-state functional magnetic resonance imaging (fMRI) requires unraveling its interrelations across spatial, temporal, and frequency domains. To comprehensively analyze four-dimensional (4D) fMRI data, we propose the Spatiotemporal Frequency Graph…

Cited by 0SourceScholar
2025

AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

ICML 2025spotlight

Fine-grained steering of language model outputs is essential for safety and reliability. Prompting and finetuning are widely used to achieve these goals, but interpretability researchers have proposed a variety of representation-based techniques as well, including sparse autoencoders (SAEs), linear…

2025

Balancing Privacy and Performance: A Many-in-One Approach for Image Anonymization

AAAI 2025technical

The effective utilization of data through Deep Neural Networks (DNNs) has profoundly influenced various aspects of society. The growing demand for high-quality, particularly personalized, data has spurred research efforts to prevent data leakage and protect privacy in recent years. Early privacy-pre…

Cited by 0SourcePDFScholar
2025

Boosting Lightweight Camouflaged Object Detection with Multi-Scale Context and Boundary Awareness

ICASSP 2025accepted

To adapt to the resource-limited environment, this study introduces the lightweight boundary-aware camouflaged object detection(COD) network LMABnet. We enhance the feature representation capability of the lightweight network through a multi-scale feature fusion architecture, while effectively avoid…

Cited by 0SourceScholar
2025

CCIN: Compositional Conflict Identification and Neutralization for Composed Image Retrieval

CVPR 2025highlight

Composed Image Retrieval (CIR) is a multi-modal task that seeks to retrieve target images by harmonizing a reference image with a modified instruction. A key challenge in CIR lies in compositional conflicts between the reference image (e.g., blue, long sleeve) and the modified instruction (e.g., gre…

2025

Controlling Thinking Speed in Reasoning Models

NeurIPS 2025spotlight

Human cognition is theorized to operate in two modes: fast, intuitive System 1 thinking and slow, deliberate System 2 thinking. While current Large Reasoning Models (LRMs) excel at System 2 thinking, their inability to perform fast thinking leads to high computational overhead and latency. In this w…

Cited by 0SourceScholar
2025

Cross-Category Subjectivity Generalization for Style-Adaptive Sketch Re-ID

ICCV 2025poster

Sketch-based person re-identification (re-ID) enables pedestrian retrieval using sketches. While recent methods have improved modality alignment between sketches and RGB images, the challenge of subjective style variation, where sketches exhibit diverse and unpredictable appearances, remains largely…

Cited by 0SourcePDFScholar
2025

Dataflow-Guided Neuro-Symbolic Language Models for Type Inference

ICML 2025poster

Language Models (LMs) are increasingly used for type inference, aiding in error detection and software development. Some real-world deployments of LMs require the model to run on local machines to safeguard the intellectual property of the source code. This setting often limits the size of the LMs…

Cited by 0SourcePDFScholar
2025

Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models

EMNLP 2025

Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RAG) systems. Ensuring its accuracy is therefore critical. But how accurate is Wikipedia, and how can we improve it?We foc

Cited by 0SourcePDFScholar
2025

Effects of Momentum in Implicit Bias of Gradient Flow for Diagonal Linear Networks

AAAI 2025technical

This paper targets on the regularization effect of momentum-based methods in regression settings and analyzes the popular diagonal linear networks to precisely characterize the implicit bias of continuous versions of heavy-ball (HB) and Nesterov's method of accelerated gradients (NAG). We show that,…

Cited by 0SourcePDFScholar
2025

Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection

ICML 2025oral

One-shot subset selection serves as an effective tool to reduce deep learning training costs by identifying an informative data subset based on the information extracted by an information extractor (IE). Traditional IEs, typically pre-trained on the target dataset, are inherently dataset-dependent.…

2025

Heavy-Ball Momentum Method in Continuous Time and Discretization Error Analysis

NeurIPS 2025poster

This paper establishes a continuous time approximation, a piece-wise continuous differential equation, for the discrete Heavy-Ball (HB) momentum method with explicit discretization error. Investigating continuous differential equations has been a promising approach for studying the discrete optimiza…

Cited by 0SourceScholar
2025

Joint Edge and Regional Depth Enhancement Network for Camouflaged Object Detection

ICASSP 2025accepted

Camouflaged object detection (COD) is a task of identifying and locating target objects that are camouflaged, masked, or confused. Research claims that depth cues can provide effective object location cues. However, depth images often contain noise interference, which may negatively affect object re…

Cited by 0SourceScholar
2025

LAMB: A Training-Free Method to Enhance the Long-Context Understanding of SSMs via Attention-Guided Token Filtering

ACL 2025short

State space models (SSMs) achieve efficient sub-quadratic compute complexity but often exhibit significant performance drops as context length increases. Recent work attributes this deterioration to an exponential decay in hidden-state memory. While token filtering has emerged as a promising remedy,…

2025

Language Pre-training Guided Masking Representation Learning for Time Series Classification

AAAI 2025technical

The representation learning of time series has a wide range of downstream tasks and applications in many practical scenarios. However, due to the complexity, spatiotemporality, and continuity of sequential stream data, compared with the representation learning of structural data such as images/video…

Cited by 0SourcePDFScholar
2025

Less is More: Improving LLM Alignment via Preference Data Selection

NeurIPS 2025spotlight

Direct Preference Optimization (DPO) has emerged as a promising approach for aligning large language models with human preferences. While prior work mainly extends DPO from the aspect of the objective function, we instead improve DPO from the largely overlooked but critical aspect of data selection.…

Cited by 0SourceScholar
2025

Multi-scale Re-weighted Attention Feature Fusion for Non-Intrusive Load Monitoring

ICASSP 2025accepted

Non-Intrusive Load Monitoring (NILM) addresses the challenge of disaggregating total energy consumption into individual appliance usage, which is essential for enhancing energy efficiency and managing smart grids. Existing methods often overlook the impact of window sizes on the separation of applia…

Cited by 0SourceScholar
2025

NeighborRetr: Balancing Hub Centrality in Cross-Modal Retrieval

CVPR 2025poster

Cross-modal retrieval aims to bridge the semantic gap between different modalities, such as visual and textual data, enabling accurate retrieval across them. Despite significant advancements with models like CLIP that align cross-modal representations, a persistent challenge remains: the hubness pro…

2025

Optimizing Personalized Federated Learning Through Adaptive Layer-Wise Learning

IJCAI 2025

Real-life deployment of federated Learning (FL) often faces non-IID data, which leads to poor accuracy and slow convergence. Personalized FL (pFL) tackles these issues by tailoring local models to individual data sources and using weighted aggregation methods for client-specific learning. However, e

2025

PiCNet: Physics-infused Convolution Network for Radar-Based Precipitation Nowcasting

ICASSP 2025accepted

Meteorological disasters, especially extreme precipitation, cause significant socioeconomic damage, highlighting the need for effective quantitative precipitation nowcasting. Existing methods, often data-driven and resource-intensive, struggle to capture the underlying physical laws of meteorology.…

Cited by 0SourceScholar
2025

Pioneering Explainable Video Fact-Checking with a New Dataset and Multi-role Multimodal Model Approach

AAAI 2025technical

Existing video fact-checking datasets often lack detailed evidence and explanations, compromising the reliability and interpretability of fact-checking methods. To address these gaps, we developed a novel dataset featuring comprehensive annotations for each news item, including veracity labels, the…

2025

ProjAttacker: A Configurable Physical Adversarial Attack for Face Recognition via Projector

CVPR 2025poster

Previous physical adversarial attacks have shown that carefully crafted perturbations can deceive face recognition systems, revealing critical security vulnerabilities. However, these attacks often struggle to impersonate multiple targets and frequently fail to bypass liveness detection. For example…

Cited by 0SourcePDFScholar
2025

ProtCLIP: Function-Informed Protein Multi-Modal Learning

AAAI 2025technical

Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-…

Cited by 2SourcePDFScholar
2025

Rethinking Cancer Gene Identification Through Graph Anomaly Analysis

AAAI 2025technical

Graph neural networks (GNNs) have shown promise in integrating protein-protein interaction (PPI) networks for identifying cancer genes in recent studies. However, due to the insufficient modeling of the biological information in PPI networks, more faithfully depiction of complex protein interaction…

2025

Self-Supervised Localized Topology Consistency for Noise-Robust Hyperspectral Image Classification

ICASSP 2025accepted

Label noise in hyperspectral image classification (HIC) can severely degrade model performance by leading to incorrect predictions and overfitting, especially as erroneous labels propagate and compound throughout the training process. To address this, we propose a robust learning framework called Se…

Cited by 0SourceScholar
2025

Single Pump-Valve Pneumatic Actuation With Continuous Flow Rate Control for Soft Robots

RA-L 2025

Pneumatic actuated soft robots attract increasing interest of the researchers due to the availability and simplicity in actuation. The soft robots driven by soft pneumatic actuators (SPAs) of various active volumes demand pneumatic systems with various range of flow rate. However, the usually bulky

Cited by 4SourceScholar
2025

Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided Diffusion

ICCV 2025poster

We introduce the concept of a subjective camera to reconstruct meaningful moments that physical cameras fail to capture. We propose Subjective Camera 1.0, a framework for reconstructing real-world scenes from readily accessible subjective readouts, i.e., textual descriptions and progressively drawn…

2025

Synergizing Multimodal Temporal Knowledge Graphs and Large Language Models for Social Relation Recognition

EMNLP 2025

Recent years have witnessed remarkable advances in Large Language Models (LLMs). However, in the task of social relation recognition, Large Language Models (LLMs) encounter significant challenges due to their reliance on sequential training data, which inherently restricts their capacity to effectiv

2025

Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer

AAAI 2025technical

Antibodies defend our health by binding to antigens with high specificity and potentiality, primarily relying on the Complementarity-Determining Region (CDR). Yet, current experimental methods of discovering new antibody CDRs are heavily time-consuming. Computational design could alleviate this burd…

2025

The Parables of the Mustard Seed and the Yeast: Extremely Low-Budget, High-Performance Nighttime Semantic Segmentation

AAAI 2025technical

Nighttime Semantic Segmentation (NSS) is essential to many cutting-edge vision applications. However, existing technologies overly rely on massive labeled data, whose annotation is time-consuming and laborious. In this paper, we pioneer a new task focusing on exploring the potential of training stra…

Cited by 0SourcePDFScholar
2025

TokenMatcher: Diverse Tokens Matching for Unsupervised Visible-Infrared Person Re-Identification

AAAI 2025technical

Unsupervised visible-infrared person re-identification (US-VI-ReID) seeks to match infrared and visible images of the same individual without the use of annotations. Current methods typically derive cross-modal correspondences through a single global feature matching process for generating pseudo la…

2025

TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection

EMNLP 2025

Rapid advances in Large Language Models (LLMs) have spurred demand for processing extended context sequences in contemporary applications. However, this progress faces two challenges: performance degradation due to sequence lengths out-of-distribution, and excessively long inference times caused by

2025

VEGAS: Towards Visually Explainable and Grounded Artificial Social Intelligence

AAAI 2025technical

Social Intelligence Queries (Social-IQ) serve as the primary multimodal benchmark for evaluating a model’s social intelligence level. While impressive multiple-choice question (MCQ) accuracy is achieved by current solutions, increasing evidence shows that they are largely, and in some cases entire…

2025

Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding

NeurIPS 2025poster

Speculative decoding improves LLM inference by generating and verifying multiple tokens in parallel, but existing systems suffer from suboptimal performance due to a mismatch between dynamic speculation and static runtime assumptions. We present Yggdrasil, a co-designed system that enables latency-o…

Cited by 0SourceScholar
2025

You Think, You ACT: The New Task of Arbitrary Text to Motion Generation

ICCV 2025poster

Text to Motion aims to generate human motions from texts. Existing settings rely on limited Action Texts that include action labels (e.g., "walk, bend"), which limits flexibility and practicability in scenarios difficult to describe directly. This paper extends limited Action Texts to arbitrary ones…

2024

A3S: A General Active Clustering Method with Pairwise Constraints

ICML 2024poster

Active clustering aims to boost the clustering performance by integrating human-annotated pairwise constraints through strategic querying. Conventional approaches with semi-supervised clustering schemes encounter high query costs when applied to large datasets with numerous classes. To address these…

2024

Bridge-IF: Learning Inverse Protein Folding with Markov Bridges

NeurIPS 2024poster

Inverse protein folding is a fundamental task in computational protein design, which aims to design protein sequences that fold into the desired backbone structures. While the development of machine learning algorithms for this task has seen significant success, the prevailing approaches, which pred…

2024

Capturing Detail Variations for Lightweight Neural Radiance Fields

ICASSP 2024accepted

Neural Radiance Fields (NeRF) has recently overhauled novel view synthesis, but it requires extensive computations for training and captures variations in detail with difficulty. In this paper, we propose a novel framework, termed CD-TDRF, to mitigate these dilemmas. CD-TDRF factorizes a density vox…

Cited by 0SourceScholar
2024

CodeIP: A Grammar-Guided Multi-Bit Watermark for Large Language Models of Code

EMNLP 2024finding

Large Language Models (LLMs) have achieved remarkable progress in code generation. It now becomes crucial to identify whether the code is AI-generated and to determine the specific model used, particularly for purposes such as protecting Intellectual Property (IP) in industry and preventing cheating…

2024

Contributing Dimension Structure of Deep Feature for Coreset Selection

AAAI 2024technical

Coreset selection seeks to choose a subset of crucial training samples for efficient learning. It has gained traction in deep learning, particularly with the surge in training dataset sizes. Sample selection hinges on two main aspects: a sample's representation in enhancing performance and the role…

2024

Crafting Personalized Agents through Retrieval-Augmented Generation on Editable Memory Graphs

EMNLP 2024main

In the age of mobile internet, user data, often referred to as memories, is continuously generated on personal devices. Effectively managing and utilizing this data to deliver services to users is a compelling research topic. In this paper, we introduce a novel task of crafting personalized agents p…

2024

Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

EMNLP 2024main

In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content. In this paper, we explore the novel challen…

2024

Ensemble Diversity Facilitates Adversarial Transferability

CVPR 2024poster

With the advent of ensemble-based attacks the transferability of generated adversarial examples is elevated by a noticeable margin despite many methods only employing superficial integration yet ignoring the diversity between ensemble models. However most of them compromise the latent value of the d…

2024

Expressiveness is Effectiveness: Self-supervised Fashion-aware CLIP for Video-to-Shop Retrieval

IJCAI 2024poster

The rise of online shopping and social media has spurred the Video-to-Shop Retrieval (VSR) task, which involves identifying fashion items (e.g., clothing) in videos and matching them with identical products provided by stores. In real-world scenarios, human movement in dynamic video scenes can cause…

Cited by 1SourcePDFScholar
2024

FedPFT: Federated Proxy Fine-Tuning of Foundation Models

IJCAI 2024poster

Adapting Foundation Models (FMs) for down- stream tasks through Federated Learning (FL) emerges a promising strategy for protecting data privacy and valuable FMs. Existing methods fine- tune FM by allocating sub-FM to clients in FL, however, leading to suboptimal performance due to insufficient tuni…

2024

Foam-Embedded Soft Robotic Joint With Inverse Kinematic Modeling by Iterative Self-Improving Learning

RA-L 2024

Soft robotic arms have gained significant attention owing to their flexibility and adaptability. Nonetheless, the instability due to their high-elasticity structure further leads to the difficulty of precise kinematic modeling and control. This letter introduces a novel solution employing foam-embed

Cited by 5SourceScholar
2024

Functional Bayesian Tucker Decomposition for Continuous-indexed Tensor Data

ICLR 2024poster

Tucker decomposition is a powerful tensor model to handle multi-aspect data. It demonstrates the low-rank property by decomposing the grid-structured data as interactions between a core tensor and a set of object representations (factors). A fundamental assumption of such decomposition is that ther…

2024

GRAPH-CONSTRAINED DIFFUSION FOR END-TO-END PATH PLANNING

ICLR 2024poster

Path planning underpins various applications such as transportation, logistics, and robotics. Conventionally, path planning is formulated with explicit optimization objectives such as distance or time. However, real-world data reveals that user intentions are hard-to-model, suggesting a need for dat…

Cited by 12SourcePDFScholar
2024

High-Performance Hydraulic Soft Robotic Control Using Continuous Flow Regulation and Partial Feedback

RA-L 2024

Hydraulic-driven soft robots have received much less attention than their pneumatic counterparts. However, the incompressibility of liquid could bring a series of desirable attributes to soft robotic control, despite the apparent disadvantage of weight addition and compliance compromise resulting fr

Cited by 7SourceScholar
2024

HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse Poses

CVPR 2024poster

We present HumanNeRF-SE a simple yet effective method that synthesizes diverse novel pose images with simple input. Previous HumanNeRF works require a large number of optimizable parameters to fit the human images. Instead we reload these approaches by combining explicit and implicit human represent…

Cited by 4SourcePDFScholar
2024

IQ-VFI: Implicit Quadratic Motion Estimation for Video Frame Interpolation

CVPR 2024poster

Advanced video frame interpolation (VFI) algorithms approximate intermediate motions between two input frames to synthesize intermediate frame. However they struggle to handle complex scenarios with curvilinear motions since they overlook the latent acceleration information between the input frames.…

Cited by 7SourcePDFScholar
2024

Introducing Compiler Semantics into Large Language Models as Programming Language Translators: A Case Study of C to x86 Assembly

EMNLP 2024finding

Compilers are complex software containing millions of lines of code, taking years to develop. This paper investigates to what extent Large Language Models (LLMs) can replace hand-crafted compilers in translating high-level programming languages to machine instructions, using C to x86 assembly as a c…

Cited by 0SourcePDFScholar
2024

Iterative Refinement of Project-Level Code Context for Precise Code Generation with Compiler Feedback

ACL 2024findings

Large Language Models (LLMs) have shown remarkable progress in automated code generation. Yet, LLM-generated code may contain errors in API usage, class, data structure, or missing project-specific information. As much of this project-specific context cannot fit into the prompts of LLMs, we must fin…

2024

Learning Based Exteroception of Soft Underwater Manipulator With Soft Actuator Network

RA-L 2024

Interactions with environmental objects can induce substantial alterations in both exteroceptive and proprioceptive signals. However, the deployment of exteroceptive sensors within underwater soft manipulators encounters numerous challenges and constraints, thereby imposing limitations on their perc

Cited by 0SourceScholar
2024

M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions

ACL 2024long

Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant memories from an external database. However, existing RAG methods typically organize all memories in a whole database, potentially limiting focus on crucial memories and introducing noise. In this paper…

Cited by 20SourcePDFScholar
2024

Multi-View Subspace Clustering With Consensus Graph Contrastive Learning

ICASSP 2024accepted

A significant challenge in multi-view clustering lies in the comprehensive extraction of consistency and complementary information from heterogeneous multi-view data. Numerous methods employ contrastive learning techniques to explore the information between views. However, the basic contrastive lear…

Cited by 0SourceScholar
2024

Open-Vocabulary Video Relation Extraction

AAAI 2024technical

A comprehensive understanding of videos is inseparable from describing the action with its contextual action-object interactions. However, many current video understanding tasks prioritize general action classification and overlook the actors and relationships that shape the nature of the action, re…

2024

Outlier-Robust Feature Selection with ℓ2, 1-Norm Minimization and Group Row-Sparsity Induced Constraints

ICASSP 2024accepted

In the realm of high-dimensional data analysis, the existence of outliers presents a substantial hurdle to the efficacy of feature selection methods that rely on the assumption of Gaussian distribution. To tackle this issue, we propose an outlier-robust feature selection method, ORFS, which combines…

Cited by 0SourceScholar
2024

Perturbation Guiding Contrastive Representation Learning for Time Series Anomaly Detection

IJCAI 2024poster

Time series anomaly detection is a critical task with applications in various domains. Due to annotation challenges, self-supervised methods have become the mainstream approach for time series anomaly detection in recent years. However, current contrastive methods categorize data perturbations int…

Cited by 2SourcePDFScholar
2024

PrefAce: Face-Centric Pretraining with Self-Structure Aware Distillation

AAAI 2024technical

Video-based facial analysis is important for autonomous agents to understand human expressions and sentiments. However, limited labeled data is available to learn effective facial representations. This paper proposes a novel self-supervised face-centric pretraining framework, called PrefAce, which l…

2024

RBI-RRT*: Efficient Sampling-based Path Planning for High-dimensional State Space

ICRA 2024poster

Sampling-based planning algorithms such as RRT have been proved to be efficient in solving path planning problems for robotic systems. Various improvements to the RRT algorithm have been presented to improve the performance of the extension and convergence of the random trees, such as Informed RRT*.…

Cited by 5SourceScholar
2024

ReFT: Representation Finetuning for Language Models

NeurIPS 2024spotlight

Parameter-efficient finetuning (PEFT) methods seek to adapt large neural models via updates to a small number of *weights*. However, much prior interpretability work has shown that *representations* encode rich semantic information, suggesting that editing representations might be a more powerful al…

2024

Revisiting Adversarial Patches for Designing Camera-Agnostic Attacks against Person Detection

NeurIPS 2024poster

Physical adversarial attacks can deceive deep neural networks (DNNs), leading to erroneous predictions in real-world scenarios. To uncover potential security risks, attacking the safety-critical task of person detection has garnered significant attention. However, we observe that existing attack met…

Cited by 1SourcePDFScholar
2024

Soft Robotic Proprioception Enhancement via 3D-Printed Differential Optical Waveguide Design

RA-L 2024

Soft robots undergo complex deformations during actuation and interaction due to the flexibility and compliance of their soft materials. This characteristic presents challenges in proprioception, particularly in characterizing their spatial deformations. Soft optical waveguide sensors have emerged a

Cited by 7SourceScholar
2024

Sparse Multi-Relational Graph Convolutional Network for Multi-type Object Trajectory Prediction

IJCAI 2024poster

Object trajectory prediction is a hot research issue with wide applications in video surveillance and autonomous driving. The previous studies consider the interaction sparsity mainly among the pedestrians instead of multi-type of objects, which brings new types of interactions and consequently supe…

Cited by 1SourcePDFScholar
2024

SynSP: Synergy of Smoothness and Precision in Pose Sequences Refinement

CVPR 2024poster

Predicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily co…

2024

The Implicit Bias of Gradient Descent toward Collaboration between Layers: A Dynamic Analysis of Multilayer Perceptions

NeurIPS 2024poster

The implicit bias of gradient descent has long been considered the primary mechanism explaining the superior generalization of over-parameterized neural networks without overfitting, even when the training error is zero. However, the implicit bias toward adversarial robustness has rarely been consid…

Cited by 0SourcePDFScholar
2024

Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration

ICML 2024poster

Attention is a fundamental component behind the remarkable achievements of large language models (LLMs). However, our current understanding of the attention mechanism, especially regarding how attention distributions are established, remains limited. Inspired by recent studies that explore the prese…

2024

When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models

ICML 2024poster

Autoregressive Large Language Models (LLMs) have achieved impressive performance in language tasks but face two significant bottlenecks: (1) quadratic complexity in the attention module as the number of tokens increases, and (2) limited efficiency due to the sequential processing nature of autoregre…

2024

Zero-shot Object Counting with Good Exemplars

ECCV 2024poster

"Zero-shot object counting (ZOC) aims to enumerate objects in images using only the names of object classes during testing, without the need for manual annotations. However, a critical challenge in current ZOC methods lies in their inability to identify high-quality exemplars effectively. This defic…

2024

pyvene: A Library for Understanding and Improving PyTorch Models via Interventions

NAACL 2024system demonstrations

Interventions on model-internal states are fundamental operations in many areas of AI, including model editing, steering, robustness, and interpretability. To facilitate such research, we introduce pyvene, an open-source Python library that supports customizable interventions on a range of different…

2023

A Strong Underwater Soft Manipulator With Planarly-Bundled Actuators and Accurate Position Control

RA-L 2023

Soft robotic manipulators have inherent advantages in underwater applications, as they generate motion by deforming seamless muscles rather than having rotational joints or sliding cylinders, as well as having excellent passive adaptability. However, limited by insufficient structural stiffness, ach

Cited by 7SourceScholar
2023

Background-Weakening Consistency Regularization for Semi-Supervised Video Action Detection

ICASSP 2023accepted

Consistency-based techniques have produced state-of-the-art results in semi-supervised action detection. When the model false detects the dynamic information in the background as an action, spatio-temporal consistency calculations can hardly reflect this false detection result. We consider weakening…

Cited by 0SourceScholar
2023

Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic Segmentation

ICASSP 2023accepted

While enlightening progress has been made recently in single-target domain adaptive semantic segmentation (ST-DASS), the multi-peak distributed multi-target domain cannot be directly aligned well with the single-peak distributed source domain. As a result, it is impossible for existing methods to ha…

Cited by 0SourceScholar
2023

DisCo: Distilled Student Models Co-training for Semi-supervised Text Mining

EMNLP 2023long main

Many text mining models are constructed by fine-tuning a large deep pre-trained language model (PLM) in downstream tasks. However, a significant challenge that arises nowadays is how to maintain performance when we use a lightweight model with limited labeled samples. We present DisCo, a semi-super…

Cited by 0SourcecodeScholar
2023

Don't Ignore Alienation and Marginalization: Correlating Fraud Detection

IJCAI 2023poster

The anonymity of online networks makes tackling fraud increasingly costly. Thanks to the superiority of graph representation learning, graph-based fraud detection has made significant progress in recent years. However, upgrading fraudulent strategies produces more advanced and difficult scams. One c…

Cited by 6SourcePDFScholar
2023

Dynamic Tensor Decomposition via Neural Diffusion-Reaction Processes

NeurIPS 2023spotlight

Tensor decomposition is an important tool for multiway data analysis. In practice, the data is often sparse yet associated with rich temporal information. Existing methods, however, often under-use the time information and ignore the structural knowledge within the sparsely observed tensor entries.…

2023

FedGS: Federated Graph-Based Sampling with Arbitrary Client Availability

AAAI 2023technical

While federated learning has shown strong results in opti- mizing a machine learning model without direct access to the original data, its performance may be hindered by in- termittent client availability which slows down the conver- gence and biases the final learned model. There are significant ch…

2023

From Generation to Suppression: Towards Effective Irregular Glow Removal for Nighttime Visibility Enhancement

IJCAI 2023poster

Most existing Low-Light Image Enhancement (LLIE) methods are primarily designed to improve brightness in dark regions, which suffer from severe degradation in nighttime images. However, these methods have limited exploration in another major visibility damage, the glow effects in real night scenes.…

Cited by 5SourcePDFScholar
2023

Good Is Bad: Causality Inspired Cloth-Debiasing for Cloth-Changing Person Re-Identification

CVPR 2023poster

Entangled representation of clothing and identity (ID)-intrinsic clues are potentially concomitant in conventional person Re-IDentification (ReID). Nevertheless, eliminating the negative impact of clothing on ID remains challenging due to the lack of theory and the difficulty of isolating the exact…

2023

HOTCOLD Block: Fooling Thermal Infrared Detectors with a Novel Wearable Design

AAAI 2023technical

Adversarial attacks on thermal infrared imaging expose the risk of related applications. Estimating the security of these systems is essential for safely deploying them in the real world. In many cases, realizing the attacks in the physical space requires elaborate special perturbations. These solut…

2023

Multilateral Semantic Relations Modeling for Image Text Retrieval

CVPR 2023poster

Image-text retrieval is a fundamental task to bridge vision and language by exploiting various strategies to fine-grained alignment between regions and words. This is still tough mainly because of one-to-many correspondence, where a set of matches from another modality can be accessed by a random qu…

Cited by 32SourcePDFScholar
2023

Origami Folding Enhances Modularity and Mechanical Efficiency of Soft Actuators

ICRA 2023poster

Soft robots have long been attractive to robotic engineers due to their remarkable dexterity; however, reports that standardize soft actuators into modularized off-shelf devices akin to rigid robots are still rare, and the mechanical efficiency of existing designs is still limited. This work identif…

Cited by 5SourceScholar
2023

Prototypical Residual Networks for Anomaly Detection and Localization

CVPR 2023poster

Anomaly detection and localization are widely used in industrial manufacturing for its efficiency and effectiveness. Anomalies are rare and hard to collect and supervised models easily over-fit to these seen anomalies with a handful of abnormal samples, producing unsatisfactory performance. On the o…

Cited by 84SourcePDFScholar
2023

Revisiting Domain-Adaptive 3D Object Detection by Reliable, Diverse and Class-balanced Pseudo-Labeling

ICCV 2023poster

Unsupervised domain adaptation (DA) with the aid of pseudo labeling techniques has emerged as a crucial approach for domain-adaptive 3D object detection. While effective, existing DA methods suffer from a substantial drop in performance when applied to a multi-class training setting, due to the co-e…

Cited by 30PDFcodeScholar
2023

Scratch Each Other's Back: Incomplete Multi-Modal Brain Tumor Segmentation via Category Aware Group Self-Support Learning

ICCV 2023poster

Although Magnetic Resonance Imaging (MRI) is very helpful for brain tumor segmentation and discovery, it often lacks some modalities in clinical practice. As a result, degradation of prediction performance is inevitable. According to current implementations, different modalities are considered to be…

Cited by 19PDFcodeScholar
2023

Spatial-Temporal Graph Learning with Adversarial Contrastive Adaptation

ICML 2023poster

Spatial-temporal graph learning has emerged as the state-of-the-art solution for modeling structured spatial-temporal data in learning region representations for various urban sensing tasks (e.g., crime forecasting, traffic flow prediction). However, most existing models are vulnerable to the qualit…

2023

Store and Fetch Immediately: Everything Is All You Need for Space-Time Video Super-resolution

AAAI 2023technical

Existing space-time video super-resolution (ST-VSR) methods fail to achieve high-quality reconstruction since they fail to fully explore the spatial-temporal correlations, long-range components in particular. Although the recurrent structure for ST-VSR adopts bidirectional propagation to aggregate i…

2023

Streaming Factor Trajectory Learning for Temporal Tensor Decomposition

NeurIPS 2023poster

Practical tensor data is often along with time information. Most existing temporal decomposition approaches estimate a set of fixed factors for the objects in each tensor mode, and hence cannot capture the temporal evolution of the objects' representation. More important, we lack an effective approa…

2023

Unsupervised Feature Selection with self-Weighted and ℓ2,0-Norm Constraint

ICASSP 2023accepted

At data mining field, it is a fundamental problem to dispose of high-dimensional data. Many existing unsupervised methods select features by manifold learning or exploring spectral analysis, thus preserving the intrinsic structure of raw data. But most of them follow an assumption that all features…

Cited by 0SourceScholar
2022

AutoIP: A United Framework to Integrate Physics into Gaussian Processes

ICML 2022spotlight

Physical modeling is critical for many modern science and engineering applications. From a data science or machine learning perspective, where more domain-agnostic, data-driven models are pervasive, physical knowledge {—} often expressed as differential equations {—} is valuable in that it is comple…

2022

Balanced Contrastive Learning for Long-Tailed Visual Recognition

CVPR 2022poster

Real-world data typically follow a long-tailed distribution, where a few majority categories occupy most of the data while most minority categories contain a limited number of samples. Classification models minimizing cross-entropy struggle to represent and classify the tail classes. Although the pr…

Cited by 259PDFcodeScholar
2022

Both Style and Fog Matter: Cumulative Domain Adaptation for Semantic Foggy Scene Understanding

CVPR 2022oral

Although considerable progress has been made in semantic scene understanding under clear weather, it is still a tough problem under adverse weather conditions, such as dense fog, due to the uncertainty caused by imperfect observations. Besides, difficulties in collecting and labeling foggy images hi…

Cited by 64PDFScholar
2022

Colorization for In Situ Marine Plankton Images

ECCV 2022poster

"Underwater imaging with red-NIR light illumination can avoid phototropic aggregation-induced observational deviation of marine plankton abundance under white light illumination, but this will lead to the loss of critical color information in the collected grayscale images, which is non-preferable t…

Cited by 2SourcePDFScholar
2022

Context-Enhanced Stereo Transformer

ECCV 2022poster

"Stereo depth estimation is of great interest for computer vision research. However, existing methods struggles to generalize and predict reliably in hazardous regions, such as large uniform regions. To overcome these limitations, we propose Context Enhanced Path (CEP). CEP improves the generalizati…

2022

DANet: Image Deraining via Dynamic Association Learning

IJCAI 2022poster

Rain streaks and background components in a rainy input are highly correlated, making the deraining task a composition of the rain streak removal and background restoration. However, the correlation of these two components is barely considered, leading to unsatisfied deraining results. To this end,…

Cited by 21SourcePDFScholar
2022

DISP6D: Disentangled Implicit Shape and Pose Learning for Scalable 6D Pose Estimation

ECCV 2022poster

"Scalable 6D pose estimation for rigid objects from RGB images aims at handling multiple objects and generalizing to novel objects. Building on a well-known auto-encoding framework to cope with object symmetry and the lack of labeled training data, we achieve scalability by disentangling the latent…

2022

Deep Learning-driven Front-Following within Close Proximity: a Hands-Free Control Model on a Smart Walker

ICRA 2022poster

With the ever-increasing elderly population, elder walking assistance is in strong demand. Instead of receiving assistance from a human carer, a smart walker can bring an elder user a more convenient and autonomous walking experience. Towards intelligent and safe walking assistance, we propose a clo…

Cited by 7SourceScholar
2022

Deep Multi-Fidelity Active Learning of High-Dimensional Outputs

AISTATS 2022poster

Many applications, such as in physical simulation and engineering design, demand we estimate functions with high-dimensional outputs. To reduce the expensive cost of generating training examples, we usually choose several fidelities to enable a cost/quality trade-off. In this paper, we consider the…

2022

Degrade Is Upgrade: Learning Degradation for Low-Light Image Enhancement

AAAI 2022technical

Low-light image enhancement aims to improve an image's visibility while keeping its visual naturalness. Different from existing methods, which tend to accomplish the relighting task directly, we investigate the intrinsic degradation and relight the low-light image while refining the details and colo…

2022

ELMA: Energy-Based Learning for Multi-Agent Activity Forecasting

AAAI 2022technical

This paper describes an energy-based learning method that predicts the activities of multiple agents simultaneously. It aims to forecast both upcoming actions and paths of all agents in a scene based on their past activities, which can be jointly formulated by a probabilistic model over time. Learni…

Cited by 6SourcePDFScholar
2022

Explicitly Modeling Importance and Coherence for Timeline Summarization

ICASSP 2022accepted

Timeline summarization (TLS) identifies major events and generates short summaries on how the event evolves in a period of time. Existing timeline summarization methods generate summaries by considering the coverage and diversity of the content and temporized information but ignore the importance an…

Cited by 0SourceScholar
2022

Exploring the Impact of Negative Samples of Contrastive Learning: A Case Study of Sentence Embedding

ACL 2022findings

Contrastive learning is emerging as a powerful technique for extracting knowledge from unlabeled data. This technique requires a balanced mixture of two ingredients: positive (similar) and negative (dissimilar) samples. This is typically achieved by maintaining a queue of negative samples during tra…

2022

Kinematic Analysis of Soft Continuum Manipulators Based on Sparse Workspace Mapping

RA-L 2022

Soft robots, with advantages of high adaptability to the environment, relatively easy and simple fabrication process as well as promising performances, have been thoroughly investigated and widely applied lately, the superiority of which has been proved in areas such as medicine, industry, daily lif

Cited by 9SourceScholar
2022

Multi-Dimensional Proprioception and Stiffness Tuning for Soft Robotic Joints

ICRA 2022poster

Proprioception and variable stiffness are two trending topics in soft robotics research. The former could endow soft robots with the ability to perceive the environment as well as their internal states without the need of dedicated sensors, while the latter could strengthen the otherwise excessive c…

Cited by 5SourceScholar
2022

Nonparametric Embeddings of Sparse High-Order Interaction Events

ICML 2022spotlight

High-order interaction events are common in real-world applications. Learning embeddings that encode the complex relationships of the participants from these events is of great importance in knowledge mining and predictive tasks. Despite the success of existing approaches, e.g. Poisson tensor factor…

Cited by 2SourcePDFScholar
2022

Nonparametric Sparse Tensor Factorization with Hierarchical Gamma Processes

ICML 2022spotlight

We propose a nonparametric factorization approach for sparsely observed tensors. The sparsity does not mean zero-valued entries are massive or dominated. Rather, it implies the observed entries are very few, and even fewer with the growth of the tensor; this is ubiquitous in practice. Compared with…

Cited by 8SourcePDFScholar
2022

Rainy WCity: A Real Rainfall Dataset with Diverse Conditions for Semantic Driving Scene Understanding

IJCAI 2022poster

Scene understanding in adverse weather conditions (e.g. rainy and foggy days) has drawn increasing attention, arising some specific benchmarks and algorithms. However, scene segmentation under rainy weather is still challenging and under-explored due to the following limitations on the datasets and…

Cited by 32SourcePDFScholar
2022

Spatial-Temporal Space Hand-in-Hand: Spatial-Temporal Video Super-Resolution via Cycle-Projected Mutual Learning

CVPR 2022poster

Spatial-Temporal Video Super-Resolution (ST-VSR) aims to generate super-resolved videos with higher resolution (HR) and higher frame rate (HFR). Quite intuitively, pioneering two-stage based methods complete ST-VSR directly combining two sub-tasks: Spatial Video Super-Resolution (S-VSR) and Temporal…

Cited by 44PDFcodeScholar
2022

VCD: View-Constraint Disentanglement for Action Recognition

ICASSP 2022accepted

Action recognition is a hot topic in computer vision due to its wide range of applications in urban surveillance. Although some methods are more advanced from an invariant view perspective, those approaches do not perform well for the viewpoint change. To address this issue, one possible solution is…

Cited by 0SourceScholar
2022

Vertebraic Soft Robotic Joint Design With Twisting and Antagonism

RA-L 2022

The soft robotic manipulators attract extensive interest of researchers due to its conformity to the unstructured environment, safe-interaction with human and fragile objects. The movement of the soft manipulator often include elongation, contraction, 2-DOF rotations due to the parallelly arranged f

Cited by 16SourceScholar
2022

Visual-tactile Sensing for Real-time Liquid Volume Estimation in Grasping

IROS 2022poster

We propose a deep visuo-tactile model for real-time estimation of the liquid inside a deformable container in a proprioceptive way. We fuse two sensory modalities, i.e., the raw visual inputs from the RGB camera and the tactile cues from our specific tactile sensor without any extra sensor calibrati…

Cited by 16SourceScholar
2021

FMA-ETA: Estimating Travel Time Entirely Based on FFN with Attention

ICASSP 2021accepted

Estimated time of arrival (ETA) is one of the most important services in intelligent transportation systems (ITS) and becomes a challenging spatial-temporal (ST) data mining task in recent years. Nowadays, deep learning based methods, specifically recurrent neural networks (RNN) based ones are adapt…

Cited by 0SourceScholar
2021

Fast Local Representation Learning with Adaptive Anchor Graph

ICASSP 2021accepted

Dimension reduction is an effective technology to embed data with high dimension to lower dimension space, where Linear Discriminant Analysis (LDA), one of representative methods, only works with Gaussian distribution data. However, in order to solve non-Gaussian issue that only one cluster cannot w…

Cited by 0SourceScholar
2021

Federated Learning with Fair Averaging

IJCAI 2021poster

Fairness has emerged as a critical problem in federated learning (FL). In this work, we identify a cause of unfairness in FL -- conflicting gradients with large differences in the magnitudes. To address this issue, we propose the federated fair averaging (FedFV) algorithm to mitigate potential confl…

2021

Image Inpainting Guided by Coherence Priors of Semantics and Textures

CVPR 2021poster

Existing inpainting methods have achieved promising performance in recovering defected images of specific scenes. However, filling holes involving multiple semantic categories remains challenging due to the obscure semantic boundaries and the mixture of different semantic textures. In this paper, we…

Cited by 120PDFScholar
2021

Learning to Attack Real-World Models for Person Re-identification via Virtual-Guided Meta-Learning

AAAI 2021technical

Recent advances in person re-identification (re-ID) have led to impressive retrieval accuracy. However, existing re-ID models are challenged by the adversarial examples crafted by adding quasi-imperceptible perturbations. Moreover, re-ID systems face the domain shift issue that training and testing…

2021

Location Predicts You: Location Prediction via Bi-direction Speculation and Dual-level Association

IJCAI 2021poster

Location prediction is of great importance in location-based applications for the construction of the smart city. To our knowledge, existing models for location prediction focus on the users' preference on POIs from the perspective of the human side. However, modeling users' interests from the histo…

Cited by 0SourcePDFScholar
2021

Modularized Interaction Network for Named Entity Recognition

ACL 2021long

Although the existing Named Entity Recognition (NER) models have achieved promising performance, they suffer from certain drawbacks. The sequence labeling-based NER models do not perform well in recognizing long entities as they focus only on word-level information, while the segment-based NER model…

Cited by 40SourcePDFScholar
2021

Multi-Fidelity High-Order Gaussian Processes for Physical Simulation

AISTATS 2021poster

The key task of physical simulation is to solve partial differential equations (PDEs) on discretized domains, which is known to be costly. In particular, high-fidelity solutions are much more expensive than low-fidelity ones. To reduce the cost, we consider novel Gaussian process (GP) models that le…

Cited by 17SourcePDFScholar
2021

Self-Adaptable Point Processes with Nonparametric Time Decays

NeurIPS 2021poster

Many applications involve multi-type event data. Understanding the complex influences of the events on each other is critical to discover useful knowledge and to predict future events and their types. Existing methods either ignore or partially account for these influences. Recent works use recurren…

Cited by 13SourcePDFScholar
2021

Very Important Person Localization in Unconstrained Conditions: A New Benchmark

AAAI 2021technical

This paper presents a new high-quality dataset for Very Important Person Localization (VIPLoc), named Unconstrained-7k. Generally, current datasets: 1) are limited in scale; 2) built under simple and constrained conditions, where the number of disturbing non-VIPs is not large, the scene is relativel…

2020

A High-Payload Proprioceptive Hybrid Robotic Gripper With Soft Origamic Actuators

RA-L 2020

Proprioception is the ability to perceive environmental stimulations through internal sensory organs. Enabling proprioception is critical for robots to be aware of the environmental interactions and respond appropriately, particularly for high-payload grippers to ensure safety when handling delicate

Cited by 39SourceScholar
2020

A Hybrid Underwater Manipulator System With Intuitive Muscle-Level sEMG Mapping Control

RA-L 2020

Soft-robotic manipulators, with their closed-chamber elastomeric actuators, natural water-sealing and inherent compliance, are ideal for underwater applications for compact, lightweight, and dexterous manipulation tasks. However, their low structure rigidity makes soft robots highly prone to underwa

Cited by 11SourceScholar
2020

A Proprioceptive Bellows (PB) Actuator With Position Feedback and Force Estimation

RA-L 2020

Soft robot is known for great safety in human-centered environments due to its inherent compliance. However, the compliance resulting from the soft continuum structure and viscoelastic material also induces challenges for sensing and control of soft robots. In this letter, we propose a proprioceptiv

Cited by 44SourceScholar
2020

BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning

NeurIPS 2020poster

There has recently been a surge in research in batch Deep Reinforcement Learning (DRL), which aims for learning a high-performing policy from a given dataset without additional interactions with the environment. We propose a new algorithm, Best-Action Imitation Learning (BAIL), which strives for bot…

2020

Beyond Intra-modality: A Survey of Heterogeneous Person Re-identification

IJCAI 2020poster

An efficient and effective person re-identification (ReID) system relieves the users from painful and boring video watching and accelerates the process of video analysis. Recently, with the explosive demands of practical applications, a lot of research efforts have been dedicated to heterogeneous pe…

2020

Discriminative Feature Selection via A Structured Sparse Subspace Learning Module

IJCAI 2020poster

In this paper, we first propose a novel Structured Sparse Subspace Learning S^3L module to address the long-standing subspace sparsity issue. Elicited by proposed module, we design a new discriminative feature selection method, named Subspace Sparsity Discriminant Feature Selection S^2DFS which enab…

2020

Guidance and Evaluation: Semantic-Aware Image Inpainting for Mixed Scenes

ECCV 2020poster

Completing a corrupted image with correct structures and reasonable textures for a mixed scene remains an elusive challenge. Since the missing hole in a mixed scene of a corrupted image often contains various semantic information, conventional two-stage approaches utilizing structural information of…

Cited by 154SourcePDFScholar
2020

Learning Canonical Shape Space for Category-Level 6D Object Pose and Size Estimation

CVPR 2020poster

We present a novel approach to category-level 6D object pose and size estimation. To tackle intra-class shape variations, we learn canonical shape space (CASS), a unified representation for a large variety of instances of a certain object category. In particular, CASS is modeled as the latent space…

Cited by 218PDFScholar
2020

Scalable Nonparametric Factorization for High-Order Interaction Events

AISTATS 2020poster

Interaction events among multiple entities are ubiquitous in real-world applications. Although these interactions can be naturally represented by tensors and analyzed by tensor decomposition, most existing approaches are limited to multilinear decomposition forms, and cannot estimate complex, nonlin…

2020

Universal Weighting Metric Learning for Cross-Modal Matching

CVPR 2020poster

Cross-modal matching has been a highlighted research topic in both vision and language areas. Learning appropriate mining strategy to sample and weight informative pairs is crucial for the cross-modal matching performance. However, most existing metric learning methods are developed for unimodal mat…

Cited by 113PDFcodeScholar
2020

When Pedestrian Detection Meets Nighttime Surveillance: A New Benchmark

IJCAI 2020poster

Pedestrian detection at nighttime is a crucial and frontier problem in surveillance, but has not been well explored by the computer vision and artificial intelligence communities. Most of existing methods detect pedestrians under favorable lighting conditions (e.g. daytime) and achieve promising per…

2019

A Compact Dental Robotic System Using Soft Bracing Technique

RA-L 2019

A wide range of commonly performed dental procedures, from operative caries removal, crown preparation, filling, to Orthodontia, could potentially benefit from robotic assistance or enhancement. Despite the wide applicability, dental robots have received far less research attention in comparison wit

Cited by 30SourceScholar
2019

A Joint Ordering, Pricing, and Freshness-Keeping Policy for Perishable Inventory Systems With Random Demand Over Infinite Horizon

RA-L 2019

For an inventory system of a single type perishable product where the demand is dependent in the sale price plus a random variable, and the decay rate is dependent in freshness-keeping effort plus a random variable, due to their quality deterioration the amount of available products that can be used

Cited by 9SourceScholar
2019

Cross-view Identical Part Area Alignment for Person Re-identification

ICASSP 2019accepted

Person re-identification aims to associate images captured by non-overlapping cameras. It is a challenging task because images are often in different conditions such as background clutter, illumination variation, viewpoint changes and different camera settings. Viewpoint changes and pose variations…

Cited by 0SourceScholar
2019

Design and Modeling of an Extensible Soft Robotic Arm

RA-L 2019

Soft robotic arms are receiving more and more attention for their intrinsic safety and natural compliance. Instead of traditional serialized rotary joints, soft robotic arms often have complex joints with coupled degrees of freedom like bending, rotation, and elongation, enabling them with more free

Cited by 34SourceScholar
2019

Learning to Reduce Dual-Level Discrepancy for Infrared-Visible Person Re-Identification

CVPR 2019poster

Infrared-Visible person RE-IDentification (IV-REID) is a rising task. Compared to conventional person re-identification (re-ID), IV-REID concerns the additional modality discrepancy originated from the different imaging processes of spectrum cameras, in addition to the person's appearance discrepanc…

Cited by 522PDFcodeScholar
2019

Reinforcement Learning Meets Hybrid Zero Dynamics: A Case Study for RABBIT

ICRA 2019poster

The design of feedback controllers for bipedal robots is challenging due to the hybrid nature of its dynamics and the complexity imposed by high-dimensional bipedal models. In this paper, we present a novel approach for the design of feedback controllers using Reinforcement Learning (RL) and Hybrid…

Cited by 28SourceScholar
2018

BCL-13: A 13-DOF Soft Robotic Hand for Dexterous Grasping and In-Hand Manipulation

RA-L 2018

This letter presents a dexterous soft robotic hand, BCL-13, with 4 fingers and 13 independently actuated joints capable of in-hand manipulation. The iconic dexterity is enabled by a novel soft robotic finger design with three degrees of freedom (DOFs), significantly improving over existing soft actu

Cited by 136SourceScholar
2018

Soft-Actuator-Based Robotic Joint for Safe and Forceful Interaction With Controllable Impact Response

RA-L 2018

Impact safety and response are critical challenges for robots working under dynamic environments and with close proximity to humans. State-of-the-art rigid robots and soft robots both have limitations and tradeoffs due to their characteristics. In this letter, we introduced a hybrid-antagonistic-pne

Cited by 29SourceScholar
2018

Visual Homing via Guided Locality Preserving Matching

ICRA 2018poster

This study proposes a simple yet surprisingly effective feature matching approach, termed as guided locality preserving matching (GLPM), for visual homing of panoramic images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two panoram…

Cited by 16SourceScholar