← Search

Heng Zhang

48 accepted papers

2026

FossilWriter: Learning Hypergraph World Models with Latent Narratives for Creative Story Generation

IJCAI 2026

Creative story generation has achieved notable progress with large language models. Current methods construct narratives through hierarchical planning or incremental expansion. These approaches produce structurally complete stories but offer limited support for organic narrative development. Many fi

Cited by 0Scholar
2026

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning

RSS 2026poster

Embodied AI research is undergoing a shift toward vision-centric perceptual paradigms. While massively parallel simulators have catalyzed breakthroughs in proprioception-based locomotion, their potential remains largely untapped for vision-centric tasks due to the prohibitive computational overhead …

Cited by 0SourceScholar
2026

History Doesn’t Repeat, but Its Patterns Echo: A Parallel Pairwise Negative-Sampling Framework for Temporal Link Prediction

IJCAI 2026

Temporal link prediction with temporal graph neural networks (TGNNs) is increasingly used to model spatio-temporal dependencies in temporal graphs and to forecast future interactions among entities. Existing sampling-based training methods typically rely on random negative sampling and pointwise los

Cited by 0Scholar
2026

Human-Centric Multi-Exposure Fusion: Benchmark and Bi-level Cognition Distillation Framework

CVPR 2026

Multi-Exposure Fusion (MEF) seeks to generate a single high-quality image from multiple inputs captured at different exposure levels. Despite substantial progress, most existing approaches depend on statistical metrics that poorly reflect human perceptual preferences. Electroencephalography (EEG) pr

Cited by 0SourcecodeScholar
2026

MOES-Pred: Molecular Structural Representation Learning by Adaptive Energy-Sentinel Vibration for Generalized Property Prediction

ICML 2026poster

Molecular property prediction from 3D structures is fundamentally constrained by the scarcity of labeled data. To address this challenge, researchers have adapted various self-supervised pre-training methods from computer vision and natural language processing; however, these approaches often neglec…

Cited by 0SourceScholar
2026

MachaGrasp: Morphology-Aware Cross-Embodiment Dexterous Hand Articulation Generation for Grasping

ICRA 2026poster

Dexterous grasping with multi-fingered hands remains challenging due to high-dimensional articulations and the cost of optimization-based pipelines. Existing end-to-end methods require training on large-scale datasets for specific hands, limiting their ability to generalize across different embodime…

2026

NeuralActuator: Neural Actuation Modeling for Robot Dynamics and External Force Perception

RSS 2026poster

Differentiable simulators have advanced policy learning and model-based control across diverse robotic tasks. To date, actuator dynamics remain underexplored and are a major source of sim-to-real error, especially on low-cost platforms where the linear current–torque model τ = K_tI breaks down under…

Cited by 0SourceScholar
2026

SAM2Text: Towards Prompt-Free and Multi-Resolution Video Scene Text Segmentation

CVPR 2026

We introduce a novel method for video Scene Text Segmentation (STS), a task critical for understanding dynamic visual content. Despite the success of foundation models like Segment Anything Model 2 (SAM2) in generic segmentation, their application to video STS is hindered by the reliance on external

Cited by 0SourcecodeScholar
2026

Semantic Contact Fields for Category-Level Generalizable Tool Manipulation

RSS 2026poster

Generalizing tool manipulation requires both semantic planning and precise physical control. Modern generalist robot policies, such as Vision-Language-Action (VLA) models, often lack the high-fidelity physical grounding required for contact-rich tool manipulation. Conversely, existing contact-aware …

Cited by 0SourceScholar
2026

ShieldedCode: Learning Robust Representations for Virtual Machine Protected Code

ICLR 2026poster

Large language models (LLMs) have achieved remarkable progress in code generation, yet their potential for software protection remains largely untapped. Reverse engineering continues to threaten software security, while traditional virtual machine protection (VMP) relies on rigid, rule-based tran…

Cited by 0SourceScholar
2025

"Oh! It's Fun Chatting with You!" a Humor-Aware Social Robot Chat Framework

ICRA 2025

Humor is a key element in human interactions, essential for building connections and rapport. To enhance human-robot communication, we developed a humor-aware chat framework that enables robots to deliver contextually appropriate humor. This framework takes into account the interaction environment,

Cited by 2SourceScholar
2025

ANASETC: Automatic Neural Architecture Search for Encrypted Traffic Classification

ICASSP 2025accepted

The widespread adoption of encrypted network protocols has made traffic encryption ubiquitous, creating substantial challenges for network management and security. This paper introduces a novel encrypted traffic classification system, ANASETC, which combines traffic burst features with Neural Archit…

Cited by 0SourceScholar
2025

ASCENT: Autonomous Skill Learning Toward Complex Embodied Tasks With Foundation Models

ICRA 2025

Collecting data from simulated scenarios for training robotic skills provides a safer and more controllable alternative to real-world environments. However, it demands considerable effort, including the manual construction of simulation environments, the careful design of tasks, and the challenge of

Cited by 0SourceScholar
2025

Foresee and Act Ahead: Task Prediction and Pre-Scheduling Enabled Efficient Robotic Warehousing

ICRA 2025

In warehousing systems, to enhance efficiency amid surging demand volumes, much attention has been placed on how to reasonably allocate tasks of delivery to robots. However, the labor of robots is still inevitably wasted to some extent. In this paper, we propose a pre-scheduling enhanced warehousing

Cited by 1SourceScholar
2025

HyperKAN: Hypergraph Representation Learning with Kolmogorov-Arnold Networks

ICASSP 2025accepted

Hypergraph representation learning has garnered increasing attention across various domains due to its capability to model high-order relationships. Traditional methods often rely on hypergraph neural networks (HNNs) employing messagepassing mechanisms to aggregate vertex and hyperedge features. How…

Cited by 0SourceScholar
2025

HyperSF: A Hypergraph Representation Learning Method Based on Structural Fusion

ICASSP 2025accepted

Hypergraph Neural Networks (HNNs) have recently gained attention as a powerful approach for capturing high-order correlations through hypergraph-structured encoding and learning techniques. However, despite their potential, existing HNN methods often encounter over-smoothing issues, which limit thei…

Cited by 0SourceScholar
2025

Improving Height Prediction for Vision-Based Roadside 3D Object Detection

ICASSP 2025accepted

Roadside vision-based 3D object detection is vital in many applications, such as autonomous driving. The mainstream methods enhance the accuracy of distance estimation by converting predicted height distribution into depth distribution. However, predicting object’s height in roadside perception is c…

Cited by 0SourceScholar
2025

TERL: Large-Scale Multi-Target Encirclement Using Transformer-Enhanced Reinforcement Learning

IROS 2025

Pursuit-evasion (PE) problem is a critical challenge in multi-robot systems (MRS). While reinforcement learning (RL) has shown its promise in addressing PE tasks, research has primarily focused on single-target pursuit, with limited exploration of multi-target encirclement, particularly in large-sca

Cited by 1SourcecodeScholar
2025

Unraveling the Mystery: Defending Against Jailbreak Attacks Via Unearthing Real Intention

COLING 2025main

As Large Language Models (LLMs) become more advanced, the security risks they pose also increase. Ensuring that LLM behavior aligns with human values, particularly in mitigating jailbreak attacks with elusive and implicit intentions, has become a significant challenge. To address this issue, we prop…

2024

An Octopus-Inspired-Configuration Sensor Array Concept toward Torso-Oriented Magnetic Localization Task and Simulation Verification

IROS 2024poster

In response to torso-oriented magnetic localization tasks that require the system to have interactivity and flexibility with guaranteed accuracy, a novel bio-inspired magnetic sensor array configuration is proposed in this paper. Precisely, the ideas of the natural characteristics of octopus flexibl…

Cited by 0SourceScholar
2024

Data-Driven Modeling of Ground Effect For UAV Landing on a Vertical Oscillating Platform

IROS 2024poster

Landing on a vertically oscillating platform poses a significant challenge for multi-rotor unmanned aerial vehicle (UAVs) due to the time-varying ground effect (GE). In this work, we formulated a data-driven GE dynamic model that accurately describes the complex interactions between UAVs and both st…

Cited by 0SourceScholar
2024

Decomposing Semantic Shifts for Composed Image Retrieval

AAAI 2024technical

Composed image retrieval is a type of image retrieval task where the user provides a reference image as a starting point and specifies a text on how to shift from the starting point to the desired target image. However, most existing methods focus on the composition learning of text and reference im…

2024

Exploring Region-Word Alignment in Built-in Detector for Open-Vocabulary Object Detection

CVPR 2024poster

Open-vocabulary object detection aims to detect novel categories that are independent from the base categories used during training. Most modern methods adhere to the paradigm of learning vision-language space from a large-scale multi-modal corpus and subsequently transferring the acquired knowledge…

Cited by 6SourcePDFScholar
2024

Extending Implicit Neural Representations for Text-to-Image Generation

ICASSP 2024accepted

Implicit neural representations (INRs) have demonstrated their effectiveness in continuous modeling for image signals. However, INRs typically operate in a continuous space, which makes it difficult to integrate the discrete symbols and structures inherent in human language. Despite this, text featu…

Cited by 0SourceScholar
2024

SRL-VIC: A Variable Stiffness-Based Safe Reinforcement Learning for Contact-Rich Robotic Tasks

RA-L 2024

Reinforcement learning (RL) has emerged as a promising paradigm in complex and continuous robotic tasks, however, safe exploration has been one of the main challenges, especially in contact-rich manipulation tasks in unstructured environments. Focusing on this issue, we propose <bold xmlns:mml="http

Cited by 25SourceScholar
2024

Toward Universal and Scalable Road Graph Partitioning for Efficient Multi-Robot Path Planning

IROS 2024

To date, multi-robot path planning has primarily been addressed by centralized solvers, typically aiming to maintain optimality. However, given its NP-hard nature, directly applying existing solvers in large and complex scenarios proves inefficient. A promising alternative lies in adopting a divide-

Cited by 1SourceScholar
2024

Traffic Flow Learning Enhanced Large-Scale Multi-Robot Cooperative Path Planning Under Uncertainties

ICRA 2024poster

Robotic systems with hundreds or even thousands of robots are widely implemented in logistic and industrial applications. In such systems, cooperative path planning is of great importance, as local congestion and motion conflict may greatly degrade system performance, especially in the presence of u…

Cited by 3SourceScholar
2023

Dancing in the Dark: A Benchmark towards General Low-light Video Enhancement

ICCV 2023poster

Low-light video enhancement is a challenging task with broad applications. However, current research in this area is limited by the lack of high-quality benchmark datasets. To address this issue, we design a camera system and collect a high-quality low-light video dataset with multiple exposures and…

Cited by 21PDFcodeScholar
2023

Exploring Temporal Concurrency for Video-Language Representation Learning

ICCV 2023poster

Paired video and language data is naturally temporal concurrency, which requires the modeling of the temporal dynamics within each modality and the temporal alignment across modalities simultaneously. However, most existing video-language representation learning methods only focus on discrete semant…

Cited by 4PDFcodeScholar
2023

Implicit Neural Field Guidance for Teleoperated Robot-assisted Surgery

ICRA 2023poster

Teleoperated techniques enable remote human-robot interaction and have been widely accepted in robot-assisted surgeries. However, it is still hard to guarantee the safety of teleoperated surgery due to the imperfect input commands limited by remote perception, preventing teleoperated surgery from be…

Cited by 5SourceScholar
2023

Incorporating Factuality Inference to Identify Document-level Event Factuality

ACL 2023findings

Document-level Event Factuality Identification (DEFI) refers to identifying the degree of certainty that a specific event occurs in a document. Previous studies on DEFI failed to link the document-level event factuality with various sentence-level factuality values in the same document. In this pape…

2023

Modeling Video As Stochastic Processes for Fine-Grained Video Representation Learning

CVPR 2023highlight

A meaningful video is semantically coherent and changes smoothly. However, most existing fine-grained video representation learning methods learn frame-wise features by aligning frames across videos or exploring relevance between multiple views, neglecting the inherent dynamic process of each video.…

2023

Refining 6-DoF Grasps with Context-Specific Classifiers

IROS 2023poster

In this work, we present GraspFlow, a refinement approach for generating context-specific grasps. We formulate the problem of grasp synthesis as a sampling problem: we seek to sample from a context-conditioned probability distribution of successful grasps. However, this target distribution is unknow…

Cited by 2SourcecodeScholar
2022

Design and Analysis of a Long-range Magnetic Actuated and Guided Endoscope for Uniport VATS

ICRA 2022poster

This paper presents a long-range magnetic actuated and guided endoscope for uniport video-assisted thoracic surgery (VATS). In VATS, the incision is quite narrow and part of the chest wall may be very thick. So, the magnetic endoscope system is required to produce sufficient attractive force at a co…

Cited by 5SourceScholar
2022

Design and Control of a Highly Redundant Rigid-flexible Coupling Robot to Assist the COVID-19 Oropharyngeal-Swab Sampling

RA-L 2022

The outbreak of novel coronavirus pneumonia (COVID-19) has caused mortality and morbidity worldwide. Oropharyngeal-swab (OP-swab) sampling is widely used for the diagnosis of COVID-19 in the world. To avoid the clinical staff from being affected by the virus, we developed a 9-degree-of-freedom (DOF)

Cited by 52SourceScholar
2022

Document-level Event Factuality Identification via Machine Reading Comprehension Frameworks with Transfer Learning

COLING 2022main

Document-level Event Factuality Identification (DEFI) predicts the factuality of a specific event based on a document from which the event can be derived, which is a fundamental and crucial task in Natural Language Processing (NLP). However, most previous studies only considered sentence-level task…

Cited by 9SourcePDFScholar
2022

Multi-Robot Object Transport Motion Planning With a Deformable Sheet

RA-L 2022

Using a deformable sheet to handle objects is convenient and found in many practical applications. For object manipulation through a deformable sheet that is held by multiple mobile robots, it is a challenging task to model the object-sheet interactions. We present a computational model and algorith

Cited by 31SourceScholar
2021

Design and Implementation of a Novel, Intrinsically Safe Rigid-Flexible Coupling Manipulator for COVID-19 Oropharyngeal Swab Sampling

ICRA 2021poster

Driven by the SARS-CoV-2 pandemic, demand for oropharyngeal swab sampling (OP-swabs) is surging. However, medical staff can easily become infected by the virus during the sampling process. In an effort to combat this, we developed a novel, intrinsically safe rigid- flexible coupling (RFC) manipulato…

Cited by 10SourceScholar
2021

Less Is More: Domain Adaptation with Lottery Ticket for Reading Comprehension

EMNLP 2021finding

In this paper, we propose a simple few-shot domain adaptation paradigm for reading comprehension. We first identify the lottery subnetwork structure within the Transformer-based source domain model via gradual magnitude pruning. Then, we only fine-tune the lottery subnetwork, a small fraction of the…