← Search

Bo He

24 accepted papers

2026

FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding

CVPR 2026

This paper presents FluxMem, a training-free framework for efficient streaming video understanding. FluxMem adaptively compresses redundant visual memory through a hierarchical, two-stage design: (1) a Temporal Adjacency Selection (TAS) module removes redundant visual tokens across adjacent frames,

Cited by 0SourcecodeScholar
2026

Hard-Constrained Graph Generation with Discrete-Projection Diffusion

ICML 2026poster

Diffusion models have achieved remarkable success in graph generation, but enforcing hard constraints on generated graphs remains challenging, limiting their deployment in constraint-critical applications. Existing approaches either fail to guarantee strict constraint satisfaction or are limited to …

Cited by 0SourceScholar
2026

M-LoRA: Efficient Serving for Concurrent LoRA Adapters with Memory-Aware Speculative Scheduler on Single GPU

IJCAI 2026

Low-Rank Adaptation (LoRA) is a popular approach that enables large language models (LLMs) to quickly adapt to domain-specific tasks by adding lightweight trainable adapters. Existing multi-LoRA serving systems typically exploit parameter sharing to serve hundreds of LoRA models with a single base m

Cited by 0Scholar
2026

NeRV-Diffusion: Diffuse Implicit Neural Representation for Video Synthesis

ICLR 2026poster

We present NeRV-Diffusion, an implicit latent video diffusion model that synthesizes videos via generating neural network weights. The generated weights can be rearranged as the parameters of a convolutional neural network, which forms an implicit neural representation (INR), and decodes into videos…

Cited by 0SourceScholar
2026

Transferring Policy of Offline Reinforcement Learning From Hybrid Dataset to Real World via Progressive Neural Network

RA-L 2026

Offline reinforcement learning (Offline RL) provides a compelling solution for applying RL in high-risk or resourceconstrained real-world domains such as healthcare, autonomous driving, and robotic manipulation, where online exploration can be unsafe or impractical. However, Offline RL faces critica

Cited by 0SourceScholar
2026

Transferring Policy of Offline Reinforcement Learning from Hybrid Dataset to Real World Via Progressive Neural Network

ICRA 2026poster

Offline reinforcement learning (Offline RL) provides a compelling solution for applying RL in high-risk or resource-constrained real-world domains such as healthcare, autonomous driving, and robotic manipulation. However, Offline RL faces critical challenges arising from limited data coverage and po…

Cited by 0SourceScholar
2026

VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding

ICML 2026poster

Recent advancements in Video Large Language Models (Video LLMs) have demonstrated impressive results, yet existing approaches handle either temporal or spatial dimension in isolation, struggling in the analysis of complex events that require spatial-temporal integration. To bridge this gap, we propo…

Cited by 5SourceScholar
2025

Rethinking Smoothness for Fast and Adaptable Entity Alignment Decoding

NAACL 2025findings

Entity alignment (EA) is crucial for integrating multi-source knowledge graphs (KGs), aiming to identify equivalent entities across different graphs. However, most existing EA decoding methods rely on both entity and relation embeddings, limiting their generalizability and efficiency, especially in…

2025

Spy Inside: Scalable Verification of Dependable Transformers for Event Time Series Systems

ICASSP 2025accepted

Event time series appear in many software scenarios and are a necessary data type in data analytics systems. Transformers are the preferred type of sequential neural network for advanced analytics on event time series, particularly due to their significant contributions to the recent surge of large…

Cited by 0SourceScholar
2024

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

CVPR 2024poster

With the success of large language models (LLMs) integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However existing LLM-based large multimodal models (e.g. Video-LLaMA VideoChat) can only take in a limited number of frames for s…

2024

OmniViD: A Generative Framework for Universal Video Understanding

CVPR 2024poster

The core of video understanding tasks such as recognition captioning and tracking is to automatically detect objects or actions in a video and analyze their temporal evolution. Despite sharing a common goal different tasks often rely on distinct model architectures and annotation formats. In contras…

2024

Opposition Movement of the Little Finger Enhances the Grasping Function of Robotic Hands

RA-L 2024

Most contemporary hand robots only consider the palmar opposition as the thumb opposition, while ignoring the important role of the opposition of little finger. Consequently, the back of the hand is considered as a rigid, immovable structure. This letter investigated the influence of little finger o

Cited by 2SourceScholar
2024

Visual Representation of the Compactness of a Stephenson-II Six-Bar Linkage Exoskeleton Using Solution Region Synthesis Theory

RA-L 2024

With the increase of sensors and actuators, the size and weight of intelligent robots are becoming larger and bulkier. Especially for wearable robots, excessively large exoskeleton dimensions will reduce the portability and the wearing experience for patients. In this paper, we propose an evaluation

Cited by 4SourceScholar
2023

Align and Attend: Multimodal Summarization With Dual Contrastive Losses

CVPR 2023poster

The goal of multimodal summarization is to extract the most important information from different modalities to form summaries. Unlike unimodal summarization, the multimodal summarization task explicitly leverages cross-modal information to help generate more reliable and high-quality summaries. Howe…

2023

An Anthropomorphic Robotic Hand With a Soft-Rigid Hybrid Structure and Positive- Negative Pneumatic Actuation

RA-L 2023

Anthropomorphic robotic hands are seeking to achieve key features such as multi-degree-of-freedom motion ability, bi-directional actuation, high adaptability, and sufficient stiffness. In this research, we propose a 10 active degrees-of-freedom anthropomorphic robotic hand with a soft-rigid hybrid s

Cited by 22SourceScholar
2023

Chop & Learn: Recognizing and Generating Object-State Compositions

ICCV 2023poster

Recognizing and generating object-state compositions has been a challenging task, especially when generalizing to unseen compositions. In this paper, we study the task of cutting objects in different styles and the resulting object state changes. We propose a new benchmark suite Chop & Learn, to acc…

Cited by 18PDFcodeScholar
2023

Towards Scalable Neural Representation for Diverse Videos

CVPR 2023poster

Implicit neural representations (INR) have gained increasing attention in representing 3D scenes and images, and have been recently applied to encode videos (e.g., NeRV, E-NeRV). While achieving promising results, existing INR-based methods are limited to encoding a handful of short videos (e.g., se…

Cited by 45SourcePDFScholar
2022

ASM-Loc: Action-Aware Segment Modeling for Weakly-Supervised Temporal Action Localization

CVPR 2022poster

Weakly-supervised temporal action localization aims to recognize and localize action segments in untrimmed videos given only video-level action labels for training. Without the boundary information of action segments, existing methods mostly rely on multiple instance learning (MIL), where the predic…

Cited by 120PDFcodeScholar
2022

Affective Behavior Learning for Social Robot Haru with Implicit Evaluative Feedback

IROS 2022poster

We propose a human-in-the-loop reinforcement learning mechanism to help robots learn emotional behavior. Unlike the previous methods of providing explicit feedback via pressing keyboard buttons or mouse clicks, we provide a more natural way for ordinary people to train social robots how to perform s…

Cited by 4SourceScholar
2022

Fusing Topology Optimization and Pseudo-Rigid-Body Method For the Development of a Finger Exoskeleton

RA-L 2022

Robotic hand exoskeletons can assist people who suffer from hand.functional disabilities caused by a stroke. However, currently existing hand exoskeletons remain inadequate in terms of user-friendly design, lightweight structure, and accurate modeling of hand motion. In this study, a large displacem

Cited by 19SourceScholar
2022

Learning Semantic Correspondence with Sparse Annotations

ECCV 2022poster

"Finding dense semantic correspondence is a fundamental problem in computer vision, which remains challenging in complex scenes due to background clutter, extreme intra-class variation, and a severe lack of ground truth. In this paper, we aim to address the challenge of label sparsity in semantic co…

2021

Developing of A Rigid-Compliant Finger Joint Exoskeleton Using Topology Optimization Method

ICRA 2021poster

Robotic hand exoskeletons can provide assistance to people who suffer from hand functional disability or spinal cord injury (SCI). However, the current hand exoskeletons remain challenging with respect to having a user-friendly design that satisfies human motion with a lightweight structure. Here we…

Cited by 6SourceScholar
2021

NeRV: Neural Representations for Videos

NeurIPS 2021poster

We propose a novel neural representation for videos (NeRV) which encodes videos in neural networks. Unlike conventional representations that treat videos as frame sequences, we represent videos as neural networks taking frame index as input. Given a frame index, NeRV outputs the corresponding RGB i…

2021

Shaping Progressive Net of Reinforcement Learning for Policy Transfer with Human Evaluative Feedback

IROS 2021poster

Deep reinforcement learning has achieved significant success in many fields, but will confront sampling efficiency and safety problems when applying to robot control in the real world. Sim-to-real transfer learning was proposed to make use of samples in the simulation and overcome the gap between si…

Cited by 9SourceScholar