← Search

hongyu wang

18 accepted papers

2026

JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials

ICML 2026poster

Discovering atom-level phenomena requires molecular dynamics (MD) simulations with ab initio accuracy. Machine learning interatomic potentials (MLIPs) enable stable, high-accuracy MD simulations, and their models exhibit scaling-law trends similar to large language models. However, the lack of scala…

Cited by 0SourceScholar
2026

Joint Navigation and Manipulation Planning with 3D Interaction Chains

ICML 2026poster

Open-vocabulary mobile manipulation (OVMM) requires long-horizon navigation in unseen environments and object-centric manipulation. Most existing methods treat navigation and manipulation as separate stages, which can yield navigation endpoints that are poor for manipulation or manipulation-friendly…

Cited by 0SourceScholar
2026

MatRIS: Toward Reliable and Efficient Pretrained Machine Learning Interaction Potentials

ICLR 2026poster

Universal MLIPs (uMLIPs) demonstrate broad applicability across diverse material systems and have emerged as a powerful and transformative paradigm in chemical and computational materials science. Equivariant uMLIPs achieve state-of-the-art accuracy in a wide range of benchmarks by incorporating equ…

Cited by 0SourceScholar
2026

USE: A Unified Model for Universal Sound Separation and Extraction

AAAI 2026technical

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require precisely specified clues to achieve optimal performance. This

Cited by 0SourcePDFScholar
2026

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension

ICML 2026poster

Existing LLM test-time scaling laws emphasize the emergence of self-reflective behaviors through extended reasoning length. Nevertheless, this vertical scaling strategy often encounters plateaus in exploration as the model becomes locked into specific thinking pattern. By shifting from depth to para…

Cited by 0SourceScholar
2025

Bitnet.cpp: Efficient Edge Inference for Ternary LLMs

ACL 2025long

The advent of 1-bit large language models (LLMs), led by BitNet b1.58, has spurred interest in ternary LLMs. Despite this, research and practical applications focusing on efficient edge inference for ternary LLMs remain scarce. To bridge this gap, we introduce Bitnet.cpp, an inference system optimiz…

2025

DDN-SLAM: Real Time Dense Dynamic Neural Implicit SLAM

RA-L 2025

SLAM systems based on NeRF have demonstrated superior performance in rendering quality and scene reconstruction for static environments compared to traditional dense SLAM. However, they encounter tracking drift and mapping errors in real-world scenarios with dynamic interferences. To address these i

Cited by 46SourceScholar
2025

Dy3DGS-SLAM: Monocular 3D Gaussian Splatting SLAM for Dynamic Environments

ICRA 2025

Current Simultaneous Localization and Mapping (SLAM) methods based on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting excel in reconstructing static 3D scenes but struggle with tracking and reconstruction in dynamic environments, such as real-world scenes with moving elements. Existing NeRF-b

Cited by 12SourceScholar
2025

Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation

IROS 2025

Zero-shot generalization across various robots, tasks and environments remains a significant challenge in robotic manipulation. Policy code generation methods use executable code to connect high-level task descriptions and low-level action sequences, leveraging the generalization capabilities of lar

Cited by 5SourcecodeScholar
2025

ST-TAR: An Efficient Spatio-Temporal Learning Framework for Traffic Accident Risk Forecasting

IJCAI 2025

Traffic accidents represent a significant concern due to their devastating consequences. The ability to predict future traffic accident risks is of key importance to accident prevention activities in transportation systems. Although existing studies have made substantial efforts to model spatio-temp

2025

STG-Avatar: Animatable Human Avatars via Spacetime Gaussian

IROS 2025

Realistic animatable human avatars from monocular videos are crucial for advancing human-robot interaction and enhancing immersive virtual experiences. While recent research on 3DGS-based human avatars has made progress, it still struggles with accurately representing detailed features of non-rigid

Cited by 6SourcecodeScholar
2025

TESTN: A Triad-Enhanced Spatio-Temporal Network for Multi-Temporal POI Relationship Inference

IJCAI 2025

Multi-temporal Point-of-Interest (POI) relationship inference aims to identify evolving relationships among locations over time, providing critical insights for location-based services. While existing studies have made substantial efforts to model relationships with custom-designed graph neural netw

2024

PREFER: Prompt Ensemble Learning via Feedback-Reflect-Refine

AAAI 2024technical

As an effective tool for eliciting the power of Large Language Models (LLMs), prompting has recently demonstrated unprecedented abilities across a variety of complex tasks. To further improve the performance, prompt ensemble has attracted substantial interest for tackling the hallucination and insta…

2024

Temporal Adaptive RGBT Tracking with Modality Prompt

AAAI 2024technical

RGBT tracking has been widely used in various fields such as robotics, surveillance processing, and autonomous driving. Existing RGBT trackers fully explore the spatial information between the template and the search region and locate the target based on the appearance matching results. However, the…

Cited by 32SourcePDFScholar
2023

Magneto: A Foundation Transformer

ICML 2023poster

A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name ''Transformers'', the above areas use different implementations for better performance, e.g., Post-LayerNorm for BERT, and Pre-LayerNorm for GPT and vision Transformers.…

Cited by 12SourcePDFScholar
2017

Amulet: Aggregating Multi-Level Convolutional Features for Salient Object Detection

ICCV 2017poster

Fully convolutional neural networks (FCNs) have shown outstanding performance in many dense labeling problems. One key pillar of these successes is mining relevant information from features in convolutional layers. However, how to better aggregate multi-level convolutional feature maps for salient o…

Cited by 1024PDFScholar
2017

Learning Uncertain Convolutional Features for Accurate Saliency Detection

ICCV 2017poster

Deep convolutional neural networks (CNNs) have delivered superior performance in many computer vision tasks. In this paper, we propose a novel deep fully convolutional network model for accurate salient object detection. The key contribution of this work is to learn deep uncertain convolutional feat…

Cited by 447PDFScholar