← Search

Haonan Zhang

18 accepted papers

2026

ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking

ICML 2026poster

Reinforcement learning (RL) has advanced LLM agents on verifiable tasks but remains challenging for open-ended tasks with vast solution spaces (e.g., complex travel planning). Lacking objective ground truth, current RL algorithms rely on reward models assigning scalar scores to individual responses.…

Cited by 0SourceScholar
2026

Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking

AAAI 2026technical

3D multi-object tracking is a critical and challenging task in the field of autonomous driving. A common paradigm relies on modeling individual object motion, e.g., Kalman filters, to predict trajectories. While effective in simple scenarios, this approach often struggles in crowded environments or

Cited by 0SourcePDFScholar
2026

Reference-Free Meta-Learning for Generalized Implicit Neural Representation in Efficient MRI Reconstruction

ICML 2026poster

Implicit Neural Representation (INR) has emerged as a powerful paradigm for continuous MRI reconstruction. However, standard unsupervised INR requires time-consuming optimization from scratch for each scan, hindering clinical deployment. This work presents IPOD, a Reference-Free Meta-Learning framew…

Cited by 0SourceScholar
2026

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models

CVPR 2026

Video Large Language Models (VLLMs) demonstrate strong video understanding but suffer from inefficiency due to redundant visual tokens. Existing pruning primary targets intra-frame spatial redundancy or prunes inside the LLM with shallow-layer overhead, yielding suboptimal spatiotemporal reduction a

Cited by 0SourcecodeScholar
2025

Bipolar Self-attention for Spiking Transformers

NeurIPS 2025spotlight

Harnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through…

Cited by 0SourceScholar
2025

D2S: Towards Efficient Sparse 3D Object Detection via Dense to Sparse Knowledge Distillation

ICASSP 2025accepted

LiDAR-based 3D object detection is widely used in high-level autonomous driving schemes. However, the cumbersome modules in most 3D detectors lead to substantial computational overhead. Despite knowledge distillation (KD) is an effective approach for compressing models, previous methods cannot be ex…

Cited by 0SourceScholar
2025

ETDE-Net: An End-to-End Time-Domain Enhancement Network for LPI Radar Signals

ICASSP 2025accepted

Low probability of intercept (LPI) radar signals are widely used in modern electromagnetic warfare due to their exceptional anti-interception capabilities. A defining characteristic of LPI radar signals is their low peak power, which makes them highly susceptible to being masked by additive white Ga…

Cited by 0SourceScholar
2025

MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

ACL 2025finding

The development of Multimodal Large Language Models (MLLMs) has seen significant progress, driven by increasing demands across various fields (e.g., multimodal agents, embodied intelligence). While model-driven approaches aim to enhance MLLM capabilities through diverse architectures, their performa…

Cited by 0SourcePDFScholar
2025

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction

ACL 2025long

Role-Playing Agents (RPAs), benefiting from large language models, is an emerging interactive AI system that simulates roles or characters with diverse personalities. However, existing methods primarily focus on mimicking dialogues among roles in textual form, neglecting the role’s voice traits (e.g…

2025

OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis

NeurIPS 2025poster

Recent advancements in omnimodal learning have significantly improved understanding and generation across images, text, and speech, yet these developments remain predominantly confined to proprietary models. The lack of high-quality omnimodal datasets and the challenges of real-time emotional speech…

Cited by 0SourcecodeScholar
2025

Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models

ACL 2025long

Despite the widespread use of Transformer-based text embedding models in NLP tasks, surprising “sticky tokens” can undermine the reliability of embeddings. These tokens, when repeatedly inserted into sentences, pull sentence similarity toward a certain value, disrupting the normal distribution of em…

2024

CaKDP: Category-aware Knowledge Distillation and Pruning Framework for Lightweight 3D Object Detection

CVPR 2024poster

Knowledge distillation (KD) possesses immense potential to accelerate the deep neural networks (DNNs) for LiDAR-based 3D detection. However in most of prevailing approaches the suboptimal teacher models and insufficient student architecture investigations limit the performance gains. To address thes…

2024

IS-DARTS: Stabilizing DARTS through Precise Measurement on Candidate Importance

AAAI 2024technical

Among existing Neural Architecture Search methods, DARTS is known for its efficiency and simplicity. This approach applies continuous relaxation of network representation to construct a weight-sharing supernet and enables the identification of excellent subnets in just a few GPU days. However, perfo…

2023

Self-supervised Graph Disentangled Networks for Review-based Recommendation

IJCAI 2023poster

User review data is considered as auxiliary information to alleviate the data sparsity problem and improve the quality of learned user/item or interaction representations in review-based recommender systems. However, existing methods usually model user-item interactions in a holistic manner and negl…

Cited by 7SourcePDFScholar
2022

A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal Retrieval

NeurIPS 2022accept

Cross-modal retrieval aims to build correspondence between multiple modalities by learning a common representation space. Typically, an image can match multiple texts semantically and vice versa, which significantly increases the difficulty of this task. To address this problem, probabilistic embedd…

2021

GLAVNet: Global-Local Audio-Visual Cues for Fine-Grained Material Recognition

CVPR 2021poster

In this paper, we aim to recognize materials with combined use of auditory and visual perception. To this end, we construct a new dataset named GLAudio that consists of both the geometry of the object being struck and the sound captured from either modal sound synthesis (for virtual objects) or real…

Cited by 8PDFScholar