← Search

Hanyu Wang

22 accepted papers

2026

Aria: an Agent for Retrieval and Iterative Auto-Formalization via Dependency Graph

ICLR 2026poster

Accurate auto-formalization of theorem statements is essential for advancing automated discovery and verification of research-level mathematics, yet remains a major bottleneck for LLMs due to hallucinations, semantic mismatches, and their inability to synthesize new definitions. To tackle these issu…

Cited by 0SourceScholar
2026

NeRV-Diffusion: Diffuse Implicit Neural Representation for Video Synthesis

ICLR 2026poster

We present NeRV-Diffusion, an implicit latent video diffusion model that synthesizes videos via generating neural network weights. The generated weights can be rearranged as the parameters of a convolutional neural network, which forms an implicit neural representation (INR), and decodes into videos…

Cited by 0SourceScholar
2026

SEAP: Sparse Expert Activation Pruning Unlocks the Brainpower of Large Language Models

AAAI 2026technical

Pruning is a promising approach to reduce the high inference cost of large language models (LLMs), but it often comes at the expense of performance. Motivated by the "functional localization" theory in neuroscience, we hypothesize that LLMs contain task-specific expert activation paths, where specif

Cited by 0SourcePDFScholar
2026

TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles

ICASSP 2026poster

As the application of Large Language Models (LLMs) expands, the demand for reliable evaluations increases. Existing LLM evaluation benchmarks primarily rely on static datasets, making it challenging to assess model performance in dynamic interactions with users. Moreover, these benchmarks often depe…

Cited by 0SourcePDFScholar
2025

Integrating Large Language Models and Möbius Group Transformations for Temporal Knowledge Graph Embedding on the Riemann Sphere

AAAI 2025technical

The significance of Temporal Knowledge Graphs (TKGs) in Artificial Intelligence (AI) lies in their capacity to incorporate time-dimensional information, support complex reasoning and prediction, optimize decision-making processes, enhance the accuracy of recommendation systems, promote multimodal da…

Cited by 0SourcePDFScholar
2025

LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

ICLR 2025oral

We present LARP, a novel video tokenizer designed to overcome limitations in current video tokenization methods for autoregressive (AR) generative models. Unlike traditional patchwise tokenizers that directly encode local visual patches into discrete tokens, LARP introduces a holistic tokenization s…

2025

MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation System

ACL 2025long

Retrieval-Augmented Generation (RAG), while serving as a viable complement to large language models (LLMs), often overlooks the crucial aspect of text chunking within its pipeline. This paper initially introduces a dual-metric evaluation method, comprising Boundary Clarity and Chunk Stickiness, to e…

2025

Retrieval-Augmented Multilingual Citation Generation

ICASSP 2025accepted

Retrieval-augmented citation generation (RACG) helps users trust the large language model output by retrieving evidence from reliable sources. However, most current RACG research focuses on single-language tasks, particularly in English, and overlooks the need for cross-lingual evidence retrieval an…

Cited by 0SourceScholar
2025

SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model

ACL 2025long

The indexing-retrieval-generation paradigm of retrieval-augmented generation (RAG) has been highly successful in solving knowledge-intensive tasks by integrating external knowledge into large language models (LLMs). However, the incorporation of external and unverified knowledge increases the vulner…

2025

TruthFlow: Truthful LLM Generation via Representation Flow Correction

ICML 2025poster

Large language models (LLMs) are known to struggle with consistently generating truthful responses. While various representation intervention techniques have been proposed, these methods typically apply a universal representation correction vector to all input queries, limiting their effectiveness a…

Cited by 0SourcePDFScholar
2025

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

NeurIPS 2025poster

This paper presents a multimodal framework that attempts to unify visual understanding and generation within a shared discrete semantic representation. At its core is the Text-Aligned Tokenizer (TA-Tok), which converts images into discrete tokens using a text-aligned codebook projected from a large…

Cited by 0SourceScholar
2025

When Sparse Graph Representation Learning Falls into Domain Shift: Feature Augmentation for Cross-Domain Graph Meta-Learning

ICASSP 2025accepted

Graph Meta-learning methods have improved the performance of few-shot node classification by means of applying meta-learning to the data in non-Euclidean domains. However, most works focus on adopting a single domain, ignoring the fact that tasks in various domains may be distinct, which can cause o…

Cited by 0SourceScholar
2024

Controlled Text Generation for Large Language Model with Dynamic Attribute Graphs

ACL 2024findings

Controlled Text Generation (CTG) aims to produce texts that exhibit specific desired attributes. In this study, we introduce a pluggable CTG framework for Large Language Models (LLMs) named Dynamic Attribute Graphs-based controlled text generation (DATG). This framework utilizes an attribute scorer…

2024

PathRL: An End-to-End Path Generation Method for Collision Avoidance via Deep Reinforcement Learning

ICRA 2024poster

Robot navigation using deep reinforcement learning (DRL) has shown great potential in improving the performance of mobile robots. Nevertheless, most existing DRL-based navigation methods primarily focus on training a policy that directly commands the robot with low-level controls, like linear and an…

Cited by 8SourceScholar
2024

Solving General Noisy Inverse Problem via Posterior Sampling: A Policy Gradient Viewpoint

AISTATS 2024poster

Solving image inverse problems (e.g., super-resolution and inpainting) requires generating a high fidelity image that matches the given input (the low-resolution image or the masked image). By using the input image as guidance, we can leverage a pretrained diffusion generative model to solve a wide…

2023

Chop & Learn: Recognizing and Generating Object-State Compositions

ICCV 2023poster

Recognizing and generating object-state compositions has been a challenging task, especially when generalizing to unseen compositions. In this paper, we study the task of cutting objects in different styles and the resulting object state changes. We propose a new benchmark suite Chop & Learn, to acc…

Cited by 18PDFcodeScholar
2023

NIRVANA: Neural Implicit Representations of Videos With Adaptive Networks and Autoregressive Patch-Wise Modeling

CVPR 2023poster

Implicit Neural Representations (INR) have recently shown to be powerful tool for high-quality video compression. However, existing works are limiting as they do not explicitly exploit the temporal redundancy in videos, leading to a long encoding time. Additionally, these methods have fixed architec…

2023

Towards Scalable Neural Representation for Diverse Videos

CVPR 2023poster

Implicit neural representations (INR) have gained increasing attention in representing 3D scenes and images, and have been recently applied to encode videos (e.g., NeRV, E-NeRV). While achieving promising results, existing INR-based methods are limited to encoding a handful of short videos (e.g., se…

Cited by 45SourcePDFScholar
2021

Motion-Aware Robotic 3D Ultrasound

ICRA 2021poster

Robotic three-dimensional (3D) ultrasound (US) imaging has been employed to overcome the drawbacks of traditional US examinations, such as high inter-operator variability and lack of repeatability. However, object movement remains a challenge as unexpected motion decreases the quality of the 3D comp…

Cited by 29SourceScholar
2021

NeRV: Neural Representations for Videos

NeurIPS 2021poster

We propose a novel neural representation for videos (NeRV) which encodes videos in neural networks. Unlike conventional representations that treat videos as frame sequences, we represent videos as neural networks taking frame index as input. Given a frame index, NeRV outputs the corresponding RGB i…

2018

Learning 3D Keypoint Descriptors for Non-Rigid Shape Matching

ECCV 2018poster

In this paper, we present a novel deep learning framework that derives discriminative local descriptors for 3D surface shapes. In contrast to previous convolutional neural networks (CNNs) that rely on rendering multi-view images or extracting intrinsic shape properties, we parameterize the multi-sca…

Cited by 49SourcePDFScholar