← Search

Yuhang He

36 accepted papers

2026

Aurelius: Relation Aware Text-to-Audio Generation At Scale

ICLR 2026poster

We present Aurelius, a new framework that enables relation aware text-to-audio (TTA) generation research at scale. Given the lack of essential audio event and relation corpora, \emph{Aurelius} contributes a large-scale audio event corpus \emph{AudioEventSet} and another large-scale relation corpus \…

Cited by 0SourcecodeScholar
2026

From Representation to Action: A Unified Laplacian Framework for Spatial Representation and Path Planning

ICML 2026poster

Navigation in complex environments relies on internal spatial representations that guide action. While the brain employs a diverse repertoire of spatial tuning cells—including grid, place, and head-direction cells—a normative theory linking these static neural codes to the dynamic process of navigat…

Cited by 0SourceScholar
2026

GOAL: Geometrically Optimal Alignment for Continual Generalized Category Discovery

AAAI 2026technical

Continual Generalized Category Discovery (C-GCD) requires identifying novel classes from unlabeled data while retaining knowledge of known classes over time. Existing methods typically update classifier weights dynamically, resulting in forgetting and inconsistent feature alignment. We propose GOAL,

Cited by 0SourcePDFScholar
2026

Is Parameter Isolation Better for Prompt-Based Continual Learning?

CVPR 2026

Prompt-based continual learning methods effectively mitigate catastrophic forgetting. However, most existing methods assign a fixed set of prompts to each task, completely isolating knowledge across tasks and resulting in suboptimal parameter utilization. To address this, we consider the practical n

Cited by 0SourceScholar
2026

Learning Like Humans: Analogical Concept Learning for Generalized Category Discovery

CVPR 2026

Generalized Category Discovery (GCD) seeks to uncover novel categories in unlabeled data while preserving recognition of known categories, yet prevailing visual-only pipelines and the loose coupling between supervised learning and discovery often yield brittle boundaries on fine-grained, look-alike

Cited by 0SourcecodeScholar
2026

MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis

CVPR 2026

While text-to-video (T2V) generation has achieved remarkable progress in photorealism, generating intent-aligned videos that faithfully obey physics principles remains a core challenge. In this work, we systematically study Newtonian motion-controlled text-to-video generation and evaluation, emphasi

Cited by 0SourcecodeScholar
2026

Shared & Domain Self-Adaptive Experts with Frequency-Aware Discrimination for Continual Test-Time Adaptation

AAAI 2026technical

This paper focuses on the Continual Test-Time Adaptation (CTTA) task, aiming to enable an agent to continuously adapt to evolving target domains while retaining previously acquired domain knowledge for effective reuse when those domains reappear. Existing shared-parameter paradigms struggle to balan

Cited by 0SourcePDFScholar
2025

Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery

NeurIPS 2025poster

Generalized Category Discovery (GCD) focuses on classifying known categories while simultaneously discovering novel categories from unlabeled data. However, previous GCD methods face challenges due to inconsistent optimization objectives and category confusion. This leads to feature overlap and ulti…

Cited by 0SourceScholar
2025

DiffRefine: Diffusion-based Proposal Specific Point Cloud Densification for Cross-Domain Object Detection

ICCV 2025poster

The robustness of 3D object detection in large-scale outdoor point clouds degrades significantly when deployed in an unseen environment due to domain shifts. To minimize the domain gap, existing works on domain adaptive detection focuses on several factors, including point density, object shape and…

Cited by 0SourcePDFScholar
2025

DualCP: Rehearsal-Free Domain-Incremental Learning via Dual-Level Concept Prototype

AAAI 2025technical

Domain-Incremental Learning (DIL) enables vision models to adapt to changing conditions in real-world environments while maintaining the knowledge acquired from previous domains. Given privacy concerns and training time, Rehearsal-Free DIL (RFDIL) is more practical. Inspired by the incremental cogni…

Cited by 0SourcePDFScholar
2025

Dynamic Integration of Task-Specific Adapters for Class Incremental Learning

CVPR 2025poster

Non-exemplar Class Incremental Learning (NECIL) enables models to continuously acquire new classes without retraining from scratch and storing old task exemplars, addressing privacy and storage issues. However, the absence of data from earlier tasks exacerbates the challenge of catastrophic forgetti…

Cited by 2SourcePDFScholar
2025

RiTTA: Modeling Event Relations in Text-to-Audio Generation

EMNLP 2025

Existing text-to-audio (TTA) generation methods have neither systematically explored audio event relation modeling, nor proposed any new framework to enhance this capability. In this work, we systematically study audio event relation modeling in TTA generation models. We first establish a benchmark

2025

SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

NeurIPS 2025poster

Accurate spatial reasoning in outdoor environments—covering geometry, object pose, and inter-object relationships—is fundamental to downstream tasks such as mapping, motion forecasting, and high-level planning in autonomous driving. We introduce SURDS, a large-scale benchmark designed to systematica…

Cited by 0SourcecodeScholar
2025

SuLoRA: Subspace Low-Rank Adaptation for Parameter-Efficient Fine-Tuning

ACL 2025finding

As the scale of large language models (LLMs) grows and natural language tasks become increasingly diverse, Parameter-Efficient Fine-Tuning (PEFT) has become the standard paradigm for fine-tuning LLMs. Among PEFT methods, LoRA is widely adopted for not introducing additional inference overhead. Howev…

Cited by 0SourcePDFScholar
2024

BC-Prover: Backward Chaining Prover for Formal Theorem Proving

EMNLP 2024main

Despite the remarkable progress made by large language models in mathematical reasoning, interactive theorem proving in formal logic still remains a prominent challenge. Previous methods resort to neural models for proofstep generation and search. However, they suffer from exploring possible proofst…

Cited by 0SourcePDFScholar
2024

Bridge the Modality and Capability Gaps in Vision-Language Model Selection

NeurIPS 2024poster

Vision Language Models (VLMs) excel in zero-shot image classification by pairing images with textual category names. The expanding variety of Pre-Trained VLMs enhances the likelihood of identifying a suitable VLM for specific tasks. To better reuse the VLM resource and fully leverage its potential o…

2024

DYSON: Dynamic Feature Space Self-Organization for Online Task-Free Class Incremental Learning

CVPR 2024poster

In this paper we focus on a challenging Online Task-Free Class Incremental Learning (OTFCIL) problem. Different from the existing methods that continuously learn the feature space from data streams we propose a novel compute-and-align paradigm for the OTFCIL. It first computes an optimal geometry i.…

2024

Decomposing Argumentative Essay Generation via Dialectical Planning of Complex Reasoning

ACL 2024findings

Argumentative Essay Generation (AEG) is a challenging task in computational argumentation, where detailed logical reasoning and effective rhetorical skills are essential.Previous methods on argument generation typically involve planning prior to generation.However, the planning strategies in these m…

Cited by 1SourcePDFScholar
2024

Evolving Parameterized Prompt Memory for Continual Learning

AAAI 2024technical

Recent studies have demonstrated the potency of leveraging prompts in Transformers for continual learning (CL). Nevertheless, employing a discrete key-prompt bottleneck can lead to selection mismatches and inappropriate prompt associations during testing. Furthermore, this approach hinders adaptive…

2024

Non-Exemplar Domain Incremental Learning via Cross-Domain Concept Integration

ECCV 2024poster

"Existing approaches to Domain Incremental Learning (DIL) address catastrophic forgetting by storing and rehearsing exemplars from old domains. However, exemplar-based solutions are not always viable due to data privacy concerns or storage limitations. Therefore, Non-Exemplar Domain Incremental Lear…

2024

Non-exemplar Domain Incremental Object Detection via Learning Domain Bias

AAAI 2024technical

Domain incremental object detection (DIOD) aims to gradually learn a unified object detection model from a dataset stream composed of different domains, achieving good performance in all encountered domains. The most critical obstacle to this goal is the catastrophic forgetting problem, where the pe…

2024

Projecting Points to Axes: Oriented Object Detection via Point-Axis Representation

ECCV 2024oral

"This paper introduces the point-axis representation for oriented object detection, as depicted in aerial images in Figure ??, emphasizing its flexibility and geometrically intuitive nature with two key components: points and axes. 1) Points delineate the spatial extent and contours of objects, prov…

Cited by 5SourcePDFScholar
2024

Prompt-Agnostic Adversarial Perturbation for Customized Diffusion Models

NeurIPS 2024poster

Diffusion models have revolutionized customized text-to-image generation, allowing for efficient synthesis of photos from personal data with textual descriptions. However, these advancements bring forth risks including privacy breaches and unauthorized replication of artworks. Previous researches pr…

Cited by 5SourcePDFScholar
2024

SoundCount: Sound Counting from Raw Audio with Dyadic Decomposition Neural Network

AAAI 2024technical

In this paper, we study an underexplored, yet important and challenging problem: counting the number of distinct sounds in raw audio characterized by a high degree of polyphonicity. We do so by systematically proposing a novel end-to-end trainable neural network~(which we call DyDecNet, consisting o…

Cited by 2SourcePDFScholar
2024

Towards Learning Group-Equivariant Features for Domain Adaptive 3D Detection

NeurIPS 2024poster

The performance of 3D object detection in large outdoor point clouds deteriorates significantly in an unseen environment due to the inter-domain gap. To address these challenges, most existing methods for domain adaptation harness self-training schemes and attempt to bridge the gap by focusing on a…

Cited by 0SourcePDFScholar
2023

DKT: Diverse Knowledge Transfer Transformer for Class Incremental Learning

CVPR 2023poster

Deep neural networks suffer from catastrophic forgetting in class incremental learning, where the classification accuracy of old classes drastically deteriorates when the networks learn the knowledge of new classes. Many works have been proposed to solve the class incremental learning problem. Howev…

Cited by 14SourcePDFScholar
2023

Knowledge Restore and Transfer for Multi-Label Class-Incremental Learning

ICCV 2023poster

Current class-incremental learning research mainly focuses on single-label classification tasks while multi-label class-incremental learning (MLCIL) with more practical application scenarios is rarely studied. Although there have been many anti-forgetting methods to solve the problem of catastrophic…

Cited by 18PDFcodeScholar
2023

Metric-Free Exploration for Topological Mapping by Task and Motion Imitation in Feature Space

RSS 2023poster

We propose DeepExplorer, a simple and lightweight metric-free exploration method for topological mapping of unknown environments. It performs task and motion planning (TAMP) entirely in image feature space. The task planner is a recurrent network using the latest image observation sequence to halluc…

2023

Multi-body SE(3) Equivariance for Unsupervised Rigid Segmentation and Motion Estimation

NeurIPS 2023poster

A truly generalizable approach to rigid segmentation and motion estimation is fundamental to 3D understanding of articulated objects and moving scenes. In view of the closely intertwined relationship between segmentation and motion estimates, we present an SE(3) equivariant architecture and a traini…

2023

SoundSynp: Sound Source Detection from Raw Waveforms with Multi-Scale Synperiodic Filterbanks

AISTATS 2023poster

We propose synperiodic filter banks, a novel multi-scale learnable filter bank construction strategy that all filters are synchronized by their rotating periodicity. By synchronizing in a certain periodicity, we naturally get filters whose temporal length are reduced if they carry higher frequency r…

Cited by 8SourcePDFScholar
2022

A Generative Model for End-to-End Argument Mining with Reconstructed Positional Encoding and Constrained Pointer Mechanism

EMNLP 2022main

Argument mining (AM) is a challenging task as it requires recognizing the complex argumentation structures involving multiple subtasks.To handle all subtasks of AM in an end-to-end fashion, previous works generally transform AM into a dependency parsing task.However, such methods largely require com…

Cited by 7SourcePDFScholar
2021

Error-Aware Density Isomorphism Reconstruction for Unsupervised Cross-Domain Crowd Counting

AAAI 2021technical

This paper focuses on the unsupervised domain adaptation problem for video-based crowd counting, in which we use labeled data as source domain and unlabelled video data as target domain. It is challenging as there is a huge gap between the source and the target domain and no annotations of samples a…

2021

SoundDet: Polyphonic Moving Sound Event Detection and Localization from Raw Waveform

ICML 2021spotlight

We present a new framework SoundDet, which is an end-to-end trainable and light-weight framework, for polyphonic moving sound event detection and localization. Prior methods typically approach this problem by preprocessing raw waveform into time-frequency representations, which is more amenable to p…

Cited by 36SourcePDFScholar
2019

Real-Time Vehicle Detection from Short-range Aerial Image with Compressed MobileNet

ICRA 2019poster

Vehicle detection from short-range aerial image faces challenges including vehicle blocking, irrelevant object interference, motion blurring, color variation etc., leading to the difficulty to achieve high detection accuracy and real-time detection speed. In this paper, benefiting from the recent de…

Cited by 19SourceScholar