← Search

Han Jiang

25 accepted papers

2026

AMap: Distilling Future Priors for Ahead-Aware Online HD Map Construction

CVPR 2026

Online High-Definition (HD) map construction is pivotal for autonomous driving. While recent approaches leverage historical temporal fusion to improve performance, we identify a critical safety flaw in this paradigm: it is inherently "spatially backward-looking." These methods predominantly enhance

Cited by 0SourceScholar
2026

Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action Localization

AAAI 2026technical

Current Zero-Shot Temporal Action Localization (ZSTAL) methods, whether training-based or training-free ones, still predominantly rely on a single, unified query to localize an entire action. This unified representation is fundamentally ill-suited for complex real-world activities, as it fails to ca

Cited by 0SourcePDFScholar
2026

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

CVPR 2026

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only weakly modeled, restricting their applicability in domains like filmmaking and v

Cited by 0SourceScholar
2026

Memory Matters: Boosting Training-Free Zero-Shot Temporal Action Localization with a Learnable Lookup Table

CVPR 2026

Zero-Shot Temporal Action Localization (ZS-TAL) aims to classify and localize actions in untrimmed videos that are unseen during training. Existing training-based ZS-TAL methods typically rely on fine-tuning models on large-scale annotated training data. This can be impractical in real-world applica

Cited by 0SourceScholar
2026

PICACO: Pluralistic In-Context Value Alignment via Total Correlation Optimization

ICML 2026poster

In-Context Learning has shown great potential for aligning Large Language Models (LLMs) with human values, helping reduce harmful outputs and accommodate diverse preferences without costly post-training, known as *In-Context Alignment* (ICA). However, LLMs' comprehension of input prompts remains agn…

Cited by 0SourceScholar
2026

SCoA: Revisiting Domain Generalized Object Detection with Style-Conditioned Adaptation

ICML 2026poster

Domain generalized object detection (DGOD) aims to train an object detector on a single source domain and generalize it to unseen target domains. Recent advances in DGOD have increasingly exploited vision foundation models (VFMs) via parameter-efficient finetuning strategies. However, existing appro…

Cited by 0SourceScholar
2026

Stability Under Scrutiny: Benchmarking Representation Paradigms for Online HD Mapping

ICLR 2026poster

As one of the fundamental intermediate modules in autonomous driving, online high-definition (HD) maps have attracted significant attention due to their cost-effectiveness and real-time capabilities. Since vehicles always cruise in highly dynamic environments, spatial displacement of onboard sensor…

Cited by 0SourcecodeScholar
2025

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation

NeurIPS 2025poster

Recently, mobile manipulation has attracted increasing attention for enabling language-conditioned robotic control in household tasks. However, existing methods still face challenges in coordinating mobile base and manipulator, primarily due to two limitations. On the one hand, they fail to explicit…

Cited by 0SourceScholar
2025

ALARM: Safe Reinforcement Learning With Reliable Mimicry for Robust Legged Locomotion

RA-L 2025

Legged robots are supposed to traverse complicated environments, which makes it challenging to design a model-based controller due to their functional complexity. Currently, using deep reinforcement learning to improve the adaptability of robots in complex scenarios has been a major research trend.

Cited by 3SourceScholar
2025

Authentic 4D Driving Simulation with a Video Generation Model

ICCV 2025poster

Simulating driving environments in 4D is crucial for developing accurate and immersive autonomous driving systems. Despite progress in generating driving scenes, challenges in transforming views and modeling the dynamics of space and time remain. To tackle these issues, we propose a fresh methodolog…

Cited by 0SourcePDFScholar
2025

Boundary-Aware Temporal Dynamic Pseudo-Supervision Pairs Generation for Zero-Shot Natural Language Video Localization

AAAI 2025technical

Zero-shot Natural Language Video Localization (NLVL) aims to automatically generate moments and corresponding pseudo queries from raw videos for the training of the localization model without any manual annotations. Existing approaches typically produce pseudo queries as simple words, which overlook…

Cited by 0SourcePDFScholar
2025

DeTAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification

ACL 2025finding

With the widespread adoption of Large Language Models (LLMs), jailbreak attacks have become an increasingly pressing safety concern. While safety-aligned LLMs can effectively defend against normal harmful queries, they remain vulnerable to such attacks. Existing defense methods primarily rely on fin…

2025

Diffusion-based Source-biased Model for Single Domain Generalized Object Detection

ICCV 2025poster

Single domain generalized object detection aims to train an object detector on a single source domain and generalize it to any unseen domain. Although existing approaches based on data augmentation exhibit promising results, they overlook domain discrepancies across multiple augmented domains, which…

Cited by 0SourcePDFScholar
2025

PanoWan: Lifting Diffusion Video Generation Models to 360$^\circ$ with Latitude/Longitude-aware Mechanisms

NeurIPS 2025poster

Panoramic video generation enables immersive 360$^\circ$ content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-q…

Cited by 0SourceScholar
2025

Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing

ICML 2025poster

*Warning: Contains harmful model outputs.* Despite significant advancements, the propensity of Large Language Models (LLMs) to generate harmful and unethical content poses critical challenges. Measuring value alignment of LLMs becomes crucial for their regulation and responsible deployment. Althoug…

Cited by 5SourcePDFScholar
2025

VIRES: Video Instance Repainting via Sketch and Text Guided Generation

CVPR 2025poster

We introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches struggle with temporal consistency and accurate alignment with the provided sketch sequence. VIRES leverages the generat…

Cited by 0SourcePDFScholar
2024

DenseKoopman: A Plug-and-Play Framework for Dense Pedestrian Trajectory Prediction

IJCAI 2024poster

Pedestrian trajectory prediction has emerged as a core component of human-robot interaction and autonomous driving. Fast and accurate prediction of surrounding pedestrians contributes to making decisions and improves safety and efficiency. However, pedestrians’ future trajectories will interact with…

2024

MARE: Multi-Aspect Rationale Extractor on Unsupervised Rationale Extraction

EMNLP 2024main

Unsupervised rationale extraction aims to extract text snippets to support model predictions without explicit rationale annotation.Researchers have made many efforts to solve this task. Previous works often encode each aspect independently, which may limit their ability to capture meaningful interna…

2024

Multi-Level Augmentation Consistency Learning and Sample Selection for Semi-Supervised Domain Generalization

ICASSP 2024accepted

Semi-supervised domain generalization (SSDG) aims to build a domain-generalized model using partially labeled data from source domains. Mainstream SSDG methods follow the augmentation consistency in FixMatch. However, the extraction of domain-invariant features may be challenging due to the absence…

Cited by 0SourceScholar
2024

Small Multi-Rotor UAV Oriented Direct Thrust Sensor Based on Lightweight Barometers

IROS 2024poster

The multirotor unmanned aerial vehicle (UAV) requires precise control over thrust output when operating in wind-disturbed environments or executing intricate flight missions. Although current commercial force sensors offer high sensitivity and accuracy, they are often heavy and costly. These charact…

Cited by 0SourceScholar
2023

Cooperative Exploration of Heterogeneous UAVs in Mountainous Environments by Constructing Steady Communication

RA-L 2023

Unmanned aerial vehicles (UAVs) must fly at low altitudes to execute certain missions when operating in complex mountainous areas. However, in these environments, UAVs lose their line-of-sight (LOS) communication with the ground station (GS) due to the obstruction of the mountains and are unable to

Cited by 8SourceScholar
2023

Large-Scale and Multi-Perspective Opinion Summarization with Diverse Review Subsets

EMNLP 2023long findings

Opinion summarization is expected to digest larger review sets and provide summaries from different perspectives. However, most existing solutions are deficient in epitomizing extensive reviews and offering opinion summaries from various angles due to the lack of designs for information selection. T…

Cited by 0SourcecodeScholar
2023

ToViLaG: Your Visual-Language Generative Model is Also An Evildoer

EMNLP 2023long main

Recent large-scale Visual-Language Generative Models (VLGMs) have achieved unprecedented improvement in multimodal image/text generation. However, these models might also generate toxic content, e.g., offensive text and pornography images, raising significant ethical risks. Despite exhaustive studie…

Cited by 0SourcecodeScholar
2022

CHAE: Fine-Grained Controllable Story Generation with Characters, Actions and Emotions

COLING 2022main

Story generation has emerged as an interesting yet challenging NLP task in recent years. Some existing studies aim at generating fluent and coherent stories from keywords and outlines; while others attempt to control the global features of the story, such as emotion, style and topic. However, these…