← Search

Jiahao Yang

14 accepted papers

2026

Expert-Inspired Multi-Agent Coordination for Multi-Objective Molecular Optimization

AAAI 2026technical

Multi-objective molecular optimization is a fundamental yet inherently challenging task in drug discovery, as it requires simultaneously optimizing multiple, often conflicting, molecular properties. Although recent deep learning methods have shown promise, they often lack objective-specific special

Cited by 0SourcePDFScholar
2026

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation

CVPR 2026

Despite significant progress in Vision-Language Navigation (VLN), existing approaches still rely on dense RGB videos that produce excessive patch tokens and lack explicit spatial structure, resulting in substantial computational overhead and limited spatial reasoning. To address these issues, we int

Cited by 0SourceScholar
2026

Zero-source LLM Hallucination Detection with Human-like Criteria Probing

ICML 2026poster

Large language models (LLMs) often hallucinate by generating factually incorrect or unfaithful content, posing significant risks to their safe use. Detecting such hallucinations is particularly challenging under the zero-source constraint, where no model internals or external references are availabl…

Cited by 0SourceScholar
2026

ZeroSiam: An Efficient Siamese for Test-Time Entropy Optimization without Collapse

ICLR 2026poster

Test-time entropy minimization helps adapt a model to novel environments and incentivize its reasoning capability, unleashing the model's potential during inference by allowing it to evolve and improve in real-time using its own predictions. However, pure test-time entropy minimization can favor non…

Cited by 0SourceScholar
2025

Multi-Objective Molecular Design Through Learning Latent Pareto Set

AAAI 2025technical

Molecular design inherently involves the optimization of multiple conflicting objectives, such as enhancing bio-activity and ensuring synthesizability. Evaluating these objectives often requires resource-intensive computations or physical experiments. Current molecular design methodologies typically…

2025

Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection

NeurIPS 2025spotlight

AI-generated videos have achieved near-perfect visual realism (e.g., Sora), urgently necessitating reliable detection mechanisms. However, detecting such videos faces significant challenges in modeling high-dimensional spatiotemporal dynamics and identifying subtle anomalies that violate physical la…

Cited by 0SourcecodeScholar
2025

STaR: Seamless Spatial-Temporal Aware Motion Retargeting with Penetration and Consistency Constraints

ICCV 2025poster

Motion retargeting seeks to faithfully replicate the spatio-temporal motion characteristics of a source character onto a target character with a different body shape. Apart from motion semantics preservation, ensuring geometric plausibility and maintaining temporal consistency are also crucial for e…

2024

Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean Discrepancy

ICLR 2024poster

Large language models (LLMs) such as ChatGPT have exhibited remarkable performance in generating human-like texts. However, machine-generated texts (MGTs) may carry critical risks, such as plagiarism issues and hallucination information. Therefore, it is very urgent and important to detect MGTs in m…

2024

Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation

CVPR 2024highlight

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. At each navigation step the agent selects from possible candidate locations and then makes the move. For better navigation planning the lookahead exp…

2024

Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation

CoRL 2024poster

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location in 3D environments following the natural language instruction. In this field, the agent is usually trained and evaluated in the navigation simulators, lacking effective approaches for sim-to-real transfer. The VL…

Cited by 11SourcecodeScholar
2023

Detecting Adversarial Data by Probing Multiple Perturbations Using Expected Perturbation Score

ICML 2023poster

Adversarial detection aims to determine whether a given sample is an adversarial one based on the discrepancy between natural and adversarial distributions. Unfortunately, estimating or comparing two data distributions is extremely difficult, especially in high-dimension spaces. Recently, the gradie…

2023

GridMM: Grid Memory Map for Vision-and-Language Navigation

ICCV 2023poster

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. To represent the previously visited environment, most approaches for VLN implement memory using recurrent states, topological maps, or top-down seman…

Cited by 59PDFcodeScholar
2023

KERM: Knowledge Enhanced Reasoning for Vision-and-Language Navigation

CVPR 2023poster

Vision-and-language navigation (VLN) is the task to enable an embodied agent to navigate to a remote location following the natural language instruction in real scenes. Most of the previous approaches utilize the entire features or object-centric features to represent navigable candidates. However,…

2021

What If We Could Not See? Counterfactual Analysis for Egocentric Action Anticipation

IJCAI 2021poster

Egocentric action anticipation aims at predicting the near future based on past observation in first-person vision. While future actions may be wrongly predicted due to the dataset bias, we present a counterfactual analysis framework for egocentric action anticipation (CA-EAA) to enhance the capacit…

Cited by 16SourcePDFScholar