← Search

Yang Tang

18 accepted papers

2026

Aligning Deep Implicit Preferences by Learning to Reason Defensively

ICLR 2026poster

Personalized alignment is crucial for enabling Large Language Models (LLMs) to engage effectively in user-centric interactions. However, current methods face a dual challenge: they fail to infer users' deep implicit preferences (including unstated goals, semantic context and risk tolerances), and th…

Cited by 0SourcecodeScholar
2026

HCPO: Hierarchical Conductor-Based Policy Optimization in Multi-Agent Reinforcement Learning

AAAI 2026technical

In cooperative Multi-Agent Reinforcement Learning (MARL), efficient exploration is crucial for optimizing the performance of joint policy. However, existing methods often update joint policies via independent agent exploration, without coordination among agents, which inherently constrains the expre

Cited by 0SourcePDFScholar
2026

Olbedo: An Albedo and Shading Aerial Dataset for Large-Scale Outdoor Environments

CVPR 2026

Intrinsic image decomposition (IID) of outdoor scenes is crucial for relighting, editing, and understanding large-scale environments, but progress has been limited by the lack of real-world datasets with reliable albedo and shading supervision. We introduce Olbedo, a large-scale aerial dataset for o

Cited by 0SourcecodeScholar
2026

ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards

ICML 2026poster

Search agents powered by Large Language Models have demonstrated significant potential in tackling knowledge-intensive tasks. Reinforcement learning has emerged as a powerful paradigm for training these agents to perform complex, multi-step reasoning. However, prior RL-based methods often rely on sp…

Cited by 0SourceScholar
2026

SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection

CVPR 2026

Existing monocular 3D detectors typically tame the pronounced nonlinear regression of 3D bounding box through decoupled prediction paradigm, which employs multiple branches to estimate geometric center, depth, dimensions, and rotation angle separately.Although this decoupling strategy simplifies the

Cited by 0SourceScholar
2026

Scalable and General Whole-Body Control for Cross-Humanoid Locomotion

ICML 2026poster

Learning-based whole-body controllers have become a key driver for humanoid robots, yet most existing approaches require robot-specific training. In this paper, we study the problem of cross-embodiment humanoid control and show that a single policy can robustly generalize across a wide range of huma…

Cited by 0SourceScholar
2026

Universal Skeleton Understanding via Differentiable Rendering and MLLMs

ICML 2026poster

Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet remain confined to their native modalities and cannot directly process structured, non-visual data such as human skeletons. Existing methods either compress skeleton dynamics into lossy feature vectors for text al…

Cited by 0SourceScholar
2025

Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings

EMNLP 2025

Large language models (LLMs) have recently demonstrated excellent performance in text embedding tasks. Previous work usually use LoRA to fine-tune existing LLMs, which are limited by the data and training gap between LLMs and embedding models. In this work, we introduce Conan-embedding-v2, a new 1.4

Cited by 0SourcePDFScholar
2024

Self-Supervised Monocular Depth Estimation in the Dark: Towards Data Distribution Compensation

IJCAI 2024poster

Nighttime self-supervised monocular depth estimation has received increasing attention in recent years. However, using night images for self-supervision is unreliable because the photometric consistency assumption is usually violated in the videos taken under complex lighting conditions. Even with d…

Cited by 2SourcePDFScholar
2024

Unsupervised Modality Adaptation with Text-to-Image Diffusion Models for Semantic Segmentation

NeurIPS 2024poster

Despite their success, unsupervised domain adaptation methods for semantic segmentation primarily focus on adaptation between image domains and do not utilize other abundant visual modalities like depth, infrared and event. This limitation hinders their performance and restricts their application in…

2023

CMDA: Cross-Modality Domain Adaptation for Nighttime Semantic Segmentation

ICCV 2023poster

Most nighttime semantic segmentation studies are based on domain adaptation approaches and image input. However, limited by the low dynamic range of conventional cameras, images fail to capture structural details and boundary information in low-light conditions. Event cameras, as a new form of visio…

Cited by 38PDFcodeScholar
2023

GasMono: Geometry-Aided Self-Supervised Monocular Depth Estimation for Indoor Scenes

ICCV 2023poster

This paper tackles the challenges of self-supervised monocular depth estimation in indoor scenes caused by large rotation between frames and low texture. We ease the learning process by obtaining coarse camera poses from monocular sequences through multi-view geometry to deal with the former. Howeve…

Cited by 25PDFcodeScholar
2023

Multi-Dimensional Deformable Object Manipulation Using Equivariant Models

IROS 2023poster

Manipulating deformable objects, such as ropes (1D), fabrics (2D), and bags (3D), poses a significant challenge in robotics research due to their high degree of freedom in physical state and nonlinear dynamics. Compared with single-dimensional deformable objects, multi-dimensional object manipulatio…

Cited by 1SourceScholar
2023

VL-SAT: Visual-Linguistic Semantics Assisted Training for 3D Semantic Scene Graph Prediction in Point Cloud

CVPR 2023highlight

The task of 3D semantic scene graph (3DSSG) prediction in the point cloud is challenging since (1) the 3D point cloud only captures geometric structures with limited semantics compared to 2D images, and (2) long-tailed relation distribution inherently hinders the learning of unbiased prediction. Sin…

2022

Rethinking Individual Global Max in Cooperative Multi-Agent Reinforcement Learning

NeurIPS 2022accept

In cooperative multi-agent reinforcement learning, centralized training and decentralized execution (CTDE) has achieved remarkable success. Individual Global Max (IGM) decomposition, which is an important element of CTDE, measures the consistency between local and joint policies. The majority of IGM…

Cited by 34SourcePDFScholar
2021

PIAP-DF: Pixel-Interested and Anti Person-Specific Facial Action Unit Detection Net With Discrete Feedback Learning

ICCV 2021poster

Facial Action Units (AUs) are of great significance in communication. Automatic AU detection can improve the understanding of psychological conditions and emotional status. Recently, several deep learning methods have been proposed to detect AUs automatically. However, several challenges, such as po…

Cited by 33PDFScholar