← Search

YE TIAN

50 accepted papers

2026

A Relative Error-Based Evaluation Framework of Heterogeneous Treatment Effect Estimators

ICLR 2026poster

While significant progress has been made in heterogeneous treatment effect (HTE) estimation, the evaluation of HTE estimators remains underdeveloped. In this article, we propose a robust evaluation framework based on relative error, which quantifies performance differences between two HTE estimators…

Cited by 0SourcecodeScholar
2026

Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

ICML 2026poster

Text-to-Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a single-pass generation paradigm limits their ability to handle complex prompts requiring iterative refinement. To enable multi-round Reflective Visual …

Cited by 0SourceScholar
2026

FLARE: A Failure-Aware Framework for Autonomous Correction and Recovery in Visual-Language Robotic Manipulation

CVPR 2026

Vision-Language-Action Models (VLAs) have demonstrated significant promise in generalizing to complex, long-horizon robotic manipulation tasks. However, their performance remains brittle, as they are typically trained on trajectory-monotonic, failure-free demonstrations. This reliance on "perfect" d

Cited by 0SourceScholar
2026

Grasp Any Region: Prompting MLLM to Understand the Dense World

ICLR 2026poster

While Multimodal Large Language Models (MLLMs) excel at holistic understanding, they struggle with the dense world, i.e., complex scenes requiring fine-grained analysis of intricate details and object inter-relationships. Region-level MLLMs have been a promising step. However, previous attempts are…

Cited by 0SourcecodeScholar
2026

Large Vision–Language Models Get Lost in Attention

ICML 2026poster

Despite the rapid evolution of training paradigms, the decoder backbone of large vision--language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer architecture. Therefore, deciphering the distinct roles of internal modules is critical for understanding model mechani…

Cited by 0SourceScholar
2026

Parallel Multimodal Diffusion Language Models for Thinking-Aware Editing and Generation

ICLR 2026poster

While thinking-aware generation aims to improve performance on complex tasks, we identify a critical failure mode where existing sequential, autoregressive approaches can paradoxically degrade performance due to error propagation. To systematically analyze this issue, we propose ParaBench, a new be…

Cited by 0SourcecodeScholar
2026

PocketLLM: Ultimate Compression of Large Language Models via Meta Networks

AAAI 2026technical

As Large Language Models (LLMs) continue to grow in size, storing and transmitting them on edge devices becomes increasingly challenging. Traditional methods like quantization and pruning struggle to achieve extreme compression of LLMs without sacrificing accuracy. In this paper, we introduce Pocket

Cited by 0SourcePDFScholar
2026

ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response

ICML 2026poster

While passive agents merely follow instructions, proactive agents align with higher-level objectives, such as assistance and safety by continuously monitoring the environment to determine when and how to act. However, developing proactive agents is hindered by the lack of specialized resources. To a…

Cited by 0SourcecodeScholar
2026

Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models

ICLR 2026poster

The extension of diffusion models to language tasks has shown promising results, but their post-training methods remain largely unexplored. We highlight the importance of aligning a diffusion language model’s preference-inference trajectory with its post-training objective. To this end, we propose T…

Cited by 0SourcecodeScholar
2026

SAMTok: Representing Any Mask with Two Words

CVPR 2026

Pixel-wise capabilities are essential for building interactive intelligent systems. However, pixel-wise multi-modal LLMs (MLLMs) remain difficult to scale due to complex region-level encoders, specialized segmentation decoders, and incompatible training objectives. To address these challenges, we pr

Cited by 0SourcecodeScholar
2026

VMoBA: Mixture-of-Block Attention for Video Diffusion Models

ICLR 2026poster

The quadratic complexity of full attention mechanisms poses a significant bottleneck for Video Diffusion Models (VDMs) aiming to generate long-duration, high-resolution videos. While various sparse attention methods have been proposed, many are designed as training-free inference accelerators or do…

Cited by 0SourcecodeScholar
2025

A Near-Field 3D Parameter Estimation Method Based on a Symmetric Enhanced Nested Array

ICASSP 2025accepted

In this paper, a high-precision three-dimensional (3-D) near-field (NF) localization method is proposed under an underdetermined case based on a symmetric enhanced nested array (SENA). Firstly, the symmetry of the array and the fourth-order cumulant (FOC) are utilized to construct the equivalent vir…

Cited by 0SourceScholar
2025

CURE: Co-Evolving Coders and Unit Testers via Reinforcement Learning

NeurIPS 2025spotlight

Mathematical reasoning in large language models has been successfully incentivized through reinforcement learning with verifiable rewards, leading to improved one-shot precision. In this work, we turn our focus to the coding domain. Beyond one-shot precision, we highlight unit test generation as ano…

Cited by 0SourceScholar
2025

DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes

EMNLP 2025

Large Vision-Language Models (LVLMs) have achieved significant progress in tasks like visual question answering and document understanding. However, their potential to comprehend embodied environments and navigate within them remains underexplored. In this work, we first study the challenge of open-

Cited by 0SourcePDFScholar
2025

Don’t Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls

ACL 2025long

Recent advancements in tree search algorithms guided by verifiers have significantly enhanced the reasoning capabilities of large language models (LLMs), but at the cost of increased computational resources. In this work, we identify two key challenges contributing to this inefficiency: over-explora…

2025

Endowing Interpretability for Neural Cognitive Diagnosis by Efficient Kolmogorov-Arnold Networks

IJCAI 2025

Cognitive diagnosis is crucial for intelligent education because of its ability to reveal students' proficiency in knowledge concepts. Although neural network-based neural cognitive diagnosis models (CDMs) have exhibited significantly better performance than traditional models, neural cognitive diag

2025

Fourth-Order Cumulant Based 3-D Near-Field Underdetermined Parameter Estimation With Exact Spatial Propagation Model

ICASSP 2025accepted

Based on the exact spherical wavefront model, an under-determined estimation method for three-dimensional (3-D) parameters of near-field (NF) sources using L-shaped nested arrays is proposed, referred to as the cumulant algorithm. This algorithm leverages the temporal-spatial domain cumulants of NF…

Cited by 0SourceScholar
2025

GLTW: Joint Improved Graph Transformer and LLM via Three-Word Language for Knowledge Graph Completion

ACL 2025finding

Knowledge Graph Completion (KGC), which aims to infer missing or incomplete facts, is a crucial task for KGs. However, integrating the vital structural information of KGs into Large Language Models (LLMs) and outputting predictions deterministically remains challenging. To address this, we propose a…

Cited by 0SourcePDFScholar
2025

HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation

NeurIPS 2025poster

The remarkable success of the autoregressive paradigm has made significant advancement in Multimodal Large Language Models (MLLMs), with powerful models like Show-o, Transfusion and Emu3 made notable strides in unified image understanding and generation. For the first time, we uncover a common pheno…

Cited by 0SourcecodeScholar
2025

Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning

ICLR 2025oral

Reinforcement Learning with Human Feedback (RLHF) has achieved great success in aligning large language models (LLMs) with human preferences. Prevalent RLHF approaches are reward-based, following the Bradley-Terry (BT) model assumption, which may not fully capture the complexity of human preferences…

Cited by 4SourcePDFScholar
2025

LiteSearch: Efficient Tree Search with Dynamic Exploration Budget for Math Reasoning

AAAI 2025technical

Recent research suggests that tree search algorithms (e.g. Monte Carlo Tree Search) can dramatically boost LLM performance on complex mathematical reasoning tasks. However, they often require more than 10 times the computational resources of greedy decoding due to wasteful search strategies, making…

Cited by 0SourcePDFScholar
2025

MMaDA: Multimodal Large Diffusion Language Models

NeurIPS 2025poster

We introduce MMaDA, a novel class of multimodal diffusion foundation models designed to achieve superior performance across diverse domains such as textual reasoning, multimodal understanding, and text-to-image generation. The approach is distinguished by three key innovations: (i) MMaDA adopts a un…

Cited by 0SourcecodeScholar
2025

Multi-Agent Collaboration via Cross-Team Orchestration

ACL 2025finding

Large Language Models (LLMs) have significantly impacted various domains, especially through organized LLM-driven autonomous agents. A representative scenario is in software development, where agents can collaborate in a team like humans, following predefined phases to complete sub-tasks sequentiall…

2025

Multi-Agent Collaboration via Evolving Orchestration

NeurIPS 2025poster

Large language models (LLMs) have achieved remarkable results across diverse downstream tasks, but their monolithic nature restricts scalability and efficiency in complex problem-solving. While recent research explores multi-agent collaboration among LLMs, most approaches rely on static organization…

Cited by 0SourcecodeScholar
2025

OCTDiff: Bridged Diffusion Model for Portable OCT Super-Resolution and Enhancement

NeurIPS 2025spotlight

Medical imaging super-resolution is critical for improving diagnostic utility and reducing costs, particularly for low-cost modalities such as portable Optical Coherence Tomography (OCT). We propose OCTDiff, a bridged diffusion model designed to enhance image resolution and quality from portable OCT…

Cited by 0SourcecodeScholar
2025

RAM-W600: A Multi-Task Wrist Dataset and Benchmark for Rheumatoid Arthritis

NeurIPS 2025poster

Rheumatoid arthritis (RA) is a common autoimmune disease that has been the focus of research in computer-aided diagnosis (CAD) and disease monitoring. In clinical settings, conventional radiography (CR) is widely used for the screening and evaluation of RA due to its low cost and accessibility. The…

Cited by 0SourcecodeScholar
2025

Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching

ACL 2025finding

Large language models (LLMs) often struggle to provide up-to-date information due to their one-time training and the constantly evolving nature of the world. To keep LLMs current, existing approaches typically involve continued pre-training on new documents. However, they frequently face difficultie…

2025

V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and Me

ACL 2025long

Oracle Bone Script (OBS) is a vital treasure of human civilization, rich in insights from ancient societies. However, the evolution of written language over millennia complicates its decipherment. In this paper, we propose V-Oracle, an innovative framework that utilizes Large Multi-modal Models (LMM…

Cited by 0SourcePDFScholar
2024

Deep Convolution Network Based Super Resolution DOA Estimation with Toeplitz and Sparse Prior

ICASSP 2024accepted

In this paper, a deep learning (DL) based approach is investigated for direction-of-arrival (DOA) estimation, where large-scale uniform linear arrays (ULAs) and small number of samples are considered. Different from existing DL based DOA estimators, the proposed solution first exploits the Toeplitz…

Cited by 0SourceScholar
2024

Improving LLM Generations via Fine-Grained Self-Endorsement

ACL 2024findings

This work studies mitigating fact-conflicting hallucinations for large language model (LLM) at inference time.Particularly, we propose a self-endorsement framework that leverages the fine-grained fact-level comparisons across multiple sampled responses.Compared with prior ensemble methods (e.g., sel…

Cited by 2SourcePDFScholar
2024

RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models

NeurIPS 2024poster

Diffusion models have achieved remarkable advancements in text-to-image generation. However, existing models still have many difficulties when faced with multiple-object compositional generation. In this paper, we propose ***RealCompo***, a new *training-free* and *transferred-friendly* text-to-imag…

2024

Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation

ACL 2024long

Despite showing impressive abilities, large language models (LLMs) often struggle with factual inaccuracies, i.e., ”hallucinations”, even when they hold relevant knowledge. To mitigate these hallucinations, current approaches typically necessitate high-quality human factuality annotations. In this w…

Cited by 35SourcePDFScholar
2024

Three-Dimensional Spatial-Temporal Near-Field Passive Localization Based on an Exact Spatial Propagation Model

ICASSP 2024accepted

Based on the exact source-sensor spatial geometry, a three-dimensional (3-D) spatial-temporal localization algorithm for multiple near-field (NF) sources is proposed without adopting the Fresnel approximation, which simplifies the spatial phase difference by Taylors polynomial. In addition, consider…

Cited by 1SourceScholar
2024

Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

NeurIPS 2024poster

Despite the impressive capabilities of Large Language Models (LLMs) on various tasks, they still struggle with scenarios that involves complex reasoning and planning. Self-correction and self-learning emerge as viable solutions, employing strategies that allow LLMs to refine their outputs and learn…

2024

Towards the Theory of Unsupervised Federated Learning: Non-asymptotic Analysis of Federated EM Algorithms

ICML 2024poster

While supervised federated learning approaches have enjoyed significant success, the domain of unsupervised federated learning remains relatively underexplored. Several federated EM algorithms have gained popularity in practice, however, their theoretical foundations are often lacking. In this paper…

Cited by 4SourcePDFScholar
2024

VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPs

ICLR 2024poster

GNN-to-MLP distillation aims to utilize knowledge distillation (KD) to learn computationally-efficient multi-layer perceptron (student MLP) on graph data by mimicking the output representations of teacher GNN. Existing methods mainly make the MLP to mimic the GNN predictions over a few class labels.…

2024

VideoTetris: Towards Compositional Text-to-Video Generation

NeurIPS 2024poster

Diffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in object numbers. To address these limitations, we propose Video…

2024

ZERO-IG: Zero-Shot Illumination-Guided Joint Denoising and Adaptive Enhancement for Low-Light Images

CVPR 2024poster

This paper presents a novel zero-shot method for jointly denoising and enhancing real-word low-light images. The proposed method is independent of training data and noise distribution. Guided by illumination we integrate denoising and enhancing processes seamlessly enabling end-to-end training. Pair…

2023

6D Pose Estimation Based on 3D Edge Binocular Reprojection Optimization for Robotic Assembly

RA-L 2023

Accurate 6D pose estimation of object is important for robot assembly. This letter presents a novel method for achieving high precision 6D pose estimation by exploiting the reprojection of 3D edges onto binocular RGB image pairs. Our proposed method encompasses three phases: detection, pose initiali

Cited by 9SourceScholar
2023

Evolutionary Neural Architecture Search for Transformer in Knowledge Tracing

NeurIPS 2023poster

Knowledge tracing (KT) aims to trace students' knowledge states by predicting whether students answer correctly on exercises. Despite the excellent performance of existing Transformer-based KT approaches, they are criticized for the manually selected input features for fusion and the defect of singl…

2023

Quality-Similar Diversity via Population Based Reinforcement Learning

ICLR 2023poster

Diversity is a growing research topic in Reinforcement Learning (RL). Previous research on diversity has mainly focused on promoting diversity to encourage exploration and thereby improve quality (the cumulative reward), maximizing diversity subject to quality constraints, or jointly maximizing qual…

Cited by 22SourcePDFScholar
2023

Training Large-Vocabulary Neural Language Models by Private Federated Learning for Resource-Constrained Devices

ICASSP 2023accepted

Federated Learning (FL) is a technique to train models on distributed edge devices with local data samples. Differential Privacy (DP) can be applied with FL to provide a formal privacy guarantee for sensitive data on device. Our goal is to train a large neural network language model (NNLM) on comput…

Cited by 0SourceScholar
2022

Acceleration in Distributed Optimization under Similarity

AISTATS 2022poster

We study distributed (strongly convex) optimization problems over a network of agents, with no centralized nodes. The loss functions of the agents are assumed to be similar, due to statistical data similarity or otherwise. In order to reduce the number of communications to reach a solution accuracy,…

Cited by 27SourcePDFScholar
2022

Conjugate Augmented Spatial-Temporal Near-Field Sources Localization with Cross Array

ICASSP 2022accepted

A new near-field source localization method is proposed for two-dimensional (2-D) direction-of-arrival (DOA) and range estimation based on a symmetrical cross array. It first employs the conjugate symmetry property of the signal auto-correlation at different time delays to construct a conjugate augm…

Cited by 0SourceScholar
2022

Greedy when Sure and Conservative when Uncertain about the Opponents

ICML 2022spotlight

We develop a new approach, named Greedy when Sure and Conservative when Uncertain (GSCU), to competing online against unknown and nonstationary opponents. GSCU improves in four aspects: 1) introduces a novel way of learning opponent policy embeddings offline; 2) trains offline a single best response…

2021

Lightweight Dual-Task Networks For Crowd Counting In Aerial Images

ICASSP 2021accepted

As a research hotspot of computer vision, crowd counting methods have achieved success in natural images. But crowd counting in aerial images are rarely explored, and existing methods do not perform well because of the higher resolution, smaller object scale and more complex scene. Therefore, this p…

Cited by 0SourceScholar
2020

Accelerated Primal-Dual Algorithms for Distributed Smooth Convex Optimization over Networks

AISTATS 2020poster

This paper proposes a novel family of primal-dual-based distributed algorithms for smooth, convex, multi-agent optimization over networks that uses only gradient information and gossip communications. The algorithms can also employ acceleration on the computation and communications. We provide a u…

2017

Estimation of EMG signal for shoulder joint based on EEG signals for the control of upper-limb power assistance devices

ICRA 2017poster

Brain-Machine Interface (BMI) has emerged as a powerful tool for assisting disabled people and for augmenting human performance. Up so far, no studies have succeeded in the power augmentation for the multi-DOFs robot based on EEG signals, especially for the complex shoulder joint. In this work, we p…

Cited by 8SourceScholar
2015

Accurate analysis method of background ionosphere effects on Geosynchronous SAR focusing

ICASSP 2015accepted

The background ionosphere time variance within the extremely long integration time of Geosynchronous Synthetic Aperture Radar (GEO SAR) needs to be considered for GEO SAR focusing. Meanwhile, because of the curved trajectory and the very complex geometry relationship between satellite motion and ear…

Cited by 0SourceScholar