← Search

Tianyu Zhang

55 accepted papers

2026

AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems provide enhanced coverage, robustness, and collaboration compared to single-sensor

Cited by 0SourcePDFScholar
2026

CADC: Content Adaptive Diffusion-Based Generative Image Compression

CVPR 2026

Diffusion-based generative image compression has demonstrated remarkable potential for achieving realistic reconstruction at ultra-low bitrates. The key to unlocking this potential lies in making the entire compression process content-adaptive, ensuring that the encoder's representation and the deco

Cited by 0SourceScholar
2026

Contrastive Weak-to-Strong Generalization

ICML 2026poster

Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples from aligned weaker ones, without requiring human feedback or explicit reward modeling. However, its robustness and generalization are hindered by the noise and…

Cited by 0SourceScholar
2026

Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning

ICLR 2026poster

A major bottleneck in scientific discovery consists of narrowing an exponentially large set of objects, such as proteins or molecules, to a small set of promising candidates with desirable properties. While this process can rely on expert knowledge, recent methods leverage reinforcement learning (RL…

Cited by 0SourceScholar
2026

Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization

ICML 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO) have significantly advanced the reasoning capabilities of large language models. Extending these methods to multimodal settings, however, faces a critical challenge: the instability of std-based norma…

Cited by 0SourceScholar
2026

LUMIN: A Longitudinal Multi-modal Knowledge Decomposition Network for Predicting Breast Cancer Recurrence

AAAI 2026technical

Accurate prediction of breast cancer recurrence after treatment is essential for improving long-term outcomes. However, existing models are limited by three key challenges: (1) they typically rely on single-modal data, missing cross-modal interactions; (2) they analyze static snapshots, failing to c

Cited by 0SourcePDFScholar
2026

OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

ICML 2026oral

As high-quality public text approaches exhaustion, a phenomenon known as the Data Wall—LLM pre-training is shifting from more tokens to better tokens. However, existing methods either rely on heuristic static filters that ignore training dynamics, or use dynamic yet optimizer-agnostic criteria based…

Cited by 0SourceScholar
2026

On The Surprising Effectiveness of a Single Global Merging in Decentralized Learning

ICLR 2026oral

Decentralized learning provides a scalable alternative to parameter-server-based training, yet its performance is often hindered by limited peer-to-peer communication. In this paper, we study how communication should be scheduled over time to improve global generalization, including determining whe…

Cited by 0SourceScholar
2026

Pareto-Guided Optimal Transport for Multi-Reward Alignment

ICML 2026poster

Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward models remains a significant challenge. Existing multi-reward fusion approaches rely on weighted summation, which is costly to tune and insufficient for …

Cited by 0SourceScholar
2026

PreferThinker: Reasoning-based Personalized Image Preference Assessment

ICLR 2026poster

Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of reference images as prior information. Existing methods mainly focus on general preference assessment, training models with large-scale data to tackle well-defined task…

Cited by 0SourceScholar
2026

Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation

CVPR 2026

Text-to-image generation has advanced rapidly, yet it still struggles to capture the nuanced user preferences. Existing approaches typically rely on multimodal large language models to infer user preferences, but the derived prompts or latent codes rarely reflect them faithfully, leading to suboptim

Cited by 0SourceScholar
2026

Sketch-Guided Anime Hair Editing Using Multimodal Diffusion Transformer (Student Abstract)

AAAI 2026technical

Anime hair design is crucial but challenging, as it conveys personality and emotion through stylized geometry and layered structure. In this work, we propose a sketch-guided approach for intuitive control of multimodal diffusion transformers (MMDiT) to generate semantically consistent anime hairstyl

Cited by 0SourcePDFScholar
2026

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

ICML 2026poster

During conversational interactions, humans subconsciously engage in concurrent thinking while listening to a speaker. Although this internal cognitive processing may not always manifest as explicit linguistic structures, it is instrumental in formulating high-quality responses. Inspired by this cogn…

Cited by 0SourceScholar
2025

AI for Global Climate Cooperation: Modeling Global Climate Negotiations, Agreements, and Long-Term Cooperation in RICE-N

ICML 2025poster

Global cooperation on climate change mitigation is essential to limit temperature increases while supporting long-term, equitable economic growth and sustainable development. Achieving such cooperation among diverse regions, each with different incentives, in a dynamic environment shaped by complex…

2025

Advantage Alignment Algorithms

ICLR 2025oral

Artificially intelligent agents are increasingly being integrated into human decision-making: from large language model (LLM) assistants to autonomous vehicles. These systems often optimize their individual objective, leading to conflicts, particularly in general-sum games where naive reinforcement…

Cited by 0SourcePDFScholar
2025

AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding

NeurIPS 2025poster

Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared embedding space with the LLM while preserving semantic similarit…

Cited by 0SourceScholar
2025

Aligning Constraint Generation with Design Intent in Parametric CAD

ICCV 2025poster

We introduce alignment techniques from reasoning large language models (LLMs) to the task of generating engineering sketch constraints in computer-aided design (CAD) models. Engineering sketches are composed of geometric primitives (such as points and lines) connected by constraints (such as perpend…

Cited by 0SourcePDFScholar
2025

AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists

EMNLP 2025

Despite long-standing efforts in accelerating scientific discovery with AI, building AI co-scientists remains challenging due to limited high-quality data for training and evaluation. To tackle this data scarcity issue, we present AutoSDT, an automatic pipeline that collects high-quality coding task

2025

BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks

ICLR 2025poster

Multimodal AI has the potential to significantly enhance document-understanding tasks, such as processing receipts, understanding workflows, extracting data from documents, and summarizing reports. Code generation tasks that require long-structured outputs can also be enhanced by multimodality. Desp…

Cited by 0SourcePDFScholar
2025

Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment

ICCV 2025poster

Contemporary image generation systems have achieved high fidelity and superior aesthetic quality beyond basic text-image alignment. However, existing evaluation frameworks have failed to evolve in parallel. This study reveals that human preference reward models fine-tuned based on CLIP and BLIP arch…

2025

Faster and Better LLMs via Latency-Aware Test-Time Scaling

EMNLP 2025

Test-Time Scaling (TTS) has proven effective in improving the performance of Large Language Models (LLMs) during inference. However, existing research has overlooked the efficiency of TTS from a latency-sensitive perspective. Through a latency-aware evaluation of representative TTS methods, we demon

Cited by 0SourcePDFScholar
2025

Few-Shot Domain Adaptation for Learned Image Compression

AAAI 2025technical

Learned image compression (LIC) has achieved state-of-the-art rate-distortion performance, deemed promising for next-generation image compression techniques. However, pre-trained LIC models usually suffer from significant performance degradation when applied to out-of-training-domain images, implyin…

Cited by 0SourcePDFScholar
2025

LaMP-Val: Large Language Models Empower Personalized Valuation in Auction

EMNLP 2025

Auctions are a vital economic mechanism used to determine the market value of goods or services through competitive bidding within a specific framework. However, much of the current research primarily focuses on the bidding algorithms used within auction mechanisms. This often neglects the potential

2025

MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation

ICLR 2025poster

Model merging has emerged as an effective approach to combining multiple single-task models into a multitask model. This process typically involves computing a weighted average of the model parameters without additional training. Existing model-merging methods focus on improving average task accurac…

2025

Majority of the Bests: Improving Best-of-N via Bootstrapping

NeurIPS 2025poster

Sampling multiple outputs from a Large Language Model (LLM) and selecting the most frequent (Self-consistency) or highest-scoring (Best-of-N) candidate is a popular approach to achieve higher accuracy in tasks with discrete final answers. Best-of-N (BoN) selects the output with the highest reward, a…

Cited by 0SourceScholar
2025

MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE

NeurIPS 2025spotlight

Large Language Models (LLMs) have achieved remarkable success across many applications, with Mixture of Experts (MoE) models demonstrating great potential. Compared to traditional dense models, MoEs achieve better performance with less computation. Speculative decoding (SD) is a widely used techniqu…

Cited by 0SourceScholar
2025

MuPT: A Generative Symbolic Music Pretrained Transformer

ICLR 2025poster

In this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design…

Cited by 10SourcePDFScholar
2025

Neuron-Level Sequential Editing for Large Language Models

ACL 2025long

This work explores sequential model editing in large language models (LLMs), a critical task that involves modifying internal knowledge within LLMs continuously through multi-round editing, each incorporating updates or corrections to adjust the model’s outputs without the need for costly retraining…

2025

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

NeurIPS 2025poster

Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more rigorous assessment methods. To address these limitations, we intro…

Cited by 0SourceScholar
2025

STRICT: Stress-Test of Rendering Image Containing Text

EMNLP 2025

While diffusion models have revolutionized text-to-image generation with their ability to synthesize realistic and diverse scenes, they continue to struggle with generating consistent and legible text within images. This shortcoming is commonly attributed to the locality bias inherent in diffusion-b

2025

VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

ICLR 2025poster

We introduce Visual Caption Restoration (VCR), a novel vision-language task that challenges models to accurately restore partially obscured texts using pixel-level hints within images through complex reasoning. This task stems from the observation that text embedded in images intrinsically differs f…

2024

AutoCali: Enhancing AoA-based Indoor Localization through Automatic Phase Calibration

ICASSP 2024accepted

Recent advancements in WiFi indoor localization have demonstrated the potential for achieving decimeter-level accuracy based on Angle of Arrival (AoA). However, existing commercial WiFi Access Points (APs) suffer from phase offset across different antennas, which significantly degrade the performanc…

Cited by 0SourceScholar
2024

CAMEL: CAusal Motion Enhancement Tailored for Lifting Text-driven Video Editing

CVPR 2024poster

Text-driven video editing poses significant challenges in exhibiting flicker-free visual continuity while preserving the inherent motion patterns of original videos. Existing methods operate under a paradigm where motion and appearance are intricately intertwined. This coupling leads to the network…

Cited by 4SourcePDFScholar
2024

Dynamic Prompt Optimizing for Text-to-Image Generation

CVPR 2024poster

Text-to-image generative models specifically those based on diffusion models like Imagen and Stable Diffusion have made substantial advancements. Recently there has been a surge of interest in the delicate refinement of text prompts. Users assign weights or alter the injection time steps of certain…

2024

Efficient Solution to PnP Problem Based on Vision Geometry

RA-L 2024

Perspective-n-Point (PnP) problem aims to estimate pose from known 3D map points and their projections. Efficient PnP (EPnP), one of the classical PnP solvers, represents camera pose with control points, which are easier to estimate utilizing the least square (LS) formulation. However, the geometry

Cited by 9SourceScholar
2024

Expected flow networks in stochastic environments and two-player zero-sum games

ICLR 2024poster

Generative flow networks (GFlowNets) are sequential sampling models trained to match a given distribution. GFlowNets have been successfully applied to various structured object generation tasks, sampling a diverse set of high-reward objects quickly. We propose expected flow networks (EFlowNets), whi…

2024

FastPCI: Motion-Structure Guided Fast Point Cloud Frame Interpolation

ECCV 2024poster

"Point cloud frame interpolation is a challenging task that involves accurate scene flow estimation across frames and maintaining the geometry structure. Prevailing techniques often rely on pre-trained motion estimators or intensive testing-time optimization, resulting in compromised interpolation a…

2024

Kinetic-energy-optimal and Safety-guaranteed Trajectory Planning for Bridge Inspection Robot Manipulator

IROS 2024poster

Bridge inspections are essential for maintaining key infrastructure and preventing structural and functional failures. Nevertheless, traditional manual inspection techniques are plagued by laboriousness, high risk, and low efficiency. Although numerous automation inspection methods have been studied…

Cited by 0SourceScholar
2024

RD-NERF: Neural Robust Distilled Feature Fields for Sparse-View Scene Segmentation

ICASSP 2024accepted

We propose Neural Robust Distilled Feature Fields (RD-NeRF) for achieving robust 3D semantic feature distillation and 3D consistent scene segmentation with sparse-view labels. Specifically, we introduce a two-stage pipeline. In the distillation stage, we employ the pre-trained image feature extracto…

Cited by 0SourceScholar
2024

Self-supervised Scale Recovery for Decoupled Visual-inertial Odometry

RA-L 2024

Accurate localization for intelligent robots remains a significant challenge, and self-supervised visual-inertial odometry (VIO) has emerged as a promising solution. However, existing self-supervised VIO works consider inertial information as the ordinary data input, losing its ability to recover ab

Cited by 3SourceScholar
2023

Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation

ICASSP 2023accepted

Contrastive learning has shown remarkable success in the field of multimodal representation learning. In this paper, we propose a pipeline of contrastive language-audio pretraining to develop an audio representation by combining audio data with natural language descriptions. To accomplish this targe…

Cited by 0SourceScholar
2023

PHA: Patch-Wise High-Frequency Augmentation for Transformer-Based Person Re-Identification

CVPR 2023highlight

Although recent studies empirically show that injecting Convolutional Neural Networks (CNNs) into Vision Transformers (ViTs) can improve the performance of person re-identification, the rationale behind it remains elusive. From a frequency perspective, we reveal that ViTs perform worse than CNNs in…

2023

Support Generation for Robot-Assisted 3D Printing with Curved Layers

ICRA 2023poster

Robot-assisted 3D printing has drawn a lot of attention by its capability to fabricate curved layers that are optimized according to different objectives. However, the support generation algorithm based on a fixed printing direction for planar layers cannot be directly applied for curved layers as t…

Cited by 8SourceScholar
2022

Biological Sequence Design with GFlowNets

ICML 2022spotlight

Design of de novo biological sequences with desired properties, like protein and DNA sequences, often involves an active loop with several rounds of molecule ideation and expensive wet-lab evaluations. These experiments can consist of multiple stages, with increasing levels of precision and cost of…

2022

ClimateGAN: Raising Climate Change Awareness by Generating Images of Floods

ICLR 2022poster

Climate change is a major threat to humanity and the actions required to prevent its catastrophic consequences include changes in both policy-making and individual behaviour. However, taking action requires understanding its seemingly abstract and distant consequences. Projecting the potential impac…

Cited by 26SourcePDFScholar
2022

FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer

IJCAI 2022poster

Network quantization significantly reduces model inference complexity and has been widely used in real-world deployments. However, most existing quantization methods have been developed mainly on Convolutional Neural Networks (CNNs), and suffer severe degradation when applied to fully quantized visi…

2022

Spatiotemporally Enhanced Photometric Loss for Self-Supervised Monocular Depth Estimation

IROS 2022poster

Recovering depth information from a single image is a long-standing challenge, and self-supervised depth estimation methods have gradually attracted attention due to not relying on high-cost ground truth. Constructing an accurate photometric loss based on photometric consistency is crucial for these…

Cited by 7SourceScholar
2021

One Loss for All: Deep Hashing with a Single Cosine Similarity based Learning Objective

NeurIPS 2021poster

A deep hashing model typically has two main learning objectives: to make the learned binary hash codes discriminative and to minimize a quantization error. With further constraints such as bit balance and code orthogonality, it is not uncommon for existing models to employ a large number (>4) of los…

2021

Singularity-Aware Motion Planning for Multi-Axis Additive Manufacturing

RA-L 2021

Multi-axis additive manufacturing enables high flexibility of material deposition along dynamically varied directions. The Cartesian motion platforms of these machines include three parallel axes and two rotational axes. Singularity on rotational axes is a critical issue to be tackled in motion plan

Cited by 40SourcecodeScholar
2021

UnrealPerson: An Adaptive Pipeline Towards Costless Person Re-Identification

CVPR 2021poster

The main difficulty of person re-identification (ReID) lies in collecting annotated data and transferring the model across different domains. This paper presents UnrealPerson, a novel pipeline that makes full use of unreal image data to decrease the costs in both the training and deployment stages.…

Cited by 88PDFcodeScholar
2021

What If We Could Not See? Counterfactual Analysis for Egocentric Action Anticipation

IJCAI 2021poster

Egocentric action anticipation aims at predicting the near future based on past observation in first-person vision. While future actions may be wrongly predicted due to the dataset bias, we present a counterfactual analysis framework for egocentric action anticipation (CA-EAA) to enhance the capacit…

Cited by 16SourcePDFScholar
2020

Deep Polarized Network for Supervised Learning of Accurate Binary Hashing Codes

IJCAI 2020poster

This paper proposes a novel deep polarized network (DPN) for learning to hash, in which each channel in the network outputs is pushed far away from zero by employing a differentiable bit-wise hinge-like loss which is dubbed as polarization loss. Reformulated within a generic Hamming Distance Metric…

Cited by 0SourcePDFScholar
2020

Rethinking the Distribution Gap of Person Re-identification with Camera-based Batch Normalization

ECCV 2020poster

The fundamental difficulty in person re-identification (ReID) lies in learning the correspondence among individual cameras. It strongly demands costly inter-camera annotations, yet the trained models are not guaranteed to transfer well to previously unseen cameras. These problems significantly limit…