← Search

Yiming Zhang

58 accepted papers

2026

Achieving Expert-Level Agent from Foundation Model via Complexity Curriculum Reinforcement Learning with Synthetic Data

ICLR 2026poster

Large language model (LLM) agents exhibit strong mathematical problem-solving abilities and can even solve International Mathematical Olympiad (IMO) level problems with the assistance of formal proof systems. However, due to weak heuristics for auxiliary constructions, AI for geometry problem solvin…

Cited by 0SourceScholar
2026

Cyto-SSL: A Self-Supervised Pretraining Framework for Cytology Foundation Model

AAAI 2026technical

Cytological images originate from exfoliated cells, collected via liquid-based slides and digitized into whole slide images (WSIs). Unlike histological WSIs that exhibit continuous and well-structured tissue, cytological WSIs are sparse in spatial distribution and unstructured in cellular relationsh

Cited by 0SourcePDFScholar
2026

ENHANCING CROSS-VIEW GEO-LOCALIZATION GENERALIZATION VIA GLOBAL-LOCAL CONSISTENCY AND GEOMETRIC EQUIVARIANCE

ICASSP 2026poster

Cross-view geo-localization (CVGL) aims to match images of the same location captured from drastically different viewpoints. Despite recent progress, existing methods still face two key challenges: (1) achieving robustness under severe appearance variations induced by diverse UAV orientations and fi…

Cited by 0SourcePDFScholar
2026

Exploring Visual Pretraining for Learning Language Intelligence

CVPR 2026

While the most fundamental pretraining paradigm typically trains modality-specific models on their respective datasets, the Platonic Representation Hypothesis that representations eventually align across modalities as data and model scale suggests an intriguing possibility: large language models (LL

Cited by 0SourcecodeScholar
2026

GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics

CVPR 2026

This paper presents GeoAgent, a model capable of reasoning closely with humans and deriving fine-grained address conclusions. Previous RL-based methods have achieved breakthroughs in performance and interpretability but still remain concerns because of their reliance on AI-generated chain-of-thought

Cited by 0SourceScholar
2026

MODEL MERGING SCALING LAWS IN LARGE LANGUAGE MODELS

ICML 2026poster

We study empirical scaling laws for language model merging measured by cross-entropy. Despite its wide practical use, merging lacks a quantitative rule that predicts returns as we add experts or scale the model size. We identify a compact power law that links model size and expert number: the size-d…

Cited by 0SourceScholar
2026

OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment

ICLR 2026poster

Large language model (LLM) alignment faces a critical dilemma when addressing multiple human preferences: improvements in one dimension frequently come at the expense of others, creating unavoidable trade-offs between competing objectives like helpfulness and harmlessness. While prior work mainly fo…

Cited by 0SourcecodeScholar
2026

RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images

AAAI 2026technical

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing images based on free-form natural language expressions. Existing approaches are typically constrained to closed-set vocabularies, limiting their applicability in open-world scenarios. While recent attempts to leverage

Cited by 0SourcePDFScholar
2026

ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

ICML 2026poster

Current evaluations of spatial intelligence can be systematically invalid under modern vision-language model (VLM) settings. First, many benchmarks derive question-answer (QA) pairs from point-cloud-based 3D annotations originally curated for traditional 3D perception. When such annotations are trea…

Cited by 0SourceScholar
2026

SPARD: Single-step Inference with Adaptive Sampling in Residual Diffusion for Human Motion Prediction

AAAI 2026technical

The task of stochastic human motion prediction has attracted significant attention in recent years due to its wide-ranging applications in robotics, animation, and human-computer interaction. While diffusion models have demonstrated promising progress in this domain, they remain hindered by two crit

Cited by 0SourcePDFScholar
2026

ST-LLM: Spatial Transcriptomics Embedding with Large Language Models

AAAI 2026technical

Spatial transcriptomics provides unprecedented opportunities to analyze gene patterns while preserving spatial tissue architecture. However, traditional deep learning methods for spatial transcriptomics analysis face significant challenges in multi-modal data integration, spatial dependency modeling

Cited by 0SourcePDFScholar
2025

AccCtr: Accelerating Training-Free Conditional Control For Diffusion Models

IJCAI 2025

In current training-free Conditional Diffusion Models (CDM), the sampling process is steered by the gradient, which measures the discrepancy between the guidance and the condition extracted by a pre-trained condition extraction network. These methods necessitate small guidance steps, resulting in lo

Cited by 0SourcePDFScholar
2025

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity

EMNLP 2025

Large Language Models (LLMs) with extended context lengths face significant computational challenges during the pre-filling phase, primarily due to the quadratic complexity of self-attention. Existing methods typically employ dynamic pattern matching and block-sparse low-level implementations. Howev

2025

Backtracking Improves Generation Safety

ICLR 2025oral

Text generation has a fundamental limitation almost by definition: there is no taking back tokens that have been generated, even when they are clearly problematic. In the context of language model safety, when a partial unsafe generation is produced, language models by their nature tend to happily k…

Cited by 14SourcePDFScholar
2025

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

ICCV 2025poster

Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding. However, achieving high performance in zero-shot video tasks remains challenging. Traditional video processing methods rely heavily on fine-tuning to capture nuanced spatial-temporal deta…

2025

CIE: Controlling Language Model Text Generations Using Continuous Signals

EMNLP 2025

Aligning language models (LMs) with user intent is becoming increasingly relevant to enhance user experience.This calls for designing methods that can allow users to control the properties of the language that LMs generate, for example, controlling the length of the generation or the complexity of t

2025

Cognitive Bias and Reassignment: Who Can Contribute High Quality LLM Data

AAAI 2025technical

In recent years, the rapid development of Large Language Models has highlighted the urgent need for large-scale, high-quality, and diverse data. We have launched an LLM data co-creation platform aimed at bringing together a wide range of participants to contribute data. Within six months, the platfo…

Cited by 0SourcePDFScholar
2025

FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research Insights

ACL 2025finding

The automation of scientific research through large language models (LLMs) presents significant opportunities but faces critical challenges in knowledge synthesis and quality assurance. We introduce Feedback-Refined Agent Methodology (FRAME), a novel framework that enhances medical paper generation…

Cited by 0SourcePDFScholar
2025

Human-Aligned Chess With a Bit of Search

ICLR 2025poster

Chess has long been a testbed for AI's quest to match human intelligence, and in recent years, chess AI systems have surpassed the strongest humans at the game. However, these systems are *not human-aligned*; they are unable to match the skill levels of all human partners or model human-like behavio…

2025

Improving Adversarial Transferability through Channel-wise Scaling and Frequency-random Dropping

ICASSP 2025accepted

For black-box attacks, most existing attack methods exhibit weak transferability due to the significant discrepancy between substitute model and victim model. We argue that the model-specific discriminative regions are a key factor causing overfitting to the source model. However, existing model aug…

Cited by 0SourceScholar
2025

InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models

NeurIPS 2025spotlight

Model fusion combines multiple Large Language Models (LLMs) with different strengths into a more powerful, integrated model through lightweight training methods. Existing works on model fusion focus primarily on supervised fine-tuning (SFT), leaving preference alignment (PA) —a critical phase for en…

Cited by 0SourcecodeScholar
2025

InfiGFusion: Graph-on-Logits Distillation via Efficient Gromov-Wasserstein for Model Fusion

NeurIPS 2025poster

Recent advances in large language models (LLMs) have intensified efforts to fuse heterogeneous open-source models into a unified system that inherits their complementary strengths. Existing logit-based fusion methods maintain inference efficiency but treat vocabulary dimensions independently, overl…

Cited by 0SourcecodeScholar
2025

MRAG: A Modular Retrieval Framework for Time-Sensitive Question Answering

EMNLP 2025

Understanding temporal concepts and answering time-sensitive questions is crucial yet a challenging task for question-answering systems powered by large language models (LLMs). Existing approaches either update the parametric knowledge of LLMs with new facts, which is resource-intensive and often im

2025

Multi-Agent Decision Transformer for Power Control in Wireless Networks

ICASSP 2025accepted

This paper introduces a novel offline approach to power control in wireless networks using a multi-agent reinforcement learning (MARL) framework. We develop a multi-agent decision transformer method to optimize performance metrics including sum-rate or packet delay. In this distributed method, each…

Cited by 0SourceScholar
2025

Multi-Scale Finetuning for Encoder-based Time Series Foundation Models

NeurIPS 2025poster

Time series foundation models (TSFMs) demonstrate impressive zero-shot performance for time series forecasting. However, an important yet underexplored challenge is how to effectively finetune TSFMs on specific downstream tasks. While naive finetuning can yield performance gains, we argue that it fa…

Cited by 0SourcecodeScholar
2025

Persistent Pre-training Poisoning of LLMs

ICLR 2025poster

Large language models are pre-trained on uncurated text datasets consisting of trillions of tokens scraped from the Web. Prior work has shown that: (1) web-scraped pre-training datasets can be practically poisoned by malicious actors; and (2) adversaries can compromise language models after poisonin…

Cited by 3SourcePDFScholar
2025

RANKCLIP: Ranking-Consistent Language-Image Pretraining

ICCV 2025poster

Self-supervised contrastive learning models, such as CLIP, have set new benchmarks for vision-language models in many downstream tasks. However, their dependency on rigid one-to-one mappings overlooks the complex and often multifaceted relationships between and within texts and images. To this end,…

2025

RPDR: A Round-trip Prediction-Based Data Augmentation Framework for Long-Tail Question Answering

EMNLP 2025

Long-tail question answering presents significant challenges for large language models (LLMs) due to their limited ability to acquire and accurately recall less common knowledge. Retrieval-augmented generation (RAG) systems have shown great promise in mitigating this limitation by integrating extern

2025

STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

ICCV 2025poster

The use of Multimodal Large Language Models (MLLMs) as an end-to-end solution for Embodied AI and Autonomous Driving has become a prevailing trend. While MLLMs have been extensively studied for visual semantic understanding tasks, their ability to perform precise and quantitative spatial-temporal un…

Cited by 0SourcePDFScholar
2025

Towards Large-Scale In-Context Reinforcement Learning by Meta-Training in Randomized Worlds

NeurIPS 2025poster

In-Context Reinforcement Learning (ICRL) enables agents to learn automatically and on-the-fly from their interactive experiences. However, a major challenge in scaling up ICRL is the lack of scalable task collections. To address this, we propose the procedurally generated tabular Markov Decision Pro…

Cited by 0SourceScholar
2025

VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior

ICCV 2025accepted

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos…

2024

EFSA: Towards Event-Level Financial Sentiment Analysis

ACL 2024long

In this paper, we extend financial sentiment analysis (FSA) to event-level since events usually serve as the subject of the sentiment in financial text. Though extracting events from the financial text may be conducive to accurate sentiment predictions, it has specialized challenges due to the lengt…

2024

Embedding and Gradient Say Wrong: A White-Box Method for Hallucination Detection

EMNLP 2024main

In recent years, large language models (LLMs) have achieved remarkable success in the field of natural language generation. Compared to previous small-scale models, they are capable of generating fluent output based on the provided prefix or prompt. However, one critical challenge — the *hallucinati…

Cited by 1SourcePDFScholar
2024

Fast Adaptation via Prompted Data: An Efficient Cross-Domain Fine-tuning Method for Large Language Models

COLING 2024main

Large language models (LLMs) have achieved great success in a variety of natural language understanding tasks. However, domain discrepancies between the downstream task and the pre-training corpora may have hurdled LLMs to excel further in the vertical applications. Contrary to prior computational-h…

2024

FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents

EMNLP 2024main

We introduce FinDVer, a comprehensive benchmark specifically designed to evaluate the explainable claim verification capabilities of LLMs in the context of understanding and analyzing long, hybrid-content financial documents. FinDVer contains 4,000 expert-annotated examples across four subsets, each…

2024

PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models

CVPR 2024poster

Recent advancements in personalized text-to-image (T2I) models have revolutionized content creation empowering non-experts to generate stunning images with unique styles. While promising animating these personalized images with realistic motions poses significant challenges in preserving distinct st…

2024

RECOST: External Knowledge Guided Data-efficient Instruction Tuning

ACL 2024findings

In the current landscape of large language models (LLMs), the process of instruction tuning serves as an essential step. Considering the high computing power overhead, data-efficient instruction tuning was proposed to reduce the training data size in this process, aiming at selecting high-quality in…

2023

BiasX: “Thinking Slow” in Toxic Content Moderation with Explanations of Implied Social Biases

EMNLP 2023short main

Toxicity annotators and content moderators often default to mental shortcuts when making decisions. This can lead to subtle toxicity being missed, and seemingly toxic but harmless content being over-detected. We introduce BiasX, a framework that enhances content moderation setups with free-text expl…

Cited by 0SourceScholar
2023

Emergent Communication in Interactive Sketch Question Answering

NeurIPS 2023poster

Vision-based emergent communication (EC) aims to learn to communicate through sketches and demystify the evolution of human communication. Ironically, previous works neglect multi-round interaction, which is indispensable in human communication. To fill this gap, we first introduce a novel Interacti…

2022

Co-Modality Graph Contrastive Learning for Imbalanced Node Classification

NeurIPS 2022accept

Graph contrastive learning (GCL), leveraging graph augmentations to convert graphs into different views and further train graph neural networks (GNNs), has achieved considerable success on graph benchmark datasets. Yet, there are still some gaps in directly applying existing GCL methods to real-worl…

2022

MultiScan: Scalable RGBD scanning for 3D environments with articulated objects

NeurIPS 2022accept

We introduce MultiScan, a scalable RGBD dataset construction pipeline leveraging commodity mobile devices to scan indoor scenes with articulated objects and web-based semantic annotation interfaces to efficiently annotate object and part semantics and part mobility parameters. We use this pipeline t…

Cited by 30SourcePDFScholar
2022

S2 Reducer: High-Performance Sparse Communication to Accelerate Distributed Deep Learning

ICASSP 2022accepted

Distributed stochastic gradient descent (SGD) approach has been widely used in large-scale deep learning, and the gradient collective method is vital to ensure the training scalability of the distributed deep learning system. Collective communication such as AllReduce has been widely adopted for the…

Cited by 0SourceScholar
2022

Towards Unifying the Label Space for Aspect- and Sentence-based Sentiment Analysis

ACL 2022findings

The aspect-based sentiment analysis (ABSA) is a fine-grained task that aims to determine the sentiment polarity towards targeted aspect terms occurring in the sentence. The development of the ABSA task is very much hindered by the lack of annotated data. To tackle this, the prior works have studied…

2021

Adapting Meta Knowledge with Heterogeneous Information Network for COVID-19 Themed Malicious Repository Detection

IJCAI 2021poster

As cyberattacks caused by malware have proliferated during the pandemic, building an automatic system to detect COVID-19 themed malware in social coding platforms is in urgent need. The existing methods mainly rely on file content analysis while ignoring structured information among entities in soci…

2021

Distilling Meta Knowledge on Heterogeneous Graph for Illicit Drug Trafficker Detection on Social Media

NeurIPS 2021poster

Driven by the considerable profits, the crime of drug trafficking (a.k.a. illicit drug trading) has co-evolved with modern technologies, e.g., social media such as Instagram has become a popular platform for marketing and selling illicit drugs. The activities of online drug trafficking are nimble an…

2020

Argot: Generating Adversarial Readable Chinese Texts

IJCAI 2020poster

Natural language processing (NLP) models are known vulnerable to adversarial examples, similar to image processing models. Studying adversarial texts is an essential step to improve the robustness of NLP models. However, existing studies mainly focus on analyzing English texts and generating adversa…

2020

MLCVNet: Multi-Level Context VoteNet for 3D Object Detection

CVPR 2020poster

In this paper, we address the 3D object detection task by capturing multi-level contextual information with the self-attention mechanism and multi-scale feature fusion. Most existing 3D object detection methods recognize objects individually, without giving any consideration on contextual informatio…

Cited by 229PDFcodeScholar