← Search

Yong Li

144 accepted papers

2026

AdaHC: Accelerating Multi-Token Prediction with Adaptive Head Chunking with Pipeline Parallelism

ICML 2026poster

Multi-token prediction (MTP) architecture is widely adopted in LLMs. MTP blocks can be appended to the tail of model to predict additional tokens. However, when training with pipeline parallel, MTP leads to more pipeline bubbles and deteriorates the pipeline efficiency. Based on in-depth analysis of…

Cited by 0SourceScholar
2026

AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent

ICML 2026poster

In modern AI research, baseline and dataset selection is a high-stakes decision in experimental design. It operationalizes a research idea into a concrete evaluation protocol and largely determines the validity and comparability of empirical conclusions. However, making appropriate choices is increa…

Cited by 0SourceScholar
2026

Beyond Accuracy and Complexity: The Effective Information Criterion for Structurally Stable Symbolic Regression

ICML 2026poster

Symbolic regression (SR) traditionally balances accuracy and complexity, implicitly assuming that simpler formulas are structurally more rational. We argue that this assumption is insufficient: existing algorithms often exploit this metric to discover accurate and compact but structurally irrational…

Cited by 0SourceScholar
2026

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

ICML 2026poster

In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, standard evaluations rely on aggregate metrics (e.g., MSE) that conflate model capability with the intrinsic difficulty of the evaluated i…

Cited by 0SourceScholar
2026

Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration

CVPR 2026

Text-to-image generative models have achieved impressive fidelity and diversity, but can inadvertently produce unsafe or undesirable content due to implicit biases embedded in large-scale training datasets.Existing concept erasure methods, whether text-only or image-assisted, face trade-offs: textua

Cited by 0SourceScholar
2026

ChaosNexus: A Foundation Model for ODE-based Chaotic System Forecasting with Hierarchical Multi-scale Awareness

ICML 2026poster

Foundation models have shown great promise in achieving zero-shot or few-shot forecasting for ODE-based chaotic systems via large-scale pretraining. However, existing architectures often fail to capture the multi-scale temporal structures and distinct spectral characteristics of chaotic dynamics. To…

Cited by 0SourceScholar
2026

CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing

ICLR 2026poster

Understanding urban socioeconomic conditions through visual data is a challenging yet essential task for sustainable urban development and policy planning. In this work, we introduce CityLens, a comprehensive benchmark designed to evaluate the capabilities of Large Vision-Language Models (LVLMs) in…

Cited by 0SourcecodeScholar
2026

DehazeGS: Seeing Through Fog with 3D Gaussian Splatting

AAAI 2026technical

Current novel view synthesis methods are typically designed for high-quality and clean input images. However, in foggy scenes, scattering and attenuation can significantly degrade the quality of rendering. Although NeRF-based dehazing approaches have been developed, their reliance on deep fully conn

Cited by 0SourcePDFScholar
2026

DyCon: Dynamic Reasoning Control via Evolving Difficulty Modeling

ICML 2026poster

Recent advances in Large Reasoning Models (LRMs) demonstrate remarkable performance improvements by iteratively reflecting, exploring, and executing complex tasks, yet suffer from inefficiencies due to redundant reasoning, known as "overthinking". Existing methods to mitigate this issue either rely …

Cited by 0SourceScholar
2026

DynaOD: Dynamic Origin-Destination Flow Generation with Discrete-to-Continuous Temporal Semantic Modeling

IJCAI 2026

Dynamic origin-destination (OD) flow generation seeks to synthesize realistic mobility dynamics from temporal context alone, without relying on historical OD observations. A key challenge is to translate semantic temporal signals into temporally coherent OD patterns while preserving the inherent spa

Cited by 0Scholar
2026

EEG-DLite: Dataset Distillation for Efficient Large EEG Model Training

AAAI 2026technical

Large-scale EEG foundation models have shown strong generalization across a range of downstream tasks, but their training remains resource-intensive due to the volume and variable quality of EEG data. In this work, we introduce EEG-DLite, a data distillation framework that enables more efficient pre

Cited by 0SourcePDFScholar
2026

Efficient Reasoning with Balanced Thinking

ICLR 2026poster

Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they often suffer from overthinking, expending redundant computational steps on simple problems, or underthinking, failing to explore sufficient reasoning paths despite inherent capabilities. These issues lead to ineffic…

Cited by 0SourcecodeScholar
2026

FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents

ICLR 2026poster

Mobile GUI agents are becoming critical tools to improve user experience on smart devices, with multimodal large language models (MLLMs) emerging as the dominant paradigms in this domain. Current agents, however, rely on explicit human instructions, overlooking the potential to leverage the contextu…

Cited by 0SourcecodeScholar
2026

Fresco: Frequency-Spatial Consistent Optimization for Fine-Grained Head Avatar Modeling

CVPR 2026

We propose Fresco, a unified optimization pipeline designed to mitigate early over-sharpening, and cross-view drifting in head avatar reconstruction. Fresco combines a Laplacian-pyramid-based frequency curriculum with UV-space consistency regularization to progressively enhance reconstruction qualit

Cited by 0SourcecodeScholar
2026

Generative Adaptation of Dynamics to Environmental Shifts via Weight-space Diffusion

ICML 2026poster

Data-driven dynamics prediction often fails under environmental shifts, while traditional fine-tuning remains computationally prohibitive for hardware-constrained or data-scarce applications. We propose DynaDiff, a generative meta-learning framework that transitions the paradigm from gradient-based …

Cited by 0SourceScholar
2026

GeoProblem Factory: A Visual Interaction System for Solvable and Controllable Geometric Problem Generation by Leveraging Symbolic Deduction Engine

AAAI 2026technical

We propose a novel system, GeoProblem Factory, designed to effectively generate high-quality geometry problems for intelligent education. The system enables to efficiently produce batches of geometry problems for teachers and students, either to save time and manual effort or to support personalized

Cited by 0SourcePDFScholar
2026

Good-for-MDP State Reduction for Stochastic LTL Planning

AAAI 2026technical

We study stochastic planning problems in Markov Decision Processes (MDPs) with goals specified in Linear Temporal Logic (LTL). The state-of-the-art approach transforms LTL formulas into good-for-MDP (GFM) automata, which feature a restricted form of nondeterminism. These automata are then composed w

Cited by 0SourcePDFScholar
2026

History-Aware Reasoning for GUI Agents

AAAI 2026technical

Advances in Multimodal Large Language Models have significantly enhanced Graphical User Interface (GUI) automation. Equipping GUI agents with reliable episodic reasoning capabilities is essential for bridging the gap between users’ concise task descriptions and the complexities of real-world executi

Cited by 0SourcePDFScholar
2026

LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review Platform

ICML 2026poster

Literature reviews are essential to reflect the landscape of research fields. Large language models, especially deep research agents, have recently shown strong capabilities in automated literature review generation. However, it remains a challenging task to rigorously evaluate the scientific value …

Cited by 0SourceScholar
2026

LoRAGen: Structure-Aware Weight Space Learning for LoRA Generation

ICLR 2026poster

The widespread adoption of Low-Rank Adaptation (LoRA) for efficient fine-tuning of large language models has created demand for scalable parameter generation methods that can synthesize adaptation weights directly from task descriptions, avoiding costly task-specific training. We present LoRAGen, a…

Cited by 0SourcecodeScholar
2026

MarketSim: Simulating Stock Markets with Large-Scale Generative Agents

ICML 2026poster

Stock markets are one of the most complex systems in the modern world, where prices emerge from billions of decentralized interactions among heterogeneous participants in an ever-evolving information landscape. While high-fidelity simulation is important for understanding market dynamics, existing a…

Cited by 0SourceScholar
2026

ProBench: Benchmarking GUI Agents with Accurate Process Information

AAAI 2026technical

With the deep integration of artificial intelligence and interactive technology, Graphical User Interface (GUI) Agent, as the carrier connecting goal-oriented natural language and real-world devices, has received widespread attention from the community. Contemporary benchmarks aim to evaluate the co

Cited by 0SourcePDFScholar
2026

Progressive Supernet Training for Efficient Visual Autoregressive Modeling

CVPR 2026

Visual Autoregressive (VAR) models have demonstrated competitive performance with diffusion models in image generation by adopting a "next-scale" prediction paradigm that significantly reduces inference steps. However, VAR's progressive multi-scale generation leads to severe memory overhead due to K

Cited by 0SourcecodeScholar
2026

RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning

ICLR 2026poster

Despite rapid advancements in large language models (LLMs), the token-level autoregressive nature constrains their complex reasoning capabilities. To enhance LLM reasoning, inference-time techniques, including Chain/Tree/Graph-of-Thought(s), successfully improve the performance, as they are fairly c…

Cited by 0SourcecodeScholar
2026

RawMetaDiff: Unlocking Extreme Darkness from Dual-Exposure RAW with Meta-Guided Diffusion

CVPR 2026

Extreme low-light Raw image restoration remains challenging due to overwhelming noise and severe detail loss.In this paper, we exploit the potential of the dual-exposure setting for this severely ill-posed problem.Existing methods suffer from unreliable cross-exposure alignment, resulting in degrade

Cited by 0SourceScholar
2026

ReFAct: Empowering Multimodal Web Agents with Visual and Context Focusing

CVPR 2026

Multimodal Web Search Agents demonstrate a practically valuable capability by fusing information from diverse modalities (e.g., text and vision), retrieved iteratively from the internet, to address complex user queries. However, the visual modality is prone to information overload, and the noise con

Cited by 0SourceScholar
2026

Reinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMs

ICLR 2026poster

Large language models (LLMs) acquire extensive prior knowledge through large-scale pretraining and can be further enhanced via supervised fine-tuning (SFT) or reinforcement learning (RL)-based post-training. A growing body of evidence has shown that RL fine-tuning improves the capability of LLMs bey…

Cited by 0SourcecodeScholar
2026

Reward-Preserving Counterfactual State Editing for Offline Reinforcement Learning

ICML 2026poster

Transformer sequence models such as Decision Transformer can learn strong offline policies from logged trajectories, but they can suffer from causal confusion: reliance on spurious correlations that predict reward in the data but do not reflect the true causal mechanisms of the environment. We propo…

Cited by 0SourceScholar
2026

Towards Professional-Grade Financial Agents: Benchmarking, Tooling, and Structured Reasoning

ICML 2026poster

Financial reasoning requires precise execution. While Large Language Model (LLM) agents have shown encouraging progress in financial reasoning, their effectiveness in realistic financial workflows is severely hindered by the lack of holistic benchmarks and the fragility of unstructured reasoning. To…

Cited by 0SourceScholar
2026

UrbanMLLM: Joint Learning of Cross-view Imagery for Urban Understanding

ICML 2026poster

Comprehensive urban understanding requires integrating macroscopic spatial structure with fine-grained street-level semantics. However, existing urban Multimodal Large Language Models (MLLMs) primarily rely on satellite imagery, limiting their ability to capture detailed urban appearance and cross-v…

Cited by 0SourceScholar
2026

VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference

CVPR 2026

Long-context video understanding and generation pose a significant computational challenge for Transformer-based video models due to the quadratic complexity of self-attention. While existing sparse attention methods employ coarse-grained patterns to improve efficiency, they typically incur redundan

Cited by 0SourcecodeScholar
2026

WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens

CVPR 2026

Recent progress in multimodal large language models (MLLMs) has highlighted the challenge of efficiently bridging pre-trained Vision-Language Models (VLMs) with Diffusion Models. While methods using a fixed number of learnable query tokens offer computational efficiency, they suffer from task genera

Cited by 0SourceScholar
2026

WeightFlow: Learning Stochastic Dynamics via Evolving Weight of Neural Network

AAAI 2026technical

Modeling stochastic dynamics from discrete observations is a key interdisciplinary challenge. Existing methods often fail to estimate the continuous evolution of probability densities from trajectories or face the curse of dimensionality. To address these limitations, we presents a novel paradigm:

Cited by 0SourcePDFScholar
2026

iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework

ICML 2026poster

Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable environments for perception, reasoning, and action. Yet current research still lacks large-scale datasets and unified benchmarks to evaluate their phys…

Cited by 0SourceScholar
2025

A Large-scale Dataset and Benchmark for Commuting Origin-Destination Flow Generation

ICLR 2025poster

Commuting Origin-Destination~(OD) flows are critical inputs for urban planning and transportation, providing crucial information about the population residing in one region and working in another within an interested area. Due to the high cost of data collection, researchers have developed physical…

Cited by 0SourcePDFScholar
2025

ASER: Activation Smoothing and Error Reconstruction for Large Language Model Quantization

AAAI 2025technical

Quantization stands as a pivotal technique for large language model (LLM) serving, yet it poses significant challenges particularly in achieving effective low-bit quantization. The limited numerical mapping makes the quantized model produce a non-trivial error, bringing out intolerable performance d…

2025

AgentMove: A Large Language Model based Agentic Framework for Zero-shot Next Location Prediction

NAACL 2025long

Next location prediction plays a crucial role in various real-world applications. Recently, due to the limitation of existing deep learning methods, attempts have been made to apply large language models (LLMs) to zero-shot next location prediction task. However, they directly generate the final out…

2025

AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems

NeurIPS 2025spotlight

The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs’ advanced reasoning and role-playing capabilities to enable autonomous, adaptive decision-making. Unlike traditional recommendation approa…

Cited by 0SourcecodeScholar
2025

AgentSquare: Automatic LLM Agent Search in Modular Design Space

ICLR 2025poster

Recent advancements in Large Language Models (LLMs) have led to a rapid growth of agentic systems capable of handling a wide range of complex tasks. However, current research largely relies on manual, task-specific design, limiting their adaptability to novel tasks. In this paper, we introduce a new…

2025

Analyzing and Modeling LLM Response Lengths with Extreme Value Theory: Anchoring Effects and Hybrid Distributions

EMNLP 2025

We present a statistical framework for modeling and controlling large language model (LLM) response lengths using extreme value theory. Analyzing 14,301 GPT-4o responses across temperature and prompting conditions, with cross-validation on Qwen and DeepSeek architectures, we demonstrate that verbosi

Cited by 0SourcePDFScholar
2025

Automated Decision-Making on Networks with LLMs through Knowledge-Guided Evolution

IJCAI 2025

Effective decision-making on networks often relies on learning from graph-structured data, where Graph Neural Networks (GNNs) play a central role, but they take efforts to configure and tune. In this demo, we propose LLMNet, showing how to design GNN automated through Large Language Models. Our syst

2025

Automated Fine-Grained Mixture-of-Experts Quantization

ACL 2025finding

The Mixture of Experts (MoE) architecture enables efficient model scaling through conditional computation, where only subset of parameters are activated per input. However, this distributed architecture poses unprecedented challenges for model compression, as conventional quantization methods optimi…

2025

Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization

NeurIPS 2025poster

Large Vision-Language Models (LVLMs) have shown impressive performance across multi-modal tasks by encoding images into thousands of tokens. However, the large number of image tokens results in significant computational overhead, and the use of dynamic high-resolution inputs further increases this b…

Cited by 0SourcecodeScholar
2025

CD^2: Constrained Dataset Distillation for Few-Shot Class-Incremental Learning

IJCAI 2025

Few-shot class-incremental learning (FSCIL) receives significant attention from the public to perform classification continuously with a few training samples, which suffers from the key catastrophic forgetting problem. Existing methods usually employ an external memory to store previous knowledge an

Cited by 0SourcePDFScholar
2025

CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration

ICML 2025poster

Large language models (LLMs) have demonstrated strong capabilities in handling long-context tasks, but processing such long contexts remains challenging due to the substantial memory requirements and inference latency. In this work, we discover that certain attention heads exhibit sequential consist…

Cited by 0SourcePDFScholar
2025

CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space

EMNLP 2025

Embodied Question Answering (EQA) has primarily focused on indoor environments, leaving the complexities of urban settings—spanning environment, action, and perception—largely unexplored. To bridge this gap, we introduce CityEQA, a new task where an embodied agent answers open-vocabulary questions t

2025

CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory

ACL 2025long

Aerial vision-and-language navigation (VLN) — requiring drones to interpret natural language instructions and navigate complex urban environments — emerges as a critical embodied AI challenge that bridges human-robot interaction, 3D spatial reasoning, and real-world deployment. Although existing gro…

2025

Context-Aware Sentiment Forecasting via LLM-based Multi-Perspective Role-Playing Agents

ACL 2025long

User sentiment on social media reveals underlying social trends, crises, and needs. Researchers have analyzed users’ past messages to track the evolution of sentiments and reconstruct sentiment dynamics. However, predicting the imminent sentiment response of users to ongoing events remains understud…

2025

Defining and Evaluating Visual Language Models’ Basic Spatial Abilities: A Perspective from Psychometrics

ACL 2025long

The Theory of Multiple Intelligences underscores the hierarchical nature of cognitive capabilities. To advance Spatial Artificial Intelligence, we pioneer a psychometric framework defining five Basic Spatial Abilities (BSAs) in Visual Language Models (VLMs): Spatial Perception, Spatial Relation, Spa…

Cited by 0SourcePDFScholar
2025

Diffusion Transformers as Open-World Spatiotemporal Foundation Models

NeurIPS 2025poster

The urban environment is characterized by complex spatio-temporal dynamics arising from diverse human activities and interactions. Effectively modeling these dynamics is essential for understanding and optimizing urban systems. In this work, we introduce UrbanDiT, a foundation model for open-world u…

Cited by 0SourcecodeScholar
2025

Diffusion-based Identity-Preserving Facial Privacy Protection

ICASSP 2025accepted

The efficacy of facial recognition systems that utilize deep learning techniques has led to significant concerns over privacy, since they possess the capability to facilitate unauthorized monitoring of individuals in the digital realm. Current techniques for improving privacy are ineffective in prod…

Cited by 0SourceScholar
2025

Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition

ICASSP 2025accepted

Very low-resolution face recognition is challenging due to the serious loss of informative facial details in resolution degradation. Recent approaches based on knowledge distillation provide an effective solution by distilling knowledge from a well-trained teacher for high-resolution face recognitio…

Cited by 0SourceScholar
2025

Efficient Long Context Fine-tuning with Chunk Flow

ICML 2025poster

Long context fine-tuning of large language models(LLMs) involves training on datasets that are predominantly composed of short sequences and a small proportion of longer sequences. However, existing approaches overlook this long-tail distribution and employ training strategies designed specifically…

Cited by 0SourcePDFScholar
2025

Equivalence of Closed Chains to Open Chains: Virtual Decomposition Control Combined With Adaptive RBF Neural Network for Hydraulic Robot Legs

RA-L 2025

The joints of the hydraulic robot, driven by linear cylinders, form triangular closed-chain structures composed of the cylinders and passive rotational joints. This configuration complicates the complete dynamic modeling and increases the system's nonlinearity. To simplify the modeling process, conv

Cited by 1SourceScholar
2025

LLM-Explorer: A Plug-in Reinforcement Learning Policy Exploration Enhancement Driven by Large Language Models

NeurIPS 2025spotlight

Policy exploration is critical in reinforcement learning (RL), where existing approaches include $\epsilon$-greedy, Gaussian process, etc. However, these approaches utilize preset stochastic processes and are indiscriminately applied in all kinds of RL tasks without considering task-specific feature…

Cited by 0SourcecodeScholar
2025

Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query Optimization

ICML 2025poster

Fine-grained hashing has become a powerful solution for rapid and efficient image retrieval, particularly in scenarios requiring high discrimination between visually similar categories. To enable each hash bit to correspond to specific visual attributes, we propose a novel method that harnesses lear…

Cited by 0SourcePDFScholar
2025

MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector

AAAI 2025technical

The increasing parameters and expansive dataset of large lan- guage models (LLMs) highlight the urgent demand for a technical solution to audit the underlying privacy risks and copyright issues associated with LLMs. Existing studies have partially addressed this need through an exploration of the pr…

2025

Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning

EMNLP 2025

Large language models (LLMs) possess extensive world knowledge, including geospatial knowledge, which has been successfully applied to various geospatial tasks such as mobility prediction and social indicator prediction. However, LLMs often generate inaccurate geospatial knowledge, leading to geospa

2025

Open-Set Living Need Prediction with Large Language Models

ACL 2025finding

Living needs are the needs people generate in their daily lives for survival and well-being. On life service platforms like Meituan, user purchases are driven by living needs, making accurate living need predictions crucial for personalized service recommendations. Traditional approaches treat this…

Cited by 0SourcePDFScholar
2025

OpenCarbon: A Contrastive Learning-based Cross-Modality Neural Approach for High-Resolution Carbon Emission Prediction Using Open Data

IJCAI 2025

Accurately estimating high-resolution carbon emissions is crucial for effective emission governance and mitigation planning. While conventional methods for precise carbon accounting are hindered by substantial data collection efforts, the rise of open data and advanced learning techniques offers a p

2025

PID-controlled Langevin Dynamics for Faster Sampling on Generative Models

NeurIPS 2025poster

Langevin dynamics sampling suffers from extremely low generation speed, fundamentally limited by numerous fine-grained iterations to converge to the target distribution. We introduce PID-controlled Langevin Dynamics (PIDLD), a novel sampling acceleration algorithm that reinterprets the sampling proc…

Cited by 0SourcecodeScholar
2025

Planning with Linear Temporal Logic Specifications: Handling Quantifiable and Unquantifiable Uncertainty

ICRA 2025

This work studies the planning problem for robotic systems under both quantifiable and unquantifiable uncertainty. The objective is to enable the robotic systems to optimally fulfill high-level tasks specified by Linear Temporal Logic (LTL) formulas. To capture both types of uncertainty in a unified

Cited by 3SourcecodeScholar
2025

Predicting the Energy Landscape of Stochastic Dynamical System via Physics-informed Self-supervised Learning

ICLR 2025poster

Energy landscapes play a crucial role in shaping dynamics of many real-world complex systems. System evolution is often modeled as particles moving on a landscape under the combined effect of energy-driven drift and noise-induced diffusion, where the energy governs the long-term motion of the partic…

2025

PychoAgent: Psychology-driven LLM Agents for Explainable Panic Prediction on Social Media during Sudden Disaster Events

EMNLP 2025

Accurately predicting public panic sentiment on social media is crucial for proactive governance and crisis management. Current efforts on this problem face three main challenges: lack of finely annotated data hinders emotion prediction studies, unmodeled risk perception causes prediction inaccuraci

2025

Re-Attentional Controllable Video Diffusion Editing

AAAI 2025technical

Editing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. Recent studies have explored and exploited large-scale text-to-image diffusion models for text-guided video editing, resu…

2025

Reinforcement Learning with Adaptive Reward Modeling for Expensive-to-Evaluate Systems

ICML 2025poster

Training reinforcement learning (RL) agents requires extensive trials and errors, which becomes prohibitively time-consuming in systems with costly reward evaluations. To address this challenge, we propose adaptive reward modeling (AdaReMo) which accelerates RL training by decomposing the complicate…

2025

RoboScape: Physics-informed Embodied World Model

NeurIPS 2025spotlight

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models exhibit limited physical awareness, particularly in modelin…

Cited by 0SourcecodeScholar
2025

SAM-Aware Graph Prompt Reasoning Network for Cross-Domain Few-Shot Segmentation

AAAI 2025technical

The primary challenge of cross-domain few-shot segmentation (CD-FSS) is the domain disparity between the training and inference phases, which can exist in either the input data or the target classes. Previous models struggle to learn feature representations that generalize to various unknown domains…

2025

Satellites Reveal Mobility: A Commuting Origin-destination Flow Generator for Global Cities

NeurIPS 2025poster

Commuting Origin-destination (OD) flows, capturing daily population mobility of citizens, are vital for sustainable development across cities around the world. However, it is challenging to obtain the data due to the high cost of travel surveys and privacy concerns. Surprisingly, we find that satell…

Cited by 0SourcecodeScholar
2025

Semantic Alignment and Reinforcement for Data-Free Quantization of Vision Transformers

ICCV 2025poster

Data-free quantization (DFQ) enables model quantization without accessing real data, addressing concerns regarding data security and privacy. With the growing adoption of Vision Transformers (ViTs), DFQ for ViTs has garnered significant attention. However, existing DFQ methods exhibit two limitation…

2025

Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data Scheduling

NeurIPS 2025poster

Long-context supervised fine-tuning (Long-SFT) plays a vital role in enhancing the performance of large language models (LLMs) on long-context tasks. To smoothly adapt LLMs to long-context scenarios, this process typically entails training on mixed datasets containing both long and short sequences.…

Cited by 0SourceScholar
2025

Solving MDPs with LTLf+ and PPLTL+ Temporal Objectives

IJCAI 2025

The temporal logics LTLf+ and PPLTL+ have recently been introduced to express objectives over infinite traces. These logics are appealing because they match the expressive power of LTL on infinite traces while enabling efficient DFA-based techniques, which have been crucial to the scalability of rea

Cited by 0SourcePDFScholar
2025

Sparse Diffusion Autoencoder for Test-time Adapting Prediction of Complex Systems

NeurIPS 2025poster

Predicting the behavior of complex systems is critical in many scientific and engineering domains, and hinges on the model’s ability to capture their underlying dynamics. Existing methods encode the intrinsic dynamics of high-dimensional observations through latent representations and predict autore…

Cited by 0SourceScholar
2025

Symbolic regression via MDLformer-guided search: from minimizing prediction error to minimizing description length

ICLR 2025poster

Symbolic regression, a task discovering the formula best fitting the given data, is typically based on the heuristical search. These methods usually update candidate formulas to obtain new ones with lower prediction errors iteratively. However, since formulas with similar function shapes may have co…

2025

Synergistic Tensor and Pipeline Parallelism

NeurIPS 2025poster

In the machine learning system, the hybrid model parallelism combining tensor parallelism (TP) and pipeline parallelism (PP) has become the dominant solution for distributed training of Large Language Models~(LLMs) and Multimodal LLMs (MLLMs). However, TP introduces significant collective communicat…

Cited by 0SourcecodeScholar
2025

TOPP-DWR: Time-Optimal Path Parameterization of Differential-Driven Wheeled Robots Considering Piecewise-Constant Angular Velocity Constraints

IROS 2025

Differential-driven wheeled robots (DWR) represent the quintessential type of mobile robots and find extensive applications across the robotic field. Most high-performance control approaches for DWR explicitly utilize the linear and angular velocities of the trajectory as control references. However

Cited by 1SourceScholar
2025

Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models

ICML 2025poster

Chain-of-Thought (CoT) technique has proven effective in improving the performance of large language models (LLMs) on complex reasoning tasks. However, the performance gains are inconsistent across different tasks, and the underlying mechanism remains a long-standing research question. In this work,…

2025

TrajAgent: An LLM-Agent Framework for Trajectory Modeling via Large-and-Small Model Collaboration

NeurIPS 2025poster

Trajectory modeling, which includes research on trajectory data pattern mining and future prediction, has widespread applications in areas such as life services, urban transportation, and public administration. Numerous methods have been proposed to address specific problems within trajectory modeli…

Cited by 0SourcecodeScholar
2025

Treasures in Discarded Weights for LLM Quantization

AAAI 2025technical

In recent years, large language models (LLMs) have developed rapidly and revolutionized natural language processing. However, high storage overhead and computing costs limit LLM deployment in resource-constrained environments. Quantization algorithms can effectively compress LLMs and accelerate infe…

Cited by 0SourcePDFScholar
2025

Unveiling the Power of Noise Priors: Enhancing Diffusion Models for Mobile Traffic Prediction

IJCAI 2025

Accurate prediction of mobile traffic,i.e., network traffic from cellular base stations, is crucial for optimizing network performance and supporting urban development. However, the non-stationary nature of mobile traffic, driven by human activity and environmental changes, leads to both regular pat

2025

UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence

ICCV 2025poster

Urban research involves a wide range of scenarios and tasks that require the understanding of multi-modal data, such as structured geospatial data, trajectory data, satellite image data, and street view image data. Current methods often focus on specific data types and lack a unified framework in ur…

2025

UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

ACL 2025long

Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored. We introduce a benchmark to evaluate whether video-large language models (Video-LLMs) can naturally process continuous first-person v…

Cited by 0SourcePDFScholar
2024

EconAgent: Large Language Model-Empowered Agents for Simulating Macroeconomic Activities

ACL 2024long

The advent of artificial intelligence has led to a growing emphasis on data-driven modeling in macroeconomics, with agent-based modeling (ABM) emerging as a prominent bottom-up simulation paradigm. In ABM, agents (*e.g.*, households, firms) interact within a macroeconomic environment, collectively g…

2024

Estimating On-Road Transportation Carbon Emissions from Open Data of Road Network and Origin-Destination Flow Data

AAAI 2024technical

Accounting for over 20% of the total carbon emissions, the precise estimation of on-road transportation carbon emissions is crucial for carbon emission monitoring and efficient mitigation policy formulation. However, existing estimation methods typically depend on hard-to-collect individual statisti…

2024

From Pixels to Progress: Generating Road Network from Satellite Imagery for Socioeconomic Insights in Impoverished Areas

IJCAI 2024poster

The Sustainable Development Goals (SDGs) aim to resolve societal challenges, such as eradicating poverty and improving the lives of vulnerable populations in impoverished areas. Those areas rely on road infrastructure construction to promote accessibility and economic development. Although publicly…

2024

GLOP: Learning Global Partition and Local Construction for Solving Large-Scale Routing Problems in Real-Time

AAAI 2024technical

The recent end-to-end neural solvers have shown promise for small-scale routing problems but suffered from limited real-time scaling-up performance. This paper proposes GLOP (Global and Local Optimization Policies), a unified hierarchical framework that efficiently scales toward large-scale routing…

2024

Giving Control Back to Models: Enabling Offensive Language Detection Models to Autonomously Identify and Mitigate Biases

EMNLP 2024finding

The rapid development of social media has led to an increase in online harassment and offensive speech, posing significant challenges for effective content moderation. Existing automated detection models often exhibit a bias towards predicting offensive speech based on specific vocabulary, which not…

Cited by 0SourcePDFScholar
2024

HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction

NeurIPS 2024poster

Citation networks are critical infrastructures of modern science, serving as intricate webs of past literature and enabling researchers to navigate the knowledge production system. To mine information hiding in the link space of such networks, predicting which previous papers (candidates) will a new…

2024

Long-tailed Object Detection Pretraining: Dynamic Rebalancing Contrastive Learning with Dual Reconstruction

NeurIPS 2024poster

Pre-training plays a vital role in various vision tasks, such as object recognition and detection. Commonly used pre-training methods, which typically rely on randomized approaches like uniform or Gaussian distributions to initialize model parameters, often fall short when confronted with long-taile…

Cited by 1SourcePDFScholar
2024

Long-term Detection and Monitory of Chinese Urban Village Using Satellite Imagery

IJCAI 2024poster

Urban villages are areas filled with rural-like improvised structures in Chinese cities, usually housing the most vulnerable groups. Under the guidance of the Sustainable Development Goals (SDGs), the Chinese government initiated renewal and redevelopment projects, underscoring the meticulous mapp…

2024

MTA: A Lightweight Multilingual Text Alignment Model for Cross-Language Visual Word Sense Disambiguation

ICASSP 2024accepted

Visual Word Sense Disambiguation (Visual-WSD), as a sub-task of fine-grained image-text retrieval, requires a high level of language-vision understanding to capture and exploit the nuanced relationships between text and visual features. However, the cross-linguistic background only with limited cont…

Cited by 0SourceScholar
2024

Masked Face Recognition with Generative-to-Discriminative Representations

ICML 2024spotlight

Masked face recognition is important for social good but challenged by diverse occlusions that cause insufficient or inaccurate representations. In this work, we propose a unified deep network to learn generative-to-discriminative representations for facilitating masked face recognition. To this end…

Cited by 4SourcePDFScholar
2024

Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration

NeurIPS 2024poster

Membership Inference Attacks (MIA) aim to infer whether a target data record has been utilized for model training or not. Existing MIAs designed for large language models (LLMs) can be bifurcated into two types: reference-free and reference-based attacks. Although reference-based attacks appear prom…

2024

Patch-Aware Sample Selection for Efficient Masked Image Modeling

AAAI 2024technical

Nowadays sample selection is drawing increasing attention. By extracting and training only on the most informative subset, sample selection can effectively reduce the training cost. Although sample selection is effective in conventional supervised learning, applying it to Masked Image Modeling (MIM)…

Cited by 4SourcePDFScholar
2024

PolCLIP: A Unified Image-Text Word Sense Disambiguation Model via Generating Multimodal Complementary Representations

ACL 2024long

Word sense disambiguation (WSD) can be viewed as two subtasks: textual word sense disambiguation (Textual-WSD) and visual word sense disambiguation (Visual-WSD). They aim to identify the most semantically relevant senses or images to a given context containing ambiguous target words. However, existi…

2024

Social Physics Informed Diffusion Model for Crowd Simulation

AAAI 2024technical

Crowd simulation holds crucial applications in various domains, such as urban planning, architectural design, and traffic arrangement. In recent years, physics-informed machine learning methods have achieved state-of-the-art performance in crowd simulation but fail to model the heterogeneity and mul…

2024

Spatio-Temporal Few-Shot Learning via Diffusive Neural Network Generation

ICLR 2024poster

Spatio-temporal modeling is foundational for smart city applications, yet it is often hindered by data scarcity in many cities and regions. To bridge this gap, we propose a novel generative pre-training framework, GPD, for spatio-temporal few-shot learning with urban knowledge transfer. Unlike conve…

2024

UV-SAM: Adapting Segment Anything Model for Urban Village Identification

AAAI 2024technical

Urban villages, defined as informal residential areas in or around urban centers, are characterized by inadequate infrastructures and poor living conditions, closely related to the Sustainable Development Goals (SDGs) on poverty, adequate housing, and sustainable cities. Traditionally, governments h…

2024

VulnerabilityMap: An Open Framework for Mapping Vulnerability among Urban Disadvantaged Populations in the United States

IJCAI 2024poster

Cities are crucibles of numerous opportunities, but also hotbeds of inequality. The plight of disadvantaged populations who are ``left behind'' within urban environments has been an increasingly pressing concern, which poses substantial threats to the realization of the UN SDG agenda. However, a com…

2023

DeepACO: Neural-enhanced Ant Systems for Combinatorial Optimization

NeurIPS 2023poster

Ant Colony Optimization (ACO) is a meta-heuristic algorithm that has been successfully applied to various Combinatorial Optimization Problems (COPs). Traditionally, customizing ACO for a specific problem requires the expert design of knowledge-driven heuristics. In this paper, we propose DeepACO, a…

2023

DynaMS: Dyanmic Margin Selection for Efficient Deep Learning

ICLR 2023poster

The great success of deep learning is largely driven by training over-parameterized models on massive datasets. To avoid excessive computation, extracting and training only on the most informative subset is drawing increasing attention. Nevertheless, it is still an open question how to select such a…

Cited by 5SourcePDFScholar
2023

Efficient Hyper-parameter Optimization with Cubic Regularization

NeurIPS 2023poster

As hyper-parameters are ubiquitous and can significantly affect the model performance, hyper-parameter optimization is extremely important in machine learning. In this paper, we consider a sub-class of hyper-parameter optimization problems, where the hyper-gradients are not available. Such problems…

Cited by 2SourcePDFScholar
2023

Learning Symbolic Models for Graph-structured Physical Mechanism

ICLR 2023poster

Graph-structured physical mechanisms are ubiquitous in real-world scenarios, thus revealing underneath formulas is of great importance for scientific discovery. However, classical symbolic regression methods fail on this task since they can only handle input-output pairs that are not graph-structure…

Cited by 14SourcePDFScholar
2023

PateGail: A Privacy-Preserving Mobility Trajectory Generator with Imitation Learning

AAAI 2023technical

Generating human mobility trajectories is of great importance to solve the lack of large-scale trajectory data in numerous applications, which is caused by privacy concerns. However, existing mobility trajectory generation methods still require real-world human trajectories centrally collected as th…

2023

Revisit Finetuning strategy for Few-Shot Learning to Transfer the Emdeddings

ICLR 2023poster

Few-Shot Learning (FSL) aims to learn a simple and effective bias on limited novel samples. Recently, many methods have been focused on re-training a randomly initialized linear classifier to adapt it to the novel features extracted by the pre-trained feature extractor (called Linear-Probing-based m…

2023

Style Projected Clustering for Domain Generalized Semantic Segmentation

CVPR 2023poster

Existing semantic segmentation methods improve generalization capability, by regularizing various images to a canonical feature space. While this process contributes to generalization, it weakens the representation inevitably. In contrast to existing methods, we instead utilize the difference betwee…

Cited by 40SourcePDFScholar
2023

Unidirectional-Road-Network-Based Global Path Planning for Cleaning Robots in Semi-Structured Environments

ICRA 2023poster

Practical global path planning is critical for commercializing cleaning robots working in semi-structured environments. In the literature, global path planning methods for free space usually focus on path length and neglect the traffic rule constraints of the environments, which leads to high-freque…

Cited by 3SourceScholar
2022

A Formal Model for Multiagent Q-Learning Dynamics on Regular Graphs

IJCAI 2022poster

Modeling the dynamics of multi-agent learning has long been an important research topic. The focus of previous research has been either on 2-agent settings or well-mixed infinitely large agent populations. In this paper, we consider the scenario where n Q-learning agents locate on regular graphs, su…

Cited by 36SourcePDFScholar
2022

Efficient Hyper-parameter Search for Knowledge Graph Embedding

ACL 2022long

While hyper-parameters (HPs) are important for knowledge graph (KG) learning, existing methods fail to search them efficiently. To solve this problem, we first analyze the properties of different HPs and measure the transfer ability from small subgraph to the full graph. Based on the analysis, we pr…

2022

MAPDP: Cooperative Multi-Agent Reinforcement Learning to Solve Pickup and Delivery Problems

AAAI 2022technical

Cooperative Pickup and Delivery Problem (PDP), as a variant of the typical Vehicle Routing Problems (VRP), is an important formulation in many real-world applications, such as on-demand delivery, industrial warehousing, etc. It is of great importance to efficiently provide high-quality solutions of…

Cited by 60SourcePDFScholar
2022

Regularized Latent Space Exploration for Discriminative Face Super-Resolution

ICASSP 2022accepted

Learning face super-resolution models is challenged in many practical scenarios where high-resolution and low-resolution face pairs usually are difficult to collect for training examples. Recent self-supervised approach provides a feasible solution by using low-resolution faces to guide the generati…

Cited by 0SourceScholar
2022

Revisiting and Advancing Chinese Natural Language Understanding with Accelerated Heterogeneous Knowledge Pre-training

EMNLP 2022industry

Recently, knowledge-enhanced pre-trained language models (KEPLMs) improve context-aware representations via learning from structured relations in knowledge bases, and/or linguistic knowledge from syntactic or dependency analysis. Unlike English, there is a lack of high-performing open-source Chinese…

2022

Towards Real-World HDRTV Reconstruction: A Data Synthesis-Based Approach

ECCV 2022poster

"Existing deep learning based HDRTV reconstruction methods assume one kind of tone mapping operators (TMOs) as the degradation procedure to synthesize SDRTV-HDRTV pairs for supervised training. In this paper, we argue that, although traditional TMOs exploit efficient dynamic range compression priors…

2021

AttnMove: History Enhanced Trajectory Recovery via Attentional Network

AAAI 2021technical

A considerable amount of mobility data has been accumulated due to the proliferation of location-based service. Nevertheless, compared with mobility data from transportation systems like the GPS module in taxis, this kind of data is commonly sparse in terms of individual trajectories in the sense th…

2021

Consistent Instance False Positive Improves Fairness in Face Recognition

CVPR 2021poster

Demographic bias is a significant challenge in practical face recognition systems. Several methods have been proposed to reduce the bias, which rely on accurate demographic annotations. However, such annotations are usually not available in real scenarios. Moreover, these methods are explicitly desi…

Cited by 66PDFcodeScholar
2021

Learning Normal Dynamics in Videos With Meta Prototype Network

CVPR 2021poster

Frame reconstruction (current or future frames) based on Auto-Encoder (AE) is a popular method for video anomaly detection. With models trained on the normal data, the reconstruction errors of anomalous scenes are usually much larger than those of normal ones. Previous methods introduced the memory…

Cited by 219PDFcodeScholar
2021

Progressive Feature Interaction Search for Deep Sparse Network

NeurIPS 2021poster

Deep sparse networks (DSNs), of which the crux is exploring the high-order feature interactions, have become the state-of-the-art on the prediction task with high-sparsity features. However, these models suffer from low computation efficiency, including large model size and slow model inference, whi…

Cited by 16SourcePDFScholar
2021

Real-Time Image Enhancer via Learnable Spatial-Aware 3D Lookup Tables

ICCV 2021poster

Recently, deep learning-based image enhancement algorithms achieved state-of-the-art (SOTA) performance on several publicly available datasets. However, most existing methods fail to meet practical requirements either for visual perception or for computation efficiency, especially for high-resolutio…

Cited by 98PDFScholar
2021

SDD-FIQA: Unsupervised Face Image Quality Assessment With Similarity Distribution Distance

CVPR 2021poster

In recent years, Face Image Quality Assessment (FIQA) has become an indispensable part of the face recognition system to guarantee the stability and reliability of recognition performance in an unconstrained scenario. For this purpose, the FIQA method should consider both the intrinsic property and…

Cited by 140PDFcodeScholar
2021

Synthesizing Good-Enough Strategies for LTLf Specifications

IJCAI 2021poster

We consider the problem of synthesizing good-enough (GE)-strategies for linear temporal logic (LTL) over finite traces or LTLf for short. The problem of synthesizing GE-strategies for an LTL formula φ over infinite traces reduces to the problem of synthesizing winning strategies for the formula (∃O…

2020

A Sequential Convolution Network for Population Flow Prediction with Explicitly Correlation Modelling

IJCAI 2020poster

Population flow prediction is one of the most fundamental components in many applications from urban management to transportation schedule. It is challenging due to the complicated spatial-temporal correlation.While many studies have been done in recent years, they fail to simultaneously and effecti…

Cited by 0SourcePDFScholar
2020

Multi-View Joint Graph Representation Learning for Urban Region Embedding

IJCAI 2020poster

The increasing amount of urban data enable us to investigate urban dynamics, assist urban planning, and eventually, make our cities more livable and sustainable. In this paper, we focus on learning an embedding space from urban data for urban regions. For the first time, we propose a multi-view join…

2020

Simplify and Robustify Negative Sampling for Implicit Collaborative Filtering

NeurIPS 2020poster

Negative sampling approaches are prevalent in implicit collaborative filtering for obtaining negative labels from massive unlabeled data. As two major concerns in negative sampling, efficiency and effectiveness are still not fully achieved by recent works that use complicate structures and overlook ri…

2019

Self-Supervised Representation Learning From Videos for Facial Action Unit Detection

CVPR 2019oral

In this paper, we aim to learn discriminative representation for facial action unit (AU) detection from large amount of videos without manual annotations. Inspired by the fact that facial actions are the movements of facial muscles, we depict the movements as the transformation between two face imag…

Cited by 129PDFcodeScholar