← Search

Wenlong Zhang

45 accepted papers

2026

Deep Research Arena: The First Exam of LLMs’ Research Abilities via Seminar-Grounded Tasks

AAAI 2026technical

Deep research agents have attracted growing attention for their potential to orchestrate multi-stage research workflows, spanning literature synthesis, methodological design, and empirical verification. Despite these strides, evaluating their research capability faithfully is rather challenging due

Cited by 0SourcePDFScholar
2026

Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents

ICLR 2026poster

Earth observation (EO) is essential for understanding the evolving states of the Earth system. Although recent MLLMs have advanced EO research, they still lack the capability to tackle complex tasks that require multi-step reasoning and the use of domain-specific tools. Agent-based methods offer a…

Cited by 0SourcecodeScholar
2026

EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language Models

ICLR 2026poster

Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or cover isolated subdomains, lacking holistic evaluation…

Cited by 0SourceScholar
2026

Eigen-1: Scientific Reasoning through Adaptive Multi-Agent Refinement and Monitor-based RAG

ICLR 2026poster

Large language models (LLMs) have recently shown strong progress on scientific reasoning, yet two major bottlenecks remain. First, explicit retrieval fragments reasoning, imposing a hidden tool tax of extra tokens and steps. Second, multi-agent pipelines often dilute strong solutions by averaging ac…

Cited by 0SourcecodeScholar
2026

FinPercep-RM: A Fine-grained Reward Model and Co-evolutionary Curriculum for RL-based Real-world Super-Resolution

CVPR 2026

Inspired by the success of Reinforcement Learning with Human Feedback (RLHF) in image generation, recent work has adapted reward-based learning to image super-resolution (ISR) by using Image Quality Assessment (IQA) models as rewards. However, existing IQA models typically output only a single globa

Cited by 0SourcecodeScholar
2026

Omni-Weather: Unified Multimodal Foundation Model for Weather Generation and Understanding

ICLR 2026poster

Weather modeling requires both accurate prediction and mechanistic interpretation, yet existing methods treat these goals in isolation, separating generation from understanding. To address this gap, we present Omni-Weather, the first multimodal foundation model that unifies weather generation and un…

Cited by 0SourcecodeScholar
2026

PICABench: How Far are We from Physical Realistic Image Editing?

ICLR 2026poster

Image editing has achieved remarkable progress recently. Modern editing models could already follow complex instructions to manipulate the original content. However, beyond completing the editing instructions, the accompanying physical effects are the key to the generation realism. For example, remo…

Cited by 0SourcecodeScholar
2026

SynWeather: Weather Observation Data Synthesis Across Multiple Regions and Variables via a General Diffusion Transformer

AAAI 2026technical

With the advancement of meteorological instruments, abundant data has become available. However, due to instruments’ intrinsic limitations such as environmental sensitivity and orbital constraints, raw data often suffer from temporal or spatial gaps, making it urgent to leverage data synthesis tech

Cited by 0SourcePDFScholar
2026

Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing Understanding

ICML 2026poster

Multimodal reasoning for ultra-high-resolution (UHR) remote sensing (RS) is usually bottlenecked by visual evidence acquisition: the model necessities localizing tiny task-relevant regions in massive pixel spaces. While Agentic Reinforcement Learning with Verifiable Rewards (RLVR) using zoom-in tool…

Cited by 0SourceScholar
2026

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

ICML 2026spotlight

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image features remains limited. In this work, we present UniPercept-Bench, a unified fr…

Cited by 0SourceScholar
2025

Align-DA: Align Score-based Atmospheric Data Assimilation with Multiple Preferences

NeurIPS 2025poster

Data assimilation (DA) aims to estimate the full state of a dynamical system by combining partial and noisy observations with a prior model forecast, commonly referred to as the background. In atmospheric applications, this problem is fundamentally ill-posed due to the sparsity of observations relat…

Cited by 0SourceScholar
2025

DAWP: A framework for global observation forecasting via Data Assimilation and Weather Prediction in satellite observation space

NeurIPS 2025poster

Weather prediction is a critical task for human society, where impressive progress has been made by training artificial intelligence weather prediction (AIWP) methods with reanalysis data. However, reliance on reanalysis data limits the AIWPs with shortcomings, including data assimilation biases an…

Cited by 0SourceScholar
2025

Decouple to Reconstruct: High Quality UHD Restoration via Active Feature Disentanglement and Reversible Fusion

ICCV 2025poster

Ultra-high-definition (UHD) image restoration often faces computational bottlenecks and information loss due to its extremely high resolution. Existing studies based on Variational Autoencoders (VAE) improve efficiency by transferring the image restoration process from pixel space to latent space. H…

Cited by 0SourcePDFScholar
2025

Design, Contact Modeling, and Collision-Inclusive Planning of a Dual-Stiffness Aerial RoboT (DART)

ICRA 2025

Collision-resilient quadrotors have gained significant attention given their potential for operating in cluttered environments and leveraging impacts to perform agile maneuvers. However, existing designs are typically single-mode: either safeguarded by propeller guards that prevent deformation or de

Cited by 0SourceScholar
2025

DiffSR: Learning Radar Reflectivity Synthesis via Diffusion Model from Satellite Observations

ICASSP 2025accepted

Weather radar data synthesis can fill in data for areas where ground observations are missing. Existing methods often employ reconstruction-based approaches with MSE loss to reconstruct radar data from satellite observation. However, such methods lead to over-smoothing, which hinders the generation…

Cited by 0SourceScholar
2025

Latent Harmony: Synergistic Unified UHD Image Restoration via Latent Space Regularization and Controllable Refinement

NeurIPS 2025poster

Ultra-High Definition (UHD) image restoration struggles to balance computational efficiency and detail retention. While Variational Autoencoders (VAEs) offer improved efficiency by operating in the latent space, with the Gaussian variational constraint, this compression preserves semantics but sacri…

Cited by 0SourceScholar
2025

Personalized Reinforcement Learning Control of Soft Robotic Exosuit for Assisting Human Normative Walking with Reduced Effort

IROS 2025

Wearable lower limb robots are promising technologies to assist human locomotion. Soft robotic exosuits introduce a promising solution for reducing muscle effort and metabolic cost as they are lightweight, transparent and inherently safe. However, it is challenging to effectively control such soft r

Cited by 0SourceScholar
2025

RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts

NeurIPS 2025poster

Quality analysis of weather forecasts is an essential topic in meteorology. Although traditional score-based evaluation metrics can quantify certain forecast errors, they are still far from meteorological experts in terms of descriptive capability, interpretability, and understanding of dynamic evol…

Cited by 0SourceScholar
2025

Reinforcement Learning Control of a Physical Robot Device for Assisted Human Walking without a Simulator

ICML 2025poster

This study presents an innovative reinforcement learning (RL) control approach to facilitate soft exosuit-assisted human walking. Our goal is to address the ongoing challenges in developing reliable RL-based methods for controlling physical devices. To overcome key obstacles—such as limited data, th…

Cited by 0SourcePDFScholar
2025

Risk-Sensitive Theory of Mind: Coordinating with Agents of Unknown Bias using Cumulative Prospect Theory

ICML 2025poster

Humans are often modeled as rational actors by interactive agents when they are in fact frequently observed to make biased decisions. This erroneous assumption may cause an agent’s model of the human to fail, especially when interaction occurs in bias-inducing settings that prompt risky decisions. T…

Cited by 0SourcePDFScholar
2025

Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning

NeurIPS 2025poster

Scientific discoveries increasingly rely on complex multimodal reasoning based on information-intensive scientific data and domain-specific expertise. Empowered by expert-level scientific benchmarks, scientific Multimodal Large Language Models (MLLMs) hold the potential to significantly enhance this…

Cited by 0SourceScholar
2025

Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation

NeurIPS 2025poster

Recent studies have demonstrated the importance of high-quality visual representations in image generation and have highlighted the limitations of generative models in image understanding. As a generative paradigm originally designed for natural language, autoregressive models face similar challenge…

Cited by 0SourceScholar
2024

GRIDS: Grouped Multiple-Degradation Restoration with Image Degradation Similarity

ECCV 2024poster

"Traditional single-task image restoration methods excel in handling specific degradation types but struggle with multiple degradations. To address this limitation, we propose Grouped Restoration with Image Degradation Similarity (GRIDS), a novel approach that harmonizes the competing objectives inh…

Cited by 5SourcePDFScholar
2024

Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid Modeling

NeurIPS 2024poster

Data-driven artificial intelligence (AI) models have made significant advancements in weather forecasting, particularly in medium-range and nowcasting. However, most data-driven weather forecasting models are black-box systems that focus on learning data mapping rather than fine-grained physical evo…

2024

SEAL: A Framework for Systematic Evaluation of Real-World Super-Resolution

ICLR 2024spotlight

Real-world Super-Resolution (Real-SR) methods focus on dealing with diverse real-world images and have attracted increasing attention in recent years. The key idea is to use a complex and high-order degradation model to mimic real-world degradations. Although they have achieved impressive results i…

2024

UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation

EMNLP 2024main

The fashion domain encompasses a variety of real-world multimodal tasks, including multimodal retrieval and multimodal generation. The rapid advancements in artificial intelligence generated content, particularly in technologies like large language models for text generation and diffusion models for…

2023

Approximating Discontinuous Nash Equilibrial Values of Two-Player General-Sum Differential Games

ICRA 2023poster

Finding Nash equilibrial policies for two-player differential games requires solving Hamilton-Jacobi-Isaacs (HJI) PDEs. Self-supervised learning has been used to approximate solutions of such PDEs while circumventing the curse of dimensionality. However, this method fails to learn discontinuous PDE…

Cited by 8SourceScholar
2023

Design, Characterization and Control of a Whole-body Grasping and Perching (WHOPPEr) Drone

IROS 2023poster

Flying robots can exploit perching abilities to position themselves on strategically-chosen locations and monitor the areas of interest from a critical vantage point. Moreover, they can significantly extend their battery life by turning off the propulsion systems when carrying out a surveillance mis…

Cited by 2SourceScholar
2023

Implementation of a Cosserat Rod-Based Configuration Tracking Controller on a Multi-Segment Soft Robotic Arm

IROS 2023poster

Controlling soft continuum robotic arms is challenging due to their hyper-redundancy and dexterity. In this paper we experimentally demonstrate, for the first time, closed-loop control of the configuration space variables of a soft robotic arm, composed of independently controllable segments, using…

Cited by 2SourceScholar
2023

Implicit Projection: Improving Team Situation Awareness for Tacit Human-Robot Interaction via Virtual Shadows

IROS 2023poster

Fluent teaming is characterized by tacit interaction without explicit communication. Such interaction requires team situation awareness (TSA) to facilitate. However, existing approaches often rely on explicit communication (such as visual projection) to support TSA, resulting in a paradox. In this p…

Cited by 4SourceScholar
2023

Real-World Image Super-Resolution as Multi-Task Learning

NeurIPS 2023poster

In this paper, we take a new look at real-world image super-resolution (real-SR) from a multi-task learning perspective. We demonstrate that the conventional formulation of real-SR can be viewed as solving multiple distinct degradation tasks using a single shared model. This poses a challenge known…

2023

Recon: Reducing Conflicting Gradients From the Root For Multi-Task Learning

ICLR 2023poster

A fundamental challenge for multi-task learning is that different tasks may conflict with each other when they are solved jointly, and a cause of this phenomenon is conflicting gradients during optimization. Recent works attempt to mitigate the influence of conflicting gradients by directly altering…

2022

Bounded Rational Game-theoretical Modeling of Human Joint Actions with Incomplete Information

IROS 2022poster

As humans and robots start to collaborate in close proximity, robots are tasked to perceive, comprehend, and anticipate human partners' actions, which demands a predictive model to describe how humans collaborate with each other in joint actions. Previous studies either simplify the collaborative ta…

Cited by 5SourceScholar
2022

Probabilistic Consensus on Feature Distribution for Multi-Robot Systems With Markovian Exploration Dynamics

RA-L 2022

In this letter, we present a consensus-based decentralized multi-robot approach to reconstruct a discrete distribution of features, modeled as an occupancy grid map, that represent information contained in a bounded planar 2D environment, such as visual cues used for navigation or semantic labels as

Cited by 4SourceScholar
2021

Overcoming Catastrophic Forgetting in Incremental Few-Shot Learning by Finding Flat Minima

NeurIPS 2021spotlight

This paper considers incremental few-shot learning, which requires a model to continually recognize new categories with only a few examples provided. Our study shows that existing methods severely suffer from catastrophic forgetting, a well-known problem in incremental learning, which is aggravated…

2021

When Shall I Be Empathetic? The Utility of Empathetic Parameter Estimation in Multi-Agent Interactions

ICRA 2021poster

Human-robot interactions (HRI) can be modeled as differential games with incomplete information, where each agent holds private reward parameters. Due to the open challenge in finding perfect Bayesian equilibria of such games, existing studies often decouple the belief and physical dynamics by itera…

Cited by 11SourceScholar
2020

Design and Control of SQUEEZE: A Spring-augmented QUadrotor for intEractions with the Environment to squeeZE-and-fly

IROS 2020poster

This paper presents the design and control of a novel quadrotor with a variable geometry to physically interact with cluttered environments and fly through narrow gaps and passageways. This compliant quadrotor with passive morphing capabilities is designed using torsional springs at every arm hinge…

Cited by 29SourceScholar
2020

Predictive Modeling of Periodic Behavior for Human-Robot Symbiotic Walking

ICRA 2020poster

We propose in this paper Periodic Interaction Primitives - a probabilistic framework that can be used to learn compact models of periodic behavior. Our approach extends existing formulations of Interaction Primitives to periodic movement regimes, i.e., walking. We show that this model is particularl…

Cited by 11SourceScholar
2020

Towards Untethered Soft Pneumatic Exosuits Using Low-Volume Inflatable Actuator Composites and a Portable Pneumatic Source

RA-L 2020

The application of pneumatic soft robots is limited by factors such as operational pressure and air flow rates. Pneumatic soft robots are typically tethered in nature due to the high energy costs for actuation as well as the lack of portable pneumatic sources capable of providing high pressures and

Cited by 25SourceScholar
2019

Design, Characterization, and Mechanical Programming of Fabric-Reinforced Textile Actuators for a Soft Robotic Hand

IROS 2019poster

In this paper, we present the design, fabrication, and evaluation of robust, fabric-reinforced textile actuators, which are capable of performing a variety and combination of motions, such as axial extension, radial expansion, bending, and twisting along its central axis. A simple fabrication proced…

Cited by 20SourceScholar
2019

Fabric Soft Poly-Limbs for Physical Assistance of Daily Living Tasks

ICRA 2019poster

This paper presents the design and development of a highly articulated, continuum, wearable, fabric-based Soft Poly-Limb (fSPL). This fabric soft arm acts as an additional limb that provides the wearer with mobile manipulation assistance through the use of soft actuators made with high-strength infl…

Cited by 66SourceScholar
2019

How Shall I Drive? Interaction Modeling and Motion Planning towards Empathetic and Socially-Graceful Driving

ICRA 2019poster

While intelligence of autonomous vehicles (AVs) has significantly advanced in recent years, accidents involving AVs suggest that these autonomous systems lack gracefulness in driving when interacting with human drivers. In the setting of a two-player game, we propose model predictive control based o…

Cited by 20SourceScholar
2019

RankSRGAN: Generative Adversarial Networks With Ranker for Image Super-Resolution

ICCV 2019oral

Generative Adversarial Networks (GAN) have demonstrated the potential to recover realistic details for single image super-resolution (SISR). To further improve the visual quality of super-resolved results, PIRM2018-SR Challenge employed perceptual metrics to assess the perceptual quality, such as PI…

Cited by 410PDFcodeScholar
2018

Optimal Collision-Free Robot Trajectory Generation Based on Time Series Prediction of Human Motion

RA-L 2018

In this letter, we propose that the joint motion of a human worker doing repetitive work could be predicted using a time series model. With a motion capture system, the elbow joint rotation data are collected and used to fit an autoregressive model. An online parameter adaptation algorithm is employ

Cited by 41SourceScholar