← Search

Yiming Sun

15 accepted papers

2026

CHAM-net: A Contrastive Hierarchical Adaptive Meta-network for Robust Global Methane Flux Prediction

IJCAI 2026

Methane is a potent greenhouse gas that significantly contributes to global warming. However, accurately estimating global methane emissions and consumption remains challenging due to the complex interactions among environmental drivers that may vary across spatial and temporal scales. Prior data-dr

Cited by 0Scholar
2026

CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion

AAAI 2026technical

Infrared and visible image fusion generates all-weather perception-capable images by combining complementary modalities, enhancing environmental awareness for intelligent unmanned systems. Existing methods either focus on pixel-level fusion while overlooking downstream task adaptability or implicitl

Cited by 0SourcePDFScholar
2026

SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse

AAAI 2026technical

Despite Video Large Language Models (Video-LLMs) having rapidly advanced in recent years, perceptual hallucinations pose a substantial safety risk, which severely restricts their real-world applicability. While several methods for hallucination mitigation have been proposed, they often compromise th

Cited by 0SourcePDFScholar
2025

A Combined Intrusion Strategy Based on Apollonius Circle for Multiple Mobile Robots in Attack-Defense Scenario

RA-L 2025

The multi-agent attack-defense game has become a hot issue in recent years. However, it is still a challenge to design an efficient intrusion strategy when the intruder has a limited detection range. In this letter, a combined intrusion strategy based on Apollonius circle is proposed, which decompos

Cited by 7SourceScholar
2025

Detect-and-Guide: Self-regulation of Diffusion Models for Safe Text-to-Image Generation via Guideline Token Optimization

CVPR 2025poster

Text-to-image diffusion models have achieved state-of-the-art results in synthesis tasks; however, there is a growing concern about their potential misuse in creating harmful content. To mitigate these risks, post-hoc model intervention techniques, such as concept unlearning and safety guidance, hav…

Cited by 2SourcePDFScholar
2025

Multi-Scale Graph Learning for Anti-Sparse Downscaling

AAAI 2025technical

Water temperature can vary substantially even across short distances within the same sub-watershed. Accurate prediction of stream water temperature at fine spatial resolutions (i.e., fine scales, ≤ 1 km) enables precise interventions to maintain water quality and protect aquatic habitats. Although s…

Cited by 0SourcePDFScholar
2025

Task-Gated Multi-Expert Collaboration Network for Degraded Multi-Modal Image Fusion

ICML 2025poster

Multi-modal image fusion aims to integrate complementary information from different modalities to enhance perceptual capabilities in applications such as rescue and security. However, real-world imaging often suffers from degradation issues, such as noise, blur, and haze in visible imaging, as well…

2025

Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline

AAAI 2025technical

Large Visual Language Models (LVLMs) have achieved remarkable success in vision tasks. However, the significant differences between industrial and natural scenes make applying LVLMs challenging. Existing LVLMs rely on user-provided prompts to segment objects. This often leads to suboptimal performan…

2024

ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language Model

NeurIPS 2024poster

Visual object tracking aims to locate a targeted object in a video sequence based on an initial bounding box. Recently, Vision-Language~(VL) trackers have proposed to utilize additional natural language descriptions to enhance versatility in various applications. However, VL trackers are still infer…

Cited by 6SourcePDFScholar
2024

Dynamic Brightness Adaptation for Robust Multi-modal Image Fusion

IJCAI 2024poster

Infrared and visible image fusion aim to integrate modality strengths for visually enhanced, informative images. Visible imaging in real-world scenarios is susceptible to dynamic environmental brightness fluctuations, leading to texture degradation. Existing fusion methods lack robustness against su…

2024

One-shot Active Learning Based on Lewis Weight Sampling for Multiple Deep Models

ICLR 2024poster

Active learning (AL) for multiple target models aims to reduce labeled data querying while effectively training multiple models concurrently. Existing AL algorithms often rely on iterative model training, which can be computationally expensive, particularly for deep models. In this paper, we propose…

Cited by 4SourcePDFScholar
2023

Multi-Modal Gated Mixture of Local-to-Global Experts for Dynamic Image Fusion

ICCV 2023poster

Infrared and visible image fusion aims to integrate comprehensive information from multiple sources to achieve superior performances on various practical tasks, such as detection, over that of a single modality. However, most existing methods directly combined the texture details and object contrast…

Cited by 47PDFcodeScholar
2022

Online Active Regression

ICML 2022oral

Active regression considers a linear regression problem where the learner receives a large number of data points but can only observe a small number of labels. Since online algorithms can deal with incremental training data and take advantage of low computational cost, we consider an online extensio…

Cited by 10SourcePDFScholar
2021

Single Pass Entrywise-Transformed Low Rank Approximation

ICML 2021spotlight

In applications such as natural language processing or computer vision, one is given a large $n \times n$ matrix $A = (a_{i,j})$ and would like to compute a matrix decomposition, e.g., a low rank approximation, of a function $f(A) = (f(a_{i,j}))$ applied entrywise to $A$. A very important special ca…

Cited by 4SourcePDFScholar