← Search

HAO CHEN

316 accepted papers

2026

A Multi-Level Similarity Approach for Single-View Object Grasping: Matching, Planning, and Fine-Tuning

ICRA 2026poster

Grasping unknown objects from a single view has remained a challenging topic in robotics due to the uncertainty of partial observation. Recent advances in large-scale models have led to benchmark solutions such as GraspNet-1Billion. However, such learning-based approaches still face a critical limit…

2026

A Reinforcement Learning Based FEM Solver for Accelerating Contact-Influenced Simulation of Continuum Robots

RA-L 2026

Continuum robots exhibit exceptional flexibility and multi-degree-of-freedom maneuverability, offering significant advantages for navigating confined luminal spaces. However, rapid simulation of their contact-influenced behavior remains challenging. This paper presents RLFEM, an innovative finite el

Cited by 0SourceScholar
2026

A Survey of Artificial Intelligence in Endoscopic Surgery Workflow: From Perception to Surgical Support

IJCAI 2026

Endoscopic surgery demands continuous real-time visual decision-making under severe constraints, including a limited field of view, motion blur, and dynamically deforming anatomy. These factors impose substantial cognitive load on surgeons and motivate the integration of artificial intelligence (AI)

Cited by 0Scholar
2026

ACTIVE-o3 : Empowering MLLMs with Active Perception via Pure Reinforcement Learning

ICML 2026poster

Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of efficient perception and decision-making in humans and advanced embodied agents. With the rise of Multimodal Large Language M…

Cited by 0SourceScholar
2026

AIR-DR: Adaptive Image Retargeting with Instance Relocation and Dual-guidance Repainting

AAAI 2026technical

Image retargeting aims to adjust the aspect ratio of images to accommodate various display devices. While existing methods consider both foreground semantics and background inpainting, their Seam-carving-based framework is inherently destructive, often compromising the structural integrity of foregr

Cited by 0SourcePDFScholar
2026

Active Multi-source Domain Adaptation for Multimodal Fake News Detection

AAAI 2026technical

Multimodal fake news detection plays a crucial role in combating online misinformation. The inherent domain diversity of news in the real world has driven the development of cross-domain detection methods. However, these detection methods either suffer from significant performance degradation due to

Cited by 0SourcePDFScholar
2026

AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation

CVPR 2026

Recent advances in subject-driven video generation with large diffusion models have enabled personalized content synthesis conditioned on user-provided subjects. However, existing methods lack fine-grained temporal control over subject appearance and disappearance, which are essential for applicatio

Cited by 0SourceScholar
2026

Awakening Visual Reasoning: Mitigating Post-Training Failure in Vision-Text Compression

ICML 2026poster

Vision-Text Compression (VTC) offers a scalable path for long-context multimodal modeling by rendering textual data into dense visual tokens. While recent Vision-Language Models (VLMs) demonstrate high decoding fidelity (OCR) on such inputs, they exhibit a severe reasoning gap: models that reason ro…

Cited by 0SourceScholar
2026

Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-Resolution

CVPR 2026

Recent diffusion-based one-step methods have shown remarkable progress in the field of image super-resolution, yet they remain constrained by three critical limitations: (1) inferior fidelity performance caused by the information loss from compression encoding of low-quality (LQ) inputs; (2) insuffi

Cited by 0SourcecodeScholar
2026

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models

ICML 2026poster

Large Language Models(LLMs) have revolutionized text generation and multimodal perception, but their capabilities in 3D content generation remain underexplored. Existing methods compromise by producing either low-resolution meshes or coarse structural proxies, failing to capture fine-grained geometr…

Cited by 0SourceScholar
2026

ConSurv: Multimodal Continual Learning for Survival Analysis

AAAI 2026technical

Survival prediction of cancers is crucial for clinical practice, as it informs mortality risks and influences treatment plans. However, a static model trained on a single dataset fails to adapt to the dynamically evolving clinical environment and continuous data streams, limiting its practical utili

Cited by 0SourcePDFScholar
2026

DSSG: Dual-Stream Semantic Guidance for Source-Fully-Free Adaptation of Vision-Language Models

IJCAI 2026

Source-Fully-Free Domain Adaptation (SFF-DA) has emerged as a strategic paradigm to adapt Vision-Language Models (VLMs) without any access to source data or task-specific source models. However, we identify a critical "Dual Semantic Drift" that hinders this process: static drift caused by the stagna

Cited by 0Scholar
2026

Dynamic Stream Network for Combinatorial Explosion Problem in Deformable Medical Image Registration

CVPR 2026

Combinatorial explosion problem caused by dual inputs presents a critical challenge in Deformable Medical Image Registration (DMIR). Since DMIR processes two images simultaneously as input, the combination relationships between features grow exponentially, ultimately the model considers more irrelev

Cited by 0SourcecodeScholar
2026

EMFormer: Efficient Multi-Scale Transformer for Accumulative Context Weather Forecasting

ICML 2026poster

Long-term weather forecasting is critical for socioeconomic planning and disaster preparedness. While recent approaches employ finetuning to extend prediction horizons, they remain constrained by the issues of catastrophic forgetting, error accumulation, and high training overhead. To address these …

Cited by 0SourceScholar
2026

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching

CVPR 2026

Wide-baseline matching (WBM) requires integrating geometric understanding, viewpoint changes, fine-grained perception, and occlusion reasoning, making it a challenging testbed for spatial reasoning in multimodal large language models (MLLMs) deployed in physical environments. However, current MLLMs

Cited by 0SourceScholar
2026

Event-Based Motion Deblurring Using Task-Oriented 3D Gaussian Event Representations

CVPR 2026

Event-based motion deblurring has attracted increasing attention, as the high temporal resolution of event cameras provides motion cues unavailable to conventional RGB sensors, thereby enabling more effective deblurring. In real-world scenes, motion blur is often complex and nonlinear, with differen

Cited by 0SourceScholar
2026

Exploiting Low-Dimensional Manifold of Features for Few-shot Whole Slide Image Classification

ICLR 2026poster

Few-shot Whole Slide Image (WSI) classification is severely hampered by overfitting. We argue that this is not merely a data-scarcity issue but a fundamentally geometric problem. Grounded in the manifold hypothesis, our analysis shows that features from pathology foundation models exhibit a low-dime…

Cited by 0SourcecodeScholar
2026

Exploring Spatial Intelligence from a Generative Perspective

CVPR 2026

Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective. We ask whether modern generative or unified multimodal models also possess generative spatial intelligence (GSI)--the ability to respect and manipu

Cited by 0SourcecodeScholar
2026

From Manuals to Actions: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation

CVPR 2026

Vision-Language-Action (VLA) models have recently emerged, demonstrating strong generalization in robotic scene understanding and manipulation. However, when confronted with long-horizon tasks that require defined goal states, such as LEGO assembly or object rearrangement, existing VLA models still

Cited by 0SourceScholar
2026

From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment

ICML 2026poster

Adapting Large Language Models (LLMs) to specialized domains typically incurs high data and computational overhead. While prior efficiency efforts have largely treated data selection and parameter-efficient fine-tuning as isolated processes, our empirical analysis suggests they may be intrinsically …

Cited by 0SourceScholar
2026

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert

ICML 2026poster

Vision-language models demonstrate strong reasoning and planning abilities, yet grounding these predictions into precise robot actions remains a central challenge. Existing Vision-Language-Action methods typically entangle reasoning and action generation, leading to limited generalization and costly…

Cited by 0SourceScholar
2026

Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer Analysis

CVPR 2026

Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their application to medical diagnosis remains limited by the lack of comprehensive and structured datasets that capture real clinical workflows. To advance the devel

Cited by 0SourceScholar
2026

HPS: Hyperspherical Parameter Sharing for Efficient Multi-Agent Reinforcement Learning

ICML 2026poster

Parameter Sharing (PS) is widely used to improve efficiency in Multi-Agent Reinforcement Learning (MARL), but it can limit behavioral diversity and degrade performance. This limitation stems from gradient conflicts among agents on shared weights, which hinders effective policy learning. To fully cha…

Cited by 0SourceScholar
2026

ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning

ICLR 2026poster

The rapid advancement of text-to-image (T2I) models has increased the need for reliable human preference modeling, a demand further amplified by recent progress in reinforcement learning for preference alignment. However, existing approaches typically quantify the quality of a generated image using…

Cited by 0SourceScholar
2026

Know More, Know Clearer: A Meta-Cognitive Framework for Knowledge Augmentation in Large Language Models

ICML 2026spotlight

Knowledge augmentation has significantly enhanced the performance of Large Language Models (LLMs) in knowledge-intensive tasks. However, existing methods typically operate on the simplistic premise that model performance equates with internal knowledge, overlooking the knowledge-confidence gaps that…

Cited by 0SourceScholar
2026

Knowledge-Enhanced Explainable Prompting for Vision-Language Models

AAAI 2026technical

Large-scale vision-language models (VLMs) embedded with expansive representations and visual concepts have showcased significant potential in image and text understanding. Efficiently adapting VLMs such as CLIP to downstream tasks like few-shot image classification has garnered growing attention, wi

Cited by 0SourcePDFScholar
2026

KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning

ICLR 2026poster

Knowledge editing and machine unlearning are two popular approaches for large language models (LLMs) to stay up-to-date. However, the knowledge updating mechanism of LLMs remains largely unexplored due to insufficient, isolated, and small-scale evaluation. For instance, are LLMs similar to humans in…

Cited by 0SourcecodeScholar
2026

LLM Collaborative Filtering: User-Item Graph as New Language

AAAI 2026technical

In collaborative filtering, learning effective embeddings for users and items from interaction data remains a central challenge. While recent efforts leverage large language models (LLMs) to enhance collaborative filtering, two critical limitations persist: (1) Efficiency: LLM-based inference is sig

Cited by 0SourcePDFScholar
2026

LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

ICML 2026spotlight

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future observations prior to execution. However, explicit reasoning typically incurs non-negligible inference latency, which const…

Cited by 0SourceScholar
2026

Learning Patient-Specific Disease Dynamics With Latent Flow Matching For Longitudinal Imaging Generation

ICLR 2026poster

Understanding disease progression is a central clinical challenge with direct implications for early diagnosis and personalized treatment. While recent generative approaches have attempted to model progression, key mismatches remain: disease dynamics are inherently continuous and monotonic, yet late…

Cited by 0SourceScholar
2026

Linear Causal Representation Learning by Topological Ordering, Pruning, and Disentanglement

ICML 2026spotlight

Causal representation learning (CRL) has garnered increasing interests from the causal inference and artificial intelligence community, due to its capability of disentangling potentially complex data-generating mechanism into causally interpretable latent features, by leveraging the heterogeneity of…

Cited by 0SourceScholar
2026

LinearRAG: Linear Graph Retrieval Augmented Generation on Large-scale Corpora

ICLR 2026poster

Retrieval-Augmented Generation (RAG) is widely used to mitigate hallucinations of Large Language Models (LLMs) by leveraging external knowledge. While effective for simple queries, traditional RAG systems struggle with large-scale, unstructured corpora where information is fragmented. Recent advance…

Cited by 0SourcecodeScholar
2026

MLA: A Multisensory Language–Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation

ICRA 2026poster

Vision-language-action models (VLAs) have shown generalization capabilities in robotic manipulation tasks by inheriting from vision-language models (VLMs) and learning action generation. Most VLA models focus on interpreting vision and language to generate actions, whereas robots must perceive and i…

2026

NeRV-Diffusion: Diffuse Implicit Neural Representation for Video Synthesis

ICLR 2026poster

We present NeRV-Diffusion, an implicit latent video diffusion model that synthesizes videos via generating neural network weights. The generated weights can be rearranged as the parameters of a convolutional neural network, which forms an implicit neural representation (INR), and decodes into videos…

Cited by 0SourceScholar
2026

ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks

AAAI 2026technical

Language-guided long-horizon mobile manipulation has long been a grand challenge in embodied semantic reasoning, generalizable manipulation, and adaptive locomotion. Three fundamental limitations hinder progress: First, although large language models have shown promise in enhancing spatial reasoning

Cited by 0SourcePDFScholar
2026

One-Shot Flow, Any-Time Frame: A Bidirectional Warping Framework for Event-Based Video Frame Interpolation

CVPR 2026

Video Frame Interpolation (VFI) is a crucial task in video processing. Flow-based methods, despite their success, are constrained by a fundamental dilemma: forward warping is efficient but prone to artifacts, while backward warping yields higher quality at a significant computational cost, especiall

Cited by 0SourcecodeScholar
2026

PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation

CVPR 2026

Hand-object interaction (HOI) reconstruction and synthesis are becoming central to embodied AI and AR/VR. Yet, despite rapid progress, existing HOI generation research remains fragmented across three disjoint tracks: (1) pose-only synthesis that predicts MANO trajectories without producing pixels; (

Cited by 0SourcecodeScholar
2026

Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality

CVPR 2026

Video face swapping is crucial in film and entertainment production, where achieving high fidelity and temporal consistency over long and complex video sequences remains a significant challenge. Inspired by recent advances in reference-guided image editing, we explore whether rich visual attributes

Cited by 0SourcecodeScholar
2026

RegionReasoner: Region-Grounded Multi-Round Visual Reasoning

ICLR 2026poster

Large vision-language models have achieved remarkable progress in visual reasoning, yet most existing systems rely on single-step or text-only reasoning, limiting their ability to iteratively refine understanding across multiple visual contexts. To address this limitation, we introduce a new multi-r…

Cited by 0SourcecodeScholar
2026

Reinforced Rate Control for Neural Video Compression via Inter-Frame Rate–Distortion Awareness

AAAI 2026technical

Neural video compression (NVC) has demonstrated superior compression efficiency, yet effective rate control remains a significant challenge due to complex temporal dependencies. Existing rate control schemes typically leverage frame content to capture distortion interactions, overlooking inter-frame

Cited by 0SourcePDFScholar
2026

SPR: A Structured Prompt Refinement Network for Modality Missing

ICML 2026poster

Prompt learning has recently emerged as a novel, parameter-efficient paradigm to tackle the missing modalities challenge. However, existing prompting methods often overlook the internal structural information within prompt vectors, limiting their effectiveness in guiding frozen backbone models under…

Cited by 0SourceScholar
2026

STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting

CVPR 2026

To gain finer regional forecasts, many works have explored the regional integration from the global atmosphere, e.g., by solving boundary equations in physics-based methods or cropping regions from global forecasts in data-driven methods. However, the effectiveness of these methods is often constrai

Cited by 0SourcecodeScholar
2026

Scalable Vision-Language-Action Model Pretraining for Robotic Dexterous Manipulation with Real-Life Human Activity Videos

ICRA 2026poster

This paper presents an approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as dexterous robot end-effector, we show that "in-the-wild" egocentric human videos wit…

Cited by 0Scholar
2026

SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training

ICLR 2026poster

Recent advances in diffusion-based video restoration (VR) demonstrate significant improvement in visual quality, yet yield a prohibitive computational cost during inference. While several distillation-based approaches have exhibited the potential of one-step image restoration, extending existing app…

Cited by 0SourcecodeScholar
2026

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

CVPR 2026

Long video understanding (LVU) remains a core challenge in multimodal learning. Although recent vision-language models (VLMs) have made notable progress, existing benchmarks mainly focus on either fine-grained perception or coarse summarization, offering limited insight into temporal understanding o

Cited by 0SourceScholar
2026

StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation

CVPR 2026

A fundamental challenge in embodied intelligence is developing expressive and compact state representations for efficient world modeling and decision making. However, existing methods often fail to achieve this balance, yielding representations that are either overly redundant or lacking in task-cri

Cited by 0SourceScholar
2026

Swimming under Constraints: A Safe Reinforcement Learning Framework for Quadrupedal Bio-Inspired Propulsion

ICRA 2026poster

Bio-inspired aquatic propulsion offers high thrust and maneuverability but is prone to destabilizing forces such as lift fluctuations, which are further amplified by six-degree-of-freedom (6-DoF) fluid coupling. We formulate quadrupedal swimming as a constrained optimization problem that maximizes f…

2026

TINKER: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization

ICLR 2026poster

We introduce TINKER, a novel framework for high-fidelity 3D editing without any per-scene finetuning, where only a single edited image (one-shot) or a few edited images (few-shot) are required as input. Unlike prior techniques that demand extensive per-scene optimization to ensure multi-view consist…

Cited by 0SourcecodeScholar
2026

TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks

ICLR 2026poster

Multi-step reasoning tasks like mathematical problem solving are vulnerable to cascading failures where a single incorrect step leads to complete solution breakdown. Current LLM routing methods assign entire queries to one model, treating all reasoning steps as equal. We propose TRIM (Targeted Routi…

Cited by 0SourceScholar
2026

The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation

CVPR 2026

A reliable reward function is essential for reinforcement learning (RL) in image generation. Most current RL approaches depend on pre-trained preference models that output scalar rewards to approximate human preferences. However, these rewards often fail to capture human perception and are vulnerabl

Cited by 0SourcecodeScholar
2026

These Magic Moments: Differentiable Uncertainty Quantification of Radiance Field Models

ICRA 2026poster

Uncertainty quantification is crucial for autonomous systems, enabling safe and robust decision making in tasks ranging from active perception to robotic planning. This paper introduces a novel approach to quantify uncertainty for radiance fields by deriving pixel-wise moment expressions from the re…

2026

Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language Models

ICLR 2026poster

Diffusion large language models (dLLMs) generate text through iterative denoising, yet current decoding strategies discard rich intermediate predictions in favor of the final output. Our work here reveals a critical phenomenon, temporal oscillation, where correct answers often emerge in the middle p…

Cited by 0SourceScholar
2026

Transforming Weather Data from Pixel to Latent Space

ICML 2026oral

The increasing impact of climate change and extreme weather events has spurred growing interest in deep learning for weather research. However, existing studies often rely on weather data in pixel space, which presents several challenges such as smooth outputs in model outputs, limited applicability…

Cited by 0SourceScholar
2026

TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them

ICLR 2026poster

The adoption of Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks. We identify two fundamental types of inconsistencies: (1) \textit{Score-Comparison Inconsistency}, where lower-rated responses outperform high…

Cited by 0SourcecodeScholar
2026

Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs

ICLR 2026poster

Although unified MLLMs aim to unify generation and understanding, they are considered to exhibit an internal gap, with understanding outperforming generation. Through large‑scale evaluation across multiple MLLMs and tasks, we confirm the widespread non‑unification of MLLMs, and demonstrate that it i…

Cited by 0SourceScholar
2026

UniGame: Turning a Unified Multimodal Model Into Its Own Adversary

CVPR 2026

Unified Multimodal Models (UMMs) have shown impressive performance in both understanding and generation with a single architecture. However, UMMs still exhibit a fundamental inconsistency: understanding favors compact embeddings, whereas generation favors reconstruction-rich representations. This st

Cited by 0SourcecodeScholar
2026

Unifying Diffusion and Autoregression for Generalizable Vision-Language-Action Model

ICLR 2026poster

A central objective of manipulation policy design is to enable robots to comprehend human instructions and predict generalized actions in unstructured environments. Recent autoregressive vision-language-action (VLA) approaches discretize actions into bins to exploit the pretrained reasoning and gene…

Cited by 0SourceScholar
2026

Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation

CVPR 2026

Feed-forward visual geometry estimation has recently made rapid progress. However, an important gap remains: multi-frame models usually produce better cross-frame consistency, yet they often underperform strong per-frame methods on single-frame accuracy. This observation motivates our systematic inv

Cited by 0SourcecodeScholar
2026

You Don’t Need Pre-Built Graphs for RAG: Retrieval Augmented Generation with Adaptive Reasoning Structures

AAAI 2026technical

Large language models (LLMs) often suffer from hallucination, generating factually incorrect statements when handling questions beyond their knowledge and perception. Retrieval-augmented generation (RAG) addresses this by retrieving query-relevant contexts from knowledge bases to support LLM reasoni

Cited by 0SourcePDFScholar
2026

Zero-Shot Recognition of Test Tube Types by Automatically Collecting and Labeling RGB Data

ICRA 2026poster

This work presents a method for automatically detecting and recognizing test tube types in a rack. It leverages automatic segmentation, clustering, and labeling processes to eliminate the need for explicitly preparing training data. These processes are addressed by using combined global prediction a…

Cited by 0SourceScholar
2025

3DS-VLA: A 3D Spatial-Aware Vision Language Action Model for Robust Multi-Task Manipulation

CoRL 2025poster

Recently, 2D vision-language-action (VLA) models have made significant strides in multi-task manipulation. However, these models struggle to reason about 3D spatial relationships from 2D image inputs. Although an increasing number of 3D approaches explicitly integrate 3D information, they encounter…

Cited by 0SourceScholar
2025

A CVAE Combined With Diffusion Mechanism to Pedestrian Trajectory Prediction

RA-L 2025

Pedestrian trajectory prediction has always been a key issue in engineering fields such as autonomous driving. Diffusion model can be used to solve the problem of multimodal pedestrian trajectory generation, but the excessive denoising steps in the model result in slow inference speed for this type

Cited by 1SourceScholar
2025

A Denoising Pre-training Framework for Accelerating Novel Material Discovery

AAAI 2025technical

Crystal materials play an important role in the development of society. The discovery of new materials is critical to achieving sustainable development goals (SDGs), such as climate change mitigation, affordable and clean energy, and fostering innovation in industry and infrastructure. Recent advanc…

Cited by 0SourcePDFScholar
2025

A Survey of Pathology Foundation Model: Progress and Future Directions

IJCAI 2025

Computational pathology, which involves analyzing whole slide images for automated cancer diagnosis, relies on multiple instance learning, where performance depends heavily on the feature extractor and aggregator. Recent Pathology Foundation Models (PFMs), pretrained on large-scale histopathology da

2025

A Survey on Foundation Language Models for Single-cell Biology

ACL 2025long

The recent advancements in language models have significantly catalyzed progress in computational biology. A growing body of research strives to construct unified foundation models for single-cell biology, with language models serving as the cornerstone. In this paper, we systematically review the d…

Cited by 0SourcePDFScholar
2025

ALPS: Attention Localization and Pruning Strategy for Efficient Adaptation of Large Language Models

ACL 2025finding

Aligning general-purpose large language models (LLMs) to downstream tasks often incurs significant training adjustment costs. Prior research has explored various avenues to enhance alignment efficiency, primarily through minimal-data training or data-driven activations to identify key attention head…

2025

Accelerated Quasi-Static FEM for Real-Time Modeling of Continuum Robots with Multiple Contacts and Large Deformation

ICRA 2025

Continuum robots offer high flexibility and multiple degrees of freedom, making them ideal for navigating narrow lumens. However, accurately modeling their behavior under large deformations and frequent environmental contacts remains challenging. Current methods for solving the deformation of these

Cited by 2SourceScholar
2025

Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generation

EMNLP 2025

Recent decoding methods improve the factuality of large language models (LLMs) by refining how the next token is selected during generation. These methods typically operate at the token level, leveraging internal representations to suppress superficial patterns. Nevertheless, LLMs remain prone to ha

2025

Adaptive Grasping of Moving Objects in Dense Clutter via Global-to-Local Detection and Static-to-Dynamic Planning

ICRA 2025

Robotic grasping is facing a variety of real-world uncertainties caused by non-static object states, unknown object properties, and cluttered object arrangements. The difficulty of grasping increases with the presence of more uncertainties, where commonly used learning-based approaches struggle to p

Cited by 0SourceScholar
2025

BLEND: Behavior-guided Neural Population Dynamics Modeling via Privileged Knowledge Distillation

ICLR 2025poster

Modeling the nonlinear dynamics of neuronal populations represents a key pursuit in computational neuroscience. Recent research has increasingly focused on jointly modeling neural activity and behavior to unravel their interconnections. Despite significant efforts, these approaches often necessitate…

2025

Beyond Zero Initialization: Investigating the Impact of Non-Zero Initialization on LoRA Fine-Tuning Dynamics

ICML 2025poster

Low-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method. In standard LoRA layers, one of the matrices, $A$ or $B$, is initialized to zero, ensuring that fine-tuning starts from the pretrained model. However, there is no theoretical support for this practice. In this paper…

2025

Boltzmann-Aligned Inverse Folding Model as a Predictor of Mutational Effects on Protein-Protein Interactions

ICLR 2025spotlight

Predicting the change in binding free energy ($\Delta \Delta G$) is crucial for understanding and modulating protein-protein interactions, which are critical in drug design. Due to the scarcity of experimental $\Delta\Delta G$ data, existing methods focus on pre-training, while neglecting the impo…

2025

CAARMA: Class Augmentation with Adversarial Mixup Regularization

EMNLP 2025

Speaker verification is a typical zero-shot learning task, where inference of unseen classes is performed by comparing embeddings of test instances to known examples. The models performing inference must hence naturally generate embeddings that cluster same-class instances compactly, while maintaini

2025

CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency

EMNLP 2025

Instruction tuning is vital for aligning large language models (LLMs) with human intent, but current methods typically rely on costly human-annotated seed data or powerful external teacher models. While instruction back-translation techniques reduce this dependency, they remain fundamentally tethere

Cited by 0SourcePDFScholar
2025

Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks

CVPR 2025poster

Pre-trained vision-language models (VLMs) have showcased remarkable performance in image and natural language understanding, such as image captioning and response generation. As the practical applications of VLMs become increasingly widespread, their potential safety and robustness issues raise conc…

2025

ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation

EMNLP 2025

Retrieval-Augmented Generation (RAG) augments Large Language Models (LLMs) with external knowledge to improve factuality. However, existing RAG systems frequently underutilize the retrieved documents, failing to extract and integrate the key clues needed to support faithful and interpretable reasoni

2025

Conditional Visual Autoregressive Modeling for Pathological Image Restoration

ICCV 2025poster

Pathological image has been recognized as the gold standard for cancer diagnosis for more than a century. However, some internal regions of pathological images may inevitably exhibit various degradation issues, including low resolution, image blurring, and image noising, which will affect disease di…

2025

Context Matters: Query-aware Dynamic Long Sequence Modeling of Gigapixel Images

ICML 2025poster

Whole slide image (WSI) analysis presents significant computational challenges due to the massive number of patches in gigapixel images. While transformer architectures excel at modeling long-range correlations through self-attention, their quadratic computational complexity makes them impractical f…

2025

Controllable Traffic Simulation through LLM-Guided Hierarchical Reasoning and Refinement

IROS 2025

Evaluating autonomous driving systems in complex and diverse traffic scenarios through controllable simulation is essential to ensure their safety and reliability. However, existing traffic simulation methods face challenges in their controllability. To address this, we propose a novel diffusion-bas

Cited by 1SourceScholar
2025

DGCPL: Dual Graph Distillation for Concept Prerequisite Relation Learning

IJCAI 2025

Concept prerequisite relations determine the learning order of knowledge concepts in one domain, which has an important impact on teachers' course design and students' personalized learning. Current research usually predicts concept prerequisite relations from the perspective of knowledge, and rarel

2025

DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks

NeurIPS 2025spotlight

This paper's primary objective is to develop a robust generalist perception model capable of addressing multiple tasks under constraints of computational resources and limited training data. We leverage text-to-image diffusion models pre-trained on billions of images and successfully introduce our D…

Cited by 0SourcecodeScholar
2025

DIDS: Domain Impact-aware Data Sampling for Large Language Model Training

EMNLP 2025

Large language models (LLMs) are commonly trained on multi-domain datasets, where domain sampling strategies significantly impact model performance due to varying domain importance across downstream tasks. Existing approaches for optimizing domain-level sampling strategies struggle with maintaining

2025

DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation

AAAI 2025technical

Monocular camera calibration is a key precondition for numerous 3D vision applications. Despite considerable advancements, existing methods often hinge on specific assumptions and struggle to generalize across varied real-world scenarios, and the performance is limited by insufficient training data.…

Cited by 4SourcePDFScholar
2025

DiffSR: Learning Radar Reflectivity Synthesis via Diffusion Model from Satellite Observations

ICASSP 2025accepted

Weather radar data synthesis can fill in data for areas where ground observations are missing. Existing methods often employ reconstruction-based approaches with MSE loss to reconstruct radar data from satellite observation. However, such methods lead to over-smoothing, which hinders the generation…

Cited by 0SourceScholar
2025

Distilled Prompt Learning for Incomplete Multimodal Survival Prediction

CVPR 2025poster

The integration of multimodal data including pathology images and gene profiles is widely applied in precise survival prediction. Despite recent advances in multimodal survival models, collecting complete modalities for multimodal fusion still poses a significant challenge, hindering their applicati…

2025

Dual-Interrelated Diffusion Model for Few-Shot Anomaly Image Generation

CVPR 2025poster

The performance of anomaly inspection in industrial manufacturing is constrained by the scarcity of anomaly data. To overcome this challenge, researchers have started employing anomaly generation approaches to augment the anomaly dataset. However, existing anomaly generation methods suffer from limi…

2025

DynamicID: Zero-Shot Multi-ID Image Personalization with Flexible Facial Editability

ICCV 2025poster

Recent advances in text-to-image generation have driven interest in generating personalized human images that depict specific identities from reference images. Although existing methods achieve high-fidelity identity preservation, they are generally limited to single-ID scenarios and offer insuffici…

Cited by 0SourcePDFScholar
2025

EPA: Boosting Event-based Video Frame Interpolation with Perceptually Aligned Learning

NeurIPS 2025poster

Event cameras, with their capacity to provide high temporal resolution information between frames, are increasingly utilized for video frame interpolation (VFI) in challenging scenarios characterized by high-speed motion and significant occlusion. However, prevalent issues of blur and distortion wit…

Cited by 0SourceScholar
2025

ESEG: Event-Based Segmentation Boosted by Explicit Edge-Semantic Guidance

AAAI 2025technical

Event-based semantic segmentation (ESS) has attracted researchers' attention recently, as event cameras can solve problems such as under/over-exposure or motion blur that are difficult for RGB cameras to handle. However, event data are noisy and sparse, resulting in difficulties for the model to loc…

2025

EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights

CVPR 2025poster

Traffic Anomaly Understanding (TAU) is essential for improving public safety and transportation efficiency by enabling timely detection and response to incidents. Beyond existing methods, which rely largely on visual data, we propose to consider audio cues, a valuable source that offers strong hints…

2025

Enforcing Hard Linear Constraints in Deep Learning Models with Decision Rules

NeurIPS 2025poster

Deep learning models are increasingly deployed in safety-critical tasks where predictions must satisfy hard constraints, such as physical laws, fairness requirements, or safety limits. However, standard architectures lack built-in mechanisms to enforce such constraints, and existing approaches based…

Cited by 0SourceScholar
2025

Evaluating Program Semantics Reasoning with Type Inference in System $F$

NeurIPS 2025poster

Large Language Models (LLMs) are increasingly integrated into the software engineering ecosystem. Their test-time compute reasoning capabilities promise significant potential in understanding program logic and semantics beyond mere token recognition. However, current benchmarks evaluating reasoning…

Cited by 0SourceScholar
2025

Exploring the Choice Behavior of Large Language Models

ACL 2025finding

Large Language Models (LLMs) are increasingly deployed as human assistants across various domains where they help to make choices. However, the mechanisms behind LLMs’ choice behavior remain unclear, posing risks in safety-critical situations. Inspired by the intrinsic and extrinsic motivation frame…

Cited by 0SourcePDFScholar
2025

FOCUS: Knowledge-enhanced Adaptive Visual Compression for Few-shot Whole Slide Image Classification

CVPR 2025poster

Few-shot learning presents a critical solution for cancer diagnosis in computational pathology (CPath), addressing fundamental limitations in data availability, particularly the scarcity of expert annotations and patient privacy constraints. A key challenge in this paradigm stems from the inherent d…

2025

Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning

NeurIPS 2025poster

Generalized policy and execution efficiency constitute the two critical challenges in robotic manipulation. While recent foundation policies benefit from the common-sense reasoning capabilities of internet-scale pretrained vision-language models (VLMs), they often suffer from low execution frequency…

Cited by 0SourcecodeScholar
2025

Framer: Interactive Frame Interpolation

ICLR 2025poster

We propose Framer for interactive frame interpolation, which targets producing smoothly transitioning frames between two images as per user creativity. Concretely, besides taking the start and end frames as inputs, our approach supports customizing the transition process by tailoring the trajectory…

2025

From Pretraining to Pathology: How Noise Leads to Catastrophic Inheritance in Medical Models

NeurIPS 2025poster

Foundation models pretrained on web-scale data drive contemporary transfer learning in vision, language, and multimodal tasks. Recent work shows that mild label noise in these corpora may lift in-distribution accuracy yet sharply reduce out-of-distribution generalization, an effect known as catastro…

Cited by 0SourceScholar
2025

FuzzAug: Data Augmentation by Coverage-guided Fuzzing for Neural Test Generation

EMNLP 2025

Testing is essential to modern software engineering for building reliable software.Given the high costs of manually creating test cases,automated test case generation, particularly methods utilizing large language models,has become increasingly popular.These neural approaches generate semantically m

2025

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling

NeurIPS 2025poster

Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers,…

Cited by 0SourcecodeScholar
2025

GameGen-X: Interactive Open-world Game Video Generation

ICLR 2025poster

We introduce GameGen-$\mathbb{X}$, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos. This model facilitates high-quality, open-domain generation by approximating various game elements, such as innovative charact…

2025

Generalizable Object Keypoint Localization from Generative Priors

CVPR 2025poster

Generalizable object keypoint localization is a fundamental computer vision task in understanding the object structure. It is challenging for existing keypoint localization methods because their limited training data cannot provide generalizable shape and semantic cues, leading to inferior performan…

Cited by 0SourcePDFScholar
2025

ImageFolder: Autoregressive Image Generation with Folded Tokens

ICLR 2025poster

Image tokenizers are crucial for visual generative models, \eg, diffusion models (DMs) and autoregressive (AR) models, as they construct the latent representation for modeling. Increasing token length is a common approach to improve image reconstruction quality. However, tokenizers with longer token…

2025

Know Where You Are From: Event-Based Segmentation via Spatio-Temporal Propagation

AAAI 2025technical

Event cameras have gained attention in segmentation due to their higher temporal resolution and dynamic range compared to traditional cameras. However, they struggle with issues like lack of color perception and triggering only at motion edges, making it hard to distinguish objects with similar cont…

2025

LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior

ICLR 2025oral

We present LARP, a novel video tokenizer designed to overcome limitations in current video tokenization methods for autoregressive (AR) generative models. Unlike traditional patchwise tokenizers that directly encode local visual patches into discrete tokens, LARP introduces a holistic tokenization s…

2025

LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization

EMNLP 2025

Large language models (LLMs) have demonstrated impressive capabilities in reasoning with the emergence of reasoning models like OpenAI-o1 and DeepSeek-R1. Recent research focuses on integrating reasoning capabilities into the realm of retrieval-augmented generation (RAG) via outcome-supervised reinf

2025

Learn to Swim: Data-Driven LSTM Hydrodynamic Model for Quadruped Robot Gait Optimization

ICRA 2025

This paper presents a Long Short-Term Memory network-based Fluid Experiment Data-Driven model (FED-LSTM) for predicting unsteady, nonlinear hydrodynamic forces on the underwater quadruped robot we constructed. Trained on experimental data from leg force and body drag tests conducted in both a recirc

Cited by 2SourceScholar
2025

Learning Concept Prerequisite Relation via Global Knowledge Relation Optimization

AAAI 2025technical

Learning concept prerequisite relations helps better master and build a logically coherent knowledge structure. Many studies use graph neural networks to create heterogeneous knowledge networks that enhance concept representations. However, different types of relations in these networks can influenc…

2025

LongTableBench: Benchmarking Long-Context Table Reasoning across Real-World Formats and Domains

EMNLP 2025

We introduce LongTableBench , a benchmark for evaluating long-context reasoning over semi-structured tables across diverse formats, tasks, and domains. It comprises 5,950 QA instances spanning 7 table formats (e.g., Markdown, HTML, SQL), 18 domains, and input lengths up to 128K tokens, including mul

2025

MM-Tracker: Motion Mamba for UAV-platform Multiple Object Tracking

AAAI 2025technical

Multiple object tracking (MOT) from unmanned aerial vehicle (UAV) platforms requires efficient motion modeling. This is because UAV-MOT faces both local object motion and global camera motion. Motion blur also increases the difficulty of detecting large moving objects. Previous UAV motion modeling a…

2025

MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models

IJCAI 2025

Text-to-image diffusion models can generate high-quality images but lack fine-grained control of visual concepts, limiting their creativity. Thus, we introduce component-controllable personalization, a new task that enables users to customize and reconfigure individual components within concepts. Th

Cited by 0SourcePDFScholar
2025

Making RALM Robust to Irrelevant Contexts via Layer Knowledge Guided Attention

ACL 2025finding

Retrieval-augmented language models (RALMs) aim to incorporate external knowledge to address the issues of factual hallucination and knowledge obsolescence faced by large language models (LLMs). Inevitably, the retrieved passages based on similarity search may be irrelevant to the given question, an…

2025

Masked Autoencoders Are Effective Tokenizers for Diffusion Models

ICML 2025spotlight

Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that i…

Cited by 8SourcePDFScholar
2025

MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequences

ICLR 2025poster

Recent advancements in video generation have primarily leveraged diffusion models for short-duration content. However, these approaches often fall short in modeling complex narratives and maintaining character consistency over extended periods, which is essential for long-form video production like…

Cited by 24SourcePDFScholar
2025

OSV: One Step is Enough for High-Quality Image to Video Generation

CVPR 2025poster

Video diffusion models have shown great potential in generating high-quality videos, making them an increasingly popular focus. However, their inherent iterative nature leads to substantial computational and time costs. Although techniques such as consistency distillation and adversarial training ha…

Cited by 10SourcePDFScholar
2025

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

NeurIPS 2025poster

Long-horizon video-audio reasoning and fine-grained pixel understanding impose conflicting requirements on omnimodal models: dense temporal coverage demands many low-resolution frames, whereas precise grounding calls for high-resolution inputs. We tackle this trade-off with a two-system architecture…

Cited by 0SourcecodeScholar
2025

On Fairness of Unified Multimodal Large Language Model for Image Generation

NeurIPS 2025poster

Unified multimodal large language models (U-MLLMs) have demonstrated impressive performance in end-to-end visual understanding and generation tasks. However, compared to generation-only systems (e.g., Stable Diffusion), the unified architecture of U-MLLMs introduces new risks of propagating demograp…

Cited by 0SourceScholar
2025

PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model

ICML 2025poster

Multi-objective test-time alignment aims to adapt large language models (LLMs) to diverse multi-dimensional user preferences during inference while keeping LLMs frozen. Recently, GenARM (Xu et al., 2025) first independently trains Autoregressive Reward Models (ARMs) for each preference dimension wi…

2025

PEACE: Empowering Geologic Map Holistic Understanding with MLLMs

CVPR 2025poster

Geologic map, as a fundamental diagram in geology science, provides critical insights into the structure and composition of Earth's subsurface and surface. These maps are indispensable in various fields, including disaster assessment, resource exploration, and civil engineering. Despite their signif…

2025

POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction

ICCV 2025poster

Recent approaches to 3D reconstruction in dynamic scenes primarily rely on the integration of separate geometry estimation and matching modules, where the latter plays a critical role in distinguishing dynamic regions and mitigating the interference caused by moving objects. Furthermore, the matchin…

2025

ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation

NeurIPS 2025poster

Large language models (LLMs) integrated with retrieval-augmented generation (RAG) have improved factuality by grounding outputs in external evidence. However, they remain susceptible to unfaithful generation, where outputs contradict retrieved context despite its relevance and accuracy. Existing app…

Cited by 0SourcecodeScholar
2025

PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training

ICLR 2025spotlight

This paper aims to address the challenge of hallucinations in Multimodal Large Language Models (MLLMs) particularly for dense image captioning tasks. To tackle the challenge, we identify the current lack of a metric that finely measures the caption quality in concept level. We hereby introduce HalF…

Cited by 0SourcePDFScholar
2025

Point Cloud Upsampling Using Conditional Diffusion Module with Adaptive Noise Suppression

CVPR 2025poster

Point cloud upsampling can improve the quality of the initial point cloud, significantly enhancing the performance of downstream tasks such as classification and segmentation. Existing methods mostly focus on generating the geometric details of point clouds, neglecting noise suppression. To address…

2025

RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards

ICLR 2025poster

Retrieval-Augmented Generation (RAG) has proven its effectiveness in mitigating hallucinations in Large Language Models (LLMs) by retrieving knowledge from external resources. To adapt LLMs for the RAG systems, current approaches use instruction tuning to optimize LLMs, improving their ability to ut…

2025

Reinforced Lifelong Editing for Language Models

ICML 2025poster

Large language models (LLMs) acquire information from pre-training corpora, but their stored knowledge can become inaccurate or outdated over time. Model editing addresses this challenge by modifying model parameters without retraining, and prevalent approaches leverage hypernetworks to generate the…

2025

Rethinking the Bias of Foundation Model under Long-tailed Distribution

ICML 2025poster

Long-tailed learning has garnered increasing attention due to its practical significance. Among the various approaches, the fine-tuning paradigm has gained considerable interest with the advent of foundation models. However, most existing methods primarily focus on leveraging knowledge from these mo…

Cited by 0SourcePDFScholar
2025

Revisiting Convolution Architecture in the Realm of DNA Foundation Models

ICLR 2025poster

In recent years, A variety of methods based on Transformer and state space model (SSM) architectures have been proposed, advancing foundational DNA language models. However, there is a lack of comparison between these recent approaches and the classical architecture—convolutional networks (CNNs)—on…

Cited by 0SourcePDFScholar
2025

Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology

NeurIPS 2025poster

Pre-trained encoders for offline feature extraction followed by multiple instance learning (MIL) aggregators have become the dominant paradigm in computational pathology (CPath), benefiting cancer diagnosis and prognosis. However, performance limitations arise from the absence of encoder fine-tuning…

Cited by 0SourcecodeScholar
2025

Robotic In Situ Measurement of Multiple Intracellular Physical Parameters Based on Three-micropipettes System

IROS 2025

Physical parameters of the intracellular environment such as mass density, intracellular pressure and elasticity have significant effects on the physiological activities of the cell and intracellular operation results. However, the significantly different measurement principles of the above paramete

Cited by 0SourceScholar
2025

Role-aware Multi-agent Reinforcement Learning for Coordinated Emergency Traffic Control

NeurIPS 2025poster

Emergency traffic control presents an increasingly critical challenge, requiring seamless coordination among emergency vehicles, regular vehicles, and traffic lights to ensure efficient passage for all vehicles. Existing models primarily only focus on traffic light control, leaving emergency and reg…

Cited by 0SourceScholar
2025

SDP-CROWN: Efficient Bound Propagation for Neural Network Verification with Tightness of Semidefinite Programming

ICML 2025spotlight

Neural network verifiers based on linear bound propagation scale impressively to massive models but can be surprisingly loose when neuron coupling is crucial. Conversely, semidefinite programming (SDP) verifiers capture inter-neuron coupling naturally, but their cubic complexity restricts them to on…

Cited by 0SourcePDFScholar
2025

Satellite Observations Guided Diffusion Model for Accurate Meteorological States at Arbitrary Resolution

CVPR 2025highlight

Accurate acquisition of surface meteorological conditions at arbitrary locations holds significant importance for weather forecasting and climate simulation. Meteorological states derived from satellite observations are often provided in the form of low-resolution grid fields. If spatial interpolati…

2025

Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data

ICCV 2025poster

AI for tumor segmentation is limited by the lack of large, voxel-wise annotated datasets, which are hard to create and require medical experts. In our proprietary JHH dataset of 3,000 annotated pancreatic tumor scans, we found that AI performance stopped improving after 1,500 scans. With synthetic d…

2025

SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems

ACL 2025finding

The rapid advancement of Large Multi-modal Models (LMMs) has enabled their application in scientific problem-solving, yet their fine-grained capabilities remain under-explored. In this paper, we introduce SciVerse, a multi-modal scientific evaluation benchmark to thoroughly assess LMMs across 5,735…

2025

Seeing the Unseen: Composing Outliers for Compositional Zero-Shot Learning

IJCAI 2025

Compositional zero-shot learning (CZSL) is to recognize unseen attribute-object compositions by learning from seen compositions. The distribution shift between unseen compositions and seen compositions poses challenges to CZSL models, especially when test images are mixed with both seen and unseen c

Cited by 0SourcePDFScholar
2025

SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories

CVPR 2025poster

While MLLMs have demonstrated adequate image understanding capabilities, they still struggle with pixel-level comprehension, limiting their practical applications. Current evaluation tasks like VQA and visual grounding remain too coarse to assess fine-grained pixel comprehension accurately. Though s…

2025

Self-cross Feature based Spiking Neural Networks for Efficient Few-shot Learning

ICML 2025poster

Deep neural networks (DNNs) excel in computer vision tasks, especially, few-shot learning (FSL), which is increasingly important for generalizing from limited examples. However, DNNs are computationally expensive with scalability issues in real world. Spiking Neural Networks (SNNs), with their even…

Cited by 0SourcePDFScholar
2025

SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer

CVPR 2025poster

Efficient image tokenization with high compression ratios remains a critical challenge for training generative models.We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple codewords into each latent token, substantially increasing the re…

2025

Structure-Guided Large Language Models for Text-to-SQL Generation

ICML 2025poster

Recent advancements in large language models (LLMs) have shown promise in bridging the gap between natural language queries and database management systems, enabling users to interact with databases without the background of SQL. However, LLMs often struggle to fully exploit and comprehend the user…

Cited by 0SourcePDFScholar
2025

SurfaceSplat: Connecting Surface Reconstruction and Gaussian Splatting

ICCV 2025poster

Surface reconstruction and novel view rendering from sparse-view images are challenging. Signed Distance Function (SDF)-based methods struggle with fine details, while 3D Gaussian Splatting (3DGS)-based approaches lack global geometry coherence. We propose a novel hybrid method that combines both st…

2025

SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset

NeurIPS 2025poster

Code-switching (CS) is the alternating use of two or more languages within a conversation or utterance, often influenced by social context and speaker identity. This linguistic phenomenon poses challenges for Automatic Speech Recognition (ASR) systems, which are typically designed for a single langu…

Cited by 0SourcecodeScholar
2025

TC–RAG: Turing–Complete RAG’s Case study on Medical LLM Systems

ACL 2025long

In the pursuit of enhancing domain-specific Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) emerges as a promising solution to mitigate issues such as hallucinations, outdated knowledge, and limited expertise in highly specialized queries. However, existing approaches to RAG fall…

2025

TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings

AAAI 2025technical

Currently, inspired by the success of vision-language models (VLMs), an increasing number of researchers are focusing on improving VLMs and have achieved promising results. However, most existing methods concentrate on optimizing the connector and enhancing the language model component, while neglec…

Cited by 4SourcePDFScholar
2025

Text-Attributed Graph Learning with Coupled Augmentations

COLING 2025main

Modeling text-attributed graphs is a well-known problem due to the difficulty of capturing both the text attribute and the graph structure effectively. Existing models often focus on either the text attribute or the graph structure, potentially neglecting the other aspect. This is primarily because…

Cited by 0SourcePDFScholar
2025

Time Series Supplier Allocation via Deep Black-Litterman Model

AAAI 2025technical

As a typical problem of Spatiotemporal Resource Management, Time Series Supplier Allocation (TSSA) poses a complex NP-hard challenge, aimed at refining future order dispatching strategies to satisfy the trade-off between demands and maximum supply. The Black-Litterman (BL) model, which comes from fi…

2025

Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs

NeurIPS 2025poster

Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. However, engineering design poses a fundamentally different challenge for large language models (LLMs) compared with tradi…

Cited by 0SourceScholar
2025

Towards Loss-Resilient Image Coding for Unstable Satellite Networks

AAAI 2025technical

Geostationary Earth Orbit (GEO) satellite communication demonstrates significant advantages in emergency short burst data services. However, unstable satellite networks, particularly those with frequent packet loss, present a severe challenge to accurate image transmission. To address it, we propose…

2025

Unified Open-World Segmentation with Multi-Modal Prompts

ICCV 2025poster

In this work, we present COSINE, a unified open-world segmentation model that Consolidates Open-vocabulary Segmentation and IN-context sEgmentation with multi-modal prompts (e.g., text and image). COSINE exploits foundation models to extract representations for an input image and corresponding multi…

2025

Unleashing Hour-Scale Video Training for Long Video-Language Understanding

NeurIPS 2025spotlight

Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hou…

Cited by 0SourceScholar
2025

UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI

ICCV 2025poster

We introduce UnrealZoo, a collection of over 100 photo-realistic 3D virtual worlds built on Unreal Engine, designed to reflect the complexity and variability of open-world environments. We also provide a rich variety of playable entities, including humans, animals, robots, and vehicles for embodied…

2025

VA-MoE: Variables-Adaptive Mixture of Experts for Incremental Weather Forecasting

ICCV 2025poster

This paper presents Variables-Adaptive Mixture of Experts (VA-MoE), a novel framework for incremental weather forecasting that dynamically adapts to evolving spatiotemporal patterns in real-time data. Traditional weather prediction models often struggle with exorbitant computational expenditure and…

2025

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

IROS 2025

We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tracking systems in recovering from tracking failure. Our approach combines the off-the-shelf active tracking methods with

Cited by 3SourceScholar
2025

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

NeurIPS 2025poster

This paper presents a multimodal framework that attempts to unify visual understanding and generation within a shared discrete semantic representation. At its core is the Text-Aligned Tokenizer (TA-Tok), which converts images into discrete tokens using a text-aligned codebook projected from a large…

Cited by 0SourceScholar
2025

VisualEDU: A Benchmark for Assessing Coding and Visual Comprehension through Educational Problem-Solving Video Generation

EMNLP 2025

Generating logically coherent video from text (T2V) for reasoning-intensive tasks like mathematical problem-solving presents a significant challenge for Vision-Language Models (VLMs). Therefore, we introduce VisualEDU, a benchmark based on Manim package to rigorously evaluate VLM capabilities in pro

2025

WeatherGFM: Learning a Weather Generalist Foundation Model via In-context Learning

ICLR 2025poster

The Earth's weather system involves intricate weather data modalities and diverse weather understanding tasks, which hold significant value to human life. Existing data-driven models focus on single weather understanding tasks (e.g., weather forecasting). While these models have achieved promising…

2025

What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?

ICLR 2025poster

Extensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I diffusion models for a few dense perception tasks. However, sever…

2025

Zero-Shot Recognition of Test Tube Types by Automatically Collecting and Labeling RGB Data

RA-L 2025

This work presents a method for automatically detecting and recognizing test tube types in a rack. It leverages automatic segmentation, clustering, and labeling processes to eliminate the need for explicitly preparing training data. These processes are addressed by using combined global prediction a

Cited by 0SourceScholar
2025

kNN-CL: Enhancing Continual Learning with Nearest Neighbor Retrieval

ICASSP 2025accepted

Continual learning aims to learn new tasks sequentially without forgetting previously acquired knowledge. However, catastrophic forgetting remains a significant challenge. In this paper, we introduce kNN-CL, a simple yet effective approach that harnesses k-nearest neighbors (kNN) to mitigate forgett…

Cited by 0SourceScholar
2024

360+x: A Panoptic Multi-modal Scene Understanding Dataset

CVPR 2024poster

Human perception of the world is shaped by a multitude of viewpoints and modalities. While many existing datasets focus on scene understanding from a certain perspective (e.g. egocentric or third-person views) our dataset offers a panoptic perspective (i.e. multiple viewpoints with multiple data mod…

2024

A Dynamic GCN with Cross-Representation Distillation for Event-Based Learning

AAAI 2024technical

Recent advances in event-based research prioritize sparsity and temporal precision. Approaches learning sparse point-based representations through graph CNNs (GCN) become more popular. Yet, these graph techniques hold lower performance than their frame-based counterpart due to two issues: (i) Biased…

Cited by 8SourcePDFScholar
2024

A General Framework for Learning from Weak Supervision

ICML 2024poster

Weakly supervised learning generally faces challenges in applicability to various scenarios with diverse weak supervision and in scalability due to the complexity of existing algorithms, thereby hindering the practical deployment. This paper introduces a general framework for learning from weak supe…

2024

A Motion-aware Spatio-temporal Graph for Video Salient Object Ranking

NeurIPS 2024poster

Video salient object ranking aims to simulate the human attention mechanism by dynamically prioritizing the visual attraction of objects in a scene over time. Despite its numerous practical applications, this area remains underexplored. In this work, we propose a graph model for video salient object…

2024

A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image Classifications

ICASSP 2024accepted

Although current data augmentation methods are successful to alleviate the data insufficiency, conventional augmentation are primarily intra-domain while advanced generative adversarial networks (GANs) generate images remaining uncertain, particularly in small-scale datasets. In this paper, we propo…

Cited by 0SourceScholar
2024

A Simple Image Segmentation Framework via In-Context Examples

NeurIPS 2024poster

Recently, there have been explorations of generalist segmentation models that can effectively tackle a variety of image segmentation tasks within a unified in-context learning framework. However, these methods still struggle with task ambiguity in in-context segmentation, as not all in-context examp…

2024

AgentReview: Exploring Peer Review Dynamics with LLM Agents

EMNLP 2024main

Peer review is fundamental to the integrity and advancement of scientific publication. Traditional methods of peer review analyses often rely on exploration and statistics of existing peer review data, which do not adequately address the multivariate nature of the process, account for the latent var…

2024

BEVLOC: End-to-End 6-DoF Localization Via Cross-Modality Correlation Under Bird's Eye View

ICASSP 2024accepted

Accurate ego-centric localization assumes a paramount significance in the domain of autonomous driving. However, traditional methods for camera-LiDAR map localization rely on perspective projection to create a unified representation, which often falls short due to challenges such as occlusion and th…

Cited by 0SourceScholar
2024

Better Zero-Shot Reasoning with Role-Play Prompting

NAACL 2024long

Modern large language models (LLMs) exhibit a remarkable capacity for role-playing, enabling them to embody not only human characters but also non-human entities. This versatility allows them to simulate complex human-like interactions and behaviors within various contexts, as well as to emulate spe…

2024

BronchoCopilot: Towards Autonomous Robotic Bronchoscopy via Multimodal Reinforcement Learning

IROS 2024poster

Bronchoscopy plays a significant role in the early diagnosis and treatment of lung diseases. This process demands physicians to maneuver the flexible endoscope for reaching distal lesions, particularly requiring substantial expertise when examining the airways of the upper lung lobe. With the develo…

Cited by 1SourceScholar
2024

Code Representation Pre-training with Complements from Program Executions

EMNLP 2024industry

Language models for natural language processing have been grafted onto programming language modeling for advancing code intelligence. Although it can be represented in the text format, code is syntactically more rigorous, as it is designed to be properly compiled or interpreted to perform a set of b…

Cited by 6SourcePDFScholar
2024

CompeteAI: Understanding the Competition Dynamics of Large Language Model-based Agents

ICML 2024oral

Large language models (LLMs) have been widely used as agents to complete different tasks, such as personal assistance or event planning. Although most of the work has focused on cooperation and collaboration between agents, little work explores *competition*, another important mechanism that promote…

2024

Completing Visual Objects via Bridging Generation and Segmentation

ICML 2024poster

This paper presents a novel approach to object completion, with the primary goal of reconstructing a complete object from its partially visible components. Our method, named MaskComp, delineates the completion process through iterative stages of generation and segmentation. In each iteration, the ob…

Cited by 6SourcePDFScholar
2024

Cost-efficient Knowledge-based Question Answering with Large Language Models

NeurIPS 2024poster

Knowledge-based question answering (KBQA) is widely used in many scenarios that necessitate domain knowledge. Large language models (LLMs) bring opportunities to KBQA, while their costs are significantly higher and absence of domain-specific knowledge during pre-training. We are motivated to combine…

Cited by 8SourcePDFScholar
2024

Customising General Large Language Models for Specialised Emotion Recognition Tasks

ICASSP 2024accepted

The advent of large language models (LLMs) has gained tremendous attention over the past year. Previous studies have shown the astonishing performance of LLMs not only in other tasks but also in emotion recognition in terms of accuracy, universality, explanation, robustness, few/zero-shot learning,…

Cited by 0SourceScholar
2024

De novo Protein Design Using Geometric Vector Field Networks

ICLR 2024spotlight

Advances like protein diffusion have marked revolutionary progress in $\textit{de novo}$ protein design, a central topic in life science. These methods typically depend on protein structure encoders to model residue backbone frames, where atoms do not exist. Most prior encoders rely on atom-wise fea…

2024

DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data

CVPR 2024poster

Instance segmentation is data-hungry and as model capacity increases data scale becomes crucial for improving the accuracy. Most instance segmentation datasets today require costly manual annotation limiting their data scale. Models trained on such data are prone to overfitting on the training set e…

2024

Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL

ECCV 2024poster

"Embodied visual tracking is to follow a target object in dynamic 3D environments using an agent’s egocentric vision. This is a vital and challenging skill for embodied agents. However, existing methods suffer from inefficient training and poor generalization. In this paper, we propose a novel frame…

Cited by 5SourcePDFScholar
2024

Enhancing Explainable Rating Prediction through Annotated Macro Concepts

ACL 2024long

Generating recommendation reasons for recommendation results is a long-standing problem because it is challenging to explain the underlying reasons for recommending an item based on user and item IDs. Existing models usually learn semantic embeddings for each user and item, and generate the reasons…

Cited by 6SourcePDFScholar
2024

FNP: Fourier Neural Processes for Arbitrary-Resolution Data Assimilation

NeurIPS 2024poster

Data assimilation is a vital component in modern global medium-range weather forecasting systems to obtain the best estimation of the atmospheric state by combining the short-term forecast and observations. Recently, AI-based data assimilation approaches have attracted increasing attention for their…

2024

Fine Tuning Out-of-Vocabulary Item Recommendation with User Sequence Imagination

NeurIPS 2024spotlight

Recommending out-of-vocabulary (OOV) items is a challenging problem since the in-vocabulary (IV) items have well-trained behavioral embeddings but the OOV items only have content features. Current OOV recommendation models often generate 'makeshift' embeddings for OOV items from content features and…

Cited by 3SourcePDFScholar
2024

Floating Anchor Diffusion Model for Multi-motif Scaffolding

ICML 2024poster

Motif scaffolding seeks to design scaffold structures for constructing proteins with functions derived from the desired motif, which is crucial for the design of vaccines and enzymes. Previous works approach the problem by inpainting or conditional generation. Both of them can only scaffold motifs w…

2024

FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition

CVPR 2024poster

Benefiting from large-scale pre-trained text-to-image (T2I) generative models impressive progress has been achieved in customized image generation which aims to generate user-specified concepts. Existing approaches have extensively focused on single-concept customization and still encounter challeng…

2024

Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid Modeling

NeurIPS 2024poster

Data-driven artificial intelligence (AI) models have made significant advancements in weather forecasting, particularly in medium-range and nowcasting. However, most data-driven weather forecasting models are black-box systems that focus on learning data mapping rather than fine-grained physical evo…

2024

Generative Active Learning for Long-tailed Instance Segmentation

ICML 2024poster

Recently, large-scale language-image generative models have gained widespread attention and many works have utilized generated data from these models to further enhance the performance of perception tasks. However, not all generated data can positively impact downstream models, and these methods do…

2024

High Accuracy Device Localization in Indoor Mmwave Networks Exploiting Channel Sparsity and Virtual Anchor Mapping

ICASSP 2024accepted

In this paper, we propose a novel indoor localization algorithm that exploits the angle and delay information of the sparse channel paths at mmWave. We consider that the user and the access point (AP) are not perfectly synchronized, which results in an unknown clock offset for the estimated delays.…

Cited by 0SourceScholar
2024

Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label Configurations

NeurIPS 2024poster

Learning with reduced labeling standards, such as noisy label, partial label, and supplementary unlabeled data, which we generically refer to as imprecise label, is a commonplace challenge in machine learning tasks. Previous methods tend to propose specific designs for every emerging imprecise label…

2024

Improved Generation of Adversarial Examples Against Safety-aligned LLMs

NeurIPS 2024poster

Adversarial prompts (or say, adversarial examples) generated using gradient-based methods exhibit outstanding performance in performing automatic jailbreak attacks against safety-aligned LLMs. Nevertheless, due to the discrete nature of texts, the input gradient of LLMs struggles to precisely reflec…

2024

KnowGPT: Knowledge Graph based Prompting for Large Language Models

NeurIPS 2024poster

Large Language Models (LLMs) have demonstrated remarkable capabilities in many real-world applications. Nonetheless, LLMs are often criticized for their tendency to produce hallucinations, wherein the models fabricate incorrect statements on tasks beyond their knowledge and perception. To alleviate…

Cited by 12SourcePDFScholar
2024

Knowledge-to-SQL: Enhancing SQL Generation with Data Expert LLM

ACL 2024findings

Generating accurate SQL queries for user questions (text-to-SQL) has been a long-standing challenge since it requires a deep understanding of both the user’s question and the corresponding database schema in order to retrieve the desired content accurately. Existing methods rely on the comprehensive…

2024

LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning

ACL 2024findings

Large Language Models (LLMs), such as LLaMA and T5, have shown exceptional performance across various tasks through fine-tuning. Although low-rank adaption (LoRA) has emerged to cheaply fine-tune these LLMs on downstream tasks, their deployment is still hindered by the vast model scale and computati…