← Search

Hao Zhang

367 accepted papers

2026

A Recursive Decomposition Framework for Causal Structure Learning in the Presence of Latent Variables

ICML 2026oral

Constraint-based causal discovery is widely used for learning causal structures, but heavy reliance on conditional independence (CI) testing makes it computationally expensive in high-dimensional settings. To mitigate this limitation, many divide-and-conquer frameworks have been proposed, but most a…

Cited by 0SourceScholar
2026

Adapting Execution-Time Objectives for Multi-Robot Policies via Collaborative Flow Policy Guidance

RSS 2026poster

Multi-robot teams are increasingly gaining attention due to their ability to scale up in terms of task workloads and complexities. However, existing approaches struggle with three key limitations: the inability of unimodal policies to capture multi-modal joint strategies, the rigidity of fixed polic…

Cited by 0SourceScholar
2026

Attn-QAT: 4-Bit Attention With Quantization-Aware Training

ICML 2026poster

Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the main obstacle due to FP4's tiny dynamic range and attention's heavy-tailed activations. This paper presents the first systematic study of 4-bit quantization-awa…

Cited by 0SourceScholar
2026

Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning

AAAI 2026technical

Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning utilizing rule-based rewards. However, the explicit reasoning process has not yet yielded su

Cited by 0SourcePDFScholar
2026

Bayesian Decomposition and Semantic Completion for Few-shot Semantic Segmentation

CVPR 2026

Few-shot Semantic Segmentation (FSS) aims to segment objects of novel categories given only a handful of labeled examples. However, existing methods often rely on complex category-specific modeling, resulting in high computational cost and limited generalization under low-data regimes. To address th

Cited by 0SourceScholar
2026

Bidirectional Noise Injection: Enhancing Diffusion Models via Coordinated Input-Output Perturbation

AAAI 2026technical

Diffusion models have demonstrated remarkable success in image generation, yet a persistent challenge remains: the bias between model predictions and the target distribution. In this paper, we propose a Bidirectional Noise Injection framework for enhancing diffusion models, implemented via Coordinat

Cited by 0SourcePDFScholar
2026

BiomedCCPL: Causal Conditional Prompt Learning for Biomedical Vision-Language Models

CVPR 2026

Vision-language models (VLMs) have demonstrated strong potential for adapting to downstream biomedical tasks with limited training samples. However, their generalization to unseen classes within the same dataset remains limited, as the image-text alignment semantics often rely on spurious cues prese

Cited by 0SourcecodeScholar
2026

C^2FG: Control Classifier-Free Guidance via Score Discrepancy Analysis

CVPR 2026

Classifier-Free Guidance (CFG) is a cornerstone of modern conditional diffusion models, yet its reliance on the fixed or heuristic dynamic guidance weight is predominantly empirical and overlooks the inherent dynamics of the diffusion process. In this paper, we provide a rigorous theoretical analysi

Cited by 0SourceScholar
2026

Collaborative Planning with Concurrent Synchronization for Operationally Constrained UAV-UGV Teams

ICRA 2026poster

Collaborative planning under operational constraints is an essential capability for heterogeneous robot teams tackling complex large-scale real-world tasks. Unmanned Aerial Vehicles (UAVs) offer rapid environmental coverage, but flight time is often limited by energy constraints, whereas Unmanned Gr…

2026

CycleChemist: A Dual-Pronged Machine Learning Framework for Organic Photovoltaic Discovery

AAAI 2026technical

Organic photovoltaic (OPV) materials offer a promising pathway for sustainable energy generation. However, their development is hindered by the challenge of identifying high-performance donor-acceptor pairs with optimal power conversion efficiencies (PCEs). Most existing design strategies focus excl

Cited by 0SourcePDFScholar
2026

Diff-NAT: Better Naturalistic and Aggressive Adversarial Attacks via Class-Optimized Diffusion for Object Detection

AAAI 2026technical

Recent advances in naturalistic physical adversarial patch generation show great promise in protecting personal privacy against detector-based malicious surveillance while remaining inconspicuous to human observers. In this work, we present the first systematic categorization and in-depth re-examina

Cited by 0SourcePDFScholar
2026

Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing

ICLR 2026poster

Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs for text generation, with the potential to decode multiple tokens in a single iteration. However, none of the existing open-source dLLMs have achieved superior inference speed over AR LLMs of…

Cited by 0SourcecodeScholar
2026

DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing

ICLR 2026poster

With the rapid progress of video generation, demand for customized video editing is surging, where subject swapping constitutes a key component yet remains under-explored. Prevailing swapping approaches either specialize in narrow domains—such as human-body animation or hand-object interaction—or re…

Cited by 0SourceScholar
2026

Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed

ICML 2026poster

Diffusion language models (dLMs) have emerged as a promising paradigm enabling parallel generation, but their learning efficiency lags behind that of autoregressive (AR) language models when trained from scratch. To this end, we study AR-to-dLM conversion, which transforms pretrained AR models into …

Cited by 0SourceScholar
2026

Fast and Accurate Causal Parallel Decoding using Jacobi Forcing

ICML 2026poster

Multi-token generation has emerged as a promising paradigm for accelerating language model inference, with the diffusion Large Language Models (dLLMs) as the most notable approach recently. Popular dLLMs like SDAR and Fast-dLLM v2 are post-trained on pre-trained AR models to minimize training cost w…

Cited by 0SourceScholar
2026

Fast-dLLM v2: Efficient Block-Diffusion LLM

ICLR 2026poster

Autoregressive (AR) large language models (LLMs) have achieved remarkable performance across a wide range of natural language tasks, yet their inherent sequential decoding limits inference efficiency. In this work, we propose Fast-dLLM v2, a carefully designed block diffusion language model (dLLM) t…

Cited by 0SourcecodeScholar
2026

Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding

ICLR 2026poster

Diffusion-based large language models (Diffusion LLMs) have shown promise for non-autoregressive text generation. However, the practical inference speed of open-sourced Diffusion LLMs often lags behind autoregressive models due to the lack of Key-Value (KV) Cache and quality degradation when decodin…

Cited by 0SourceScholar
2026

Invariant Feature Learning for Counterfactual Watch-time Prediction in Video Recommendation

AAAI 2026technical

Video recommendation systems heavily rely on user watch time feedback, making accurate watch time prediction a crucial task. However, this task inherently suffers from bias, as recommendation models tend to favor long-duration videos to maximize watch time. This issue, known as duration bias in the

Cited by 0SourcePDFScholar
2026

Learning Forward Looking Adaptation to Dynamic Payloads for Quadruped Locomotion Via Physics-Informed Neural Networks

ICRA 2026poster

Payload-adaptive locomotion is an essential capability for quadruped robots operating in real-world scenarios, particularly when tasked with transporting dynamic payloads. Existing approaches face fundamental limitations: reactive adaptation strategies respond too slowly to sudden payload changes, w…

Cited by 0Scholar
2026

Learning Human-Robot Collaboration via Heterogeneous-Agent Lyapunov Policy Optimization

ICML 2026oral

To improve generalization and resilience in human–robot collaboration (HRC), robots must handle the combinatorial diversity of human behaviors and contexts, motivating multi-agent reinforcement learning (MARL). However, inherent heterogeneity between robots and humans creates a rationality gap (RG) …

Cited by 0SourceScholar
2026

Local Covariate Selection for Average Causal Effect Estimation without Pretreatment and Causal Sufficiency Assumptions

ICML 2026spotlight

Causal effect estimation is a fundamental task in many scientific fields. Selecting appropriate covariates for adjustment is crucial for obtaining unbiased causal effects. However, most existing methods either rely on learning the global causal structure, assume the absence of latent variables, or i…

Cited by 0SourceScholar
2026

MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement

CVPR 2026

This paper focuses on a highly practical scenario: how to continue benefiting from the advantages of multi-modal image fusion under harsh conditions when only visible imaging sensors are available. To achieve this goal, we propose a novel concept of single image fusion, which extends conventional da

Cited by 0SourcecodeScholar
2026

MatchMask: Mask-Centric Generative Data Augmentation for Label-Scarce Semantic Segmentation

CVPR 2026

Current semantic segmentation models are very data-hungry and require massive costly pixel-wise human annotations. Generative data augmentation, which scales the train set using generative models, provides a potential remedy. In this paper, we propose MatchMask, a novel mask-centric generative data

Cited by 0SourceScholar
2026

MoLoRA: Boosting LLM-based End-to-end Speech Translation with Mixture of Low-rank Experts

AAAI 2026technical

Recently, End-to-End Speech Translation (E2E-ST) methods leveraging large language models (LLMs) have demonstrated strong generalization capabilities and excellent scalability by integrating pre-trained speech encoders with LLMs, where Low-Rank Adaptation (LoRA) is commonly used for parameter-effici

Cited by 0SourcePDFScholar
2026

Multi-Object System Identification from Videos

ICLR 2026poster

We introduce the challenging problem of multi-object system identification from videos, for which prior methods are ill-suited due to their focus on single-object scenes or discrete material classification with a fixed set of material prototypes. To address this, we propose MOSIV, a new framework th…

Cited by 0SourceScholar
2026

No outlier channels but with outlier blocks

ICLR 2026poster

With the rapid scaling of large language models, achieving efficient compression while maintaining model performance has become a critical challenge. To address the limitations of existing non-uniform quantization methods, which typically rely on fixed codebooks and require costly optimization, we p…

Cited by 0SourceScholar
2026

Occlusion-Robust Relative Pose Estimation for Multi-Robot Systems Via Geometric-Aware Diffusion Matching

ICRA 2026poster

Relative pose estimation is crucial for coordinated multi-robot navigation. However, robots in close proximity often face intra-team occlusions, where teammates partially block each other's field of view, while dynamic environments further introduce environmental occlusions. Classical relative pose …

Cited by 0Scholar
2026

PKR-QA: A Benchmark for Procedural Knowledge Reasoning with Knowledge Module Learning

AAAI 2026technical

We introduce PKR-QA (Procedural Knowledge Reasoning Question Answering), a new benchmark for question answering over procedural tasks that require structured reasoning. PKR-QA is constructed semi-automatically using a procedural knowledge graph (PKG), which encodes task-specific knowledge across div

Cited by 0SourcePDFScholar
2026

Powerful and Theoretically Guaranteed Independence Testing on Heterogeneous Federated Clients

ICML 2026poster

In this paper, we present a novel federated independence testing method that addresses both theoretical and practical challenges arising from client heterogeneity. We begin by revisiting existing federated independence testing methods and showing why they fail to provide valid guarantees or maintain…

Cited by 0SourceScholar
2026

ProSAR: Prototype-Guided Semantic Augmentation and Refinement for Time Series Contrastive Learning

ICML 2026poster

Contrastive learning has advanced the representation learning across domains, yet its success relies on data augmentations that preserve semantic contents while providing the view diversities. Multivariate time series, however, are inherently noisy, non-stationary, and lack such intuitive semantic c…

Cited by 0SourceScholar
2026

Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

ICML 2026poster

Token-level reweighting is a simple yet effective mechanism for controlling supervised fine-tuning, but common indicators are largely one-dimensional: the ground-truth probability reflects downstream alignment, while token entropy reflects intrinsic uncertainty induced by the pre-training prior. Ign…

Cited by 0SourceScholar
2026

ReCoFuse: Ultra-Robust Image Fusion via Restorative Multi-Modal Diffusion Reciprocal Coupling

CVPR 2026

Existing methods following the integrated hard-regression or decoupling optimization paradigms exhibit limited fusion performance under complex degradations. To address these paradigm-level shortcomings, we propose ReCoFuse, an ultra-robust image fusion framework based on restorative multi-modal dif

Cited by 0SourcecodeScholar
2026

RigMo: Unifying Rig and Motion Learning for Generative Animation

CVPR 2026

Despite significant progress in 4D generation, rig and motion--the core structural and dynamic components of animation--are typically modeled as separate problems. Existing pipelines rely on ground-truth skeletons and skinning weights for motion generation and treat auto-rigging as an independent pr

Cited by 0SourceScholar
2026

Robust Fusion Controller: Degradation-Aware Image Fusion with Fine-Grained Language Instructions

AAAI 2026technical

Current image fusion methods struggle to adapt to real-world environments encompassing diverse degradations with spatially varying characteristics. To address this challenge, we propose a robust fusion controller (RFC) capable of achieving degradation-aware image fusion through fine-grained language

Cited by 0SourcePDFScholar
2026

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

ICLR 2026oral

We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720×1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong text-video alignment at a remarkably fast speed, deployable on RTX 5090 GPU. Tw…

Cited by 0SourcecodeScholar
2026

SGPFeat: Semantic and Geometric Priors for Multi-modal Image Matching

AAAI 2026technical

Multi-modal image matching is a fundamental task in multi-view and multi-modal image processing. Its key challenge lies in extracting features that remain consistent despite drastic appearance variations across modalities. However, the learning of the feature is hindered by the scarcity and the inac

Cited by 0SourcePDFScholar
2026

ShieldedCode: Learning Robust Representations for Virtual Machine Protected Code

ICLR 2026poster

Large language models (LLMs) have achieved remarkable progress in code generation, yet their potential for software protection remains largely untapped. Reverse engineering continues to threaten software security, while traditional virtual machine protection (VMP) relies on rigid, rule-based tran…

Cited by 0SourceScholar
2026

Stop Mixing Things Up! BISCUIT Teaches Vision-Language Models to Learn New Concepts from Images on the Spot

AAAI 2026technical

Vision-Language Models (VLMs) have achieved impressive performance across various tasks, but often struggle to apply newly introduced visual concepts during inference. A common failure pattern is what we call Mixing Things Up: VLMs frequently confuse concept names, resulting in vague descriptions an

Cited by 0SourcePDFScholar
2026

Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections

ICML 2026oral

Multimodal agents offer a compelling path to automating complex document-intensive workflows, yet a critical question remains: do these architectures demonstrate genuine strategic reasoning, or simply conduct stochastic trial-and-error search? To address this, we introduce Agentic Document VQA, a be…

Cited by 0SourceScholar
2026

Streaming Covariate Balancing via Discrepancy-Based Feature Coresets

ICML 2026poster

Real-time estimation of average treatment effects (ATE) in streaming observational data poses two key challenges: strict memory constraints that preclude storing the full data history, and distributional shifts in both treatment assignment and outcome-generating process. Existing methods either requ…

Cited by 0SourceScholar
2026

Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMs

ICLR 2026poster

Multi-Agent System (MAS) and Reinforcement Learning (RL) are both widely adopted to improve large language model (LLM) agentic performance. MAS strengthens task-specialized performance via role-based orchestration; RL leverages environment rewards to train stronger policies, such as Group Relative P…

Cited by 0SourcecodeScholar
2026

TINY BUT MIGHTY: A SOFTWARE-HARDWARE CO- DESIGN APPROACH FOR EFFICIENT MULTIMODAL IN- FERENCE ON BATTERY-POWERED SMALL DEVICES

ICLR 2026poster

Large Multimodal Models (LMMs) are inherently modular, consisting of vision and audio encoders, projectors, and large language models. Yet, they are almost always executed monolithically, which underutilizes the heterogeneous accelera- tors (NPUs, GPUs, DSPs) in modern SoCs and leads to high end-to-…

Cited by 0SourceScholar
2026

Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Framework

CVPR 2026

Existing mobile devices are constrained by compact optical designs, such as small apertures, which make it difficult to produce natural, optically realistic bokeh effects. Although recent learning-based methods have shown promising results, they still struggle with photos captured under high digital

Cited by 0SourcecodeScholar
2026

Tracking Control of Biomimetic Wave-Spiral Robot Based on Deep Reinforcement Learning

RA-L 2026

Drawing inspiration from piscine locomotion strategies, biomimetic underwater robotic systems demonstrate enhanced operational efficiency and stealth capabilities when executing target tracking missions within dynamic aquatic environments. However, challenges related to their operational speed and c

Cited by 0SourceScholar
2026

Video-KTR: Reinforcing Video Reasoning via Key Token Attribution

ICLR 2026poster

Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models (MLLMs), yet existing video reasoning methods often rely on coarse sequence-level rewards or single-factor token selection. Such approaches neglect fine-grained links among visual input…

Cited by 5SourcecodeScholar
2026

VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion

CVPR 2026

Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due to the scarcity of large-scale multi-sensor video datasets, l

Cited by 0SourcecodeScholar
2026

When Drafts Evolve: Speculative Decoding Meets Online Learning

ICML 2026poster

Speculative decoding has emerged as a widely adopted paradigm for accelerating large language model inference, where a lightweight draft model rapidly generates candidate tokens that are then verified in parallel by a larger target model. However, due to limited model capacity, drafts often struggle…

Cited by 0SourceScholar
2026

d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation

ICML 2026poster

Diffusion large language models (dLLMs) offer capabilities beyond those of autoregressive (AR) LLMs, such as parallel decoding and random-order generation. However, realizing these benefits in practice is non-trivial, as dLLMs inherently face an *accuracy-parallelism trade-off*. Despite increasing i…

Cited by 0SourceScholar
2026

lmgame-Bench: How Good are LLMs at Playing Games?

ICLR 2026poster

Playing video games requires perception, reasoning, memory, and long-horizon planning—exactly the faculties expected of modern large language and vision–language models (LLMs/VLMs). We introduce LMGame-Bench, a benchmark built on six popular games spanning platformer, puzzle, and narrative games thr…

Cited by 0SourcecodeScholar
2025

3DMolFormer: A Dual-channel Framework for Structure-based Drug Discovery

ICLR 2025poster

Structure-based drug discovery, encompassing the tasks of protein-ligand docking and pocket-aware 3D drug design, represents a core challenge in drug discovery. However, no existing work can deal with both tasks to effectively leverage the duality between them, and current methods for each task are…

2025

A Dual-Circuit Magnetic Actuation System for Multi-Robot Collaboration in Large-Scale Medical Environments

RA-L 2025

Untethered miniature robots, after ultra-long-distance transportation to the lesion by continuum robots, can further deliver drugs to the deep fine tissues. Specifically, the magnetic steering continuum robot with the follower-the-leader manner enhances the safety of channel construction, while the

Cited by 3SourceScholar
2025

A Fast Saturation Based Dehazing Framework with Accelerated Convolution and Attention Block

ICASSP 2025accepted

Real-time image dehazmg is crucial for applications such as autonomous driving, surveillance, and remote sensing, where haze can significantly reduce visibility. However, many deep learning algorithms are hindered by large model sizes, making real-time processing difficult to achieve. Several fast a…

Cited by 0SourceScholar
2025

Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger

ACL 2025long

Large language models (LLMs) have shown remarkable emergent capabilities, transforming the execution of functional tasks by leveraging external tools for complex problems that require specialized processing or up-to-date data. While existing research expands LLMs access to diverse tools (e.g., progr…

Cited by 0SourcePDFScholar
2025

Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning

AAAI 2025technical

Image Aesthetic Assessment (IAA) is a vital and intricate task that entails analyzing and assessing an image's aesthetic values, and identifying its highlights and areas for improvement. Traditional methods of IAA often concentrate on a single aesthetic task and suffer from inadequate labeled datase…

Cited by 0SourcePDFScholar
2025

Alleviating Hallucinations in Large Language Models through Multi-Model Contrastive Decoding and Dynamic Hallucination Detection

NeurIPS 2025poster

Despite their outstanding performance in numerous applications, large language models (LLMs) remain prone to hallucinations, generating content inconsistent with their pretraining corpora. Currently, almost all contrastive decoding approaches alleviate hallucinations by introducing a model susceptib…

Cited by 0SourceScholar
2025

Analyzing LLMs’ Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations

ACL 2025long

While understanding the knowledge boundaries of LLMs is crucial to prevent hallucination, research on the knowledge boundaries of LLMs has predominantly focused on English. In this work, we present the first study to analyze how LLMs recognize knowledge boundaries across different languages by probi…

2025

ArcPro: Architectural Programs for Structured 3D Abstraction of Sparse Points

CVPR 2025highlight

We introduce ArcPro, a novel learning framework built on architectural programs to recover structured 3D abstractions from highly sparse and low-quality point clouds. Specifically, we design a domain-specific language (DSL) to hierarchically represent building structures as a program, which can be e…

Cited by 0SourcePDFScholar
2025

Bandwidth-Adaptive Spatiotemporal Correspondence Identification for Collaborative Perception

ICRA 2025

Correspondence identification (CoID) is an essential capability in multi-robot collaborative perception, which enables a group of robots to consistently refer to the same objects within their respective fields of view. In real-world applications, such as connected autonomous driving, vehicles face c

Cited by 1SourcecodeScholar
2025

Boosting Vision State Space Model with Fractal Scanning

AAAI 2025technical

Recently, foundational models have significantly advanced in different tasks, accompanied by Transformer as the general backbone. However, Transformer's quadratic complexity poses challenges for handling longer sequences and higher resolution images, which may limit foundational models further devel…

Cited by 0SourcePDFScholar
2025

CoIR: A Comprehensive Benchmark for Code Information Retrieval Models

ACL 2025long

Despite the substantial success of Information Retrieval (IR) in various NLP tasks, most IR systems predominantly handle queries and corpora in natural language, neglecting the domain of code retrieval. Code retrieval is critically important yet remains under-explored, with existing methods and benc…

2025

Coordinated Multi-Robot Navigation with Formation Adaptation

ICRA 2025

Coordinated multi-robot navigation is an essential ability for a team of robots operating in diverse environments. Robot teams often need to maintain specific formations, such as wedge formations, to enhance visibility, positioning, and efficiency during fast movement. However, complex environments

Cited by 4SourceScholar
2025

CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Control

ACL 2025finding

Retrieval-augmented generation (RAG) has emerged as a promising solution for mitigating hallucinations of large language models (LLMs) with retrieved external knowledge. Adaptive RAG enhances this approach by enabling dynamic retrieval during generation, activating retrieval only when the query exce…

2025

Data-Driven Selection of Instrumental Variables for Additive Nonlinear, Constant Effects Models

ICML 2025poster

We consider the problem of selecting instrumental variables from observational data, a fundamental challenge in causal inference. Existing methods mostly focus on additive linear, constant effects models, limiting their applicability in complex real-world scenarios. In this paper, we tackle a more…

Cited by 0SourcePDFScholar
2025

Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Models

CVPR 2025poster

Concept Bottleneck Models (CBMs) try to make the decision-making process transparent by exploring an intermediate concept space between the input image and the output prediction. Existing CBMs just learn coarse-grained relations between the whole image and the concepts, less considering local image…

Cited by 1SourcePDFScholar
2025

Efficient Constraint-based Window Causal Graph Discovery in Time Series with Multiple Time Lags

IJCAI 2025

We address the identification of direct causes in time series with multiple time lags, and propose a constraint-based window causal graph discovery method. A key advantage of our method is that the number of required conditional independence (CI) tests scales quadratically with the number of sub-ser

Cited by 0SourcePDFScholar
2025

Efficiently Scaling LLM Reasoning Programs with Certaindex

NeurIPS 2025poster

Test-time reasoning algorithms such as chain-of-thought, self-consistency, and MCTS enhance LLM problem-solving but can wastefully generate many tokens without improving accuracy. At the same time, we observe that these algorithms exhibit answer stabilization: their intermediate solutions often ceas…

Cited by 36SourcecodeScholar
2025

Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification

CVPR 2025poster

Concept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequentl…

2025

Fast Video Generation with Sliding Tile Attention

ICML 2025poster

Diffusion Transformers (DiTs) with 3D full attention power state-of-the-art video generation, but suffer from prohibitive compute cost -- when generating just a 5-second 720P video, attention alone takes 800 out of 950 seconds of total inference time. This paper introduces sliding tile attention (ST…

Cited by 5SourcePDFScholar
2025

Faster Video Diffusion with Trainable Sparse Attention

NeurIPS 2025poster

Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at both traini…

Cited by 0SourcecodeScholar
2025

FedSMU: Communication-Efficient and Generalization-Enhanced Federated Learning through Symbolic Model Updates

ICML 2025poster

The significant communication overhead and client data heterogeneity have posed an important challenge to current federated learning (FL) paradigm. Existing compression-based and optimization-based FL algorithms typically focus on addressing either the model compression challenge or the data heterog…

Cited by 0SourcePDFScholar
2025

FineReason: Evaluating and Improving LLMs’ Deliberate Reasoning through Reflective Puzzle Solving

ACL 2025long

Many challenging reasoning tasks require not just rapid, intuitive responses, but a more deliberate, multi-step approach. Recent progress in large language models (LLMs) highlights an important shift from the “System 1” way of quick reactions to the “System 2” style of reflection-and-correction prob…

2025

Frame-Voyager: Learning to Query Frames for Video Large Language Models

ICLR 2025poster

Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it impractical to input entire videos. Existing frame selection approaches, such as uniform frame sampling and text-frame r…

Cited by 8SourcePDFScholar
2025

FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes

CVPR 2025poster

We propose FreeSim, a camera simulation method for driving scenes via 3D Gaussian Splatting and diffusion-based image generation. FreeSim emphasizes high-quality rendering from viewpoints beyond the recorded ego trajectories. In such viewpoints, previous methods have unacceptable degradation because…

Cited by 8SourcePDFScholar
2025

GALA: Geometry-Aware Local Adaptive Grids for Detailed 3D Generation

ICLR 2025poster

We propose GALA, a novel representation of 3D shapes that (i) excels at capturing and reproducing complex geometry and surface details, (ii) is computationally efficient, and (iii) lends itself to 3D generative modelling with modern, diffusion-based schemes. The key idea of GALA is to exploit both t…

2025

GameArena: Evaluating LLM Reasoning through Live Computer Games

ICLR 2025poster

Evaluating the reasoning abilities of large language models (LLMs) is challenging. Existing benchmarks often depend on static datasets, which are vulnerable to data contamination and may get saturated over time, or on binary live human feedback that conflates reasoning with other abilities. As the m…

Cited by 2SourcePDFScholar
2025

GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning

EMNLP 2025

Recent advancements in reinforcement learning (RL) have enhanced the reasoning abilities of large language models (LLMs), yet the impact on multimodal LLMs (MLLMs) is limited. Particularly in vision-intensive tasks like geometric reasoning, MLLMs hallucinate frequently, leading to inaccurate reasoni

2025

Geometry and Force-Informed Robotic Assembly with Small Relative Initial Deviations for Circular Electrical Connectors

ICRA 2025

Circular electrical connectors (CECs) have a wide range of applications in scenarios that require reliable connections. However, sockets are often located in narrow scenes with random spatial orientations, complex lighting conditions, and obstructions from cables, making it difficult to accurately l

Cited by 0SourceScholar
2025

HEATS: A Hierarchical Framework for Efficient Autonomous Target Search with Mobile Manipulators

IROS 2025

Utilizing robots for autonomous target search in complex and unknown environments can greatly improve the efficiency of search and rescue missions. However, existing methods have shown inadequate performance due to hardware platform limitations, inefficient viewpoint selection strategies, and conser

Cited by 3SourceScholar
2025

High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity

ICLR 2025poster

In the realm of high-resolution (HR), fine-grained image segmentation, the primary challenge is balancing broad contextual awareness with the precision required for detailed object delineation, capturing intricate details and the finest edges of objects. Diffusion models, trained on vast datasets co…

2025

IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A

ICCV 2025poster

Existing human motion Q&A methods rely on explicit program execution, where the requirement for manually defined functional modules may limit the scalability and adaptability. To overcome this, we propose an implicit program-guided motion reasoning (IMoRe) framework that unifies reasoning across mul…

2025

Identifying Causal Mechanism Shifts Under Additive Models with Arbitrary Noise

IJCAI 2025

In many real-world scenarios, the goal is to identify variables whose causal mechanisms change across related datasets. For example, detecting abnormal root nodes in manufacturing, and identifying key genes that influence cancer by analyzing differences in gene regulatory mechanisms between healthy

Cited by 0SourcePDFScholar
2025

Improving Open-vocabulary Video Visual Relation Detection with Decomposed Prompt Learning and Relation Adjustment

ICASSP 2025accepted

Open-vocabulary video visual relation detection (VidVRD) expands the scope of detecting object relations in videos to include unseen categories. It marks considerable advancement in recognizing novel relations solely by training on a base set, thus extending the frontiers of automated video understa…

Cited by 0SourceScholar
2025

LBI-FL: Low-Bit Integerized Federated Learning with Temporally Dynamic Bit-Width Allocation

ICML 2025poster

Federated learning (FL) is greatly challenged by the communication bottleneck and computation limitation on clients. Existing methods based on quantization for FL cannot simultaneously reduce the uplink and downlink communication cost and mitigate the computation burden on clients. To address this p…

Cited by 0SourcePDFScholar
2025

LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

ICLR 2025spotlight

Visual instruction tuning has made considerable strides in enhancing the capabilities of Large Multimodal Models (LMMs). However, existing open LMMs largely focus on single-image tasks, their applications to multi-image scenarios remains less explored. Additionally, prior LMM research separately tac…

Cited by 0SourcePDFScholar
2025

Learning Adaptive Lighting via Channel-Aware Guidance

ICML 2025poster

Learning lighting adaptation is a crucial step in achieving good visual perception and supporting downstream vision tasks. Current research often addresses individual light-related challenges, such as high dynamic range imaging and exposure correction, in isolation. However, we identify shared funda…

2025

Local Identifying Causal Relations in the Presence of Latent Variables

ICML 2025spotlight

We tackle the problem of identifying whether a variable is the cause of a specified target using observational data. State-of-the-art causal learning algorithms that handle latent variables typically rely on identifying the global causal structure, often represented as a partial ancestral graph (PAG…

Cited by 0SourcePDFScholar
2025

Local Learning for Covariate Selection in Nonparametric Causal Effect Estimation with Latent Variables

NeurIPS 2025poster

Estimating causal effects from nonexperimental data is a fundamental problem in many fields of science. A key component of this task is selecting an appropriate set of covariates for confounding adjustment to avoid bias. Most existing methods for covariate selection often assume the absence of laten…

Cited by 0SourceScholar
2025

Long-form Hallucination Detection with Self-elicitation

ACL 2025finding

While Large Language Models (LLMs) have exhibited impressive performance in generating long-form content, they frequently present a hazard of producing factual inaccuracies or hallucinations. An effective strategy to mitigate this hazard is to leverage off-the-shelf LLMs to detect hallucinations aft…

Cited by 0SourcePDFScholar
2025

Loss-Aware Curriculum Learning for Chinese Grammatical Error Correction

ICASSP 2025accepted

Chinese grammatical error correction (CGEC) aims to detect and correct errors in the input Chinese sentences. Recently, Pre-trained Language Models (PLMS) have been employed to improve the performance. However, current approaches ignore that correction difficulty varies across different instances an…

Cited by 0SourceScholar
2025

MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection

ICLR 2025poster

KV cache has become a *de facto* technique for the inference of large language models (LLMs), where tensors of shape (layer number, head number, sequence length, feature dimension) are introduced to cache historical information for self-attention. As the size of the model and data grows, the KV cac…

Cited by 3SourcePDFScholar
2025

MetaMixSpeech: Meta Task Augmentation for Low-Resource Speech Recognition

EMNLP 2025

Meta-learning has proven to be a powerful paradigm for effectively improving the performance of low-resource speech recognition by learning generalizable knowledge across multiple tasks. However, multilingual meta learning also faces challenges such as task overfitting and learner overfitting, there

Cited by 0SourcePDFScholar
2025

MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity Recognition

NeurIPS 2025poster

Human Activity Recognition (HAR) with wearable sensors is challenged by limited interpretability, which significantly impacts cross-dataset generalization. To address this challenge, we propose Motion-Primitive Transformer (MoPFormer), a novel self-supervised framework that enhances interpretability…

Cited by 0SourceScholar
2025

Multi-Task Dense Predictions via Unleashing the Power of Diffusion

ICLR 2025poster

Diffusion models have exhibited extraordinary performance in dense prediction tasks. However, there are few works exploring the diffusion pipeline for multi-task dense predictions. In this paper, we unlock the potential of diffusion models in solving multi-task dense predictions and propose a novel…

2025

Non-Overlap-Aware Egocentric Pose Estimation for Collaborative Perception in Connected Autonomy

IROS 2025

Egocentric pose estimation is a fundamental capability for multi-robot collaborative perception in connected autonomy, such as connected autonomous vehicles. During multi-robot operations, a robot needs to know the relative pose between itself and its teammates with respect to its own coordinates. H

Cited by 0SourceScholar
2025

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

CVPR 2025poster

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding and generation abilities. While the progress in unified mode…

2025

PALMBENCH: A COMPREHENSIVE BENCHMARK OF COMPRESSED LARGE LANGUAGE MODELS ON MOBILE PLATFORMS

ICLR 2025poster

Deploying large language models (LLMs) locally on mobile devices is advantageous in scenarios where transmitting data to remote cloud servers is either undesirable due to privacy concerns or impractical due to network connection. Recent advancements have facilitated the local deployment of LLMs. How…

Cited by 2SourcePDFScholar
2025

PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling

ICCV 2025poster

Skinning and rigging are fundamental components in animation, articulated object reconstruction, motion transfer, and 4D generation. Existing approaches predominantly rely on Linear Blend Skinning (LBS), due to its simplicity and differentiability. However, LBS introduces artifacts such as volume lo…

Cited by 0SourcePDFScholar
2025

Planning with Multi-Constraints via Collaborative Language Agents

COLING 2025main

The rapid advancement of neural language models has sparked a new surge of intelligent agent research. Unlike traditional agents, large language model-based agents (LLM agents) have emerged as a promising paradigm for achieving artificial general intelligence (AGI) due to their superior reasoning an…

2025

Predicting Compact Phrasal Rewrites with Large Language Models for ASR Post Editing

ICASSP 2025accepted

Large Language Models (LLMs) excel at rewriting tasks such as text style transfer and grammatical error correction. While there is considerable overlap between the inputs and outputs in these tasks, the decoding cost still increases with output length, regardless of the amount of overlap. By leverag…

Cited by 0SourceScholar
2025

Preference Alignment Improves Language Model-Based TTS

ICASSP 2025accepted

Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts. Further optimization can be achieved through preference alignment algorithms, which adjust LMs to align with the preferences of reward models, enhancing…

Cited by 0SourceScholar
2025

Projection-Manifold Regularized Latent Diffusion for Robust General Image Fusion

NeurIPS 2025poster

This study proposes PDFuse, a robust, general training-free image fusion framework built on pre-trained latent diffusion models with projection–manifold regularization. By redefining fusion as a diffusion inference process constrained by multiple source images, PDFuse can adapt to varied image modal…

Cited by 0SourcecodeScholar
2025

RAPID: Efficient Retrieval-Augmented Long Text Generation with Writing Planning and Information Discovery

ACL 2025finding

Generating knowledge-intensive and comprehensive long texts, such as encyclopedia articles, remains significant challenges for Large Language Models. It requires not only the precise integration of facts but also the maintenance of thematic coherence throughout the article. Existing methods, such as…

2025

RINA: Rapid Introspective Neural Adaptation for Out-of-Distribution Payload Configurations on Quadruped Robots

ICRA 2025

Adaptive locomotion is a fundamental capability for quadruped robots, particularly in real-world scenarios when they must transport novel or out-of-distribution (O.O.D.) payloads across diverse terrains. Previous learning-based methods often tightly couple a locomotion controller's learned parameter

Cited by 0SourceScholar
2025

ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning

EMNLP 2025

Reasoning-based large language models have excelled in mathematics and programming, yet their potential in knowledge-intensive medical question answering remains underexplored and insufficiently validated in clinical contexts. To bridge this gap, we introduce ReasonMed , the largest medical reasonin

2025

Reversal of Thought: Enhancing Large Language Models with Preference-Guided Reverse Reasoning Warm-up

ACL 2025long

Large language models (LLMs) have shown remarkable performance in reasoning tasks but face limitations in mathematical and complex logical reasoning. Existing methods to improve LLMs’ logical capabilities either involve traceable or verifiable logical sequences that generate more reliable responses…

2025

Reverse Modeling in Large Language Models

NAACL 2025short

Humans are accustomed to reading and writing in a forward manner, and this natural bias extends to text understanding in auto-regressive large language models (LLMs). This paper investigates whether LLMs, like humans, struggle with reverse modeling, specifically with reversed text inputs. We found t…

Cited by 0SourcePDFScholar
2025

SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models

NeurIPS 2025poster

Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires mastering two core capabilities: video referring understanding, which captures the semantics of video regions, and video…

Cited by 0SourceScholar
2025

SDMatte: Grafting Diffusion Models for Interactive Matting

ICCV 2025poster

Recent interactive matting methods have demonstrated satisfactory performance in capturing the primary regions of objects, but they fall short in extracting fine-grained details in edge regions. Diffusion models trained on billions of image-text pairs, demonstrate exceptional capability in modeling…

2025

SERENA: A Unified Stochastic Recursive Variance Reduced Gradient Framework for Riemannian Non-Convex Optimization

ICML 2025poster

Recently, the expansion of Variance Reduction (VR) to Riemannian stochastic non-convex optimization has attracted increasing interest. Inspired by recursive momentum, we first introduce Stochastic Recursive Variance Reduced Gradient (SRVRG) algorithm and further present Stochastic Recursive Gradient…

Cited by 0SourcePDFScholar
2025

SafetyQuizzer: Timely and Dynamic Evaluation on the Safety of LLMs

NAACL 2025long

With the expansion of the application of Large Language Models (LLMs), concerns about their safety have grown among researchers. Numerous studies have demonstrated the potential risks of LLMs generating harmful content and have proposed various safety assessment benchmarks to evaluate these risks. H…

2025

Scalable Trajectory-User Linking with Dual-Stream Representation Networks

AAAI 2025technical

Trajectory-user linking (TUL) aims to match anonymous trajectories to the most likely users who generated them, offering benefits for a wide range of real-world spatio-temporal applications. However, existing TUL methods are limited by high model complexity and poor learning of the effective represe…

2025

Scaling Language-centric Omnimodal Representation Learning

NeurIPS 2025poster

Recent multimodal embedding approaches leveraging multimodal large language models (MLLMs) fine-tuned with contrastive learning (CL) have shown promising results, yet the underlying reasons behind their superiority remain underexplored. This work argues that a crucial advantage of MLLM-based approac…

Cited by 0SourcecodeScholar
2025

Scaling Long Context Training Data by Long-Distance Referrals

ICLR 2025poster

Training large language models for long context understanding faces the challenge of data shortage. Previous data engineering approaches mechanically concatenate short documents, which may create many pseudo long documents but raise concerns about data quality. In this paper, we study the core attri…

Cited by 0SourcePDFScholar
2025

Self-Reflective Perceptual Adaptation for Robust Ground Navigation in Unstructured Off-Road Environments

ICRA 2025

Autonomous ground robots navigating unstructured off-road environments face perceptual challenges, such as sensor obscuration or failure, which can lead to inaccurate perception or navigation failures. While robot adaptation has recently gained increasing attention, self-reflective robot adaptation,

Cited by 0SourceScholar
2025

Sensitivity-LoRA : Low-Load Sensitivity-Based Fine-Tuning for Large Language Models

EMNLP 2025

Large Language Models (LLMs) have transformed both everyday life and scientific research. However, adapting LLMs from general-purpose models to specialized tasks remains challenging, particularly in resource-constrained environments. Low-Rank Adaptation (LoRA), a prominent method within Parameter-Ef

Cited by 0SourcePDFScholar
2025

SolEval: Benchmarking Large Language Models for Repository-level Solidity Smart Contract Generation

EMNLP 2025

Large language models (LLMs) have transformed code generation.However, most existing approaches focus on mainstream languages such as Python and Java, neglecting the Solidity language, the predominant programming language for Ethereum smart contracts.Due to the lack of adequate benchmarks for Solidi

2025

Spatial-Temporal Graph Contrastive Learning with Decreasing Masks for Traffic Flow Forecasting

IROS 2025

In recent years, Contrastive learning has shown great potential in traffic flow prediction tasks. However, existing contrastive learning methods have difficulties in dealing with missing data and noise, and it is difficult to fully capture local and global correlations by relying on a single contras

Cited by 0SourceScholar
2025

Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation

NeurIPS 2025spotlight

We present Stable Part Diffusion 4D (SP4D), a framework for generating paired RGB and kinematic part videos from monocular inputs. Unlike conventional part segmentation methods that rely on appearance-based semantic cues, SP4D learns to produce kinematic parts --- structural components aligned with…

Cited by 0SourceScholar
2025

Subteaming and Adaptive Formation Control for Coordinated Multi-Robot Navigation

CoRL 2025poster

Coordinated multi-robot navigation is essential for robots to operate as a team in diverse environments. During navigation, robot teams usually need to maintain specific formations, such as circular formations to protect human teammates at the center. However, in complex scenarios such as narrow c…

Cited by 0SourceScholar
2025

Sugar-Coated Poison: Benign Generation Unlocks Jailbreaking

EMNLP 2025

With the increasingly deep integration of large language models (LLMs) across diverse domains, the effectiveness of their safety mechanisms is encountering severe challenges. Currently, jailbreak attacks based on prompt engineering, which induce models to generate potentially harmful content, have b

2025

Towards Automatic Sampling of User Behaviors for Sequential Recommender Systems

IJCAI 2025

Sequential recommender systems (SRS) have gained increasing popularity due to their remarkable proficiency in capturing dynamic user preferences. In the current setup of SRS, a common configuration is to uniformly consider each historical behavior as a positive interaction. However, this setting has

2025

Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling

EMNLP 2025

Direct Prompt Injection (DPI) attacks pose a critical security threat to Large Language Models (LLMs) due to their low barrier of execution and high potential damage. To address the impracticality of existing white-box/gray-box methods and the poor transferability of black-box methods, we propose an

Cited by 0SourcePDFScholar
2025

UAV-DETR: Efficient End-to-End Object Detection for Unmanned Aerial Vehicle Imagery

IROS 2025

Unmanned aerial vehicle object detection (UAV-OD) has been widely used in various scenarios. However, most existing UAV-OD algorithms rely on manually designed components, which require extensive tuning. End-to-end models that do not depend on such manually designed components are mainly designed fo

Cited by 54SourcecodeScholar
2024

"DECOLLAGE: 3D Detailization by Controllable, Localized, and Learned Geometry Enhancement"

ECCV 2024poster

"We present a 3D modeling method which enables end-users to refine or detailize 3D shapes using machine learning, expanding the capabilities of AI-assisted 3D content creation. Given a coarse voxel shape (e.g., one produced with a simple box extrusion tool or via generative modeling), a user can dir…

Cited by 2SourcePDFScholar
2024

A Robust Mutual-Reinforcing Framework for 3D Multi-Modal Medical Image Fusion Based on Visual-Semantic Consistency

AAAI 2024technical

This work proposes a robust 3D medical image fusion framework to establish a mutual-reinforcing mechanism between visual fusion and lesion segmentation, achieving their double improvement. Specifically, we explore the consistency between vision and semantics by sharing feature fusion modules. Throug…

2024

Accurate and Efficient Loop Closure Detection With Deep Binary Image Descriptor and Augmented Point Cloud Registration

IROS 2024poster

Loop Closure Detection (LCD) is an essential component of Simultaneous Localization and Mapping (SLAM), helping to correct drift errors, facilitate map merging, or both by identifying previously observed scenes. Despite its importance, traditional LCD algorithms based on single sensor such as camera…

Cited by 0SourceScholar
2024

Active Coarse-to-Fine Segmentation of Moveable Parts from Real Images

ECCV 2024poster

"We introduce the first active learning (AL) model for high-accuracy instance segmentation of parts from RGB images of real indoor scenes. Specifically, our goal is to obtain fully validated segmentation results by humans while minimizing manual effort. To this end, we employ a transformer that util…

2024

AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models

EMNLP 2024finding

Mixture of experts (MoE) has become the standard for constructing production-level large language models (LLMs) due to its promise to boost model capacity without causing significant overheads. Nevertheless, existing MoE methods usually enforce a constant top-k routing for all tokens, which is argua…

2024

Advancing Acoustic Howling Suppression Through Recursive Training of Neural Networks

ICASSP 2024accepted

In this paper, we introduce a novel training framework designed to comprehensively address the acoustic howling issue by examining its fundamental formation process. This framework integrates a neural network (NN) module into the closed-loop system during training with signals generated recursively…

Cited by 0SourceScholar
2024

An Instruction Tuning-Based Contrastive Learning Framework for Aspect Sentiment Quad Prediction with Implicit Aspects and Opinions

EMNLP 2024finding

Aspect sentiment quad prediction (ASQP) is crucial in aspect-based sentiment analysis (ABSA). It involves identifying a text’s aspect,sentiment, opinion, and category. Existing methods have insufficiently explored how to effectively leverage the knowledge of pre-trainedlanguage models (PLMs) to hand…

2024

Bayesian Activity Detection for Massive Connectivity in Cell-Free IoT Networks

ICASSP 2024accepted

Activity detection is an important task in the next generation Internet-of-things (IoT) networks. Existing algorithms mostly require precise information about the network, such as large-scale fading, noise variance, and small-scale fading statistics. Acquiring such information would take a significa…

Cited by 0SourceScholar
2024

Beta-Tuned Timestep Diffusion Model

ECCV 2024poster

"Diffusion models have received a lot of attention in the field of generation due to their ability to produce high-quality samples. However, several recent studies indicate that treating all distributions equally in diffusion model training is sub-optimal. In this paper, we conduct an in-depth theor…

Cited by 12SourcePDFScholar
2024

Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

ICML 2024poster

Autoregressive decoding of large language models (LLMs) is memory bandwidth bounded, resulting in high latency and significant wastes of the parallel processing power of modern accelerators. Existing methods for accelerating LLM decoding often require a draft model (e.g., speculative decoding), whic…

2024

CRAYM: Neural Field Optimization via Camera RAY Matching

NeurIPS 2024poster

We introduce camera ray matching (CRAYM) into the joint optimization of camera poses and neural fields from multi-view images. The optimized field, referred to as a feature volume, can be “probed” by the camera rays for novel view synthesis (NVS) and 3D geometry reconstruction. One key reason for ma…

Cited by 1SourcePDFScholar
2024

CSL: Class-Agnostic Structure-Constrained Learning for Segmentation Including the Unseen

AAAI 2024technical

Addressing Out-Of-Distribution (OOD) Segmentation and Zero-Shot Semantic Segmentation (ZS3) is challenging, necessitating segmenting unseen classes. Existing strategies adapt the class-agnostic Mask2Former (CA-M2F) tailored to specific tasks. However, these methods cater to singular tasks, demand tr…

Cited by 13SourcePDFScholar
2024

Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding

ICLR 2024poster

Table-based reasoning with large language models (LLMs) is a promising direction to tackle many table understanding tasks, such as table-based question answering and fact verification. Compared with generic reasoning, table-based reasoning requires the extraction of underlying semantics from both fr…

Cited by 107SourcePDFScholar
2024

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

ICML 2024poster

Large Language Models (LLMs) have unlocked new capabilities and applications; however, evaluating the alignment with human preferences still poses significant challenges. To address this issue, we introduce Chatbot Arena, an open platform for evaluating LLMs based on human preferences. Our methodolo…

Cited by 554SourcePDFScholar
2024

Clarifying the Behavior and the Difficulty of Adversarial Training

AAAI 2024technical

Adversarial training is usually difficult to optimize. This paper provides conceptual and analytic insights into the difficulty of adversarial training via a simple theoretical study, where we derive an approximate dynamics of a recursive multi-step attack in a simple setting. Despite the simplicity…

Cited by 0SourcePDFScholar
2024

ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models

NeurIPS 2024spotlight

Multi-turn visual conversation is an important ability of real-world AI assistants. However, the related evaluation benchmark is missed. This paper presents ConvBench, a multi-turn conversation benchmark with hierarchical capabilities ablation evaluation for Large Vision-Language Models (LVLMs). Co…

2024

Cross-Scale Domain Adaptation with Comprehensive Information for Pansharpening

IJCAI 2024poster

Deep learning-based pansharpening methods typically use simulated data at the reduced-resolution scale for training. It limits their performance when generalizing the trained model to the full-resolution scale due to incomprehensive information utilization of panchromatic (PAN) images at the full-re…

2024

DPA-Net: Structured 3D Abstraction from Sparse Views via Differentiable Primitive Assembly

ECCV 2024poster

"We present a differentiable rendering framework to learn structured 3D abstractions in the form of primitive assemblies from sparse RGB images capturing a 3D object. By leveraging differentiable volume rendering, our method does not require 3D supervision. Architecturally, our network follows the g…

Cited by 4SourcePDFScholar
2024

DVD: Dynamic Contrastive Decoding for Knowledge Amplification in Multi-Document Question Answering

EMNLP 2024main

Large language models (LLMs) are widely used in question-answering (QA) systems but often generate information with hallucinations. Retrieval-augmented generation (RAG) offers a potential remedy, yet the uneven retrieval quality and irrelevant contents may distract LLMs.In this work, we address thes…

2024

Data Adaptive Traceback for Vision-Language Foundation Models in Image Classification

AAAI 2024technical

Vision-language foundation models have been incredibly successful in a wide range of downstream computer vision tasks using adaptation methods. However, due to the high cost of obtaining pre-training datasets, pairs with weak image-text correlation in the data exist in large numbers. We call them we…

Cited by 1SourcePDFScholar
2024

Deep Unfolded Network with Intrinsic Supervision for Pan-Sharpening

AAAI 2024technical

Existing deep pan-sharpening methods lack the learning of complementary information between PAN and MS modalities in the intermediate layers, and exhibit low interpretability due to their black-box designs. To this end, an interpretable deep unfolded network with intrinsic supervision for pan-sharpe…

2024

Dispel Darkness for Better Fusion: A Controllable Visual Enhancer based on Cross-modal Conditional Adversarial Learning

CVPR 2024poster

We propose a controllable visual enhancer named DDBF which is based on cross-modal conditional adversarial learning and aims to dispel darkness and achieve better visible and infrared modalities fusion. Specifically a guided restoration module (GRM) is firstly designed to enhance weakened informatio…

2024

Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model

ACL 2024findings

The detection of machine-generated text, especially from large language models (LLMs), is crucial in preventing serious social problems resulting from their misuse. Some methods train dedicated detectors on specific datasets but fall short in generalizing to unseen test data, while other zero-shot o…

Cited by 24SourcePDFScholar
2024

Efficient LLM Scheduling by Learning to Rank

NeurIPS 2024poster

In Large Language Model (LLM) inference, the output length of an LLM request is typically regarded as not known a priori. Consequently, most LLM serving systems employ a simple First-come-first-serve (FCFS) scheduling strategy, leading to Head-Of-Line (HOL) blocking and reduced throughput and servic…

2024

Efficiently Learning Significant Fourier Feature Pairs for Statistical Independence Testing

NeurIPS 2024poster

We propose a novel method to efficiently learn significant Fourier feature pairs for maximizing the power of Hilbert-Schmidt Independence Criterion~(HSIC) based independence tests. We first reinterpret HSIC in the frequency domain, which reveals its limited discriminative power due to the inability…

Cited by 0SourcePDFScholar
2024

Enhanced Deep Reinforcement Learning for Parcel Singulation in Non-Stationary Environments

ICASSP 2024accepted

In the rapidly expanding logistics sector, parcel singulation has emerged as a significant bottleneck. To address this, we propose an automated parcel singulator utilizing a sparse actuator array, which presents an optimal balance between cost and efficiency, albeit requiring a sophisticated control…

Cited by 0SourceScholar
2024

Enhancing Reinforcement Learning via Causally Correct Input Identification and Targeted Intervention

ICASSP 2024accepted

Causal confusion, characterized by the learning of spurious correlations, detrimentally affects the generalization and effectiveness of reinforcement learning (RL) algorithms, especially in environments without latent confounders often encountered in robot autonomous navigation tasks. This study add…

Cited by 0SourceScholar
2024

EpiGEN: An Efficient Multi-Api Code GENeration Framework under Enterprise Scenario

COLING 2024main

In recent years, Large Language Models (LLMs) have demonstrated exceptional performance in code-generation tasks. However, under enterprise scenarios where private APIs are pre-built, general LLMs often fail to meet expectations. Existing approaches are confronted with drawbacks of high resource con…

Cited by 0SourcePDFScholar
2024

Explaining Generalization Power of a DNN Using Interactive Concepts

AAAI 2024technical

This paper explains the generalization power of a deep neural network (DNN) from the perspective of interactions. Although there is no universally accepted definition of the concepts encoded by a DNN, the sparsity of interactions in a DNN has been proved, i.e., the output score of a DNN can be well…

Cited by 18SourcePDFScholar
2024

Exploiting Hybrid Policy in Reinforcement Learning for Interpretable Temporal Logic Manipulation

IROS 2024

Reinforcement Learning (RL) based methods have been increasingly explored for robot learning. However, RL based methods often suffer from low sampling efficiency in the exploration phase, especially for long-horizon manipulation tasks, and generally neglect the semantic information from the task lev

Cited by 1SourcecodeScholar
2024

Glance, Focus and Refinement Network for Remote Sensing Change Detection

ICASSP 2024accepted

Existing change detection (CD) methods often directly fuse the multi-level features from bi-temporal remote sensing images without discriminatively considering each pixel's importance. Despite the demonstrated success, unselectively mixing the features degrades the model's performance to effectively…

Cited by 0SourceScholar
2024

Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

ECCV 2024poster

"In this paper, we develop an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category names or referring expressions. The key solution of open-set object detection i…

2024

Improving Adversarial Energy-Based Model via Diffusion Process

ICML 2024poster

Generative models have shown strong generation ability while efficient likelihood estimation is less explored. Energy-based models (EBMs) define a flexible energy function to parameterize unnormalized densities efficiently but are notorious for being difficult to train. Adversarial EBMs introduce a…

Cited by 5SourcePDFScholar
2024

Improving Generalization in Federated Learning with Model-Data Mutual Information Regularization: A Posterior Inference Approach

NeurIPS 2024poster

Most of existing federated learning (FL) formulation is treated as a point-estimate of models, inherently prone to overfitting on scarce client-side data with overconfident decisions. Though Bayesian inference can alleviate this issue, a direct posterior inference at clients may result in biased loc…

Cited by 0SourcePDFScholar
2024

InferCept: Efficient Intercept Support for Augmented Large Language Model Inference

ICML 2024poster

Large language models are increasingly integrated with external environments, tools, and agents like ChatGPT plugins to extend their capability beyond language-centric tasks. However, today's LLM inference systems are designed for standalone LLMs. They treat each external interaction as the end of L…

2024

Interfacing Foundation Models' Embeddings

NeurIPS 2024poster

Foundation models possess strong capabilities in reasoning and memorizing across modalities. To further unleash the power of foundation models, we present FIND, a generalized interface for aligning foundation models' embeddings with unified image and dataset-level understanding spanning modality and…

2024

LEEPS: Learning End-to-End Legged Perceptive Parkour Skills on Challenging Terrains

IROS 2024poster

Empowering legged robots with agile maneuvers is a great challenge. While existing works have proposed diverse control-based and learning-based methods, it remains an open problem to endow robots with animal-like perception and athleticism. Towards this goal, we develop an End-to-End Legged Percepti…

Cited by 0SourceScholar
2024

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

ECCV 2024poster

"This paper presents (), a general-purpose multimodal assistant trained using an end-to-end approach that systematically expands the capabilities of large multimodal models (LMMs). maintains a skill repository that contains a wide range of vision and vision-language pre-trained models (tools), and i…

2024

LMDX: Language Model-based Document Information Extraction and Localization

ACL 2024findings

Large Language Models (LLM) have revolutionized Natural Language Processing (NLP), improving state-of-the-art and exhibiting emergent capabilities across various tasks. However, their application in extracting information from visually rich documents, which is at the core of many document processing…

2024

LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

ICLR 2024spotlight

Studying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications. In this paper, we introduce LMSYS-Chat-1M, a large-scale dataset containing one million real-world conversations with 25 state-of-the-…

2024

Learning Adaptive Kernels for Statistical Independence Tests

AISTATS 2024poster

We propose a novel framework for kernel-based statistical independence tests that enable adaptatively learning parameterized kernels to maximize test power. Our framework can effectively address the pitfall inherent in the existing signal-to-noise ratio criterion by modeling the change of the null d…

2024

Learning Implicit Representation for Reconstructing Articulated Objects

ICLR 2024poster

3D Reconstruction of moving articulated objects without additional information about object structure is a challenging problem. Current methods overcome such challenges by employing category-specific skeletal models. Consequently, they do not generalize well to articulated objects in the wild. We tr…

2024

Learning for Dynamic Subteaming and Voluntary Waiting in Heterogeneous Multi-Robot Collaborative Scheduling

ICRA 2024poster

Coordinating heterogeneous robots is essential for autonomous multi-robot teaming. To execute a set of dependent tasks as quickly as possible, and to complete tasks that cannot be addressed by individual robots, it is necessary to form subteams that can collaboratively finish the tasks. It is also a…

Cited by 1SourceScholar
2024

MC-indexing: Effective Long Document Retrieval via Multi-view Content-aware Indexing

EMNLP 2024finding

Long document question answering (DocQA) aims to answer questions from long documents over 10k words. They usually contain content structures such as sections, sub-sections, and paragraph demarcations. However, the indexing methods of long documents remain under-explored, while existing systems gene…

2024

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

ICML 2024poster

Large Vision-Language Models (LVLMs) show significant strides in general-propose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in t…

Cited by 84SourcePDFScholar
2024

MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs

NeurIPS 2024poster

Large language models (LLMs) have shown increasing capability in problem-solving and decision-making, largely based on the step-by-step chain-of-thought reasoning processes. However, evaluating these reasoning abilities has become increasingly challenging. Existing outcome-based benchmarks are begin…

Cited by 14SourcePDFScholar
2024

MRFS: Mutually Reinforcing Image Fusion and Segmentation

CVPR 2024poster

This paper proposes a coupled learning framework to break the performance bottleneck of infrared-visible image fusion and segmentation called MRFS. By leveraging the intrinsic consistency between vision and semantics it emphasizes mutual reinforcement rather than treating these tasks as separate iss…

2024

MeaCap: Memory-Augmented Zero-shot Image Captioning

CVPR 2024poster

Zero-shot image captioning (IC) without well-paired image-text data can be categorized into two main types: training-free and text-only-training methods. While both types integrate pre-trained vision-language models such as CLIP for image-text similarity evaluation and a pre-trained language model (…

2024

Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length

NeurIPS 2024poster

The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and state space models exist, they empirically underperform Transformers in pretraining efficiency and downstream task accur…

2024

Meta-Adapter for Self-Supervised Speech Models: A Solution to Low-Resource Speech Recognition Challenges

COLING 2024main

Self-supervised models have demonstrated remarkable performance in speech processing by learning latent representations from large amounts of unlabeled data. Although these models yield promising results on low-resource languages, the computational expense of fine-tuning all model parameters is proh…

Cited by 0SourcePDFScholar
2024

Multi-Task Dense Prediction via Mixture of Low-Rank Experts

CVPR 2024poster

Previous multi-task dense prediction methods based on the Mixture of Experts (MoE) have received great performance but they neglect the importance of explicitly modeling the global relations among all tasks. In this paper we present a novel decoder-focused method for multi-task dense prediction call…

2024

MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving

ICML 2024poster

Large language models (LLMs) have demonstrated remarkable performance, and organizations are racing to serve LLMs of varying sizes as endpoints for use-cases like chat, programming and search. However, efficiently serving multiple LLMs poses significant challenges for existing approaches due to vary…

2024

Online Speculative Decoding

ICML 2024poster

Speculative decoding is a pivotal technique to accelerate the inference of large language models (LLMs) by employing a smaller draft model to predict the target model's outputs. However, its efficacy can be limited due to the low predictive accuracy of the draft model, particularly when faced with d…

2024

Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language Models

EMNLP 2024main

Code retrieval aims to identify code from extensive codebases that semantically aligns with a given query code snippet. Collecting a broad and high-quality set of query and code pairs is crucial to the success of this task. However, existing data collection methods struggle to effectively balance sc…

Cited by 5SourcePDFScholar
2024

PVitNet: An Effective Approach for Android Malware Detection Using Pyramid Feature Processing and Vision Transformer

ICASSP 2024accepted

This presents a significant challenge for detecting and combating malicious software. Users often grant software permissions unknowingly, exposing their devices to risks such as unauthorized access, file manipulation, and malware propagation. Traditional detection algorithms relying on limited permi…

Cited by 0SourceScholar
2024

PointTFA: Training-Free Clustering Adaption for Large 3D Point Cloud Models

IJCAI 2024poster

The success of contrastive learning models like CLIP, known for aligning 2D image-text pairs, has inspired the development of triplet alignment for Large 3D Point Cloud Models (3D-PCM). Examples like ULIP integrate images, text, and point clouds into a unified semantic space. However, despite showin…

2024

Quaternion-Based Optimal Interpolation of Similarity Transformations for Multi-Agent Formation

RA-L 2024

This paper addresses the challenge of optimal motion interpolation in multi-agent formation control. The primary goal is to generate trajectories of similarity transformations that minimize various metrics, such as distance traveled, kinetic energy consumption, and overall smoothness. The quaternion

Cited by 2SourceScholar
2024

Revisiting Single Image Reflection Removal In the Wild

CVPR 2024poster

This research focuses on the issue of single-image reflection removal (SIRR) in real-world conditions examining it from two angles: the collection pipeline of real reflection pairs and the perception of real reflection locations. We devise an advanced reflection collection pipeline that is highly ad…

2024

S3O: A Dual-Phase Approach for Reconstructing Dynamic Shape and Skeleton of Articulated Objects from Single Monocular Video

ICML 2024poster

Reconstructing dynamic articulated objects from a singular monocular video is challenging, requiring joint estimation of shape, motion, and camera parameters from limited views. Current methods typically demand extensive computational resources and training time, and require additional human annotat…

2024

SAFNet: Selective Alignment Fusion Network for Efficient HDR Imaging

ECCV 2024poster

"Multi-exposure High Dynamic Range (HDR) imaging is a challenging task when facing truncated texture and complex motion. Existing deep learning-based methods have achieved great success by either following the alignment and fusion pipeline or utilizing attention mechanism. However, the large computa…

2024

Segment and Recognize Anything at Any Granularity

ECCV 2024poster

"In this work, we introduce , an augmented image segmentation foundation for segmenting and recognizing anything at desired granularities. Compared to the foundational segmentation model SAM [?], our model has two unique advantages: (i) granularity-controllability in that the model can produce segme…

Cited by 214SourcePDFScholar
2024

Segmented Safety Docking Control for Mobile Self-Reconfigurable Robots

IROS 2024poster

Mobile self-reconfigurable robots (MSRRs), as a novel multi-robot system with flexible configurations and task adaptability, hold promising applications in unstructured task environments. However, existing autonomous docking strategies are primarily applied in laboratory settings and face numerous c…

Cited by 0SourceScholar
2024

SiCP: Simultaneous Individual and Cooperative Perception for 3D Object Detection in Connected and Automated Vehicles

IROS 2024poster

Cooperative perception for connected and automated vehicles is traditionally achieved through the fusion of feature maps from two or more vehicles. However, the absence of feature maps shared from other vehicles can lead to a significant decline in 3D object detection performance for cooperative per…

Cited by 6SourcecodeScholar