← Search

Yong Liu

313 accepted papers

2026

AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection

AAAI 2026technical

Universal visual anomaly detection aims to identify anomalies from novel or unseen vision domains without additional fine-tuning, which is critical in open scenarios. Recent studies have demonstrated that pre-trained vision-language models like CLIP exhibit strong generalization with just zero or a

Cited by 0SourcePDFScholar
2026

CAMEL: Confidence-Gated Reflection for Reward Modeling

ICML 2026poster

Reward models play a fundamental role in aligning large language models with human preferences. Existing methods predominantly follow two paradigms: scalar discriminative preference models, which are efficient but lack interpretability, and generative judging models, which offer richer reasoning at …

Cited by 0SourceScholar
2026

Can Recommender Systems Teach Themselves? A Recursive Self-Improving Framework with Fidelity Control

ICML 2026poster

The scarcity of high-quality training data presents a fundamental bottleneck to scaling machine learning models. This challenge is particularly acute in recommendation systems, where extreme sparsity in user interactions leads to rugged optimization landscapes and poor generalization. We propose the…

Cited by 0SourceScholar
2026

ClimaOoD: Improving Anomaly Segmentation via Physically Realistic Synthetic Data

CVPR 2026

Anomaly segmentation seeks to detect and localize unknown or out-of-distribution (OoD) objects that fall outside predefined semantic classes--a capability essential for safe autonomous driving. However, the scarcity and limited diversity of anomaly data severely constrain model generalization in ope

Cited by 0SourceScholar
2026

Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning

ICLR 2026poster

Chain-of-Thought (CoT) training has markedly advanced the reasoning capabilities of large language models (LLMs), yet the mechanisms by which CoT training enhances generalization remain inadequately understood. In this work, we demonstrate that compositional generalization is fundamental: models sys…

Cited by 0SourcecodeScholar
2026

Don’t Start Over: A Cost-Effective Framework for Migrating Personalized Prompts Between LLMs

AAAI 2026technical

Personalization in Large Language Models (LLMs) often relies on user-specific soft prompts. However, these prompts become obsolete when the foundation model is upgraded, necessitating costly, full-scale retraining. To overcome this limitation, we propose the Prompt-level User Migration Adapter (PUMA

Cited by 0SourcePDFScholar
2026

DynBridge: Bridging Imagination and Control through Interaction Dynamics for Robot Manipulation

CVPR 2026

Recent generative models allow robots to generate future visual outcomes for action guidance, yet most still address imagination and control independently, resulting in visually coherent rollouts but physically inconsistent behaviors. While structural priors enhance spatial grounding, these methods

Cited by 0SourceScholar
2026

Evoking User Memory: Personalizing LLM via Recollection-Familiarity Adaptive Retrieval

ICLR 2026poster

Personalized large language models (LLMs) rely on memory retrieval to incorporate user-specific histories, preferences, and contexts. Existing approaches either overload the LLM by feeding all the user's past memory into the prompt, which is costly and unscalable, or simplify retrieval into a one-sh…

Cited by 0SourcecodeScholar
2026

FINMCP-BENCH: BENCHMARKING LLM AGENTS FOR REAL-WORLD FINANCIAL TOOL USE UNDER THE MODEL CONTEXT PROTOCOL

ICASSP 2026poster

This paper introduces \textbf{FinMCP-Bench}, a novel benchmark for evaluating large language models (LLMs) in solving real-world financial problems through tool invocation of financial model context protocols. FinMCP-Bench contains 613 samples spanning 10 main scenarios and 33 sub-scenarios, featuri…

Cited by 0SourcePDFScholar
2026

FOCUS: Efficient Keyframe Selection for Long Video Understanding

ICLR 2026poster

Multimodal large language models (MLLMs) represent images and video frames as visual tokens. Scaling from single images to hour-long videos, however, inflates the token budget far beyond practical limits. Popular pipelines therefore either uniformly subsample or apply keyframe selection with retriev…

Cited by 0SourcecodeScholar
2026

High Probability Bounds for Non-Convex Stochastic Optimization with Momentum

ICLR 2026poster

Stochastic gradient descent with momentum (SGDM) is widely used in machine learning, yet high-probability learning bounds for SGDM in non-convex settings remain scarce. In this paper, we provide high-probability convergence bounds and generalization bounds for SGDM. First, we establish such bounds f…

Cited by 0SourceScholar
2026

IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment

ICLR 2026poster

Instruction-guided video editing has emerged as a rapidly advancing research direction, offering new opportunities for intuitive content transformation while also posing significant challenges for systematic evaluation. Existing video editing benchmarks fail to support the evaluation of instruction-…

Cited by 0SourcecodeScholar
2026

It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks

ICML 2026poster

Time series foundation models (TSFMs) are revolutionizing the forecasting landscape from specific dataset modeling to generalizable task evaluation. However, we contend that existing benchmarks exhibit common limitations in four dimensions: constrained data composition dominated by reused legacy sou…

Cited by 0SourceScholar
2026

LLM-Oriented Token-Adaptive Knowledge Distillation

AAAI 2026technical

Knowledge Distillation (KD) is a key technique for compressing Large-scale Language Models (LLMs), but prevailing logit-based methods employ static strategies misaligned with the student’s dynamic learning process. By treating all tokens indiscriminately with a fixed temperature, these methods resul

Cited by 0SourcePDFScholar
2026

Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR

ICLR 2026poster

Large Language Models (LLMs) have recently improved mathematical reasoning through Reinforcement Learning with Verifiable Reward (RLVR). However, existing RLVR algorithms require large query budgets, making annotation costly. We investigate whether fewer but more informative queries can yield simila…

Cited by 0SourcecodeScholar
2026

LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation

ICLR 2026poster

Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and background elements. However, precise face–attribute alignment across subjects remains challenging, as existing methods…

Cited by 0SourcecodeScholar
2026

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

ICLR 2026poster

Chain-of-Thought (CoT) reasoning has proven effective in enhancing large language models by encouraging step-by-step intermediate reasoning, and recent advances have extended this paradigm to Multimodal Large Language Models (MLLMs). In the medical domain, where diagnostic decisions depend on nuance…

Cited by 0SourceScholar
2026

Note2Chat: Improving LLMs for Multi-Turn Clinical History Taking Using Medical Notes

AAAI 2026technical

Effective clinical history taking is a foundational yet underexplored component of clinical reasoning. While large language models (LLMs) have shown promise on static benchmarks, they often fall short in dynamic, multi-turn diagnostic settings that require iterative questioning and hypothesis refine

Cited by 0SourcePDFScholar
2026

OptMark: Robust Multi-bit Diffusion Watermarking via Inference Time Optimization

AAAI 2026technical

Watermarking diffusion-generated images is crucial for copyright protection and user tracking. However, current diffusion watermarking methods face significant limitations: zero-bit watermarking systems lack the capacity for large-scale user tracking, while multi-bit methods are highly sensitive to

Cited by 0SourcePDFScholar
2026

PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training

CVPR 2026

Open-Set Object Detection (OSOD) enables recognition of novel categories beyond fixed classes but faces challenges in aligning text representations with complex visual concepts and the scarcity of image-text pairs for rare categories. This results in suboptimal performance in specialized domains or

Cited by 0SourcecodeScholar
2026

PHPFND: Detecting Fake News via Post-Hoc Processing of LLMs Hallucination

AAAI 2026technical

Large Language Models (LLMs) perform excellently in fake news detection tasks, but their outputs are often accompanied by hallucinations, i.e., generated content that is contradictory to facts. Previous studies have mostly mitigated hallucinations through prompt design. However, this paper reveals t

Cited by 0SourcePDFScholar
2026

Personalize Before Retrieve: LLM-based Personalized Query Expansion for User-Centric Retrieval

AAAI 2026technical

Retrieval-Augmented Generation (RAG) critically depends on effective query expansion to retrieve relevant information. However, existing expansion methods adopt uniform strategies that overlook user-specific semantics, ignoring individual expression styles, preferences, and historical context. In pr

Cited by 0SourcePDFScholar
2026

Put the Space of LoRA Initialization to the Extreme to Preserve Pre-trained Knowledge

AAAI 2026technical

Low-Rank Adaptation (LoRA) is the leading parameter-efficient fine-tuning method for Large Language Models (LLMs), but it still suffers from catastrophic forgetting. Recent work has shown that specialized LoRA initialization can alleviate catastrophic forgetting. There are currently two approaches t

Cited by 0SourcePDFScholar
2026

SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion Training

CVPR 2026

Denoising-based diffusion transformers, despite their strong generation performance, suffer from inefficient training convergence. Existing methods addressing this issue, such as REPA (relying on external representation encoders) or SRA (requiring dual-model setups), inevitably incur heavy computati

Cited by 0SourceScholar
2026

Sketch-Based Low-Rank Model Merging with Shared Circulant Transforms

ICML 2026poster

Merging multiple low-rank adapters (LoRA) provides a practical route to scaling multi-task learning and deployment more efficiently than full-model weight merging, while avoiding reliance on task-specific training data. However, most existing approaches either treat LoRA updates as dense weight delt…

Cited by 0SourceScholar
2026

SwarmNav: Swarm Robotics Navigation in Dynamic and Dense Environments Via Reinforcement Learning

ICRA 2026poster

Collision avoidance and navigation in dynamic and dense environments remain highly challenging for swarm robotics. To address this, we propose SwarmNav, a novel goal-region amplification navigation policy that leverages LiDAR-based position data to generate velocity commands guiding robots toward th…

Cited by 0Scholar
2026

Transform Trained Transformer for Accelerating Native 4K Video Generation

ICML 2026poster

Native 4K (2176$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer…

Cited by 0SourceScholar
2026

Vision-Centric 4D Occupancy Forecasting and Planning Via Implicit Residual World Models

ICRA 2026poster

End-to-end autonomous driving systems increasingly rely on vision-centric world models to understand and predict their environment. However, a common ineffectiveness in these models is the full reconstruction of future scenes, which expends significant capacity on redundantly modeling static backgro…

2025

ADePT: Adaptive Decomposed Prompt Tuning for Parameter-Efficient Fine-tuning

ICLR 2025poster

Prompt Tuning (PT) enables the adaptation of Pre-trained Large Language Models (PLMs) to downstream tasks by optimizing a small amount of soft virtual tokens, which are prepended to the input token embeddings. Recently, Decomposed Prompt Tuning (DePT) has demonstrated superior adaptation capabilitie…

2025

Action Detail Matters: Refining Video Recognition with Local Action Queries

CVPR 2025poster

Video action recognition involves interpreting both global context and specific details to accurately identify actions. While previous models are effective at capturing spatiotemporal features, they often lack a focused representation of key action details. To address this, we introduce \nameo, a fr…

Cited by 0SourcePDFScholar
2025

AdaO2B: Adaptive Online to Batch Conversion for Out-of-Distribution Generalization

AAAI 2025technical

Online to batch conversion involves constructing a new batch learner by utilizing a series of models generated by an existing online learning algorithm, for achieving generalization guarantees under i.i.d assumption. However, when applied to real-world streaming applications such as streaming recomm…

Cited by 0SourcePDFScholar
2025

AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated excellent performance in video understanding but suffer from degraded effectiveness when processing long videos due to fixed-length contexts and weaknesses in modeling long-term dependencies. Retrieval-Augmented Generation (RAG) technology c…

Cited by 0SourcecodeScholar
2025

Adaptive Schema-aware Event Extraction with Retrieval-Augmented Generation

EMNLP 2025

Event extraction (EE) is a fundamental task in natural language processing (NLP) that involves identifying and extracting event information from unstructured text. Effective EE in real-world scenarios requires two key steps: selecting appropriate schemas from hundreds of candidates and executing the

2025

Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger

ACL 2025long

Large language models (LLMs) have shown remarkable emergent capabilities, transforming the execution of functional tasks by leveraging external tools for complex problems that require specialized processing or up-to-date data. While existing research expands LLMs access to diverse tools (e.g., progr…

Cited by 0SourcePDFScholar
2025

Benchmarking Retrieval-Augmented Multimomal Generation for Document Question Answering

NeurIPS 2025poster

Document Visual Question Answering (DocVQA) faces dual challenges in processing lengthy multimodal documents (text, images, tables) and performing cross-modal reasoning. Current document retrieval-augmented generation (DocRAG) methods remain limited by their text-centric approaches, frequently missi…

Cited by 0SourcecodeScholar
2025

CFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion

ICASSP 2025accepted

Video summarization, by selecting the most informative and/or user-relevant parts of original videos to create concise summary videos, has high research value and consumer demand in today’s video proliferation era. Multi-modal video summarization that accomodates user input has become a research hot…

Cited by 0SourceScholar
2025

Can LLMs Outshine Conventional Recommenders? A Comparative Evaluation

NeurIPS 2025poster

Integrating large language models (LLMs) into recommender systems has created new opportunities for improving recommendation quality. However, a comprehensive benchmark is needed to thoroughly evaluate and compare the recommendation capabilities of LLMs with traditional recommender systems. In this…

Cited by 0SourcecodeScholar
2025

CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation

ICCV 2025poster

The newly proposed Generalized Referring Expression Segmentation (GRES) amplifies the formulation of classic RES by involving complex multiple/non-target scenarios. Recent approaches address GRES by directly extending the well-adopted RES frameworks with object-existence identification. However, the…

Cited by 0SourcePDFScholar
2025

Contrastive Pre-Training and Post-Tuning for Heterogeneous Graph Learning

ICASSP 2025accepted

In recent years, the field of heterogeneous graph learning has garnered significant interest. Various efforts have been made towards learning heterogeneous graph representations, such as designing meta-paths to mine implicit graph knowledge or directly applying Graph Neural Networks (GNNs) for graph…

Cited by 0SourceScholar
2025

CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Control

ACL 2025finding

Retrieval-augmented generation (RAG) has emerged as a promising solution for mitigating hallucinations of large language models (LLMs) with retrieved external knowledge. Adaptive RAG enhances this approach by enabling dynamic retrieval during generation, activating retrieval only when the query exce…

2025

Decentralized Federated Learning with Model Caching on Mobile Agents

AAAI 2025technical

Federated Learning (FL) trains a shared model using data and computation power on distributed agents coordinated by a central server. Decentralized FL (DFL) utilizes local model exchange and aggregation between agents to reduce the communication and computation overheads on the central server. Howev…

2025

Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning

NeurIPS 2025poster

Large reasoning models (LRMs) have demonstrated impressive capabilities in complex problem-solving, yet their internal reasoning mechanisms remain poorly understood. In this paper, we investigate the reasoning trajectories of LRMs from an information-theoretic perspective. By tracking how mutual in…

Cited by 0SourcecodeScholar
2025

Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation

AAAI 2025technical

In the domain of computer vision, Parameter-Efficient Tuning (PET) is increasingly replacing the traditional paradigm of pre-training followed by full fine-tuning. PET is particularly favored for its effectiveness in large foundation models, as it streamlines transfer learning costs and optimizes ha…

2025

Do not Abstain! Identify and Solve the Uncertainty

ACL 2025long

Despite the widespread application of Large Language Models (LLMs) across various domains, they frequently exhibit overconfidence when encountering uncertain scenarios, yet existing solutions primarily rely on evasive responses (e.g., “I don’t know”) overlooks the opportunity of identifying and addr…

Cited by 0SourcePDFScholar
2025

DreamLight: Towards Harmonious and Consistent Image Relighting

NeurIPS 2025poster

We introduce a model named DreamLight for universal image relighting in this work, which can seamlessly composite subjects into a new background while maintaining aesthetic uniformity in terms of lighting and color tone. The background can be specified by natural images (image-based relighting) or g…

Cited by 0SourceScholar
2025

DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving

ICCV 2025poster

This paper introduces DriveArena, the first high-fidelity closed-loop simulation system designed for driving agents navigating real-world scenarios. DriveArena comprises two core components: Traffic Manager, a traffic simulator capable of generating realistic traffic flow on any global street map, a…

Cited by 0SourcePDFScholar
2025

Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving

AAAI 2025technical

World models envision potential future states based on various ego actions. They embed extensive knowledge about the driving environment, facilitating safe and scalable autonomous driving. Most existing methods primarily focus on either data generation or the pretraining paradigms of world models. U…

Cited by 4SourcePDFScholar
2025

DynaMind: Reasoning over Abstract Video Dynamics for Embodied Decision-Making

ICML 2025poster

Integrating natural language instructions and visual perception with decision-making is a critical challenge for embodied agents. Existing methods often struggle to balance the conciseness of language commands with the richness of video content. To bridge the gap between modalities, we propose extra…

Cited by 0SourcePDFScholar
2025

Efficient Learning of A Unified Policy For Whole-body Manipulation and Locomotion Skills

IROS 2025

Equipping quadruped robots with manipulators provides unique loco-manipulation capabilities, enabling diverse practical applications. This integration creates a more complex system that has increased difficulties in modeling and control. Reinforcement learning (RL) offers a promising solution to add

Cited by 3SourceScholar
2025

Efficient Multi-Robot Task and Path Planning in Large-Scale Cluttered Environments

RA-L 2025

As the potential of multi-robot systems continues to be explored and validated across various real-world applications, such as package delivery, search and rescue, and autonomous exploration, the need to improve the efficiency and quality of task and path planning has become increasingly urgent, par

Cited by 4SourceScholar
2025

FLARE: Fast Large-Scale Autonomous Exploration Guided by Unknown Regions

RA-L 2025

Autonomous exploration is a critical foundation for unmanned aerial vehicle (UAV) applications such as search and rescue. However, existing methods typically focus only on known spaces or frontiers without considering unknown regions or providing further guidance for the global path, which results i

Cited by 2SourceScholar
2025

Federated Residual Low-Rank Adaptation of Large Language Models

ICLR 2025poster

Low-Rank Adaptation (LoRA) presents an effective solution for federated fine-tuning of Large Language Models (LLMs), as it substantially reduces communication overhead. However, a straightforward combination of FedAvg and LoRA results in suboptimal performance, especially under data heterogeneity. W…

Cited by 0SourcePDFScholar
2025

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

ICCV 2025poster

Benefiting from the advances in large language models and cross-modal alignment, existing multimodal large language models have achieved prominent performance in image and short video understanding. However, the understanding of long videos is still challenging, as their long-context nature results…

2025

FreEformer: Frequency Enhanced Transformer for Multivariate Time Series Forecasting

IJCAI 2025

This paper presents FreEformer, a simple yet effective model that leverages a Frequency Enhanced Transformer for multivariate time series forecasting. Our work is based on the assumption that the frequency spectrum provides a global perspective on the composition of series across various frequencies

2025

Gaussian-LIC: Real-Time Photo-Realistic SLAM with Gaussian Splatting and LiDAR-Inertial-Camera Fusion

ICRA 2025

In this paper, we present a real-time photo-realistic SLAM method based on marrying Gaussian Splatting with LiDAR-Inertial-Camera SLAM. Most existing radiance-field-based SLAM systems mainly focus on bounded indoor environments, equipped with RGB-D or RGB sensors. However, they are prone to decline

Cited by 28SourcecodeScholar
2025

GroundingFace: Fine-grained Face Understanding via Pixel Grounding Multimodal Large Language Model

CVPR 2025highlight

Multimodal Language Learning Models (MLLMs) have shown remarkable performance in image understanding, generation, and editing, with recent advancements achieving pixel-level grounding with reasoning. However, these models for common objects struggle with fine-grained face understanding. In this work…

Cited by 0SourcePDFScholar
2025

Hash-GS: Anchor-Based 3D Gaussian Splatting with Multi-Resolution Hash Encoding for Efficient Scene Reconstruction

ICRA 2025

Realistic 3D object and scene reconstruction is pivotal in advancing fields such as world model simulation and embodied intelligence. In this paper, we introduce Hash-GS, a storage-efficient method for large-scale scene reconstruction using anchor-based 3D Gaussian Splatting (3DGS). The vanilla 3DGS

Cited by 1SourceScholar
2025

HyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual Perceiver

CVPR 2025poster

This paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs). Despite significant progress in current unified segmentation methods, limitations in adaptation to both image and video scenarios, as…

2025

InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models

ICCV 2025poster

Boosted by Multi-modal Large Language Models (MLLMs), text-guided universal segmentation models for the image and video domains have made rapid progress recently. However, these methods are often developed separately for specific domains, overlooking the similarities in task settings and solutions a…

2025

IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis

AAAI 2025technical

Zero-shot Referring Image Segmentation (RIS) identifies the instance mask that best aligns with a specified referring expression without training and fine-tuning, significantly reducing the labor-intensive annotation process. Despite achieving commendable results, previous CLIP-based models have a c…

2025

KGN-Pro: Keypoint-Based Grasp Prediction through Probabilistic 2D-3D Correspondence Learning

IROS 2025

High-level robotic manipulation tasks demand flexible 6-DoF grasp estimation to serve as a basic function. Previous approaches either directly generate grasps from point-cloud data, suffering from challenges with small objects and sensor noise, or infer 3D information from RGB images, which introduc

Cited by 0SourcecodeScholar
2025

L2COcc: Lightweight Camera-Centric Semantic Scene Completion via Distillation of LiDAR Model

IROS 2025

Semantic Scene Completion (SSC) constitutes a pivotal element in autonomous driving perception systems, tasked with inferring the 3D semantic occupancy of a scene from sensory data. To improve accuracy, prior research has implemented various computationally demanding and memory-intensive 3D operatio

Cited by 3SourcecodeScholar
2025

L2Calib: SE (3)-Manifold Reinforcement Learning for Robust Extrinsic Calibration with Degenerate Motion Resilience

IROS 2025

Extrinsic calibration is essential for multi-sensor fusion, existing methods rely on structured targets or fully-excited data, limiting real-world applicability. Online calibration further suffers from weak excitation, leading to unreliable estimates. To address these limitations, we propose a reinf

Cited by 0SourcecodeScholar
2025

LITE: A Learning-Integrated Topological Explorer for Multi-Floor Indoor Environments

IROS 2025

This work focuses on multi-floor indoor exploration, which remains an open area of research. Compared to traditional methods, recent learning-based explorers have demonstrated significant potential due to their robust environmental learning and modeling capabilities, but most are restricted to 2D en

Cited by 0SourceScholar
2025

LLaVA-KD: A Framework of Distilling Multimodal Large Language Models

ICCV 2025poster

The success of Large Language Models (LLMs) has inspired the development of Multimodal Large Language Models (MLLMs) for unified understanding of vision and language. However, the increasing model size and computational complexity of large-scale MLLMs (l-MLLMs) limit their use in resource-constraine…

2025

Learning Symmetric Legged Locomotion via State Distribution Symmetrization

IROS 2025

Morphological symmetry is a fundamental characteristic of legged animals and robots. Most existing Deep Reinforcement Learning approaches for legged locomotion neglect to exploit this inherent symmetry, often producing unnatural and suboptimal behaviors such as dominant legs or non-periodic gaits. T

Cited by 0SourceScholar
2025

LiCROcc: Teach Radar for Accurate Semantic Occupancy Prediction Using LiDAR and Camera

RA-L 2025

Semantic Scene Completion (SSC) is pivotal in autonomous driving perception, frequently confronted with the complexities of weather and illumination changes. The long-term strategy involves fusing multi-modal information to bolster the system's robustness. Radar, increasingly utilized for 3D target

Cited by 18SourceScholar
2025

LiT: Delving into a Simple Linear Diffusion Transformer for Image Generation

ICCV 2025poster

In this paper, we investigate how to convert a pre-trained Diffusion Transformer (DiT) into a linear DiT, as its simplicity, parallelism, and efficiency for image generation. Through detailed exploration, we offer a suite of ready-to-use solutions, ranging from linear attention design to optimizatio…

Cited by 0SourcePDFScholar
2025

Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter Tuning

AAAI 2025technical

Personalized federated learning (PFL) studies effective model personalization to address the data heterogeneity issue among clients in traditional federated learning (FL). Existing PFL approaches mainly generate personalized models by relying solely on the clients' latest updated models while ignori…

Cited by 0SourcePDFScholar
2025

MARF: Cooperative Multi-Agent Path Finding with Reinforcement Learning and Frenet Lattice in Dynamic Environments

ICRA 2025

Multi-agent path finding (MAPF) in dynamic and complex environments is a highly challenging task. Recent research has focused on the scalability of agent numbers or the complexity of the environment. Usually, they disregard the agents' physical constraints or use a differential-driven model. However

Cited by 1SourceScholar
2025

MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training

ICML 2025poster

Large-batch training has become a cornerstone in accelerating the training of deep neural networks, yet it poses challenges in optimization and generalization. Existing optimizers like AdamW present performance degradation during language models' large-batch training, due to the information bottlen…

2025

MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection

ICLR 2025poster

In the field of industrial inspection, Multimodal Large Language Models (MLLMs) have a high potential to renew the paradigms in practical applications due to their robust language capabilities and generalization abilities. However, despite their impressive problem-solving skills in many domains, MLL…

2025

MMDocIR: Benchmarking Multimodal Retrieval for Long Documents

EMNLP 2025

Multimodal document retrieval aims to identify and retrieve various forms of multimodal content, such as figures, tables, charts, and layout information from extensive documents. Despite its increasing popularity, there is a notable lack of a comprehensive and robust benchmark to effectively evaluat

Cited by 0SourcePDFScholar
2025

OARecon: Object-Aware Viewpoint Augmentation for Indoor Compositional Reconstruction

ICASSP 2025accepted

Real-world scenes likely involve repetitive objects indicating that the reconstruction of the target object can be supplemented by the views of other identical objects. However, traditional 3D reconstruction methods do not take this a priori knowledge into account and fail to make full use of the av…

Cited by 0SourceScholar
2025

OLinear: A Linear Model for Time Series Forecasting in Orthogonally Transformed Domain

NeurIPS 2025poster

This paper presents $\mathbf{OLinear}$, a $\mathbf{linear}$-based multivariate time series forecasting model that operates in an $\mathbf{o}$rthogonally transformed domain. Recent forecasting models typically adopt the temporal forecast (TF) paradigm, which directly encode and decode time series in…

Cited by 0SourcecodeScholar
2025

On the Importance of Language-driven Representation Learning for Heterogeneous Federated Learning

ICLR 2025poster

Non-Independent and Identically Distributed (Non-IID) training data significantly challenge federated learning (FL), impairing the performance of the global model in distributed frameworks. Inspired by the superior performance and generalizability of language-driven representation learning in centra…

Cited by 0SourcePDFScholar
2025

P-Law: Predicting Quantitative Scaling Law with Entropy Guidance in Large Recommendation Models

NeurIPS 2025poster

With the growing size of data and models in Large Recommendation Models, the time required for debugging has become increasingly prohibitive, underscoring the urgent need for effective guidance in parameter configuration. The Scaling Law (SL) offers analogous guidance in the Sequential Language doma…

Cited by 0SourcecodeScholar
2025

PatchScaler: An Efficient Patch-Independent Diffusion Model for Image Super-Resolution

ICCV 2025poster

While diffusion models significantly improve the perceptual quality of super-resolved images, they usually require a large number of sampling steps, resulting in high computational costs and long inference times. Recent efforts have explored reasonable acceleration schemes by reducing the number of…

2025

Planning with Multi-Constraints via Collaborative Language Agents

COLING 2025main

The rapid advancement of neural language models has sparked a new surge of intelligent agent research. Unlike traditional agents, large language model-based agents (LLM agents) have emerged as a promising paradigm for achieving artificial general intelligence (AGI) due to their superior reasoning an…

2025

RAPID: Efficient Retrieval-Augmented Long Text Generation with Writing Planning and Information Discovery

ACL 2025finding

Generating knowledge-intensive and comprehensive long texts, such as encyclopedia articles, remains significant challenges for Large Language Models. It requires not only the precise integration of facts but also the maintenance of thematic coherence throughout the article. Existing methods, such as…

2025

REEF: Representation Encoding Fingerprints for Large Language Models

ICLR 2025oral

Protecting the intellectual property of open-source Large Language Models (LLMs) is very important, because training LLMs costs extensive computational resources and data. Therefore, model owners and third parties need to identify whether a suspect model is a subsequent development of the victim mod…

2025

Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning

ICML 2025poster

Test-time scaling, which is also often referred to as *slow-thinking*, has been demonstrated to enhance multi-step reasoning in large language models (LLMs). However, despite its widespread utilization, the mechanisms underlying slow-thinking methods remain poorly understood. This paper explores the…

2025

Revisiting Chain-of-Thought Prompting: Zero-shot Can Be Stronger than Few-shot

EMNLP 2025

In-Context Learning (ICL) is an essential emergent ability of Large Language Models (LLMs), and recent studies introduce CoT to exemplars of ICL to enhance the reasoning capability, especially in mathematics tasks. However, given the continuous advancement of model capabilities, it remains unclear w

Cited by 0SourcePDFScholar
2025

Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL

ACL 2025finding

As large language models advance toward superhuman performance, ensuring their alignment with human values and abilities grows increasingly complex. Weak-to-strong generalization offers a promising approach by leveraging predictions from weaker models to guide stronger systems, but its effectiveness…

Cited by 0SourcePDFScholar
2025

Reward Mixology: Crafting Hybrid Signals for Reinforcement Learning Driven In-Context Learning

EMNLP 2025

In-context learning (ICL) performance heavily relies on the quality and ordering of demonstrations. Iterative selection (IS) is a promising approach to address this issue, but existing IS methods face two key challenges: the oversimplification of process reward signals that guide intermediate steps

Cited by 0SourcePDFScholar
2025

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes

CVPR 2025poster

Reference Audio-Visual Segmentation (Ref-AVS) aims to provide a pixel-wise scene understanding in Language-aided Audio-Visual Scenes (LAVS). This task requires the model to continuously segment objects referred to by text and audio from a video. Previous dual-modality methods always fail due to the…

Cited by 0SourcePDFScholar
2025

SARO: Space-Aware Robot System for Terrain Crossing via Vision-Language Model

ICRA 2025

The application of vision-language models (VLMs) has achieved impressive success in various robotics tasks. However, there are few explorations for foundation models used in quadruped robot navigation through terrains in 3D environments. We introduce SARO (Space-Aware Robot System for Terrain Crossi

Cited by 5SourcecodeScholar
2025

SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin

EMNLP 2025

Enhancing the numerical and logical reasoning capabilities of Large Language Models (LLMs) has become a prominent research focus. Existing approaches exhibit notable limitations: inference-phase techniques, such as Chain of Thought, depend on prompt engineering and pretrained knowledge; sentence-lev

Cited by 0SourcePDFScholar
2025

SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed Graphs

NeurIPS 2025poster

Large-scale pre-trained models have revolutionized Natural Language Processing (NLP) and Computer Vision (CV), showcasing remarkable cross-domain generalization abilities. However, in graph learning, models are typically trained on individual graph datasets, limiting their capacity to transfer knowl…

Cited by 0SourceScholar
2025

SeedLoRA: A Fusion Approach to Efficient LLM Fine-Tuning

ICML 2025poster

Despite Low-Rank Adaptation (LoRA)'s popularity for fine-tuning large models, it often exhibits a noticeable performance gap compared to full fine-tuning, particularly in complex tasks such as mathematical reasoning and code generation. Motivated by this discrepancy, we propose a novel fusion approa…

Cited by 0SourcePDFScholar
2025

Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning

NeurIPS 2025poster

While fine-tuning large language models (LLMs) for specific tasks often yields impressive results, it comes at the cost of memory inefficiency due to back-propagation in gradient-based training. Memory-efficient Zeroth-order (MeZO) optimizers, recently proposed to address this issue, only require fo…

Cited by 0SourceScholar
2025

Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing Topology

AAAI 2025technical

Zeroth-order (ZO) optimization as the gradient-free method has become a powerful tool when the first-order gradient is unavailable or expensive to obtain, especially in decentralized learning scenarios where data and computational resources are distributed across multiple clients. There have been ma…

Cited by 0SourcePDFScholar
2025

Stability and Sharper Risk Bounds with Convergence Rate $\tilde{O}(1/n^2)$

NeurIPS 2025poster

Prior work (Klochkov \& Zhivotovskiy, 2021) establishes at most $O\left(\log (n)/n\right)$ excess risk bounds via algorithmic stability for strongly-convex learners with high probability. We show that under the similar common assumptions — Polyak-Lojasiewicz condition, smoothness, and Lipschitz cont…

Cited by 0SourceScholar
2025

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation

ICCV 2025poster

Open-vocabulary segmentation aims to achieve segmentation of arbitrary categories given unlimited text inputs as guidance. To achieve this, recent works have focused on developing various technical routes to exploit the potential of large-scale pre-trained vision-language models and have made signif…

Cited by 0SourcePDFScholar
2025

Sundial: A Family of Highly Capable Time Series Foundation Models

ICML 2025oral

We introduce Sundial, a family of native, flexible, and scalable time series foundation models. To predict the next-patch's distribution, we propose a TimeFlow Loss based on flow-matching, which facilitates native pre-training of Transformers on continuous-valued time series without discrete tokeniz…

2025

Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization

ICLR 2025poster

Superalignment, where humans act as weak supervisors for superhuman models, has become a crucial problem with the rapid development of Large Language Models (LLMs). Recent work has preliminarily studied this problem by using weak models to supervise strong models, and discovered that weakly supervis…

2025

TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion Generation

CVPR 2025poster

Human-human motion generation is essential for understanding humans as social beings. Current methods fall into two main categories: single-person-based methods and separate modeling-based methods. To delve into this field, we abstract the overall generation process into a general framework MetaMoti…

Cited by 0SourcePDFScholar
2025

The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models

ACL 2025long

Ensuring awareness of fairness and privacy in Large Language Models (LLMs) is critical. Interestingly, we discover a counter-intuitive trade-off phenomenon that enhancing an LLM’s privacy awareness through Supervised Fine-Tuning (SFT) methods significantly decreases its fairness awareness with thous…

2025

Theoretical Insights into Fine-Tuning Attention Mechanism: Generalization and Optimization

IJCAI 2025

Large Language Models (LLMs), built on Transformer architectures, exhibit remarkable generalization across a wide range of tasks. However, fine-tuning these models for specific tasks remains resource-intensive due to their extensive parameterization. In this paper, we explore two remarkable phenomen

Cited by 0SourcePDFScholar
2025

Timer-XL: Long-Context Transformers for Unified Time Series Forecasting

ICLR 2025poster

We present Timer-XL, a causal Transformer for unified time series forecasting. To uniformly predict multidimensional time series, we generalize next token prediction, predominantly adopted for 1D token sequences, to multivariate next token prediction. The paradigm formulates various forecasting task…

2025

ToolACE: Winning the Points of LLM Function Calling

ICLR 2025poster

Function calling significantly extends the application boundary of large language models (LLMs), where high-quality and diverse training data is critical for unlocking this capability. However, collecting and annotating real function-calling data is challenging, while synthetic data from existing pi…

Cited by 23SourcePDFScholar
2025

Towards Auto-Regressive Next-Token Prediction: In-context Learning Emerges from Generalization

ICLR 2025poster

Large language models (LLMs) have demonstrated remarkable in-context learning (ICL) abilities. However, existing theoretical analysis of ICL primarily exhibits two limitations: \textbf{(a) Limited \textit{i.i.d.} Setting.} Most studies focus on supervised function learning tasks where prompts are co…

Cited by 0SourcePDFScholar
2025

Towards Reward Fairness in RLHF: From a Resource Allocation Perspective

ACL 2025long

Rewards serve as proxies for human preferences and play a crucial role in Reinforcement Learning from Human Feedback (RLHF). However, if these rewards are inherently imperfect, exhibiting various biases, they can adversely affect the alignment of large language models (LLMs). In this paper, we colle…

2025

Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective

ICLR 2025poster

Synthetic data has become a pivotal resource in post-training tasks for large language models (LLMs) due to the scarcity of high-quality, specific data. While various methods have been developed to generate synthetic data, there remains a discernible gap between the practical effects of synthetic da…

2025

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

NeurIPS 2025poster

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. % The growing demand for video applications sets higher requirements for high-quality video generation models. % For example, the generation of m…

Cited by 0SourcecodeScholar
2025

VPCI: Self-Supervised Visual Prompt-Guided Cross-Domain Interactive Image Fusion Framework

ICASSP 2025accepted

Image fusion combines information from multi-modality images to produce high-quality fused images with enhanced clarity, contrast, and informativeness. However, limited ground truth fusion data lead to difficulties in effectively training these fusion models. Moreover, current studies lack of fine-g…

Cited by 0SourceScholar
2025

VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering

AAAI 2025technical

Albeit progress has been made in Composed Image Retrieval (CIR), we empirically find that a certain percentage of failure retrieval results are not consistent with their relative captions. To address this issue, this work provides a Visual Question Answering (VQA) perspective to boost the performanc…

2025

X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability

NeurIPS 2025poster

Diffusion models are advancing autonomous driving by enabling realistic data synthesis, predictive end-to-end planning, and closed-loop simulation, with a primary focus on temporally consistent generation. However, large-scale 3D scene generation requiring spatial coherence remains underexplored. In…

Cited by 0SourceScholar
2024

A Coarse-to-Fine Place Recognition Approach using Attention-guided Descriptors and Overlap Estimation

ICRA 2024poster

Place recognition is a challenging but crucial task in robotics. Current description-based methods may be limited by representation capabilities, while pairwise similarity-based methods require exhaustive searches, which is time-consuming. In this paper, we present a novel coarse-to-fine approach to…

Cited by 1SourcecodeScholar
2024

A Distributed Pipeline for Collaborative Pursuit in the Target Guarding Problem

RA-L 2024

The target guarding problem (TGP) is a classical combat game where pursuers aim to capture evaders to protect a territory from intrusion. This paper proposes a distributed pipeline for multi-pursuer multi-evader TGP with the capability to accommodate varying numbers of evaders and criteria for succe

Cited by 8SourceScholar
2024

A Multimodal, Multi-Task Adapting Framework for Video Action Recognition

AAAI 2024technical

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance a…

Cited by 17SourcePDFScholar
2024

A Robotic-centric Paradigm for 3D Human Tracking Under Complex Environments Using Multi-modal Adaptation

IROS 2024poster

The goal of this paper is to strike a feasible tracking paradigm that can make 3D human trackers applicable on robot platforms and enable more high-level tasks. Till now, two fundamental problems haven’t been adequately addressed. One is the computational cost lightweight enough for robotic deployme…

Cited by 0SourceScholar
2024

ASWT-SGNN: Adaptive Spectral Wavelet Transform-Based Self-Supervised Graph Neural Network

AAAI 2024technical

Graph Comparative Learning (GCL) is a self-supervised method that combines the advantages of Graph Convolutional Networks (GCNs) and comparative learning, making it promising for learning node representations. However, the GCN encoders used in these methods rely on the Fourier transform to learn fix…

Cited by 7SourcePDFScholar
2024

An Aggregation-Free Federated Learning for Tackling Data Heterogeneity

CVPR 2024poster

The performance of Federated Learning (FL) hinges on the effectiveness of utilizing knowledge from distributed datasets. Traditional FL methods adopt an aggregate-then-adapt framework where clients update local models based on a global model aggregated by the server from the previous training round.…

Cited by 42SourcePDFScholar
2024

AutoTimes: Autoregressive Time Series Forecasters via Large Language Models

NeurIPS 2024poster

Foundation models of time series have not been fully developed due to the limited availability of time series corpora and the underexploration of scalable pre-training. Based on the similar sequential formulation of time series and natural language, increasing research demonstrates the feasibility o…

2024

BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays

NeurIPS 2024poster

Medical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and facilitate adapting task-specific models to new setups using fe…

2024

Beyond Prototypes: Semantic Anchor Regularization for Better Representation Learning

AAAI 2024technical

One of the ultimate goals of representation learning is to achieve compactness within a class and well-separability between classes. Many outstanding metric-based and prototype-based methods following the Expectation-Maximization paradigm, have been proposed for this objective. However, they inevita…

2024

Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection

CVPR 2024poster

Video Moment Retrieval (MR) and Highlight Detection (HD) have attracted significant attention due to the growing demand for video analysis. Recent approaches treat MR and HD as similar video grounding problems and address them together with transformer-based architecture. However we observe that the…

2024

CKT-RCM: Clip-Based Knowledge Transfer and Relational Context Mining for Unbiased Panoptic Scene Graph Generation

ICASSP 2024accepted

Panoptic Scene Graph (PSG) generation aims to generate a scene graph representing pairwise relationship between objects within an image. Its use of pixel-wise segmentation mask and inclusion of background regions in relationship inference make it quickly become a popular approach. However, it has an…

Cited by 0SourceScholar
2024

Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving

NeurIPS 2024poster

Autonomous driving has advanced significantly due to sensors, machine learning, and artificial intelligence improvements. However, prevailing methods struggle with intricate scenarios and causal relationships, hindering adaptability and interpretability in varied environments. To address the above p…

2024

Convolutional Spectral Kernel Learning with Generalization Guarantees (Abstract Reprint)

AAAI 2024technical

Kernel methods are powerful tools to capture nonlinear patterns behind given data but often lead to poor performance on complicated tasks compared to convolutional neural networks. The reason is that kernel methods are still shallow and fully connected models, failing to reveal hierarchical features…

Cited by 0SourcePDFScholar
2024

DCS: Debiased Contrastive Learning with Weak Supervision for Time Series Classification

ICASSP 2024accepted

Self-supervised contrastive learning (SSCL) has performed excellently on time series classification tasks. Most SSCL- based classification algorithms generate positive and negative samples in the time or frequency domains, focusing on mining similarities between them. However, two issues are not wel…

Cited by 0SourceScholar
2024

Decentralized Riemannian Conjugate Gradient Method on the Stiefel Manifold

ICLR 2024poster

The conjugate gradient method is a crucial first-order optimization method that generally converges faster than the steepest descent method, and its computational cost is much lower than that of second-order methods. However, while various types of conjugate gradient methods have been studied in Euc…

Cited by 11SourcePDFScholar
2024

ETAS: Zero-Shot Transformer Architecture Search via Network Trainability and Expressivity

ACL 2024findings

Transformer Architecture Search (TAS) methods aim to automate searching for the optimal Transformer architecture configurations for a given task. However, they are impeded by the prohibitive cost of evaluating Transformer architectures. Recently, several Zero-Shot TAS methods have been proposed to m…

Cited by 1SourcePDFScholar
2024

Enhancing In-Context Learning Performance with just SVD-Based Weight Pruning: A Theoretical Perspective

NeurIPS 2024poster

Pre-trained large language models (LLMs) based on Transformer have demonstrated striking in-context learning (ICL) abilities. With a few demonstration input-label pairs, they can predict the label for an unseen input without any parameter updates. In this paper, we show an exciting phenomenon that…

2024

Face Adapter for Pre-Trained Diffusion Models with Fine-Grained ID and Attribute Control

ECCV 2024poster

"Current face reenactment and swapping methods mainly rely on GAN frameworks, but recent focus has shifted to pre-trained diffusion models for their superior generation capabilities. However, training these models is resource-intensive, and the results have not yet achieved satisfactory performance…

Cited by 27SourcePDFScholar
2024

GFMAE: Self-Supervised GNN-Free Masked Autoencoders

ICASSP 2024accepted

Generative self-supervised learning, represented by graph autoencoders (GAEs), has begun to exhibit significant potential in addressing graph tasks. However, GAEs often rely on Graph Neural Networks (GNNs) for encoding and decoding, this can pose a computation challenge due to the inherent complexit…

Cited by 0SourceScholar
2024

HeterGCL: Graph Contrastive Learning Framework on Heterophilic Graph

IJCAI 2024poster

Graph Contrastive Learning (GCL) has attracted significant research attention due to its self-supervised ability to learn robust node representations. Unfortunately, most methods primarily focus on homophilic graphs, rendering them less effective for heterophilic graphs. In addition, the complexity…

2024

Hierarchical Search-Based Cooperative Motion Planning

IROS 2024poster

Cooperative path planning, a crucial aspect of multi-agent systems research, serves a variety of sectors, including military, agriculture, and industry. Many existing algorithms, however, come with certain limitations, such as simplified kinematic models and inadequate support for multiple group sce…

Cited by 0SourcecodeScholar
2024

High-Dimensional Analysis for Generalized Nonlinear Regression: From Asymptotics to Algorithm

AAAI 2024technical

Overparameterization often leads to benign overfitting, where deep neural networks can be trained to overfit the training data but still generalize well on unseen data. However, it lacks a generalized asymptotic framework for nonlinear regressions and connections to conventional complexity notions.…

2024

LORS: Low-rank Residual Structure for Parameter-Efficient Network Stacking

CVPR 2024highlight

Deep learning models particularly those based on transformers often employ numerous stacked structures which possess identical architectures and perform similar functions. While effective this stacking paradigm leads to a substantial increase in the number of parameters pos- ing challenges for pract…

2024

Learning Multi-Scale Video-Text Correspondence for Weakly Supervised Temporal Article Gronding

AAAI 2024technical

Weakly Supervised temporal Article Grounding (WSAG) is a challenging and practical task in video understanding. Specifically, given a video and a relevant article, whose sentences are at different semantic scales, WSAG aims to localize corresponding video segments for all “groundable” sentences. Com…

Cited by 1SourcePDFScholar
2024

Learning Safe Locomotion for Quadrupedal Robots by Derived-Action Optimization

IROS 2024poster

Deep reinforcement learning controllers with exteroception have enabled quadrupedal robots to traverse terrain robustly. However, most of these controllers heavily depend on complex reward functions and suffer from poor convergence. This work proposes a novel learning framework called derived-action…

Cited by 0SourceScholar
2024

LiteGrasp: A Light Robotic Grasp Detection via Semi-Supervised Knowledge Distillation

RA-L 2024

Grasping detection from single images in robotic applications poses a significant challenge. While contemporary deep learning techniques excel, their success often hinges on large annotated datasets and intricate network architectures. In this letter, we present LiteGrasp, a novel semi-supervised li

Cited by 2SourceScholar
2024

MC-indexing: Effective Long Document Retrieval via Multi-view Content-aware Indexing

EMNLP 2024finding

Long document question answering (DocQA) aims to answer questions from long documents over 10k words. They usually contain content structures such as sections, sub-sections, and paragraph demarcations. However, the indexing methods of long documents remain under-explored, while existing systems gene…

2024

MaxQ: Multi-Axis Query for N:M Sparsity Network

CVPR 2024poster

N:M sparsity has received increasing attention due to its remarkable performance and latency trade-off compared with structured and unstructured sparsity. However existing N:M sparsity methods do not differentiate the relative importance of weights among blocks and leave important weights underappre…

2024

Monocular Event-Inertial Odometry with Adaptive decay-based Time Surface and Polarity-aware Tracking

IROS 2024poster

Event cameras have garnered considerable attention due to their advantages over traditional cameras in low power consumption, high dynamic range, and no motion blur. This paper proposes a monocular event-inertial odometry incorporating an adaptive decay kernel-based time surface with polarity-aware…

Cited by 2SourceScholar
2024

Multi-modal 3D Human Tracking for Robots in Complex Environment with Siamese Point-Video Transformer

ICRA 2024poster

Tracking a specific person in 3D scene is gaining momentum due to its numerous applications in robotics. Currently, most 3D trackers focus on driving scenarios with neglected jitter and uncomplicated surroundings, which results in their severe degeneration in complex environments, especially on jolt…

Cited by 3SourceScholar
2024

Open-Vocabulary Segmentation with Semantic-Assisted Calibration

CVPR 2024poster

This paper studies open-vocabulary segmentation (OVS) through calibrating in-vocabulary and domain-biased embedding space with generalized contextual prior of CLIP. As the core of open-vocabulary understanding alignment of visual content with the semantics of unbounded text has become the bottleneck…

Cited by 32SourcePDFScholar
2024

OvSW: Overcoming Silent Weights for Accurate Binary Neural Networks

ECCV 2024poster

"Binary Neural Networks (BNNs) have been proven to be highly effective for deploying deep neural networks on mobile and embedded platforms. Most existing works focus on minimizing quantization errors, improving representation ability, or designing gradient approximations to alleviate gradient mismat…

2024

Paint3D: Paint Anything 3D with Lighting-Less Texture Diffusion Models

CVPR 2024poster

This paper presents Paint3D a novel coarse-to-fine generative framework that is capable of producing high-resolution lighting-less and diverse 2K UV texture maps for untextured 3D meshes conditioned on text or image inputs. The key challenge addressed is generating high-quality textures without embe…

2024

RLPeri: Accelerating Visual Perimetry Test with Reinforcement Learning and Convolutional Feature Extraction

AAAI 2024technical

Visual perimetry is an important eye examination that helps detect vision problems caused by ocular or neurological conditions. During the test, a patient's gaze is fixed at a specific location while light stimuli of varying intensities are presented in central and peripheral vision. Based on the pa…

Cited by 0SourcePDFScholar
2024

RadarCam-Depth: Radar-Camera Fusion for Depth Estimation with Learned Metric Scale

ICRA 2024poster

We present a novel approach for metric dense depth estimation based on the fusion of a single-view image and a sparse, noisy Radar point cloud. The direct fusion of heterogeneous Radar and image data, or their encodings, tends to yield dense depth maps with significant artifacts, blurred boundaries,…

Cited by 11SourcecodeScholar
2024

SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

CVPR 2024poster

Multimodal Visual Object Tracking (VOT) has recently gained significant attention due to its robustness. Early research focused on fully fine-tuning RGB-based trackers which was inefficient and lacked generalized representation due to the scarcity of multimodal data. Therefore recent studies have ut…

2024

Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image Editing

NeurIPS 2024poster

Text-guided diffusion models have significantly advanced image editing, enabling high-quality and diverse modifications driven by text prompts. However, effective editing requires inverting the source image into a latent space, a process often hindered by prediction errors inherent in DDIM inversion…

2024

Semi-Supervised Learning for Visual Bird’s Eye View Semantic Segmentation

ICRA 2024poster

Visual bird’s eye view (BEV) semantic segmentation helps autonomous vehicles understand the surrounding environment only from front-view (FV) images, including static elements (e.g., roads) and dynamic elements (e.g., vehicles, pedestrians). However, the high cost of annotation procedures of full-su…

Cited by 4SourcecodeScholar
2024

Sentence-level Prompts Benefit Composed Image Retrieval

ICLR 2024spotlight

Composed image retrieval (CIR) is the task of retrieving specific images by using a query that involves both a reference image and a relative caption. Most existing CIR models adopt the late-fusion strategy to combine visual and language features. Besides, several approaches have also been suggested…

2024

Skip-Step Contrastive Predictive Coding for Time Series Anomaly Detection

ICASSP 2024accepted

Self-supervised learning (SSL) shows impressive performance in many tasks lacking sufficient labels. In this paper, we study SSL in time series anomaly detection (TSAD) by incorporating the characteristics of time series data. Specifically, we build an anomaly detection algorithm consisting of globa…

Cited by 0SourceScholar
2024

Solving Homogeneous and Heterogeneous Cooperative Tasks with Greedy Sequential Execution

ICLR 2024spotlight

Cooperative multi-agent reinforcement learning (MARL) is extensively used for solving complex cooperative tasks, and value decomposition methods are a prevalent approach for this domain. However, these methods have not been successful in addressing both homogeneous and heterogeneous tasks simultaneo…

Cited by 2SourcePDFScholar
2024

Structured Optimal Brain Pruning for Large Language Models

EMNLP 2024main

The massive parameters and computational demands hinder the widespread application of Large Language Models (LLMs). Network pruning provides a practical solution to this problem. However, existing pruning works for LLMs mainly focus on unstructured pruning or necessitate post-pruning fine-tuning. Th…

Cited by 1SourcePDFScholar
2024

TimeXer: Empowering Transformers for Time Series Forecasting with Exogenous Variables

NeurIPS 2024poster

Deep models have demonstrated remarkable performance in time series forecasting. However, due to the partially-observed nature of real-world applications, solely focusing on the target of interest, so-called endogenous variables, is usually insufficient to guarantee accurate forecasting. Notably, a…

2024

Timer: Generative Pre-trained Transformers Are Large Time Series Models

ICML 2024poster

Deep learning has contributed remarkably to the advancement of time series analysis. Still, deep models can encounter performance bottlenecks in real-world data-scarce scenarios, which can be concealed due to the performance saturation with small models on current benchmarks. Meanwhile, large models…

2024

Timestep-Aware Correction for Quantized Diffusion Models

ECCV 2024poster

"Diffusion models have marked a significant breakthrough in the synthesis of semantically coherent images. However, their extensive noise estimation networks and the iterative generation process limit their wider application, particularly on resource-constrained platforms like mobile devices. Existi…

Cited by 4SourcePDFScholar
2024

Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models

ACL 2024findings

Ensuring the trustworthiness of large language models (LLMs) is crucial. Most studies concentrate on fully pre-trained LLMs to better understand and improve LLMs’ trustworthiness. In this paper, to reveal the untapped potential of pre-training, we pioneer the exploration of LLMs’ trustworthiness dur…

2024

Towards Understanding How Transformers Learn In-context Through a Representation Learning Lens

NeurIPS 2024poster

Pre-trained large language models based on Transformers have demonstrated remarkable in-context learning (ICL) abilities. With just a few demonstration examples, the models can implement new tasks without any parameter updates. However, it is still an open question to understand the mechanism of ICL…

Cited by 2SourcePDFScholar
2024

Tuning-Free Image Customization with Image and Text Guidance

ECCV 2024poster

"Despite significant advancements in image customization with diffusion models, current methods still have several limitations: 1) unintended changes in non-target areas when regenerating the entire image; 2) guidance solely by a reference image or text descriptions; and 3) time-consuming fine-tunin…

2024

Universal Segmentation at Arbitrary Granularity with Language Instruction

CVPR 2024poster

This paper aims to achieve universal segmentation of arbitrary semantic level. Despite significant progress in recent years specialist segmentation approaches are limited to specific tasks and data distribution. Retraining a new model for adaptation to new scenarios or settings takes expensive compu…

Cited by 16SourcePDFScholar
2024

Unsupervised Continual Anomaly Detection with Contrastively-Learned Prompt

AAAI 2024technical

Unsupervised Anomaly Detection (UAD) with incremental training is crucial in industrial manufacturing, as unpredictable defects make obtaining sufficient labeled data infeasible. However, continual learning methods primarily rely on supervised annotations, while the application in UAD is limited due…

2024

WaveNet: Tackling Non-stationary Graph Signals via Graph Spectral Wavelets

AAAI 2024technical

In the existing spectral GNNs, polynomial-based methods occupy the mainstream in designing a filter through the Laplacian matrix. However, polynomial combinations factored by the Laplacian matrix naturally have limitations in message passing (e.g., over-smoothing). Furthermore, most existing spectra…

2024

iTransformer: Inverted Transformers Are Effective for Time Series Forecasting

ICLR 2024spotlight

The recent boom of linear forecasting models questions the ongoing passion for architectural modifications of Transformer-based forecasters. These forecasters leverage Transformers to model the global dependencies over temporal tokens of time series, with each token formed by multiple variates of th…

2023

A Unified BEV Model for Joint Learning of 3D Local Features and Overlap Estimation

ICRA 2023poster

Pairwise point cloud registration is a critical task for many applications, which heavily depends on finding correct correspondences from the two point clouds. However, the low overlap between input point clouds causes the registration to fail easily, leading to mistaken overlapping and mismatched c…

Cited by 5SourcecodeScholar
2023

AdaCM: Adaptive ColorMLP for Real-Time Universal Photo-Realistic Style Transfer

AAAI 2023technical

Photo-realistic style transfer aims at migrating the artistic style from an exemplar style image to a content image, producing a result image without spatial distortions or unrealistic artifacts. Impressive results have been achieved by recent deep models. However, deep neural network based methods…

Cited by 4SourcePDFScholar
2023

Adaptive Assignment for Geometry Aware Local Feature Matching

CVPR 2023poster

The detector-free feature matching approaches are currently attracting great attention thanks to their excellent performance. However, these methods still struggle at large-scale and viewpoint variations, due to the geometric inconsistency resulting from the application of the mutual nearest neighbo…

2023

Boosting Few-shot Action Recognition with Graph-guided Hybrid Matching

ICCV 2023poster

Class prototype construction and matching are core aspects of few-shot action recognition. Previous methods mainly focus on designing spatiotemporal relation modeling modules or complex temporal alignment algorithms. Despite the promising results, they ignored the value of class prototype constructi…

Cited by 37PDFcodeScholar
2023

Coco-LIC: Continuous-Time Tightly-Coupled LiDAR-Inertial-Camera Odometry Using Non-Uniform B-Spline

RA-L 2023

In this letter, we propose an efficient continuous-time LiDAR-Inertial-Camera Odometry, utilizing non-uniform B-splines to tightly couple measurements from the LiDAR, IMU, and camera. In contrast to uniform B-spline-based continuous-time methods, our non-uniform B-spline approach offers significant

Cited by 39SourcecodeScholar
2023

Consistency of Multiple Kernel Clustering

ICML 2023poster

Consistency plays an important role in learning theory. However, in multiple kernel clustering (MKC), the consistency of kernel weights has not been sufficiently investigated. In this work, we fill this gap with a non-asymptotic analysis on the consistency of kernel weights of a novel method termed…

Cited by 9SourcePDFScholar
2023

Diverse Data Augmentation with Diffusions for Effective Test-time Prompt Tuning

ICCV 2023poster

Benefiting from prompt tuning, recent years have witnessed the promising performance of pre-trained vision-language models, e.g., CLIP, on versatile downstream tasks. In this paper, we focus on a particular setting of learning adaptive prompts on the fly for each test sample from an unseen new domai…

Cited by 93PDFcodeScholar
2023

FG-Depth: Flow-Guided Unsupervised Monocular Depth Estimation

ICRA 2023poster

The great potential of unsupervised monocular depth estimation has been demonstrated by many works due to low annotation cost and impressive accuracy comparable to supervised methods. To further improve the performance, recent works mainly focus on designing more complex network structures and explo…

Cited by 8SourceScholar
2023

Fair Scratch Tickets: Finding Fair Sparse Networks Without Weight Training

CVPR 2023poster

Recent studies suggest that computer vision models come at the risk of compromising fairness. There are extensive works to alleviate unfairness in computer vision using pre-processing, in-processing, and post-processing methods. In this paper, we lead a novel fairness-aware learning paradigm for in-…

2023

Generalization Bounds for Federated Learning: Fast Rates, Unparticipating Clients and Unbounded Losses

ICLR 2023poster

In {federated learning}, the underlying data distributions may be different across clients. This paper provides a theoretical analysis of generalization error of {federated learning}, which captures both heterogeneity and relatedness of the distributions. In particular, we assume that the heterogene…

Cited by 18SourcePDFScholar
2023

Generative Gradient Inversion via Over-Parameterized Networks in Federated Learning

ICCV 2023poster

Federated learning has gained recognitions as a secure approach for safeguarding local private data in collaborative learning. But the advent of gradient inversion research has posed significant challenges to this premise by enabling a third-party to recover groundtruth images via gradients. While p…

Cited by 13PDFcodeScholar
2023

Geo-Localization With Transformer-Based 2D-3D Match Network

RA-L 2023

This letter presents a novel method for geographical localization by registering satellite maps with LiDAR point clouds. This method includes a Transformer-based 2D-3D matching network called D-GLSNet that directly matches the LiDAR point clouds and satellite images through end-to-end learning. With

Cited by 18SourceScholar
2023

Global Knowledge Calibration for Fast Open-Vocabulary Segmentation

ICCV 2023poster

Recent advancements in pre-trained vision-language models, such as CLIP, have enabled the segmentation of arbitrary concepts solely from textual inputs, a process commonly referred to as open-vocabulary semantic segmentation (OVS). However, existing OVS techniques confront a fundamental challenge: t…

Cited by 55PDFScholar
2023

High-Fidelity Generalized Emotional Talking Face Generation With Multi-Modal Emotion Space Learning

CVPR 2023poster

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing to handle unseen emotion styles due to limited semantics. T…

Cited by 46SourcePDFScholar
2023

Koopa: Learning Non-stationary Time Series Dynamics with Koopman Predictors

NeurIPS 2023poster

Real-world time series are characterized by intrinsic non-stationarity that poses a principal challenge for deep forecasting models. While previous models suffer from complicated series variations induced by changing temporal distribution, we tackle non-stationary time series with modern Koopman the…

2023

Large Scale Pursuit-Evasion Under Collision Avoidance Using Deep Reinforcement Learning

IROS 2023poster

This paper examines a pursuit-evasion game (PEG) involving multiple pursuers and evaders. The decentralized pursuers aim to collaborate to capture the faster evaders while avoiding collisions. The policies of all agents are learning-based and are subjected to kinematic constraints that are specific…

Cited by 5SourceScholar
2023

Learning Federated Visual Prompt in Null Space for MRI Reconstruction

CVPR 2023poster

Federated Magnetic Resonance Imaging (MRI) reconstruction enables multiple hospitals to collaborate distributedly without aggregating local data, thereby protecting patient privacy. However, the data heterogeneity caused by different MRI protocols, insufficient local training data, and limited commu…

2023

Learning Global-aware Kernel for Image Harmonization

ICCV 2023poster

Image harmonization aims to solve the visual inconsistency problem in composited images by adaptively adjusting the foreground pixels with the background as references. Existing methods employ local color transformation or region matching between foreground and background, which neglects powerful pr…

Cited by 9PDFScholar
2023

Learning To Measure the Point Cloud Reconstruction Loss in a Representation Space

CVPR 2023poster

For point cloud reconstruction-related tasks, the reconstruction losses to evaluate the shape differences between reconstructed results and the ground truths are typically used to train the task networks. Most existing works measure the training loss with point-to-point distance, which may introduce…

Cited by 7SourcePDFScholar
2023

MFF-Net: Towards Efficient Monocular Depth Completion With Multi-Modal Feature Fusion

RA-L 2023

Remarkable progress has been achieved by current depth completion approaches, which produce dense depth maps from sparse depth maps and corresponding color images. However, the performances of these approaches are limited due to the insufficient feature extractions and fusions. In this work, we prop

Cited by 38SourceScholar
2023

MHCCL: Masked Hierarchical Cluster-Wise Contrastive Learning for Multivariate Time Series

AAAI 2023technical

Learning semantic-rich representations from raw unlabeled time series data is critical for downstream tasks such as classification and forecasting. Contrastive learning has recently shown its promising representation learning capability in the absence of expert annotations. However, existing contras…

2023

NeRF-Loc: Visual Localization with Conditional Neural Radiance Field

ICRA 2023poster

We propose a novel visual re-localization method based on direct matching between the implicit 3D descriptors and the 2D image with transformer. A conditional neural radiance field(NeRF) is chosen as the 3D scene representation in our pipeline, which supports continuous 3D descriptors generation and…

Cited by 42SourcecodeScholar
2023

Next POI Recommendation with Dynamic Graph and Explicit Dependency

AAAI 2023technical

Next Point-Of-Interest (POI) recommendation plays an important role in various location-based services. Its main objective is to predict the user's next interested POI based on her previous check-in information. Most existing methods directly use users' historical check-in trajectories to construct…

2023

PANet: LiDAR Panoptic Segmentation with Sparse Instance Proposal and Aggregation

IROS 2023poster

Reliable LiDAR panoptic segmentation (LPS), including both semantic and instance segmentation, is vital for many robotic applications, such as autonomous driving. This work proposes a new LPS framework named PANet to eliminate the dependency on the offset branch and improve the performance on large…

Cited by 5SourcecodeScholar
2023

RICO: Regularizing the Unobservable for Indoor Compositional Reconstruction

ICCV 2023poster

Recently, neural implicit surfaces have become popular for multi-view reconstruction. To facilitate practical applications like scene editing and manipulation, some works extend the framework with semantic masks input for the object-compositional reconstruction rather than the holistic perspective.…

Cited by 13PDFcodeScholar
2023

RIFormer: Keep Your Vision Backbone Effective but Removing Token Mixer

CVPR 2023poster

This paper studies how to keep a vision backbone effective while removing token mixers in its basic building blocks. Token mixers, as self-attention for vision transformers (ViTs), are intended to perform information communication between different spatial tokens but suffer from considerable computa…

Cited by 37SourcePDFScholar