← Search

YiFan Zhang

147 accepted papers

2026

AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors

AAAI 2026technical

As Convolutional Neural Networks (CNNs) continue to gain traction in deep learning, Winograd convolution has emerged as a key algorithm to enhance computational efficiency. Although ARM-based CPUs are increasingly prevalent in mobile devices, embedded systems and HPC servers, existing 2D Winograd co

Cited by 0SourcePDFScholar
2026

AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models

ICLR 2026poster

The rapid development and widespread adoption of Audio Large Language Models (ALLMs) require a rigorous assessment of their trustworthiness. However, existing evaluation frameworks, primarily designed for text, are not equipped to handle the unique vulnerabilities introduced by audio’s acoustic prop…

Cited by 0SourcecodeScholar
2026

BaseReward: A Strong Baseline for Multimodal Reward Model

ICLR 2026poster

The rapid advancement of Multimodal Large Language Models (MLLMs) has made aligning them with human preferences a critical challenge. Reward Models (RMs) are a core technology for achieving this goal, but a systematic guide for building state-of-the-art Multimodal Reward Models (MRMs) is currently l…

Cited by 0SourceScholar
2026

CoGenSAM: Codebook-Interactive Generative Labeling for Adapting SAM to Crack Segmentation

AAAI 2026technical

The goal of this work is to adapt Segment Anything Models (SAM) into crack segmentation tasks via automatic label generation, thus eliminating manual annotation cost. In this regard, an intuitive approach is to extract edges of crack samples and generate labels via the dilation and erosion processes

Cited by 0SourcePDFScholar
2026

CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human

ICRA 2026poster

In this work, we present CollabVLA, a self-reflective vision-language-action framework that transforms a standard visuomotor policy into a collaborative assistant. CollabVLA tackles key limitations of prior VLAs, including domain overfitting, non-interpretable reasoning, and the high latency of auxi…

2026

DRAMA: Next-Gen Dynamic Orchestration for Resilient Multi-Agent Ecosystems in Flux

CVPR 2026

Embodied Multi-Agent Systems have proven highly effective in addressing complex tasks through coordinated collaboration among heterogeneous agents. However, real-world environments and task specifications are inherently dynamic, exhibiting frequent changes, uncertainty, and variability. Despite thes

Cited by 0SourceScholar
2026

DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous Driving

AAAI 2026technical

In autonomous driving, vision-centric 3D object detection recognizes and localizes 3D objects from RGB images. However, due to high annotation costs and diverse outdoor scenes, training data often fails to cover all possible test scenarios, known as the out-of-distribution (OOD) issue. Training-free

Cited by 0SourcePDFScholar
2026

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning

RSS 2026poster

Extrinsic dexterity leverages environmental contact to overcome the limitations of prehensile manipulation. However, achieving such dexterity in cluttered scenes remains challenging and underexplored, as it requires selectively exploiting contact among multiple interacting objects with inherently co…

Cited by 0SourceScholar
2026

Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction

AAAI 2026technical

Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competitive performance. However, these pipelines rely on implicit modeling that uses fra

Cited by 0SourcePDFScholar
2026

Generalized Boundary FDR Control under Arbitrary Dependence: An Approach on Closure Principle

ICML 2026poster

False discovery rate (FDR) is a cornerstone of modern multiple testing. However, it often fails to guarantee the reliability of ``marginal" discoveries that lie at the boundary of the rejection set, which are often crucial in high-precision applications. While recent works (Soloff et al., 2024; Xian…

Cited by 0SourceScholar
2026

Group Representational Position Encoding

ICLR 2026poster

We present GRAPE (Group RepresentAtional Position Encoding), a unified framework for positional encoding based on group actions. GRAPE brings together two families of mechanisms: (i) multiplicative rotations (Multiplicative GRAPE) in $\operatorname{SO}(d)$ and (ii) additive logit biases (Additive GR…

Cited by 0SourcecodeScholar
2026

HyMTRL: A Hybrid Multi-Task Reinforcement Learning Framework via Phased Policy Evolution

ICML 2026poster

Multi-task reinforcement learning (MTRL) aims to improve sample efficiency by sharing knowledge across related tasks, but it often suffers from asynchronous learning progress caused by inherent differences in task difficulty. This imbalance places substantial representational strain on the shared cr…

Cited by 0SourceScholar
2026

InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition

ICML 2026poster

Upweighting high-quality data in LLM pretraining often improves performance, but in data-limited regimes, especially under overtraining, stronger upweighting increases repetition and can degrade performance. However, standard scaling laws do not reliably extrapolate across mixture recipes or under r…

Cited by 0SourceScholar
2026

LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments

ICRA 2026poster

Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an agent to navigate unseen environments based on natural language instructions without any prior training. Current methods face a critical trade-off: either rely on environment-specific waypoint predictors that li…

2026

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining

ICML 2026poster

Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit inductive bias for locality, leading to redundant modeling of sequence-int…

Cited by 0SourceScholar
2026

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

CVPR 2026

Large multimodal models (LMMs) have shown great potential for video reasoning with textual Chain-of-Thought. However, they remain vulnerable to hallucinations, especially when processing long-form videos where evidence is sparse and temporally dispersed. Inspired by how humans comprehend long videos

Cited by 50SourcecodeScholar
2026

MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance

AAAI 2026technical

Recent advances in instruction-based image editing have shown remarkable progress. However, existing methods remain limited to relatively simple editing operations, hindering real-world applications that require complex and compositional instructions. In this work, we address these limitations from

Cited by 0SourcePDFScholar
2026

MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current emotional benchmarks remain limited, as it is still unknown: (a)…

Cited by 0SourcecodeScholar
2026

MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models

ICLR 2026poster

Unified Multimodal Large Language Models (U-MLLMs) have garnered considerable interest for their ability to seamlessly integrate generation and comprehension tasks. However, existing research lacks a unified evaluation standard, often relying on isolated benchmarks to assess these capabilities. More…

Cited by 0SourceScholar
2026

Medical thinking with multiple images

ICLR 2026poster

Large language models and vision-language models score high on many medical QA benchmarks; however, real-world clinical reasoning remains challenging because cases often involve multiple images and require cross-view fusion. We present MedThinkVQA, a benchmark that asks models to think with multiple…

Cited by 0SourcecodeScholar
2026

On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning

ICLR 2026poster

Policy gradient algorithms have been successfully applied to enhance the reasoning capabilities of large language models (LLMs). KL regularization is ubiquitous, yet the design surface, choice of KL direction (forward vs. reverse), normalization (normalized vs. unnormalized), and estimator ($k_1/k_2…

Cited by 0SourcecodeScholar
2026

Preference-Modulated Structural Attention for Multi-Objective Combinatorial Optimization

ICML 2026poster

Recent decomposition-based approaches have achieved significant success in Multi-Objective Combinatorial Optimization (MOCO). However,existing methods typically rely exclusively on node-centric representations, failing to capture the complementary representations provided by edge features for proble…

Cited by 0SourceScholar
2026

ProOPF: Benchmarking and Improving LLMs for Professional-Grade Power Systems Optimization Modeling

ICML 2026poster

Growing renewable penetration introduces substantial uncertainty into power system operations, necessitating frequent adaptation of dispatch objectives and constraints and challenging expertise-intensive, near-real-time modeling workflows. Large Language Models (LLMs) provide a promising avenue for …

Cited by 0SourceScholar
2026

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

ICLR 2026poster

Multimodal Reward Models (MRMs) play a crucial role in enhancing the performance of Multimodal Large Language Models (MLLMs). While recent advancements have primarily focused on improving the model structure and training data of MRMs, there has been limited exploration into the effectiveness of long…

Cited by 0SourcecodeScholar
2026

RefineEvo: Planning-Guided Heuristic Evolution with Bidirectional Experience

ICML 2026poster

Automatic Heuristic Design (AHD) has emerged as a transformative approach for solving combinatorial optimization problems. While recent Large Language Model (LLM)-based methods have shown promise, they predominantly rely on fixed evolutionary operators and struggle to effectively accumulate and reus…

Cited by 0SourceScholar
2026

SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search

AAAI 2026technical

Large Language Models (LLMs) often falter at complex planning tasks that require exploration and self-correction, as their linear reasoning process struggles to recover from early mistakes. While search algorithms like Monte Carlo Tree Search (MCTS) can explore alternatives, they are often ineffecti

Cited by 0SourcePDFScholar
2026

Towards Human-Like Robot Handwriting via Contour-Aware Generation

CVPR 2026

Empowering machines to simulate human handwriting is a promising research direction. Most existing methods, however, primarily focus on reproducing the writing trajectory to capture the overall character structure, while neglecting the critical aspect of stroke contour modeling. Consequently, these

Cited by 0SourceScholar
2026

Translation Heads: Unveiling Attention's Role in LLM Multilingual Translation

ICLR 2026poster

Recently, large language models (LLMs) have made remarkable progress, with multilingual capability emerging as a core foundational strengths. However, the internal mechanisms by which these models perform translation remain incompletely understood. In this paper, we elucidate the relationship betwee…

Cited by 0SourceScholar
2026

Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer

ICML 2026poster

Algorithm extraction aims to synthesize executable programs directly from models trained on algorithmic tasks, enabling *de novo* algorithm discovery without relying on human-written code. However, applying this paradigm to Transformer is hindered by representation entanglement (e.g., superposition)…

Cited by 0SourceScholar
2025

A Comprehensive Evaluation on Event Reasoning of Large Language Models

AAAI 2025technical

Event reasoning is a fundamental ability that underlies many applications. It requires event schema knowledge to perform global reasoning and needs to deal with the diversity of the inter-event relations and the reasoning paradigms. The extent to which LLMs excel in event reasoning across various re…

2025

AI-Driven Virtual Teacher for Enhanced Educational Efficiency: Leveraging Large Pretrain Models for Autonomous Error Analysis and Correction

AAAI 2025technical

Students frequently make mistakes while solving mathematical problems, and traditional error correction methods are both time-consuming and labor-intensive. This paper introduces an innovative Virtual AI Teacher system designed to autonomously analyze and correct student Errors (VATE). Leveraging ad…

2025

ALIC: Adaptive Fusion Entropy Model for Learned Image Compression

ICASSP 2025accepted

Recently, learned image compression algorithms have achieved significant performance. The entropy model is crucial for improving the rate-distortion performance by estimating the probability distribution of latent representation. In this paper, we propose an adaptive fusion entropy model for learned…

Cited by 0SourceScholar
2025

Augmenting Math Word Problems via Iterative Question Composing

AAAI 2025technical

Despite the advancements in large language models (LLMs) for mathematical reasoning, solving competition-level math problems remains a significant challenge, especially for open-source LLMs without external tools. We introduce the MMIQC dataset, comprising a mixture of processed web data and synthet…

2025

Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts

ACL 2025finding

We present Autonomous Data Selection (AutoDS), a method that leverages base language models as zero-shot “generative classifiers” to automatically curate high-quality mathematical texts. Unlike prior approaches that require human annotations or training a dedicated data filter, AutoDS relies solely…

Cited by 0SourcePDFScholar
2025

Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment

ICML 2025poster

Modeling human preferences is crucial for aligning foundation models with human values. Traditional reward modeling methods, such as the Bradley-Terry (BT) reward model, fall short in expressiveness, particularly in addressing intransitive preferences. In this paper, we introduce \emph{preference em…

2025

Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line Generation

ICCV 2025poster

Existing handwritten text generation methods primarily focus on isolated words. However, realistic handwritten text demands attention not only to individual words but also to the relationships between them, such as vertical alignment and horizontal spacing. Therefore, generating entire text line eme…

2025

Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks

ICLR 2025spotlight

Generative Flow Networks (GFlowNets) are a novel class of generative models designed to sample from unnormalized distributions and have found applications in various important tasks, attracting great research interest in their training algorithms. In general, GFlowNets are trained by fitting the for…

Cited by 1SourcePDFScholar
2025

Bi-Level Knowledge Transfer for Multi-Task Multi-Agent Reinforcement Learning

NeurIPS 2025poster

Multi-Agent Reinforcement Learning (MARL) has achieved remarkable success in various real-world scenarios, but its high cost of online training makes it impractical to learn each task from scratch. To enable effective policy reuse, we consider the problem of zero-shot generalization from offline da…

Cited by 0SourceScholar
2025

CONSTRUCTA: Automating Commercial Construction Schedules in Fabrication Facilities with Large Language Models

NAACL 2025industry

Automating planning with LLMs presents transformative opportunities for traditional industries, yet remains underexplored. In commercial construction, the complexity of automated scheduling often requires manual intervention to ensure precision. We propose CONSTRUCTA, a novel framework leveraging LL…

Cited by 0SourcePDFScholar
2025

Chatbot To Help Patients Understand Their Health

EMNLP 2025

Patients must possess the knowledge necessary to actively participate in their care. To this end, we developed NoteAid-Chatbot, a conversational AI designed to help patients better understand their health through a novel framework of learning as conversation. We introduce a new learning paradigm tha

2025

Computational Complexity of Planning for Recursive Primitive Task Networks: Selective Action Nullification with State Preservation

IJCAI 2025

This paper investigates fundamental aspects of Hierarchical Task Network (HTN) planning by systematically exploring recursive arrangements of primitive task networks. Working within a general framework that aligns with recently identified ACKERMANN-complete HTN problems, we map the computational com

Cited by 0SourcePDFScholar
2025

DAMA: Data- and Model-aware Alignment of Multi-modal LLMs

ICML 2025poster

Direct Preference Optimization (DPO) has shown effectiveness in aligning multi-modal large language models (MLLM) with human preferences. However, existing methods exhibit an imbalanced responsiveness to the data of varying hardness, tending to overfit on the easy-to-distinguish data while underfit…

Cited by 0SourcePDFScholar
2025

DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation

CVPR 2025poster

In autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of training data. Once distribution shifts occur between training and test data, existing methods often suffer from perform…

2025

Dynamic Behavior Cloning With Temporal Feature Prediction: Enhancing Robotic Arm Manipulation in Moving Object Tasks

RA-L 2025

In numerous real-world applications, the ability to accurately perceive and respond to dynamic changes in the environment, while also maintaining the flexibility to transfer learned skills across different tasks, is crucial for the effective operation of robotic arms. Behavior cloning is particularl

Cited by 5SourceScholar
2025

ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive Programming

ACL 2025long

While recent research increasingly emphasizes the value of human-LLM collaboration in competitive programming and proposes numerous empirical methods, a comprehensive understanding remains elusive due to the fragmented nature of existing studies and their use of diverse, application-specific human f…

2025

Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models Pretraining

NeurIPS 2025poster

Large language models (LLMs) have become integral to a wide range of applications worldwide, driving an unprecedented global demand for effective multilingual capabilities. Central to achieving robust multilingual performance is the strategic allocation of language proportions within training corpor…

Cited by 0SourceScholar
2025

Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State Tracking

ACL 2025long

Chain-of-thought (CoT) significantly enhances the performance of large language models (LLMs) across a wide range of tasks, and prior research shows that CoT can theoretically increase expressiveness. However, there is limited mechanistic understanding of the algorithms that Transformer+CoT can lear…

2025

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations

EMNLP 2025

Large language models (LLMs) have demonstrated promising performance on medical benchmarks; however, their ability to perform medical calculations, a crucial aspect of clinical decision-making, remains underexplored and poorly evaluated. Existing benchmarks often assess only the final answer with a

Cited by 0SourcePDFScholar
2025

GOEN: Guided Obstacle Endpoint Navigation for Real-Time Collision-Free Path Planning in Unstructured Environments

IROS 2025

We present GOEN, an advanced navigation and path planning framework specifically engineered to tackle the complexities of dynamic and unstructured environments through real-time 3D pointcloud processing. Our approach integrates pointcloud downsampling, collision risk assessment, and obstacle endpoin

Cited by 0SourceScholar
2025

Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval

EMNLP 2025

Although Contrastive Language-Image Pre-training (CLIP) exhibits strong performance across diverse vision tasks, its application to person representation learning faces two critical challenges: (i) the scarcity of large-scale annotated vision-language data focused on person-centric images, and (ii)

2025

Improving Multimodal Social Media Popularity Prediction via Selective Retrieval Knowledge Augmentation

AAAI 2025technical

Understanding and predicting the popularity of online User-Generated Content (UGC) is critical for various social and recommendation systems. Existing efforts have focused on extracting predictive features and using pre-trained deep models to learn and fuse multimodal UGC representations. However, t…

2025

Learning to Extrapolate and Adjust: Two-Stage Meta-Learning for Concept Drift in Online Time Series Forecasting

IJCAI 2025

The inherent non-stationarity of time series in practical applications poses significant challenges for accurate forecasting. This paper tackles the concept drift problem where the underlying distribution or environment of time series changes. To better describe the characteristics and effectively m

2025

Learning to Generalize without Bias for Open-Vocabulary Action Recognition

ICCV 2025poster

Leveraging the effective visual-text alignment and static generalizability from CLIP, recent video learners adopt CLIP initialization with further regularization or recombination for generalization in open-vocabulary action recognition in-context. However, due to the static bias of CLIP, such video…

2025

LoRaDA: Low-Rank Direct Attention Adaptation for Efficient LLM Fine-tuning

EMNLP 2025

As the parameter size of language models becomes extremely large, fine-tuning them with limited resources has become a challenging task. Latest advancements in parameter-efficient fine-tuning (PEFT) techniques allow for adjustments to only a minor fraction of the parameters of these LLMs. Yet, most

Cited by 0SourcePDFScholar
2025

MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes

ICCV 2025poster

4D Gaussian Splatting (4DGS) has recently emerged as a promising technique for capturing complex dynamic 3D scenes with high fidelity. It utilizes a 4D Gaussian representation and a GPU-friendly rasterizer, enabling rapid rendering speeds. Despite its advantages, 4DGS faces significant challenges, n…

2025

MEGAD: A Memory-Efficient Framework for Large-Scale Attributed Graph Anomaly Detection

IJCAI 2025

Graph anomaly detection (GAD), with its ability to accurately identify anomalous patterns in graph data, plays a vital role in areas such as network security, social media platforms, and fraud detection. Graph autoencoder-based methods are widely used for GAD due to their efficiency and effectivenes

2025

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

ICML 2025poster

Existing efforts to align multimodal large language models (MLLMs) with human preferences have only achieved progress in narrow areas, such as hallucination reduction, but remain limited in practical applicability and generalizability. To this end, we introduce **MM-RLHF**, a dataset containing **12…

Cited by 13SourcePDFScholar
2025

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

ICLR 2025poster

Comprehensive evaluation of Multimodal Large Language Models (MLLMs) has recently garnered widespread attention in the research community. However, we observe that existing benchmarks present several common barriers that make it difficult to measure the significant challenges that models face in the…

Cited by 41SourcePDFScholar
2025

MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video conte…

Cited by 0SourceScholar
2025

MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models

EMNLP 2025

Memes have emerged as a popular form of multimodal online communication, where their interpretation heavily depends on the specific context in which they appear. Current approaches predominantly focus on isolated meme analysis, either for harmful content detection or standalone interpretation, overl

Cited by 0SourcePDFScholar
2025

MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining

NeurIPS 2025poster

Data quality is a critical driver of large language model performance, yet existing model-based selection methods focus almost exclusively on English, neglecting other languages that are essential in the training mix for multilingual LLMs. We introduce MuRating, a scalable framework that transfers h…

Cited by 0SourceScholar
2025

Poison-splat: Computation Cost Attack on 3D Gaussian Splatting

ICLR 2025spotlight

3D Gaussian splatting (3DGS), known for its groundbreaking performance and efficiency, has become a dominant 3D representation and brought progress to many 3D vision tasks. However, in this work, we reveal a significant security vulnerability that has been largely overlooked in 3DGS: the computation…

2025

Position: Trustworthy AI Agents Require the Integration of Large Language Models and Formal Methods

ICML 2025poster

Large Language Models (LLMs) have emerged as a transformative AI paradigm, profoundly influencing broad aspects of daily life. Despite their remarkable performance, LLMs exhibit a fundamental limitation: hallucination—the tendency to produce misleading outputs that appear plausible. This inherent…

Cited by 0SourcePDFScholar
2025

RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models

ACL 2025long

This work introduces RARE (Retrieval-Augmented Reasoning Enhancement), a versatile extension to the mutual reasoning framework (rStar), aimed at enhancing reasoning accuracy and factual integrity across large language models (LLMs) for complex, knowledge-intensive tasks such as medical and commonsen…

2025

SPO: Self Preference Optimization with Self Regularization

EMNLP 2025

Direct Preference Optimization (DPO) is a widely used offline preference optimization algorithm that enhances the simplicity and training stability of reinforcement learning through reward function reparameterization from PPO. Recently, SimPO (Simple Preference Optimization) and CPO (Contrastive Pre

Cited by 0SourcePDFScholar
2025

STRAP: Spatio-Temporal Pattern Retrieval for Out-of-Distribution Generalization

NeurIPS 2025poster

Spatio-Temporal Graph Neural Networks (STGNNs) have emerged as a powerful tool for modeling dynamic graph-structured data across diverse domains. However, they often fail to generalize in Spatio-Temporal Out-of-Distribution (STOOD) scenarios, where both temporal dynamics and spatial structures evolv…

Cited by 0SourceScholar
2025

Tactile-Guided Robotic Ultrasound: Mapping Preplanned Scan Paths for Intercostal Imaging

IROS 2025

Medical ultrasound (US) imaging is widely used in clinical examinations due to its portability, real-time capability, and radiation-free nature. To address inter- and intra-operator variability, robotic ultrasound systems have gained increasing attention. However, their application in challenging in

Cited by 1SourceScholar
2025

Tensor Product Attention Is All You Need

NeurIPS 2025spotlight

Scaling language models to handle longer input sequences typically necessitates large key-value (KV) caches, resulting in substantial memory overhead during inference. In this paper, we propose Tensor Product Attention (TPA), a novel attention mechanism that uses tensor decompositions to represent q…

Cited by 0SourcecodeScholar
2025

TopNet: Transformer-Efficient Occupancy Prediction Network for Octree-Structured Point Cloud Geometry Compression

CVPR 2025poster

Efficient Point Cloud Geometry Compression (PCGC) with a lower bits per point (BPP) and higher peak signal-to-noise ratio (PSNR) is essential for the transportation of large-scale 3D data. Although octree-based entropy models can reduce BPP without introducing geometry distortion, existing CNN-based…

2025

Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs

NeurIPS 2025poster

Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. However, engineering design poses a fundamentally different challenge for large language models (LLMs) compared with tradi…

Cited by 0SourceScholar
2025

Training LLMs for Optimization Modeling via Iterative Data Synthesis and Structured Validation

EMNLP 2025

Large Language Models (LLMs) have revolutionized various domains but encounter substantial challenges in tackling optimization modeling tasks for Operations Research (OR), particularly when dealing with complex problem. In this work, we propose Step-Opt-Instruct, a framework that augments existing d

2025

V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and Me

ACL 2025long

Oracle Bone Script (OBS) is a vital treasure of human civilization, rich in insights from ancient societies. However, the evolution of written language over millennia complicates its decipherment. In this paper, we propose V-Oracle, an innovative framework that utilizes Large Multi-modal Models (LMM…

Cited by 0SourcePDFScholar
2025

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

NeurIPS 2025spotlight

Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in multimodal dialogue systems, and implementing high-performance in bot…

Cited by 0SourcecodeScholar
2025

We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

ACL 2025long

Visual mathematical reasoning, as a fundamental visual reasoning ability, has received widespread attention from the Large Multimodal Models (LMMs) community. Existing benchmarks mainly focus more on the end-to-end performance, but neglect the underlying principles of knowledge acquisition and gener…

2025

Wheel-Legged SLAM: Indoor LiDAR-Inertial SLAM Integrating Kinematic Model of Wheel-Legged Robots

RA-L 2025

SLAM is the key technique for localization and surrounding perception in indoor environments. However, the dynamic posture adjustments of wheel-legged robots cast new challenges that affect the accuracy of localization. Therefore, this letter presents the Wheel-Legged SLAM, a novel indoor SLAM metho

Cited by 3SourceScholar
2025

e-GAI: e-value-based Generalized $\alpha$-Investing for Online False Discovery Rate Control

ICML 2025poster

Online multiple hypothesis testing has attracted a lot of attention in many applications, e.g., anomaly status detection and stock market price monitoring. The state-of-the-art generalized $\alpha$-investing (GAI) algorithms can control online false discovery rate (FDR) on p-values only under specif…

Cited by 0SourcePDFScholar
2024

A Study on the Calibration of In-context Learning

NAACL 2024long

Accurate uncertainty quantification is crucial for the safe deployment of machine learning models, and prior research has demonstrated improvements in the calibration of modern language models (LMs). We study in-context learning (ICL), a prevalent method for adapting static LMs through tailored prom…

2024

Contrastive Learning is Spectral Clustering on Similarity Graph

ICLR 2024poster

Contrastive learning is a powerful self-supervised learning method, but we have a limited theoretical understanding of how it works and why it works. In this paper, we prove that contrastive learning with the standard InfoNCE loss is equivalent to spectral clustering on the similarity graph. Using t…

2024

DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models

EMNLP 2024main

The performance of Large Language Models (LLMs) is substantially influenced by the pretraining corpus, which consists of vast quantities of unsupervised data processed by the models. Despite its critical role in model performance, ensuring the quality of this data is challenging due to its sheer vol…

2024

End-To-End Spatially-Constrained Multi-Perspective Fine-Grained Image Captioning

ICASSP 2024accepted

The perspective of captions in fine-grained image captioning crucially impacts people’s perception and understanding of the image. However, existing methods often overlook this aspect, resulting in captions that struggle to accurately convey the image’s hierarchical and spatial information. In this…

Cited by 0SourceScholar
2024

Evaluating Step-by-Step Reasoning through Symbolic Verification

NAACL 2024findings

Pre-trained language models (LMs) have shown remarkable reasoning performance using explanations or chain-of-thoughts (CoT)) for in-context learning. On the other hand, these reasoning tasks are usually presumed to be more approachable for symbolic programming. To understand the mechanism of reasoni…

2024

Fine-grained Image-to-LiDAR Contrastive Distillation with Visual Foundation Models

NeurIPS 2024poster

Contrastive image-to-LiDAR knowledge transfer, commonly used for learning 3D representations with synchronized images and point clouds, often faces a self-conflict dilemma. This issue arises as contrastive losses unintentionally dissociate features of unmatched points and pixels that share semantic…

2024

HGCN2SP: Hierarchical Graph Convolutional Network for Two-Stage Stochastic Programming

ICML 2024poster

Two-stage Stochastic Programming (2SP) is a standard framework for modeling decision-making problems under uncertainty. While numerous methods exist, solving such problems with many scenarios remains challenging. Selecting representative scenarios is a practical method for accelerating solutions. Ho…

Cited by 2SourcePDFScholar
2024

Implicit Concept Removal of Diffusion Models

ECCV 2024poster

"Text-to-image (T2I) diffusion models often inadvertently generate unwanted concepts such as watermarks and unsafe images. These concepts, termed “implicit concepts”, can be unintentionally learned during training and then be generated uncontrollably during inference. Existing removal methods still…

2024

Information Flow in Self-Supervised Learning

ICML 2024poster

In this paper, we conduct a comprehensive analysis of two dual-branch (Siamese architecture) self-supervised learning approaches, namely Barlow Twins and spectral contrastive learning, through the lens of matrix mutual information. We prove that the loss functions of these methods implicitly optimiz…

2024

Intrinsic Action Tendency Consistency for Cooperative Multi-Agent Reinforcement Learning

AAAI 2024technical

Efficient collaboration in the centralized training with decentralized execution (CTDE) paradigm remains a challenge in cooperative multi-agent systems. We identify divergent action tendencies among agents as a significant obstacle to CTDE's training efficiency, requiring a large number of training…

Cited by 3SourcePDFScholar
2024

MEEL: Multi-Modal Event Evolution Learning

ACL 2024findings

Multi-modal Event Reasoning (MMER) endeavors to endow machines with the ability to comprehend intricate event relations across diverse data modalities. MMER is fundamental and underlies a wide broad of applications. Despite extensive instruction fine-tuning, current multi-modal large language models…

2024

Matrix Information Theory for Self-Supervised Learning

ICML 2024poster

The maximum entropy encoding framework provides a unified perspective for many non-contrastive learning methods like SimSiam, Barlow Twins, and MEC. Inspired by this framework, we introduce Matrix-SSL, a novel approach that leverages matrix information theory to interpret the maximum entropy encodin…

Cited by 18SourcePDFScholar
2024

MonoTTA: Fully Test-Time Adaptation for Monocular 3D Object Detection

ECCV 2024poster

"Monocular 3D object detection (Mono 3Det) aims to identify 3D objects from a single RGB image. However, existing methods often assume training and test data follow the same distribution, which may not hold in real-world test scenarios. To address the out-of-distribution (OOD) problems, we explore a…

Cited by 3SourcePDFScholar
2024

One-Shot Diffusion Mimicker for Handwritten Text Generation

ECCV 2024poster

"Existing handwritten text generation methods often require more than ten handwriting samples as style references. However, in practical applications, users tend to prefer a handwriting generation model that operates with just a single reference sample for its convenience and efficiency. This approa…

2024

Position: What Can Large Language Models Tell Us about Time Series Analysis

ICML 2024poster

Time series analysis is essential for comprehending the complexities inherent in various real-world systems and applications. Although large language models (LLMs) have recently made significant strides, the development of artificial general intelligence (AGI) equipped with time series analysis capa…

Cited by 36SourcePDFScholar
2024

Segment Any Event Streams via Weighted Adaptation of Pivotal Tokens

CVPR 2024poster

In this paper we delve into the nuanced challenge of tailoring the Segment Anything Models (SAMs) for integration with event data with the overarching objective of attaining robust and universal object segmentation within the event-centric domain. One pivotal issue at the heart of this endeavor is t…

2024

Towards Automated Chinese Ancient Character Restoration: A Diffusion-Based Method with a New Dataset

AAAI 2024technical

Automated Chinese ancient character restoration (ACACR) remains a challenging task due to its historical significance and aesthetic complexity. Existing methods are constrained by non-professional masks and even overfitting when training on small-scale datasets, which hinder their interdisciplinary…

2023

AdaNPC: Exploring Non-Parametric Classifier for Test-Time Adaptation

ICML 2023poster

Many recent machine learning tasks focus to develop models that can generalize to unseen distributions. Domain generalization (DG) has become one of the key topics in various fields. Several literatures show that DG can be arbitrarily hard without exploiting target domain information. To address thi…

2023

Asynchronous Event Processing with Local-Shift Graph Convolutional Network

AAAI 2023technical

Event cameras are bio-inspired sensors that produce sparse and asynchronous event streams instead of frame-based images at a high-rate. Recent works utilizing graph convolutional networks (GCNs) have achieved remarkable performance in recognition tasks, which model event stream as spatio-temporal gr…

Cited by 2SourcePDFScholar
2023

Bidirectional Propagation for Cross-Modal 3D Object Detection

ICLR 2023poster

Recent works have revealed the superiority of feature-level fusion for cross-modal 3D object detection, where fine-grained feature propagation from 2D image pixels to 3D LiDAR points has been widely adopted for performance improvement. Still, the potential of heterogeneous feature propagation betwee…

Cited by 2SourcePDFScholar
2023

Disentangling Writer and Character Styles for Handwriting Generation

CVPR 2023poster

Training machines to synthesize diverse handwritings is an intriguing task. Recently, RNN-based methods have been proposed to generate stylized online Chinese characters. However, these methods mainly focus on capturing a person's overall writing style, neglecting subtle style inconsistencies betwee…

2023

Expanding Small-Scale Datasets with Guided Imagination

NeurIPS 2023poster

The power of DNNs relies heavily on the quantity and quality of training data. However, collecting and annotating data on a large scale is often expensive and time-consuming. To address this issue, we explore a new task, termed dataset expansion, aimed at expanding a ready-to-use small dataset by au…

2023

Free Lunch for Domain Adversarial Training: Environment Label Smoothing

ICLR 2023poster

A fundamental challenge for machine learning models is how to generalize learned models for out-of-distribution (OOD) data. Among various approaches, exploiting invariant features by Domain Adversarial Training (DAT) received widespread attention. Despite its success, we observe training instability…

2023

Knowledge Distillation with Active Exploration and Self-Attention Based Inter-Class Variation Transfer for Image Segmentation

ICASSP 2023accepted

Knowledge distillation (KD) aims to distill the knowledge from a more extensive deep neural network into a small net-work without losing validity. This paper proposes a novel approach with active exploration and passive transfer (AEPT) and self-attention-based inter-class feature variation (AIFV) di…

Cited by 0SourceScholar
2023

On the Data-Efficiency with Contrastive Image Transformation in Reinforcement Learning

ICLR 2023poster

Data-efficiency has always been an essential issue in pixel-based reinforcement learning (RL). As the agent not only learns decision-making but also meaningful representations from images. The line of reinforcement learning with data augmentation shows significant improvements in sample-efficiency.…

2023

On-the-Fly Adapting Code Summarization on Trainable Cost-Effective Language Models

NeurIPS 2023poster

Deep learning models are emerging to summarize source code to comment, facilitating tasks of code documentation and program comprehension. Scaled-up large language models trained on large open corpus have achieved good performance in such tasks. However, in practice, the subject code in one ce…

Cited by 9SourcePDFScholar
2023

OneNet: Enhancing Time Series Forecasting Models under Concept Drift by Online Ensembling

NeurIPS 2023poster

Online updating of time series forecasting models aims to address the concept drifting problem by efficiently updating forecasting models based on streaming data. Many algorithms are designed for online time series forecasting, with some exploiting cross-variable dependency while others assume indep…

2023

QD-BEV : Quantization-aware View-guided Distillation for Multi-view 3D Object Detection

ICCV 2023poster

Multi-view 3D detection based on BEV (bird-eye-view) has recently achieved significant improvements. However, the huge memory consumption of state-of-the-art models makes it hard to deploy them on vehicles, and the non-trivial latency will affect the real-time perception of streaming applications. D…

Cited by 11PDFScholar
2023

TOFG: A Unified and Fine-Grained Environment Representation in Autonomous Driving

ICRA 2023poster

In autonomous driving, an accurate understanding of environment, e.g., the vehicle-to-vehicle and vehicle-to-lane interactions, plays a critical role in many driving tasks such as trajectory prediction and motion planning. Environment information comes from high-definition (HD) map and historical tr…

Cited by 2SourceScholar
2023

Towards Stable Test-time Adaptation in Dynamic Wild World

ICLR 2023top-5%

Test-time adaptation (TTA) has shown to be effective at tackling distribution shifts between training and testing data by adapting a given model on test samples. However, the online model updating of TTA may be unstable and this is often a key obstacle preventing existing TTA methods from being depl…

2023

Trade-off Between Efficiency and Consistency for Removal-based Explanations

NeurIPS 2023poster

In the current landscape of explanation methodologies, most predominant approaches, such as SHAP and LIME, employ removal-based techniques to evaluate the impact of individual features by simulating various scenarios with specific features omitted. Nonetheless, these methods primarily emphasize effi…

2023

Unleash the Potential of Image Branch for Cross-modal 3D Object Detection

NeurIPS 2023poster

To achieve reliable and precise scene understanding, autonomous vehicles typically incorporate multiple sensing modalities to capitalize on their complementary attributes. However, existing cross-modal 3D detectors do not fully utilize the image domain information to address the bottleneck issues of…

2022

AutoMS: Automatic Model Selection for Novelty Detection with Error Rate Control

NeurIPS 2022accept

Given an unsupervised novelty detection task on a new dataset, how can we automatically select a ''best'' detection model while simultaneously controlling the error rate of the best model? For novelty detection analysis, numerous detectors have been proposed to detect outliers on a new unseen datase…

2022

Efficient Test-Time Model Adaptation without Forgetting

ICML 2022spotlight

Test-time adaptation provides an effective means of tackling the potential distribution shift between model training and inference, by dynamically updating the model at test time. This area has seen fast progress recently, at the effectiveness of handling test shifts. Nonetheless, prior methods stil…

2022

How Well Does Self-Supervised Pre-Training Perform with Streaming Data?

ICLR 2022poster

Prior works on self-supervised pre-training focus on the joint training scenario, where massive unlabeled data are assumed to be given as input all at once, and only then is a learner trained. Unfortunately, such a problem setting is often impractical if not infeasible since many real-world tasks re…

Cited by 39SourcePDFScholar
2022

MENet: A Memory-Based Network with Dual-Branch for Efficient Event Stream Processing

ECCV 2022poster

"Event cameras are bio-inspired sensors that asynchronously capture per-pixel brightness change and trigger a stream of events instead of frame-based images. Each event stream is generally split into multiple sliding windows for subsequent processing. However, most existing event-based methods ignor…

Cited by 1SourcePDFScholar
2022

Not All Points Are Equal: Learning Highly Efficient Point-Based Detectors for 3D LiDAR Point Clouds

CVPR 2022oral

We study the problem of efficient object detection of 3D LiDAR point clouds. To reduce the memory and computational cost, existing point-based pipelines usually adopt task-agnostic random sampling or farthest point sampling to progressively downsample input point clouds, despite the fact that not al…

Cited by 389PDFcodeScholar
2022

PKD: General Distillation Framework for Object Detectors via Pearson Correlation Coefficient

NeurIPS 2022accept

Knowledge distillation(KD) is a widely-used technique to train compact models in object detection. However, there is still a lack of study on how to distill between heterogeneous detectors. In this paper, we empirically find that better FPN features from a heterogeneous teacher detector can help the…

2022

Prototype-Guided Continual Adaptation for Class-Incremental Unsupervised Domain Adaptation

ECCV 2022poster

"This paper studies a new, practical but challenging problem, called Class-Incremental Unsupervised Domain Adaptation (CI-UDA), where the labeled source domain contains all classes, but the classes in the unlabeled target domain increase sequentially. This problem is challenging due to two difficult…

2022

Self-Supervised Aggregation of Diverse Experts for Test-Agnostic Long-Tailed Recognition

NeurIPS 2022accept

Existing long-tailed recognition methods, aiming to train class-balanced models from long-tailed data, generally assume the models would be evaluated on the uniform test class distribution. However, practical test class distributions often violate this assumption (e.g., being either long-tailed or e…

2022

Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive Learning

AAAI 2022technical

We study the unsupervised representation learning for the semantic segmentation task. Different from previous works that aim at providing unsupervised pre-trained backbones for segmentation models which need further supervised fine-tune, here, we focus on providing representation that is only traine…

Cited by 11SourcePDFScholar
2022

Unsupervised Visual Representation Learning by Synchronous Momentum Grouping

ECCV 2022poster

"In this paper, we propose a genuine group-level contrastive visual representation learning method whose linear evaluation performance on ImageNet surpasses the vanilla supervised learning. Two mainstream unsupervised learning schemes are the instance-level contrastive framework and clustering-based…

Cited by 36SourcePDFScholar
2021

AdaSGN: Adapting Joint Number and Model Size for Efficient Skeleton-Based Action Recognition

ICCV 2021poster

Existing methods for skeleton-based action recognition mainly focus on improving the recognition accuracy, whereas the efficiency of the model is rarely considered. Recently, there are some works trying to speed up the skeleton modeling by designing light-weight modules. However, in addition to the…

Cited by 67PDFcodeScholar
2021

AdaXpert: Adapting Neural Architecture for Growing Data

ICML 2021spotlight

In real-world applications, data often come in a growing manner, where the data volume and the number of classes may increase dynamically. This will bring a critical challenge for learning: given the increasing data volume or the number of classes, one has to instantaneously adjust the neural model…

2021

No Fear of Heterogeneity: Classifier Calibration for Federated Learning with Non-IID Data

NeurIPS 2021poster

A central challenge in training classification models in the real-world federated system is learning with non-IID data. To cope with this, most of the existing works involve enforcing regularization in local optimization or improving the model aggregation scheme at the server. Other works also share…

Cited by 423SourcePDFScholar
2021

Source-free Domain Adaptation via Avatar Prototype Generation and Adaptation

IJCAI 2021poster

We study a practical domain adaptation task, called source-free unsupervised domain adaptation (UDA) problem, in which we cannot access source domain data due to data privacy issues but only a pre-trained source model and unlabeled target data are available. This task, however, is very difficult du…

2021

StablePose: Learning 6D Object Poses From Geometrically Stable Patches

CVPR 2021poster

We introduce the concept of geometric stability to the problem of 6D object pose estimation and propose to learn pose inference based on geometrically stable patches extracted from observed 3D point clouds. According to the theory of geometric stability analysis, a minimal set of three planar/cylind…

Cited by 46PDFScholar
2021

Unleashing the Power of Contrastive Self-Supervised Visual Models via Contrast-Regularized Fine-Tuning

NeurIPS 2021poster

Contrastive self-supervised learning (CSL) has attracted increasing attention for model pre-training via unlabeled data. The resulted CSL models provide instance-discriminative visual features that are uniformly scattered in the feature space. During deployment, the common practice is to directly f…

2020

Decoupling GCN with DropGraph Module for Skeleton-Based Action Recognition

ECCV 2020poster

In skeleton-based action recognition, graph convolutional networks (GCNs) have achieved remarkable success. Nevertheless, how to efficiently model the spatial-temporal skeleton graph without introducing extra computation burden is a challenging problem for industrial deployment. In this paper, we re…

2020

Interpretable Complex-Valued Neural Networks for Privacy Protection

ICLR 2020poster

Previous studies have found that an adversary attacker can often infer unintended input information from intermediate-layer features. We study the possibility of preventing such adversarial inference, yet without too much accuracy degradation. We propose a generic method to revise the neural network…

Cited by 45SourceScholar
2020

Relation-Aware Transformer for Portfolio Policy Learning

IJCAI 2020poster

Portfolio selection is an important yet challenging task in AI for FinTech. One of the key issues is how to represent the non-stationary price series of assets in a portfolio, which is important for portfolio decisions. The existing methods, however, fall short of capturing: 1) the complicated seq…

2020

Skeleton-Based Action Recognition With Shift Graph Convolutional Network

CVPR 2020oral

Action recognition with skeleton data is attracting more attention in computer vision. Recently, graph convolutional networks (GCNs), which model the human body skeletons as spatiotemporal graphs, have obtained remarkable performance. However, the computational complexity of GCN-based methods are pr…

Cited by 1008PDFScholar
2020

TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training Model

CVPR 2020oral

Multi-object tracking is a fundamental vision problem that has been studied for a long time. As deep learning brings excellent performances to object detection algorithms, Tracking by Detection (TBD) has become the mainstream tracking framework. Despite the success of TBD, this two-step method is to…

Cited by 344PDFcodeScholar
2019

A Multimodal Soft Crawling-Climbing Robot with the Controllable Horizontal Plane to Slope Transition

IROS 2019poster

Most of the existing soft locomotive robots are capable of moving on horizontal planes and small-angled slopes, but few of them can accomplish the large-angled slope climbing or wall climbing. We introduce an inchworm inspired soft crawling-climbing robot capable of continuous motion from a horizont…

Cited by 24SourceScholar
2019

Multi-marginal Wasserstein GAN

NeurIPS 2019poster

Multiple marginal matching problem aims at learning mappings to match a source domain to multiple target domains and it has attracted great attention in many applications, such as multi-domain image translation. However, addressing this problem has two critical challenges: (i) Measuring the multi-ma…

2019

Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition

CVPR 2019poster

In skeleton-based action recognition, graph convolutional networks (GCNs), which model the human body skeletons as spatiotemporal graphs, have achieved remarkable performance. However, in existing GCN-based methods, the topology of the graph is set manually, and it is fixed over all layers and input…

Cited by 2123PDFcodeScholar
2018

Training Binary Weight Networks via Semi-Binary Decomposition

ECCV 2018poster

Recently binary weight networks have attracted lots of attentions due to their high computational efficiency and small parameter size. Yet they still suffer from large accuracy drops because of their limited representation capacity. In this paper, we propose a novel semi-binary decomposition method…

Cited by 23SourcePDFScholar
2018

Two-Step Quantization for Low-Bit Neural Networks

CVPR 2018poster

Every bit matters in the hardware design of quantized neural networks. However, extremely-low-bit representation usually causes large accuracy drop. Thus, how to train extremely-low-bit neural networks with high accuracy is of central importance. Most existing network quantization approaches learn t…

Cited by 167SourcePDFScholar
2017

Egocentric Gesture Recognition Using Recurrent 3D Convolutional Neural Networks With Spatiotemporal Transformer Modules

ICCV 2017spotlight

Gesture is a natural interface in interacting with wearable devices such as VR/AR helmet and glasses. The main challenge of gesture recognition in egocentric vision arises from the global camera motion caused by the spontaneous head movement of the device wearer. In this paper, we address the proble…

Cited by 122PDFScholar