← Search

Lei CHEN

136 accepted papers

2026

Beyond Model Base Retrieval: Weaving Knowledge to Master Fine-grained Neural Network Design

ICML 2026poster

Designing high-performance neural networks for new tasks requires balancing optimization quality with search efficiency. Current methods fail to achieve this balance: neural architectural search is computationally expensive, while model retrieval often yields suboptimal static checkpoints. To resolv…

Cited by 0SourceScholar
2026

Beyond Tokens: Enhancing RTL Quality Estimation via Structural Graph Learning

ICML 2026poster

Estimating the quality of register transfer level (RTL) designs is crucial in the electronic design automation (EDA) workflow, as it enables instant feedback on key performance metrics like area and delay without the need for time-consuming logic synthesis. While recent approaches have leveraged lar…

Cited by 0SourceScholar
2026

Block-wise Adaptive Caching for Accelerating Diffusion Policy

ICLR 2026poster

Diffusion Policy has demonstrated strong visuomotor modeling capabilities, but its high computational cost renders it impractical for real-time robotic control. Despite huge redundancy across repetitive denoising steps, existing diffusion acceleration techniques fail to generalize to Diffusion Polic…

Cited by 0SourcecodeScholar
2026

Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation

ICLR 2026poster

While reinforcement learning (RL) has proven highly effective for general reasoning in vision-language models, its application to tasks requiring deep understanding of information-rich images and structured output generation remains underexplored. Chart-to-code generation exemplifies this challenge,…

Cited by 0SourcecodeScholar
2026

CONTINUUM: Restoring the Contiguous Tensor Abstraction Efficiently for Dynamic AI Workloads via Hardware Virtualization

ICML 2026spotlight

Emerging LLM workloads demand extreme mem- ory agility. However, state-of-the-art inference systems (e.g., vLLM) rely on software-defined paging, which sacrifices the contiguous tensor abstraction. This rigid interface exposes fragmen- tation complexity to developers, imposing a se- vere engineering…

Cited by 0SourceScholar
2026

Context-Driven Incremental Compression for Multi-Turn Dialogue Generation

ICML 2026poster

Modern conversational agents condition on an ever-growing dialogue history at each turn, incurring redundant attention and encoding costs that grow with conversation length. Naive truncation or summarization degrades fidelity, while existing context compressors lack cross-turn memory sharing or revi…

Cited by 0SourceScholar
2026

Evolving Graph Structured Programs for Circuit Generation with Large Language Models

ICLR 2026poster

Logic synthesis (LS), which aims to generate a *compact* logic circuit graph with minimized size while *accurately* satisfying a given functionality, plays an important role in chip design. However, existing LS methods struggle to balance circuit structure compactness and functional accuracy, often…

Cited by 0SourceScholar
2026

Fast Low-light Enhancement and Deblurring for 3D Dark Scenes

ICASSP 2026poster

Novel view synthesis from low-light, noisy, and motion-blurred imagery remains a valuable and challenging task. Current volumetric rendering methods struggle with compound degradation, and sequential 2D preprocessing introduces artifacts due to interdependencies. In this work, we introduce FLED-GS,…

Cited by 0SourcePDFScholar
2026

IncreFA: Breaking the Static Wall of Generative Model Attribution

CVPR 2026

As AI generative models evolve at unprecedented speed, image attribution has become a moving target. New diffusion, adversarial and autoregressive generators appear almost monthly, making existing watermark, classifier and inversion methods obsolete upon release. The core problem lies not in model r

Cited by 0SourcecodeScholar
2026

KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image Representation

CVPR 2026

Cross-modal biomedical signals such as pathology and genomics can provide richer and more robust semantic guidance for medical image representation learning. However, the availability of such guidance remains limited, as privacy constraints and acquisition costs severely restrict access to medical i

Cited by 0SourceScholar
2026

Knowledge Graph Guided Heterogeneity-Informed Diffusion Model for Spatio-Temporal Generation

AAAI 2026technical

Spatio-temporal data generation aims to synthesize realistic urban data across graph nodes by learning spatial and temporal dependencies. This task plays a crucial role in urban planning by enabling the simulation of unobserved nodes. However, existing approaches face critical limitations that time

Cited by 0SourcePDFScholar
2026

Learning from Scoring Disagreements: Contrastive Error Mining for Efficient and Robust LLM-based Assessment

AAAI 2026technical

Automated grading of student responses still faces numerous challenges, particularly when dealing with complex and ambiguous answers. In particular, large models are prone to scoring bias when handling uncertain responses, and few-shot reasoning methods often lack stability, which limits their appli

Cited by 0SourcePDFScholar
2026

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds

ICLR 2026poster

Multimodal large language models are progressively advancing toward multimodal agents that can proactively execute tasks. Existing research on multimodal agents primarily targets either GUI or embodied scenarios, corresponding to interactions within 2D virtual world and 3D physical world, respective…

Cited by 0SourceScholar
2026

Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR

CVPR 2026

Reading text from images or scanned documents via OCR models has been a longstanding focus of researchers. Intuitively, text reading is perceived as a straightforward perceptual task, and existing work primarily focuses on constructing enriched data engineering to enhance SFT capabilities. In this w

Cited by 0SourcecodeScholar
2026

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning

CVPR 2026

The misuse of AI-driven video generation technologies has raised serious social concerns, highlighting the urgent need for reliable AI-generated video detectors. However, most existing methods are limited to binary classification and lack the necessary explanations for human interpretation. In this

Cited by 0SourcecodeScholar
2026

TOWARDS PRIVACY-PRESERVING FINE-GRAINED VISUAL CLASSIFICATION VIA HIERARCHICAL LEARNING FROM LABEL PROPORTIONS

ICASSP 2026poster

In recent years, Fine-Grained Visual Classification (FGVC) has achieved impressive recognition accuracy, despite minimal inter-class variations. However, existing methods heavily rely on instance-level labels, making them impractical in privacy-sensitive scenarios such as medical image analysis. Thi…

Cited by 0SourcePDFScholar
2026

Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models

AAAI 2026technical

Text-guided image inpainting aims to inpaint masked image regions based on a textual prompt while preserving the background. Although diffusion-based methods have become dominant, their property of modeling the entire image in latent space makes it challenging for the results to align well with prom

Cited by 0SourcePDFScholar
2026

Transformer-Based Hierarchical Reinforcement Learning for Sequential Decision-Making in Swarm Confrontation

ICRA 2026poster

Hierarchical Reinforcement Learning (HRL) is a potent paradigm for addressing long-horizon sequential decision-making in swarm confrontation. However, its strategic capabilities are often bottlenecked by high-level policies that struggle to reason over the dynamic, variable-sized observations of oth…

Cited by 0Scholar
2026

TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution

ICML 2026poster

Effectively scaling GUI automation is essential for computer-use agents (CUAs); however, existing work primarily focuses on scaling GUI grounding rather than the more crucial GUI planning, which requires more sophisticated data collection. In reality, the exploration process of a CUA across apps/des…

Cited by 0SourceScholar
2026

URICA: A Uniformity Region Affine Identifier Capture Algorithm for Arbitrary Region Retrieval in Pathology Images

CVPR 2026

Whole slide image (WSI) region retrieval remains an open challenge in computational pathology, as existing methods struggle to represent and preserve information of all possible regions. Current approaches that rely on fixed-size patches or slide-level retrieval are misaligned with real clinical wor

Cited by 0SourceScholar
2026

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection

CVPR 2026

In recent years, significant progress has been made in both image generation and generated image detection. Despite their rapid, yet largely independent, development, these two fields have evolved distinct architectural paradigms: the former predominantly relies on generative networks, while the lat

Cited by 0SourcecodeScholar
2026

UniRTL: Unifying Code and Graph for Robust RTL Representation Learning

ICML 2026poster

Developing effective representations for register transfer level (RTL) designs is crucial for accelerating the hardware design workflow. Existing approaches, however, typically rely on a single data modality, either the RTL code or its associated graph-based representation, limiting the expressivene…

Cited by 0SourceScholar
2026

VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution

ICLR 2026poster

Recent advancements in visual autoregressive models (VAR) have demonstrated their effectiveness in image generation, highlighting their potential for real-world image super-resolution (Real-ISR). However, adapting VAR for ISR presents critical challenges. The next-scale prediction mechanism, constra…

Cited by 0SourcecodeScholar
2026

VoG: Enhancing LLM Reasoning through Stepwise Verification on Knowledge Graphs

ICLR 2026poster

Large Language Models (LLMs) excel at various reasoning tasks but still encounter challenges such as hallucination and factual inconsistency in knowledge-intensive tasks, primarily due to a lack of external knowledge and factual verification. These challenges could be mitigated by leveraging knowled…

Cited by 0SourceScholar
2025

A Graph Enhanced Symbolic Discovery Framework For Efficient Logic Optimization

ICLR 2025poster

The efficiency of Logic Optimization (LO) has become one of the key bottlenecks in chip design. To prompt efficient LO, previous studies propose using a key scoring function to predict and prune a large number of ineffective nodes of the LO heuristics. However, the existing scoring functions struggl…

Cited by 0SourcePDFScholar
2025

A Selective Learning Method for Temporal Graph Continual Learning

ICML 2025poster

Node classification is a key task in temporal graph learning (TGL). Real-life temporal graphs often introduce new node classes over time, but existing TGL methods assume a fixed set of classes. This assumption brings limitations, as updating models with full data is costly, while focusing only on ne…

Cited by 0SourcePDFScholar
2025

AgentPose: Progressive Distribution Alignment via Feature Agent for Human Pose Distillation

ICASSP 2025accepted

Pose distillation is widely adopted to reduce model size in human pose estimation. However, existing methods primarily emphasize the transfer of teacher knowledge while often neglecting the performance degradation resulted from the curse of capacity gap between teacher and student. To address this i…

Cited by 0SourceScholar
2025

AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios

NAACL 2025long

Large language models (LLMs) are increasingly leveraged to empower autonomous agents to simulate human beings in various fields of behavioral research. However, evaluating their capacity to navigate complex social interactions remains a challenge. Previous studies face limitations due to insufficien…

2025

AttentionPredictor: Temporal Patterns Matter for KV Cache Compression

NeurIPS 2025poster

With the development of large language models (LLMs), efficient inference through Key-Value (KV) cache compression has attracted considerable attention, especially for long-context generation. To compress the KV cache, recent methods identify critical KV tokens through static modeling of attention s…

Cited by 0SourcecodeScholar
2025

Automate Strategy Finding with LLM in Quant Investment

EMNLP 2025

We present a novel three-stage framework leveraging Large Language Models (LLMs) within a risk-aware multi-agent system for automate strategy finding in quantitative finance. Our approach addresses the brittleness of traditional deep learning models in financial applications by: employing prompt-eng

Cited by 0SourcePDFScholar
2025

Bidirectional Task-Motion Planning Based on Hierarchical Reinforcement Learning for Strategic Confrontation

IROS 2025

In swarm robotics, confrontation scenarios, including strategic confrontations, require efficient decision– making that integrates discrete commands and continuous actions. Traditional task and motion planning methods separate decision–making into two layers, but their unidirectional structure fails

Cited by 0SourceScholar
2025

Circuit Transformer: A Transformer That Preserves Logical Equivalence

ICLR 2025poster

Implementing Boolean functions with circuits consisting of logic gates is fundamental in digital computer design. However, the implemented circuit must be exactly equivalent, which hinders generative neural approaches on this task due to their occasionally wrong predictions. In this study, we introd…

2025

Computing Circuits Optimization via Model-Based Circuit Genetic Evolution

ICLR 2025poster

Optimizing computing circuits such as multipliers and adders is a fundamental challenge in modern integrated circuit design. Recent efforts propose formulating this optimization problem as a reinforcement learning (RL) proxy task, offering a promising approach to search high-speed and area-efficient…

Cited by 4SourcePDFScholar
2025

D3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection

ICCV 2025poster

The emergence of visual autoregressive (AR) models has revolutionized image generation while presenting new challenges for synthetic image detection. Unlike previous GAN or diffusion-based methods, AR models generate images through discrete token prediction, exhibiting both marked improvements in im…

2025

Design, Modeling and Control of a Novel Jet Vectoring Backpack

RA-L 2025

This paper presents the design and development of a jet vectoring backpack (JVP) - a single-person vertical take-off and landing (VTOL) aircraft specifically designed for emergency and disaster relief. Unlike existing electrically powered personal flying devices, the JVP employs turbojet engines as

Cited by 3SourceScholar
2025

Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers

ICLR 2025poster

Large language models have been successful at tasks involving basic forms of in-context reasoning, such as generating coherent language, as well as storing vast amounts of knowledge. At the core of the Transformer architecture behind such models are feed-forward and attention layers, which are often…

Cited by 0SourcePDFScholar
2025

Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language Models

ACL 2025finding

Large Language Models (LLMs) have demonstrated impressive capabilities in natural language processing tasks, such as text generation and semantic understanding. However, their performance on numerical reasoning tasks, such as basic arithmetic, numerical retrieval, and magnitude comparison, remains s…

2025

FADE: Frequency-Aware Diffusion Model Factorization for Video Editing

CVPR 2025poster

Recent advancements in diffusion frameworks have significantly enhanced video editing, achieving high fidelity and strong alignment with textual prompts. However, conventional approaches using image diffusion models fall short in handling video dynamics, particularly for challenging temporal edits l…

2025

FuXi-Ocean: A Global Ocean Forecasting System with Sub-Daily Resolution

NeurIPS 2025oral

Accurate, high-resolution ocean forecasting is crucial for maritime operations and environmental monitoring. While traditional numerical models are capable of producing sub-daily, eddy-resolving forecasts, they are computationally intensive and face challenges in maintaining accuracy at fine spatial…

Cited by 0SourceScholar
2025

GoRA: Gradient-driven Adaptive Low Rank Adaptation

NeurIPS 2025poster

Low-Rank Adaptation (LoRA) is a crucial method for efficiently fine-tuning large language models (LLMs), with its effectiveness influenced by two key factors: rank selection and weight initialization. While numerous LoRA variants have been proposed to improve performance by addressing one of these a…

Cited by 0SourcecodeScholar
2025

InstaRevive: One-Step Image Enhancement via Dynamic Score Matching

ICLR 2025poster

Image enhancement finds wide-ranging applications in real-world scenarios due to complex environments and the inherent limitations of imaging devices. Recent diffusion-based methods yield promising outcomes but necessitate prolonged and computationally intensive iterative sampling. In response, we p…

Cited by 0SourcePDFScholar
2025

KAA: Kolmogorov-Arnold Attention for Enhancing Attentive Graph Neural Networks

ICLR 2025poster

Graph neural networks (GNNs) with attention mechanisms, often referred to as attentive GNNs, have emerged as a prominent paradigm in advanced GNN models in recent years. However, our understanding of the critical process of scoring neighbor nodes remains limited, leading to the underperformance of m…

2025

KERAG: Knowledge-Enhanced Retrieval-Augmented Generation for Advanced Question Answering

EMNLP 2025

Retrieval-Augmented Generation (RAG) mitigates hallucination in Large Language Models (LLMs) by incorporating external data, with Knowledge Graphs (KGs) offering crucial information for question answering. Traditional Knowledge Graph Question Answering (KGQA) methods rely on semantic parsing, which

2025

Leader360V: A Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment

NeurIPS 2025poster

360 video captures the complete surrounding scenes with the ultra-large field of view of 360x180. This makes 360 scene understanding tasks, *e.g.*, segmentation and tracking, crucial for appications, such as autonomous driving, robotics. With the recent emergence of foundation models, the community…

Cited by 0SourcecodeScholar
2025

N-ForGOT: Towards Not-forgetting and Generalization of Open Temporal Graph Learning

ICLR 2025poster

Temporal Graph Neural Networks (TGNNs) lay emphasis on capturing node interactions over time but often overlook evolution in node classes and dynamic data distributions triggered by the continuous emergence of new class labels, known as the open-set problem. This problem poses challenges for existin…

Cited by 0SourcePDFScholar
2025

Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers

CVPR 2025poster

Recent advancements in diffusion models, particularly the architectural transformation from UNet-based models to Diffusion Transformers (DiTs), significantly improve the quality and scalability of image and video generation. However, despite their impressive capabilities, the substantial computation…

2025

Robust Low-Light Human Pose Estimation through Illumination-Texture Modulation

ICASSP 2025accepted

As critical visual details become obscured, the low visibility and high ISO noise in extremely low-light images pose a significant challenge to human pose estimation. Current methods fail to provide high-quality representations due to reliance on pixel-level enhancements that compromise semantics an…

Cited by 4SourceScholar
2025

SpaceGNN: Multi-Space Graph Neural Network for Node Anomaly Detection with Extremely Limited Labels

ICLR 2025poster

Node Anomaly Detection (NAD) has gained significant attention in the deep learning community due to its diverse applications in real-world scenarios. Existing NAD methods primarily embed graphs within a single Euclidean space, while overlooking the potential of non-Euclidean spaces. Besides, to ad…

2025

Structuring Benchmark into Knowledge Graphs to Assist Large Language Models in Retrieving and Designing Models

ICLR 2025poster

In recent years, the design and transfer of neural network models have been widely studied due to their exceptional performance and capabilities. However, the complex nature of datasets and the vast architecture space pose significant challenges for both manual and automated algorithms in creating h…

Cited by 0SourcePDFScholar
2024

A Circuit Domain Generalization Framework for Efficient Logic Synthesis in Chip Design

ICML 2024spotlight

Logic Synthesis (LS) plays a vital role in chip design. A key task in LS is to simplify circuits---modeled by directed acyclic graphs (DAGs)---with functionality-equivalent transformations. To tackle this task, many LS heuristics apply transformations to subgraphs---rooted at each node on an input D…

2024

Benchmarking PtO and PnO Methods in the Predictive Combinatorial Optimization Regime

NeurIPS 2024poster

Predictive combinatorial optimization, where the parameters of combinatorial optimization (CO) are unknown at the decision-making time, is the precise modeling of many real-world applications, including energy cost-aware scheduling and budget allocation on advertising. Tackling such a problem usuall…

Cited by 1SourcecodeScholar
2024

CRAG - Comprehensive RAG Benchmark

NeurIPS 2024poster

Retrieval-Augmented Generation (RAG) has recently emerged as a promising solution to alleviate Large Language Model (LLM)’s deficiency in lack of knowledge. Existing RAG datasets, however, do not adequately represent the diverse and dynamic nature of real-world Question Answering (QA) tasks. To brid…

2024

Compositional Generalization for Multi-Label Text Classification: A Data-Augmentation Approach

AAAI 2024technical

Despite significant advancements in multi-label text classification, the ability of existing models to generalize to novel and seldom-encountered complex concepts, which are compositions of elementary ones, remains underexplored. This research addresses this gap. By creating unique data splits acros…

2024

Demonstration Data-Driven Parameter Adjustment for Trajectory Planning in Highly Constrained Environments

RA-L 2024

Trajectory planning in highly constrained environments is crucial for robotic navigation. Classical algorithms are widely used for their interpretability, generalization, and system robustness. However, these algorithms often require parameter retuning when adapting to new scenarios. To address this

Cited by 2SourceScholar
2024

EG-NAS: Neural Architecture Search with Fast Evolutionary Exploration

AAAI 2024technical

Differentiable Architecture Search (DARTS) has achieved a rapid search for excellent architectures by optimizing architecture parameters through gradient descent. However, this efficiency comes with a significant challenge: the risk of premature convergence to local optima, resulting in subpar perfo…

2024

Fine-Tuning Graph Neural Networks by Preserving Graph Generative Patterns

AAAI 2024technical

Recently, the paradigm of pre-training and fine-tuning graph neural networks has been intensively studied and applied in a wide range of graph mining tasks. Its success is generally attributed to the structural consistency between pre-training and downstream datasets, which, however, does not hold…

2024

Learning Multi-Scale Video-Text Correspondence for Weakly Supervised Temporal Article Gronding

AAAI 2024technical

Weakly Supervised temporal Article Grounding (WSAG) is a challenging and practical task in video understanding. Specifically, given a video and a relevant article, whose sentences are at different semantic scales, WSAG aims to localize corresponding video segments for all “groundable” sentences. Com…

Cited by 1SourcePDFScholar
2024

Narrative Action Evaluation with Prompt-Guided Multimodal Interaction

CVPR 2024poster

In this paper we investigate a new problem called narrative action evaluation (NAE). NAE aims to generate professional commentary that evaluates the execution of an action. Unlike traditional tasks such as score-based action quality assessment and video captioning involving superficial sentences NAE…

2024

R2AG: Incorporating Retrieval Information into Retrieval Augmented Generation

EMNLP 2024finding

Retrieval augmented generation (RAG) has been applied in many scenarios to augment large language models (LLMs) with external documents provided by retrievers. However, a semantic gap exists between LLMs and retrievers due to differences in their training objectives and architectures. This misalignm…

2024

Refining and Synthesis: A Simple yet Effective Data Augmentation Framework for Cross-Domain Aspect-based Sentiment Analysis

ACL 2024findings

Aspect-based Sentiment Analysis (ABSA) is extensively researched in the NLP community, yet related models face challenges due to data sparsity when shifting to a new domain. Hence, data augmentation for cross-domain ABSA has attracted increasing attention in recent years. However, two key points hav…

Cited by 2SourcePDFScholar
2024

Retrieval-based Disentangled Representation Learning with Natural Language Supervision

ICLR 2024spotlight

Disentangled representation learning remains challenging as the underlying factors of variation in the data do not naturally exist. The inherent complexity of real-world data makes it unfeasible to exhaustively enumerate and encapsulate all its variations within a finite set of factors. However, it…

Cited by 9SourcePDFScholar
2024

SMARTCAL: An Approach to Self-Aware Tool-Use Evaluation and Calibration

EMNLP 2024industry

The tool-use ability of Large Language Models (LLMs) has a profound impact on a wide range of applications. However, LLMs’ self-awareness and self-control capability in appropriately using tools remains understudied. The problem is consequential as it alarms a potential risk of degraded performance…

2024

Towards Fair Graph Federated Learning via Incentive Mechanisms

AAAI 2024technical

Graph federated learning (FL) has emerged as a pivotal paradigm enabling multiple agents to collaboratively train a graph model while preserving local data privacy. Yet, current efforts overlook a key issue: agents are self-interested and would hesitant to share data without fair and satisfactory i…

2024

Towards Next-Generation Logic Synthesis: A Scalable Neural Circuit Generation Framework

NeurIPS 2024poster

Logic Synthesis (LS) aims to generate an optimized logic circuit satisfying a given functionality, which generally consists of circuit translation and optimization. It is a challenging and fundamental combinatorial optimization problem in integrated circuit design. Traditional LS approaches rely on…

Cited by 5SourcePDFScholar
2024

What Factors Influence LLMs’ Judgments? A Case Study on Question Answering

COLING 2024main

Large Language Models (LLMs) are now being considered as judges of high efficiency to evaluate the quality of answers generated by candidate models. However, their judgments may be influenced by complex scenarios and inherent biases, raising concerns about their reliability. This study aims to bridg…

Cited by 3SourcePDFScholar
2023

A Simple and Effective Framework for Strict Zero-Shot Hierarchical Classification

ACL 2023short

In recent years, large language models (LLMs) have achieved strong performance on benchmark tasks, especially in zero or few-shot settings. However, these benchmarks often do not adequately address the challenges posed in the real-world, such as that of hierarchical classification. In order to addre…

Cited by 11SourcePDFScholar
2023

BHE-DARTS: Bilevel Optimization Based on Hypergradient Estimation for Differentiable Architecture Search

ICASSP 2023accepted

In this paper, we propose a stochastic bilevel optimization approach based on a hypergradient estimator, called BHE- DARTS, as a remedy for this issue that it is easy to search for locally optimal structures rather than globally optimal ones in Differentiable Architecture Search (DARTS) bilevel opti…

Cited by 0SourceScholar
2023

CAT: A Contextualized Conceptualization and Instantiation Framework for Commonsense Reasoning

ACL 2023long

Commonsense reasoning, aiming at endowing machines with a human-like ability to make situational presumptions, is extremely challenging to generalize. For someone who barely knows about “meditation,” while is knowledgeable about “singing,” he can still infer that “meditation makes people relaxed” fr…

2023

Contrastive Domain Adaptation Via Delimitation Discriminator

ICASSP 2023accepted

Unsupervised domain adaptation aims to transfer the knowledge learned from the labeled source domain to the unlabeled target domain, thereby improving the classification performance of the target domain. Recent methods use contrastive learning to optimize this task, however, these methods only focus…

Cited by 5SourceScholar
2023

Disentangling Cognitive Diagnosis with Limited Exercise Labels

NeurIPS 2023poster

Cognitive diagnosis is an important task in intelligence education, which aims at measuring students’ proficiency in specific knowledge concepts. Given a fully labeled exercise-concept matrix, most existing models focused on mining students' response records for cognitive diagnosis. Despite their su…

2023

Efficient Video Action Detection with Token Dropout and Context Refinement

ICCV 2023poster

Streaming video clips with large-scale video tokens impede vision transformers (ViTs) for efficient recognition, especially in video action detection where sufficient spatiotemporal representations are required for precise actor identification. In this work, we propose an end-to-end framework for ef…

Cited by 26PDFcodeScholar
2023

Flora: Dual-Frequency LOss-Compensated ReAl-Time Monocular 3D Video Reconstruction

AAAI 2023technical

In this work, we propose a real-time monocular 3D video reconstruction approach named Flora for reconstructing delicate and complete 3D scenes from RGB video sequences in an end-to-end manner. Specifically, we introduce a novel method with two main contributions. Firstly, the proposed feature aggreg…

2023

Joint Feature and Differentiable $ k $-NN Graph Learning using Dirichlet Energy

NeurIPS 2023poster

Feature selection (FS) plays an important role in machine learning, which extracts important features and accelerates the learning process. In this paper, we propose a deep FS method that simultaneously conducts feature selection and differentiable $ k $-NN graph learning based on the Dirichlet Ene…

Cited by 4SourcePDFScholar
2023

Learning to Describe for Predicting Zero-shot Drug-Drug Interactions

EMNLP 2023long main

Adverse drug-drug interactions (DDIs) can compromise the effectiveness of concurrent drug administration, posing a significant challenge in healthcare. As the development of new drugs continues, the potential for unknown adverse effects resulting from DDIs becomes a growing concern. Traditional…

Cited by 0SourcecodeScholar
2023

Loan Fraud Users Detection in Online Lending Leveraging Multiple Data Views

AAAI 2023technical

In recent years, online lending platforms have been becoming attractive for micro-financing and popular in financial industries. However, such online lending platforms face a high risk of failure due to the lack of expertise on borrowers' creditworthness. Thus, risk forecasting is important to avoid…

Cited by 5SourcePDFScholar
2023

Noise2Info: Noisy Image to Information of Noise for Self-Supervised Image Denoising

ICCV 2023accepted

Unsupervised image denoising has been proposed to alleviate the widespread noise problem without requiring clean images. Existing works mainly follow the self-supervised way, which tries to reconstruct each pixel x of noisy images without the knowledge of x. More recently, some pioneer works further…

2023

Sancus: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural Networks (Extended Abstract)

IJCAI 2023poster

Graph neural networks (GNNs) have emerged due to their success at modeling graph data. Yet, it is challenging for GNNs to efficiently scale to large graphs. Thus, distributed GNNs come into play. To avoid communication caused by expensive data movement between workers, we propose SANCUS, a staleness…

Cited by 79SourcePDFScholar
2023

Skip-Plan: Procedure Planning in Instructional Videos via Condensed Action Space Learning

ICCV 2023poster

In this paper, we propose Skip-Plan, a condensed action space learning method for procedure planning in instructional videos. Current procedure planning methods all stick to the state-action pair prediction at every timestep and generate actions adjacently. Although it coincides with human intuition…

Cited by 12PDFScholar
2023

SwinRDM: Integrate SwinRNN with Diffusion Model towards High-Resolution and High-Quality Weather Forecasting

AAAI 2023technical

Data-driven medium-range weather forecasting has attracted much attention in recent years. However, the forecasting accuracy at high resolution is unsatisfactory currently. Pursuing high-resolution and high-quality weather forecasting, we develop a data-driven model SwinRDM which integrates an impro…

Cited by 55SourcePDFScholar
2023

Towards a Unified Conversational Recommendation System: Multi-task Learning via Contextualized Knowledge Distillation

EMNLP 2023long main

In Conversational Recommendation System (CRS), an agent is asked to recommend a set of items to users within natural language conversations. To address the need for both conversational capability and personalized recommendations, prior works have utilized separate recommendation and dialogue modules…

Cited by 0SourcecodeScholar
2023

Universal Prompt Tuning for Graph Neural Networks

NeurIPS 2023poster

In recent years, prompt tuning has sparked a research surge in adapting pre-trained models. Unlike the unified pre-training strategy employed in the language field, the graph field exhibits diverse pre-training strategies, posing challenges in designing appropriate prompt-based tuning methods for gr…

2023

Weighted Contrastive Learning With False Negative Control to Help Long-tailed Product Classification

ACL 2023industry

Item categorization (IC) aims to classify product descriptions into leaf nodes in a categorical taxonomy, which is a key technology used in a wide range of applications. Along with the fact that most datasets often has a long-tailed distribution, classification performances on tail labels tend to be…

Cited by 3SourcePDFScholar
2022

A Progressive Framework for Role-Aware Rumor Resolution

COLING 2022main

Existing works on rumor resolution have shown great potential in recognizing word appearance and user participation. However, they ignore the intrinsic propagation mechanisms of rumors and present poor adaptive ability when unprecedented news emerges. To exploit the fine-grained rumor diffusion patt…

2022

A Universal PINNs Method for Solving Partial Differential Equations with a Point Source

IJCAI 2022poster

In recent years, deep learning technology has been used to solve partial differential equations (PDEs), among which the physics-informed neural networks (PINNs)method emerges to be a promising method for solving both forward and inverse PDE problems. PDEs with a point source that is expressed as a D…

Cited by 12SourcePDFScholar
2022

Beyond Homophily: Structure-aware Path Aggregation Graph Neural Network

IJCAI 2022poster

Graph neural networks (GNNs) have been intensively studied in various real-world tasks. However, the homophily assumption of GNNs' aggregation function limits their representation learning ability in heterophily graphs. In this paper, we shed light on the path level patterns in graphs that can exp…

2022

Bridge-Prompt: Towards Ordinal Action Understanding in Instructional Videos

CVPR 2022poster

Action recognition models have shown a promising capability to classify human actions in short video clips. In a real scenario, multiple correlated human actions commonly occur in particular orders, forming semantically meaningful human activities. Conventional action recognition approaches focus on…

Cited by 88PDFcodeScholar
2022

DGraph: A Large-Scale Financial Dataset for Graph Anomaly Detection

NeurIPS 2022accept

Graph Anomaly Detection (GAD) has recently become a hot research spot due to its practicability and theoretical value. Since GAD emphasizes the application and the rarity of anomalous samples, enriching the varieties of its datasets is fundamental. Thus, this paper present DGraph, a real-world dynam…

Cited by 97SourcePDFScholar
2022

Developing Prefix-Tuning Models for Hierarchical Text Classification

EMNLP 2022industry

Hierarchical text classification (HTC) is a key problem and task in many industrial applications, which aims to predict labels organized in a hierarchy for given input text. For example, HTC can group the descriptions of online products into a taxonomy or organizing customer reviews into a hierarchy…

2022

Hybrid Weighting Loss for Precipitation Nowcasting from Radar Images

ICASSP 2022accepted

Precipitation nowcasting is gaining increasing attention in the signal processing community. Existing deep learning-based studies focus on designing an effective model architecture, neglecting the influence of the severe imbalanced distribution of rainfall data that can compromise the predictive acc…

Cited by 8SourceScholar
2022

Hyperlink-induced Pre-training for Passage Retrieval in Open-domain Question Answering

ACL 2022long

To alleviate the data scarcity problem in training question answering systems, recent works propose additional intermediate pre-training for dense passage retrieval (DPR). However, there still remains a large discrepancy between the provided upstream signals and the downstream question-passage relev…

2022

IS-MVSNet: Importance Sampling-Based MVSNet

ECCV 2022poster

"This paper presents a novel coarse-to-fine multi-view stereo (MVS) algorithm called importance-sampling-based MVSNet (IS-MVSNet) to address a crucial problem of limited depth resolution adopted by current learning-based MVS methods. We proposed an importance-sampling module for sampling candidate d…

2022

Meta-Auto-Decoder for Solving Parametric Partial Differential Equations

NeurIPS 2022accept

Many important problems in science and engineering require solving the so-called parametric partial differential equations (PDEs), i.e., PDEs with different physical parameters, boundary conditions, shapes of computation domains, etc. Recently, building learning-based numerical solvers for parametr…

Cited by 44SourcePDFScholar
2022

Tenrec: A Large-scale Multipurpose Benchmark Dataset for Recommender Systems

NeurIPS 2022accept

Existing benchmark datasets for recommender systems (RS) either are created at a small scale or involve very limited forms of user feedback. RS models evaluated on such datasets often lack practical values for large-scale real-world applications. In this paper, we describe Tenrec, a novel and publ…

2022

Uncertainty-Aware Representation Learning for Action Segmentation

IJCAI 2022poster

In this paper, we propose an uncertainty-aware representation Learning (UARL) method for action segmentation. Most existing action segmentation methods exploit continuity information of the action period to predict frame-level labels, which ignores the temporal ambiguity of the transition region bet…

Cited by 17SourcePDFScholar
2021

A User-Adaptive Layer Selection Framework for Very Deep Sequential Recommender Models

AAAI 2021technical

Sequential recommender systems (SRS) have become a research hotspot in recent studies. Because of the requirement in capturing user's dynamic interests, sequential neural network based recommender models often need to be stacked with more hidden layers (e.g., up to 100 layers) compared with standard…

Cited by 12SourcePDFScholar
2021

Align Voting Behavior with Public Statements for Legislator Representation Learning

ACL 2021long

Ideology of legislators is typically estimated by ideal point models from historical records of votes. It represents legislators and legislation as points in a latent space and shows promising results for modeling voting behavior. However, it fails to capture more specific attitudes of legislators t…

2021

AutoGEL: An Automated Graph Neural Network with Explicit Link Information

NeurIPS 2021poster

Recently, Graph Neural Networks (GNNs) have gained popularity in a variety of real-world scenarios. Despite the great success, the architecture design of GNNs heavily relies on manual labor. Thus, automated graph neural network (AutoGNN) has attracted interest and attention from the research communi…

2021

EPP-MVSNet: Epipolar-Assembling Based Depth Prediction for Multi-View Stereo

ICCV 2021poster

In this paper, we proposed EPP-MVSNet, a novel deep learning network for 3D reconstruction from multi-view stereo (MVS). EPP-MVSNet can accurately aggregate features at high resolution to a limited cost volume with an optimal depth range, thus, leads to effective and efficient 3D construction. Disti…

Cited by 137PDFScholar
2021

MixSeq: Connecting Macroscopic Time Series Forecasting with Microscopic Time Series Data

NeurIPS 2021poster

Time series forecasting is widely used in business intelligence, e.g., forecast stock market price, sales, and help the analysis of data trend. Most time series of interest are macroscopic time series that are aggregated from microscopic data. However, instead of directly modeling the macroscopic ti…

Cited by 21SourcePDFScholar
2021

MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions

ICCV 2021poster

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions. This paper aims to present a new multi-person dataset of spat…

Cited by 124PDFcodeScholar
2021

THOR, Trace-based Hardware-driven Layer-Oriented Natural Gradient Descent Computation

AAAI 2021technical

It is well-known that second-order optimizer can accelerate the training of deep neural networks, however, the huge computation cost of second-order optimization makes it impractical to apply in real practice. In order to reduce the cost, many methods have been proposed to approximate a second-order…

Cited by 9SourcePDFScholar
2020

Globally optimal consensus maximization for robust visual inertial localization in point and line map

IROS 2020poster

Map based visual inertial localization is a crucial step to reduce the drift in state estimation of mobile robots. The underlying problem for localization is to estimate the pose from a set of 3D-2D feature correspondences, of which the main challenge is the presence of outliers, especially in chang…

Cited by 6SourceScholar
2020

Interstellar: Searching Recurrent Architecture for Knowledge Graph Embedding

NeurIPS 2020spotlight

Knowledge graph (KG) embedding is well-known in learning representations of KGs. Many models have been proposed to learn the interactions between entities and relations of the triplets. However, long-term information among multiple triplets is also important to KG. In this work, based on the relatio…

2020

Modeling Evolution of Message Interaction for Rumor Resolution

COLING 2020main

Previous work for rumor resolution concentrates on exploiting time-series characteristics or modeling topology structure separately. However, how local interactive pattern affects global information assemblage has not been explored. In this paper, we attempt to address the problem by learning evolut…

2020

PG-GSQL: Pointer-Generator Network with Guide Decoding for Cross-Domain Context-Dependent Text-to-SQL Generation

COLING 2020main

Text-to-SQL is a task of translating utterances to SQL queries, and most existing neural approaches of text-to-SQL focus on the cross-domain context-independent generation task. We pay close attention to the cross-domain context-dependent text-to-SQL generation task, which requires a model to depend…

2020

Piggyback GAN: Efficient Lifelong Learning for Image Conditioned Generation

ECCV 2020poster

Humans accumulate knowledge in a lifelong fashion. Modern deep neural networks, on the other hand, are susceptible to catastrophic forgetting: when adapted to perform new tasks, they often fail to preserve their performance on previously learned tasks. Given a sequence of tasks, a naive approach add…

Cited by 47SourcePDFScholar
2020

SPL-MLL: Selecting Predictable Landmarks for Multi-Label Learning

ECCV 2020poster

Although significant progress achieved, multi-label classification is still challenging due to the complexity of correlations among different labels. Furthermore, modeling the relationships between input and some (dull) classes further increases the difficulty of accurately predicting all possible l…

Cited by 12SourcePDFScholar
2020

Simultaneous Arrival Matching for New Spatial Crowdsourcing Platforms

IJCAI 2020poster

In recent years, 3D spatial crowdsourcing platforms become popular, in which users and workers travel together to their assigned workplaces for services, such as InterestingSport and Nanguache. A typical problem over 3D spatial crowdsourcing platforms is to match users with suitable workers and work…

Cited by 0SourcePDFScholar
2019

2-Entity RANSAC for robust visual localization in changing environment

IROS 2019poster

Visual localization has attracted considerable attention due to its low-cost and stable sensor, which is desired in many applications, such as autonomous driving, inspection robots and unmanned aerial vehicles. However, current visual localization methods still struggle with environmental changes ac…

Cited by 11SourceScholar
2019

L2 Learners' Emotion Production in Video Dubbing Practices

ICASSP 2019accepted

Video dubbing is a new type of language learning practice. Because of the fun it brings into learning, video dubbing mobile applications have become quite popular. During video dubbing, learners not only mimic characters' pronunciations but also other voicing characteristics, e.g., emotions. In this…

Cited by 4SourceScholar
2019

Lifelong GAN: Continual Learning for Conditional Image Generation

ICCV 2019poster

Lifelong learning is challenging for deep neural networks due to their susceptibility to catastrophic forgetting. Catastrophic forgetting occurs when a trained network is not able to maintain its ability to accomplish previously learned tasks when it is trained to perform new tasks. We study the pro…

Cited by 250PDFScholar
2019

On the equivalence between graph isomorphism testing and function approximation with GNNs

NeurIPS 2019poster

Graph neural networks (GNNs) have achieved lots of success on graph-structured data. In light of this, there has been increasing interest in studying their representation power. One line of work focuses on the universal approximation of permutation-invariant functions by certain classes of GNNs, and…

2016

Exploring deep learning architectures for automatically grading non-native spontaneous speech

ICASSP 2016accepted

We investigate two deep learning architectures reported to have superior performance in ASR over the conventional GMM system, with respect to automatic speech scoring. We use an approximately 800-hour large-vocabulary non-native spontaneous English corpus to build three ASR systems. One system is in…

Cited by 0SourceScholar