← Search

Wei Wu

170 accepted papers

2026

Diffusion Guided Chain-of-Vision for Large Autoregressive Vision Models

CVPR 2026

Chain-of-Thought (CoT) has recently shown encouraging progress in the vision language model. However, the pure-vision CoT (i.e., chain-of-vision) has been underexplored in visual in-context learning. In this paper, we introduce Diffusion Guided Chain-of-Vision, which integrates an explicit chain-of-

Cited by 0SourcecodeScholar
2026

DriveMamba: Task-Centric Scalable State Space Model for Efficient End-to-End Autonomous Driving

ICLR 2026poster

Recent advances towards End-to-End Autonomous Driving (E2E-AD) focus on integrating modular designs into a unified framework for joint optimization. Most of these advances follow a sequential paradigm (i.e., perception-prediction-planning) based on separable Transformer decoders and rely on dense BE…

Cited by 0SourceScholar
2026

EgoFSD: Ego-Centric Fully Sparse Paradigm with Uncertainty Denoising and Iterative Refinement for End-To-End Self-Driving

ICRA 2026poster

Current End-to-End Autonomous Driving (E2E-AD) methods resort to unifying modular designs for various tasks (e.g. perception, prediction and planning). Although optimized with a fully differentiable framework in a planning-oriented manner, existing end-to-end driving systems lacking ego-centric desi…

Cited by 0Scholar
2026

Explicit Modeling of Causal Factors and Confounders for Image Classification

AAAI 2026technical

Causal inference has emerged as a promising approach for identifying decisive semantic factors and eliminating spurious correlations in visual representation learning. However, most existing methods rely on latent, data-driven confounder modeling, normally attributing the source of bias to backgroun

Cited by 0SourcePDFScholar
2026

Introducing Decomposed Causality with Spatiotemporal Object-Centric Representation for Video Classification

AAAI 2026technical

Video classification requires event-level representations of objects and their interactions. Existing methods typically rely on data-driven approaches, which either learn such features from whole frames or object-centric visual regions. Therefore, the modeling of spatiotemporal interactions among ob

Cited by 0SourcePDFScholar
2026

MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning

AAAI 2026technical

Outcome-based reinforcement learning has made notable advances in training language models (LMs) for reasoning. However, without explicit incentives and controls, this paradigm has limitations and instability in eliciting high-quality reasoning trajectories with diverse actions—particularly for mode

Cited by 0SourcePDFScholar
2026

ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding

ICLR 2026poster

Autoregressive models (ARMs) are hindered by slow sequential inference. While masked diffusion models (MDMs) offer a parallel alternative, they suffer from critical drawbacks: high computational overhead from precluding Key-Value (KV) caching, and incoherent generation arising from learning dependen…

Cited by 0SourcecodeScholar
2026

SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

ICML 2026poster

Reinforcement learning (RL) has become a key paradigm for training software engineering (SWE) agents, yet its practical accessibility and scalability is often constrained by container-based execution frameworks used for environment isolation. As the number of task instances increases, pre-cached con…

Cited by 0SourceScholar
2026

Scaling Prompt Synthesis for Large Language Model Reasoning

ICML 2026poster

Large language models (LLMs) are evolving from conversational systems into strong reasoners for tasks such as Olympiad mathematics and competitive programming. While scaling parameters and test-time computation has driven progress, a key bottleneck is the lack of high-quality training problems: huma…

Cited by 0SourceScholar
2026

Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models

ICLR 2026poster

Effectively processing long contexts is a critical challenge for language models. While standard Transformers are limited by quadratic complexity and poor length extrapolation, alternative architectures like sliding window attention and state space models sacrifice the ability to effectively utilize…

Cited by 0SourcecodeScholar
2026

Virne: A Comprehensive Benchmark for RL-based Network Resource Allocation in NFV

ICLR 2026poster

Resource allocation (RA) is critical to efficient service deployment in Network Function Virtualization (NFV), a transformative networking paradigm. This task is termed NFV-RA. Recently, deep Reinforcement Learning (RL)-based methods have been showing promising potential to address this combinatoria…

Cited by 0SourcecodeScholar
2025

A Survey on Personalized Alignment—The Missing Piece for Large Language Models in Real-World Applications

ACL 2025finding

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their transition to real-world applications reveals a critical limitation: the inability to adapt to individual preferences while maintaining alignment with universal human values. Current alignment techniques adopt a one-si…

Cited by 0SourcePDFScholar
2025

BLEND: Behavior-guided Neural Population Dynamics Modeling via Privileged Knowledge Distillation

ICLR 2025poster

Modeling the nonlinear dynamics of neuronal populations represents a key pursuit in computational neuroscience. Recent research has increasingly focused on jointly modeling neural activity and behavior to unravel their interconnections. Despite significant efforts, these approaches often necessitate…

2025

Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning

CVPR 2025poster

Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have developed benchmarks specifically tailored for detailed captions to measure their accuracy and comprehensiveness. In this p…

2025

Beyond Node-Centric Modeling: Sketching Signed Networks with Simplicial Complexes

NeurIPS 2025poster

Signed networks can reflect more complex connections through positive and negative edges, and cost-effective signed network sketching can significantly benefit an important link sign prediction task in the era of big data. Existing signed network embedding algorithms mainly learn node representation…

Cited by 0SourceScholar
2025

Beyond Online Sampling: Bridging Offline-to-Online Alignment via Dynamic Data Transformation for LLMs

EMNLP 2025

While Direct Preference Optimization (DPO) eliminates complex reward modeling in aligning large language models (LLMs) with human preferences, its online variant faces significant efficiency bottlenecks due to costly real-time preference sampling and the reward model annotation. We propose a novel f

2025

Causal Inference over Visual-Semantic-Aligned Graph for Image Classification

AAAI 2025technical

Incorporating tagging information to regularize the representation learning of images usually leads to improved performance in image classification by aligning the visual features with the textual ones of higher discriminative power. Existing methods typically follow the predictive approach, which u…

Cited by 0SourcePDFScholar
2025

CodePlan: Unlocking Reasoning Potential in Large Language Models by Scaling Code-form Planning

ICLR 2025poster

Despite the remarkable success of large language models (LLMs) on traditional natural language processing tasks, their planning ability remains a critical bottleneck in tackling complex multi-step reasoning tasks. Existing approaches mainly rely on prompting or task-specific fine-tuning, often suffe…

Cited by 3SourcePDFScholar
2025

DriveScape: High-Resolution Driving Video Generation by Multi-View Feature Fusion

CVPR 2025poster

Recent advancements in generative models offer promising solutions for synthesizing realistic driving videos, aiding in training autonomous driving perception models. However, existing methods often struggle with high-resolution multi-view generation, mainly due to the significant memory and computa…

Cited by 0SourcePDFScholar
2025

DynaAct: Large Language Model Reasoning with Dynamic Action Spaces

NeurIPS 2025poster

In modern sequential decision-making systems, the construction of an optimal candidate action space is critical to efficient inference. However, existing approaches either rely on manually defined action spaces that lack scalability or utilize unstructured spaces that render exhaustive search comput…

Cited by 0SourcecodeScholar
2025

Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling

ICML 2025poster

Despite the success of Transformers, handling longer contexts remains challenging due to the limited length generalization and quadratic complexity of self-attention, which often requires post-training with a larger attention window, significantly increasing computational and memory costs. In this p…

Cited by 0SourcePDFScholar
2025

Empowering Vision Transformers with Multi-Scale Causal Intervention for Long-Tailed Image Classification

IJCAI 2025

Causal inference has emerged as a promising approach to mitigate long-tail classification by handling the biases introduced by class imbalance. However, along with the change of advanced backbone models from Convolutional Neural Networks (CNNs) to Visual Transformers (ViT), existing causal models ma

Cited by 0SourcePDFScholar
2025

FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection Transformers

ICCV 2025poster

Detecting 3D objects accurately from multi-view 2D images is a challenging yet essential task in the field of autonomous driving. Current methods resort to integrating depth prediction to recover the spatial information for object query decoding, which necessitates explicit supervision from LiDAR po…

Cited by 0SourcePDFScholar
2025

Generalized and Invariant Single-Neuron In-Vivo Activity Representation Learning

NeurIPS 2025poster

In computational neuroscience, models representing single-neuron in-vivo activity have become essential for understanding the functional identities of individual neurons. These models, such as implicit representation methods based on Transformer architectures, contrastive learning frameworks, and va…

Cited by 0SourceScholar
2025

GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-based Transformer

ICCV 2025poster

Lidar-based 3D detection is one of the most popular research fields in autonomous driving. 3D detectors typically detect specific targets in a scene according to the pattern formed by the spatial distribution of point clouds. However, existing voxel-based methods usually adopt MLP and global pooling…

2025

Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory Access

NeurIPS 2025poster

A key advantage of Recurrent Neural Networks (RNNs) over Transformers is their linear computational and space complexity enables faster training and inference for long sequences. However, RNNs are fundamentally unable to randomly access historical context, and simply integrating attention mechanisms…

Cited by 0SourceScholar
2025

InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation

ICCV 2025poster

Autonomous driving relies on robust models trained on high-quality, large-scale multi-view driving videos for tasks like perception and planning. While world models offer a cost-effective solution for generating realistic driving videos, they struggle to maintain instance-level temporal consistency…

Cited by 0SourcePDFScholar
2025

KARMA: Leveraging Multi-Agent LLMs for Automated Knowledge Graph Enrichment

NeurIPS 2025spotlight

Maintaining comprehensive and up-to-date knowledge graphs (KGs) is critical for modern AI systems, but manual curation struggles to scale with the rapid growth of scientific literature. This paper presents KARMA, a novel framework employing multi-agent large language models (LLMs) to automate KG enr…

Cited by 0SourcecodeScholar
2025

LEMMA: Learning from Errors for MatheMatical Advancement in LLMs

ACL 2025finding

Large language models (LLMs) have demonstrated remarkable reasoning capability in solving mathematical problems. However, existing approaches primarily focus on improving the quality of correct training data, e.g., distilling high-quality correct solutions from advanced models, neglecting the value…

2025

Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes

ICML 2025poster

Large language models (LLMs) have achieved remarkable success, yet aligning their generations with human preferences remains a critical challenge. Existing approaches to preference modeling often rely on an explicit or implicit reward function, overlooking the intricate and multifaceted nature of hu…

Cited by 2SourcePDFScholar
2025

Learning Visual Generative Priors without Text

CVPR 2025poster

Although text-to-image (T2I) models have recently thrived as visual generative priors, their reliance on high-quality text-image pairs makes scaling up expensive. We argue that grasping the cross-modality alignment is not a necessity for a sound visual generative prior, whose focus should be on text…

Cited by 1SourcePDFScholar
2025

Neuron Platonic Intrinsic Representation From Dynamics Using Contrastive Learning

ICLR 2025poster

The Platonic Representation Hypothesis posits that behind different modalities of data (what we sense or detect), there exists a universal, modality-independent representation of reality. Inspired by this, we treat each neuron as a system, where we can detect the neuron’s multi-segment activity data…

Cited by 0SourcePDFScholar
2025

OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer Models

ICCV 2025poster

Diffusion models have emerged as a powerful paradigm for generative tasks such as image synthesis and video generation, with Transformer architectures further enhancing performance. However, the high computational cost of diffusion Transformers--stemming from a large number of sampling steps and com…

Cited by 0SourcePDFScholar
2025

PromptCoT: Synthesizing Olympiad-level Problems for Mathematical Reasoning in Large Language Models

ACL 2025finding

The ability of large language models to solve complex mathematical problems has progressed significantly, particularly for tasks requiring advanced reasoning. However, the scarcity of sufficiently challenging problems, particularly at the Olympiad level, hinders further advancements. In this work, w…

2025

ProtCLIP: Function-Informed Protein Multi-Modal Learning

AAAI 2025technical

Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-…

Cited by 2SourcePDFScholar
2025

Real-time Whole-body Motion Planning Based on Optimized NMPC in Static and Dynamic Environments for Mobile Manipulator

IROS 2025

Recently, the research on mobile manipulators has attracted increasing attention. Ensuring that mobile manipulators can meet obstacle avoidance constraints and efficiently accomplish assigned tasks in dynamic environments remains a significant challenge. To address this issue, this paper proposes an

Cited by 0SourceScholar
2025

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

NeurIPS 2025poster

As textual reasoning with large language models (LLMs) has advanced significant, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods primarily approach multimodal reasoning in a straightforward, text-ce…

Cited by 0SourcecodeScholar
2025

RoboScape: Physics-informed Embodied World Model

NeurIPS 2025spotlight

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models exhibit limited physical awareness, particularly in modelin…

Cited by 0SourcecodeScholar
2025

RoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured Environments

CVPR 2025poster

Reliable embodied perception from an egocentric perspective is challenging yet essential for autonomous navigation technology of intelligent mobile agents. With the growing demand of social robotics, near-field scene understanding becomes an important research topic in the areas of egocentric percep…

2025

Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

ICML 2025poster

Long-form video processing fundamentally challenges vision-language models (VLMs) due to the high computational costs of handling extended temporal sequences. Existing token pruning and feature merging methods often sacrifice critical temporal dependencies or dilute semantic information. We introduc…

2025

Semantic-Space-Intervened Diffusive Alignment for Visual Classification

IJCAI 2025

Cross-modal alignment is an effective approach to improving visual classification. Existing studies typically enforce a one-step mapping that uses deep neural networks to project the visual features to mimic the distribution of textual features. However, they typically face difficulties in finding s

Cited by 0SourcePDFScholar
2025

Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer

AAAI 2025technical

Antibodies defend our health by binding to antigens with high specificity and potentiality, primarily relying on the Complementarity-Determining Region (CDR). Yet, current experimental methods of discovering new antibody CDRs are heavily time-consuming. Computational design could alleviate this burd…

2025

TIETracker: A CLIP-based RGB-T Tracking via Feature Interaction and Semantic Enhancement

IROS 2025

The goal of RGB-T tracking is to enhance the accuracy and robustness by leveraging the complementary features of RGB and TIR modalities in complex scenarios. Previous methods have overlooked the power of semantic features in extracting valuable information from different modalities and improving int

Cited by 0SourceScholar
2025

Theoretical Benefit and Limitation of Diffusion Language Model

NeurIPS 2025poster

Diffusion language models have emerged as a new approach for text generation. By enabling the parallel sampling of multiple tokens in each diffusion step, they appear to offer a more efficient alternative to auto-regressive models. However, our observations show that current open-sourced diffusion l…

Cited by 0SourceScholar
2025

TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection

EMNLP 2025

Rapid advances in Large Language Models (LLMs) have spurred demand for processing extended context sequences in contemporary applications. However, this progress faces two challenges: performance degradation due to sequence lengths out-of-distribution, and excessively long inference times caused by

2025

Towards Doctor-Like Reasoning: Medical RAG Fusing Knowledge with Patient Analogy through Textual Gradients

NeurIPS 2025poster

Existing medical RAG systems mainly leverage knowledge from medical knowledge bases, neglecting the crucial role of experiential knowledge derived from similar patient cases - a key component of human clinical reasoning. To bridge this gap, we propose DoctorRAG, a RAG framework that emulates doctor-…

Cited by 0SourceScholar
2025

Understanding Parametric and Contextual Knowledge Reconciliation within Large Language Models

NeurIPS 2025spotlight

Retrieval-Augmented Generation (RAG) provides additional contextual knowledge to complement the parametric knowledge in Large Language Models (LLMs). These two knowledge interweave to enhance the accuracy and timeliness of LLM responses. However, the internal mechanisms by which LLMs utilize thes…

Cited by 0SourceScholar
2025

UniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object Detection

CVPR 2025poster

Recent advances in LiDAR 3D detection have demonstrated the effectiveness of Transformer-based frameworks in capturing the global dependencies from point cloud spaces, which serialize the 3D voxels into the flattened 1D sequence for iterative self-attention. However, the spatial structure of 3D voxe…

Cited by 0SourcePDFScholar
2024

AMOR: A Recipe for Building Adaptable Modular Knowledge Agents Through Process Feedback

NeurIPS 2024poster

The notable success of large language models (LLMs) has sparked an upsurge in building language agents to complete various complex tasks. We present AMOR, an agent framework based on open-source LLMs, which reasons with external knowledge bases and adapts to specific domains through human supervisio…

2024

Augmenting Transformers with Recursively Composed Multi-grained Representations

ICLR 2024poster

We present ReCAT, a recursive composition augmented Transformer that is able to explicitly model hierarchical syntactic structures of raw texts without relying on gold trees during both learning and inference. Existing research along this line restricts data to follow a hierarchical tree structure…

2024

Bridge-IF: Learning Inverse Protein Folding with Markov Bridges

NeurIPS 2024poster

Inverse protein folding is a fundamental task in computational protein design, which aims to design protein sequences that fold into the desired backbone structures. While the development of machine learning algorithms for this task has seen significant success, the prevailing approaches, which pred…

2024

From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis

EMNLP 2024main

We explore multi-step reasoning in vision-language models (VLMs). The problem is challenging, as reasoning data consisting of multiple steps of visual and language processing are barely available. To overcome the challenge, we first introduce a least-to-most visual reasoning paradigm, which interlea…

2024

Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale

ACL 2024long

A syntactic language model (SLM) incrementally generates a sentence with its syntactic tree in a left-to-right manner.We present Generative Pretrained Structured Transformers (GPST), an unsupervised SLM at scale capable of being pre-trained from scratch on raw texts with high parallelism. GPST circu…

2024

HoloVIC: Large-scale Dataset and Benchmark for Multi-Sensor Holographic Intersection and Vehicle-Infrastructure Cooperative

CVPR 2024poster

Vehicle-to-everything (V2X) is a popular topic in the field of Autonomous Driving in recent years. Vehicle-infrastructure cooperation (VIC) becomes one of the important research area. Due to the complexity of traffic conditions such as blind spots and occlusion it greatly limits the perception capab…

Cited by 32SourcePDFScholar
2024

Language-Image Pre-training with Long Captions

ECCV 2024poster

"Language-image pre-training largely relies on how precisely and thoroughly a text describes its paired image. In practice, however, the contents of an image can be so rich that well describing them requires lengthy captions (e.g., with 10 sentences), which are usually missing in existing datasets.…

2024

LoTLIP: Improving Language-Image Pre-training for Long Text Understanding

NeurIPS 2024poster

In this work, we empirically confirm that the key reason causing such an issue is that the training images are usually paired with short captions, leaving certain tokens easily overshadowed by salient tokens. Towards this problem, our initial attempt is to relabel the data with long captions, howeve…

2024

Mixture-of-Modules: Reinventing Transformers as Dynamic Assemblies of Modules

EMNLP 2024main

Is it always necessary to compute tokens from shallow to deep layers in Transformers? The continued success of vanilla Transformers and their variants suggests an undoubted “yes”. In this work, however, we attempt to break the depth-ordered convention by proposing a novel architecture dubbed mixture…

2024

SMART: Scalable Multi-agent Real-time Motion Generation via Next-token Prediction

NeurIPS 2024poster

Data-driven autonomous driving motion generation tasks are frequently impacted by the limitations of dataset size and the domain gap between datasets, which precludes their extensive application in real-world scenarios. To address this issue, we introduce SMART, a novel autonomous driving motion gen…

2024

SwiftPillars: High-Efficiency Pillar Encoder for Lidar-Based 3D Detection

AAAI 2024technical

Lidar-based 3D Detection is one of the significant components of Autonomous Driving. However, current methods over-focus on improving the performance of 3D Lidar perception, which causes the architecture of networks becoming complicated and hard to deploy. Thus, the methods are difficult to apply in…

Cited by 4SourcePDFScholar
2024

Tackling Uncertain Correspondences for Multi-Modal Entity Alignment

NeurIPS 2024poster

Recently, multi-modal entity alignment has emerged as a pivotal endeavor for the integration of Multi-Modal Knowledge Graphs (MMKGs) originating from diverse data sources. Existing works primarily focus on fully depicting entity features by designing various modality encoders or fusion approaches. H…

Cited by 5SourcePDFScholar
2024

“In-Dialogues We Learn”: Towards Personalized Dialogue Without Pre-defined Profiles through In-Dialogue Learning

EMNLP 2024main

Personalized dialogue systems have gained significant attention in recent years for their ability to generate responses in alignment with different personas. However, most existing approaches rely on pre-defined personal profiles, which are not only time-consuming and labor-intensive to create but a…

Cited by 2SourcePDFScholar
2023

A Hybrid Quadratic Programming Framework for Real-Time Embedded Safety-Critical Control

ICRA 2023poster

We present a new framework for implementing real-time embedded safety-critical controllers which utilizes hybrid computing to address the issue of limited computational resources, a problem that is particularly prevalent in microrobotics. In our approach, the nominal stabilizing control algorithm is…

Cited by 9SourceScholar
2023

ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation

ICCV 2023poster

We present a GAN-based Transformer for general action-conditioned 3D human motion generation, including not only single-person actions but also multi-person interactive actions. Our approach consists of a powerful Action-conditioned motion TransFormer (ActFormer) under a GAN training scheme, equippe…

Cited by 74PDFScholar
2023

Dual Path Modeling for Semantic Matching by Perceiving Subtle Conflicts

ICASSP 2023accepted

Transformer-based pre-trained models have achieved great improvements in semantic matching. However, existing models still suffer from insufficient ability to capture subtle differences. The modification, addition and deletion of words in sentence pairs may make it difficult for the model to predict…

Cited by 0SourceScholar
2023

Fusion or Defusion? Flexible Vision-and-Language Pre-Training

ACL 2023findings

Existing approaches in the vision-and-language pre-training (VLP) paradigm mainly deploy either fusion-based encoders or dual-encoders, failing to achieve both effectiveness and efficiency in downstream multimodal tasks. In this paper, we build a flexible VLP model by incorporating cross-modal fusio…

Cited by 2SourcePDFScholar
2023

GROVE: A Retrieval-augmented Complex Story Generation Framework with A Forest of Evidence

EMNLP 2023long findings

Conditional story generation is significant in human-machine interaction, particularly in producing stories with complex plots. While Large language models (LLMs) perform well on multiple NLP tasks, including story generation, it is challenging to generate stories with both complex and creative plot…

Cited by 0SourceScholar
2023

Guide and Select: A Transformer-Based Multimodal Fusion Method for Points of Interest Description Generation

ICASSP 2023accepted

The task of Points of Interest (POI) description generation aims to generate an objective and informative description for a given POI based on POI-related information. High-quality descriptions can better guide users and improve the performance of POI-related recommendation systems. A practical POI…

Cited by 0SourceScholar
2023

Intent-aware Recommendation via Disentangled Graph Contrastive Learning

IJCAI 2023poster

Graph neural network (GNN) based recommender systems have become one of the mainstream trends due to the powerful learning ability from user behavior data. Understanding the user intents from behavior data is the key to recommender systems, which poses two basic requirements for GNN-based recommende…

2023

Let Me Check the Examples: Enhancing Demonstration Learning via Explicit Imitation

ACL 2023short

Demonstration learning aims to guide the prompt prediction by providing answered demonstrations in the few shot settings. Despite achieving promising results, existing work only concatenates the answered examples as demonstrations to the prompt template (including the raw context) without any additi…

2023

LidarGait: Benchmarking 3D Gait Recognition With Point Clouds

CVPR 2023poster

Video-based gait recognition has achieved impressive results in constrained scenarios. However, visual cameras neglect human 3D structure information, which limits the feasibility of gait recognition in the 3D wild world. Instead of extracting gait features from images, this work explores precise 3D…

2023

Local and Global: Temporal Question Answering via Information Fusion

IJCAI 2023poster

Many models that leverage knowledge graphs (KGs) have recently demonstrated remarkable success in question answering (QA) tasks. In the real world, many facts contained in KGs are time-constrained thus temporal KGQA has received increasing attention. Despite the fruitful efforts of previous models i…

Cited by 18SourcePDFScholar
2023

Low-Complexity Acoustic Echo Cancellation with Neural Kalman Filtering

ICASSP 2023accepted

The Kalman filter has been adopted in acoustic echo cancellation due to its robustness to double-talk, fast convergence, and good steady-state performance. The performance of Kalman filter is closely related to the estimation accuracy of the state noise covariance and the observation noise covarianc…

Cited by 0SourceScholar
2023

MAPO: Boosting Large Language Model Performance with Model-Adaptive Prompt Optimization

EMNLP 2023long findings

Prompt engineering, as an efficient and effective way to leverage Large Language Models (LLM), has drawn a lot of attention from the research community. The existing research primarily emphasizes the importance of adapting prompts to specific tasks, rather than specific LLMs. However, a good prompt…

Cited by 0SourceScholar
2023

MD-VQA: Multi-Dimensional Quality Assessment for UGC Live Videos

CVPR 2023poster

User-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers before being distributed to end-users. Because of the flourishing…

2023

Multi-Task Transformer with Relation-Attention and Type-Attention for Named Entity Recognition

ICASSP 2023accepted

Named entity recognition (NER) is an important research problem in natural language processing. There are three types of NER tasks, including flat, nested and discontinuous entity recognition. Most previous sequential labeling models are task-specific, while recent years have witnessed the rising of…

Cited by 0SourceScholar
2023

PreQuant: A Task-agnostic Quantization Approach for Pre-trained Language Models

ACL 2023findings

While transformer-based pre-trained language models (PLMs) have dominated a number of NLP applications, these models are heavy to deploy and expensive to use. Therefore, effectively compressing large-scale PLMs becomes an increasingly important problem. Quantization, which represents high-precision…

Cited by 8SourcePDFScholar
2023

RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank

ACL 2023long

Unsupervised sentence representation learning is one of the fundamental problems in natural language processing with various downstream applications. Recently, contrastive learning has been widely adopted which derives high-quality sentence representations by pulling similar semantics closer and pus…

2023

Regularized Mask Tuning: Uncovering Hidden Knowledge in Pre-Trained Vision-Language Models

ICCV 2023poster

Prompt tuning and adapter tuning have shown great potential in transferring pre-trained vision-language models (VLMs) to various downstream tasks. In this work, we design a new type of tuning method, termed as regularized mask tuning, which masks the network parameters through a learnable selection.…

Cited by 12PDFScholar
2023

Seen to Unseen: Exploring Compositional Generalization of Multi-Attribute Controllable Dialogue Generation

ACL 2023long

Existing controllable dialogue generation work focuses on the single-attribute control and lacks generalization capability to out-of-distribution multiple attribute combinations. In this paper, we explore the compositional generalization for multi-attribute controllable dialogue generation where a m…

2023

T5-SR: A Unified Seq-to-Seq Decoding Strategy for Semantic Parsing

ICASSP 2023accepted

Translating natural language queries into SQLs in a seq2seq manner has attracted much attention recently. However, compared with abstract-syntactic-tree-based SQL generation, seq2seq semantic parsers face much more challenges, including poor quality on schematical information prediction and poor sem…

Cited by 0SourceScholar
2023

Time-Aware Multiway Adaptive Fusion Network for Temporal Knowledge Graph Question Answering

ICASSP 2023accepted

Knowledge graphs (KGs) have received increasing attention due to its wide applications on natural language processing. However, its use case on temporal question answering (QA) has not been well-explored. Most of existing methods are developed based on pre-trained language models, which might not be…

Cited by 0SourceScholar
2022

An Effective and Efficient Entity Alignment Decoding Algorithm via Third-Order Tensor Isomorphism

ACL 2022long

Entity alignment (EA) aims to discover the equivalent entity pairs between KGs, which is a crucial step for integrating multi-source KGs.For a long time, most researchers have regarded EA as a pure graph representation learning task and focused on improving graph encoders while paying little attenti…

2022

Backbone Is All Your Need: A Simplified Architecture for Visual Object Tracking

ECCV 2022poster

"Exploiting a general-purpose neural architecture to replace hand-wired designs or inductive biases has recently drawn extensive interest. However, existing tracking approaches rely on customized sub-modules and need prior knowledge for architecture selection, hindering the development of tracking i…

2022

CLOWER: A Pre-trained Language Model with Contrastive Learning over Word and Character Representations

COLING 2022main

Pre-trained Language Models (PLMs) have achieved remarkable performance gains across numerous downstream tasks in natural language understanding. Various Chinese PLMs have been successively proposed for learning better Chinese language representation. However, most current models use Chinese charact…

2022

CQG: A Simple and Effective Controlled Generation Framework for Multi-hop Question Generation

ACL 2022long

Multi-hop question generation focuses on generating complex questions that require reasoning over multiple pieces of information of the input passage. Current models with state-of-the-art performance have been able to generate the correct questions corresponding to the answers. However, most models…

2022

Cross Domain Object Detection by Target-Perceived Dual Branch Distillation

CVPR 2022poster

Cross domain object detection is a realistic and challenging task in the wild. It suffers from performance degradation due to large shift of data distributions and lack of instance-level annotations in the target domain. Existing approaches mainly focus on either of these two difficulties, even thou…

Cited by 90PDFcodeScholar
2022

Disentangled Knowledge Transfer for OOD Intent Discovery with Unified Contrastive Learning

ACL 2022short

Discovering Out-of-Domain(OOD) intents is essential for developing new skills in a task-oriented dialogue system. The key challenge is how to transfer prior IND knowledge to OOD clustering. Different from existing work based on shared intent representation, we propose a novel disentangled knowledge…

2022

Domain-Oriented Prefix-Tuning: Towards Efficient and Generalizable Fine-tuning for Zero-Shot Dialogue Summarization

NAACL 2022long

The most advanced abstractive dialogue summarizers lack generalization ability on new domains and the existing researches for domain adaptation in summarization generally rely on large-scale pre-trainings. To explore the lightweight fine-tuning methods for domain adaptation of dialogue summarization…

2022

Ensemble Multi-Relational Graph Neural Networks

IJCAI 2022poster

It is well established that graph neural networks (GNNs) can be interpreted and designed from the perspective of optimization objective. With this clear optimization objective, the deduced GNNs architecture has sound theoretical foundation, which is able to flexibly remedy the weakness of GNNs. Howe…

2022

GNN-encoder: Learning a Dual-encoder Architecture via Graph Neural Networks for Dense Passage Retrieval

EMNLP 2022finding

Recently, retrieval models based on dense representations are dominant in passage retrieval tasks, due to their outstanding ability in terms of capturing semantics of input text compared to the traditional sparse vector space models. A common practice of dense retrieval models is to exploit a dual-e…

Cited by 4SourcePDFScholar
2022

Generalized Intent Discovery: Learning from Open World Dialogue System

COLING 2022main

Traditional intent classification models are based on a pre-defined intent set and only recognize limited in-domain (IND) intent classes. But users may input out-of-domain (OOD) queries in a practical dialogue system. Such OOD queries can provide directions for future improvement. In this paper, we…

2022

Improving Semantic Matching through Dependency-Enhanced Pre-trained Model with Adaptive Fusion

EMNLP 2022finding

Transformer-based pre-trained models like BERT have achieved great progress on Semantic Sentence Matching. Meanwhile, dependency prior knowledge has also shown general benefits in multiple NLP tasks. However, how to efficiently integrate dependency prior structure into pre-trained models to better m…

2022

Incorporating Dynamic Semantics into Pre-Trained Language Model for Aspect-based Sentiment Analysis

ACL 2022findings

Aspect-based sentiment analysis (ABSA) predicts sentiment polarity towards a specific aspect in the given sentence. While pre-trained language models such as BERT have achieved great success, incorporating dynamic semantic changes into ABSA remains challenging. To this end, in this paper, we propose…

Cited by 84SourcePDFScholar
2022

Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification

ACL 2022long

Tuning pre-trained language models (PLMs) with task-specific prompts has been a promising approach for text classification. Particularly, previous studies suggest that prompt-tuning has remarkable superiority in the low-data scenario over the generic fine-tuning methods with extra classifiers. The c…

2022

L-Tracing: Fast Light Visibility Estimation on Neural Surfaces by Sphere Tracing

ECCV 2022poster

"We introduce a highly efficient light visibility estimation method, called L-Tracing, for reflectance factorization on neural implicit surfaces. Light visibility is indispensable for modeling shadows and specular of high quality on object’s surface. For neural implicit representations, former metho…

Cited by 11SourcePDFScholar
2022

Learning Video Representations of Human Motion From Synthetic Data

CVPR 2022poster

In this paper, we take an early step towards video representation learning of human actions with the help of largescale synthetic videos, particularly for human motion representation enhancement. Specifically, we first introduce an automatic action-related video synthesis pipeline based on a photore…

Cited by 17PDFScholar
2022

Learning to Express in Knowledge-Grounded Conversation

NAACL 2022long

Grounding dialogue generation by extra knowledge has shown great potentials towards building a system capable of replying with knowledgeable and engaging responses. Existing studies focus on how to synthesize a response with proper knowledge, yet neglect that the same knowledge could be expressed di…

2022

Making Parameter-efficient Tuning More Efficient: A Unified Framework for Classification Tasks

COLING 2022main

Large pre-trained language models (PLMs) have demonstrated superior performance in industrial applications. Recent studies have explored parameter-efficient PLM tuning, which only updates a small amount of task-specific parameters while achieving both high efficiency and comparable performance again…

2022

Making Pretrained Language Models Good Long-tailed Learners

EMNLP 2022main

Prompt-tuning has shown appealing performance in few-shot classification by virtue of its capability in effectively exploiting pre-trained knowledge. This motivates us to check the hypothesis that prompt-tuning is also a promising choice for long-tailed classification, since the tail classes are int…

2022

Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language Models

NeurIPS 2022accept

Despite the great success of pre-trained language models (PLMs) in a large set of natural language processing (NLP) tasks, there has been a growing concern about their security in real-world applications. Backdoor attack, which poisons a small number of training samples by inserting backdoor trigger…

2022

Multi-Thread CTAEA-Based Workstation Reconfiguration for Multi-Stage Automobile Engine Flow Shop Considering Performance Deterioration

RA-L 2022

In the automobile engine flow shop (AEFS), the equipment has more chance to work for a long time, so the manufacturing performance may deteriorate rapidly, thus causing operating unbalance and inefficiency. To tackle this problem at minimum cost, this study proposes a multi-thread constrained two-ar

Cited by 6SourceScholar
2022

PATS: Sensitivity-aware Noisy Learning for Pretrained Language Models

EMNLP 2022main

A wide range of NLP tasks benefit from the fine-tuning of pretrained language models (PLMs). However, a number of redundant parameters which contribute less to the downstream task are observed in a directly fine-tuned model. We consider the gap between pretraining and downstream tasks hinders the tr…

2022

PlugAT: A Plug and Play Module to Defend against Textual Adversarial Attack

COLING 2022main

Adversarial training, which minimizes the loss of adversarially perturbed examples, has received considerable attention. However, these methods require modifying all model parameters and optimizing the model from scratch, which is parameter inefficient and unfriendly to the already deployed models.…

2022

Retrieval Enhanced Segment Generation Neural Network for Task-Oriented Dialogue Systems

ICASSP 2022accepted

For task-oriented dialogue systems, Natural Language Generation (NLG) is the last and vital step which aims at generating an appropriate response according to the dialogue act (DA). While end-to-end neural networks have achieved promising performances on this task, the existing models still struggle…

Cited by 0SourceScholar
2022

Revisit Overconfidence for OOD Detection: Reassigned Contrastive Learning with Adaptive Class-dependent Threshold

NAACL 2022long

Detecting Out-of-Domain (OOD) or unknown intents from user queries is essential in a task-oriented dialog system. A key challenge of OOD detection is the overconfidence of neural models. In this paper, we comprehensively analyze overconfidence and classify it into two perspectives: over-confident OO…

2022

Robust Lottery Tickets for Pre-trained Language Models

ACL 2022long

Recent works on Lottery Ticket Hypothesis have shown that pre-trained language models (PLMs) contain smaller matching subnetworks(winning tickets) which are capable of reaching accuracy comparable to the original models. However, these tickets are proved to be notrobust to adversarial examples, and…

2022

Searching for Optimal Subword Tokenization in Cross-domain NER

IJCAI 2022poster

Input distribution shift is one of the vital problems in unsupervised domain adaptation (UDA). The most popular UDA approaches focus on domain-invariant representation learning, trying to align the features from different domains into a similar feature distribution. However, these approaches ignore…

2022

Structural Bias for Aspect Sentiment Triplet Extraction

COLING 2022main

Structural bias has recently been exploited for aspect sentiment triplet extraction (ASTE) and led to improved performance. On the other hand, it is recognized that explicitly incorporating structural bias would have a negative impact on efficiency, whereas pretrained language models (PLMs) can alre…

2022

TANet: Thread-Aware Pretraining for Abstractive Conversational Summarization

NAACL 2022findings

Although pre-trained language models (PLMs) have achieved great success and become a milestone in NLP, abstractive conversational summarization remains a challenging but less studied task. The difficulty lies in two aspects. One is the lack of large-scale conversational summary data. Another is that…

2022

Target-Relevant Knowledge Preservation for Multi-Source Domain Adaptive Object Detection

CVPR 2022oral

Domain adaptive object detection (DAOD) is a promising way to alleviate performance drop of detectors in new scenes. Albeit great effort made in single source domain adaptation, a more generalized task with multiple source domains remains not being well explored, due to knowledge degradation during…

Cited by 31PDFScholar
2022

Temporal Complementarity-Guided Reinforcement Learning for Image-to-Video Person Re-Identification

CVPR 2022poster

Image-to-video person re-identification aims to retrieve the same pedestrian as the image-based query from a video-based gallery set. Existing methods treat it as a cross-modality retrieval task and learn the common latent embeddings from image and video modalities, which are both less effective and…

Cited by 17PDFScholar
2022

UniNL: Aligning Representation Learning with Scoring Function for OOD Detection via Unified Neighborhood Learning

EMNLP 2022main

Detecting out-of-domain (OOD) intents from user queries is essential for avoiding wrong operations in task-oriented dialogue systems. The key challenge is how to distinguish in-domain (IND) and OOD intents. Previous methods ignore the alignment between representation learning and scoring function, l…

2022

VIRT: Improving Representation-based Text Matching via Virtual Interaction

EMNLP 2022main

Text matching is a fundamental research problem in natural language understanding. Interaction-based approaches treat the text pair as a single sequence and encode it through cross encoders, while representation-based models encode the text pair independently with siamese or dual encoders. Interacti…

Cited by 8SourcePDFScholar
2022

Watch the Neighbors: A Unified K-Nearest Neighbor Contrastive Learning Framework for OOD Intent Discovery

EMNLP 2022main

Discovering out-of-domain (OOD) intent is important for developing new skills in task-oriented dialogue systems. The key challenges lie in how to transfer prior in-domain (IND) knowledge to OOD clustering, as well as jointly learn OOD representations and cluster assignments. Previous methods suffer…

2022

XPrompt: Exploring the Extreme of Prompt Tuning

EMNLP 2022main

Prompt tuning learns soft prompts to condition the frozen Pre-trained Language Models (PLMs) for performing downstream tasks in a parameter-efficient manner. While prompt tuning has gradually reached the performance level of fine-tuning as the model scale increases, there is still a large performanc…

Cited by 39SourcePDFScholar
2021

A Survey on Response Selection for Retrieval-based Dialogues

IJCAI 2021poster

Building an intelligent dialogue system capable of naturally and coherently conversing with humans has been a long-standing goal of artificial intelligence. In the past decade, with the development of machine/deep learning technology and the explosive growth of available conversation data in social…

Cited by 36SourcePDFScholar
2021

ASAP: A Chinese Review Dataset Towards Aspect Category Sentiment Analysis and Rating Prediction

NAACL 2021long

Sentiment analysis has attracted increasing attention in e-commerce. The sentiment polarities underlying user reviews are of great value for business intelligence. Aspect category sentiment analysis (ACSA) and review rating prediction (RP) are two essential tasks to detect the fine-to-coarse sentime…

2021

BSN++: Complementary Boundary Regressor with Scale-Balanced Relation Modeling for Temporal Action Proposal Generation

AAAI 2021technical

Generating human action proposals in untrimmed videos is an important yet challenging task with wide applications. Current methods often suffer from the noisy boundary locations and the inferior quality of confidence scores used for proposal retrieving. In this paper, we present BSN++, a new framewo…

Cited by 143SourcePDFScholar
2021

Capturing Event Argument Interaction via A Bi-Directional Entity-Level Recurrent Decoder

ACL 2021long

Capturing interactions among event arguments is an essential step towards robust event argument extraction (EAE). However, existing efforts in this direction suffer from two limitations: 1) The argument role type information of contextual entities is mainly utilized as training signals, ignoring the…

2021

ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation Transfer

ACL 2021long

Learning high-quality sentence representations benefits a wide range of natural language processing tasks. Though BERT-based pre-trained language models achieve high performance on many downstream tasks, the native derived sentence representations are proved to be collapsed and thus produce a poor p…

2021

Context-Aware Graph Convolution Network for Target Re-identification

AAAI 2021technical

Most existing re-identification methods focus on learning robust and discriminative features with deep convolution networks. However, many of them consider content similarity separately and fail to utilize the context information of the query and gallery sets, e.g. probe-gallery and gallery-gallery…

Cited by 31SourcePDFScholar
2021

Correlation-Aware Heuristic Search for Intelligent Virtual Machine Provisioning in Cloud Systems

AAAI 2021technical

The optimization of resource is crucial for the operation of public cloud systems such as Microsoft Azure, as well as servers dedicated to the workloads of large customers such as Microsoft 365. Those optimization tasks often need to take unknown parameters into consideration and can be formulated a…

2021

Dance Revolution: Long-Term Dance Generation with Music via Curriculum Learning

ICLR 2021poster

Dancing to music is one of human's innate abilities since ancient times. In machine learning research, however, synthesizing dance movements from music is a challenging problem. Recently, researchers synthesize human motion sequences through autoregressive models like recurrent neural network (RNN).…

Cited by 159SourcePDFScholar
2021

Explaining A Black-box By Using A Deep Variational Information Bottleneck Approach

AAAI 2021technical

Interpretable machine learning has gained much attention recently. Briefness and comprehensiveness are necessary in order to provide a large amount of information concisely when explaining a black-box decision system. However, existing interpretable machine learning methods fail to consider briefnes…

2021

Improving Document Representations by Generating Pseudo Query Embeddings for Dense Retrieval

ACL 2021long

Recently, the retrieval models based on dense representations have been gradually applied in the first stage of the document retrieval tasks, showing better performance than traditional sparse vector space models. To obtain high efficiency, the basic structure of these models is Bi-encoder in most c…

2021

Improving Event Detection by Exploiting Label Hierarchy

ICASSP 2021accepted

Event types are hierarchical, yet most existing methods for event detection classify candidate triggers into fine-grained event types directly, without considering the rich semantic correlations in the hierarchy of event types. To fully utilize such information to improve the detection of fine-grain…

Cited by 0SourceScholar
2021

Incorporating Convolution Designs Into Visual Transformers

ICCV 2021poster

Motivated by the success of Transformers in natural language processing (NLP) tasks, there exist some attempts (e.g., ViT and DeiT) to apply Transformers to the vision domain. However, pure Transformer architectures often require a large amount of training data or extra supervision to obtain compara…

Cited by 659PDFcodeScholar
2021

Large-Scale Relation Learning for Question Answering over Knowledge Bases with Pre-trained Language Models

EMNLP 2021main

The key challenge of question answering over knowledge bases (KBQA) is the inconsistency between the natural language questions and the reasoning paths in the knowledge base (KB). Recent graph-based KBQA methods are good at grasping the topological structure of the graph but often ignore the textual…

2021

Learning Statistical Texture for Semantic Segmentation

CVPR 2021poster

Existing semantic segmentation works mainly focus on learning the contextual information in high-level semantic features with CNNs. In order to maintain a precise boundary, low-level texture features are directly skip-connected into the deeper layers. Nevertheless, texture features are not only abou…

Cited by 177PDFcodeScholar
2021

Open Domain Dialogue Generation with Latent Images

AAAI 2021technical

We consider grounding open domain dialogues with images. Existing work assumes that both an image and a textual context are available, but image-grounded dialogues by nature are more difficult to obtain than textual dialogues. Thus, we propose learning a response generation model with both image-gro…

2021

PULNS: Positive-Unlabeled Learning with Effective Negative Sample Selector

AAAI 2021technical

Positive-unlabeled learning (PU learning) is an important case of binary classification where the training data only contains positive and unlabeled samples. The current state-of-the-art approach for PU learning is the cost-sensitive approach, which casts PU learning as a cost-sensitive classificati…

2021

Spatial-Temporal Correlation and Topology Learning for Person Re-Identification in Videos

CVPR 2021poster

Video-based person re-identification aims to match pedestrians from video sequences across non-overlapping camera views. The key factor for video person re-identification is to effectively exploit both spatial and temporal clues from video sequences. In this work, we propose a novel Spatial-Temporal…

Cited by 81PDFScholar
2021

Task-Oriented Clustering for Dialogues

EMNLP 2021finding

A reliable clustering algorithm for task-oriented dialogues can help developer analysis and define dialogue tasks efficiently. It is challenging to directly apply prior normal text clustering algorithms for task-oriented dialogues, due to the inherent differences between them, such as coreference, o…

2021

Temporal Context Aggregation Network for Temporal Action Proposal Refinement

CVPR 2021poster

Temporal action proposal generation aims to estimate temporal intervals of actions in untrimmed videos, which is a challenging yet important task in the video understanding field. The proposals generated by current methods still suffer from inaccurate temporal boundaries and inferior confidence used…

Cited by 166PDFScholar
2021

Virtual Data Augmentation: A Robust and General Framework for Fine-tuning Pre-trained Models

EMNLP 2021main

Recent works have shown that powerful pre-trained language models (PLM) can be fooled by small perturbations or intentional attacks. To solve this issue, various data augmentation techniques are proposed to improve the robustness of PLMs. However, it is still challenging to augment semantically rele…

2020

Adaptive Dilated Network With Self-Correction Supervision for Counting

CVPR 2020poster

The counting problem aims to estimate the number of objects in images. Due to large scale variation and labeling deviations, it remains a challenging task. The static density map supervised learning framework is widely used in existing methods, which uses the Gaussian kernel to generate a density ma…

Cited by 207PDFScholar
2020

Class-wise Dynamic Graph Convolution for Semantic Segmentation

ECCV 2020poster

Recent works have made great progress in semantic segmentation by exploiting contextual information in a local or global manner with dilated convolutions, pyramid pooling or self-attention mechanism. In order to avoid potential misleading contextual information aggregation in previous work, we propo…

Cited by 105SourcePDFScholar
2020

Description Based Text Classification with Reinforcement Learning

ICML 2020poster

The task of text classification is usually divided into two stages: text feature extraction and classification. In this standard formalization, categories are merely represented as indexes in the label vocabulary, and the model lacks for explicit instructions on what to classify. Inspired by the cur…

Cited by 73SourcePDFScholar
2020

Hierarchical Feature Embedding for Attribute Recognition

CVPR 2020poster

Attribute recognition is a crucial but challenging task due to viewpoint changes, illumination variations and appearance diversities, etc. Most of previous work only consider the attribute-level feature embedding, which might perform poorly in complicated heterogeneous conditions. To address this pr…

Cited by 61PDFScholar
2020

Intelligent Virtual Machine Provisioning in Cloud Computing

IJCAI 2020poster

Virtual machine (VM) provisioning is a common and critical problem in cloud computing. In industrial cloud platforms, there are a huge number of VMs provisioned per day. Due to the complexity and resource constraints, it needs to be carefully optimized to make cloud platforms effectively utilize the…

2020

Low-Resource Knowledge-Grounded Dialogue Generation

ICLR 2020poster

Responding with knowledge has been recognized as an important capability for an intelligent conversational agent. Yet knowledge-grounded dialogues, as training data for learning such a response generation model, are difficult to obtain. Motivated by the challenge in practice, we consider knowledge-g…

Cited by 117SourceScholar
2020

Retrieve, Program, Repeat: Complex Knowledge Base Question Answering via Alternate Meta-learning

IJCAI 2020poster

A compelling approach to complex question answering is to convert the question to a sequence of actions, which can then be executed on the knowledge base to yield the answer, aka the programmer-interpreter approach. Use similar training questions to the test question, meta-learning enables the progr…

2020

Zero-Resource Knowledge-Grounded Dialogue Generation

NeurIPS 2020poster

While neural conversation models have shown great potentials towards generating informative and engaging responses via introducing external knowledge, learning such a model often requires knowledge-grounded dialogues that are difficult to obtain. To overcome the data challenge and reduce the cost of…

2019

Feedback Network for Image Super-Resolution

CVPR 2019poster

Recent advances in image super-resolution (SR) explored the power of deep learning to achieve a better reconstruction performance. However, the feedback mechanism, which commonly exists in human visual system, has not been fully exploited in existing deep learning based image SR methods. In this pap…

Cited by 1053PDFcodeScholar
2019

Glyce: Glyph-vectors for Chinese Character Representations

NeurIPS 2019poster

It is intuitive that NLP tasks for logographic languages like Chinese should benefit from the use of the glyph information in those languages. However, due to the lack of rich pictographic evidence in glyphs and the weak generalization ability of standard computer vision models on character data, a…

2019

Multi-scale Vehicle Re-identification Using Self-adapting Label Smoothing Regularization

ICASSP 2019accepted

Vehicle re-identification (re-id) plays an important role in intelligent surveillance. Since difference vehicle models may have similar appearances, together with the problem of image scale variations, the vehicle re-id remains long-term challenging. We present a novel multi-scale vehicle re-id fram…

Cited by 0SourceScholar
2019

Online Hyper-Parameter Learning for Auto-Augmentation Strategy

ICCV 2019poster

Data augmentation is critical to the success of modern deep learning techniques. In this paper, we propose Online Hyper-parameter Learning for Auto-Augmentation (OHL-Auto-Aug), an economical solution that learns the augmentation policy distribution along with network training. Unlike previous method…

Cited by 109PDFScholar
2019

STM: SpatioTemporal and Motion Encoding for Action Recognition

ICCV 2019poster

Spatiotemporal and motion features are two complementary and crucial information for video action recognition. Recent state-of-the-art methods adopt a 3D CNN stream to learn spatiotemporal features and another flow stream to learn motion features. In this work, we aim to efficiently encode these two…

Cited by 556PDFScholar
2019

Selective Sensor Fusion for Neural Visual-Inertial Odometry

CVPR 2019poster

Deep learning approaches for Visual-Inertial Odometry (VIO) have proven successful, but they rarely focus on incorporating robust fusion strategies for dealing with imperfect input sensory data. We propose a novel end-to-end selective sensor fusion framework for monocular VIO, which fuses monocular…

Cited by 192PDFcodeScholar
2019

SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks

CVPR 2019oral

Siamese network based trackers formulate tracking as convolutional feature cross-correlation between target template and searching region. However, Siamese trackers still have accuracy gap compared with state-of-the-art algorithms and they cannot take advantage of feature from deep networks, such as…

Cited by 2748PDFScholar
2018

Distractor-aware Siamese Networks for Visual Object Tracking

ECCV 2018poster

Recently, Siamese networks have drawn great attention in visual tracking community because of their balanced accuracy and speed. However, features used in most Siamese tracking approaches can only discriminate foreground from the non-semantic backgrounds. The semantic backgrounds are always consider…

2018

End-to-End Flow Correlation Tracking With Spatial-Temporal Attention

CVPR 2018poster

Discriminative correlation filters (DCF) with deep convolutional features have achieved favorable performance in recent tracking benchmarks. However, most of existing DCF trackers only consider appearance features of current frame, and hardly benefit from motion and inter-frame information. The lack…

2018

High Performance Visual Tracking With Siamese Region Proposal Network

CVPR 2018poster

Visual object tracking has been a fundamental topic in recent years and many deep learning based trackers have achieved state-of-the-art performance on multiple benchmarks. However, most of these trackers can hardly get top performance with real-time speed. In this paper, we propose the Siamese regi…

Cited by 3306SourcePDFScholar
2018

Orthogonality-Promoting Distance Metric Learning: Convex Relaxation and Theoretical Analysis

ICML 2018oral

Distance metric learning (DML), which learns a distance metric from labeled "similar" and "dissimilar" data pairs, is widely utilized. Recently, several works investigate orthogonality-promoting regularization (OPR), which encourages the projection vectors in DML to be close to being orthogonal, to…

Cited by 33SourcePDFScholar
2018

PointCNN: Convolution On X-Transformed Points

NeurIPS 2018poster

We present a simple and general framework for feature learning from point cloud. The key to the success of CNNs is the convolution operator that is capable of leveraging spatially-local correlation in data represented densely in grids (e.g. images). However, point cloud are irregular and unordered,…

2018

Practical Block-Wise Neural Network Architecture Generation

CVPR 2018poster

Convolutional neural networks have gained a remarkable success in computer vision. However, most usable network architectures are hand-crafted and usually require expertise and elaborate design. In this paper, we provide a block-wise network generation pipeline called BlockQNN which automatically bu…

Cited by 650SourcePDFScholar