← Search

Ge Li

104 accepted papers

2026

DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt

AAAI 2026technical

Large Vision-Language Models (LVLMs) have achieved impressive progress across various applications but remain vulnerable to malicious queries. Existing safety alignment approaches typically fail to resist malicious queries while preserving utility on benign ones effectively. To address these challen

Cited by 0SourcePDFScholar
2026

Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models

ICLR 2026poster

Data contamination poses a significant threat to the reliable evaluation of Large Language Models (LLMs). This issue arises when benchmark samples may inadvertently appear in training sets, compromising the validity of reported performance. While detection methods have been developed for the pre-tra…

Cited by 0SourcecodeScholar
2026

ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting

CVPR 2026

Lifting 2D open-vocabulary understanding into 3D Gaussian Splatting (3DGS) scenes is a critical challenge. Mainstream methods, built on an embedding paradigm, suffer from three key flaws: (i) geometry-semantic inconsistency, where points, rather than objects, serve as the semantic basis, limiting se

Cited by 0SourceScholar
2026

Large Language Model Unlearning for Source Code

AAAI 2026technical

While Large Language Models (LLMs) excel at code generation, their inherent tendency toward verbatim memorization of training data introduces critical risks like copyright infringement, insecurity emission, and deprecated API utilization, etc. A straightforward yet promising defense is unlearning, i

Cited by 0SourcePDFScholar
2026

Layered 4D-Rotor Gaussian Splatting: A Compressed Representation for Long Dynamic Scenes

CVPR 2026

We address the challenge of reconstructing long dynamic scenes from multi-view videos in a storage-efficient manner. Recent advances in Gaussian Splatting and its extensions to dynamic scenes have demonstrated impressive visual quality, but remain limited to short duration (<10 s), large storage siz

Cited by 0SourceScholar
2026

MoRe-ERL: Learning Motion Residuals Using Episodic Reinforcement Learning

ICRA 2026poster

We propose MoRe-ERL, a framework that combines Episodic Reinforcement Learning (ERL) and residual learning, which refines preplanned reference trajectories into safe, feasible, and efficient task-specific trajectories. This framework is general enough to incorporate into arbitrary ERL methods and mo…

2026

PAWS: Preference Learning with Advantage-Weighted Segments

ICML 2026poster

Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods typically train utility functions on trajectory or segment-level preferences while relying on per-step utility estimates…

Cited by 0SourceScholar
2026

Video Spatial Reasoning with Object-Centric 3D Rollout

AAAI 2026technical

Recent advances in Multi-modal Large Language Models (MLLMs) have showcased remarkable capabilities in vision-language understanding. However, enabling robust video spatial reasoning—the ability to comprehend object locations, orientations, and inter-object relationships in dynamic 3D scenes—remains

Cited by 0SourcePDFScholar
2025

BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning

NeurIPS 2025poster

We present the B-spline Encoded Action Sequence Tokenizer (BEAST), a novel action tokenizer that encodes action sequences into compact discrete or continuous tokens using B-splines. In contrast to existing action tokenizers based on vector quantization or byte pair encoding, BEAST requires no separ…

Cited by 0SourceScholar
2025

Benchmarking Long-Context Language Models on Long Code Understanding

ACL 2025long

Current advanced long-context language models offer great potential for real-world software engineering applications. However, progress in this critical domain remains hampered by a fundamental limitation: the absence of a rigorous evaluation framework for long code understanding. To gap this obstac…

Cited by 0SourcePDFScholar
2025

CodeDPO: Aligning Code Models with Self Generated and Verified Source Code

ACL 2025long

Code generation models have shown significant potential for programming tasks. However, existing training methods like supervised fine-tuning face key limitations: they do not effectively teach models to prioritize correct over incorrect solutions in ambiguous situations, nor do they effectively opt…

Cited by 0SourcePDFScholar
2025

ContextHOI: Spatial Context Learning for Human-Object Interaction Detection

AAAI 2025technical

Spatial contexts, such as the backgrounds and surroundings, are considered critical in Human-Object Interaction (HOI) recognition, especially when the instance-centric foreground is blurred or occluded. Recent advancements in HOI detectors are usually built upon detection transformer pipelines. Whil…

Cited by 1SourcePDFScholar
2025

DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking Heads

ICCV 2025poster

In this work, we investigate the generation of high-fidelity, audio-driven 3D Gaussian talking heads from monocular videos. We present DGTalker, an innovative framework designed for real-time, high-fidelity, and 3D-aware talking head synthesis. By leveraging Gaussian generative priors and treating t…

Cited by 0SourcePDFScholar
2025

DIME: Diffusion-Based Maximum Entropy Reinforcement Learning

ICML 2025poster

Maximum entropy reinforcement learning (MaxEnt-RL) has become the standard approach to RL due to its beneficial exploration properties. Traditionally, policies are parameterized using Gaussian distributions, which significantly limits their representational capacity. Diffusion-based policies offer a…

Cited by 0SourcePDFScholar
2025

Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow

ICCV 2025poster

Long-form video understanding has always been a challenging problem due to the significant redundancy in both temporal and spatial contents. This challenge is further exacerbated by the limited context length of Multimodal Large Language Models (MLLMs). To address this issue, many previous works hav…

Cited by 0SourcePDFScholar
2025

Focused-DPO: Enhancing Code Generation Through Focused Preference Optimization on Error-Prone Points

ACL 2025finding

Code generation models have shown significant potential for automating programming tasks. However, the challenge of generating accurate and reliable code persists due to the highly complex and long-reasoning nature of the task. Even state-of-the-art models often fail in code generation due to small…

Cited by 0SourcePDFScholar
2025

IRIS: An Immersive Robot Interaction System

CoRL 2025poster

This paper introduces IRIS, an Immersive Robot Interaction System leveraging Extended Reality (XR). Existing XR-based systems enable efficient data collection but are often challenging to reproduce and reuse due to their specificity to particular robots, objects, simulators, and environments. IRIS a…

Cited by 0SourceScholar
2025

Improving Formal Reasoning of Transformer with State Stack

NeurIPS 2025poster

The Transformer architecture has emerged as a landmark advancement within the broad field of artificial intelligence, effectively catalyzing the advent of large language models (LLMs). However, despite its remarkable capabilities and the substantial progress it has facilitated, the Transformer archi…

Cited by 0SourceScholar
2025

LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs

ACL 2025long

Detecting tricky bugs in plausible programs, those that pass existing test suites yet still contain bugs, remains a significant challenge in software testing. To address this problem, we propose TrickCatcher, an LLM-powered approach to generating test cases for uncovering bugs in plausible programs.…

2025

Learning Semantic Facial Descriptors for Accurate Face Animation

ICASSP 2025accepted

Face animation is a challenging task. Existing model-based methods (utilizing 3DMMs or landmarks) often result in a model-like reconstruction effect, which doesn't effectively preserve identity. Conversely, model-free approaches face challenges in attaining a decoupled and semantically rich feature…

Cited by 0SourceScholar
2025

MUSE: Mamba Is Efficient Multi-scale Learner for Text-video Retrieval

AAAI 2025technical

Text-Video Retrieval (TVR) aims to align and associate relevant video content with corresponding natural language queries. Most existing TVR methods are based on large-scale pre-trained vision-language models (e.g., CLIP). However, due to CLIP's inherent plain structure, few TVR methods explore the…

2025

Orchestrating the Symphony of Prompt Distribution Learning for Human-Object Interaction Detection

AAAI 2025technical

Human-object interaction (HOI) detectors with popular query-transformer architecture have achieved promising performance. However, accurately identifying uncommon visual patterns and distinguishing between ambiguous HOIs continue to be difficult for them. We observe that these difficulties may arise…

Cited by 1SourcePDFScholar
2025

Point Cloud Semantic Segmentation with Sparse and Inhomogeneous Annotations

AAAI 2025technical

Utilizing uniformly distributed sparse annotations, weakly supervised learning alleviates the heavy reliance on fine-grained annotations in point cloud semantic segmentation tasks. However, few works discuss the inhomogeneity of sparse annotations, albeit it is common in real-world scenarios. Theref…

2025

PointMapPolicy: Structured Point Cloud Processing for Multi-Modal Imitation Learning

NeurIPS 2025poster

Robotic manipulation systems benefit from complementary sensing modalities, where each provides unique environmental information. Point clouds capture detailed geometric structure, while RGB images provide rich semantic context. Current point cloud methods struggle to capture fine-grained detail, es…

Cited by 0SourcecodeScholar
2025

Reasoning is Periodicity? Improving Large Language Models Through Effective Periodicity Modeling

NeurIPS 2025poster

Periodicity, as one of the most important basic characteristics, lays the foundation for facilitating structured knowledge acquisition and systematic cognitive processes within human learning paradigms. However, the potential flaws of periodicity modeling in Transformer affect the learning efficienc…

Cited by 0SourceScholar
2025

Rethinking Repetition Problems of LLMs in Code Generation

ACL 2025long

With the advent of neural language models, the performance of code generation has been significantly boosted. However, the problem of repetitions during the generation process continues to linger. Previous work has primarily focused on content repetition, which is merely a fraction of the broader re…

2025

SATURN: SAT-based Reinforcement Learning to Unleash LLMs Reasoning

NeurIPS 2025spotlight

How to design reinforcement learning (RL) tasks that effectively unleash the reasoning capability of large language models (LLMs) remains an open question. Existing RL tasks (e.g., math, programming, and constructing reasoning tasks) suffer from three key limitations: (1) Scalability. They rely heav…

Cited by 0SourceScholar
2025

Stochasticity-aware No-Reference Point Cloud Quality Assessment

IJCAI 2025

The evolution of point cloud processing algorithms necessitates an accurate assessment for their quality. Previous works consistently regard point cloud quality assessment (PCQA) as a MOS regression problem and devise a deterministic mapping, ignoring the stochasticity in generating MOS from subject

Cited by 0SourcePDFScholar
2025

TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement Learning

ICLR 2025spotlight

This work introduces Transformer-based Off-Policy Episodic Reinforcement Learning (TOP-ERL), a novel algorithm that enables off-policy updates in the ERL framework. In ERL, policies predict entire action trajectories over multiple time steps instead of single actions at every time step. These trajec…

2025

Trustworthy Robot Behavior Tree Generation Based on Multi-Source Heterogeneous Knowledge Graph

ICRA 2025

In robotics, the design of robot behavior trees generally requires roboticists to comprehensively and customizable consider all the relevant factors including the robot hardware capabilities, task descriptions, etc, posing great challenges for design quality and efficiency. The mainstream practice o

Cited by 0SourceScholar
2024

BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

CVPR 2024poster

The recent progress in Large Language Models (LLM) has spurred various advancements in image-language conversation agents while how to build a proficient video-based dialogue system is still under exploration. Considering the extensive scale of LLM and visual backbone minimal GPU memory is left for…

2024

CableInspect-AD: An Expert-Annotated Anomaly Detection Dataset

NeurIPS 2024poster

Machine learning models are increasingly being deployed in real-world contexts. However, systematic studies on their transferability to specific and critical applications are underrepresented in the research literature. An important example is visual anomaly detection (VAD) for robotic power line in…

2024

CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges

ACL 2024long

Large Language Models (LLMs) have shown promise in automated code generation but typically excel only in simpler tasks such as generating standalone code units. However, real-world software development often involves complex code repositories with complex dependencies and extensive documentation. To…

2024

DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories

ACL 2024findings

How to evaluate the coding abilities of Large Language Models (LLMs) remains an open question. We find that existing benchmarks are poorly aligned with real-world code repositories and are insufficient to evaluate the coding abilities of LLMs.To address the knowledge gap, we propose a new benchmark…

2024

Distribution Guidance Network for Weakly Supervised Point Cloud Semantic Segmentation

NeurIPS 2024poster

Despite alleviating the dependence on dense annotations inherent to fully supervised methods, weakly supervised point cloud semantic segmentation suffers from inadequate supervision signals. In response to this challenge, we introduce a novel perspective that imparts auxiliary constraints by regulat…

Cited by 2SourcePDFScholar
2024

Efficient Point Cloud Attribute Compression Framework using Attribute-Guided Graph Fourier Transform

ICASSP 2024accepted

The Graph Fourier Transform (GFT) has achieved remarkable success in point cloud attribute compression due to its adaptability in handling irregular signals. However, the conventional graph-based attribute compression method mostly relies on geometry information to construct the Laplace matrix. In t…

Cited by 0SourceScholar
2024

Enhancing Code Generation Performance of Smaller Models by Distilling the Reasoning Ability of LLMs

COLING 2024main

Large Language Models (LLMs) have recently made significant advances in code generation through the ‘Chain-of-Thought’ prompting technique. This technique empowers the model to autonomously devise “solution plans” to tackle intricate programming challenges, thereby improving its performance in code…

2024

EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations

NeurIPS 2024poster

How to evaluate Large Language Models (LLMs) in code generation remains an open question. Many benchmarks have been proposed, but they have two limitations, i.e., data leakage and lack of domain-specific evaluation. The former hurts the fairness of benchmarks, and the latter hinders practitioners f…

Cited by 7SourcePDFScholar
2024

Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models

ACL 2024findings

Recent statements about the impressive capabilities of large language models (LLMs) are usually supported by evaluating on open-access benchmarks. Considering the vast size and wide-ranging sources of LLMs’ training data, it could explicitly or implicitly include test data, leading to LLMs being mor…

2024

HiRoPE: Length Extrapolation for Code Models Using Hierarchical Position

ACL 2024long

Addressing the limitation of context length in large language models for code-related tasks is the primary focus of this paper. Existing LLMs are constrained by their pre-trained context lengths, leading to performance issues in handling long complex code sequences. Inspired by how human programmers…

2024

Hot or Cold? Adaptive Temperature Sampling for Code Generation with Large Language Models

AAAI 2024technical

Recently, Large Language Models (LLMs) have shown impressive abilities in code generation. However, existing LLMs' decoding strategies are designed for Natural Language (NL) generation, overlooking the differences between NL and programming languages (PL). Due to this oversight, a better decoding st…

2024

Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection

CVPR 2024poster

Open vocabulary object detection (OVD) aims at seeking an optimal object detector capable of recognizing objects from both base and novel categories. Recent advances leverage knowledge distillation to transfer insightful knowledge from pre-trained large-scale vision-language models to the task of ob…

Cited by 16SourcePDFScholar
2024

Less Is More: Label Recommendation for Weakly Supervised Point Cloud Semantic Segmentation

AAAI 2024technical

Weak supervision has proven to be an effective strategy for reducing the burden of annotating semantic segmentation tasks in 3D space. However, unconstrained or heuristic weakly supervised annotation forms may lead to suboptimal label efficiency. To address this issue, we propose a novel label recom…

Cited by 20SourcePDFScholar
2024

MaIL: Improving Imitation Learning with Selective State Space Models

CoRL 2024poster

This work introduces Mamba Imitation Learning (MaIL), a novel imitation learning (IL) architecture that offers a computationally efficient alternative to state-of-the-art (SoTA) Transformer policies. Transformer-based policies have achieved remarkable results due to their ability in handling human-r…

Cited by 7SourceScholar
2024

Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning

ICLR 2024poster

Current advancements in reinforcement learning (RL) have predominantly focused on learning step-based policies that generate actions for each perceived state. While these methods efficiently leverage step information from environmental interaction, they often ignore the temporal correlation between…

2024

PACE: Improving Prompt with Actor-Critic Editing for Large Language Model

ACL 2024findings

Large language models (LLMs) have showcased remarkable potential across various tasks by conditioning on prompts. However, the quality of different human-written prompts leads to substantial discrepancies in LLMs’ performance, and improving prompts usually necessitates considerable human effort and…

Cited by 14SourcePDFScholar
2024

RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter

ACL 2024findings

Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries. To date, most of the state-of-the-art TVR methods learn image-to-video transfer learning based on the large-scale pre-trained vision-language models (e.g., CLIP). However, fully fine-tuning these pre-train…

Cited by 17SourcePDFScholar
2024

ScanPCGC: Learning-Based Lossless Point Cloud Geometry Compression using Sequential Slice Representation

ICASSP 2024accepted

The efficient storage and transportation requirements of point clouds promote the development of point cloud compression algorithms. In this paper, we develop a novel point cloud geometry compression using sequential slice representation. Unlike the limited contexts in conventional codecs and other…

Cited by 0SourceScholar
2024

Variational Distillation of Diffusion Policies into Mixture of Experts

NeurIPS 2024poster

This work introduces Variational Diffusion Distillation (VDD), a novel method that distills denoising diffusion policies into Mixtures of Experts (MoE) through variational inference. Diffusion Models are the current state-of-the-art in generative modeling due to their exceptional ability to accurate…

2023

Birder: Communication-Efficient 1-bit Adaptive Optimizer for Practical Distributed DNN Training

NeurIPS 2023poster

Various gradient compression algorithms have been proposed to alleviate the communication bottleneck in distributed learning, and they have demonstrated effectiveness in terms of high compression ratios and theoretical low communication complexity. However, when it comes to practically training mod…

Cited by 1SourcePDFScholar
2023

Causality Compensated Attention for Contextual Biased Visual Recognition

ICLR 2023poster

Visual attention does not always capture the essential object representation desired for robust predictions. Attention modules tend to underline not only the target object but also the common co-occurring context that the module thinks helpful in the training. The problem is rooted in the confoundin…

Cited by 23SourcePDFScholar
2023

Improving Graph Representation for Point Cloud Segmentation via Attentive Filtering

CVPR 2023poster

Recently, self-attention networks achieve impressive performance in point cloud segmentation due to their superiority in modeling long-range dependencies. However, compared to self-attention mechanism, we find graph convolutions show a stronger ability in capturing local geometry information with le…

Cited by 43SourcePDFScholar
2023

LIO-PPF: Fast LiDAR-Inertial Odometry via Incremental Plane Pre-Fitting and Skeleton Tracking

IROS 2023poster

As a crucial infrastructure of intelligent mobile robots, LiDAR-Inertial odometry (LIO) provides the basic capability of state estimation by tracking LiDAR scans. The high-accuracy tracking generally involves the k\text{NN}k\text{NN} search, which is used with minimizing the point-to-plane distance.…

Cited by 7SourcecodeScholar
2023

Null-Space Diffusion Sampling for Zero-Shot Point Cloud Completion

IJCAI 2023poster

Point cloud completion aims at estimating the complete data of objects from degraded observations. Despite existing completion methods achieving impressive performances, they rely heavily on degraded-complete data pairs for supervision. In this work, we propose a novel framework named Null-Space Dif…

Cited by 11SourcePDFScholar
2023

ProDMP: A Unified Perspective on Dynamic and Probabilistic Movement Primitives

RA-L 2023

Movement Primitives (MPs) are a well-known concept to represent and generate modular trajectories. MPs can be broadly categorized into two types: (a) dynamics-based approaches that generate smooth trajectories from any initial state, e. g., Dynamic Movement Primitives (DMPs), and (b) probabilistic a

Cited by 57SourceScholar
2023

Revisiting Temporal Modeling for CLIP-Based Image-to-Video Knowledge Transferring

CVPR 2023poster

Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention for their potential to improve visual representation learning in the video domain. In this paper, based on the CLIP model…

2023

Surface-Sampling Based Objective Quality Assessment Metrics for Meshes

ICASSP 2023accepted

In this paper, we prove that it is feasible to perform mesh quality assessment by sampling it into point cloud. We propose a general and efficient surface-sampling based framework that can deal with various types and levels of distortions with less complexity. In this method, the original and distor…

Cited by 0SourceScholar
2022

Contextual Debiasing for Visual Recognition With Causal Mechanisms

CVPR 2022poster

As a common problem in the visual world, contextual bias means the recognition may depend on the co-occurrence context rather than the objects themselves, which is even more severe in multi-label tasks due to multiple targets and the absence of location. Although some studies have focused on tacklin…

Cited by 46PDFcodeScholar
2022

DKNAS: A Practical Deep Keypoint Extraction Framework Based on Neural Architecture Search

ICRA 2022poster

Keypoint extraction including both keypoint detection and description is a fundamental step in a wide range of geometric multimedia applications. In recent years, many learning-based approaches for keypoint extraction emerge and achieve promising results. However, they usually design network archite…

Cited by 1SourceScholar
2022

Fine-Tuning Pre-Trained Language Models Effectively by Optimizing Subnetworks Adaptively

NeurIPS 2022accept

Large-scale pre-trained language models have achieved impressive results on a wide range of downstream tasks recently. However, fine-tuning an extremely large-scale pre-trained language model on limited target datasets is often plagued by overfitting and representation degradation. In this paper, we…

2022

JE2NET: Joint Exploitation and Exploration in Reinforcement Learning Based Image Restoration

ICASSP 2022accepted

Previous reinforcement learning (RL) based image restoration studies typically train RL agents to search for recovery tools from a constructed toolset and iteratively recover images. However, we argue that these agents rely on pre-trained RL models with fixed-length paths for restoration, which perf…

Cited by 0SourceScholar
2022

Local Surface Descriptor for Geometry and Feature Preserved Mesh Denoising

AAAI 2022technical

3D meshes are widely employed to represent geometry structure of 3D shapes. Due to limitation of scanning sensor precision and other issues, meshes are inevitably affected by noise, which hampers the subsequent applications. Convolultional neural networks (CNNs) achieve great success in image proces…

Cited by 10SourcePDFScholar
2022

Neural Texture Extraction and Distribution for Controllable Person Image Synthesis

CVPR 2022oral

We deal with the controllable person image synthesis task which aims to re-render a human from a reference image with explicit control over body pose and appearance. Observing that person images are highly structured, we propose to generate desired images by extracting and distributing semantic enti…

Cited by 94PDFcodeScholar
2022

OctAttention: Octree-Based Large-Scale Contexts Model for Point Cloud Compression

AAAI 2022technical

In point cloud compression, sufficient contexts are significant for modeling the point cloud distribution. However, the contexts gathered by the previous voxel-based methods decrease when handling sparse point clouds. To address this problem, we propose a multiple-contexts deep learning framework ca…

2022

Rethinking Positional Encoding in Tree Transformer for Code Representation

EMNLP 2022main

Transformers are now widely used in code representation, and several recent works further develop tree Transformers to capture the syntactic structure in source code. Specifically, novel tree positional encodings have been proposed to incorporate inductive bias into Transformer.In this work, we prop…

2022

SKFlow: Learning Optical Flow with Super Kernels

NeurIPS 2022accept

Optical flow estimation is a classical yet challenging task in computer vision. One of the essential factors in accurately predicting optical flow is to alleviate occlusions between frames. However, it is still a thorny problem for current top-performing optical flow estimation methods due to insuff…

2022

Self-Supervised Arbitrary-Scale Point Clouds Upsampling via Implicit Neural Representation

CVPR 2022poster

Point clouds upsampling is a challenging issue to generate dense and uniform point clouds from the given sparse input. Most existing methods either take the end-to-end supervised learning based manner, where large amounts of pairs of sparse input and dense ground-truth are exploited as supervision i…

Cited by 63PDFcodeScholar
2021

CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation

NeurIPS 2021poster

Benchmark datasets have a significant impact on accelerating research in programming language tasks. In this paper, we introduce CodeXGLUE, a benchmark dataset to foster machine learning research for program understanding and generation. CodeXGLUE includes a collection of 10 tasks across 14 datasets…

Cited by 981SourcecodeScholar
2021

Integrating Tree Path in Transformer for Code Representation

NeurIPS 2021poster

Learning distributed representation of source code requires modelling its syntax and semantics. Recent state-of-the-art models leverage highly structured source code representations, such as the syntax trees and paths therein. In this paper, we investigate two representative path encoding methods sh…

2021

Nested Error Map Generation Network for No-Reference Image Quality Assessment

ICASSP 2021accepted

We propose a multi-task learning neural network for No-Reference image quality assessment (NR-IQA). The pro-posed architecture consists of a backbone feature extractor, a nested multi-task generative module and a quality regression module. We adopt a coarse-to-fine strategy to predict objective erro…

Cited by 0SourceScholar
2021

PIRenderer: Controllable Portrait Image Generation via Semantic Neural Rendering

ICCV 2021poster

Generating portrait images by controlling the motions of existing faces is an important task of great consequence to social media industries. For easy use and intuitive control, semantically meaningful and fully disentangled parameters should be used as modifications. However, many existing techniqu…

Cited by 252PDFcodeScholar
2021

SSD-GAN: Measuring the Realness in the Spatial and Spectral Domains

AAAI 2021technical

This paper observes that there is an issue of high frequencies missing in the discriminator of standard GAN, and we reveal it stems from downsampling layers employed in the network architecture. This issue makes the generator lack the incentive from the discriminator to learn high-frequency content…

2021

Specializing Versatile Skill Libraries using Local Mixture of Experts

CoRL 2021poster

A long-cherished vision in robotics is to equip robots with skills that match the versatility and precision of humans. For example, when playing table tennis, a robot should be capable of returning the ball in various ways while precisely placing it at the desired location. A common approach to mod…

Cited by 41SourcecodeScholar
2021

Structure-Transformed Texture-Enhanced Network for Person Image Synthesis

ICCV 2021poster

Pose-guided virtual try-on task aims to modify the fashion item based on pose transfer task. These two tasks that belong to person image synthesis have strong correlations and similarities. However, existing methods treat them as two individual tasks and do not explore correlations between them. Mor…

Cited by 8PDFScholar
2020

C3DVQA: Full-Reference Video Quality Assessment with 3D Convolutional Neural Network

ICASSP 2020accepted

Traditional video quality assessment (VQA) methods evaluate localized picture quality and video score is predicted by temporally aggregating frame scores. However, video quality exhibits different characteristics from static image quality due to the existence of temporal masking effects. In this pap…

Cited by 0SourceScholar
2020

Deep Image Spatial Transformation for Person Image Generation

CVPR 2020poster

Pose-guided person image generation is to transform a source person image to a target pose. This task requires spatial manipulations of source data. However, Convolutional Neural Networks are limited by the lack of ability to spatially transform the inputs. In this paper, we propose a differentiable…

Cited by 231PDFcodeScholar
2020

NLocalSAT: Boosting Local Search with Solution Prediction

IJCAI 2020poster

The Boolean satisfiability problem (SAT) is a famous NP-complete problem in computer science. An effective way for solving a satisfiable SAT problem is the stochastic local search (SLS). However, in this method, the initialization is assigned in a random manner, which impacts the effectiveness of SL…

2020

ROIMIX: Proposal-Fusion Among Multiple Images for Underwater Object Detection

ICASSP 2020accepted

Generic object detection algorithms have proven their excellent performance in recent years. However, object detection on underwater datasets is still less explored. In contrast to generic datasets, underwater images usually have color shift and low contrast; sediment would cause blurring in underwa…

Cited by 0SourceScholar
2020

Regression Before Classification for Temporal Action Detection

ICASSP 2020accepted

Action classification combined with location regression is a widely-utilized mechanism in existing temporal action detection methods. However, there exists an inconsistency problem between locations and categories of action instances in this mechanism. More specifically, while the location of the pr…

Cited by 0SourceScholar
2019

AttPool: Towards Hierarchical Feature Representation in Graph Convolutional Networks via Attention Mechanism

ICCV 2019poster

Graph convolutional networks (GCNs) are potentially short of the ability to learn hierarchical representation for graph embedding, which holds them back in the graph classification task. Here, we propose AttPool, which is a novel graph pooling module based on attention mechanism, to remedy the probl…

Cited by 87PDFcodeScholar
2019

BLP - Boundary Likelihood Pinpointing Networks for Accurate Temporal Action Localization

ICASSP 2019accepted

Despite tremendous progress achieved in temporal action detection, state-of-the-art methods still suffer from the sharp performance deterioration when localizing the starting and ending temporal action boundaries. Although most methods apply boundary regression paradigm to tackle this problem, we ar…

Cited by 0SourceScholar
2019

Boundary Information Matters More: Accurate Temporal Action Detection with Temporal Boundary Network

ICASSP 2019accepted

Temporal action detection in untrimmed videos is an important yet challenging task. How to locate complex actions accurately is still an open question due to the ambiguous boundaries between action instances and the background. Recently a newly proposed work exploits Structured Segment Networks (SSN…

Cited by 0SourceScholar
2019

Graph Convolutional Label Noise Cleaner: Train a Plug-And-Play Action Classifier for Anomaly Detection

CVPR 2019poster

Video anomaly detection under weak labels is formulated as a typical multiple-instance learning problem in previous works. In this paper, we provide a new perspective, i.e., a supervised learning task under noisy labels. In such a viewpoint, as long as cleaning away label noise, we can directly appl…

Cited by 590PDFcodeScholar
2019

Multi-mapping Image-to-Image Translation via Learning Disentanglement

NeurIPS 2019poster

Recent advances of image-to-image translation focus on learning the one-to-many mapping from two aspects: multi-modal translation and multi-domain translation. However, the existing methods only consider one of the two perspectives, which makes them unable to solve each other's problem. To address t…

2019

Multi-step Self-attention Network for Cross-modal Retrieval Based on a Limited Text Space

ICASSP 2019accepted

Cross-modal retrieval has been recently proposed to find an appropriate subspace where the similarity among different modalities, such as image and text, can be directly measured. In this paper, we propose Multi-step Self-Attention Network (MSAN) to perform cross-modal retrieval in a limited text sp…

Cited by 0SourceScholar
2019

StructureFlow: Image Inpainting via Structure-Aware Appearance Flow

ICCV 2019poster

Image inpainting techniques have shown significant improvements by using deep neural networks recently. However, most of them may either fail to reconstruct reasonable structures or restore fine-grained textures. In order to solve this problem, in this paper, we propose a two-stage model which split…

Cited by 455PDFcodeScholar
2019

Why Do Neural Dialog Systems Generate Short and Meaningless Replies? a Comparison between Dialog and Translation

ICASSP 2019accepted

This paper addresses the question: In neural dialog systems, why do sequence-to-sequence (Seq2Seq) neural networks generate short and meaningless replies for open-domain response generation? We conjecture that in a dialog system, due to the randomness of spoken language, there may be multiple equall…

Cited by 0SourceScholar
2017

ORGB: Offset correction in RGB color space for illumination-robust image processing

ICASSP 2017accepted

Single materials have colors which form straight lines in RGB space. However, in severe shadow cases, those lines do not intersect the origin, which is inconsistent with the description of most literature. This paper is concerned with the detection and correction of the offset between the intersecti…

Cited by 0SourceScholar