← Search

yang song

130 accepted papers

2026

FOCA: Frequency-Oriented Cross-Domain Forgery Detection, Localization and Explanation via Multi-Modal Large Language Model

ICASSP 2026poster

Advances in image tampering techniques, particularly generative models, pose significant challenges to media verification, digital forensics, and public trust. Existing image forgery detection and localization (IFDL) methods suffer from two key limitations: over-reliance on semantic content while ne…

Cited by 0SourcePDFScholar
2026

Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation

ICML 2026poster

Reverse Chain-of-Thought Generation (RCG) synthesizes reasoning traces from query-answer pairs, but runs the risk of producing post-hoc rationalizations: when models can see the answer during generation, the answer serves as a cognitive anchor that shapes the entire explanation. We formalize this ph…

Cited by 0SourceScholar
2026

NEMO: Execution-Aware Optimization Modeling via Autonomous Coding Agents

ICML 2026poster

In this paper, we present **NEMO**, a system that translates **N**atural-language descriptions of decision problems into formal **E**xecutable **M**athematical **O**ptimization implementations, operating collaboratively with users or autonomously. Existing approaches typically rely on specialized la…

Cited by 0SourceScholar
2026

PositionIC: Unified Position and Identity Consistency for Image Customization

CVPR 2026

Recent subject-driven image customization excels in fidelity, yet fine-grained instance-level spatial control remains an elusive challenge, hindering real-world applications. This limitation stems from two factors: a scarcity of scalable, position-annotated datasets, and the entanglement of identity

Cited by 0SourcecodeScholar
2026

Principles2Plan: LLM-Guided System for Operationalising Ethical Principles into Plans

AAAI 2026technical

Ethical awareness is critical for robots operating in human environments, yet existing automated planning tools provide little support. Manually specifying ethical rules is labour-intensive and highly context-specific. We present Principles2Plan, an interactive research prototype demonstrating how a

Cited by 0SourcePDFScholar
2026

Relative Advantage Debiasing for Watch-Time Prediction in Short-Video Recommendation

AAAI 2026technical

Watch time is widely used as a proxy for user satisfaction in video recommendation platforms. However, raw watch times are influenced by confounding factors such as video duration, popularity, and individual user behaviors, potentially distorting preference signals and resulting in biased recommenda

Cited by 0SourcePDFScholar
2026

SpeeDiff: Scalable Pixel-Anchored End-to-End Latent Diffusion Model

CVPR 2026

We present Scalable Pixel-anchored End-to-end Diffusion (SpeeDiff), a latent diffusion method that jointly trains the VAE and the diffusion model from scratch. In principle, joint training allows the diffusion loss gradient to directly guide the VAE encoder, encouraging the formation of a generation

Cited by 0SourceScholar
2026

iFire AI: AI-powered Wildfire Simulation and 3D Immersive Visualisation

IJCAI 2026

Wildfires, especially extreme wildfires, cause irreversible damage to ecosystems, human lives and economies globally. To reduce such losses, understanding wildfires is crucial for effective preparedness. This research proposal introduces iFire AI, a collaborative project aimed at developing world's

Cited by 0Scholar
2026

iLoRA: Bayesian Low-Rank Adaptation with Latent Interaction Graphs for Microbiome Diagnosis

ICML 2026poster

Reliable microbiome-based diagnosis is critical for precision medicine at scale in inflammatory diseases, yet current post-training pipelines in LLMs often overlook the interaction structure that governs microbial ecosystems. In inflammatory bowel disease (IBD), disease signals arise not only from s…

Cited by 0SourceScholar
2025

BiCo-Fusion: Bidirectional Complementary LiDAR-Camera Fusion for Semantic- and Spatial-Aware 3D Object Detection

RA-L 2025

3D object detection is an important task that has been widely applied in autonomous driving. To perform this task, a new trend is to fuse multi-modal inputs, i.e., LiDAR and camera. Under such a trend, recent methods fuse these two modalities by unifying them in the same 3D space. However, during di

Cited by 16SourceScholar
2025

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model

ICML 2025poster

Since the debut of DPO, it has been shown that aligning a target LLM with human preferences via the KL-constrained RLHF loss is mathematically equivalent to a special kind of reward modeling task. Concretely, the task requires: 1) using the target LLM to parameterize the reward model, and 2) tuning…

Cited by 1SourcePDFScholar
2025

Enhancing Change Detection in Remote Sensing: Integrating Synthetic Data with Semi-Supervised Learning

ICASSP 2025accepted

Change detection (CD) in remote sensing is a crucial yet challenging task, particularly due to the labor-intensive nature of labeling bi-temporal images. We introduce a novel framework that leverages synthetic datasets, style transfer, and semi-supervised learning to enhance CD model performance whi…

Cited by 0SourceScholar
2025

Following the Autoregressive Nature of LLM Embeddings via Compression and Alignment

EMNLP 2025

A new trend uses LLMs as dense text encoders via contrastive learning. However, since LLM embeddings predict the probability distribution of the next token, they are inherently generative and distributive, conflicting with contrastive learning, which requires embeddings to capture full-text semantic

2025

GRPose: Learning Graph Relations for Human Image Generation with Pose Priors

AAAI 2025technical

Recent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. However, existing efforts are still struggling to generate high-quality images with consistent pose alignment, resulting in unsatisfactory output. In this…

2025

GVPO: Group Variance Policy Optimization for Large Language Model Post-Training

NeurIPS 2025poster

Post-training plays a crucial role in refining and aligning large language models to meet specific tasks and human preferences. While recent advancements in post-training techniques, such as Group Relative Policy Optimization (GRPO), leverage increased sampling with relative reward scoring to achiev…

Cited by 0SourceScholar
2025

Granularity-Adaptive Spatial Evidence Tokenization for Video Question Answering

AAAI 2025technical

Video question answering plays a vital role in computer vision, and recent advances in large language models have further propelled the development of this field. However, existing video question answering techniques often face limitations in grasping fine-grained video content in spatial dimensions…

Cited by 0SourcePDFScholar
2025

Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing

CVPR 2025poster

Diffusion models have recently achieved success in solving Bayesian inverse problems with learned data priors. Current methods build on top of the diffusion sampling process, where each denoising step makes small modifications to samples from the previous step. However, this process struggles to cor…

2025

Interpretable Image Classification via Non-parametric Part Prototype Learning

CVPR 2025poster

Classifying images with an interpretable decision-making process is a long-standing problem in computer vision. In recent years, Prototypical Part Networks has gained traction as an approach for self-explainable neural networks, due to their ability to mimic human visual reasoning by providing expla…

2025

KG-Agent: An Efficient Autonomous Agent Framework for Complex Reasoning over Knowledge Graph

ACL 2025long

In this paper, we aim to improve the reasoning ability of large language models(LLMs) over knowledge graphs(KGs) to answer complex questions. Inspired by existing methods that design the interaction strategy between LLMs and KG, we propose an autonomous LLM-based agent framework, called KG-Agent, wh…

2025

LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert Space

NeurIPS 2025poster

Image geolocalization is a fundamental yet challenging task, aiming at inferring the geolocation on Earth where an image is taken. State-of-the-art methods employ either grid-based classification or gallery-based image-location retrieval, whose spatial generalizability significantly suffers if the s…

Cited by 0SourceScholar
2025

Lock on Target! Precision Unlearning via Directional Control

EMNLP 2025

The unlearning method aims at effectively removing harmful, sensitive, or outdated knowledge without costly retraining the model. However, existing methods suffer from two critical limitations: (1) collateral forgetting, where erasing target data inadvertently removes related but desirable knowledge

Cited by 0SourcePDFScholar
2025

MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny Objects

CVPR 2025poster

We present MANTA, a visual-text anomaly detection dataset for tiny objects. The visual component comprises over 137.3K images across 38 object categories spanning five typical domains, of which 8.6K images are labeled as anomalous with pixel-level annotations. Each image is captured from five distin…

Cited by 1SourcePDFScholar
2025

Making Transformer Decoders Better Differentiable Indexers

ICLR 2025poster

Retrieval aims to find the top-k items most relevant to a query/user from a large dataset. Traditional retrieval models represent queries/users and items as embedding vectors and use Approximate Nearest Neighbor (ANN) search for retrieval. Recently, researchers have proposed a generative-based retri…

Cited by 0SourcePDFScholar
2025

Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format Alignment

ICLR 2025poster

Adapting large language models (LLMs) to specialized domains typically requires domain-specific corpora for continual pre-training to facilitate knowledge memorization and related instructions for fine-tuning to apply this knowledge. However, this method may lead to inefficient knowledge memorizatio…

Cited by 1SourcePDFScholar
2025

Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance

IJCAI 2025

Multimodal pathology-genomic analysis has become increasingly prominent in cancer survival prediction. However, existing studies mainly utilize multi-instance learning to aggregate patch-level features, neglecting the information loss of contextual and hierarchical details within pathology images. F

2025

PEACE: Empowering Geologic Map Holistic Understanding with MLLMs

CVPR 2025poster

Geologic map, as a fundamental diagram in geology science, provides critical insights into the structure and composition of Earth's subsurface and surface. These maps are indispensable in various fields, including disaster assessment, resource exploration, and civil engineering. Despite their signif…

2025

Preference-Oriented Supervised Fine-Tuning: Favoring Target Model over Aligned Large Language Models

AAAI 2025technical

Alignment, endowing a pre-trained Large language model (LLM) with the ability to follow instructions, is crucial for its real-world applications. Conventional supervised fine-tuning (SFT) methods formalize it as causal language modeling typically with a cross-entropy objective, requiring a large am…

2025

Prototype-Based Image Prompting for Weakly Supervised Histopathological Image Segmentation

CVPR 2025poster

Weakly supervised image segmentation with image-level labels has drawn attention due to the high cost of pixel-level annotations. Traditional methods using Class Activation Maps (CAMs) often highlight only the most discriminative regions, leading to incomplete masks. Recent approaches that introduce…

2025

Pyramidal Flow Matching for Efficient Video Generative Modeling

ICLR 2025poster

Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational de…

2025

RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement

NAACL 2025long

Existing large language models (LLMs) show exceptional problem-solving capabilities but might struggle with complex reasoning tasks. Despite the successes of chain-of-thought and tree-based search methods, they mainly depend on the internal knowledge of LLMs to search over intermediate reasoning ste…

2025

ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability

ICLR 2025spotlight

Retrieval-Augmented Generation (RAG) models are designed to incorporate external knowledge, reducing hallucinations caused by insufficient parametric (internal) knowledge. However, even with accurate and relevant retrieved content, RAG models can still produce hallucinations by generating outputs th…

Cited by 8SourcePDFScholar
2025

RecFlow: An Industrial Full Flow Recommendation Dataset

ICLR 2025poster

Industrial recommendation systems (RS) rely on the multi-stage pipeline to balance effectiveness and efficiency when delivering items from a vast corpus to users. Existing RS benchmark datasets primarily focus on the exposure space, where novel RS algorithms are trained and evaluated. However, when…

2025

Salvaging the Overlooked: Leveraging Class-Aware Contrastive Learning for Multi-Class Anomaly Detection

ICCV 2025poster

For anomaly detection (AD), early approaches often train separate models for individual classes, yielding high performance but posing challenges in scalability and resource management. Recent efforts have shifted toward training a single model capable of handling multiple classes. However, directly…

Cited by 0SourcePDFScholar
2025

SynerGuard: A Robust Framework for Point Cloud Classification via Local Geometry and Spatial Topology

ICRA 2025

Point cloud recognition models are known to be vulnerable to adversarial attacks. The state-of-the-art defense solutions either focus on partial features of the point cloud, limiting their effectiveness, or rely heavily on known adversarial examples, reducing their generalizability, while others, li

Cited by 2SourceScholar
2025

Trigger3:Refining Query Correction via Adaptive Model Selector

AAAI 2025technical

In search scenarios, user experience can be hindered by erroneous queries due to typos, voice errors, or knowledge gaps. Therefore, query correction is crucial for search engines. Current correction models, usually small models trained on specific data, often struggle with queries beyond their train…

2024

AV4GAInsp: An Efficient Dual-Camera System for Identifying Defective Kernels of Cereal Grains

RA-L 2024

Grain Appearance Inspection (GAI) is a pre-requisite for grain quality determination, providing guidance for grain processing, storage, and trade. GAI is routinely performed by trained inspectors who are required to visually inspect cereal grains for each individual kernel. Since grain kernels (e.g.

Cited by 10SourceScholar
2024

Coupling Graph Neural Networks with Fractional Order Continuous Dynamics: A Robustness Study

AAAI 2024technical

In this work, we rigorously investigate the robustness of graph neural fractional-order differential equation (FDE) models. This framework extends beyond traditional graph neural (integer-order) ordinary differential equation (ODE) models by implementing the time-fractional Caputo derivative. Utiliz…

Cited by 6SourcePDFScholar
2024

Decoupled Optimisation for Long-Tailed Visual Recognition

AAAI 2024technical

When training on a long-tailed dataset, conventional learning algorithms tend to exhibit a bias towards classes with a larger sample size. Our investigation has revealed that this biased learning tendency originates from the model parameters, which are trained to disproportionately contribute to the…

Cited by 6SourcePDFScholar
2024

DistilVPR: Cross-Modal Knowledge Distillation for Visual Place Recognition

AAAI 2024technical

The utilization of multi-modal sensor data in visual place recognition (VPR) has demonstrated enhanced performance compared to single-modal counterparts. Nonetheless, integrating additional sensors comes with elevated costs and may not be feasible for systems that demand lightweight operation, there…

2024

Enhancing Job Recommendation through LLM-Based Generative Adversarial Networks

AAAI 2024technical

Recommending suitable jobs to users is a critical task in online recruitment platforms. While existing job recommendation methods encounter challenges such as the low quality of users' resumes, which hampers their accuracy and practical effectiveness.With the rapid development of large language mode…

Cited by 60SourcePDFScholar
2024

Federated Adaptation for Foundation Model-based Recommendations

IJCAI 2024poster

With the recent success of large language models, particularly foundation models with generalization abilities, applying foundation models for recommendations becomes a new paradigm to improve existing recommendation systems. It becomes a new open challenge to enable the foundation model to capture…

2024

Formalisation and Evaluation of Properties for Consequentialist Machine Ethics

IJCAI 2024poster

As artificial intelligence (AI) technologies continue to influence our daily lives, there has been a growing need to ensure that AI enabled decision making systems adhere to principles expected of human decision makers. This need has given rise to the area of Machine Ethics. We formalise several eth…

Cited by 1SourcePDFScholar
2024

Fully Decoupling Trajectory and Scene Encoding for Lightweight Heatmap-Oriented Trajectory Prediction

RA-L 2024

Recently, heatmap-oriented approaches have demonstrated their state-of-the-art performance in pedestrian trajectory prediction by exploiting scene information from input images before running the encoder. To align the image and trajectory information, existing methods centre the scene images to agen

Cited by 6SourceScholar
2024

Fully Distributed, Flexible Compositional Visual Representations via Soft Tensor Products

NeurIPS 2024poster

Since the inception of the classicalist vs. connectionist debate, it has been argued that the ability to systematically combine symbol-like entities into compositional representations is crucial for human intelligence. In connectionist systems, the field of disentanglement has gained prominence for…

2024

Mixture of In-Context Experts Enhance LLMs' Long Context Awareness

NeurIPS 2024poster

Many studies have revealed that large language models (LLMs) exhibit uneven awareness of different contextual positions. Their limited context awareness can lead to overlooking critical information and subsequent task failures. While several approaches have been proposed to enhance LLMs' context awa…

2024

PosDiffNet: Positional Neural Diffusion for Point Cloud Registration in a Large Field of View with Perturbations

AAAI 2024technical

Point cloud registration is a crucial technique in 3D computer vision with a wide range of applications. However, this task can be challenging, particularly in large fields of view with dynamic objects, environmental noise, or other perturbations. To address this challenge, we propose a model called…

2024

RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance

NeurIPS 2024poster

Customizing diffusion models to generate identity-preserving images from user-provided reference images is an intriguing new problem. The prevalent approaches typically require training on extensive domain-specific images to achieve identity preservation, which lacks flexibility across different use…

2024

Refining Airway Segmentation Through Breakage Filling and Leakage Reduction Using Point Clouds

IROS 2024poster

Bronchoscopy reveals air passages and internal tissues for accurate diagnosis of various lung diseases. Robot-assisted bronchoscopy using an airway tree model can help path planning before surgery and navigation during surgery. In airway tree modeling, though volumetric deep learning methods have ac…

Cited by 0SourceScholar
2024

Rewiring the Transformer with Depth-Wise LSTMs

COLING 2024main

Stacking non-linear layers allows deep neural networks to model complicated functions, and including residual connections in Transformer layers is beneficial for convergence and performance. However, residual connections may make the model “forget” distant layers and fail to fuse information from pr…

Cited by 2SourcePDFScholar
2024

Spontts: Modeling and Transferring Spontaneous Style for TTS

ICASSP 2024accepted

Spontaneous speaking style exhibits notable differences from other speaking styles due to various spontaneous phenomena (e.g., filled pauses, prolongation) and substantial prosody variation (e.g., diverse pitch and duration variation, occasional non-verbal speech like a smile), posing challenges to…

Cited by 0SourceScholar
2024

Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image Generation

CVPR 2024poster

Vanilla text-to-image diffusion models struggle with generating accurate human images commonly resulting in imperfect anatomies such as unnatural postures or disproportionate limbs. Existing methods address this issue mostly by fine-tuning the model with extra images or adding additional controls --…

Cited by 9SourcePDFScholar
2024

Unleashing the Potential of Fractional Calculus in Graph Neural Networks with FROND

ICLR 2024spotlight

We introduce the FRactional-Order graph Neural Dynamical network (FROND), a new continuous graph neural network (GNN) framework. Unlike traditional continuous GNNs that rely on integer-order differential equations, FROND employs the Caputo fractional derivative to leverage the non-local properties o…

2024

Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization

ICML 2024oral

In light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos. Compared to static images, video poses unique challenges for effective large-scale pre-training due to the modeling of its…

2024

Your Career Path Matters in Person-Job Fit

AAAI 2024technical

We are again confronted with one of the most vexing aspects of the advancement of technology: automation and AI technology cause the devaluation of human labor, resulting in unemployment. With this background, automatic person-job fit systems are promising solutions to promote the employment rate. T…

2023

Adversarial Robustness in Graph Neural Networks: A Hamiltonian Approach

NeurIPS 2023spotlight

Graph neural networks (GNNs) are vulnerable to adversarial perturbations, including those that affect both node features and graph topology. This paper investigates GNNs derived from diverse neural flows, concentrating on their connection to various stability notions such as BIBO stability, Lyapunov…

2023

Graph Neural Convection-Diffusion with Heterophily

IJCAI 2023poster

Graph neural networks (GNNs) have shown promising results across various graph learning tasks, but they often assume homophily, which can result in poor performance on heterophilic graphs. The connected nodes are likely to be from different classes or have dissimilar features on heterophilic graphs.…

2023

HypLiLoc: Towards Effective LiDAR Pose Regression With Hyperbolic Fusion

CVPR 2023poster

LiDAR relocalization plays a crucial role in many fields, including robotics, autonomous driving, and computer vision. LiDAR-based retrieval from a database typically incurs high computation storage costs and can lead to globally inaccurate pose estimations if the database is too sparse. On the othe…

2023

HyperTraj: Towards Simple and Fast Scene-Compliant Endpoint Conditioned Trajectory Prediction

IROS 2023poster

An important task in trajectory prediction is to model the uncertainty of agents' motions, which requires the system to propose multiple plausible future trajectories for agents based on their past movements. Recently, many approaches have been developed following an endpointconditioned deep learnin…

Cited by 2SourceScholar
2023

Maximizing Spatio-Temporal Entropy of Deep 3D CNNs for Efficient Video Recognition

ICLR 2023poster

3D convolution neural networks (CNNs) have been the prevailing option for video recognition. To capture the temporal information, 3D convolutions are computed along the sequences, leading to cubically growing and expensive computations. To reduce the computational cost, previous methods resort to ma…

2023

Node Embedding from Neural Hamiltonian Orbits in Graph Neural Networks

ICML 2023poster

In the graph node embedding problem, embedding spaces can vary significantly for different data types, leading to the need for different GNN model types. In this paper, we model the embedding update of a node feature as a Hamiltonian orbit over time. Since the Hamiltonian orbits generalize the expon…

2022

A Right Invariant Extended Kalman Filter for Object Based SLAM

RA-L 2022

With the recent advance of deep learning based object recognition and estimation, it is possible to consider object level SLAM where the pose of each object is estimated in the SLAM process. In this letter, based on a novel Lie group structure, a right invariant extended Kalman filter (RI-EKF) for o

Cited by 37SourceScholar
2022

Density Ratio Estimation via Infinitesimal Classification

AISTATS 2022poster

Density ratio estimation (DRE) is a fundamental machine learning technique for comparing two probability distributions. However, existing methods struggle in high-dimensional settings, as it is difficult to accurately compare probability distributions based on finite samples. In this work we propose…

2022

GeoDiff: A Geometric Diffusion Model for Molecular Conformation Generation

ICLR 2022oral

Predicting molecular conformations from molecular graphs is a fundamental problem in cheminformatics and drug discovery. Recently, significant progress has been achieved with machine learning approaches, especially with deep generative models. Inspired by the diffusion process in classical non-equil…

2022

GrainSpace: A Large-Scale Dataset for Fine-Grained and Domain-Adaptive Recognition of Cereal Grains

CVPR 2022poster

Cereal grains are a vital part of human diets and are important commodities for people's livelihood and international trade. Grain Appearance Inspection (GAI) serves as one of the crucial steps for the determination of grain quality and grain stratification for proper circulation, storage and food p…

Cited by 25PDFcodeScholar
2022

Graph-Based Spatial Transformer With Memory Replay for Multi-Future Pedestrian Trajectory Prediction

CVPR 2022poster

Pedestrian trajectory prediction is an essential and challenging task for a variety of real-life applications such as autonomous driving and robotic motion planning. Besides generating a single future path, predicting multiple plausible future paths is becoming popular in some recent work on traject…

Cited by 92PDFcodeScholar
2022

On the Robustness of Graph Neural Diffusion to Topology Perturbations

NeurIPS 2022accept

Neural diffusion on graphs is a novel class of graph neural networks that has attracted increasing attention recently. The capability of graph neural partial differential equations (PDEs) in addressing common hurdles of graph neural networks (GNNs), such as the problems of over-smoothing and bottlen…

2022

SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

ICLR 2022poster

Guided image synthesis enables everyday users to create and edit photo-realistic images with minimum effort. The key challenge is balancing faithfulness to the user inputs (e.g., hand-drawn colored strokes) and realism of the synthesized images. Existing GAN-based methods attempt to achieve such bal…

2022

Solving Inverse Problems in Medical Imaging with Score-Based Generative Models

ICLR 2022poster

Reconstructing medical images from partial measurements is an important inverse problem in Computed Tomography (CT) and Magnetic Resonance Imaging (MRI). Existing solutions based on machine learning typically train a model to directly map measurements to medical images, leveraging a training dataset…

2021

Accelerating Feedforward Computation via Parallel Nonlinear Equation Solving

ICML 2021spotlight

Feedforward computation, such as evaluating a neural network or sampling from an autoregressive model, is ubiquitous in machine learning. The sequential nature of feedforward computation, however, requires a strict order of execution and cannot be easily accelerated with parallel computing. To enabl…

2021

Anytime Sampling for Autoregressive Models via Ordered Autoencoding

ICLR 2021poster

Autoregressive models are widely used for tasks such as image and audio generation. The sampling process of these models, however, does not allow interruptions and cannot adapt to real-time computational resources. This challenge impedes the deployment of powerful autoregressive models, which involv…

2021

CSDI: Conditional Score-based Diffusion Models for Probabilistic Time Series Imputation

NeurIPS 2021poster

The imputation of missing values in time series has many applications in healthcare and finance. While autoregressive models are natural candidates for time series imputation, score-based diffusion models have recently outperformed existing counterparts including autoregressive models in many tasks…

2021

Error-Correcting Output Codes with Ensemble Diversity for Robust Learning in Neural Networks

AAAI 2021technical

Though deep learning has been applied successfully in many scenarios, malicious inputs with human-imperceptible perturbations can make it vulnerable in real applications. This paper proposes an error-correcting neural network (ECNN) that combines a set of binary classifiers to combat adversarial exa…

2021

Estimating High Order Gradients of the Data Distribution by Denoising

NeurIPS 2021poster

The first order derivative of a data density can be estimated efficiently by denoising score matching, and has become an important component in many applications, such as image generation and audio synthesis. Higher order derivatives provide additional local information about the data distribution a…

Cited by 51SourcePDFScholar
2021

Exploiting Edge-Oriented Reasoning for 3D Point-Based Scene Graph Analysis

CVPR 2021poster

Scene understanding is a critical problem in computer vision. In this paper, we propose a 3D point-based scene graph generation (SGGpoint) framework to effectively bridge perception and reasoning to achieve scene understanding via three sequential stages, namely scene graph construction, reasoning,…

Cited by 65PDFScholar
2021

Imitation with Neural Density Models

NeurIPS 2021poster

We propose a new framework for Imitation Learning (IL) via density estimation of the expert's occupancy measure followed by Maximum Occupancy Entropy Reinforcement Learning (RL) using the density as a reward. Our approach maximizes a non-adversarial model-free RL objective that provably lower bounds…

Cited by 15SourcePDFScholar
2021

Improved Autoregressive Modeling with Distribution Smoothing

ICLR 2021oral

While autoregressive models excel at image compression, their sample quality is often lacking. Although not realistic, generated images often have high likelihood according to the model, resembling the case of adversarial examples. Inspired by a successful adversarial defense method, we incorporate…

Cited by 23SourcePDFScholar
2021

Improved Word Sense Disambiguation with Enhanced Sense Representations

EMNLP 2021finding

Current state-of-the-art supervised word sense disambiguation (WSD) systems (such as GlossBERT and bi-encoder model) yield surprisingly good results by purely leveraging pre-trained language models and short dictionary definitions (or glosses) of the different word senses. While concise and intuitiv…

2021

Learning Energy-Based Models by Diffusion Recovery Likelihood

ICLR 2021poster

While energy-based models (EBMs) exhibit a number of desirable properties, training and sampling on high-dimensional datasets remains challenging. Inspired by recent progress on diffusion probabilistic models, we present a diffusion recovery likelihood method to tractably learn and sample from a seq…

2021

Maximum Likelihood Training of Score-Based Diffusion Models

NeurIPS 2021spotlight

Score-based diffusion models synthesize samples by reversing a stochastic process that diffuses data to noise, and are trained by minimizing a weighted combination of score matching losses. The log-likelihood of score-based diffusion models can be tractably computed through a connection to continuou…

2021

PhotoChat: A Human-Human Dialogue Dataset With Photo Sharing Behavior For Joint Image-Text Modeling

ACL 2021long

We present a new human-human dialogue dataset - PhotoChat, the first dataset that casts light on the photo sharing behavior in online messaging. PhotoChat contains 12k dialogues, each of which is paired with a user photo that is shared during the conversation. Based on this dataset, we propose two t…

2021

Score-Based Generative Modeling through Stochastic Differential Equations

ICLR 2021oral

Creating noise from data is easy; creating data from noise is generative modeling. We present a stochastic differential equation (SDE) that smoothly transforms a complex data distribution to a known prior distribution by slowly injecting noise, and a corresponding reverse-time SDE that transforms th…

2021

Stable Neural ODE with Lyapunov-Stable Equilibrium Points for Defending Against Adversarial Attacks

NeurIPS 2021poster

Deep neural networks (DNNs) are well-known to be vulnerable to adversarial attacks, where malicious human-imperceptible perturbations are included in the input to the deep network to fool it into making a wrong classification. Recent studies have demonstrated that neural Ordinary Differential Equati…

Cited by 111SourcePDFScholar
2021

Walk in the Cloud: Learning Curves for Point Clouds Shape Analysis

ICCV 2021poster

Discrete point cloud objects lack sufficient shape descriptors of 3D geometries. In this paper, we present a novel method for aggregating hypothetical curves in point clouds. Sequences of connected points (curves) are initially grouped by taking guided walks in the point clouds, and then subsequentl…

Cited by 367PDFcodeScholar
2020

Diversity can be Transferred: Output Diversification for White- and Black-box Attacks

NeurIPS 2020poster

Adversarial attacks often involve random perturbations of the inputs drawn from uniform or Gaussian distributions, e.g. to initialize optimization-based white-box attacks or generate update directions in black-box attacks. These simple perturbations, however, could be sub-optimal as they are agnosti…

2020

Efficient Learning of Generative Models via Finite-Difference Score Matching

NeurIPS 2020poster

Several machine learning applications involve the optimization of higher-order derivatives (e.g., gradients of gradients) during training, which can be expensive with respect to memory and computation even with automatic differentiation. As a typical example in generative modeling, score matching~(S…

2020

Improving 3D Object Detection through Progressive Population Based Augmentation

ECCV 2020poster

Data augmentation has been widely adopted for object detection in 3D point clouds. However, all previous related efforts have focused on manually designing specific data augmentation methods for individual architectures. In this work, we present the first attempt to automate the design of data augme…

Cited by 94SourcePDFScholar
2020

Permutation Invariant Graph Generation via Score-Based Generative Modeling

AISTATS 2020poster

Learning generative models for graph-structured data is challenging because graphs are discrete, combinatorial, and the underlying data distribution is invariant to the ordering of nodes. However, most of the existing generative models for graphs are not invariant to the chosen ordering, which might…

2020

Training Deep Energy-Based Models with f-Divergence Minimization

ICML 2020poster

Deep energy-based models (EBMs) are very flexible in distribution parametrization but computationally challenging because of the intractable partition function. They are typically trained via maximum likelihood, using contrastive divergence to approximate the gradient of the KL divergence between da…

2020

Unsupervised Instance Segmentation in Microscopy Images via Panoptic Domain Adaptation and Task Re-Weighting

CVPR 2020poster

Unsupervised domain adaptation (UDA) for nuclei instance segmentation is important for digital pathology, as it alleviates the burden of labor-intensive annotation and domain shift across datasets. In this work, we propose a Cycle Consistency Panoptic Domain Adaptive Mask R-CNN (CyC-PDAM) architectu…

Cited by 98PDFcodeScholar
2019

Class-Balanced Loss Based on Effective Number of Samples

CVPR 2019poster

With the rapid increase of large-scale, real-world datasets, it becomes critical to address the problem of long-tailed data distribution (i.e., a few classes account for most of the data, while most classes are under-represented). Existing solutions typically adopt class re-balancing strategies such…

Cited by 3212PDFcodeScholar
2019

Efficient Graph Generation with Graph Recurrent Attention Networks

NeurIPS 2019poster

We propose a new family of efficient and expressive deep generative models of graphs, called Graph Recurrent Attention Networks (GRANs). Our model generates graphs one block of nodes and associated edges at a time. The block size and sampling stride allow us to trade off sample quality for efficienc…

2019

MintNet: Building Invertible Neural Networks with Masked Convolutions

NeurIPS 2019poster

We propose a new way of constructing invertible neural networks by combining simple building blocks with a novel set of composition rules. This leads to a rich set of invertible architectures, including those similar to ResNets. Inversion is achieved with a locally convergent iterative procedure tha…

2019

Sliced Score Matching: A Scalable Approach to Density and Score Estimation

UAI 2019poster

Score matching is a popular method for estimating unnormalized statistical models. However, it has been so far limited to simple, shallow models or low-dimensional data, due to the difficulty of computing the Hessian of log-density functions. We show this difficulty can be mitigated by projecting th…

2018

Constructing Unrestricted Adversarial Examples with Generative Models

NeurIPS 2018poster

Adversarial examples are typically constructed by perturbing an existing data point within a small matrix norm, and current defense methods are focused on guarding against this type of attack. In this paper, we propose a new class of adversarial examples that are synthesized entirely from scratch us…

2018

Large Scale Fine-Grained Categorization and Domain-Specific Transfer Learning

CVPR 2018poster

Transferring the knowledge learned from large scale datasets (e.g., ImageNet) via fine-tuning offers an effective solution for domain-specific fine-grained visual categorization (FGVC) tasks (e.g., recognizing bird species or car make & model). In such scenarios, data annotation often calls for spec…

Cited by 656SourcePDFScholar
2018

No-Reference Hdr Image Quality Assessment Method Based on Tensor Space

ICASSP 2018accepted

The full-reference image quality assessment (IQA) method are limited in practical applications. Here we propose a no-reference quality assessment method for high dynamic range (HDR) images based on tensor space. First, the tensor decomposition is used to generate three feature maps of an HDR image,…

Cited by 0SourceScholar
2018

PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples

ICLR 2018poster

Adversarial perturbations of normal images are usually imperceptible to humans, but they can seriously confuse state-of-the-art machine learning models. What makes them so special in the eyes of image classifiers? In this paper, we show empirically that adversarial examples mainly lie in the low pro…

2018

The INaturalist Species Classification and Detection Dataset

CVPR 2018poster

Existing image classification datasets used in computer vision tend to have a uniform distribution of images across object categories. In contrast, the natural world is heavily imbalanced, as some species are more abundant and easier to photograph than others. To encourage further progress in challe…

2017

Locally-Transferred Fisher Vectors for Texture Classification

ICCV 2017poster

Texture classification has been extensively studied in computer vision. Recent research shows that the combination of Fisher vector (FV) encoding and convolutional neural network (CNN) provides significant improvement in texture classification over the previous feature representation methods. Howeve…

Cited by 75PDFScholar
2017

Speed/Accuracy Trade-Offs for Modern Convolutional Object Detectors

CVPR 2017spotlight

The goal of this paper is to serve as a guide for selecting a detection architecture that achieves the right speed/memory/accuracy balance for a given application and platform. To this end, we investigate various ways to trade accuracy for speed and memory usage in modern convolutional object detect…

Cited by 3693PDFcodeScholar
2016

Improving the Robustness of Deep Neural Networks via Stability Training

CVPR 2016poster

In this paper we address the issue of output instability of deep neural networks: small perturbations in the visual input can significantly distort the feature embeddings and output of a neural network. Such instability affects many deep architectures with state-of-the-art performance on a wide rang…

Cited by 823PDFScholar
2016

Training Deep Neural Networks via Direct Loss Minimization

ICML 2016poster

Supervised training of deep neural nets typically relies on minimizing cross-entropy. However, in many domains, we are interested in performing well on metrics specific to the application. In this paper we propose a direct loss minimization approach to train deep neural networks, which provably mini…

2015

Determining the number of correlated signals between two data sets using PCA-CCA when sample support is extremely small

ICASSP 2015accepted

This paper is concerned with determining the number of correlated signals between two data sets when the number of samples from these data sets is extremely small. In such a scenario, a principal component analysis (PCA) preprocessing step is commonly performed before applying canonical correlation…

Cited by 0SourceScholar
2015

Fusing Subcategory Probabilities for Texture Classification

CVPR 2015poster

Texture, as a fundamental characteristic of objects, has attracted much attention in computer vision research. Performance of texture classification is however still lacking for some challenging cases, largely due to the high intra-class variation and low inter-class distinction. To tackle these iss…

Cited by 22SourcePDFScholar
2015

Learning Semantic Relationships for Better Action Retrieval in Images

CVPR 2015poster

Human actions capture a wide variety of interactions between people and objects. As a result, the set of possible actions is extremely large and it is difficult to obtain sufficient training examples for all actions. However, we could compensate for this sparsity in supervision by leveraging the ric…

Cited by 150SourcePDFScholar