← Search

NAN DING

19 accepted papers

2026

AquaSplatting: A Hybrid 3D Representation for Robust Underwater Scene Reconstruction via Dual-Branch Rendering

AAAI 2026technical

While 3D Gaussian Splatting (3DGS) excels at real-time rendering of standard scenes, it struggles to reconstruct underwater environments due to severe challenges such as light scattering, color attenuation, and sparse coverage of Gaussian kernels in far-field aqueous regions. To address this, we int

Cited by 0SourcePDFScholar
2026

FILD-Nav:Vision-and-Language Navigation with Instruction Landmark Features in Continuous Environments

IJCAI 2026

Vision-and-language navigation (VLN) requires agents to follow natural language instructions to navigate autonomously in continuous environments. However, existing approaches often lack high-level semantic guidance in waypoint prediction and explicit language–landmark alignment in cross-modal planni

Cited by 0Scholar
2025

Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluations

EMNLP 2025

Auto-evaluating language models (LMs), *i.e*., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and the cost associated with it. But this presents a paradox: how can we trust the grader LM, which is presumably weaker than the candidate LM, to a

Cited by 0SourcePDFScholar
2024

A Complete Method for the 3D Reconstruction of Axonal Pathways from 2 Orthogonal 3D OCT Images of the Lamina Cribrosa

ICASSP 2024accepted

The lamina cribrosa is a 3D mesh-like structure within the optical nerve head, consisting of pores through which all the nerve fibers from the retina pass to join the brain. Alterations and damages in this structure have proven to be linked with glaucoma, the second leading global cause of blindness…

Cited by 0SourceScholar
2024

CausalLM is not optimal for in-context learning

ICLR 2024poster

Recent empirical evidence indicates that transformer based in-context learning performs better when using a prefix language model (prefixLM), in which in-context samples can all attend to each other, compared to causal language models (causalLM), which use auto-regressive attention that prohibits in…

2024

One Model to Drift Them All: Physics-Informed Conditional Diffusion Model for Driving at the Limits

CoRL 2024poster

Enabling autonomous vehicles to reliably operate at the limits of handling— where tire forces are saturated — would improve their safety, particularly in scenarios like emergency obstacle avoidance or adverse weather conditions. However, unlocking this capability is challenging due to the task's dyn…

Cited by 8SourceScholar
2023

Improving Robust Generalization by Direct PAC-Bayesian Bound Minimization

CVPR 2023highlight

Recent research in robust optimization has shown an overfitting-like phenomenon in which models trained against adversarial attacks exhibit higher robustness on the training set compared to the test set. Although previous work provided theoretical explanations for this phenomenon using a robust PAC-…

Cited by 8SourcePDFScholar
2023

PaLI: A Jointly-Scaled Multilingual Language-Image Model

ICLR 2023top-5%

Effective scaling and a flexible task interface enable large language models to excel at many tasks. We present PaLI, a model that extends this approach to the joint modeling of language and vision. PaLI generates text based on visual and textual inputs, and with this interface performs many vision,…

2022

All You May Need for VQA are Image Captions

NAACL 2022long

Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation. In this paper, we propose a method that automatically derives VQA examples at volume, by leveraging the abundance of existing image-captio…

2022

Explicit Role Interaction Network for Event Argument Extraction

EMNLP 2022finding

Event argument extraction is a challenging subtask of event extraction, aiming to identify and assign roles to arguments under a certain event. Existing methods extract arguments of each role independently, ignoring the relationship between different roles. Such an approach hinders the model from le…

2022

PACTran: PAC-Bayesian Metrics for Estimating the Transferability of Pretrained Models to Classification Tasks

ECCV 2022poster

"With the increasing abundance of pretrained models in recent years, the problem of selecting the best pretrained checkpoint for a particular downstream classification task has been gaining increased attention. Although several methods have recently been proposed to tackle the selection problem (e.g…

2021

Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning

NeurIPS 2021poster

Despite recent advances in its theoretical understanding, there still remains a significant gap in the ability of existing PAC-Bayesian theories on meta-learning to explain performance improvements in the few-shot learning setting, where the number of training examples in the target tasks is severel…

Cited by 35SourcePDFScholar
2021

Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

CVPR 2021poster

The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected with overrestrictive requirements inherited from their original target tasks (e.g…

Cited by 1186PDFcodeScholar
2021

Do Transformer Modifications Transfer Across Implementations and Applications?

EMNLP 2021main

The research community has proposed copious modifications to the Transformer architecture since it was introduced over three years ago, relatively few of which have seen widespread adoption. In this paper, we comprehensively evaluate many of these modifications in a shared experimental setting that…

2016

Stochastic Gradient MCMC with Stale Gradients

NeurIPS 2016poster

Stochastic gradient MCMC (SG-MCMC) has played an important role in large-scale Bayesian learning, with well-developed theoretical convergence properties. In such applications of SG-MCMC, it is becoming increasingly popular to employ distributed systems, where stochastic gradients are computed based…

Cited by 34SourcePDFScholar
2015

Embedding Inference for Structured Multilabel Prediction

NeurIPS 2015poster

A key bottleneck in structured output prediction is the need for inference during training and testing, usually requiring some form of dynamic programming. Rather than using approximate inference or tailoring a specialized inference method for a particular structure---standard responses to the scal…

Cited by 23SourcePDFScholar
2015

On the Convergence of Stochastic Gradient MCMC Algorithms with High-Order Integrators

NeurIPS 2015poster

Recent advances in Bayesian learning with large-scale data have witnessed emergence of stochastic gradient MCMC algorithms (SG-MCMC), such as stochastic gradient Langevin dynamics (SGLD), stochastic gradient Hamiltonian MCMC (SGHMC), and the stochastic gradient thermostat. While finite-time converge…

Cited by 214SourcePDFScholar