Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks
Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be key for pushing AI beyond merely scaling current designs. **Method.** We present a method to optimize a transformer ar…