Dust: Pretraining Transformers Without Backpropagation
Posted by E-Reverance |an hour ago |2 commentsapi 22 minutes ago[1 more]
It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?
derin-picment 22 minutes ago