A Research Group Reports a 243M-Parameter Transformer Trained Without Backpropagation — type0 | type0