Comparing memory mechanisms across RNNs, Transformers and SSMs
A community post compares architectures through working memory: RNNs compress history into a recurrent state with roughly O(N) state versus O(N²) parameters, Transformers store past representations in a growing KV cache while weights stay frozen, and selective SSMs like Mamba return to fixed-size recurrent memory with input-dependent retention.