You must log in or # to comment.
Interesting, and impressive to just whip up a new learning architecture single handedly like this ! There was a paper on how to inject memory into the last layer(s) of a normal transformer architecture. Seemed a bit ‘hacky’ but not sure which is better. No matter, In a few months we’ll have several architectures and techniques to choose from, that all replace/augment primitive RAG systems, so the missing parts of the artificial brain is being developed everywhere.
This really does seem like the next frontier. If a model can learn and adopt on the fly, it becomes far more useful. And I suspect that small models could become tuned to do specific tasks a lot better than a generic large model.



