I think everyone here might find this paper interesting. This certainly is more datapoints for not just offline memory, state control etc, but that even with all that you are just extending the collapse out further, because you are still just optimising what you are putting back into the LLM to handle. Now, I know that means for the current/near term using these methods and extending them with offline non LLM based methods. But I also personally think this means ultimately no matter what you do, you are at the mercy of the power of the model regardless.
It’s a very good paper, with some really nice data. Definetly worth a read.