What to watch
- Persistent, user-controlled memory in assistants
- Long context versus retrieval trade-offs as context windows grow
- Evaluation of retrieval quality in production
The state of play
Models know only their training data and their context window. Retrieval-augmented generation fills the gap by fetching relevant material at query time. Long context windows reduced the need for elaborate retrieval in small cases, but retrieval remains essential for large, changing or permissioned data. Agents add a new requirement: memory that persists across sessions.
What works today
- Question answering over a company’s documents with citations.
- Hybrid search (keywords plus embeddings) with reranking.
- Exposing data sources to assistants through standard connectors.
What doesn’t yet
- Memory that is selective, correctable and transparent to the user.
- Reliable retrieval over messy, contradictory or out-of-date sources.
Open questions
Should memory live in the model’s weights, in external stores, or in both?
Evidence