News · 2026-10-02
Cognee: Persistent memory for AI agents
A technical introduction to Cognee as a persistent memory layer for AI agents, combining knowledge graphs, semantic retrieval, sessions, and structured memory beyond a single context window.

What it is
Cognee is an open-source memory platform designed to provide persistent memory for AI agents.
An LLM works with information in its current context. An agent may also need knowledge learned in another session, from a document, or during an earlier task. Repeatedly injecting all of that history into the context window does not scale well.
Cognee moves selected information into an external memory layer. Its documentation describes a system that can turn text into entities, relationships, and searchable chunks; code can be represented through symbols and dependencies. Retrieval can select graph, vector, or code context for the current task.
User
↓
Agent
↓
Working Context ←── Recall ── Cognee Memory
│ ↑
└──────── Remember ────────────┘The model context becomes working memory. Cognee becomes persistent memory. The goal is not simply to enlarge a context window; it is to retrieve the relevant information when it is needed.
The memory lifecycle
Cognee's current high-level memory API is organized around four operations: remember, recall, improve, and forget. These describe a lifecycle rather than only storage and search.
REMEMBER
↓
RECALL
↓
IMPROVE
↓
FORGETRemember
remember stores information that should remain available beyond the immediate interaction. The agent can finish the session without requiring the fact to remain inside its prompt.
await cognee.remember(
"The project uses PostgreSQL for production.",
dataset_name="project_memory",
)Recall
A later task can retrieve relevant context rather than loading an entire memory store. This makes the context window a temporary workspace populated for the current task.
results = await cognee.recall(
"Which database does the project use?",
datasets=["project_memory"],
)Improve
improve enriches memory, applies feedback, and can bridge useful session knowledge into longer-lived graph structures. Memory can therefore have a lifecycle instead of treating every interaction as permanent knowledge.
Forget
Persistent systems also need deletion. forget can remove selected stored information or datasets. This matters because memory can become outdated, incorrect, duplicated, irrelevant, or no longer appropriate to retain.
More than vector search
Traditional retrieval-augmented generation often indexes document chunks as embeddings and retrieves by similarity. Cognee also represents entities and relationships, so those connections can be part of the stored knowledge.
Project A
│ uses
▼
Cognee
│ provides memory
▼
Agent B
│ belongs to
▼
Workspace CSemantic similarity and graph relationships solve different parts of retrieval. A graph can express relationships directly; it does not guarantee that extraction or retrieval is correct.
Session memory and long-term memory
Cognee documents session-scoped memory alongside durable graph-backed memory. Useful session knowledge can later be promoted into longer-lived structures, which means not everything the agent sees needs to become permanent memory.
Agent session
│
▼
Session memory
│ useful?
▼
Improve / sync
│
▼
Long-term memory
│
▼
Future sessionsRunning Cognee locally
Current Cognee documentation describes local text-memory workflows using local extraction and embedding models without a cloud LLM API key. The GLiNER extra provides the local extraction model used in that path; LLM-dependent stages still require an appropriate model configuration.
uv pip install "cognee[gliner]"import asyncio
import cognee
async def main():
await cognee.remember(
"Cognee provides persistent memory for AI agents.",
dataset_name="agent_memory",
)
results = await cognee.recall(
"What provides persistent memory?",
datasets=["agent_memory"],
)
print(results)
asyncio.run(main())Connecting memory to an agent
Memory becomes useful when it sits inside the agent loop: retrieve relevant memory before reasoning and persist selected information after a task. Cognee exposes SDK, REST, MCP, and integration surfaces around this memory workflow.
┌───────────────┐
│ Cognee │
│ Memory │
└───────┬───────┘
│
recall │ remember
│
▼
User ───────────────► Agent
│
▼
LLM
│
▼
ToolsThis creates a loop closer to Recall → Reason → Act → Verify → Remember than Prompt → Answer → Forget everything.
Core technical characteristics
- Persistent memory — information can survive beyond an individual model interaction or agent session.
- Knowledge graphs — entities and relationships can be explicit parts of memory rather than existing only inside text chunks.
- Semantic retrieval — embeddings provide similarity-based retrieval alongside graph structures.
- Session-aware memory — temporary session knowledge can remain separate from permanent memory and later be promoted when appropriate.
- Datasets — memory can be scoped so applications do not search every stored item indiscriminately.
- Local execution — current Cognee documentation describes local text-memory workflows with local extraction and embedding models.
- Multiple integration surfaces — Cognee exposes SDK, REST, MCP, and integration surfaces.
- Explicit lifecycle — remember, recall, improve, and forget make memory management an application concern.
Why it matters
The difficult questions in AI memory are not only about storage. They are what an agent should remember, when it should remember it, how it should be represented, what should be retrieved for a task, and when old information should be updated or forgotten.
An agent that remembers everything can become worse rather than better: more stored information creates more opportunities for stale, irrelevant, or incorrect context to enter future reasoning. The aim is useful memory, supported by a deliberate write policy and evaluation of retrieval quality, latency, correctness, and downstream task outcomes.
From tools to memory
The Agentic Engineering series now connects two practical components: browser execution and persistent memory.
AI AGENT
│
┌────────┴────────┐
│ │
▼ ▼
MEMORY TOOLS
│ │
Cognee Playwright CLI
│ │
▼ ▼
Persistent Browser
Knowledge ExecutionAn agent that can act but cannot preserve useful knowledge repeatedly rediscovers its world. An agent that remembers without a clear write policy creates a different problem. Persistent memory is infrastructure, not a guarantee of better reasoning.
Limitations and open questions
Memory does not automatically improve reasoning. Incorrect or irrelevant memories can degrade future outputs just as easily as useful memories can help them.
Retrieval remains imperfect. Relevant memories can be missed and irrelevant memories can be retrieved; graph structure and semantic similarity do not eliminate that risk.
Graph construction can introduce errors. An entity or relationship can look structured and authoritative while originating from an incorrect extraction.
Write policy matters. Persisting every conversation, tool result, and intermediate thought is rarely desirable in a production agent.
Memory requires identity boundaries. Multi-user and multi-workspace systems must ensure that retrieved memory belongs to the correct user, workspace, or application context.
Forgetting must actually work. Long-lived systems need reliable deletion and lifecycle policies, particularly for user or organizational information.
Evaluation is difficult. The meaningful outcome is whether memory improves downstream performance, not how many memories were stored.
