Cortex (local models)
Note: these are small local models — fast and free, but not 100% accurate. Expect occasional misfires: a memory saved when it shouldn’t have been, a paraphrase missed by the dedup head, a contradiction flagged on agreeing statements. You can correct any of them in chat (“forget that”, “that wasn’t a preference”) and the heads keep improving as the models are retrained.
Cortex is the bundled set of local models OpenEnsemble runs in-process. They handle small, frequent reasoning tasks so the install works offline and doesn’t burn cloud tokens for every internal decision.
What it includes
openensemble-reason-v3— a SmolLM2-based GGUF reasoning model used for tiny classifications, agent memory updates, content gating, and similar internal calls.nomic-embed-text-v1— the embedding model used for memory recall, search, and similarity.- Plan model —
openensemble-plan-360m-v2, a SmolLM2-360M GGUF that parses scheduling intent (“every Monday at 9am”). A related 360M extract variant powers the local cognition slot-filling tier.
All three run via node-llama-cpp on CPU. No GPU required. They’re loaded the first time they’re needed and stay resident.
See and correct remembered context
New chat answers show Memories used as context when stored memories were supplied to the model. Expand it to see the remembered text and why it was included, such as a pinned rule, relevant fact, or past conversation. This identifies the context the model received; it does not prove which memory determined the answer.
Source conversation opens the original conversation excerpt when OE recorded a source link and that conversation is still available. Older memories may have no recorded source. The same source controls are available in Memory Control and Run Inspector.
Use This is outdated to correct a memory for future answers. OE checks that another screen has not changed it first. Previous answers keep the context they originally received, and a correction preserves the original source conversation. Memory details are private to the owning profile.
Search remembered information
Ask “Show me my food preferences” or “What do you know about me?” to search saved facts and notes. Refine the topic or request another page if needed; one short result is not a complete inventory. Search uses local Nomic and, when compatible and enabled, Cortex. The chat model receives bounded matching excerpts to explain.
Memory search and supplied memory context share a 1,000-byte output allowance per turn, with at most two searches. Normal conversation and other tools have separate limits. Use the source controls and This is outdated to correct an entry you no longer want used.
What you don’t need to do
- You don’t need to download anything separately — the GGUFs ship with OpenEnsemble.
- You don’t need to set up Ollama or LM Studio for these specific tasks. (You may still want them for user-facing chat models — see LLM providers.)
Performance check
oe bench
Run that on the install to see tokens/sec and memory footprint for the reason and embed models. If it’s slow on your hardware, see “Swapping providers” below.
Swapping providers
In config.json:
"cortex": {
"reasonProvider": "auto", // built-in | ollama | lmstudio
"lmstudioUrl": "http://127.0.0.1:1234"
}
auto— use the bundled GGUF.ollama— call your local Ollama instead. Faster on a GPU; needs Ollama running.lmstudio— call LM Studio. Make sure JIT model loading is enabled in LM Studio, otherwise non-loaded models 404.
When Cortex is slow
The bundled reason model is small but everything is CPU-bound. If you’re on a constrained box, the most-felt slowness is:
- Memory writes after long chats (every chat persists summaries via Cortex)
- Schedule parsing when creating tasks
- Agent classification in the Coordinator
All of those are cacheable. The first one of each is slow; subsequent calls in the same session are fast.
Privacy
Bundled Cortex inference runs on the OE host. If you select another runtime, requests go to its configured endpoint, which may be another machine. Retrieved memories supplied to a cloud chat model leave the server as part of that chat context. See Which model does what.