Memory

A memory you can check.

Jelly remembers what you saw, said and promised, across months of conversations. We measure how well it recalls that on a public benchmark, publish every answer it gave, and tell you exactly how the test was run.

What Jelly remembers

One memory of every conversation you are already in.

Chats, email, meetings and the pages you read become dated moments Jelly can find again — by the words someone used or by what they meant — and every answer shows the moment it came from.

Every moment keeps its day.So "when did she ask?", "what changed?" and "what did I say first?" have answers, not guesses.
Words and meaning.It finds the exact phrase someone used, and the message that meant the same thing in other words.
The newest value, and the one it replaced.When a budget, a date or a decision changes, Jelly answers with the current one and knows it moved.
An answer shows its source.Each one names the message and the day it came from. When nothing in memory matches, it says so rather than filling the gap.
It stays on your Mac.The memory is kept on your own machine, encrypted. See Disclosure for what is sent to a model and when.
What budget did I settle on for the new laptop?↵ ask
$1,300. You raised it from $1,200 on Friday 11 September, after Priya said the 32 GB model was worth it.
Based on
WhatsApp Priya Nair
ok, $1,300 for the 32 GB · Fri 11 Sep
Notes Laptop shortlist
budget $1,200 · Wed 2 Sep
How Jelly remembers

Two kinds of memory, the way yours works.

You remember a Wednesday afternoon, and separately you know who Dana is. Jelly keeps the same two apart: what happened, and what it has come to know — and each one helps it read the other.

Episodic memory

What happened, and when.

Every stretch of your day is kept as an episode: what you were working on, who you spoke with, for how long, and what you asked Jelly while you were there.

  • Go back to any day and open what was said.
  • Ask "what was I doing when Dana messaged?" and land on the moment.
  • Forget any episode, and it is gone.

Wednesday 23 September

4 stretches
09:1042 min
Onboarding flow redesign

Reworked the sign-up screens and wrote the copy for the new welcome step.

WorkOnboarding revampFigma, NotionWhat was said
11:3025 min
Weekly sync with Dana

Agreed the new sign-up flow ships first. Dana asked for the onboarding update by Thursday.

MeetingOnboarding revampGoogle MeetWhat was said
16:128 min
Dana on WhatsApp

Dana confirmed Thursday and asked for the sign-up flow first. No reply yet.

MessagesWhatsAppHide what was said
YouWhen is the onboarding update due?
JellyThursday. Dana asked on WhatsApp at 4:12 PM, with the new sign-up flow first.
Semantic memory

What is true, and how sure it is.

From those episodes Jelly builds up what it knows: the people in your work, the projects, the things you are deciding — and where each one stands right now.

  • Every belief shows how firmly it is held, and the conversations it came from.
  • When something changes, the current state moves with it.
  • Confirm what is right, correct what is not. Your word wins.
Onboarding revamp project

The new sign-up and welcome flow for the web app. You lead it; Dana reviews.

Currently: waiting on your update — Dana wants it by Thursday

revised 23 Sept, 16:20
Dana Whitfield

Reviews your onboarding work. Talks with you most on WhatsApp and in the weekly Meet.

manager ●●○✓✕seen in 3 conversations ▾

Dana: I'll sign off on the onboarding copy before it goes to your team — Slack, 15 Sept, 10:04

Dana: as your manager I'd rather we ship the sign-up flow first — Google Meet, 23 Sept, 11:41

Dana: can we get the onboarding update out by Thursday? — WhatsApp, 23 Sept, 16:12

Works at
Northwind ●●●✕you told Jelly this
New laptop thing

The 32 GB model you are buying.

Currently: budget $1,300 (raised from $1,200 on 11 Sep)

revised 23 Sept, 18:55
Together

Ask "what budget did I settle on?" Semantic memory knows the answer is $1,300 and that it changed. Episodic memory takes you to the Friday conversation where it did.

The benchmark

BEAM: ten ways a memory gets tested.

BEAM (ICLR 2026) builds long conversations — twenty of about 100,000 tokens each in the split we ran — and asks 400 questions across ten abilities, from recalling a stated fact to noticing that you contradicted yourself. We ran all 400 once.

AbilityJelly past.devJelly · past.dev
Preference following98.8100.0
Abstention98.897.5
Event ordering100.0100.0
Summarization96.8100.0
Temporal reasoning96.398.1
Instruction following90.095.6
Information extraction84.790.0
Knowledge updates75.687.5
Contradiction resolution73.878.8
Multi-session reasoning61.170.9

Both on the same protocol and the same answering model. past.dev's figures are from its benchmarks page. Event ordering is scored by how well the order matches; every other ability by rubric.

Where it sits

Every published BEAM-100K result.

Each figure is that team's own published run, with its own models. Only past.dev's used the same protocol and answering model as ours, so read the others as indicative.

Checked against each source on 29 September 2026. Hindsight's earlier post gave 73.4%; its benchmarks page now gives 75.0%, which we show. Mem0 publishes BEAM results at 1M and 10M tokens only.

Method and disclosure

Exactly how the 87.6% was produced.

A benchmark number means little without the conditions behind it. These are all of them.

What was run

  • Dataset: BEAM's 100K split — 20 conversations, 400 questions, two per ability per conversation — from the authors' repository at commit b2da22e, licensed CC BY-SA 4.0.
  • Memory: each conversation stored in Jelly as the app would store it, then each question answered from what Jelly retrieved: about 8,000 tokens of evidence per question.
  • Answering and judging: gpt-5.6-luna with reasoning off for both, the same as past.dev.
  • Scoring: each question scores 0 to 1; the headline is the mean over all 400.
  • Runs: one. A repeat would move the headline by about two points either way.

What you should know

  • The protocol gives the answering model hints. We used the answer and judge prompts ExaBase published for BEAM, which past.dev also uses, so our number compares with theirs. Those prompts pass parts of the dataset's own annotations to the answering model — the points a summary should cover, a hint for date arithmetic, the topics to put in order. That is why event ordering and temporal reasoning score near 100% on this board.
  • On BEAM's own, stricter prompts, Jelly scored 46.1% (a run before the two retrieval improvements below). The BEAM paper's own method scored 35.8% on those prompts with a different model.
  • Two retrieval improvements were checked against this split. We changed how Jelly ranks keyword matches and how much weight it gives to recency after measuring evidence recall on these 400 questions, and confirmed on a separate benchmark (LongMemEval) that everyday recall did not get worse. They are general changes, not tuned to particular questions, but they were not validated on a held-out split.
  • Other teams' numbers are theirs. Different models, judges and set-ups; see the sources above.
Check it yourself

Every answer, public.

The results repository has what you need to re-score our run without trusting us.

In the repository
  • Every one of the 400 answers Jelly gave, keyed to BEAM's question ids.
  • Every judge verdict behind every score.
  • A script that recomputes 87.57% from the verdicts, with no API calls.
  • A script that re-judges every answer with the published judge prompts and your own API key.
Not in it
  • Jelly's retrieval code. The answers show what it recalled; how it recalls stays ours.
  • The dataset and the third-party prompts. The scripts download them from their authors and check them.