Reworked the sign-up screens and wrote the copy for the new welcome step.
Jelly remembers what you saw, said and promised, across months of conversations. We measure how well it recalls that on a public benchmark, publish every answer it gave, and tell you exactly how the test was run.
Chats, email, meetings and the pages you read become dated moments Jelly can find again — by the words someone used or by what they meant — and every answer shows the moment it came from.
ok, $1,300 for the 32 GB · Fri 11 Sepbudget $1,200 · Wed 2 SepYou remember a Wednesday afternoon, and separately you know who Dana is. Jelly keeps the same two apart: what happened, and what it has come to know — and each one helps it read the other.
Every stretch of your day is kept as an episode: what you were working on, who you spoke with, for how long, and what you asked Jelly while you were there.
Reworked the sign-up screens and wrote the copy for the new welcome step.
Agreed the new sign-up flow ships first. Dana asked for the onboarding update by Thursday.
Dana confirmed Thursday and asked for the sign-up flow first. No reply yet.
From those episodes Jelly builds up what it knows: the people in your work, the projects, the things you are deciding — and where each one stands right now.
The new sign-up and welcome flow for the web app. You lead it; Dana reviews.
Currently: waiting on your update — Dana wants it by Thursday
Reviews your onboarding work. Talks with you most on WhatsApp and in the weekly Meet.
Dana: I'll sign off on the onboarding copy before it goes to your team — Slack, 15 Sept, 10:04
Dana: as your manager I'd rather we ship the sign-up flow first — Google Meet, 23 Sept, 11:41
Dana: can we get the onboarding update out by Thursday? — WhatsApp, 23 Sept, 16:12
The 32 GB model you are buying.
Currently: budget $1,300 (raised from $1,200 on 11 Sep)
Ask "what budget did I settle on?" Semantic memory knows the answer is $1,300 and that it changed. Episodic memory takes you to the Friday conversation where it did.
BEAM (ICLR 2026) builds long conversations — twenty of about 100,000 tokens each in the split we ran — and asks 400 questions across ten abilities, from recalling a stated fact to noticing that you contradicted yourself. We ran all 400 once.
Both on the same protocol and the same answering model. past.dev's figures are from its benchmarks page. Event ordering is scored by how well the order matches; every other ability by rubric.
Each figure is that team's own published run, with its own models. Only past.dev's used the same protocol and answering model as ours, so read the others as indicative.
Checked against each source on 29 September 2026. Hindsight's earlier post gave 73.4%; its benchmarks page now gives 75.0%, which we show. Mem0 publishes BEAM results at 1M and 10M tokens only.
A benchmark number means little without the conditions behind it. These are all of them.
b2da22e, licensed CC BY-SA 4.0.gpt-5.6-luna with reasoning off for both, the same as past.dev.The results repository has what you need to re-score our run without trusting us.