Agent Memory

that works with

your existing Search Stack

Haystack EU · Berlin, Main Stage · Wed 16 September 2026, 16:00
Paul-Louis Nech · Algolia
Intro/skeleton 3

At least I remembered the interleaving tip.

Intro/skeleton 4

Remember everything?

memory
User prefers Vue.js over React for new repos
technical2026-08-12
Memento — don't believe his lies
Intro/skeleton 6
panel cue 1

This year I briefed my agents.
Every. Single. Conversation.

Remembering not enough is bad.

Intro/skeleton 7
a ChatGPT system prompt
[system]
…
model set context:
  2025-04-11: user prefers concise answers…
  2025-04-12: user is building a Vue app…
  2025-04-14: user lives in Paris…
  … 40+ lines, every single turn …

The opposite is true too:
bringing everything up is bad as well.

Charlie Day at his string-covered conspiracy board, overwhelmed by too many pinned notes
Intro/skeleton 5

Remembering nothing is bad.

Remembering everything
is bad too.

Intro/skeleton 8+9

False positives.

False negatives.

It's an F1 score. ✨

Memory is a Retrieval Problem

Haystack EU · Berlin, Main Stage · Wed 16 September 2026, 16:00
Paul-Louis Nech · Algolia
Our turf/scaffold

BTW hi 👋

Paul-Louis Nech · Staff ML Engineer, AlgoliaAlgolia
We do Search.
(and Recommendations)
(and more)
Agent Studio is the agentic harness on top.
Paul-Louis Nech livecoding
Act 1 of 4

Our turf

Our turf9 minComponents12 minSide by side5 minConclusion7 min
Our turf/skeleton 11

Agents are new.
Retrieval under constraints
is our turf.

Our turf/skeleton 12

Budget 1: latency

A test-card image painting in from the top in bands, like a slow dial-up load
Our turf/skeleton 13

Budget 2: tokens 💸

Token bill · estimated
memories in the prompt≈ tokens× a 20-turn chat455011k101.4k28k304.2k83k
You pay for every over-fetched memory, on every single turn.
Our turf/skeleton 14+14b
panel cue 2

Budget 3: attention

Our needles · 2025 vs 2026
20252026needle lost at20knever128k costs youa wrong answer75–134 s

In 2025 it was impossible.
Now it's just stupid expensive.

Our turf/skeleton 15a

Where benchmarks give you a baseline

thus a first impression

  • QR code linking to the LoCoMo paper on arXivLoCoMoarXiv 2402.17753
  • QR code linking to the LongMemEval paper on arXivLongMemEvalarXiv 2410.10813
Our turf/LoCoMo, added 09-16

One answer, four turns, three sessions

one question, verbatim
Q    What activities does Melanie partake in?

A pottery, camping, painting, swimming

ev D1:12 D1:18 D5:4 D9:1

No string match scores that answer.

QR code linking to the LoCoMo paper on arXivLoCoMoarXiv 2402.17753

LoCoMo, as released
conversations10sessions each19–32questions1,986
Our turf/added 09-16

Benchmarks tell a partial picture

What the score hides
  • ten conversations, 25k tokens at most
  • one question in five has no answer
  • full context beat the memory systems
A blog post titled Lies, Damn Lies, and Statistics: Is Mem0 Really SOTA in Agent Memory?

first two measured in locomo10.json on 09-16 · third reported by Zep, May 2025

Our turf/LongMemEval, added 09-16

What LongMemEval asks

What it asks · n=500
multi-session133temporal-reasoning133knowledge-update78single-session-user70single-session-assistant56single-session-preference30
one question, verbatim
Q   What degree did I graduate with?

A Business Administration

One of 54 sessions holds it.
and one it must refuse
Q   What is the name of my hamster?

A You did not mention this information.
You mentioned your cat Luna but not
your hamster.

Our turf/added 09-16

“Lies, Damned Lies, and Benchmarks”

Benchmark Lies
  • Models chase benchmarks.
  • Benchmarks chase users.
  • People cherry-pick for their favorite narrative.
  • Benchmarks have been broken for a while.
Their own accuracy against cost chart, with a frontier line

“De-emphasizing benchmarks
even when you're ahead.”

typesafe.ai/blog/antibenchmaxxing · their words and their chart, on their own model release

Our turf/skeleton 15b

…and where your shopper doesn't care.

and has different success metrics

Our turf/skeleton 16
panel cue 3

You already have query → results.

Act 2 of 4

Components

Our turf9 minComponents12 minSide by side5 minConclusion7 min
Components/skeleton 17

① The query.

user_input
query
four ways to build it
  • verbatim?
  • with topics?
  • saved triggers?
  • LLM rewrite?
Components/skeleton 18
panel cue 4

Baseline: just the last N.

Components/skeleton 18b
panel cue 8

One conversation. Five queries.

  • tiredthe last N memories— no query at all
  • wiredthe user, verbatim"what haven't i tasted yet?"
  • +widen with topics… · topics: food, preferences
  • +phrases saved at write timerecallTriggers: "coffee not tried yet"
  • hiredLLM rewrite, one round-trip"specialty coffee Paris untried"
Same ask. The query is a design decision, not an inheritance.
Components/skeleton 18c

The query you wrote
at write time.

memory    "… been to every specialty coffee bar on their Paris list
           except Substance in the 3rd …"

triggers  "what coffee haven't i tried yet"
           "specialty coffee, 3rd arrondissement"
           "places still on my list"
Triggers: simplified example.
Write the query when you write the memory.
Components/skeleton 19

…especially with topics.

preferencesshoppingfoodwork healthtravelentertainmentfamily hobbiesgoalshistoryfeedback complaintspraisetechnical learningschedulefinance
closed: our real 18, one set, filterable
  • specialty coffee
  • monday closures
  • palak paneer
  • co-ferments
  • 3rd arrondissement
  • vegetarian mains
  • natural wine
  • queue tolerance
  • goûter
  • filter over espresso
open: the corpus writes its own
Components/skeleton 20

② Retrieval.

query
hits
levers you already tune
  • ranking formula
  • facets & filters
  • optional scoring
  • rules · synonyms
  • typo tolerance
  • Docs
Components/skeleton 21

Same words: fine.
Same meaning: not so much.

Recall@1 · n=625/644/644
the memory's own words81%rephrased34%intent, no shared words9.5%same index. same records.
Components/skeleton 21b

Three settings that moved recall.

Twelve swept · n=1913
settingΔ recall@10removeStopWords=false−35.4 ppsearchableAttributes=[text]−5.4 ppoptionalWords dropped+2.9 ppthe other ninenoise
Components/skeleton 22

③ Formatting.

hits
prompt
what you hand the model
  • attributesToRetrieve
  • snippeting
  • record format?
Components/skeleton 23

JSON? TOON? YAML? Markdown?

json
{"hits": [{"objectID": "d0cd1026e8ae51ef342eddaa108faeed2a1f6e2d0386dd1756a75e7fcfba34e2", "text": "User completed their undergraduate degree in Computer Science from UCLA.", "rawExtract": "I completed my undergrad in CS from UCLA, which has a great reputation in the industry.", "memoryType": "semantic", "createdAt": "1788961514", "keywords": "UCLA, Computer Science, undergraduate, bachelor degree…
toon
hits[4]{objectID,text,rawExtract,memoryType,createdAt,keywords,topics,recallTriggers}:
  d0cd1026e8ae51ef342eddaa108faeed2a1f6e2d0386dd1756a75e7fcfba34e2,User completed their undergraduate degree in Computer Science from UCLA.,"I co…
same 4 memories, ~11% fewer tokens
Components/skeleton 23b

Measure it.

Answer accuracy % · n=60 per arm
json53.3toon58.3 (+5.0)yaml56.7markdown_kv56.7not significant
Components/skeleton 24

Query · Retrieval · Formatting.

panel cue 6
Act 3 of 4

Side by side

Our turf9 minComponents12 minSide by side5 minConclusion7 min
Side by side/skeleton 25
panel cue 6

Same question. Three agents.

Side by side/tradeoff table, added 09-14

Four seasons of Memory

our own stage rig · one question · September 2026 · tokens exact, clocks indicative
secondstokens infacts recalledno memory2.9650the last 5 turns5.43451memory as a tool18.19,1841preload + preflight7.53453

Tools-only: 27× tokens, yet less context.

Act 4 of 4

Conclusion

Our turf9 minComponents12 minSide by side5 minConclusion7 min
Conclusion/skeleton 27

Some players go pure vector.
Sub-200 ms. Impressive!

Conclusion/skeleton 28

First: make it work on what you already have.

Conclusion/skeleton 30

Forgetting is retrieval too.

I asked the agent to compact its own memory.
91% of deletions destroyed unique anchors.
Shakira, from the video for her track Remember to Forget
Conclusion/skeleton 31
CTO@5

Your Turn: Measure It

Read 5 conversations.
Read 20 memories.
Search twice yourself.
Retrieved, irrelevant?
Relevant, not retrieved?

Memory is Retrieval,
and retrieval is our domain.

Retrieved memory records beside a judgment card scoring 0.77 passed, with expected-versus-got chips
Conclusion/skeleton 31b

Memory is just search.

  • A list of memory retrieval experiments, each with its own score
  • A dataset of cases, each with an expected tool, query and filters
  • Retrieved records beside a judgment card scoring 0.77 passed
  • One case end to end: the question, the execution, the judging receipts
Conclusion/skeleton 32

Thanks.

QR code linking to these slides at alg.li/haystack26
Grab these Slides
QR code linking to Paul-Louis Nech on LinkedIn
Let’s Connect