Same responses. Two policies.
Matches / 20Loading repeatability results…
Should your AI keep that thought?
We use Jev to test whether a memory is supported,
useful over time, and faithful to what was said.
Open source · Recorded examples need no key
Mean Jev response
—msClient-observed request timePreview agreement
—%Against our expected labelsIntended memories saved
—Preview policy · unique casesEnglish evaluations
—One evaluation per distinct case100 agent-authored synthetic cases; no independent human annotation. These are admission recommendations, not verified facts or stored memories.
Loading repeatability results…
Includes network time and the first request. No outliers removed. Not a throughput test or speed guarantee.
Does the original passage actually support the proposed memory?
Will this matter beyond the current conversation?
Are proposals, conditions and uncertainty preserved?
Start with a recorded case. Then test a memory of your own.
Explore examplesReal recorded results. No key, no new API calls.
Evaluate your textRun locally with your TypeSafe key. One API call per evaluation.
Choose a recorded example above,
or evaluate your own source and candidate.
Same Jev response. Different confidence thresholds.
Baseline remains the default; preview is experimental.
Confidence is the provider's estimate, not a guarantee of correctness.
Try your own synthetic examples. Share the expected decision
and the actual result so others can reproduce it.
This lab explores what deserves to become a memory.
Cairn Memory is the memory project; this demo does not write to it.