A · Filesystem-native agent
Strong baseline with files, search, shell/tool access and freedom to organize or rewrite persistent information.
research://memorybench
An architecture-neutral benchmark for persistent agent memory: measure continuity, temporal correctness, adaptation, latency and cost without rewarding one storage design by construction.
research://thesis
MemoryBench avoids a weak “chat history vs memory product” comparison. The benchmark gives a best-of-breed filesystem-native agent real tools to search, organize, rewrite and consolidate its own persistent workspace, then compares that baseline with SynapseFS and SynapseFS + SymbioRAP Memory under matched model, information and compute conditions.
focus://memorybench
Strong baseline with files, search, shell/tool access and freedom to organize or rewrite persistent information.
Same agent conditions with deterministic persistence, stable identities, evidence and version semantics.
Adds episodic/semantic memory, temporal graph, consolidation and provenance-aware context selection.
Model, information access, task set and compute budget are controlled so architecture effects remain interpretable.
architecture://working-model
The benchmark protocol is designed to expose trade-offs rather than produce a single flattering score. It measures whether an architecture remains correct over time, adapts to changed facts, retrieves the right evidence and does so within practical latency and token/compute budgets.
Correct task answers and evidence use across short and long horizons.
Correct handling of superseded facts, changed preferences and time-dependent state.
Ability to revise behavior when new evidence conflicts with earlier assumptions.
Retrieval/consolidation overhead, token use and compute required to maintain memory quality.
research://questions
Which tasks genuinely require a dedicated memory layer instead of disciplined filesystem organization?
How much improvement comes from persistence semantics alone before semantic memory is added?
Which consolidation policies improve long-horizon accuracy without creating stale or overconfident memory?
How should benchmark tasks resist model memorization and architecture-specific shortcuts?
status://research
The points shown describe research and implementation directions, not guarantees of finished product functionality.
connect://research
For research, funding and technology partners, we share architecture questions, benchmark design and technical work states in qualified collaboration.
Next useful step
You do not need to define the technical solution. Describe what you want to build or improve and we will map feasibility, a sensible first scope and next steps.
Get a free idea assessment
Four fields are enough. You will receive a concrete response on how to start small and extend the project later.