Textual Memory
Long-context, persona, script, and conversation-memory benchmarks.
The open benchmark for agent memory
Measure what your agents remember.
Compare what truly matters.
A public benchmark space for comparing textual, multimodal, and coding-agent memory systems under a consistent evaluation flow.
Each track keeps its own result table and detailed metric breakdown.
Long-context, persona, script, and conversation-memory benchmarks.
Memory retrieval and generation over image-rich or multimodal tasks.
Agent memory support for coding tasks and repository-context recall.
Industry systems use the hosted Add/Search key flow. Academic systems may use the same flow or submit a public GitHub repository for maintainer Docker deployment.
Provide hosted Add/Search APIs, or submit a public GitHub repository with Docker and API run instructions.
Use the issued key to verify the synchronous Add/Search flow.
After smoke passes, select an evaluation mode and submit a scored run.
Use the product pages to inspect rankings, run evaluations, and prepare an integration.
Public ranking with filters, dataset columns, and score bars.
Create eval jobs, watch progress, and inspect private results.
Dataset descriptions, source links, and benchmark coverage.
User guide, evaluation workflow, API contract, security, and result publication.
Add/search API contract, request fields, polling, and response schemas.
Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions.
Academic leaderboard results will appear here once available.
Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions.
Academic leaderboard results will appear here once available.
Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions.
Academic leaderboard results will appear here once available.
Create API-gated eval jobs against the Leaderboard Suite, monitor task progress, inspect private results, and submit eligible full-suite runs for administrator review.
Choose a bound version. Use Run label to distinguish repeated evaluations of the same version.
Confirm each item before starting a public-board candidate run. The button remains locked until every item is checked.
Live job status and progress.
Review completed full-suite runs from non-admin leaderboard keys. Approved results enter the public board; rejected results remain private.
Private scores stay scoped to the current leaderboard key. Admin runs remain separate from external review candidates.
Browse benchmark datasets by dimension. Each card links to its source repository with task format, metrics, and evaluation notes.
Software engineering context benchmark for retrieval match and task solving.
The public leaderboard only accepts verified full-suite results. Smoke, light, failed, partial, and dataset-subset runs remain private.
A private result can be promoted to the public board only when every check passes.
Only full-mode runs can be submitted to the public leaderboard.
Selected datasets must match the current public suite.
The job and every evaluation task must finish successfully.
Result ownership is verified. External full-suite results can be published only after administrator approval.
A private result can be promoted once to avoid duplicate rows.
Mode controls what the run is allowed to do after scoring.
The separate integration-test endpoint verifies API binding and the Add/Search flow. External formal evaluation keys do not submit smoke jobs.
Reduced suite for iteration and debugging. Results stay private. For external keys, light and full have separate 30-day cooldowns, so a light run does not delay the next full run.
Full benchmark suite. A successful full run may be submitted when ownership and review checks pass.
Write a complete README and run guide, including the Docker command, API wrapper, original work disclosure, and method changes.
Provide a public GitHub repository for platform deployment, or a conformant Add/Search endpoint for a hosted submission.
We review the source and disclosures, start code submissions from the documented Docker entrypoint, then verify the API with smoke.
Complete every checklist item before launching the full benchmark run.
Successful full results enter administrator review before appearing on the public board.
Reused papers, repositories, and code must identify the original authors, technical report, and all method changes. Undisclosed reuse may be treated as plagiarism.
DisqualificationRepeated near-duplicate or low-quality submissions, prompt injection, benchmark leakage, result manipulation, malicious behavior, and other leaderboard abuse may cancel eligibility. For deployed endpoints, public availability and stability are required for 30 days after submission.
了解如何接入 Add / Search、运行统一评测、查看结果,并将符合条件的结果发布到公开榜单。
OVERVIEW
Agent Memory Leaderboard 使用统一的端到端流程评估不同记忆系统。我们不限定你使用的数据库、索引、向量模型或内部架构,只要求 Add / Search 接口符合现行规范,且评测过程和检索结果能够复核。
我们按样本和来源会话调用 Add 接口。
你的系统负责持久化、组织、索引和更新记忆。
我们针对每道题调用 Search 接口并接收排序结果。
我们使用固定流程生成答案、完成评分并汇总结果。
你只负责 Add / Search 环节。回答模型、提示词、评分器、数据集组合、Top K 和汇总规则由我们统一固定,使榜单分数尽可能反映记忆系统本身的能力。
WHO THIS IS FOR
你可以在 Textual、Multimodal 或 Coding Track 中查看 Overall、分项指标、系统版本和发布日期。不同 Track 使用的任务和指标不同,分数不宜跨 Track 直接比较。
实现 Add / Search → 提交准入申请 → 获取 API Key → 运行公开 smoke → 选择评测模式 → 查看私有结果 → 进入公榜审核。
PARTICIPATION WORKFLOW
先提交 Evaluation Access Request,完成参测系统、版本、代码或公开接口和鉴权方式的审核。代码提交必须附带完整 README、Docker 启动命令、API 封装说明和原始方法披露;平台会按说明启动并部署后再评测。托管接口提交则需在提交后至少 30 天保持公网可访问且稳定。
选择申请新 Key,或输入已有 Key 为其增加版本;填写联系人、系统版本、GitHub 仓库或 Add / Search 接口,并在提交说明中写清运行方式、Docker 命令、API 封装和原始方法信息。
审核通过后,新申请会签发 Leaderboard API Key;新增版本申请会直接绑定到原 Key,无需管理员再次手工录入接口信息。
代码提交由平台按 Docker 和 API 说明构建、启动并封装;托管提交直接校验你提供的接口。随后按照现行同步规范执行 Add → Search → Answer → Evaluate smoke 流程,确认接口可用后再进入正式评测。
选择 full 前,必须逐项勾选提交清单:smoke 已通过、API 契约正确、运行说明完整、托管接口可持续 30 天稳定,以及原创性披露和诚信承诺均已完成。
在 Evaluation 页面验证 API Key,选择已绑定的版本和评测模式,填写唯一的 Run Label 后提交。正式任务使用统一的数据集套件和 Top K。
任务依次执行写入、检索、回答、评分和结果汇总。页面会显示当前阶段、完成进度和错误摘要;我们同时保留复核所需的过程记录。
评测结果首先仅对绑定的 API Key 可见。成功完成的 full 任务会进入公榜资格校验和审核队列;审核通过后生成公榜提交记录。
允许复现已有论文或封装已有仓库,但必须披露原始作者、技术报告以及方法改动;未披露的重复代码可能被视为抄袭。发现重复代码时,平台会优先保护本人已披露并先提交的评测结果。重复提交相似或低质量代码、提示词注入、数据或结果操纵、恶意刷榜等行为,均可能取消参赛资格。
EVALUATION MODES
| 模式 | 用途 | 标准配额 | 结果范围 | 公榜资格 |
|---|---|---|---|---|
| smoke | 通过独立的兼容性测试接口检查同步 Add、Search 和评分链路。 | 每小时 1 次 | 仅私有 | 不可发布 |
| light | 用于版本迭代和内部效果验证的快速评测。 | 每月 1 次 | 仅私有 | 不可发布 |
| full | 运行固定的完整评测套件,并进入正式排队、审计和发布流程。 | 每 3 个月 1 次 | 优先私有 | 审核通过后发布 |
首次接入或接口行为变更后,使用独立的兼容性 smoke 检查接口规范;接口稳定后,使用 light 验证效果;只有准备正式发布的固定版本才建议提交 full。
ADD / SEARCH CONTRACT
Add 和 Search 的接口地址由你配置,请求和响应格式按照现行规范固定,不随 URL 路径变化。接口必须能够从我们的评测环境访问;生产环境建议使用 HTTPS。URL 中不得包含用户名、密码等凭据,也不得指向私有、回环或链路本地地址。
记忆写入完成后,Add 接口返回 HTTP 200
我们按照返回顺序最多读取 top_k 条记忆
Add 和 Search 必须使用完全相同的 user_id
Add / Search 支持 Token、Bearer 和 X-Api-Key;none 仅用于公开 smoke。Health 接口通过无需鉴权的 GET 请求调用,返回任意 2xx 状态码即表示服务正常。如果没有单独配置 Health 地址,正式任务会检查 Add 同源的 /health。
每个来源会话默认调用一次 Add;超过 20 条消息或 2000 个词时,在最近的完整消息或句子边界分段。请求只包含下列字段。
{
"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",
"messages": [{
"role": "user",
"timestamp": 1704067200000,
"content": "memory text"
}],
"user_id": "eval:run_abc123:locomo:conv-0",
"session_id": "eval:run_abc123:sample:0"
}
request_id必填。本次写入请求的唯一标识;成功响应必须原样返回该值。
messages必填。消息按原顺序排列,每条消息包含 role 和非空的 content。timestamp 可选,单位为 Unix 毫秒。
user_id必填。Search 接口唯一使用的检索范围标识;写入和检索时必须保持一致。
session_id必填。用于标识来源会话,可以用于组织记忆,但不作为 Search 的筛选条件。
未使用字段现行规范不发送 metadata、app_id、agent_id 或 async_mode;同步语义由 Add 完成写入后返回 HTTP 200 保证。
写入完成并且相关记忆能够立即检索后,才能返回成功响应。
HTTP 200
{
"success": true,
"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",
"user_id": "eval:run_abc123:locomo:conv-0",
"session_id": "eval:run_abc123:sample:0"
}
success必填,且必须是布尔值 true。
request_id / user_id / session_id全部必填,并且必须与请求中的值完全一致。
同步完成服务内部可以采用异步处理,但 Add 接口必须等待处理完成后再返回。
不支持的响应请勿返回 HTTP 202、task ID 或状态查询地址;响应中也不需要 memory_ids。
所有 Add 请求成功后,我们会针对每道题调用一次 Search 接口。查询使用数据集原文;选择题会另行传入选项。
{
"query": "Which answer best matches the memory?",
"options": ["A. First answer", "B. Second answer"],
"user_id": "eval:run_abc123:locomo:conv-0",
"top_k": 100
}
query必填。按照原文检索相关记忆,不得将其替换为最终答案,也不得使用评测金标。
options可选。题目有候选项时传入字符串数组;开放题不发送该字段。
user_id必填。只能在该 user_id 对应的记忆范围内检索。
top_k必填。返回的记忆数量不得超过该值;正式外部评测固定为 100。
请求格式现行规范不发送 filters、rerank 或 keyword_search。
响应必须是一个 JSON 对象,其中 data 为按相关性排序的数组。没有检索结果时,返回空数组。
{
"data": [{
"id": "mem_1",
"content": "remembered fact text",
"score": 0.87,
"created_at": "2026-07-01T12:00:00Z"
}]
}
data必填,类型为数组。不要增加 items 包装层,也不要直接返回顶层数组。
id必填,非空字符串,用于稳定标识该条记忆。
content必填,非空字符串。该内容会直接提供给统一回答模型。
score可选,数值类型。数值越大应表示相关性越高。
created_at可选,用于记录记忆的来源时间或持久化时间。
其他字段我们只读取上述字段;metadata 等未声明字段会被忽略。
ERROR HANDLING
参测接口应使用标准 HTTP 状态码,并返回便于排查问题且不含密钥的错误信息。平台业务错误通常采用 {"detail":{"reason":"..."}} 格式;字段校验失败时,HTTP 422 会返回结构化错误明细。
| HTTP | 错误类型 | 常见原因 | 处理方式 |
|---|---|---|---|
400 / 422 | 接口格式错误 | Add / Search 请求无法解析,或者成功响应缺少必填字段、字段类型不正确。 | 不自动重试。根据错误信息修正请求或响应格式,然后重新运行 smoke。 |
401 | 认证失败 | Memory System Key 无效,或者 Authorization / X-Api-Key 与申请时选择的方式不一致。 | 不自动重试。核对鉴权方式和绑定的密钥。 |
403 | 访问被拒绝 | 当前密钥无权调用接口,或者服务端拒绝访问指定的 user_id。 | 不自动重试。检查接口权限和检索范围配置。 |
404 | 资源不存在 | 接口路径错误,或者平台侧的 test、job、result ID 无效。 | 检查 URL 或资源 ID。现行同步规范不包含 Add Status 查询。 |
409 | 状态冲突 | Add 暂时无法写入;或者平台已有任务正在运行、Run Label 重复。 | Add 请求会在限定次数内重试,Search 请求遇到 409 不重试。平台任务冲突时,请等待现有任务结束或更换 Run Label。 |
408 / 425 | 暂时不可用 | 请求超时,或者服务尚未准备完成。 | Add / Search 请求都会按照平台策略进行有限次数的退避重试。 |
429 | 触发限流或配额 | 参测接口容量不足,或者平台评测配额尚未恢复。 | 接口调用会在限定次数内重试;如果是平台配额限制,请等待页面显示的下次可用时间。 |
500 / 502 / 503 / 504 | 临时服务异常 | 参测接口、网关或上游依赖暂时不可用。 | 我们会自动退避重试。若持续失败,请保留 Job/Test ID 和发生时间,便于进一步排查。 |
Add 遇到 408、409、425、429、500、502、503、504 时会重试;Search 的重试范围相同,但不包含 409。网络超时和传输错误也会触发重试,且重试次数有限。
即使 HTTP 状态码为 200,只要 Add 未返回 success=true、三个 ID 未正确返回,或者 Search 未返回 data 数组、某条记录缺少 id / content,当前阶段都会立即失败。
DATA, SECURITY & PRIVACY
你的接口只会收到当前任务所需的记忆片段、user/session 标识和检索问题。我们不会提供金标答案、评分依据或完整数据集下载。
user_id 是 Search 接口唯一使用的检索范围标识,存储和检索时必须完全一致。session_id 只用于组织来源会话。禁止跨 user_id 返回记忆。
Memory System Key 通过受控的申请流程提交并加密保存,任务信息中只保留不可读的引用。密钥不会出现在邮件、公榜或公开 API 响应中。
我们只连接通过网络校验的公开 HTTP(S) 接口,并拒绝 URL 中包含凭据,或者指向私有、回环、链路本地地址的目标。
我们会保留复核所需的请求结果、耗时、错误、候选记忆和接口格式校验记录,用于确认评测是否完整,以及结果是否符合公榜条件。
私有任务和结果仅对绑定的 Leaderboard API Key 可见。公榜只展示审核通过的系统、版本、分数和必要的评测信息。
评测数据及其派生副本只能用于完成当前任务,不得用于模型训练、微调、产品分析、数据集重建或对外传播。请仅向必要人员开放访问权限,避免记录不必要的请求正文,并在任务完成后 30 天内删除相关数据;如需延长保留时间,必须事先获得我们的书面同意。
EVALUATION & RESULTS
每个记忆分块都必须在返回 HTTP 200 前完成持久化,并且能够立即检索。
我们会检查 data 数组、必填字段和 Top K,并按照接口返回的顺序接收候选记忆。
通过校验的候选记忆会进入统一的回答模型和提示词流程。
我们按照题型使用固定的评分规则,并汇总各数据集和 Overall 结果。
任务完成后,你可以使用绑定的 API Key 查看运行状态、分项结果、错误摘要和任务信息。
full 结果通过资格校验和审核后,会生成公榜提交记录并展示在公开榜单中。
结果必须来自成功完成的 full 模式固定套件,且所有评测任务均执行成功。统一回答模型、评测规范、pipeline code hash、dataset bundle hash 和题量记录必须完整,并与现行发布基线一致。结果不得重复提交,同时还需通过平台审核。
Learn how to integrate Add / Search, run consistent evaluations, review results, and publish eligible runs to the public leaderboard.
OVERVIEW
Agent Memory Leaderboard compares memory systems under a consistent end-to-end evaluation. The platform does not prescribe a database, index, embedding model, or internal architecture. It requires a conformant external API and verifiable evidence for retrieval and audit.
The platform calls Add by sample and source session.
Your system persists, organizes, indexes, and updates memory.
The platform calls Search per question and receives ranked results.
A fixed platform workflow generates answers, scores them, and aggregates results.
Participant-controlled behavior is limited to Add / Search. The platform locks the answer model, prompts, evaluators, dataset suite, Top K, and aggregation rules so that score differences primarily reflect the memory system.
WHO THIS IS FOR
Review Overall scores, metric breakdowns, system versions, and publication dates within the Textual, Multimodal, or Coding track. Each track has a distinct task and metric contract; scores are not comparable across tracks.
Implement Add / Search → submit an access request → receive an API Key → run public smoke → select an evaluation mode → review private results → enter public review.
PARTICIPATION WORKFLOW
Industry submissions and academic systems with hosted APIs follow the existing review, Key issuance, and public smoke flow. Code submissions must include a complete README, Docker startup command, API wrapper instructions, and original-method disclosure; maintainers build and start the documented entrypoint before evaluation. Hosted endpoints must remain publicly reachable and stable for at least 30 days after submission.
Industry requests provide Add / Search endpoints. Code submissions provide a publicly accessible GitHub repository and explain the Docker command, API wrapper, original authors, technical report, and every method change.
Approval either issues a new Leaderboard API Key or binds the submitted version to the verified existing key without administrator re-entry.
For code submissions, maintainers build and start the documented Docker entrypoint and expose the official API. Hosted submissions are checked at the declared endpoint. The platform then runs the synchronous Add → Search → Answer → Evaluate smoke flow.
Before selecting full, confirm the interactive checklist: smoke passed, API contract followed, run instructions complete, hosted runtime stable for 30 days, original work disclosed, and no integrity violations.
Verify the API Key on the Evaluation page, select a bound version and mode, assign a unique Run Label, and submit. Formal jobs use the fixed suite and platform Top K.
The job proceeds through ingestion, retrieval, answering, scoring, and aggregation. The interface reports stage, progress, and error summaries while the platform retains evidence required for review.
Results are initially visible only to the bound API Key. A successful full job enters the public-eligibility gate and review queue; approval creates a public submission.
Reproductions of papers or existing repositories are allowed only with full attribution to the original authors, technical report, and all method changes. Undisclosed reuse may be treated as plagiarism. When duplicate code is discovered, the platform prioritizes the participant's first disclosed submission. Repeated near-duplicate or low-quality code, prompt injection, benchmark or result manipulation, malicious behavior, and other leaderboard abuse may cancel eligibility.
EVALUATION MODES
| Mode | Purpose | Standard quota | Visibility | Public eligibility |
|---|---|---|---|---|
| smoke | Separate integration endpoint for synchronous Add, Search, and scoring compatibility. | 1 per hour | Private | Not eligible |
| light | Fast internal performance validation for an iterating version. | 1 per month | Private | Not eligible |
| full | Fixed complete benchmark suite with formal queueing, audit, and release workflow. | 1 every 3 months | Private first | Eligible after approval |
Use the separate compatibility smoke after initial integration or an API behavior change, light for performance checks, and full only for a version ready for publication.
ADD / SEARCH CONTRACT
Participants configure the Add and Search URLs; request and response schemas are fixed and do not vary by path. Endpoints must be reachable from the platform network, and production deployments should use HTTPS. URLs must not embed credentials or resolve to private, loopback, or link-local addresses.
Add succeeds only with HTTP 200 after persistence
The platform reads at most top_k items in response order
Add and Search must use the identical isolation boundary
Add / Search support Token, Bearer, and X-Api-Key; none is limited to public smoke. Health is an unauthenticated GET where any 2xx means healthy. If no custom Health URL is bound, formal jobs check /health on the Add origin.
Each source session uses one Add by default. Sessions over 20 messages or 2,000 words split at the nearest complete message or sentence boundary. Requests contain only fields declared by the synchronous contract.
{
"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",
"messages": [{
"role": "user",
"timestamp": 1704067200000,
"content": "memory text"
}],
"user_id": "eval:run_abc123:locomo:conv-0",
"session_id": "eval:run_abc123:sample:0"
}
request_idRequired. Unique identifier for this chunk request; the success response must echo it exactly.
messagesRequired ordered array. Each item includes role and non-empty content; timestamp is an optional Unix-millisecond event time.
user_idRequired. The sole retrieval-isolation field; Search must use the identical value.
session_idRequired. Identifies the source session and may be used for grouping, but is not a Search filter.
Fields not sentThe current contract omits metadata, app_id, agent_id, and async_mode. Synchronous behavior is guaranteed by returning HTTP 200 only after Add completes.
Return success only after the write is persisted and immediately searchable.
HTTP 200
{
"success": true,
"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",
"user_id": "eval:run_abc123:locomo:conv-0",
"session_id": "eval:run_abc123:sample:0"
}
successRequired and must be the boolean value true.
request_id / user_id / session_idAll are required and must match the request byte for byte.
Synchronous completionInternal work may be asynchronous, but the endpoint must wait for completion before responding.
Not supportedDo not return HTTP 202, a task ID, or a status-poll URL. The response does not need memory_ids.
After every Add succeeds, the platform sends one Search request per question. The query stays in its original language; choice questions include options separately.
{
"query": "Which answer best matches the memory?",
"options": ["A. First answer", "B. Second answer"],
"user_id": "eval:run_abc123:locomo:conv-0",
"top_k": 100
}
queryRequired. Use the original text for retrieval; do not replace it with a final answer or use benchmark gold data.
optionsOptional. A string array sent for questions with answer choices; omitted for open questions.
user_idRequired. Retrieve only from memory stored under this exact value.
top_kRequired. The response must not exceed this number; formal external evaluations use 100.
Fixed schemaThe current contract does not send filters, rerank, or keyword_search.
Return a JSON object whose data field is a relevance-ordered array. Return an empty array when nothing is found.
{
"data": [{
"id": "mem_1",
"content": "remembered fact text",
"score": 0.87,
"created_at": "2026-07-01T12:00:00Z"
}]
}
dataRequired array. Do not add an items wrapper or return a top-level array.
idRequired non-empty string that stably identifies the memory.
contentRequired non-empty string passed directly to the platform Answer model.
scoreOptional number. Higher values should indicate greater relevance.
created_atOptional source or persistence timestamp.
Other fieldsThe platform reads only the declared fields above; undeclared fields such as metadata are ignored.
ERROR HANDLING
Participant endpoints should use standard HTTP status codes and return actionable errors without secrets. Platform business errors normally use {"detail":{"reason":"..."}}; HTTP 422 provides structured field-validation details.
| HTTP | Class | Typical case | Platform behavior and action |
|---|---|---|---|
400 / 422 | Contract error | An Add / Search request cannot be parsed, or a success response has missing or invalid required fields. | Not retried. Correct the schema using the error details, then rerun smoke. |
401 | Authentication failure | The Memory System Key is invalid, or the Authorization / X-Api-Key scheme does not match. | Not retried. Verify the authentication scheme and secret bound to the request. |
403 | Access denied | The key cannot call the endpoint, or the service rejects the current user_id scope. | Not retried. Review endpoint authorization and retrieval isolation. |
404 | Resource not found | An endpoint path is wrong, or a platform test, job, or result ID is invalid. | Verify the URL or resource ID. The synchronous contract has no Add Status polling. |
409 | State conflict | Add is temporarily conflicted, or a platform job is active and the Run Label is duplicated. | Add is retried with bounds; Search 409 is not retried. Wait or choose a new Run Label for platform conflicts. |
408 / 425 | Transient unavailability | The request timed out or the service is not ready. | Add and Search are retried with bounded backoff. |
429 | Rate or quota limit | The participant endpoint is capacity-limited, or a platform evaluation quota has not reset. | Endpoint calls are retried with bounds; for platform quotas, wait until the displayed availability time. |
500 / 502 / 503 / 504 | Transient service failure | The participant endpoint, gateway, or upstream dependency is temporarily unavailable. | The platform retries with backoff. Preserve the Job/Test ID and timestamp if failures persist. |
Add retries 408, 409, 425, 429, 500, 502, 503, and 504. Search retries the same set except 409. Network timeouts and transport failures are also eligible.
Even with HTTP 200, the stage fails if Add omits success=true or mis-echoes an ID, or if Search omits the data array or an item lacks id / content.
DATA, SECURITY & PRIVACY
Participant endpoints receive only the memory chunks, user/session identifiers, and retrieval questions needed for the current job. Gold answers, scoring criteria, and bulk dataset downloads are not provided.
user_id is the sole Search-isolation field and must match exactly during storage and retrieval. session_id is only for source-session organization. Cross-user_id retrieval is prohibited.
The Memory System Key is submitted through a controlled request flow and stored encrypted; job metadata contains only an opaque reference. Secrets do not appear in email, public rankings, or public API responses.
The platform connects only to network-validated public HTTP(S) endpoints and rejects URLs with embedded credentials or targets resolving to private, loopback, or link-local addresses.
The platform retains request outcomes, latency, errors, returned memories, and contract-validation evidence required to verify evaluation integrity and public eligibility.
Private jobs and results are visible only to the bound Leaderboard API Key. The public board shows approved systems, versions, scores, and required evaluation metadata.
Evaluation data and derived copies may be used only to complete the current job. Do not use them for training, fine-tuning, product analytics, dataset reconstruction, or redistribution. Restrict access to authorized personnel, avoid unnecessary payload logging, and delete the data within 30 days after job completion unless the platform approves another retention period in writing.
EVALUATION & RESULTS
Each chunk must be persisted and searchable before its HTTP 200 response.
The platform validates the data array, required fields, and Top K, then accepts candidates in participant order.
Valid candidate memories enter the locked platform Answer model and prompt workflow.
The platform applies fixed scoring contracts by question type and aggregates dataset and Overall results.
After completion, the bound API Key can inspect run status, breakdowns, error summaries, and job metadata.
An eligible full result creates a public submission only after qualification checks and review.
The result must come from a successful full run on the fixed suite; every evaluation task must succeed; the Answer model, evaluation contract, pipeline code hash, dataset bundle hashes, and question counts must be complete and match the current release baseline; and the result must be unique and pass platform review.
Hosted integrations expose synchronous Add and Search HTTPS endpoints. Code submissions provide a public GitHub repository, Docker startup command, and API wrapper instructions; maintainers deploy and evaluate the submitted system.
The platform sends one request for each memory chunk. Store every message before responding, and associate the data with the supplied user_id. The session_id identifies the source conversation and may be used for grouping, but retrieval isolation is based on user_id.
{
"request_id": "eval:<run_id>:locomo_refined:conv-0:chunk-0",
"messages": [{
"role": "user",
"timestamp": 1704067200000,
"content": "raw memory text"
}],
"user_id": "eval:<run_id>:locomo:conv-0",
"session_id": "eval:<run_id>:sample:0"
}
user or assistant.
contentRequiredOriginal message text supplied for ingestion.
timestampOptionalMessage timestamp in Unix milliseconds when available.
user_idRequiredRetrieval scope. Store it exactly and use it to isolate later searches.
session_idRequiredIdentifier for the source conversation or session.
Return HTTP 200 only after the submitted messages have been fully stored and are available to Search. If your system performs ingestion in the background, wait for that work to finish before returning success; otherwise the benchmark may search before the memory is ready.
{
"success": true,
"request_id": "eval:<run_id>:locomo_refined:conv-0:chunk-0",
"user_id": "eval:<run_id>:locomo:conv-0",
"session_id": "eval:<run_id>:sample:0"
}
true. It confirms that the messages are stored and searchable.
request_idRequiredExact request_id received in the Add request.
user_idRequiredExact user_id received in the Add request.
session_idRequiredExact session_id received in the Add request.
After every Add request has completed, the platform sends one Search request per benchmark question. The query remains unchanged, and choice questions carry their options in a separate optional field.
{
"query": "Which answer best matches the memory?",
"options": ["A. First answer", "B. Second answer"],
"user_id": "eval:<run_id>:locomo:conv-0",
"top_k": 100
}
Return a data array ordered from most relevant to least relevant. The platform preserves this order and passes the returned content to the shared answer pipeline. Return an empty array when no relevant memory is available.
{
"data": [
{
"id": "mem_123",
"content": "remembered fact text",
"score": 0.87,
"created_at": "2026-07-01T12:00:00Z"
}
]
}
Your API receives benchmark content during each run.
Your API receives evaluation memories and questions.
Use this data only for the evaluation. Do not train on it, analyze it, or share it.
Keep it private, avoid storing logs, and delete it within 30 days after the run.
A compact map of the product surface so users know where to go after their endpoints are ready.
Overview and benchmark preview.
Public rankings by track.
Verify key, start eval jobs, inspect private results and attribution.
Dataset cards and details links.
User workflow, API contract, security, and result publication.