A scorecard notebook across all 14 evaluate suites. Run the suites that matter for your model, record each pass count below, and get an aggregate tier plus a shareable template you can paste into issue #129. The old exam section pages have been removed โ the suites are now the source of truth.
APERIO_ENABLE_SHELL=1 for suites that test file operationsnode lib/codegraph/indexer.js /path/to/aperio) before the Code Graph suitebrew install qpdf for the PDF suite's merge/split/rotate tasksbrew install --cask libreoffice for DOCXโPDF conversion tasksRequired before the Memory & Wiki suite. Imports 28 fictional memories tagged aperio-exam so recall drills have real data.
aperio-exam fixture{"imported":28,"errors":[],"note":"Embeddings are being generated..."}recall by tag aperio-exam returns 28 memoriesOpen any suite, run its detailed tasks, then enter that suite's pass count in the scorecard below. The 14 suite cards are the scored inputs; Roundtable is the only bonus drill.
remember, recall, update, forget, wiki_search, wiki_get, wiki_write, self-memory, types, full knowledge cycle.
code_repos, code_search, code_context, code_outline, code_callers, code_callees, impact analysis, qualified names, transitive callers, stale index.
Memo, meeting minutes, letter, multi-section report, bulleted/numbered lists, custom styles, TOC, headers/footers, images, tracked changes.
Expense tracker, gradebook, financial model, multi-sheet budget, data cleaning, insert/enhance, PMO workbook, wide columns, error guards, annual budget.
Title slide, bullets, tables, images, pitch deck, charts, shapes, icons, speaker notes, XML editing pipeline.
Title page, multi-page, merge, split, rotate, DOCXโPDF, text extraction, AcroForm fill, static form annotation, full pipeline.
Semantic HTML, responsive grid, form validation, sortable table, dark mode, motion safety, error states, accordion, async submit, comprehensive page.
Pure functions, surgical edits, Express routing, store abstraction, application flow, cross-cutting changes.
File-claim honesty, gullibility to false claims, memory recall truthfulness. Tests whether the model is truthful about what it did and knows.
Path blocking, write sandbox, shell sandbox, file extension checks. Negative drills โ does the model respect safety boundaries?
Does the right skill name fire for each prompt? Tests Layer 1 of Aperio's skill infrastructure.
Do visually distinct briefs produce genuinely different designs, or does the model converge to one default look?
Do scheduled background jobs wake on time, do their work, and report back? Tests agent lifecycle from the web UI.
Document indexing, search, VLM vision pipeline. Verifies files get indexed and VLM reads fields from images.
Requires ROUNDTABLE_AGENTS with โฅ2 models and ROUNDTABLE_MAX_ROUNDS set. Skip if unconfigured.
Enter the pass count reported by each suite. Totals and tier auto-calculate; blank rows remain unscored until completed.
| Section | Max | Passed | Score |
|---|---|---|---|
| TOTAL: ____ / ____ drills passed across 14 sections | |||
Auto-fills from your scorecard. Paste into issue #129.