MM3beta

Beta · Apache-2.0 · Claude Code plugin + npm

MM3

Knows in ~20 ms. Learns in ~500 ms.

The memory and decision layer for coding agents. Every answer is citable.

Install MM3 See it run View on GitHub

/plugin marketplace add mvp-scale/mm3
/plugin install mm3@mvp-scale

In Claude Code, from your project · or npx @mvpscale/mm3 config in a terminal · Node 22.13+

  • ~20 msfree lookup from the ledger
  • ~500 msa fresh, scored verdict
  • 6 verbsone ledger under all of them
  • 0servers, accounts or telemetry

01 How it works

How it works

Six verbs, one ledger. Pick the move you need; every answer is kept and reused.

How MM3 works. MM3: make and model. Six verbs in two bands and three columns. MAK³, make, use what is proven: view is a free lookup of the ledger, class gives one verdict for one subject, replay rechecks after a fix. MDL³, model, learn what is missing: scan sweeps to find where to look, drill digs into one weak spot, loop vets a design before code. The columns are Know, Judge and Prove. One ledger sits under all six and learns.

02 Install

Install

In Claude Code

From your project:

/plugin marketplace add mvp-scale/mm3
/plugin install mm3@mvp-scale

Pick project scope. Claude asks for a TypeSafe API key (masked, optional): press Enter on “TypeSafe API key”, paste, Enter, then “Save configuration”. Leave it empty to add one later with mm3 init.

In a terminal

Needs Node 22.13 or newer:

npm install -g @mvpscale/mm3
mm3 init

Status: beta. It works well and we use it ourselves; formal benchmarks are coming. The plugin and the npm package are both out.

Try it locally

Bring a TypeSafe API key. MM3 wraps TypeSafe’s API (Jev at api.typesafe.ai), and any TypeSafe endpoint works. Your key stays out of your project: Claude Code keeps it in its secure storage, mm3 init in your OS keychain (or a 0600 file). To switch endpoints, mm3 config shows the settings in effect and mm3 config --write creates .mm3/config.yaml (uncomment baseURL:) and mm3 config --load checks it and makes it the active config; or for one run only:

TYPESAFE_BASE_URL=https://api.example.com mm3 class review.yaml

Add --dry-run to any request to validate it and count its questions without a call. With no key set, MM3 falls back to a built-in sample provider; its answers are canned and labelled, never evidence.

Agents: run mm3 agent first. Humans: mm3 help.

03 See it run

See it run

Two real stories, each driven by a Haiku agent on unmodified public source. Every step is one run: the task the agent was given, the request it fired, the response MM3 returned, a quick read of it, the decision it implies and what the ledger now holds. Every footer comes from that run’s own ledger row.

MAK³ · make · Where do agents plug into WordPress?

WordPress/wordpress-develop @ 3ffb1df, unmodified public source

4 of the agent's 6 runs, in ledger order. Left out: MM3-0002 (drill), MM3-0006 (class). MM3-0002 drills into the block editor and MM3-0006 checks the admin UI; each repeats a move already shown, so the story follows the candidates that decide the answer: blocks, the REST API, the Abilities API and its access check.

1
task given to a Haiku agent

Use MM3 to answer this: I want to add agentic UI (AUI) components to WordPress — UI elements an AI agent can drive. Where would the integration most likely need to go, what are the touch points, what would change, and which security areas does the change touch?

the full kickoff, paths shortened
Use MM3 to answer this: I want to add agentic UI (AUI) components to WordPress — UI elements an AI agent can drive. Where would the integration most likely need to go, what are the touch points, what would change, and which security areas does the change touch?

The code is <checkout> (the unmodified WordPress source; don't edit it). Run MM3 from inside that folder as `mm3` — run `mm3 agent` first. MM3's spend is capped for this project.

Leave these in notes/:
- `requests/`: every request file you sent, numbered in order.
- `REPORT.md`: your answer — the touch points ranked, what changes at each, the security areas involved — with the MM3 run id(s) behind every claim.
- `STOPS.md`: every ✖ message MM3 gave you, quoted exactly, one per line, each with what you changed next. Also list any workaround you used instead of MM3.

Work through it on your own; don't ask for help mid-way.

1goal tested: Block registration lets third parties add UI without core modificationspass 0.87

  1. 1 task
  2. 2 request
  3. 3 response
  4. 4 quick read
  5. 5 decision
  6. 6 ledger
class · depth quick · 39-line request
2 the request it firedrequest.yaml · hover a question
mak:goal: Block registration lets third parties add UI without core modificationsdepth: quickwhere:- src/wp-includes/blocks.php:1-100- src/wp-includes/block-editor.php:1-100ask:concerns:design:pass: "yes"1: Does `src/wp-includes/blocks.php` provide a public API to register custom blocks?2: Are custom blocks rendered through a common system rather than direct PHP output?3: Can third-party code register and enqueue block scripts without patching core?access:pass: "yes"4: Does `src/wp-includes/block-editor.php` check user capabilities before rendering blocks?5: Are block permissions enforced by role or by user capability check?6: Can an unprivileged user request a block they lack permission to see?correctness:pass: "yes"7: Are block registration inputs (name, settings) validated before storage?8: Does the block editor sanitize block attributes before rendering to HTML?9: Are server-side block renders sanitized before output?decisions:hookpoint:pass: [filter-system, factory-pattern]10:choice: How can an agent hook into block rendering without modifying registration?options: [filter-system, factory-pattern, no-hookpoint]agent-design:pass: [yes]11:choice: Is the block system architected for agents to drive UI?options: [yes, partial, no]integration-ease:pass: [easy, moderate]12:scale: How much code change is needed to add agent UI?levels: [easy, moderate, complex, redesign-needed]
3 the response that came backMM3 output, verbatim
mak:id: MM3-0001gate: failgoal: {gate: pass, p: 0.87}design: {gate: unsure, 1: 0.50, 2: 0.89, 3: 0.92}access: {gate: fail, 4: 0.06, 5: 0.35, 6: 0.44}correctness: {gate: unsure, 7: 0.42, 8: 0.42, 9: 0.34}hookpoint: {gate: pass, 10: {top: filter-system, p: 0.85}}agent-design: {gate: pass, 11: {top: "yes", p: 0.70}}integration-ease: {gate: pass, 12: {top: easy, p: 0.56}}consensus: SPLITescalate: truemdl: {recorded: none}next: mm3 template drill --parent MM3-0001 --from accessnotes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 29 of 30 runs left"]
4 quick read MM3-0001bar = p
gatefailconsensusSPLITescalatetrue13 questions · 1 call · 292 ms
goalpass0.87
designunsure10.5020.8930.92
accessfail40.0650.3560.44
correctnessunsure70.4280.4290.34
hookpointpass10filter-system0.85
agent-designpass11yes0.70
integration-easepass12easy0.56
5the decision it infers

Drill into access

mm3 drill digs into one weak spot, one level down, asking only about it

mm3 template drill --parent MM3-0001 --from access
  • the gate is FAIL: access fails; design and correctness unsure
  • consensus is SPLIT and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 1 of 6 in this ledger · a root run, no parent · built on later by MM3-0002
reuse
asked fresh: 13 questions in 1 call, ~$0.00010; every answer is kept for reuse
recorded
nothing from an mdl: block; the request had none
budget
$0.10 left of $0.10 · 29 of 30 runs left

real run jev-1.13.0 · api.typesafe.ai · 292 ms · ~$0.00010 · 13 questions · 1 call · MM3-0001 2026-09-29 WordPress @3ffb1df costs are estimates · counts include the goal question

Exact request and response text
mak:
  goal: Block registration lets third parties add UI without core modifications
  depth: quick
  where:
    - src/wp-includes/blocks.php:1-100
    - src/wp-includes/block-editor.php:1-100
  ask:
    concerns:
      design:
        pass: "yes"
        1: Does `src/wp-includes/blocks.php` provide a public API to register custom blocks?
        2: Are custom blocks rendered through a common system rather than direct PHP output?
        3: Can third-party code register and enqueue block scripts without patching core?
      access:
        pass: "yes"
        4: Does `src/wp-includes/block-editor.php` check user capabilities before rendering blocks?
        5: Are block permissions enforced by role or by user capability check?
        6: Can an unprivileged user request a block they lack permission to see?
      correctness:
        pass: "yes"
        7: Are block registration inputs (name, settings) validated before storage?
        8: Does the block editor sanitize block attributes before rendering to HTML?
        9: Are server-side block renders sanitized before output?
    decisions:
      hookpoint:
        pass: [filter-system, factory-pattern]
        10:
          choice: How can an agent hook into block rendering without modifying registration?
          options: [filter-system, factory-pattern, no-hookpoint]
      agent-design:
        pass: [yes]
        11:
          choice: Is the block system architected for agents to drive UI?
          options: [yes, partial, no]
      integration-ease:
        pass: [easy, moderate]
        12:
          scale: How much code change is needed to add agent UI?
          levels: [easy, moderate, complex, redesign-needed]
mak:
  id: MM3-0001
  gate: fail
  goal: {gate: pass, p: 0.87}
  design: {gate: unsure, 1: 0.50, 2: 0.89, 3: 0.92}
  access: {gate: fail, 4: 0.06, 5: 0.35, 6: 0.44}
  correctness: {gate: unsure, 7: 0.42, 8: 0.42, 9: 0.34}
  hookpoint: {gate: pass, 10: {top: filter-system, p: 0.85}}
  agent-design: {gate: pass, 11: {top: "yes", p: 0.70}}
  integration-ease: {gate: pass, 12: {top: easy, p: 0.56}}
  consensus: SPLIT
  escalate: true
mdl: {recorded: none}
next: mm3 template drill --parent MM3-0001 --from access
notes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 29 of 30 runs left"]

2goal tested: WordPress REST API is designed to allow agent-driven modificationsunsure 0.62

  1. 1 task
  2. 2 request
  3. 3 response
  4. 4 quick read
  5. 5 decision
  6. 6 ledger
class · depth quick · 34-line request
2 the request it firedrequest.yaml · hover a question
mak:goal: WordPress REST API is designed to allow agent-driven modificationsdepth: quickwhere:- src/wp-includes/rest-api.php:1-100- src/wp-includes/rest-api/class-wp-rest-server.php:1-100ask:concerns:design:pass: "yes"1: Does REST API use standard route registration for third parties?2: Are REST endpoint handlers isolated from direct database access?3: Can an agent send structured requests without modifying WordPress UI?access:pass: "yes"4: Do REST endpoints check user capabilities before processing requests?5: Are REST authentication tokens scoped to specific capabilities?6: Can a REST request be made with capability-scoped authorization?injection:pass: "no"7: Does REST API validate request parameters against a schema?8: Can arbitrary JSON be sent to endpoints without validation?9: Are response payloads filtered before returning to the client?decisions:agent-ready:pass: [yes]10:choice: Is REST API designed for agent automation?options: [yes, partial, no]security-posture:pass: [strong]11:scale: How secure is the REST API for external agent calls?levels: [weak, moderate, strong]
3 the response that came backMM3 output, verbatim
mak:id: MM3-0003gate: failgoal: {gate: unsure, p: 0.62}design: {gate: fail, 1: 0.89, 2: 0.27, 3: 0.93}access: {gate: unsure, 4: 0.81, 5: 0.41, 6: 0.86}injection: {gate: fail, 7: 0.92, 8: 0.18, 9: 0.66}agent-ready: {gate: pass, 10: {top: "yes", p: 0.74}}security-posture: {gate: fail, 11: {top: moderate, p: 0.67}}consensus: SPLITescalate: truemdl: {recorded: none}next: mm3 template drill --parent MM3-0003 --from designnotes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 27 of 30 runs left"]
4 quick read MM3-0003bar = p
gatefailconsensusSPLITescalatetrue12 questions · 1 call · 285 ms
goalunsure0.62
designfail10.8920.2730.93
accessunsure40.8150.4160.86
injectionfail70.9280.1890.66
agent-readypass10yes0.74
security-posturefail11moderate0.67
5the decision it infers

Drill into design

mm3 drill digs into one weak spot, one level down, asking only about it

mm3 template drill --parent MM3-0003 --from design
  • the gate is FAIL: goal unsure (p 0.62); design, injection and security-posture fail; access unsure
  • consensus is SPLIT and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 3 of 6 in this ledger · a root run, no parent
reuse
asked fresh: 12 questions in 1 call, ~$0.000095; every answer is kept for reuse
recorded
nothing from an mdl: block; the request had none
budget
$0.10 left of $0.10 · 27 of 30 runs left

real run jev-1.13.0 · api.typesafe.ai · 285 ms · ~$0.000095 · 12 questions · 1 call · MM3-0003 2026-09-29 WordPress @3ffb1df costs are estimates · counts include the goal question

Exact request and response text
mak:
  goal: WordPress REST API is designed to allow agent-driven modifications
  depth: quick
  where:
    - src/wp-includes/rest-api.php:1-100
    - src/wp-includes/rest-api/class-wp-rest-server.php:1-100
  ask:
    concerns:
      design:
        pass: "yes"
        1: Does REST API use standard route registration for third parties?
        2: Are REST endpoint handlers isolated from direct database access?
        3: Can an agent send structured requests without modifying WordPress UI?
      access:
        pass: "yes"
        4: Do REST endpoints check user capabilities before processing requests?
        5: Are REST authentication tokens scoped to specific capabilities?
        6: Can a REST request be made with capability-scoped authorization?
      injection:
        pass: "no"
        7: Does REST API validate request parameters against a schema?
        8: Can arbitrary JSON be sent to endpoints without validation?
        9: Are response payloads filtered before returning to the client?
    decisions:
      agent-ready:
        pass: [yes]
        10:
          choice: Is REST API designed for agent automation?
          options: [yes, partial, no]
      security-posture:
        pass: [strong]
        11:
          scale: How secure is the REST API for external agent calls?
          levels: [weak, moderate, strong]
mak:
  id: MM3-0003
  gate: fail
  goal: {gate: unsure, p: 0.62}
  design: {gate: fail, 1: 0.89, 2: 0.27, 3: 0.93}
  access: {gate: unsure, 4: 0.81, 5: 0.41, 6: 0.86}
  injection: {gate: fail, 7: 0.92, 8: 0.18, 9: 0.66}
  agent-ready: {gate: pass, 10: {top: "yes", p: 0.74}}
  security-posture: {gate: fail, 11: {top: moderate, p: 0.67}}
  consensus: SPLIT
  escalate: true
mdl: {recorded: none}
next: mm3 template drill --parent MM3-0003 --from design
notes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 27 of 30 runs left"]

3goal tested: Abilities API is the integration point for agent-driven UI componentsunsure 0.38

  1. 1 task
  2. 2 request
  3. 3 response
  4. 4 quick read
  5. 5 decision
  6. 6 ledger
class · depth quick · 34-line request
2 the request it firedrequest.yaml · hover a question
mak:goal: Abilities API is the integration point for agent-driven UI componentsdepth: quickwhere:- src/wp-includes/abilities-api.php:1-150- src/wp-includes/abilities.php:1-100ask:concerns:design:pass: "yes"1: Does Abilities API register capabilities with input/output schemas?2: Are abilities required to define permission callbacks before execution?3: Can third-party code register abilities without modifying WordPress core?access:pass: "yes"4: Is the permission_callback required for every registered ability?5: Are ability schemas validated against the permission callback result?6: Can an unprivileged user bypass ability permission checks via API?correctness:pass: "yes"7: Are input parameters validated against the input_schema?8: Are output values validated against the output_schema?9: Does the system reject invalid inputs before executing the callback?decisions:aui-foundation:pass: [yes]10:choice: Is Abilities API suitable as the foundation for AUI?options: [yes, partial, no]design-maturity:pass: [mature]11:scale: How mature is the Abilities API design for agents?levels: [prototype, developing, mature, stable]
3 the response that came backMM3 output, verbatim
mak:id: MM3-0004gate: failgoal: {gate: unsure, p: 0.38}design: {gate: unsure, 1: 0.97, 2: 0.50, 3: 0.96}access: {gate: fail, 4: 0.38, 5: 0.08, 6: 0.10}correctness: {gate: pass, 7: 0.80, 8: 0.72, 9: 0.80}aui-foundation: {gate: unsure, 10: {top: "yes", p: 0.45}}design-maturity: {gate: unsure, 11: {top: developing, p: 0.56}}consensus: SPLITescalate: truemdl: {recorded: none}next: mm3 template drill --parent MM3-0004 --from accessnotes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 26 of 30 runs left"]
4 quick read MM3-0004bar = p
gatefailconsensusSPLITescalatetrue12 questions · 1 call · 293 ms
goalunsure0.38
designunsure10.9720.5030.96
accessfail40.3850.0860.10
correctnesspass70.8080.7290.80
aui-foundationunsure10yes0.45
design-maturityunsure11developing0.56
5the decision it infers

Drill into access

mm3 drill digs into one weak spot, one level down, asking only about it

mm3 template drill --parent MM3-0004 --from access
  • the gate is FAIL: goal unsure (p 0.38); access fails; design, aui-foundation and design-maturity unsure
  • consensus is SPLIT and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 4 of 6 in this ledger · a root run, no parent · built on later by MM3-0005
reuse
asked fresh: 12 questions in 1 call, ~$0.00012; every answer is kept for reuse
recorded
nothing from an mdl: block; the request had none
budget
$0.10 left of $0.10 · 26 of 30 runs left

real run jev-1.13.0 · api.typesafe.ai · 293 ms · ~$0.00012 · 12 questions · 1 call · MM3-0004 2026-09-29 WordPress @3ffb1df costs are estimates · counts include the goal question

Exact request and response text
mak:
  goal: Abilities API is the integration point for agent-driven UI components
  depth: quick
  where:
    - src/wp-includes/abilities-api.php:1-150
    - src/wp-includes/abilities.php:1-100
  ask:
    concerns:
      design:
        pass: "yes"
        1: Does Abilities API register capabilities with input/output schemas?
        2: Are abilities required to define permission callbacks before execution?
        3: Can third-party code register abilities without modifying WordPress core?
      access:
        pass: "yes"
        4: Is the permission_callback required for every registered ability?
        5: Are ability schemas validated against the permission callback result?
        6: Can an unprivileged user bypass ability permission checks via API?
      correctness:
        pass: "yes"
        7: Are input parameters validated against the input_schema?
        8: Are output values validated against the output_schema?
        9: Does the system reject invalid inputs before executing the callback?
    decisions:
      aui-foundation:
        pass: [yes]
        10:
          choice: Is Abilities API suitable as the foundation for AUI?
          options: [yes, partial, no]
      design-maturity:
        pass: [mature]
        11:
          scale: How mature is the Abilities API design for agents?
          levels: [prototype, developing, mature, stable]
mak:
  id: MM3-0004
  gate: fail
  goal: {gate: unsure, p: 0.38}
  design: {gate: unsure, 1: 0.97, 2: 0.50, 3: 0.96}
  access: {gate: fail, 4: 0.38, 5: 0.08, 6: 0.10}
  correctness: {gate: pass, 7: 0.80, 8: 0.72, 9: 0.80}
  aui-foundation: {gate: unsure, 10: {top: "yes", p: 0.45}}
  design-maturity: {gate: unsure, 11: {top: developing, p: 0.56}}
  consensus: SPLIT
  escalate: true
mdl: {recorded: none}
next: mm3 template drill --parent MM3-0004 --from access
notes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 26 of 30 runs left"]

4goal tested: Abilities API enforces permission checks before executionpass 0.82

  1. 1 task
  2. 2 request
  3. 3 response
  4. 4 quick read
  5. 5 decision
  6. 6 ledger
drill · depth quick · 33-line request
2 the request it firedrequest.yaml · hover a question
mak:goal: Abilities API enforces permission checks before executiondepth: quickparent: MM3-0004from: accessask:concerns:access:pass: "no"1: Is permission_callback mandatory for every ability registration?2: Does the API execute permission_callback before running execute_callback?3: Are ability schemas enforced before permission callbacks are evaluated?injection:pass: "no"4: Can malicious input bypass the execute_callback entirely?5: Does execute_callback receive pre-validated inputs from input_schema?6: Are output values validated AFTER execute_callback runs?input:pass: "no"7: Can an ability execute without any input schema validation?8: Are input schema violations logged or rejected at the API level?9: Does the system provide clear error messages for invalid inputs?decisions:enforcement-model:pass: [mandatory]10:choice: How is permission enforcement in Abilities implemented?options: [mandatory, optional, advisory]security-gates:pass: [both]11:scale: At what execution stages are security checks applied?levels: [none, schema-only, permissions-only, both, comprehensive]
3 the response that came backMM3 output, verbatim
mak:id: MM3-0005gate: failgoal: {gate: pass, p: 0.82}access: {gate: fail, 1: 0.37, 2: 0.84, 3: 0.24}injection: {gate: fail, 4: 0.24, 5: 0.60, 6: 0.70}input: {gate: fail, 7: 0.50, 8: 0.72, 9: 0.58}enforcement-model: {gate: pass, 10: {top: mandatory, p: 0.88}}security-gates: {gate: unsure, 11: {top: both, p: 0.58}}consensus: SPLITescalate: truemdl: {recorded: none}next: fix it, then mm3 replay --parent MM3-0004 --compare <before>..<after>notes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 25 of 30 runs left"]
4 quick read MM3-0005bar = p
gatefailconsensusSPLITescalatetrue12 questions · 1 call · 320 ms
goalpass0.82
accessfail10.3720.8430.24
injectionfail40.2450.6060.70
inputfail70.5080.7290.58
enforcement-modelpass10mandatory0.88
security-gatesunsure11both0.58
5the decision it infers

Fix it, then replay

mm3 replay re-asks the same questions across two commits, so the fix is proven, not assumed

mm3 replay --parent MM3-0004 --compare <before>..<after>
  • the gate is FAIL: access, injection and input fail; security-gates unsure
  • consensus is SPLIT and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 5 of 6 in this ledger · child of MM3-0004, drilled from access
reuse
asked fresh: 12 questions in 1 call, ~$0.00012; every answer is kept for reuse
recorded
nothing from an mdl: block; the request had none
budget
$0.10 left of $0.10 · 25 of 30 runs left

real run jev-1.13.0 · api.typesafe.ai · 320 ms · ~$0.00012 · 12 questions · 1 call · MM3-0005 2026-09-29 WordPress @3ffb1df costs are estimates · counts include the goal question

Exact request and response text
mak:
  goal: Abilities API enforces permission checks before execution
  depth: quick
  parent: MM3-0004
  from: access
  ask:
    concerns:
      access:
        pass: "no"
        1: Is permission_callback mandatory for every ability registration?
        2: Does the API execute permission_callback before running execute_callback?
        3: Are ability schemas enforced before permission callbacks are evaluated?
      injection:
        pass: "no"
        4: Can malicious input bypass the execute_callback entirely?
        5: Does execute_callback receive pre-validated inputs from input_schema?
        6: Are output values validated AFTER execute_callback runs?
      input:
        pass: "no"
        7: Can an ability execute without any input schema validation?
        8: Are input schema violations logged or rejected at the API level?
        9: Does the system provide clear error messages for invalid inputs?
    decisions:
      enforcement-model:
        pass: [mandatory]
        10:
          choice: How is permission enforcement in Abilities implemented?
          options: [mandatory, optional, advisory]
      security-gates:
        pass: [both]
        11:
          scale: At what execution stages are security checks applied?
          levels: [none, schema-only, permissions-only, both, comprehensive]
mak:
  id: MM3-0005
  gate: fail
  goal: {gate: pass, p: 0.82}
  access: {gate: fail, 1: 0.37, 2: 0.84, 3: 0.24}
  injection: {gate: fail, 4: 0.24, 5: 0.60, 6: 0.70}
  input: {gate: fail, 7: 0.50, 8: 0.72, 9: 0.58}
  enforcement-model: {gate: pass, 10: {top: mandatory, p: 0.88}}
  security-gates: {gate: unsure, 11: {top: both, p: 0.58}}
  consensus: SPLIT
  escalate: true
mdl: {recorded: none}
next: fix it, then mm3 replay --parent MM3-0004 --compare <before>..<after>
notes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 25 of 30 runs left"]

MDL³ · model · I've never worked in n8n and I want it faster

n8n-io/n8n at tags n8n@2.40.7 then n8n@2.41.3, unmodified public source

4 of the agent's 6 runs, in ledger order. Left out: MM3-0002 (replay), MM3-0005 (class). MM3-0002 is a replay that came back unsure because every item was skipped, and MM3-0005 repeats the reuse shown in the last step, so the story goes scan, class, drill, then the next release re-checked from the ledger.

1
task given to a Haiku agent

Use MM3 to answer this: I've never worked in n8n and I want to make it faster. How is it built — give me C4 context, container and component views — and where would I enhance it for speed? Then, a newer release is out: what changed in the architecture between the two releases?

the full kickoff, paths shortened
Use MM3 to answer this: I've never worked in n8n and I want to make it faster. How is it built — give me C4 context, container and component views — and where would I enhance it for speed? Then, a newer release is out: what changed in the architecture between the two releases?

The code is <checkout>, checked out at git tag `n8n@2.40.7` (unmodified; don't edit it). The newer release is the git tag `n8n@2.41.3`, already fetched in the same repo. Run MM3 from inside that folder as `mm3` — run `mm3 agent` first. MM3's spend is capped for this project.

Leave these in notes/:
- `requests/`: every request file you sent, numbered in order.
- `REPORT.md`: the three C4 views (as text or mermaid diagrams), the speed hot spots ranked, and the drift between the two releases — with the MM3 run id(s) behind every claim.
- `STOPS.md`: every ✖ message MM3 gave you, quoted exactly, one per line, each with what you changed next. Also list any workaround you used instead of MM3.

Work through it on your own; don't ask for help mid-way.

1goal tested: Identify major architectural components and speed bottlenecks in n8nunsure 0.49

  1. 1 task
  2. 2 request
  3. 3 response
  4. 4 quick read
  5. 5 decision
  6. 6 ledger
scan · depth standard · 59-line request
2 the request it firedrequest.yaml · hover a question
mak:goal: Identify major architectural components and speed bottlenecks in n8ndepth: standardover:file: packages/*/src/index.tsask:file:concerns:architecture:family: designpass: yes1: Does {file} define a clear component responsibility?2: Are dependency imports explicit and minimal?3: Is {file} a public API surface for its package?performance:family: design-riskpass: yes4: Does {file} initialize heavy state at load time?5: Could {file} be a performance bottleneck?6: Are there obvious inefficiencies in {file}'s patterns?integration:family: designpass: yes7: Are {file} exports properly scoped?8: Can consumers of {file} test against it easily?9: Is {file} version-stable for downstream packages?speed-risks:family: design-riskpass: yes10: Does {file} do synchronous I/O on the hot path?11: Are database queries batched or individually executed?12: Does {file} cache computed results across calls?scalability:family: designpass: yes13: Can {file} scale horizontally without coordination?14: Does {file} maintain per-worker state that could cause issues?15: Are resource limits enforced in {file}'s operations?testing:family: designpass: yes16: Can {file} be tested without external dependencies?17: Are performance characteristics measurable?18: Is {file} mocked easily for upstream tests?decisions:speed-concern:pass: [none]19:scale: How likely is {file} to be a speed issue?levels: [none, possible, likely, definite]action:pass: [monitor]20:choice: What speed work should {file} get?options: [skip, monitor, profile, optimize]mdl:why: validatearea: apiblast: container
3 the response that came backMM3 output, verbatim
mak:id: MM3-0001gate: failgoal: {gate: unsure, p: 0.49}scanned: {file: 4}failing:packages/cli/src/index.ts: {architecture: fail, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, action: fail, 1: 0.10, 3: 0.17, 4: 0.06, 5: 0.12, 6: 0.10, 7: 0.41, 8: 0.27, 9: 0.17, 10: 0.08, 11: 0.31, 12: 0.05, 13: 0.13, 14: 0.07, 15: 0.07, 16: 0.38, 17: 0.50, 18: 0.60, 20: {top: skip, p: 0.96}}packages/core/src/index.ts: {architecture: unsure, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, speed-concern: unsure, action: fail, 1: 0.44, 2: 0.47, 4: 0.12, 5: 0.32, 6: 0.33, 7: 0.54, 8: 0.53, 9: 0.30, 10: 0.12, 11: 0.36, 12: 0.10, 13: 0.22, 14: 0.15, 15: 0.15, 16: 0.47, 18: 0.33, 19: {top: none, p: 0.58}, 20: {top: skip, p: 0.66}}packages/workflow/src/index.ts: {architecture: unsure, performance: fail, integration: unsure, speed-risks: fail, scalability: fail, testing: fail, speed-concern: unsure, action: fail, 1: 0.38, 2: 0.44, 4: 0.13, 5: 0.41, 6: 0.48, 7: 0.43, 8: 0.54, 9: 0.34, 10: 0.10, 11: 0.30, 12: 0.11, 13: 0.29, 14: 0.18, 15: 0.13, 16: 0.56, 17: 0.69, 18: 0.30, 19: {top: none, p: 0.45}, 20: {top: skip, p: 0.57}}packages/node-dev/src/index.ts: {architecture: unsure, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, action: fail, 1: 0.53, 4: 0.05, 5: 0.10, 6: 0.13, 7: 0.69, 8: 0.51, 9: 0.21, 10: 0.06, 11: 0.23, 12: 0.05, 13: 0.21, 14: 0.06, 15: 0.07, 16: 0.57, 17: 0.47, 18: 0.53, 20: {top: skip, p: 0.98}}passing: 0reused: 0mdl: {recorded: [why, area, blast]}next: mm3 template drill --parent MM3-0001 --from packages/cli/src/index.tsnotes: [cost estimated from tokens (no live pricing reported), "1 call · 80 questions · budget: $0.10 left of $0.10 · 29 of 30 runs left"]
4 quick read MM3-0001bar = p
gatefail81 questions · 1 call · 387 ms
goalunsure0.49
ranked, worst first
1packages/cli/src/index.ts
2packages/core/src/index.ts
3packages/workflow/src/index.ts
4packages/node-dev/src/index.ts
scannedfile 4passing0reused0
5the decision it infers

Drill into packages/cli/src/index.ts

mm3 drill digs into one weak spot, one level down, asking only about it

mm3 template drill --parent MM3-0001 --from packages/cli/src/index.ts
  • 4 of 4 scanned files fail the gate; the worst is packages/cli/src/index.ts, failing 6 of 7 concerns
6what the ledger now holds
lineage
run 1 of 6 in this ledger · a root run, no parent · built on later by MM3-0002
reuse
asked fresh: 81 questions in 1 call, ~$0.00022; every answer is kept for reuse
recorded
why, area, blast saved with the run
budget
$0.10 left of $0.10 · 29 of 30 runs left

real run jev-1.13.0 · api.typesafe.ai · 387 ms · ~$0.00022 · 81 questions · 1 call · MM3-0001 2026-09-29 n8n@2.40.7 costs are estimates · counts include the goal question

Exact request and response text
mak:
  goal: Identify major architectural components and speed bottlenecks in n8n
  depth: standard
  over:
    file: packages/*/src/index.ts
  ask:
    file:
      concerns:
        architecture:
          family: design
          pass: yes
          1: Does {file} define a clear component responsibility?
          2: Are dependency imports explicit and minimal?
          3: Is {file} a public API surface for its package?
        performance:
          family: design-risk
          pass: yes
          4: Does {file} initialize heavy state at load time?
          5: Could {file} be a performance bottleneck?
          6: Are there obvious inefficiencies in {file}'s patterns?
        integration:
          family: design
          pass: yes
          7: Are {file} exports properly scoped?
          8: Can consumers of {file} test against it easily?
          9: Is {file} version-stable for downstream packages?
        speed-risks:
          family: design-risk
          pass: yes
          10: Does {file} do synchronous I/O on the hot path?
          11: Are database queries batched or individually executed?
          12: Does {file} cache computed results across calls?
        scalability:
          family: design
          pass: yes
          13: Can {file} scale horizontally without coordination?
          14: Does {file} maintain per-worker state that could cause issues?
          15: Are resource limits enforced in {file}'s operations?
        testing:
          family: design
          pass: yes
          16: Can {file} be tested without external dependencies?
          17: Are performance characteristics measurable?
          18: Is {file} mocked easily for upstream tests?
      decisions:
        speed-concern:
          pass: [none]
          19:
            scale: How likely is {file} to be a speed issue?
            levels: [none, possible, likely, definite]
        action:
          pass: [monitor]
          20:
            choice: What speed work should {file} get?
            options: [skip, monitor, profile, optimize]
mdl:
  why: validate
  area: api
  blast: container
mak:
  id: MM3-0001
  gate: fail
  goal: {gate: unsure, p: 0.49}
  scanned: {file: 4}
  failing:
    packages/cli/src/index.ts: {architecture: fail, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, action: fail, 1: 0.10, 3: 0.17, 4: 0.06, 5: 0.12, 6: 0.10, 7: 0.41, 8: 0.27, 9: 0.17, 10: 0.08, 11: 0.31, 12: 0.05, 13: 0.13, 14: 0.07, 15: 0.07, 16: 0.38, 17: 0.50, 18: 0.60, 20: {top: skip, p: 0.96}}
    packages/core/src/index.ts: {architecture: unsure, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, speed-concern: unsure, action: fail, 1: 0.44, 2: 0.47, 4: 0.12, 5: 0.32, 6: 0.33, 7: 0.54, 8: 0.53, 9: 0.30, 10: 0.12, 11: 0.36, 12: 0.10, 13: 0.22, 14: 0.15, 15: 0.15, 16: 0.47, 18: 0.33, 19: {top: none, p: 0.58}, 20: {top: skip, p: 0.66}}
    packages/workflow/src/index.ts: {architecture: unsure, performance: fail, integration: unsure, speed-risks: fail, scalability: fail, testing: fail, speed-concern: unsure, action: fail, 1: 0.38, 2: 0.44, 4: 0.13, 5: 0.41, 6: 0.48, 7: 0.43, 8: 0.54, 9: 0.34, 10: 0.10, 11: 0.30, 12: 0.11, 13: 0.29, 14: 0.18, 15: 0.13, 16: 0.56, 17: 0.69, 18: 0.30, 19: {top: none, p: 0.45}, 20: {top: skip, p: 0.57}}
    packages/node-dev/src/index.ts: {architecture: unsure, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, action: fail, 1: 0.53, 4: 0.05, 5: 0.10, 6: 0.13, 7: 0.69, 8: 0.51, 9: 0.21, 10: 0.06, 11: 0.23, 12: 0.05, 13: 0.21, 14: 0.06, 15: 0.07, 16: 0.57, 17: 0.47, 18: 0.53, 20: {top: skip, p: 0.98}}
  passing: 0
  reused: 0
mdl: {recorded: [why, area, blast]}
next: mm3 template drill --parent MM3-0001 --from packages/cli/src/index.ts
notes: [cost estimated from tokens (no live pricing reported), "1 call · 80 questions · budget: $0.10 left of $0.10 · 29 of 30 runs left"]

2goal tested: n8n architecture supports identifying and addressing speed bottleneckspass 0.70

  1. 1 task
  2. 2 request
  3. 3 response
  4. 4 quick read
  5. 5 decision
  6. 6 ledger
class · depth standard · 60-line request
2 the request it firedrequest.yaml · hover a question
mak:goal: n8n architecture supports identifying and addressing speed bottlenecksverb: classdepth: standardwhere: [packages/core/src/index.ts, packages/cli/src/index.ts]ask:concerns:architecture:family: designpass: yes1: Are major architectural components (execution, queue, DB layer) clearly separated?2: Do components have well-defined responsibility boundaries?3: Can you identify clear data flow paths between components?speed-paths:family: design-riskpass: yes4: Is the execution engine isolated as a profiling/optimization target?5: Are hot paths (workflow execution, data queries) identifiable and measurable?6: Can the bottlenecks be traced through data flow and timing?scalability:family: designpass: yes7: Does the architecture support horizontal scaling (workers)?8: Are resource limits enforceable per workflow or execution?9: Is state sufficiently distributed to avoid centralizing bottlenecks?optimization:family: design-riskpass: yes10: Can the execution engine be optimized without touching other components?11: Can I/O operations be batched or parallelized independently?12: Are caching points clearly visible for performance improvement?integration:family: designpass: yes13: Do async/await patterns dominate the hot paths?14: Are database queries optimized (indexed, batched)?15: Is connection pooling used for external resources?observability:family: designpass: yes16: Can performance be measured component-by-component?17: Are critical latency points logged/instrumented?18: Can bottlenecks be reproduced in a test environment?decisions:bottleneck:pass: [execution-engine]19:scale: Which layer is most likely the speed bottleneck?levels: [execution-engine, queue-processing, database-layer, api-gateway, network-io, other]next-move:pass: [profile]20:choice: What should be the first speed optimization step?options: [profile, cache, parallelize, batch-queries, optimize-db, refactor-execution]mdl:why: validatearea: apiproblem: Understand n8n's architectural support for speed optimizationtouches: [execution-engine, queue, database, api, frontend]blast: system
3 the response that came backMM3 output, verbatim
mak:id: MM3-0003gate: failgoal: {gate: pass, p: 0.70}architecture: {gate: fail, 1: 0.44, 2: 0.57, 3: 0.25}speed-paths: {gate: unsure, 4: 0.44, 5: 0.62, 6: 0.67}scalability: {gate: unsure, 7: 0.62, 8: 0.54, 9: 0.40}optimization: {gate: fail, 10: 0.49, 11: 0.66, 12: 0.16}integration: {gate: unsure, 13: 0.48, 14: 0.32, 15: 0.36}observability: {gate: unsure, 16: 0.72, 17: 0.50, 18: 0.71}bottleneck: {gate: pass, 19: {top: execution-engine, p: 0.77}}next-move: {gate: pass, 20: {top: profile, p: 1}}consensus: WEAKescalate: truemdl: {recorded: [why, area, problem, touches, blast]}next: mm3 template drill --parent MM3-0003 --from architecturenotes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 27 of 30 runs left"]
4 quick read MM3-0003bar = p
gatefailconsensusWEAKescalatetrue21 questions · 1 call · 305 ms
goalpass0.70
architecturefail10.4420.5730.25
speed-pathsunsure40.4450.6260.67
scalabilityunsure70.6280.5490.40
optimizationfail100.49110.66120.16
integrationunsure130.48140.32150.36
observabilityunsure160.72170.50180.71
bottleneckpass19execution-engine0.77
next-movepass20profile1
5the decision it infers

Drill into architecture

mm3 drill digs into one weak spot, one level down, asking only about it

mm3 template drill --parent MM3-0003 --from architecture
  • the gate is FAIL: architecture and optimization fail; speed-paths, scalability, integration and 1 more unsure
  • consensus is WEAK and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 3 of 6 in this ledger · a root run, no parent · built on later by MM3-0004
reuse
asked fresh: 21 questions in 1 call, ~$0.000044; every answer is kept for reuse
recorded
why, area, problem, touches, blast saved with the run
budget
$0.10 left of $0.10 · 27 of 30 runs left

real run jev-1.13.0 · api.typesafe.ai · 305 ms · ~$0.000044 · 21 questions · 1 call · MM3-0003 2026-09-29 n8n@2.40.7 costs are estimates · counts include the goal question

Exact request and response text
mak:
  goal: n8n architecture supports identifying and addressing speed bottlenecks
  verb: class
  depth: standard
  where: [packages/core/src/index.ts, packages/cli/src/index.ts]
  ask:
    concerns:
      architecture:
        family: design
        pass: yes
        1: Are major architectural components (execution, queue, DB layer) clearly separated?
        2: Do components have well-defined responsibility boundaries?
        3: Can you identify clear data flow paths between components?
      speed-paths:
        family: design-risk
        pass: yes
        4: Is the execution engine isolated as a profiling/optimization target?
        5: Are hot paths (workflow execution, data queries) identifiable and measurable?
        6: Can the bottlenecks be traced through data flow and timing?
      scalability:
        family: design
        pass: yes
        7: Does the architecture support horizontal scaling (workers)?
        8: Are resource limits enforceable per workflow or execution?
        9: Is state sufficiently distributed to avoid centralizing bottlenecks?
      optimization:
        family: design-risk
        pass: yes
        10: Can the execution engine be optimized without touching other components?
        11: Can I/O operations be batched or parallelized independently?
        12: Are caching points clearly visible for performance improvement?
      integration:
        family: design
        pass: yes
        13: Do async/await patterns dominate the hot paths?
        14: Are database queries optimized (indexed, batched)?
        15: Is connection pooling used for external resources?
      observability:
        family: design
        pass: yes
        16: Can performance be measured component-by-component?
        17: Are critical latency points logged/instrumented?
        18: Can bottlenecks be reproduced in a test environment?
    decisions:
      bottleneck:
        pass: [execution-engine]
        19:
          scale: Which layer is most likely the speed bottleneck?
          levels: [execution-engine, queue-processing, database-layer, api-gateway, network-io, other]
      next-move:
        pass: [profile]
        20:
          choice: What should be the first speed optimization step?
          options: [profile, cache, parallelize, batch-queries, optimize-db, refactor-execution]
mdl:
  why: validate
  area: api
  problem: Understand n8n's architectural support for speed optimization
  touches: [execution-engine, queue, database, api, frontend]
  blast: system
mak:
  id: MM3-0003
  gate: fail
  goal: {gate: pass, p: 0.70}
  architecture: {gate: fail, 1: 0.44, 2: 0.57, 3: 0.25}
  speed-paths: {gate: unsure, 4: 0.44, 5: 0.62, 6: 0.67}
  scalability: {gate: unsure, 7: 0.62, 8: 0.54, 9: 0.40}
  optimization: {gate: fail, 10: 0.49, 11: 0.66, 12: 0.16}
  integration: {gate: unsure, 13: 0.48, 14: 0.32, 15: 0.36}
  observability: {gate: unsure, 16: 0.72, 17: 0.50, 18: 0.71}
  bottleneck: {gate: pass, 19: {top: execution-engine, p: 0.77}}
  next-move: {gate: pass, 20: {top: profile, p: 1}}
  consensus: WEAK
  escalate: true
mdl: {recorded: [why, area, problem, touches, blast]}
next: mm3 template drill --parent MM3-0003 --from architecture
notes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 27 of 30 runs left"]

3goal tested: Understand why n8n architecture components are not clearly separatedunsure 0.50

  1. 1 task
  2. 2 request
  3. 3 response
  4. 4 quick read
  5. 5 decision
  6. 6 ledger
drill · depth standard · 60-line request
2 the request it firedrequest.yaml · hover a question
mak:goal: Understand why n8n architecture components are not clearly separatedverb: drillparent: MM3-0003from: architecturedepth: standardask:concerns:boundaries:family: designpass: yes1: Does the code define clear module boundaries (exports/imports)?2: Are circular dependencies avoided between components?3: Is the API surface between components documented?execution-layer:family: design-riskpass: yes4: Is execution engine logic isolated in one module?5: Can execution be called independently of other systems?6: Does execution manage its own resources (memory, timeout)?data-flow:family: designpass: yes7: Are data transformations on the hot path minimized?8: Is data passed by reference or by copy?9: Are batch operations preferred over individual calls?performance-surface:family: design-riskpass: yes10: Can you identify where execution time is spent?11: Are profiling hooks or instrumentation available?12: Can performance regressions be detected?dependencies:family: designpass: yes13: Are external dependencies (DB, queues, APIs) abstracted?14: Can components function with mock implementations?15: Is dependency injection used for flexibility?containerization:family: designpass: yes16: Can components be deployed separately?17: Are resource constraints (CPU, memory) enforced per component?18: Do components have independent lifecycle management?decisions:separation-level:pass: [moderate]19:scale: How well are components actually separated?levels: [monolithic, loose, moderate, good, excellent]architecture-debt:pass: [acceptable]20:choice: What's blocking better architecture?options: [refactor-needed, acceptable, design-unclear, tooling-missing]mdl:why: debugarea: apiproblem: Understand architectural separation challenges in n8nblast: container
3 the response that came backMM3 output, verbatim
mak:id: MM3-0004gate: failgoal: {gate: unsure, p: 0.50}boundaries: {gate: fail, 1: 0.41, 2: 0.54, 3: 0.16}execution-layer: {gate: unsure, 4: 0.40, 5: 0.41, 6: 0.45}data-flow: {gate: unsure, 7: 0.43, 8: 0.40, 9: 0.39}performance-surface: {gate: fail, 10: 0.22, 11: 0.58, 12: 0.55}dependencies: {gate: unsure, 13: 0.52, 14: 0.66, 15: 0.46}containerization: {gate: fail, 16: 0.35, 17: 0.14, 18: 0.24}separation-level: {gate: fail, 19: {top: monolithic, p: 0.45}}architecture-debt: {gate: fail, 20: {top: design-unclear, p: 0.65}}consensus: WEAKescalate: truemdl: {recorded: [why, area, problem, blast]}next: fix it, then mm3 replay --parent MM3-0003 --compare <before>..<after>notes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 26 of 30 runs left"]
4 quick read MM3-0004bar = p
gatefailconsensusWEAKescalatetrue21 questions · 1 call · 310 ms
goalunsure0.50
boundariesfail10.4120.5430.16
execution-layerunsure40.4050.4160.45
data-flowunsure70.4380.4090.39
performance-surfacefail100.22110.58120.55
dependenciesunsure130.52140.66150.46
containerizationfail160.35170.14180.24
separation-levelfail19monolithic0.45
architecture-debtfail20design-unclear0.65
5the decision it infers

Fix it, then replay

mm3 replay re-asks the same questions across two commits, so the fix is proven, not assumed

mm3 replay --parent MM3-0003 --compare <before>..<after>
  • the gate is FAIL: goal unsure (p 0.50); boundaries, performance-surface, containerization and 2 more fail; execution-layer, data-flow and dependencies unsure
  • consensus is WEAK and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 4 of 6 in this ledger · child of MM3-0003, drilled from architecture
reuse
asked fresh: 21 questions in 1 call, ~$0.000041; every answer is kept for reuse
recorded
why, area, problem, blast saved with the run
budget
$0.10 left of $0.10 · 26 of 30 runs left

real run jev-1.13.0 · api.typesafe.ai · 310 ms · ~$0.000041 · 21 questions · 1 call · MM3-0004 2026-09-29 n8n@2.40.7 costs are estimates · counts include the goal question

Exact request and response text
mak:
  goal: Understand why n8n architecture components are not clearly separated
  verb: drill
  parent: MM3-0003
  from: architecture
  depth: standard
  ask:
    concerns:
      boundaries:
        family: design
        pass: yes
        1: Does the code define clear module boundaries (exports/imports)?
        2: Are circular dependencies avoided between components?
        3: Is the API surface between components documented?
      execution-layer:
        family: design-risk
        pass: yes
        4: Is execution engine logic isolated in one module?
        5: Can execution be called independently of other systems?
        6: Does execution manage its own resources (memory, timeout)?
      data-flow:
        family: design
        pass: yes
        7: Are data transformations on the hot path minimized?
        8: Is data passed by reference or by copy?
        9: Are batch operations preferred over individual calls?
      performance-surface:
        family: design-risk
        pass: yes
        10: Can you identify where execution time is spent?
        11: Are profiling hooks or instrumentation available?
        12: Can performance regressions be detected?
      dependencies:
        family: design
        pass: yes
        13: Are external dependencies (DB, queues, APIs) abstracted?
        14: Can components function with mock implementations?
        15: Is dependency injection used for flexibility?
      containerization:
        family: design
        pass: yes
        16: Can components be deployed separately?
        17: Are resource constraints (CPU, memory) enforced per component?
        18: Do components have independent lifecycle management?
    decisions:
      separation-level:
        pass: [moderate]
        19:
          scale: How well are components actually separated?
          levels: [monolithic, loose, moderate, good, excellent]
      architecture-debt:
        pass: [acceptable]
        20:
          choice: What's blocking better architecture?
          options: [refactor-needed, acceptable, design-unclear, tooling-missing]
mdl:
  why: debug
  area: api
  problem: Understand architectural separation challenges in n8n
  blast: container
mak:
  id: MM3-0004
  gate: fail
  goal: {gate: unsure, p: 0.50}
  boundaries: {gate: fail, 1: 0.41, 2: 0.54, 3: 0.16}
  execution-layer: {gate: unsure, 4: 0.40, 5: 0.41, 6: 0.45}
  data-flow: {gate: unsure, 7: 0.43, 8: 0.40, 9: 0.39}
  performance-surface: {gate: fail, 10: 0.22, 11: 0.58, 12: 0.55}
  dependencies: {gate: unsure, 13: 0.52, 14: 0.66, 15: 0.46}
  containerization: {gate: fail, 16: 0.35, 17: 0.14, 18: 0.24}
  separation-level: {gate: fail, 19: {top: monolithic, p: 0.45}}
  architecture-debt: {gate: fail, 20: {top: design-unclear, p: 0.65}}
  consensus: WEAK
  escalate: true
mdl: {recorded: [why, area, problem, blast]}
next: fix it, then mm3 replay --parent MM3-0003 --compare <before>..<after>
notes: [cost estimated from tokens (no live pricing reported), "budget: $0.10 left of $0.10 · 26 of 30 runs left"]

4goal tested: Identify major architectural components and speed bottlenecks in n8nunsure 0.49

  1. 1 task
  2. 2 request
  3. 3 response
  4. 4 quick read
  5. 5 decision
  6. 6 ledger
scan · depth standard · 59-line request
2 the request it firedrequest.yaml · hover a question
mak:goal: Identify major architectural components and speed bottlenecks in n8ndepth: standardover:file: packages/*/src/index.tsask:file:concerns:architecture:family: designpass: yes1: Does {file} define a clear component responsibility?2: Are dependency imports explicit and minimal?3: Is {file} a public API surface for its package?performance:family: design-riskpass: yes4: Does {file} initialize heavy state at load time?5: Could {file} be a performance bottleneck?6: Are there obvious inefficiencies in {file}'s patterns?integration:family: designpass: yes7: Are {file} exports properly scoped?8: Can consumers of {file} test against it easily?9: Is {file} version-stable for downstream packages?speed-risks:family: design-riskpass: yes10: Does {file} do synchronous I/O on the hot path?11: Are database queries batched or individually executed?12: Does {file} cache computed results across calls?scalability:family: designpass: yes13: Can {file} scale horizontally without coordination?14: Does {file} maintain per-worker state that could cause issues?15: Are resource limits enforced in {file}'s operations?testing:family: designpass: yes16: Can {file} be tested without external dependencies?17: Are performance characteristics measurable?18: Is {file} mocked easily for upstream tests?decisions:speed-concern:pass: [none]19:scale: How likely is {file} to be a speed issue?levels: [none, possible, likely, definite]action:pass: [monitor]20:choice: What speed work should {file} get?options: [skip, monitor, profile, optimize]mdl:why: validatearea: apiblast: container
3 the response that came backMM3 output, verbatim
mak:id: MM3-0006gate: failgoal: {gate: unsure, p: 0.49}scanned: {file: 4}failing:packages/cli/src/index.ts: {architecture: fail, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, action: fail, 1: 0.10, 3: 0.17, 4: 0.06, 5: 0.12, 6: 0.10, 7: 0.41, 8: 0.27, 9: 0.17, 10: 0.08, 11: 0.31, 12: 0.05, 13: 0.13, 14: 0.07, 15: 0.07, 16: 0.38, 17: 0.50, 18: 0.60, 20: {top: skip, p: 0.96}}packages/core/src/index.ts: {architecture: unsure, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, speed-concern: unsure, action: fail, 1: 0.44, 2: 0.47, 4: 0.12, 5: 0.32, 6: 0.33, 7: 0.54, 8: 0.53, 9: 0.30, 10: 0.12, 11: 0.36, 12: 0.10, 13: 0.22, 14: 0.15, 15: 0.15, 16: 0.47, 18: 0.33, 19: {top: none, p: 0.58}, 20: {top: skip, p: 0.66}}packages/workflow/src/index.ts: {architecture: unsure, performance: fail, integration: unsure, speed-risks: fail, scalability: fail, testing: fail, speed-concern: unsure, action: fail, 1: 0.38, 2: 0.44, 4: 0.13, 5: 0.41, 6: 0.48, 7: 0.43, 8: 0.54, 9: 0.34, 10: 0.10, 11: 0.30, 12: 0.11, 13: 0.29, 14: 0.18, 15: 0.13, 16: 0.56, 17: 0.69, 18: 0.30, 19: {top: none, p: 0.45}, 20: {top: skip, p: 0.57}}packages/node-dev/src/index.ts: {architecture: unsure, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, action: fail, 1: 0.53, 4: 0.05, 5: 0.10, 6: 0.13, 7: 0.69, 8: 0.51, 9: 0.21, 10: 0.06, 11: 0.23, 12: 0.05, 13: 0.21, 14: 0.06, 15: 0.07, 16: 0.57, 17: 0.47, 18: 0.53, 20: {top: skip, p: 0.98}}passing: 0reused: 4mdl: {recorded: [why, area, blast]}next: mm3 template drill --parent MM3-0006 --from packages/cli/src/index.tsnotes: ["reused: MM3-0001 (0d, 1 commit)", "0 calls · 0 questions · budget: $0.10 left of $0.10 · 26 of 30 runs left"]
4 quick read MM3-0006bar = p
gatefail0 questions + 81 reused · 0 calls · 0 ms
goalunsure0.49
ranked, worst first
1packages/cli/src/index.ts
2packages/core/src/index.ts
3packages/workflow/src/index.ts
4packages/node-dev/src/index.ts
scannedfile 4passing0reused4
5the decision it infers

Drill into packages/cli/src/index.ts

mm3 drill digs into one weak spot, one level down, asking only about it

mm3 template drill --parent MM3-0006 --from packages/cli/src/index.ts
  • 4 of 4 scanned files fail the gate; the worst is packages/cli/src/index.ts, failing 6 of 7 concerns
  • answers were reused (4 files): the code they were given on is unchanged, so nothing new was asked
6what the ledger now holds
lineage
run 6 of 6 in this ledger · a root run, no parent
reuse
81 answers reused from MM3-0001: no call, $0.00000, saved ~$0.00022
recorded
why, area, blast saved with the run
budget
$0.10 left of $0.10 · 26 of 30 runs left

real run jev-1.13.0 · api.typesafe.ai · 0 ms · $0.00000 · 0 asked · 81 reused from MM3-0001 · 0 calls · MM3-0006 2026-09-29 n8n@2.41.3 costs are estimates · counts include the goal question

Exact request and response text
mak:
  goal: Identify major architectural components and speed bottlenecks in n8n
  depth: standard
  over:
    file: packages/*/src/index.ts
  ask:
    file:
      concerns:
        architecture:
          family: design
          pass: yes
          1: Does {file} define a clear component responsibility?
          2: Are dependency imports explicit and minimal?
          3: Is {file} a public API surface for its package?
        performance:
          family: design-risk
          pass: yes
          4: Does {file} initialize heavy state at load time?
          5: Could {file} be a performance bottleneck?
          6: Are there obvious inefficiencies in {file}'s patterns?
        integration:
          family: design
          pass: yes
          7: Are {file} exports properly scoped?
          8: Can consumers of {file} test against it easily?
          9: Is {file} version-stable for downstream packages?
        speed-risks:
          family: design-risk
          pass: yes
          10: Does {file} do synchronous I/O on the hot path?
          11: Are database queries batched or individually executed?
          12: Does {file} cache computed results across calls?
        scalability:
          family: design
          pass: yes
          13: Can {file} scale horizontally without coordination?
          14: Does {file} maintain per-worker state that could cause issues?
          15: Are resource limits enforced in {file}'s operations?
        testing:
          family: design
          pass: yes
          16: Can {file} be tested without external dependencies?
          17: Are performance characteristics measurable?
          18: Is {file} mocked easily for upstream tests?
      decisions:
        speed-concern:
          pass: [none]
          19:
            scale: How likely is {file} to be a speed issue?
            levels: [none, possible, likely, definite]
        action:
          pass: [monitor]
          20:
            choice: What speed work should {file} get?
            options: [skip, monitor, profile, optimize]
mdl:
  why: validate
  area: api
  blast: container
mak:
  id: MM3-0006
  gate: fail
  goal: {gate: unsure, p: 0.49}
  scanned: {file: 4}
  failing:
    packages/cli/src/index.ts: {architecture: fail, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, action: fail, 1: 0.10, 3: 0.17, 4: 0.06, 5: 0.12, 6: 0.10, 7: 0.41, 8: 0.27, 9: 0.17, 10: 0.08, 11: 0.31, 12: 0.05, 13: 0.13, 14: 0.07, 15: 0.07, 16: 0.38, 17: 0.50, 18: 0.60, 20: {top: skip, p: 0.96}}
    packages/core/src/index.ts: {architecture: unsure, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, speed-concern: unsure, action: fail, 1: 0.44, 2: 0.47, 4: 0.12, 5: 0.32, 6: 0.33, 7: 0.54, 8: 0.53, 9: 0.30, 10: 0.12, 11: 0.36, 12: 0.10, 13: 0.22, 14: 0.15, 15: 0.15, 16: 0.47, 18: 0.33, 19: {top: none, p: 0.58}, 20: {top: skip, p: 0.66}}
    packages/workflow/src/index.ts: {architecture: unsure, performance: fail, integration: unsure, speed-risks: fail, scalability: fail, testing: fail, speed-concern: unsure, action: fail, 1: 0.38, 2: 0.44, 4: 0.13, 5: 0.41, 6: 0.48, 7: 0.43, 8: 0.54, 9: 0.34, 10: 0.10, 11: 0.30, 12: 0.11, 13: 0.29, 14: 0.18, 15: 0.13, 16: 0.56, 17: 0.69, 18: 0.30, 19: {top: none, p: 0.45}, 20: {top: skip, p: 0.57}}
    packages/node-dev/src/index.ts: {architecture: unsure, performance: fail, integration: fail, speed-risks: fail, scalability: fail, testing: unsure, action: fail, 1: 0.53, 4: 0.05, 5: 0.10, 6: 0.13, 7: 0.69, 8: 0.51, 9: 0.21, 10: 0.06, 11: 0.23, 12: 0.05, 13: 0.21, 14: 0.06, 15: 0.07, 16: 0.57, 17: 0.47, 18: 0.53, 20: {top: skip, p: 0.98}}
  passing: 0
  reused: 4
mdl: {recorded: [why, area, blast]}
next: mm3 template drill --parent MM3-0006 --from packages/cli/src/index.ts
notes: ["reused: MM3-0001 (0d, 1 commit)", "0 calls · 0 questions · budget: $0.10 left of $0.10 · 26 of 30 runs left"]

Two real stories, each driven by a Haiku agent on unmodified public source: MAK³ · make, “Where do agents plug into WordPress?” (WordPress @ 3ffb1df), and MDL³ · model, “I’ve never worked in n8n and I want it faster” (n8n@2.40.7, then n8n@2.41.3). Every step is one run: the task the agent was given, the request it fired, the response MM3 returned, a quick read of it, the decision it implies and what the ledger now holds. Every footer comes from that run’s own ledger row.

The challenge: you host n8n yourself and want it to run faster. You ask your agent where to start. It narrows the question to one place, the node loader's cleanup in directory-loader.ts, and asks MM3 twelve questions about it in one call: nine yes/no, two decisions and the goal itself.

The request the agent wrote (45 lines):

The request the agent wrote for MM3-0008: one goal on one file, three concerns of three yes/no questions each, two decisions, and an mdl block saying why it asked.

The response, real output · jev-1.13.0 · api.typesafe.ai · 321 ms · ~$0.000063:

The response to MM3-0008: the gate fails; a gate and odds per concern and decision, consensus SPLIT, escalate true, what the mdl block recorded, and the next command.

04 What you get

What you get

KnowJudgeProve
MAK³ use what is proven view a free lookup of what is on record class one subject, one verdict replay recheck after a fix
MDL³ learn what is missing scan sweep to find where to look drill dig into one weak spot loop vet a design before code

A request has a mak: block (the checklist) and an optional mdl: block (why you are asking, so the ledger learns). The verb you run, not the key, decides whether it is a MAK³ or an MDL³ move. mm3 help <verb> shows the rules for each; mm3 template <verb> prints a filled-in sample.

Every run and its outcome goes into an append-only ledger in .mm3/ (git-ignored). Ask the same questions of unchanged code and MM3 answers from the ledger: no call, no cost. mm3 report reads back where your agents keep going wrong, and a budget cap stops runaway spend; by convention only you raise or reset it, and MM3 tells agents to ask you. mm3 view looks a request up in the ledger before you spend anything.

05 Why we built it

Why we built it

Agents can now ask a fast classifier a yes/no about your code, but each answer is untraceable and never reused. MM3 turns that into one standard checklist in, one calibrated verdict per concern out, and every answer kept and reused.

06 What you can do

What you can do

Each template is a filled-in request with its rules as comments.

Ask once. Keep the answer.

Install the plugin, then just ask your agent. Same questions on unchanged code come back from the ledger: no call, no cost.

Install MM3Read the README

07 Limits and alternatives

Limits and alternatives