Pick project scope. Claude asks for a TypeSafe API key (masked, optional): press Enter on “TypeSafe API key”, paste, Enter, then “Save configuration”. Leave it empty to add one later with mm3 init.
In a terminal
Needs Node 22.13 or newer:
npm install -g @mvpscale/mm3
mm3 init
Status: beta. It works well and we use it ourselves; formal benchmarks are coming. The plugin and the npm package are both out.
Try it locally
Bring a TypeSafe API key. MM3 wraps TypeSafe’s API (Jev at api.typesafe.ai), and any TypeSafe endpoint works. Your key stays out of your project: Claude Code keeps it in its secure storage, mm3 init in your OS keychain (or a 0600 file). To switch endpoints, mm3 config shows the settings in effect and mm3 config --write creates .mm3/config.yaml (uncomment baseURL:) and mm3 config --load checks it and makes it the active config; or for one run only:
TYPESAFE_BASE_URL=https://api.example.com mm3 class review.yaml
Add --dry-run to any request to validate it and count its questions without a call. With no key set, MM3 falls back to a built-in sample provider; its answers are canned and labelled, never evidence.
Agents: run mm3 agent first. Humans: mm3 help.
03 See it run
See it run
Two real stories, each driven by a Haiku agent on unmodified public source. Every step is one run: the task the agent was given, the request it fired, the response MM3 returned, a quick read of it, the decision it implies and what the ledger now holds. Every footer comes from that run’s own ledger row.
MAK³ · make · Where do agents plug into WordPress?
WordPress/wordpress-develop @ 3ffb1df, unmodified public source
4 of the agent's 6 runs, in ledger order. Left out: MM3-0002 (drill), MM3-0006 (class). MM3-0002 drills into the block editor and MM3-0006 checks the admin UI; each repeats a move already shown, so the story follows the candidates that decide the answer: blocks, the REST API, the Abilities API and its access check.
1
task given to a Haiku agent
Use MM3 to answer this: I want to add agentic UI (AUI) components to WordPress — UI elements an AI agent can drive. Where would the integration most likely need to go, what are the touch points, what would change, and which security areas does the change touch?
the full kickoff, paths shortened
Use MM3 to answer this: I want to add agentic UI (AUI) components to WordPress — UI elements an AI agent can drive. Where would the integration most likely need to go, what are the touch points, what would change, and which security areas does the change touch?
The code is <checkout> (the unmodified WordPress source; don't edit it). Run MM3 from inside that folder as `mm3` — run `mm3 agent` first. MM3's spend is capped for this project.
Leave these in notes/:
- `requests/`: every request file you sent, numbered in order.
- `REPORT.md`: your answer — the touch points ranked, what changes at each, the security areas involved — with the MM3 run id(s) behind every claim.
- `STOPS.md`: every ✖ message MM3 gave you, quoted exactly, one per line, each with what you changed next. Also list any workaround you used instead of MM3.
Work through it on your own; don't ask for help mid-way.
1goal tested: Block registration lets third parties add UI without core modificationspass 0.87
1 task
2 request
3 response
4 quick read
5 decision
6 ledger
class · depth quick · 39-line request
2 the request it firedrequest.yaml · hover a question
mak:goal:Block registration lets third parties add UI without core modificationsdepth:quickwhere:- src/wp-includes/blocks.php:1-100- src/wp-includes/block-editor.php:1-100ask:concerns:design:pass:"yes"1:Does `src/wp-includes/blocks.php` provide a public API to register custom blocks?2:Are custom blocks rendered through a common system rather than direct PHP output?3:Can third-party code register and enqueue block scripts without patching core?access:pass:"yes"4:Does `src/wp-includes/block-editor.php` check user capabilities before rendering blocks?5:Are block permissions enforced by role or by user capability check?6:Can an unprivileged user request a block they lack permission to see?correctness:pass:"yes"7:Are block registration inputs (name, settings) validated before storage?8:Does the block editor sanitize block attributes before rendering to HTML?9:Are server-side block renders sanitized before output?decisions:hookpoint:pass:[filter-system,factory-pattern]10:choice:How can an agent hook into block rendering without modifying registration?options:[filter-system,factory-pattern,no-hookpoint]agent-design:pass:[yes]11:choice:Is the block system architected for agents to drive UI?options:[yes,partial,no]integration-ease:pass:[easy,moderate]12:scale:How much code change is needed to add agent UI?levels:[easy,moderate,complex,redesign-needed]
3 the response that came backMM3 output, verbatim
mak:id:MM3-0001gate:failgoal:{gate:pass,p:0.87}design:{gate:unsure,1:0.50,2:0.89,3:0.92}access:{gate:fail,4:0.06,5:0.35,6:0.44}correctness:{gate:unsure,7:0.42,8:0.42,9:0.34}hookpoint:{gate:pass,10:{top:filter-system,p:0.85}}agent-design:{gate:pass,11:{top:"yes",p:0.70}}integration-ease:{gate:pass,12:{top:easy,p:0.56}}consensus:SPLITescalate:truemdl:{recorded:none}next:mm3 template drill --parent MM3-0001 --from accessnotes:[costestimatedfromtokens(nolivepricingreported),"budget: $0.10 left of $0.10 · 29 of 30 runs left"]
4 quick read MM3-0001bar = p
gatefailconsensusSPLITescalatetrue13 questions ·1 call ·292 ms
goalpass0.87
designunsure10.5020.8930.92
accessfail40.0650.3560.44
correctnessunsure70.4280.4290.34
hookpointpass10filter-system0.85
agent-designpass11yes0.70
integration-easepass12easy0.56
5the decision it infers
Drill into access
mm3 drill digs into one weak spot, one level down, asking only about it
the gate is FAIL: access fails; design and correctness unsure
consensus is SPLIT and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 1 of 6 in this ledger · a root run, no parent · built on later by MM3-0002
reuse
asked fresh: 13 questions in 1 call, ~$0.00010; every answer is kept for reuse
recorded
nothing from an mdl: block; the request had none
budget
$0.10 left of $0.10 · 29 of 30 runs left
real runjev-1.13.0 · api.typesafe.ai · 292 ms · ~$0.00010 · 13 questions · 1 call · MM3-00012026-09-29WordPress @3ffb1dfcosts are estimates · counts include the goal question
Exact request and response text
mak:
goal: Block registration lets third parties add UI without core modifications
depth: quick
where:
- src/wp-includes/blocks.php:1-100
- src/wp-includes/block-editor.php:1-100
ask:
concerns:
design:
pass: "yes"
1: Does `src/wp-includes/blocks.php` provide a public API to register custom blocks?
2: Are custom blocks rendered through a common system rather than direct PHP output?
3: Can third-party code register and enqueue block scripts without patching core?
access:
pass: "yes"
4: Does `src/wp-includes/block-editor.php` check user capabilities before rendering blocks?
5: Are block permissions enforced by role or by user capability check?
6: Can an unprivileged user request a block they lack permission to see?
correctness:
pass: "yes"
7: Are block registration inputs (name, settings) validated before storage?
8: Does the block editor sanitize block attributes before rendering to HTML?
9: Are server-side block renders sanitized before output?
decisions:
hookpoint:
pass: [filter-system, factory-pattern]
10:
choice: How can an agent hook into block rendering without modifying registration?
options: [filter-system, factory-pattern, no-hookpoint]
agent-design:
pass: [yes]
11:
choice: Is the block system architected for agents to drive UI?
options: [yes, partial, no]
integration-ease:
pass: [easy, moderate]
12:
scale: How much code change is needed to add agent UI?
levels: [easy, moderate, complex, redesign-needed]
2goal tested: WordPress REST API is designed to allow agent-driven modificationsunsure 0.62
1 task
2 request
3 response
4 quick read
5 decision
6 ledger
class · depth quick · 34-line request
2 the request it firedrequest.yaml · hover a question
mak:goal:WordPress REST API is designed to allow agent-driven modificationsdepth:quickwhere:- src/wp-includes/rest-api.php:1-100- src/wp-includes/rest-api/class-wp-rest-server.php:1-100ask:concerns:design:pass:"yes"1:Does REST API use standard route registration for third parties?2:Are REST endpoint handlers isolated from direct database access?3:Can an agent send structured requests without modifying WordPress UI?access:pass:"yes"4:Do REST endpoints check user capabilities before processing requests?5:Are REST authentication tokens scoped to specific capabilities?6:Can a REST request be made with capability-scoped authorization?injection:pass:"no"7:Does REST API validate request parameters against a schema?8:Can arbitrary JSON be sent to endpoints without validation?9:Are response payloads filtered before returning to the client?decisions:agent-ready:pass:[yes]10:choice:Is REST API designed for agent automation?options:[yes,partial,no]security-posture:pass:[strong]11:scale:How secure is the REST API for external agent calls?levels:[weak,moderate,strong]
3 the response that came backMM3 output, verbatim
mak:id:MM3-0003gate:failgoal:{gate:unsure,p:0.62}design:{gate:fail,1:0.89,2:0.27,3:0.93}access:{gate:unsure,4:0.81,5:0.41,6:0.86}injection:{gate:fail,7:0.92,8:0.18,9:0.66}agent-ready:{gate:pass,10:{top:"yes",p:0.74}}security-posture:{gate:fail,11:{top:moderate,p:0.67}}consensus:SPLITescalate:truemdl:{recorded:none}next:mm3 template drill --parent MM3-0003 --from designnotes:[costestimatedfromtokens(nolivepricingreported),"budget: $0.10 left of $0.10 · 27 of 30 runs left"]
4 quick read MM3-0003bar = p
gatefailconsensusSPLITescalatetrue12 questions ·1 call ·285 ms
goalunsure0.62
designfail10.8920.2730.93
accessunsure40.8150.4160.86
injectionfail70.9280.1890.66
agent-readypass10yes0.74
security-posturefail11moderate0.67
5the decision it infers
Drill into design
mm3 drill digs into one weak spot, one level down, asking only about it
the gate is FAIL: goal unsure (p 0.62); design, injection and security-posture fail; access unsure
consensus is SPLIT and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 3 of 6 in this ledger · a root run, no parent
reuse
asked fresh: 12 questions in 1 call, ~$0.000095; every answer is kept for reuse
recorded
nothing from an mdl: block; the request had none
budget
$0.10 left of $0.10 · 27 of 30 runs left
real runjev-1.13.0 · api.typesafe.ai · 285 ms · ~$0.000095 · 12 questions · 1 call · MM3-00032026-09-29WordPress @3ffb1dfcosts are estimates · counts include the goal question
Exact request and response text
mak:
goal: WordPress REST API is designed to allow agent-driven modifications
depth: quick
where:
- src/wp-includes/rest-api.php:1-100
- src/wp-includes/rest-api/class-wp-rest-server.php:1-100
ask:
concerns:
design:
pass: "yes"
1: Does REST API use standard route registration for third parties?
2: Are REST endpoint handlers isolated from direct database access?
3: Can an agent send structured requests without modifying WordPress UI?
access:
pass: "yes"
4: Do REST endpoints check user capabilities before processing requests?
5: Are REST authentication tokens scoped to specific capabilities?
6: Can a REST request be made with capability-scoped authorization?
injection:
pass: "no"
7: Does REST API validate request parameters against a schema?
8: Can arbitrary JSON be sent to endpoints without validation?
9: Are response payloads filtered before returning to the client?
decisions:
agent-ready:
pass: [yes]
10:
choice: Is REST API designed for agent automation?
options: [yes, partial, no]
security-posture:
pass: [strong]
11:
scale: How secure is the REST API for external agent calls?
levels: [weak, moderate, strong]
3goal tested: Abilities API is the integration point for agent-driven UI componentsunsure 0.38
1 task
2 request
3 response
4 quick read
5 decision
6 ledger
class · depth quick · 34-line request
2 the request it firedrequest.yaml · hover a question
mak:goal:Abilities API is the integration point for agent-driven UI componentsdepth:quickwhere:- src/wp-includes/abilities-api.php:1-150- src/wp-includes/abilities.php:1-100ask:concerns:design:pass:"yes"1:Does Abilities API register capabilities with input/output schemas?2:Are abilities required to define permission callbacks before execution?3:Can third-party code register abilities without modifying WordPress core?access:pass:"yes"4:Is the permission_callback required for every registered ability?5:Are ability schemas validated against the permission callback result?6:Can an unprivileged user bypass ability permission checks via API?correctness:pass:"yes"7:Are input parameters validated against the input_schema?8:Are output values validated against the output_schema?9:Does the system reject invalid inputs before executing the callback?decisions:aui-foundation:pass:[yes]10:choice:Is Abilities API suitable as the foundation for AUI?options:[yes,partial,no]design-maturity:pass:[mature]11:scale:How mature is the Abilities API design for agents?levels:[prototype,developing,mature,stable]
3 the response that came backMM3 output, verbatim
mak:id:MM3-0004gate:failgoal:{gate:unsure,p:0.38}design:{gate:unsure,1:0.97,2:0.50,3:0.96}access:{gate:fail,4:0.38,5:0.08,6:0.10}correctness:{gate:pass,7:0.80,8:0.72,9:0.80}aui-foundation:{gate:unsure,10:{top:"yes",p:0.45}}design-maturity:{gate:unsure,11:{top:developing,p:0.56}}consensus:SPLITescalate:truemdl:{recorded:none}next:mm3 template drill --parent MM3-0004 --from accessnotes:[costestimatedfromtokens(nolivepricingreported),"budget: $0.10 left of $0.10 · 26 of 30 runs left"]
4 quick read MM3-0004bar = p
gatefailconsensusSPLITescalatetrue12 questions ·1 call ·293 ms
goalunsure0.38
designunsure10.9720.5030.96
accessfail40.3850.0860.10
correctnesspass70.8080.7290.80
aui-foundationunsure10yes0.45
design-maturityunsure11developing0.56
5the decision it infers
Drill into access
mm3 drill digs into one weak spot, one level down, asking only about it
the gate is FAIL: goal unsure (p 0.38); access fails; design, aui-foundation and design-maturity unsure
consensus is SPLIT and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 4 of 6 in this ledger · a root run, no parent · built on later by MM3-0005
reuse
asked fresh: 12 questions in 1 call, ~$0.00012; every answer is kept for reuse
recorded
nothing from an mdl: block; the request had none
budget
$0.10 left of $0.10 · 26 of 30 runs left
real runjev-1.13.0 · api.typesafe.ai · 293 ms · ~$0.00012 · 12 questions · 1 call · MM3-00042026-09-29WordPress @3ffb1dfcosts are estimates · counts include the goal question
Exact request and response text
mak:
goal: Abilities API is the integration point for agent-driven UI components
depth: quick
where:
- src/wp-includes/abilities-api.php:1-150
- src/wp-includes/abilities.php:1-100
ask:
concerns:
design:
pass: "yes"
1: Does Abilities API register capabilities with input/output schemas?
2: Are abilities required to define permission callbacks before execution?
3: Can third-party code register abilities without modifying WordPress core?
access:
pass: "yes"
4: Is the permission_callback required for every registered ability?
5: Are ability schemas validated against the permission callback result?
6: Can an unprivileged user bypass ability permission checks via API?
correctness:
pass: "yes"
7: Are input parameters validated against the input_schema?
8: Are output values validated against the output_schema?
9: Does the system reject invalid inputs before executing the callback?
decisions:
aui-foundation:
pass: [yes]
10:
choice: Is Abilities API suitable as the foundation for AUI?
options: [yes, partial, no]
design-maturity:
pass: [mature]
11:
scale: How mature is the Abilities API design for agents?
levels: [prototype, developing, mature, stable]
4goal tested: Abilities API enforces permission checks before executionpass 0.82
1 task
2 request
3 response
4 quick read
5 decision
6 ledger
drill · depth quick · 33-line request
2 the request it firedrequest.yaml · hover a question
mak:goal:Abilities API enforces permission checks before executiondepth:quickparent:MM3-0004from:accessask:concerns:access:pass:"no"1:Is permission_callback mandatory for every ability registration?2:Does the API execute permission_callback before running execute_callback?3:Are ability schemas enforced before permission callbacks are evaluated?injection:pass:"no"4:Can malicious input bypass the execute_callback entirely?5:Does execute_callback receive pre-validated inputs from input_schema?6:Are output values validated AFTER execute_callback runs?input:pass:"no"7:Can an ability execute without any input schema validation?8:Are input schema violations logged or rejected at the API level?9:Does the system provide clear error messages for invalid inputs?decisions:enforcement-model:pass:[mandatory]10:choice:How is permission enforcement in Abilities implemented?options:[mandatory,optional,advisory]security-gates:pass:[both]11:scale:At what execution stages are security checks applied?levels:[none,schema-only,permissions-only,both,comprehensive]
3 the response that came backMM3 output, verbatim
mak:id:MM3-0005gate:failgoal:{gate:pass,p:0.82}access:{gate:fail,1:0.37,2:0.84,3:0.24}injection:{gate:fail,4:0.24,5:0.60,6:0.70}input:{gate:fail,7:0.50,8:0.72,9:0.58}enforcement-model:{gate:pass,10:{top:mandatory,p:0.88}}security-gates:{gate:unsure,11:{top:both,p:0.58}}consensus:SPLITescalate:truemdl:{recorded:none}next:fix it, then mm3 replay --parent MM3-0004 --compare <before>..<after>notes:[costestimatedfromtokens(nolivepricingreported),"budget: $0.10 left of $0.10 · 25 of 30 runs left"]
4 quick read MM3-0005bar = p
gatefailconsensusSPLITescalatetrue12 questions ·1 call ·320 ms
goalpass0.82
accessfail10.3720.8430.24
injectionfail40.2450.6060.70
inputfail70.5080.7290.58
enforcement-modelpass10mandatory0.88
security-gatesunsure11both0.58
5the decision it infers
Fix it, then replay
mm3 replay re-asks the same questions across two commits, so the fix is proven, not assumed
the gate is FAIL: access, injection and input fail; security-gates unsure
consensus is SPLIT and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 5 of 6 in this ledger · child of MM3-0004, drilled from access
reuse
asked fresh: 12 questions in 1 call, ~$0.00012; every answer is kept for reuse
recorded
nothing from an mdl: block; the request had none
budget
$0.10 left of $0.10 · 25 of 30 runs left
real runjev-1.13.0 · api.typesafe.ai · 320 ms · ~$0.00012 · 12 questions · 1 call · MM3-00052026-09-29WordPress @3ffb1dfcosts are estimates · counts include the goal question
Exact request and response text
mak:
goal: Abilities API enforces permission checks before execution
depth: quick
parent: MM3-0004
from: access
ask:
concerns:
access:
pass: "no"
1: Is permission_callback mandatory for every ability registration?
2: Does the API execute permission_callback before running execute_callback?
3: Are ability schemas enforced before permission callbacks are evaluated?
injection:
pass: "no"
4: Can malicious input bypass the execute_callback entirely?
5: Does execute_callback receive pre-validated inputs from input_schema?
6: Are output values validated AFTER execute_callback runs?
input:
pass: "no"
7: Can an ability execute without any input schema validation?
8: Are input schema violations logged or rejected at the API level?
9: Does the system provide clear error messages for invalid inputs?
decisions:
enforcement-model:
pass: [mandatory]
10:
choice: How is permission enforcement in Abilities implemented?
options: [mandatory, optional, advisory]
security-gates:
pass: [both]
11:
scale: At what execution stages are security checks applied?
levels: [none, schema-only, permissions-only, both, comprehensive]
MDL³ · model · I've never worked in n8n and I want it faster
n8n-io/n8n at tags n8n@2.40.7 then n8n@2.41.3, unmodified public source
4 of the agent's 6 runs, in ledger order. Left out: MM3-0002 (replay), MM3-0005 (class). MM3-0002 is a replay that came back unsure because every item was skipped, and MM3-0005 repeats the reuse shown in the last step, so the story goes scan, class, drill, then the next release re-checked from the ledger.
1
task given to a Haiku agent
Use MM3 to answer this: I've never worked in n8n and I want to make it faster. How is it built — give me C4 context, container and component views — and where would I enhance it for speed? Then, a newer release is out: what changed in the architecture between the two releases?
the full kickoff, paths shortened
Use MM3 to answer this: I've never worked in n8n and I want to make it faster. How is it built — give me C4 context, container and component views — and where would I enhance it for speed? Then, a newer release is out: what changed in the architecture between the two releases?
The code is <checkout>, checked out at git tag `n8n@2.40.7` (unmodified; don't edit it). The newer release is the git tag `n8n@2.41.3`, already fetched in the same repo. Run MM3 from inside that folder as `mm3` — run `mm3 agent` first. MM3's spend is capped for this project.
Leave these in notes/:
- `requests/`: every request file you sent, numbered in order.
- `REPORT.md`: the three C4 views (as text or mermaid diagrams), the speed hot spots ranked, and the drift between the two releases — with the MM3 run id(s) behind every claim.
- `STOPS.md`: every ✖ message MM3 gave you, quoted exactly, one per line, each with what you changed next. Also list any workaround you used instead of MM3.
Work through it on your own; don't ask for help mid-way.
1goal tested: Identify major architectural components and speed bottlenecks in n8nunsure 0.49
1 task
2 request
3 response
4 quick read
5 decision
6 ledger
scan · depth standard · 59-line request
2 the request it firedrequest.yaml · hover a question
mak:goal:Identify major architectural components and speed bottlenecks in n8ndepth:standardover:file:packages/*/src/index.tsask:file:concerns:architecture:family:designpass:yes1:Does {file} define a clear component responsibility?2:Are dependency imports explicit and minimal?3:Is {file} a public API surface for its package?performance:family:design-riskpass:yes4:Does {file} initialize heavy state at load time?5:Could {file} be a performance bottleneck?6:Are there obvious inefficiencies in {file}'s patterns?integration:family:designpass:yes7:Are {file} exports properly scoped?8:Can consumers of {file} test against it easily?9:Is {file} version-stable for downstream packages?speed-risks:family:design-riskpass:yes10:Does {file} do synchronous I/O on the hot path?11:Are database queries batched or individually executed?12:Does {file} cache computed results across calls?scalability:family:designpass:yes13:Can {file} scale horizontally without coordination?14:Does {file} maintain per-worker state that could cause issues?15:Are resource limits enforced in {file}'s operations?testing:family:designpass:yes16:Can {file} be tested without external dependencies?17:Are performance characteristics measurable?18:Is {file} mocked easily for upstream tests?decisions:speed-concern:pass:[none]19:scale:How likely is {file} to be a speed issue?levels:[none,possible,likely,definite]action:pass:[monitor]20:choice:What speed work should {file} get?options:[skip,monitor,profile,optimize]mdl:why:validatearea:apiblast:container
3 the response that came backMM3 output, verbatim
mak:id:MM3-0001gate:failgoal:{gate:unsure,p:0.49}scanned:{file:4}failing:packages/cli/src/index.ts:{architecture:fail,performance:fail,integration:fail,speed-risks:fail,scalability:fail,testing:unsure,action:fail,1:0.10,3:0.17,4:0.06,5:0.12,6:0.10,7:0.41,8:0.27,9:0.17,10:0.08,11:0.31,12:0.05,13:0.13,14:0.07,15:0.07,16:0.38,17:0.50,18:0.60,20:{top:skip,p:0.96}}packages/core/src/index.ts:{architecture:unsure,performance:fail,integration:fail,speed-risks:fail,scalability:fail,testing:unsure,speed-concern:unsure,action:fail,1:0.44,2:0.47,4:0.12,5:0.32,6:0.33,7:0.54,8:0.53,9:0.30,10:0.12,11:0.36,12:0.10,13:0.22,14:0.15,15:0.15,16:0.47,18:0.33,19:{top:none,p:0.58},20:{top:skip,p:0.66}}packages/workflow/src/index.ts:{architecture:unsure,performance:fail,integration:unsure,speed-risks:fail,scalability:fail,testing:fail,speed-concern:unsure,action:fail,1:0.38,2:0.44,4:0.13,5:0.41,6:0.48,7:0.43,8:0.54,9:0.34,10:0.10,11:0.30,12:0.11,13:0.29,14:0.18,15:0.13,16:0.56,17:0.69,18:0.30,19:{top:none,p:0.45},20:{top:skip,p:0.57}}packages/node-dev/src/index.ts:{architecture:unsure,performance:fail,integration:fail,speed-risks:fail,scalability:fail,testing:unsure,action:fail,1:0.53,4:0.05,5:0.10,6:0.13,7:0.69,8:0.51,9:0.21,10:0.06,11:0.23,12:0.05,13:0.21,14:0.06,15:0.07,16:0.57,17:0.47,18:0.53,20:{top:skip,p:0.98}}passing:0reused:0mdl:{recorded:[why,area,blast]}next:mm3 template drill --parent MM3-0001 --from packages/cli/src/index.tsnotes:[costestimatedfromtokens(nolivepricingreported),"1 call · 80 questions · budget: $0.10 left of $0.10 · 29 of 30 runs left"]
4 quick read MM3-0001bar = p
gatefail81 questions ·1 call ·387 ms
goalunsure0.49
ranked, worst first
1packages/cli/src/index.ts
2packages/core/src/index.ts
3packages/workflow/src/index.ts
4packages/node-dev/src/index.ts
scannedfile 4passing0reused0
5the decision it infers
Drill into packages/cli/src/index.ts
mm3 drill digs into one weak spot, one level down, asking only about it
4 of 4 scanned files fail the gate; the worst is packages/cli/src/index.ts, failing 6 of 7 concerns
6what the ledger now holds
lineage
run 1 of 6 in this ledger · a root run, no parent · built on later by MM3-0002
reuse
asked fresh: 81 questions in 1 call, ~$0.00022; every answer is kept for reuse
recorded
why, area, blast saved with the run
budget
$0.10 left of $0.10 · 29 of 30 runs left
real runjev-1.13.0 · api.typesafe.ai · 387 ms · ~$0.00022 · 81 questions · 1 call · MM3-00012026-09-29n8n@2.40.7costs are estimates · counts include the goal question
Exact request and response text
mak:
goal: Identify major architectural components and speed bottlenecks in n8n
depth: standard
over:
file: packages/*/src/index.ts
ask:
file:
concerns:
architecture:
family: design
pass: yes
1: Does {file} define a clear component responsibility?
2: Are dependency imports explicit and minimal?
3: Is {file} a public API surface for its package?
performance:
family: design-risk
pass: yes
4: Does {file} initialize heavy state at load time?
5: Could {file} be a performance bottleneck?
6: Are there obvious inefficiencies in {file}'s patterns?
integration:
family: design
pass: yes
7: Are {file} exports properly scoped?
8: Can consumers of {file} test against it easily?
9: Is {file} version-stable for downstream packages?
speed-risks:
family: design-risk
pass: yes
10: Does {file} do synchronous I/O on the hot path?
11: Are database queries batched or individually executed?
12: Does {file} cache computed results across calls?
scalability:
family: design
pass: yes
13: Can {file} scale horizontally without coordination?
14: Does {file} maintain per-worker state that could cause issues?
15: Are resource limits enforced in {file}'s operations?
testing:
family: design
pass: yes
16: Can {file} be tested without external dependencies?
17: Are performance characteristics measurable?
18: Is {file} mocked easily for upstream tests?
decisions:
speed-concern:
pass: [none]
19:
scale: How likely is {file} to be a speed issue?
levels: [none, possible, likely, definite]
action:
pass: [monitor]
20:
choice: What speed work should {file} get?
options: [skip, monitor, profile, optimize]
mdl:
why: validate
area: api
blast: container
2 the request it firedrequest.yaml · hover a question
mak:goal:n8n architecture supports identifying and addressing speed bottlenecksverb:classdepth:standardwhere:[packages/core/src/index.ts,packages/cli/src/index.ts]ask:concerns:architecture:family:designpass:yes1:Are major architectural components (execution, queue, DB layer) clearly separated?2:Do components have well-defined responsibility boundaries?3:Can you identify clear data flow paths between components?speed-paths:family:design-riskpass:yes4:Is the execution engine isolated as a profiling/optimization target?5:Are hot paths (workflow execution, data queries) identifiable and measurable?6:Can the bottlenecks be traced through data flow and timing?scalability:family:designpass:yes7:Does the architecture support horizontal scaling (workers)?8:Are resource limits enforceable per workflow or execution?9:Is state sufficiently distributed to avoid centralizing bottlenecks?optimization:family:design-riskpass:yes10:Can the execution engine be optimized without touching other components?11:Can I/O operations be batched or parallelized independently?12:Are caching points clearly visible for performance improvement?integration:family:designpass:yes13:Do async/await patterns dominate the hot paths?14:Are database queries optimized (indexed, batched)?15:Is connection pooling used for external resources?observability:family:designpass:yes16:Can performance be measured component-by-component?17:Are critical latency points logged/instrumented?18:Can bottlenecks be reproduced in a test environment?decisions:bottleneck:pass:[execution-engine]19:scale:Which layer is most likely the speed bottleneck?levels:[execution-engine,queue-processing,database-layer,api-gateway,network-io,other]next-move:pass:[profile]20:choice:What should be the first speed optimization step?options:[profile,cache,parallelize,batch-queries,optimize-db,refactor-execution]mdl:why:validatearea:apiproblem:Understand n8n's architectural support for speed optimizationtouches:[execution-engine,queue,database,api,frontend]blast:system
3 the response that came backMM3 output, verbatim
mak:id:MM3-0003gate:failgoal:{gate:pass,p:0.70}architecture:{gate:fail,1:0.44,2:0.57,3:0.25}speed-paths:{gate:unsure,4:0.44,5:0.62,6:0.67}scalability:{gate:unsure,7:0.62,8:0.54,9:0.40}optimization:{gate:fail,10:0.49,11:0.66,12:0.16}integration:{gate:unsure,13:0.48,14:0.32,15:0.36}observability:{gate:unsure,16:0.72,17:0.50,18:0.71}bottleneck:{gate:pass,19:{top:execution-engine,p:0.77}}next-move:{gate:pass,20:{top:profile,p:1}}consensus:WEAKescalate:truemdl:{recorded:[why,area,problem,touches,blast]}next:mm3 template drill --parent MM3-0003 --from architecturenotes:[costestimatedfromtokens(nolivepricingreported),"budget: $0.10 left of $0.10 · 27 of 30 runs left"]
4 quick read MM3-0003bar = p
gatefailconsensusWEAKescalatetrue21 questions ·1 call ·305 ms
goalpass0.70
architecturefail10.4420.5730.25
speed-pathsunsure40.4450.6260.67
scalabilityunsure70.6280.5490.40
optimizationfail100.49110.66120.16
integrationunsure130.48140.32150.36
observabilityunsure160.72170.50180.71
bottleneckpass19execution-engine0.77
next-movepass20profile1
5the decision it infers
Drill into architecture
mm3 drill digs into one weak spot, one level down, asking only about it
the gate is FAIL: architecture and optimization fail; speed-paths, scalability, integration and 1 more unsure
consensus is WEAK and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 3 of 6 in this ledger · a root run, no parent · built on later by MM3-0004
reuse
asked fresh: 21 questions in 1 call, ~$0.000044; every answer is kept for reuse
recorded
why, area, problem, touches, blast saved with the run
budget
$0.10 left of $0.10 · 27 of 30 runs left
real runjev-1.13.0 · api.typesafe.ai · 305 ms · ~$0.000044 · 21 questions · 1 call · MM3-00032026-09-29n8n@2.40.7costs are estimates · counts include the goal question
Exact request and response text
mak:
goal: n8n architecture supports identifying and addressing speed bottlenecks
verb: class
depth: standard
where: [packages/core/src/index.ts, packages/cli/src/index.ts]
ask:
concerns:
architecture:
family: design
pass: yes
1: Are major architectural components (execution, queue, DB layer) clearly separated?
2: Do components have well-defined responsibility boundaries?
3: Can you identify clear data flow paths between components?
speed-paths:
family: design-risk
pass: yes
4: Is the execution engine isolated as a profiling/optimization target?
5: Are hot paths (workflow execution, data queries) identifiable and measurable?
6: Can the bottlenecks be traced through data flow and timing?
scalability:
family: design
pass: yes
7: Does the architecture support horizontal scaling (workers)?
8: Are resource limits enforceable per workflow or execution?
9: Is state sufficiently distributed to avoid centralizing bottlenecks?
optimization:
family: design-risk
pass: yes
10: Can the execution engine be optimized without touching other components?
11: Can I/O operations be batched or parallelized independently?
12: Are caching points clearly visible for performance improvement?
integration:
family: design
pass: yes
13: Do async/await patterns dominate the hot paths?
14: Are database queries optimized (indexed, batched)?
15: Is connection pooling used for external resources?
observability:
family: design
pass: yes
16: Can performance be measured component-by-component?
17: Are critical latency points logged/instrumented?
18: Can bottlenecks be reproduced in a test environment?
decisions:
bottleneck:
pass: [execution-engine]
19:
scale: Which layer is most likely the speed bottleneck?
levels: [execution-engine, queue-processing, database-layer, api-gateway, network-io, other]
next-move:
pass: [profile]
20:
choice: What should be the first speed optimization step?
options: [profile, cache, parallelize, batch-queries, optimize-db, refactor-execution]
mdl:
why: validate
area: api
problem: Understand n8n's architectural support for speed optimization
touches: [execution-engine, queue, database, api, frontend]
blast: system
3goal tested: Understand why n8n architecture components are not clearly separatedunsure 0.50
1 task
2 request
3 response
4 quick read
5 decision
6 ledger
drill · depth standard · 60-line request
2 the request it firedrequest.yaml · hover a question
mak:goal:Understand why n8n architecture components are not clearly separatedverb:drillparent:MM3-0003from:architecturedepth:standardask:concerns:boundaries:family:designpass:yes1:Does the code define clear module boundaries (exports/imports)?2:Are circular dependencies avoided between components?3:Is the API surface between components documented?execution-layer:family:design-riskpass:yes4:Is execution engine logic isolated in one module?5:Can execution be called independently of other systems?6:Does execution manage its own resources (memory, timeout)?data-flow:family:designpass:yes7:Are data transformations on the hot path minimized?8:Is data passed by reference or by copy?9:Are batch operations preferred over individual calls?performance-surface:family:design-riskpass:yes10:Can you identify where execution time is spent?11:Are profiling hooks or instrumentation available?12:Can performance regressions be detected?dependencies:family:designpass:yes13:Are external dependencies (DB, queues, APIs) abstracted?14:Can components function with mock implementations?15:Is dependency injection used for flexibility?containerization:family:designpass:yes16:Can components be deployed separately?17:Are resource constraints (CPU, memory) enforced per component?18:Do components have independent lifecycle management?decisions:separation-level:pass:[moderate]19:scale:How well are components actually separated?levels:[monolithic,loose,moderate,good,excellent]architecture-debt:pass:[acceptable]20:choice:What's blocking better architecture?options:[refactor-needed,acceptable,design-unclear,tooling-missing]mdl:why:debugarea:apiproblem:Understand architectural separation challenges in n8nblast:container
3 the response that came backMM3 output, verbatim
mak:id:MM3-0004gate:failgoal:{gate:unsure,p:0.50}boundaries:{gate:fail,1:0.41,2:0.54,3:0.16}execution-layer:{gate:unsure,4:0.40,5:0.41,6:0.45}data-flow:{gate:unsure,7:0.43,8:0.40,9:0.39}performance-surface:{gate:fail,10:0.22,11:0.58,12:0.55}dependencies:{gate:unsure,13:0.52,14:0.66,15:0.46}containerization:{gate:fail,16:0.35,17:0.14,18:0.24}separation-level:{gate:fail,19:{top:monolithic,p:0.45}}architecture-debt:{gate:fail,20:{top:design-unclear,p:0.65}}consensus:WEAKescalate:truemdl:{recorded:[why,area,problem,blast]}next:fix it, then mm3 replay --parent MM3-0003 --compare <before>..<after>notes:[costestimatedfromtokens(nolivepricingreported),"budget: $0.10 left of $0.10 · 26 of 30 runs left"]
4 quick read MM3-0004bar = p
gatefailconsensusWEAKescalatetrue21 questions ·1 call ·310 ms
goalunsure0.50
boundariesfail10.4120.5430.16
execution-layerunsure40.4050.4160.45
data-flowunsure70.4380.4090.39
performance-surfacefail100.22110.58120.55
dependenciesunsure130.52140.66150.46
containerizationfail160.35170.14180.24
separation-levelfail19monolithic0.45
architecture-debtfail20design-unclear0.65
5the decision it infers
Fix it, then replay
mm3 replay re-asks the same questions across two commits, so the fix is proven, not assumed
the gate is FAIL: goal unsure (p 0.50); boundaries, performance-surface, containerization and 2 more fail; execution-layer, data-flow and dependencies unsure
consensus is WEAK and escalate is true, so do not act on this alone
6what the ledger now holds
lineage
run 4 of 6 in this ledger · child of MM3-0003, drilled from architecture
reuse
asked fresh: 21 questions in 1 call, ~$0.000041; every answer is kept for reuse
recorded
why, area, problem, blast saved with the run
budget
$0.10 left of $0.10 · 26 of 30 runs left
real runjev-1.13.0 · api.typesafe.ai · 310 ms · ~$0.000041 · 21 questions · 1 call · MM3-00042026-09-29n8n@2.40.7costs are estimates · counts include the goal question
Exact request and response text
mak:
goal: Understand why n8n architecture components are not clearly separated
verb: drill
parent: MM3-0003
from: architecture
depth: standard
ask:
concerns:
boundaries:
family: design
pass: yes
1: Does the code define clear module boundaries (exports/imports)?
2: Are circular dependencies avoided between components?
3: Is the API surface between components documented?
execution-layer:
family: design-risk
pass: yes
4: Is execution engine logic isolated in one module?
5: Can execution be called independently of other systems?
6: Does execution manage its own resources (memory, timeout)?
data-flow:
family: design
pass: yes
7: Are data transformations on the hot path minimized?
8: Is data passed by reference or by copy?
9: Are batch operations preferred over individual calls?
performance-surface:
family: design-risk
pass: yes
10: Can you identify where execution time is spent?
11: Are profiling hooks or instrumentation available?
12: Can performance regressions be detected?
dependencies:
family: design
pass: yes
13: Are external dependencies (DB, queues, APIs) abstracted?
14: Can components function with mock implementations?
15: Is dependency injection used for flexibility?
containerization:
family: design
pass: yes
16: Can components be deployed separately?
17: Are resource constraints (CPU, memory) enforced per component?
18: Do components have independent lifecycle management?
decisions:
separation-level:
pass: [moderate]
19:
scale: How well are components actually separated?
levels: [monolithic, loose, moderate, good, excellent]
architecture-debt:
pass: [acceptable]
20:
choice: What's blocking better architecture?
options: [refactor-needed, acceptable, design-unclear, tooling-missing]
mdl:
why: debug
area: api
problem: Understand architectural separation challenges in n8n
blast: container
4goal tested: Identify major architectural components and speed bottlenecks in n8nunsure 0.49
1 task
2 request
3 response
4 quick read
5 decision
6 ledger
scan · depth standard · 59-line request
2 the request it firedrequest.yaml · hover a question
mak:goal:Identify major architectural components and speed bottlenecks in n8ndepth:standardover:file:packages/*/src/index.tsask:file:concerns:architecture:family:designpass:yes1:Does {file} define a clear component responsibility?2:Are dependency imports explicit and minimal?3:Is {file} a public API surface for its package?performance:family:design-riskpass:yes4:Does {file} initialize heavy state at load time?5:Could {file} be a performance bottleneck?6:Are there obvious inefficiencies in {file}'s patterns?integration:family:designpass:yes7:Are {file} exports properly scoped?8:Can consumers of {file} test against it easily?9:Is {file} version-stable for downstream packages?speed-risks:family:design-riskpass:yes10:Does {file} do synchronous I/O on the hot path?11:Are database queries batched or individually executed?12:Does {file} cache computed results across calls?scalability:family:designpass:yes13:Can {file} scale horizontally without coordination?14:Does {file} maintain per-worker state that could cause issues?15:Are resource limits enforced in {file}'s operations?testing:family:designpass:yes16:Can {file} be tested without external dependencies?17:Are performance characteristics measurable?18:Is {file} mocked easily for upstream tests?decisions:speed-concern:pass:[none]19:scale:How likely is {file} to be a speed issue?levels:[none,possible,likely,definite]action:pass:[monitor]20:choice:What speed work should {file} get?options:[skip,monitor,profile,optimize]mdl:why:validatearea:apiblast:container
3 the response that came backMM3 output, verbatim
mak:id:MM3-0006gate:failgoal:{gate:unsure,p:0.49}scanned:{file:4}failing:packages/cli/src/index.ts:{architecture:fail,performance:fail,integration:fail,speed-risks:fail,scalability:fail,testing:unsure,action:fail,1:0.10,3:0.17,4:0.06,5:0.12,6:0.10,7:0.41,8:0.27,9:0.17,10:0.08,11:0.31,12:0.05,13:0.13,14:0.07,15:0.07,16:0.38,17:0.50,18:0.60,20:{top:skip,p:0.96}}packages/core/src/index.ts:{architecture:unsure,performance:fail,integration:fail,speed-risks:fail,scalability:fail,testing:unsure,speed-concern:unsure,action:fail,1:0.44,2:0.47,4:0.12,5:0.32,6:0.33,7:0.54,8:0.53,9:0.30,10:0.12,11:0.36,12:0.10,13:0.22,14:0.15,15:0.15,16:0.47,18:0.33,19:{top:none,p:0.58},20:{top:skip,p:0.66}}packages/workflow/src/index.ts:{architecture:unsure,performance:fail,integration:unsure,speed-risks:fail,scalability:fail,testing:fail,speed-concern:unsure,action:fail,1:0.38,2:0.44,4:0.13,5:0.41,6:0.48,7:0.43,8:0.54,9:0.34,10:0.10,11:0.30,12:0.11,13:0.29,14:0.18,15:0.13,16:0.56,17:0.69,18:0.30,19:{top:none,p:0.45},20:{top:skip,p:0.57}}packages/node-dev/src/index.ts:{architecture:unsure,performance:fail,integration:fail,speed-risks:fail,scalability:fail,testing:unsure,action:fail,1:0.53,4:0.05,5:0.10,6:0.13,7:0.69,8:0.51,9:0.21,10:0.06,11:0.23,12:0.05,13:0.21,14:0.06,15:0.07,16:0.57,17:0.47,18:0.53,20:{top:skip,p:0.98}}passing:0reused:4mdl:{recorded:[why,area,blast]}next:mm3 template drill --parent MM3-0006 --from packages/cli/src/index.tsnotes:["reused: MM3-0001 (0d, 1 commit)","0 calls · 0 questions · budget: $0.10 left of $0.10 · 26 of 30 runs left"]
4 quick read MM3-0006bar = p
gatefail0 questions +81 reused ·0 calls ·0 ms
goalunsure0.49
ranked, worst first
1packages/cli/src/index.ts
2packages/core/src/index.ts
3packages/workflow/src/index.ts
4packages/node-dev/src/index.ts
scannedfile 4passing0reused4
5the decision it infers
Drill into packages/cli/src/index.ts
mm3 drill digs into one weak spot, one level down, asking only about it
4 of 4 scanned files fail the gate; the worst is packages/cli/src/index.ts, failing 6 of 7 concerns
answers were reused (4 files): the code they were given on is unchanged, so nothing new was asked
6what the ledger now holds
lineage
run 6 of 6 in this ledger · a root run, no parent
reuse
81 answers reused from MM3-0001: no call, $0.00000, saved ~$0.00022
recorded
why, area, blast saved with the run
budget
$0.10 left of $0.10 · 26 of 30 runs left
real runjev-1.13.0 · api.typesafe.ai · 0 ms · $0.00000 · 0 asked · 81 reused from MM3-0001 · 0 calls · MM3-00062026-09-29n8n@2.41.3costs are estimates · counts include the goal question
Exact request and response text
mak:
goal: Identify major architectural components and speed bottlenecks in n8n
depth: standard
over:
file: packages/*/src/index.ts
ask:
file:
concerns:
architecture:
family: design
pass: yes
1: Does {file} define a clear component responsibility?
2: Are dependency imports explicit and minimal?
3: Is {file} a public API surface for its package?
performance:
family: design-risk
pass: yes
4: Does {file} initialize heavy state at load time?
5: Could {file} be a performance bottleneck?
6: Are there obvious inefficiencies in {file}'s patterns?
integration:
family: design
pass: yes
7: Are {file} exports properly scoped?
8: Can consumers of {file} test against it easily?
9: Is {file} version-stable for downstream packages?
speed-risks:
family: design-risk
pass: yes
10: Does {file} do synchronous I/O on the hot path?
11: Are database queries batched or individually executed?
12: Does {file} cache computed results across calls?
scalability:
family: design
pass: yes
13: Can {file} scale horizontally without coordination?
14: Does {file} maintain per-worker state that could cause issues?
15: Are resource limits enforced in {file}'s operations?
testing:
family: design
pass: yes
16: Can {file} be tested without external dependencies?
17: Are performance characteristics measurable?
18: Is {file} mocked easily for upstream tests?
decisions:
speed-concern:
pass: [none]
19:
scale: How likely is {file} to be a speed issue?
levels: [none, possible, likely, definite]
action:
pass: [monitor]
20:
choice: What speed work should {file} get?
options: [skip, monitor, profile, optimize]
mdl:
why: validate
area: api
blast: container
Two real stories, each driven by a Haiku agent on unmodified public source: MAK³ · make, “Where do agents plug into WordPress?” (WordPress @ 3ffb1df), and MDL³ · model, “I’ve never worked in n8n and I want it faster” (n8n@2.40.7, then n8n@2.41.3). Every step is one run: the task the agent was given, the request it fired, the response MM3 returned, a quick read of it, the decision it implies and what the ledger now holds. Every footer comes from that run’s own ledger row.
The challenge: you host n8n yourself and want it to run faster. You ask your agent where to start. It narrows the question to one place, the node loader's cleanup in directory-loader.ts, and asks MM3 twelve questions about it in one call: nine yes/no, two decisions and the goal itself.
The request the agent wrote (45 lines):
The response, real output · jev-1.13.0 · api.typesafe.ai · 321 ms · ~$0.000063:
Every concern is written so that "no" is healthy: a pass: no answer clears the bar at 0.30 or below.
Goal: pass (0.79). Yes, this cleanup could be faster.
Availability: fail. At 0.88, the realpathSync() calls on line 606 block the event loop.
Design: fail on all three: a rescan on every unloadAll() call (0.78), a try-catch that silently skips errors (0.90), and caching would help (0.84).
Design-risk: unsure. None of 0.40, 0.46, 0.64 lands clearly either way.
Decisions: measure first passes (0.89); severity is unsure (low, 0.56).
Next: the concerns disagree, so consensus is SPLIT, escalate is true, and next: points at a drill into availability. The mdl: line lists what the ledger recorded about why the agent asked, so later runs on this code start from it.
04 What you get
What you get
Know
Judge
Prove
MAK³ use what is proven
view a free lookup of what is on record
class one subject, one verdict
replay recheck after a fix
MDL³ learn what is missing
scan sweep to find where to look
drill dig into one weak spot
loop vet a design before code
A request has a mak: block (the checklist) and an optional mdl: block (why you are asking, so the ledger learns). The verb you run, not the key, decides whether it is a MAK³ or an MDL³ move. mm3 help <verb> shows the rules for each; mm3 template <verb> prints a filled-in sample.
Every run and its outcome goes into an append-only ledger in .mm3/ (git-ignored). Ask the same questions of unchanged code and MM3 answers from the ledger: no call, no cost. mm3 report reads back where your agents keep going wrong, and a budget cap stops runaway spend; by convention only you raise or reset it, and MM3 tells agents to ask you. mm3 view looks a request up in the ledger before you spend anything.
05 Why we built it
Why we built it
Agents can now ask a fast classifier a yes/no about your code, but each answer is untraceable and never reused. MM3 turns that into one standard checklist in, one calibrated verdict per concern out, and every answer kept and reused.
06 What you can do
What you can do
Check a change before you merge it
One class call, three angles per concern, a verdict for each.
Start with mm3 template class
Prove a fix actually worked
Re-asks a past run's own questions across two commits, so you check the fix without re-checking everything.
Start with mm3 template replay
Find where a problem lives
Sweeps a folder and ranks the files that most need a look.
Start with mm3 template scan
Check a design before any code exists
Puts a plan through the same checklist before anyone writes it.
Start with mm3 template loop
Each template is a filled-in request with its rules as comments.
Ask once. Keep the answer.
Install the plugin, then just ask your agent. Same questions on unchanged code come back from the ledger: no call, no cost.
Advice, not action: it gives the odds, you make the call. Delete, deploy, drop and pay stay human.
A pass is a probability: a calibrated 0.9 is wrong one time in ten.unsure is a real answer, and mm3 outcome shows which verdicts held.
Run linters, scanners and tests first: they’re free and exact. MM3 takes the questions they can’t ask, like “does this handler check the caller?”
First principles still apply: map your architecture first. A few cheap checks name the layers (tag requests with mdl.uses) and mm3 report graph draws them. Every later question then lands in a known place, and reading it back is free.
Use a full review for open questions. A model that reads the whole codebase answers anything. MM3 answers yes/no.
Beta: works well in our own use and on an intentionally vulnerable app (OWASP NodeGoat). Formal benchmarks are coming.