← Work / Active files

[ CASE FILE / MIGHTY-MOUSE ]

Active

Mighty Mouse

A provider-agnostic reliability harness that gives AI coding agents structured protocols and project-native verification.

Role
Creator, systems engineer, and evaluation designer
Discipline
AI workflow reliability / developer tooling
Year
2026
Strongest proof
6/10 vs 6/10 / First-attempt passes
Mighty MouseA provider-agnostic reliability harness that gives AI coding agents structured protocols and project-native verification.6/10 vs 6/10 / First-attempt passes

Agent output is not proof that a coding task is complete.

Coding agents can produce plausible explanations and still miss required files, violate scope, or stop before the project’s own checks pass. The operational problem is turning a model response into a bounded, inspectable result another person can trust.

Mighty Mouse supplies versioned complexity protocols, project-native verification, structured results, and a bounded retry contract. The model and agent platform remain replaceable; the reliability layer stays explicit.

I built the reliability layer, not the models or agent platforms.

I designed and built the harness, versioned protocols, project verifier, result schema, CLI, MCP transport, packaging, platform rules, CI, evaluation fixtures, and prospective study.

Foundation models, model runtimes, MCP itself, and platforms such as Codex and Antigravity are third-party systems. Mighty Mouse integrates with them; it does not claim their underlying work.

Declared ownership

Original reliability harness, versioned protocols, project verifier, evaluation design, packaging, CLI, MCP transport, platform rules, CI, and study execution. Foundation models, model runtimes, MCP itself, and agent platforms are third-party systems.

How the work moves.

  1. Input

    Task + complexity input

    A coding task enters with an explicit complexity level or a deterministic default.

  2. Decision

    Protocol selection

    The harness returns a versioned low, medium, or high-complexity execution contract.

  3. State

    Agent execution

    The chosen agent works inside declared scope using the project’s existing tools and conventions.

  4. Decision

    Project-native verification

    Tests, lint, build, scope, and changed-file checks run as structured verification steps.

  5. Decision

    Bounded retry loop

    Failures return specific next steps; the agent may retry without entering an unbounded correction cycle.

  6. Output

    Structured result

    Pass state, checks, warnings, and suggestions are returned in a stable human- and machine-readable form.

  7. Failure path

    Stop condition

    A passing result completes the task; exhausted retries or unresolved scope failures stop and escalate.

Decisions that changed the system.

Selected implementation record.

01 / INTERFACE

Core CLI

Protocol and verify commands expose stable JSON interfaces alongside doctor, demo, and benchmark operations.

02 / TRANSPORT

MCP server

A separate package exposes protocol and verification tools over stdio for compatible agent platforms.

03 / DISTRIBUTE

Installable packages

Core and MCP wheels and source distributions are built, archive-inspected, checksummed, and smoke-tested outside the checkout.

04 / VERIFY

CI and paired evaluation

Python 3.10–3.13 CI, packaging gates, frozen trial manifests, and blind reviews keep implementation and claims auditable.

Claims with their edges left on.

HS-7VERIFIED RECORD
6/10 vs 6/10

First-attempt passes

Mighty Mouse and the control each passed six of ten paired real-project tasks on the first attempt, with zero scope violations in both conditions.

Sample
10 paired tasks producing 20 condition runs
Boundary
Ten paired tasks from the recorded projects, agent, models, and environment only; no generalized reliability or speed improvement was demonstrated.

Sourcedata/evidence/real_project_report.md — completed prospective study

LimitThe study did not demonstrate better generalized first-pass reliability.

HS-7VERIFIED RECORD
4 vs 6

Retry rounds

Mighty Mouse used four retry rounds across the sample; the control used six.

Sample
10 paired tasks producing 20 condition runs
Boundary
Ten paired tasks from the recorded projects, agent, models, and environment only; no generalized reliability or speed improvement was demonstrated.

Sourcedata/evidence/real_project_report.md — completed prospective study

LimitTwo fewer retries in this sample is not evidence of a universal improvement.

HS-7VERIFIED RECORD
4.60 vs 4.30

Mean blind-review quality

Mighty Mouse received a 4.60 mean blind-review score versus 4.30 for the control across the ten paired tasks.

Sample
10 paired tasks producing 20 condition runs
Boundary
Ten paired tasks from the recorded projects, agent, models, and environment only; no generalized reliability or speed improvement was demonstrated.

Sourcedata/evidence/real_project_report.md — completed prospective study

LimitThe quality difference belongs to this sample and its recorded blind-review procedure.

HS-7VERIFIED RECORD
262.5s vs 229.5s

Median duration

Mighty Mouse was slower: 262.5 seconds median duration versus 229.5 seconds for the control; raw mean duration was also slower.

Sample
10 paired tasks producing 20 condition runs
Boundary
Ten paired tasks from the recorded projects, agent, models, and environment only; no generalized reliability or speed improvement was demonstrated.

Sourcedata/evidence/real_project_report.md — completed prospective study

LimitThe study did not demonstrate a generalized speed improvement.

HS-7VERIFIED RECORD
29.5%

Historical synthetic latency result

Lean reduced average latency by 29.5% while both conditions passed 15 of 15 tasks in the historical synthetic promotion suite.

Sample
15 paired tasks producing 30 condition runs
Boundary
Historical synthetic promotion suite and recorded environment only; the later prospective real-project study did not reproduce a generalized speed or first-pass reliability improvement.

Sourcedata/evidence/results/PROMOTION_NOTES.md — historical promotion validation

LimitThis historical synthetic result is bounded to that suite and must not be generalized to real project work.

The system, without the sales pitch.

Proof artifactAn 80-second, captioned walkthrough uses real v0.2.1 CLI output and a disposable Python task to show the protocol-to-verification workflow without overstating the paired-study result.

What failed stays in the record.

resolved

The synthetic suite hit a ceiling

Observed
Mighty Mouse, the historical protocol conditions, and a bare control all passed the frozen synthetic tasks, limiting what the suite could prove about reliability.
Response
Bound the 29.5% latency result to its recorded environment and ran a prospective paired study on real project tasks.
monitoring

The real-project result was mixed

Observed
First-attempt passes tied at 6/10, and Mighty Mouse was slower by both mean and median duration despite fewer retries and higher mean blind-review quality.
Response
Published the unfavorable timing and tied reliability results and stated that no generalized improvement was demonstrated.
resolved

macOS metadata contaminated package inputs

Observed
AppleDouble resource-fork files created noisy Git and archive failures on the external volume.
Response
Added package-data exclusions, source-distribution filtering, archive inspection, and explicit regression tests.
resolved

Editable installs and Python 3.10 exposed packaging gaps

Observed
Custom build-backend behavior and older-Python compatibility failed outside the original development environment.
Response
Completed editable-build hooks, added compatibility dependencies, and required clean Python 3.10–3.13 CI plus artifact-only installation smokes.

Released reliability tooling with bounded evidence.

Version 0.2.1 ships the core library, CLI, MCP transport, platform rules, verifier, CI matrix, release artifacts, and the completed ten-pair study. The current evidence supports fewer retries and higher mean blind-review quality in this sample, not a generalized first-pass reliability or speed improvement.

Open to the right work

A useful system starts with the actual problem.

Full-time product, automation, and technical-operations roles are the priority. Select workflow and creative-technology projects are open.

© 2026 John Roberts / Johnny MaconnyYBF Studios