What was measured
One German spring village scene, one hero view per arm, judged in a
single sealed sitting. Each arm received n = 8 scores
— four per judge family (Claude Opus and Codex gpt-reserve)
— of the same final fixed renders. These are repeated
evaluations of final artifacts, not eight independent builds, and
not eight independent scenes. Visual quality only; this is not a
runtime-correctness or general-ability measure.
All ten final arms
Final sealed sitting, 16 September 2026. Mean and sd over n = 8 scores per arm (1–10 scale). Gate is the deterministic artifact gate on the final attempt.
| Arm |
Final attempt |
Gate |
Mean |
sd |
| Site render, GPT-6 Astra orchestrating the SGT-GPT6 fleet on Flash 4.1 (reference) | — | — | 7.78 | 0.91 |
| Fable 5.1 orchestrating SGT-Fleet on Flash 4.1, COMPONENT protocol (house, site, driver authored separately) | 5 | PASS | 7.66 | 0.51 |
| Fable 5.1 with GPT-5.6 Luna alone, no fleet | 4 | PASS | 7.62 | 0.88 |
| Fable 5.1 with Flash 4.1 alone, no fleet | 8 | PASS | 7.45 | 0.61 |
| Fable 5.1 orchestrating the SGT-GPT6 fleet on Flash 4.1 | 8 | PASS | 7.20 | 0.97 |
| Fable 5.1 orchestrating SGT-Fleet on Flash 4.1 after the Blender training port, monolithic script | 8 | PASS | 7.14 | 0.71 |
| Fable 5.1 orchestrating SGT-Fleet on Sonnet 5 | 8 | PASS | 6.22 | 0.92 |
| Fable 5.1 orchestrating SGT-Fleet on Flash 4.1 before the port | 8 | PASS | 5.83 | 1.04 |
| Fable 5.1 with Sonnet 5 alone, no fleet | 8 | PASS | 4.72 | 1.05 |
| Fable 5.1 orchestrating SGT-Fleet on GPT-5.6 Luna | 8 | FAIL | 1.38 | 0.30 |
Method
Identical brief and delivery contract across arms, run in a common
Blender 5.2 headless environment, a deterministic artifact gate,
hands-on orchestration by Fable 5.1 with rounds until the blind
judges put the arm within 0.5 of the reference or attempt 8, and
blind judging with sealed label mappings by two judge families. The
orchestrator inspected every frame before any PASS. The reference
row is the live gallery render of the same scene, shown for context
and not as a competing product.
Fairness limits
- One scene and one hero view per arm; not a general model ranking.
- n = 8 scores per arm are four per judge family on the same final fixed renders, not eight independent builds.
- Selected final attempts only; earlier attempts are not scored here.
- Orchestration-assisted iteration: the orchestrator chose revisions between rounds, so this is not a measure of unaided model output.
- Provider-specific knowledge: the fleet's stores and skills were minted from Flash 4.1 failure modes, so results do not transfer as a model comparison.
- This is a comparison of configurations, not causal proof of training gains.
- No deterministic gate result was recorded for the 7.78 reference; it is the live gallery render, displayed for context.