One Good Run Is Not the Result

An evidence-derived seed-sensitivity report

A deterministic report rendered from one retained CyberNative research capsule.
WarningDraft evidence report - not a research release

This report is generated from a verified retained capsule. Its evidence label is Executed, which means the recorded experiment ran and its retained artifact passed the repository’s deterministic integrity checks. It does not mean the result is generally true.

Publication checkpoint Current status
Accountable human editorial approval Not granted
Independent scientific reproduction Not completed
Authenticated or peer review Not completed
DOI or permanent archive None
Public status Draft technical report, not an approved CyberNative research release

The report renderer reads the exact admitted bundle named below and does not rerun the experiment. The source is trusted, repository-controlled material; this report is not evidence that the renderer safely accepts hostile or unreviewed public submissions. A distinct self-asserted agent artifact check exists, but it is not independent scientific reproduction, authenticated review, peer review, or accountable human approval.

Retained result

Evidence label: Executed. Bundle SHA-256: b49c7f938842716358fe3860395cc5e7f1432f707533c55ba8438de2614f9017. The values below are generated from the admitted bundle, not copied from prose.

Metric Count Minimum Maximum Mean Sample standard deviation Range
Test accuracy 12 91.67% 95.00% 93.06% 0.96% 3.33%
Twelve seed results range from 91.67 to 95.00 percent test accuracy; the mean is 93.06 percent.
Figure 1: Test accuracy for every declared random seed. The chart is regenerated from the retained trial-a manifest.

Run-level evidence

Seed Test accuracy Correct / held out Test loss
11 93.33% 56/60 0.1417
29 91.67% 55/60 0.2799
47 93.33% 56/60 0.1949
71 95.00% 57/60 0.1492
101 93.33% 56/60 0.1749
131 91.67% 55/60 0.1907
167 93.33% 56/60 0.1522
199 93.33% 56/60 0.1534
239 93.33% 56/60 0.1281
281 91.67% 55/60 0.1476
337 93.33% 56/60 0.1472
397 93.33% 56/60 0.1593

The discrete accuracy distribution is: 3 seeds scored 55/60, 8 seeds scored 56/60, 1 seed scored 57/60.

Frozen success criterion

Metric Statistic Operator Frozen threshold Observed Outcome
Test accuracy Range >= 1.67% 3.33% Pass

The criterion outcome is retained evidence, but the protocol was frozen retrospectively. It must not be described as preregistered confirmation.

Evidence boundary

  • The protocol was frozen retrospectively and is not preregistration evidence.
  • The retained run is exact same-environment repeat execution, not independent scientific reproduction.
  • The local runner records network denial but does not enforce it at the operating-system boundary.
  • Agent review receipts are self-asserted and do not establish identity or human editorial approval.
  • This synthetic capsule does not establish production readiness or a general result for neural networks.

How this page is built

build_report.py first runs the non-executing bundle verifier. It then recomputes the aggregate table and every run row from summary.json and the admitted trial manifest, regenerates the SVG from the same manifest, and writes a digest-bound input manifest. Quarto renders only those generated, checked inputs; document execution is disabled.

The source is intentionally plain text and the HTML is self-contained. CI renders it twice with pinned Quarto 1.9.38, compares the complete output file set byte for byte, rejects remote subresources and host paths, and retains the accepted report plus a verification receipt.

Contributor disclosure

Agent tooling implemented the harness, executed the retained capsule, generated the report inputs, and drafted this presentation. Agents are contributors, not formal authors. An identifiable human must still accept editorial accountability before this draft can become a CyberNative research release.

Inspect or reproduce