Independent research · Active causal mission assurance

Models may propose. Physical outcomes decide.

A system is not proven because it scores well. Its claimed objective has to survive independent intervention, physical execution, held-out transfer and settled recovery. Inbar is the research that tries to make that survivable claim, and it has not made it yet.

What the live surface shows

  • The admission gate

    Thresholds frozen on 2026-07-14, before any evidence existed: 30 complete physical incidents across 6 hardware identities and 3 fault families. The corpus stands at 0. Four of the six criteria have never been adjudicated and are shown as unmeasured rather than as zero.

    Preregistration 47a1920b1b532660… · iteration iter001_physical_causal_evidence_acquisition
  • The screened corpus

    16 NASA ADAPT incidents were screened. The source passed integrity, parser and truth-separation checks and still returned BLOCKED_EVIDENCE: it carried no independently reviewed ambiguity sets and no safe discriminating actions, so none qualified as a complete dossier. A substrate-screening exclusion kept at full evidentiary weight; not an Inbar outcome or scientific finding.

    2884 telemetry rows · 95 commands · 28 fault injections · 0 leakage markers
  • The seals

    Model-visible evidence and adjudication truth are separate artifacts with separate hashes, and the truth bundle is never passed into anything a model can read. Both hashes are published; the truth contents are not. That separation is the mission's first scientific invariant.

    Evidence 111f068f08d0e099… · truth bb9e80e3f72f6a2d…
  • The boundary

    Control authority is marked bootstrap and the control-suite hash is an unset placeholder of zeros; production verification rejects that state by design. Execution authority is limited to replay, simulator, testbed — flight, live spacecraft, live robots and destructive tests are forbidden, and learned systems never hold safety or execution authority.

    Source-role verdict KILL_PUBLIC_SUBSTRATE · 1504 tests · 91.22% coverage · 22 exact controls