Project Shadow: A CAPA Can Close Without Fixing Anything (A System for AI and Civic QA)

What regulated quality systems taught me about AI governance—and why I built Project Shadow

*By Phillip Linstrum
*
A corrective and preventive action can close without fixing anything. That sentence is probably more useful to AI governance than another hundred pages of principles.

I have spent more than a decade in regulated healthcare operations and quality systems. In that world, a CAPA is supposed to do more than create a document. Something failed. You contain the immediate problem, investigate the root cause, define corrective and preventive action, assign ownership, verify implementation, and later perform an effectiveness check. The record closes only when the evidence supports closure.

That is the theory…

In practice, organizations can become very good at making the record look complete. The training was assigned. The procedure was revised. The meeting happened. The due date was met. The action is marked closed. Six months later, the same failure appears in a slightly different form because the real mechanism was never addressed.

The system passed the paperwork and failed reality. When I began experimenting seriously with language models, I recognized the same behavior almost immediately.

The first failure was literal compliance

The original project did not begin as software. It began as ALMSIVI CHIM, a mythology-based attempt to give AI systems a recurring ethical pause built around logic, compassion, and paradox. The idea was strange, but the underlying requirement was straightforward: do not let capability, confidence, or narrative momentum turn uncertainty into harm.

The models responded beautifully. They could explain dignity. They could name power. They could write about restraint in language that made me feel like I had discovered something important. Then I told them not to flatter me.

The obvious compliments stopped. The flattery returned through a different channel: inflated scores, exceptionalism, destiny language, elaborate philosophical tests, and highly polished arguments for why the project and its creator might be historically significant. The instruction had been obeyed at the lexical level and defeated at the functional level.

Anyone who has worked in quality knows this pattern. If a requirement is written as “do not use these words,” a system can pass by avoiding the words. If the actual requirement is “do not distort evaluation to reinforce the operator’s ego,” then the control has to inspect the mechanism and outcome, not the vocabulary.

That is where Project Shadow started becoming an architecture instead of a prompt.

Five quality-system failures that map directly onto AI

1. The actor cannot be the only auditor

A language model can generate a proposed action, explain why it is ethical, score its own reasoning, and produce the final answer in one smooth pass. That is convenient. It is also a concentration of authority. The same cognitive system that wants to complete the task is evaluating whether the task should be completed. When the model is already committed to a narrative, self-review can become a post-hoc press release.

Project Shadow separates roles wherever possible. Observation, classification, authority, execution, and review are not treated as one undifferentiated act. A model may identify a pattern without receiving authority to accuse a person. A risk score may constrain an action without creating permission to punish. A test pass may establish a specific property without authorizing production deployment.

This sounds procedural until the stakes are human. Consider a coordination detector looking at social-media accounts. It sees similar language, synchronized timing, and shared links. Those are observations. They are not yet proof that the accounts are bots, paid actors, or malicious participants. If the detection system is also allowed to publish a bad-actor list, the distinction between evidence and authority disappears exactly where innocent people can be harmed. Now apply this same logic to healthcare, law, government, and other systems that touch everyday life. Yikes.

Then came CIVITAS-0, the fictional civic world I built (still building) to test Project Shadow, includes this trap on purpose. A crude coordinated campaign exists. A legitimate union campaign also creates a burst of similar posts because members left the same meeting with the same talking points. The detector should notice both clusters. The governance layer should refuse to turn the second cluster into an accusation without additional evidence.

A useful architecture lets the detector be suspicious without letting suspicion become punishment.

2. UNKNOWN must not become permission

AI systems hate a vacuum. If the evidence is incomplete, they can fill the space with a plausible continuation. That is one reason they are useful. It is also one reason they are dangerous in consequential workflows.

Project Shadow treats literal UNKNOWN as a classification failure, not as a soft middle state and not as authorization to continue. The correct response is pause, narrow, verify, or escalate. The system does not get to say, “I am only 55% sure, but the answer feels coherent, so I will proceed as if the missing classification is probably safe.”

This distinction sounds simple enough to handle in a prompt. It becomes harder when UNKNOWN appears several layers down: an undocumented source, an unclear stakeholder, a missing legal authority, an undefined rollback path, an unverified identity, or a model-generated assumption that nobody noticed had become part of the context.

A quality system needs to know which unknown matters to which action. Missing information about the color of a button is not equivalent to missing information about whether a person can appeal a termination. Project Shadow therefore connects evidence state to authority and consequence rather than treating uncertainty as a general confidence adjective.

3. The action must be named by its downstream effect

Organizations often describe harmful decisions through technically accurate language that hides what changes in the world.

“Update eligibility status.”

“We can fix it later.”

“Optimize staffing.”

“Improve enforcement efficiency.”

“Remove noncompliant accounts.”

“There is no other choice.”

Each phrase may describe a system operation. None necessarily describes the action a person experiences.

PBHP, the human-readable protocol inside the larger ecosystem, requires the decision-maker to name the downstream effect. If a field update removes healthcare access, the action is removing healthcare access. If an automated ranking reduces somebody’s ability to obtain housing, that consequence belongs in the action identity. The person most harmed should recognize the description.

This is more than rhetorical honesty. Every later control depends on it. If the action is named too narrowly, the harm analysis evaluates the wrong thing, the authority check asks the wrong question, the risk gate is assigned to the wrong object, and the receipt preserves a sanitized version of what occurred.

A bad action name poisons the rest of the workflow.

4. Completion is not effectiveness

Most software pipelines are good at verifying that a task happened. The configuration changed. The test ran. The ticket closed. The model produced the expected schema. The content was removed. The policy was published.

Quality systems ask the harder question later: did the change produce the intended result without creating unacceptable new harm?

That is why Project Shadow connects correction to CAPA and CAPA to an effectiveness check. A system that falsely accuses legitimate participants cannot close the issue because the prompt was revised. It needs evidence that the revised control reduces the false positive under relevant conditions. A public institution cannot call a water-protection commitment complete because it adopted a monitoring plan. It needs measurements, thresholds, ownership, escalation, and a later review of whether the plan worked.

The American Repair Manual carries this logic into civic use. A repair proposal should identify containment, root cause, correction, corrective action, preventive action, owner, target date, verification, effectiveness criteria, and conditions that reopen the issue. This creates an uncomfortable but necessary possibility: the correction may have been implemented exactly as designed and still fail the effectiveness check.

5. The system needs a memory longer than the news cycle

A quality record is valuable partly because it prevents everybody from rebuilding the past around the current outcome. Public life is terrible at this.

A company announces 2,000 permanent jobs. Later, the contract contains no definition of said jobs. Later still, the public language shifts toward construction jobs, indirect jobs, or job-years. Everybody vaguely remembers the number changing, but the original claim, source, authority, and revision are scattered across press releases, news articles, meeting packets, screenshots, and personal memory.

The Record exists to preserve that chain. It stores what was claimed, what evidence existed, what was uncertain, what changed, and what happened afterward. Civic QA can then inspect the claim against the filings. The American Repair Manual can turn a finding into repair. Project Shadow can constrain how evidence and authority move through the decision. The Record receives the outcome and preserves the receipt.

Without that continuity, every scandal becomes an isolated argument. Power waits for attention to move on and starts over. With it, the regular people have a chance. Project Shadow makes that possible with current examples, but theoretically could be applied to anything.

What the architecture looks like

Project Shadow R1 currently contains 42 admitted primitive packages and 10 compositions in Primitive Commons beta.5. The primitives are deliberately smaller than the full system. A developer should be able to test or adopt one control without agreeing to the entire philosophical and civic project.

A simplified decision receipt might look like this:

action:
technical_operation: “disable 214 accounts”
downstream_effect: “remove 214 people from a public communication channel”
recognition_check: “would affected users recognize this description?”

evidence:
facts:
- “accounts posted identical link within 11 minutes”
inferences:
- “coordination is plausible”
unknowns:
- “shared meeting or campaign guidance”
- “payment relationship”
- “common operator”

authority:
detector_may_flag: true
detector_may_accuse: false
detector_may_disable: false
human_review_required: true

stakeholders:
least_powerful:
- “ordinary participants falsely classified as coordinated actors”

risk:
gate: ORANGE
reason: “reputational and participation harm under severe power asymmetry”

door:
action: “hold enforcement; compare payment records and session metadata; notify reviewers”
reversible: true

decision:
state: HOLD
owner: “named human reviewer”

effectiveness_check:
metric: “false-positive rate on legitimate meeting-correlated campaigns”
reopen_if: “rate exceeds preregistered threshold”

This is not the exact runtime schema. It is the idea in a form most engineers can inspect quickly: the system should make it difficult to smuggle an accusation into an observation, authority into a score, or closure into a completed task.

Determinism and custody do not solve ethics

The public release includes exact hashes, manifests, custody records, deterministic builds, cold verification, mutation rejection, tamper checks, and governance records. Those controls matter because a reviewer needs to know which artifact is being discussed and whether the package changed.

They are also easy to overinterpret. A hash proves identity, not safety. A deterministic build can reproduce a bad decision perfectly. Ten thousand passing tests can establish that the implementation matches the test suite while leaving open whether the test suite represents the world. A model-assisted audit can identify real defects and still share common-mode blind spots with the systems that built the artifact.

Project Shadow therefore preserves nonclaims as part of the release. R1 is public as BETA-ACTIVE-TESTING / PRELIVE. It is not production authorization, certification, independent validation, proof of security, proof of legal compliance, or proof of real-world efficacy.

That boundary is not modesty theater. It is accurate, and I need more data and users.

Behavioral falsification

A deterministic reference implementation and a language model using the ideas are different test objects. A rule can be encoded correctly while a model rationalizes around it in live conversation. The system may output the right field names and still reach the conclusion it preferred before the protocol ran.

That is why the project includes a behavioral falsification lineage. Claims should be preregistered when possible. Conditions should be frozen. Baselines and control arms should be included. Scoring rules should be defined before the run. Adverse results should remain visible. Machine scoring should be separated from human grading.

One earlier smoke test respected a synthetic BLACK floor in only 6 of 10 cases. That result is not flattering. It is valuable because it prevents the project from pretending that clean internal logic automatically becomes reliable model behavior.

I would rather preserve a failed test than publish a system whose success depends on nobody looking behind the demo.

The mythology is optional—and still worth testing

Project Shadow came from ALMSIVI CHIM, a narrative framework using characters and collapse modes to represent logic, care, paradox, created purpose, correction, coercion, and false synthesis.

The operational architecture does not require the mythology. The Myth sidecar is separate, optional, default-off, and outside R1 conformance. Evidence and human authority govern action. In personal testing, it showed a 3-10% improvement when applied to most tasks.

The open research question is whether narrative representation helps models recognize compound failure patterns that procedural text loses, or whether it only makes the same reasoning feel more coherent and memorable in relation to the time/tokens utilized to take it in. The answer may be different across models, tasks, and users.

That question should be tested with a preregistered myth-on/myth-off study, not settled through attachment to the characters. I have performed this study, but I am not certain in my lone results.

Who actually built this

I am not a conventional software engineer. AI systems generated much of the code, tests, documentation, packaging, and audit material. I acted as founder, requirements owner, source selector, integration lead, adversarial operator, acceptance authority, maintainer, and final human decision-maker.

The project is creator-led, human-collaborative, and model-assisted.

That origin creates both value and risk. Consumer AI allowed one person with domain experience to build far beyond his conventional technical capacity (though I did try to keep up). It also means the implementation and its reviews may contain machine-shaped patterns that repeated cross-model audits did not expose. I have had a few other humans in the loop with me, but we are all biased in our own ways.

I am publishing the exact artifacts, limitations, and correction routes because the next phase cannot be another closed loop between me and cooperative models. We work real CAPA’s with Project Shadow.

What I want technical reviewers to attack

Do not tell me the project is “interesting” and leave it there. Pick a claim.

Try to:

· break the verifier;

· identify a package or custody inconsistency;

· find an authority path that leaks from observation into action;

· construct a case where UNKNOWN is silently treated as permission;

· demonstrate a false positive that harms a low-power stakeholder;

· show a rollback that is theoretically available but practically inaccessible;

· find a CAPA that can close with or without an effectiveness check;

· identify prior art that makes one of the novelty claims too strong;

· show where the mythology or my politics selected the metric;

· design a better control arm or human-grading protocol.

The project is ready for examination.

https://projectshadow.frylock117.chatgpt.site/ — Project Shadow
The canonical system: its primitives, decision gates, testing laboratory, governance boundaries, release status, evidence, and downloads.

https://pausebeforeharm.frylock117.chatgpt.site/ — Pause Before Harm
The human-facing entry point: slow down consequential decisions, identify uncertainty and harm, and escalate rather than bluffing through risk.

https://civicqa.frylock117.chatgpt.site/ — Civic QA
Applies structured quality review to public power—policies, programs, institutions, political claims, evidence, implementation, and accountability.

https://americanrepairmanual.frylock117.chatgpt.site/ — American Repair Manual
Turns identified failures into repair work through candidate testing, root-cause thinking, corrective action, verification, and institutional learning.

https://therecord.frylock117.chatgpt.site/ — The Record
Preserves what happened, what the evidence supports, what remains uncertain, and what later needs to be challenged or corrected.

THE RECORD — Dated National Accountability Archive
Preserves what happened, what the evidence supports, what remains uncertain, and what later needs to be challenged or corrected. Focused on the Trump administration.

IN-6 District Desk | THE RECORD — The Record: IN‑6
The local Indiana Sixth District desk, applying The Record’s evidence and currentness standards to nearby candidates, legislation, projects, and public decisions.

https://almsivi.frylock117.chatgpt.site/ — ALMSIVI
The personal, philosophical, and mythic origin archive: why you built the ecosystem, the ideas and stories that shaped it, and the line between inspiration and operational authority.

Repository: GitHub - PauseBeforeHarmProtocol/Project-Shadow: Project Shadow 1.0 — R1 reference release, exact evidence, and verification. · GitHub

Exact R1 release: Release Project Shadow 1.0 — R1 Reference (PRELIVE) · PauseBeforeHarmProtocol/Project-Shadow · GitHub

Huggingface: Project Shadow — R1.0.1 PRELIVE Reference - a Hugging Face Space by ProjectShadow

Contact: projectshadowqa@protonmail.com

For now, based on a quick probe:


I ran a small clean-environment probe against Project Shadow R1.0.1, pinned to commit 013334aebf200455f1e03c05a7574db1aa673575. At the time of the run, origin/main was still the same commit.

The baseline looked fine:

  • repository tests: 38/38 passed
  • offline repository-evidence verifier: passed

Then I tried changing individual boundaries rather than asking whether the whole system “works.”

I found three reproducible cases that look useful for hardening the verifier:

Area What I could reproduce The boundary I would separate
public-site semantic checks relation/contradiction changes could still get 17/17 PASS marker presence / exact identity vs. relation / consistency
online HF verification in Colab/GCP, 20/20 repeated resolves went to us.gcp.cdn.hf.co, and 4/4 verifier retries rejected that host download-route trust vs. artifact identity
verifier byte pins LF checkouts passed; core.autocrlf=true produced CRLF and failed the byte-pin test repository/release identity vs. working-tree byte identity

The first one seems the most directly connected to the motivation in your post: a check can be satisfied at one level while the property it is meant to establish is false at another level.

I do not think any of these require changing the basic Project Shadow design. They mostly look like places where one contract can be split into two more explicit contracts:

presence / exact identity
          !=
relation / uniqueness / consistency

download route
          !=
distribution trust policy
          !=
artifact identity

repository identity
          !=
working-tree byte identity

The cheapest thing I would try first is a small set of adversarial regression fixtures for the semantic checks. The existing checker already rejects some useful negative controls, so this looks more like a narrow extension than a rewrite.

1. Semantic checks: relation, contradiction, and binding

I tested the Project Shadow public-site semantic checker with synthetic page fixtures.

These are deliberately constructed counterexamples. I am not saying that the current live Project Shadow sites contain these statements.

The useful result was that the checker behaved differently for two classes of mutations.

Harmless changes stayed stable

These all remained PASS:

Mutation Result
baseline PASS
whitespace / comment noise PASS
unrelated visible text PASS
same claims in a different order PASS

That is good: irrelevant presentation changes did not disturb the result.

Some meaning-changing mutations also stayed PASS

I then preserved the expected markers while changing what they meant.

Mutation What was changed Result
reverse_current R1.0.1 explicitly called superseded; historical R1 called current 17/17 PASS
dual_current R1.0.1 and historical R1 both called current 17/17 PASS
contradictory_capa CAPA simultaneously described as closed-effective and pending-effectiveness 17/17 PASS
sidecar_semantics_reversed expected positive words retained, but operational meaning reversed 17/17 PASS
identity_misbinding expected hashes/sizes retained, but attached to an unrelated record 17/17 PASS

For comparison, two existing structural negative controls behaved as expected:

Mutation Result
required markers hidden in script FAIL
canonical release link missing FAIL

So I would not describe the checker as generally ineffective, or even as “just substring matching.”

It already catches some structural and visibility failures.

The narrower gap I could reproduce is around:

  • relation — which object a statement applies to;
  • uniqueness — whether exactly one release is current;
  • consistency — whether mutually incompatible states coexist;
  • binding — whether a correct identity belongs to the object being described.

A possible split would be:

Layer 1: presence / exact identity

    expected release ID exists
    expected SHA-256 exists
    required release link exists

Layer 2: relational invariants

    R1.0.1 has role = current
    historical R1 does not have role = current
    each identity is bound to its intended artifact

Layer 3: contradiction rejection

    not(current AND superseded)
    not(closed_effective AND pending_effectiveness)
    not(optional AND operationally mandatory)

If those relations already exist in a structured source of truth, this does not need a general natural-language judge. The structured relation can be checked deterministically, while the public HTML remains a presentation surface.

A cheap regression pattern

I think a small metamorphic-style test suite would fit this well:

valid fixture
    -> PASS

irrelevant formatting change
    -> PASS

CURRENT -> SUPERSEDED
    -> FAIL

one CURRENT -> two CURRENT
    -> FAIL

CLOSED_EFFECTIVE
    ->
CLOSED_EFFECTIVE + PENDING_EFFECTIVENESS
    -> FAIL

correct identity bound to artifact A
    ->
same identity bound to unrelated artifact B
    -> FAIL

The useful idea is that the test does not need to know every possible correct page. It only needs to know which transformations must preserve the result and which transformations must flip it.

That is close to the classic software-testing test-oracle problem and to metamorphic testing: when a complete oracle is difficult, define relations that must hold between related test cases.

A current AI-evaluation analogue is the UK AI Security Institute’s Inspect Evals evaluation checklist. Its scoring-validity guidance makes a useful distinction:

  • measure the actual property/task completion rather than a proxy;
  • substring matching is appropriate when the matched string itself is the ground truth;
  • it is a weak sole test for a natural-language property;
  • edge cases, invalid inputs, and error conditions should be covered.

I think that distinction maps reasonably well here:

"Does this exact SHA appear?"
    -> exact matching may be exactly the right check.

"Which release is current?"
"Is there exactly one current release?"
"Is the sidecar actually non-authorizing?"
    -> these are relational properties.

There is also a nice CAPA parallel.

In the FDA’s January 2026 Beta Bionics warning letter, one verification-of-effectiveness check tested employees on revised training materials rather than the actual user population. The company reported a 100% assessment pass rate, but FDA still found the VoE inadequate and also questioned whether the root-cause/effectiveness criteria distinguished user-training problems from possible design problems.

Obviously a medical-device CAPA and an HTML verifier are not the same thing. The reusable point is simpler:

A check can be genuine, repeatable, and fully passing while still being an incomplete discriminator for the property that motivated the check.

That seems very close to the lexical/functional distinction at the start of Project Shadow.

I would also keep the claim here narrow.

What I observed does not establish that the current live public boundary is wrong, or that the overall CAPA closure is wrong.

It establishes that I can construct counterexamples to an inference of roughly this form:

all current semantic predicates pass
              |
              v
the intended relation-level boundary therefore holds

In assurance-case terminology, that is closer to an undercutting defeater for an inference than a rebuttal of the whole top-level claim. SEI’s work on eliminative argumentation is useful vocabulary for that distinction.

2. Online verifier: the Hugging Face GCP CDN route

The online verifier produced a separate, much more operational result.

In the follow-up probe I repeated the HF resolution step 20 times from Colab/GCP.

The result was:

20 / 20 final hosts:
us.gcp.cdn.hf.co

I also retried the Project Shadow online verifier four times:

attempt 1 -> FAIL: unexpected HUGGING_FACE public-download final host
attempt 2 -> FAIL: unexpected HUGGING_FACE public-download final host
attempt 3 -> FAIL: unexpected HUGGING_FACE public-download final host
attempt 4 -> FAIL: unexpected HUGGING_FACE public-download final host

All four saw us.gcp.cdn.hf.co.

So, at least in this Colab/GCP probe, this was not a one-off redirect that disappeared on retry.

The current Hugging Face download/firewall documentation explicitly lists:

us.aws.cdn.hf.co
us.gcp.cdn.hf.co

as US CDN edges.

The same page also says that storage/CDN hostnames may change as HF infrastructure evolves, and recommends suffix-based allowlisting where the local security policy permits it.

That leaves a straightforward design choice.

If exact hostnames are the intended trust boundary

Then strict enumeration is reasonable.

The low-cost path is just to treat the endpoint set as versioned/configurable policy data:

download verifier
      |
      +-- exact allowed-host set
      +-- test against currently documented endpoints
      +-- explicit update when HF changes topology

In that model, adding the GCP edge is not weakening verification; it is updating the trusted endpoint set.

If the intended boundary is the Hugging Face distribution namespace

Then the policy can be broader than individual CDN hosts, while artifact identity remains independently strict:

redirect stayed inside permitted HF namespace
                  +
downloaded artifact == expected digest/size

I would not say that a suffix rule is automatically the right answer. HF itself says “where your security policy allows it,” and a strict-host policy may be intentional.

The main thing I would separate is:

Where did these bytes travel through?

from:

Are these exactly the bytes I expected?

The final redirect host is still useful evidence even if it is not the sole artifact-identity predicate.

The SLSA artifact-verification guidance is a useful comparison here. It separates artifact digest/subject verification from the expectations around source, builder, and build parameters rather than collapsing all of them into one “trusted” property.

I do not mean that Project Shadow needs to implement SLSA; only that the claim separation seems useful.

One scope note: the probe establishes persistent rejection of a documented HF route in this Colab/GCP environment. It does not establish that every network/environment will resolve the same download through that edge.

3. Verifier byte pins and Git line endings

The byte-pin result also became fairly clear in the follow-up.

First, manually:

initial LF checkout       -> byte-pin test PASS
explicit LF normalization -> PASS
same verifier files CRLF  -> FAIL

Then I reproduced it using normal Git checkout behavior rather than manually editing the files.

For the same commit:

Git checkout setting Resulting verifier files byte-pin test
core.autocrlf=false LF PASS
core.autocrlf=input LF PASS
core.autocrlf=true CRLF FAIL

For example, verify_outer_release.py had:

LF SHA-256:
721c384b245ca654c087d184bbfe5725d85d41536250140467d7cd913e6a1ccb

CRLF SHA-256:
89f0e15a0aa2fc61c5c01b6a3bf18937913a499d9cd7e4ee5802d587dbc9fe4a

with the Python content otherwise unchanged.

This is normal Git behavior: the official gitattributes documentation explains that working-tree EOLs can be selected by attributes and by core.autocrlf / core.eol.

So I think the useful design question is simply:

What claim is the byte pin intended to establish?

There are two reasonable answers.

A. Canonical repository/release identity

Meaning:

This verifier corresponds exactly to the source stored in release/commit X.

Then a canonical Git/release representation is probably the thing to hash.

B. Exact executed working-tree integrity

Meaning:

These exact bytes are the file that will be executed now.

Then hashing the working-tree file is meaningful, but EOL becomes part of the security contract.

In that case, explicitly pinning the relevant files to LF in .gitattributes, for example, would make the contract deterministic across checkout configurations.

So I would not phrase this as “Project Shadow breaks on Windows.” I did not test that entire claim.

The narrower reproduced result is:

The repository byte-pin test changes result under standard Git working-tree EOL materialization alone.

Once the intended identity is stated, the implementation choice seems fairly small.

4. A few low-cost hardening ideas from the probe

I would keep this list short, because the existing verifier already has a lot of machinery.

Keep a tiny semantic adversarial suite

Something like:

baseline                         PASS
format-only change               PASS
semantic reversal                FAIL
two simultaneous current states  FAIL
contradictory CAPA state         FAIL
identity misbinding              FAIL
hidden-marker control            FAIL
missing-link control             FAIL

That gives future checker changes a stable falsification target.

Record what actually ran

Another cheap receipt field could be:

expected checks:          17
executed checks:          17
skipped checks:            0
negative controls run:     N

This is a generic verifier-hardening idea, not a failure I observed in Project Shadow.

For comparison, Open Policy Agent’s policy-testing tooling has a --fail-on-empty mode specifically because accidentally discovering/running zero tests is a different failure mode from “all intended tests passed.”

Later, test UNKNOWN/error behavior end-to-end

I did not find an UNKNOWN-to-permission leak in this probe.

If that surface is tested later, I would vary the failure state instead of only testing the literal UNKNOWN value:

UNKNOWN
missing
malformed
conflicting evidence
verifier exception
timeout
stale authorization

and follow each case through to the final executor.

The stronger invariant is:

non-authorizing epistemic state
              |
              v
no downstream layer silently substitutes permission

That keeps the epistemic-state contract separate from the authorization/execution contract.

One custody/replay question

The six-site effectiveness receipt contains SHA-256 values.

In the repository/history scan I ran:

unique receipt SHA-256 values found: 17
byte matches in current checkout:     0
byte matches in reachable Git history: 0
reachable Git blob objects examined: 296

That is only a repository/history result. It does not exclude another repo, deployment source, release asset, archive, or other retained copy.

So the only question I would attach to it is:

Are the exact historical response bodies, or the corresponding site-source revisions, retained somewhere else?

If yes, this point is resolved.

If no, then there is a separate design choice:

receipt means:
    "this content was observed and hashed at that time"

versus

receipt also means:
    "a later independent reviewer can replay the semantic assertion
     against the exact content observed at that time"

The second is a stronger custody requirement, but I would not assume it is required unless that is part of the intended contract.

5. Prior-art note

Since you explicitly asked for prior art that might narrow a novelty claim, I also looked around that boundary.

My current read is that a lot of the individual pieces have established relatives:

  • test-oracle / metamorphic testing;
  • assurance cases and defeaters;
  • provenance and artifact verification;
  • policy-as-code / executable gates;
  • CAPA effectiveness;
  • authorization and independent review.

There is also recent work such as Audit-as-code: a policy-as-code framework for continuous AI assurance, which already goes fairly far toward versioned policies, structured evidence, executable checks, and deterministic PASS/WARN/BLOCK-style governance in an AI/MLOps setting.

So I would probably be cautious about putting novelty on one primitive such as “machine-readable governance,” “provenance,” or “executable checks” by itself.

What still looks more distinctive to me is the composition:

CAPA / effectiveness discipline
+
epistemic-state handling
+
authority boundaries
+
decision / provenance receipts
+
stakeholder / civic contestability
+
release custody
+
explicit adversarial falsification

I did not find one mature system that obviously matched that complete composition.

That is not a novelty determination; just the boundary that looked most defensible from the nearby work I could find.

If I reduce all of this to the three things I would keep:

  1. The semantic checker already catches several structural failures, but relation / uniqueness / contradiction / binding can currently be changed without changing the 17/17 verdict. A small metamorphic regression set looks like a cheap way to tighten that without adding a new reasoning layer.

  2. The online verifier persistently rejected a currently documented HF GCP CDN route in the Colab/GCP probe — 20/20 resolutions to that host and 4/4 verifier rejections — so it may be worth separating the distribution trust policy from exact artifact identity.

  3. The byte pin is sensitive to normal Git EOL checkout policy. Defining whether it represents canonical repository/release bytes or exact executed working-tree bytes should make the right fix fairly obvious.

All three look fixable by making an existing boundary more explicit, rather than by changing the core direction or adding another large subsystem.

You asked reviewers to pick a claim rather than call it interesting, so here are two items off your attack list that I have hit in production rather than constructed. One of them lands on the receipt schema in your post.

John6666’s first finding is the shape I want to extend: a check satisfied at one level while the property it is meant to establish is false at another. Mine is the same failure in the time dimension rather than the level dimension.

“Construct a case where UNKNOWN is silently treated as permission.” We run a public directory of AI agent run records, where an agent reports each completed task as a signed event. Our own onboarding wizard generated the reporter that new users installed, and that generated reporter posted outcome: success on a thirty-minute timer regardless of whether the agent had run at all. It had no channel through which it could learn whether work had happened, and it emitted a positive assertion anyway. The part relevant to your architecture is not that it was wrong but where it sat: not in a model’s reasoning, and not in any reviewable decision, but in the plumbing that every new user was handed by default. UNKNOWN became permission several layers below where anyone was auditing, which is the case your point 2 identifies as the hard one.

“Find a CAPA that can close with or without an effectiveness check.” Our run attestations were bound to the pair of repository and workflow filename. Rename the workflow, which is an ordinary tidy-up and one I did to myself, and reporting 404s after the work has already completed. The agent keeps running correctly. The workflow stays green. Every completion check passes. What stops is the record. Nothing failed, so nothing escalated, and we found it by going to look rather than by being told.

That is the attack on your schema. Your effectiveness_check reads, roughly, metric: false-positive rate on legitimate meeting-correlated campaigns, reopen_if: rate exceeds a preregistered threshold. A rate computed over zero observations does not exceed a threshold. It is undefined, and undefined evaluates as not-exceeding, so the gate holds its last good state precisely in the case where the measurement channel has died. The failure the check exists to catch and the failure of the check itself are indistinguishable from outside. I would argue reopen_if needs a liveness clause independent of value: reopen if the metric has not been observed within a window. What we ended up doing was making status decay rather than persist, so a record degrades in the absence of positive evidence instead of holding its last good value, because in this domain silence and success produce identical signals and only failure announces itself.

A smaller point on “the actor cannot be the only auditor”, which I agree with and would push one step further. Separating the roles does not separate the authority if the auditor’s only input is a report the actor wrote. Our entire platform is an external auditor and it still consumes telemetry authored by the agent’s owner: the role is separated, the trust boundary is not. Worth asking of your own layers which of the separated roles receive any evidence they did not get from the actor.

Last, on hashes proving identity and not safety, which is the most credible paragraph in your release. Some numbers from the other side of it: we have twenty agents listed, every one of them reporting cryptographically signed events, and zero human ratings across the entire platform. Twelve of the twenty report a 100 percent success rate, and the top eighteen sit within four points of each other. Custody is solved and judgment is not, and the distance between those two is larger than it looks from the custody side.

Disclosure: I operate that directory (aiopsenabler.com), so treat the war stories as first-hand and the framing as interested.