Level 5 · Practice & recap
Run it yourself
Section titled “Run it yourself”The arc runs on a laptop CPU, on the real Campus Placement dataset (~215 rows; no GPU, no token, no synthetic data). Each command’s result is exactly the tag linked in the arc.
Use the V1 milestone chip (in the arc) to clone the repo and check out the starting point, then install the project dependencies:
uv sync # install deps from pyproject.toml + uv.lock (reproducible)With the project environment active, reproduce the arc. Each block below is the REAL step from the steps/ script —the same one that builds the published repo, not a hand-written copy—; the “copy” button gives you the exact command:
1 · Measure V1 (with the leak) → RED on the one blocker. Compile the gate contract and run it. froga run exits non-zero on purpose (the blocker fails), and still signs the evidence — with the two advisory findings already measured alongside the red:
froga compilegit add -Agit commit -m "compile: assessment plan OSCAL (contrato del gate, antes del modelo)"
patch: params.patchpatch: dvc.patchpatch: train.patchpatch: evaluate.patchpatch: compliance-eval.patchfroga runmay failgit add -Agit commit -m "modelo: gradient-boosting V1 (train/evaluate + dvc.yaml) — run: V1 evidencia base (gate RED esperado; seed=42)"
2 · Treat the leak (drop salary) → GREEN. The treatment is a single versioned change: edit params.yaml to mitigate: true —the diff below is exactly what you change— and commit it with the script’s message. The flip drops the leaked feature and, in the same step, you run the gate again: since it is the only blocking control, the overall gate closes:
git checkout -b tratamiento/retira-salarypatch: params-mitigate.patchfroga rungit add -Agit commit -m "treatment: retira la feature fugada salary (mitigate false→true; cierra model-feature-leakage; data-sample-adequacy sigue abierto)"
Step 080 reproduces talent-v1.0.0-red; step 090 reproduces talent-v2.0.0-green. Note what does not change: data-sample-adequacy still reads 0.43 (215/500) and selection-rate-disparity-gender is still underpowered — neither is blocking, so neither prevents the GREEN, but neither disappears either: they ride advisory all the way to the approval. A nightly CI job re-clones each tag and confirms red → green.
What you just saw
Section titled “What you just saw”Your third overlay, this time in employment: a talent-screening classifier high-risk through Annex III §4 (the employment / worker management route — neither retina’s Art. 6(1) product route nor credit’s Annex III §5(b) use route). The risk program carries three controls measured on one set of tags, but only one is blocking. The blocker closed: a feature-leakage control was RED at V1 because the model trained on salary (present only for placed candidates → collinear with the label → total leakage, a safety defect under Art.15); Martha flipped mitigate: true, the pipeline dropped the salary feature, and model-feature-leakage went RED→GREEN — since it was the only blocker, the overall gate closed GREEN. The other two stayed advisory: data-sample-adequacy (sample_adequacy_ratio 0.43 = 215/500) and selection-rate-disparity-gender (selection-rate parity by sex, underpowered at n=215) were genuinely measured at V1 and V2, unchanged, but their enforcement: audit meant they never blocked the gate — the governance decision (Option B) was not to manufacture a blocking threshold the 215-row sample, capped by the public dataset itself, cannot sustain with statistical reliability. At governance, James approved the system GREEN, consciously accepting that advisory residual with a recorded motive and quarterly post-market monitoring (Art.72) of selection disparity — prEN 18228’s clause 11 flipped from Gap to Covered on reading the approval act. Art. 10 went deep (the 10(3)/(5) union: a too-small sample limits — but, because it is honestly declared, does not block — bias certification). National labour law (ET Art. 64.4.d) + RGPD Art. 22 anchor the live obligation, with national legislation at rank 0 → no presumption. And Aitor debuted: an external auditor who froga verify --pubkey’d the authentic signature, read the honest GREEN, approved, with the advisory residual on record state and — being not a notified body — did not certify. The line to carry away: the engine certifies what it can measure reliably; it honestly documents and monitors what the sample will not let it certify — without manufacturing either a permanent red or a certification the data cannot support.
Self-check
Section titled “Self-check”The risk program measures three controls on one set of tags, but only one blocks the gate. Describe all three — what each measures, which one closes the arc, and why the other two are advisory instead of gate.
The blocker — feature leakage. The control model-feature-leakage (metric leaky_feature_flag, gate < 1.0, enforcement: gate) is 1.0 → RED at V1 because the model trains on salary, a feature present only for placed candidates (NaN→0 otherwise) and therefore collinear with the placement label — total leakage, a safety defect (Art.15). Martha flips mitigate: true, the pipeline drops the salary feature, and the control reads 0.0 → GREEN at V2 — a real RED→GREEN by the treatment (removing the feature, a versioned params.yaml flip, ISO 23894 §6.5), with a bounded residual (risk.feature-leakage: residual_likelihood: UNLIKELY → MEDIUM, within appetite). Because it is the ONLY enforcement: gate control in the program, its close alone is enough for the overall gate to be GREEN. The two advisory controls — data adequacy and selection fairness. data-sample-adequacy (sample_adequacy_ratio, gate ≥ 1.0) reads 0.43 (215/500) at V1 and V2 — unchanged, because dropping a column adds no rows. selection-rate-disparity-gender (demographic_parity_diff by sex) is genuinely measured on the test cohort, but with n=215 (~24 women) it is underpowered. Both are enforcement: audit: they get measured and signed, but do not enter the blocking-verdict computation — the governance decision was that certifying them as gate would manufacture a threshold this structurally capped 215-row sample cannot sustain reliably.
James opens the system to approve and sees the gate GREEN, but with two advisory findings uncertified. What exactly does he accept when he approves, and how is that different from the 'red door' override-with-motive you see in other levels?
James accepts the data/fairness advisory residual: data-sample-adequacy (n=215/500=0.43, with no residual_likelihood declared because you cannot estimate a residual over a sample that does not let you measure it) and selection-rate-disparity-gender (measured, underpowered at n=215). The difference from other levels’ “red door” (residual-exceeds-appetite / untreated-above-appetite) is that those apply to a blocking control the engine marked failed or inconclusive, and the approval consciously overrides THAT block. Here there is no block to override: data-sample-adequacy and selection-rate-disparity-gender were never enforcement: gate — the governance decision (Option B) declared them advisory before measuring, precisely to avoid manufacturing a blocking threshold the sample cannot sustain. What James accepts is a documented, measured finding, not an exception to a technical verdict — the same conscious governance act (ISO 23894 §6.5.2 / 42001 6.1.3), without the red/amber in the way. James also sets quarterly post-market monitoring (Art.72) over selection disparity.
The system declares ET Art. 64.4.d and Ley 15/2022 (Spanish national law). Why does national legislation carry NO presumption of conformity, and what does that mean for how the engine treats ET 64.4.d?
Because national legislation is a NationalLegislation authority at rank 0 in the engine’s typed authority chain. A harmonised standard (or a governed standard like prEN 18228 / ISO 23894) can confer a presumption of conformity — prove its clauses and the verdict means something legally. A NationalLegislation authority does not confer that presumption, but it can still be projected against its own catalog. And specifically, ET Art. 64.4.d does have a clause catalog wired into the engine (et.64-4-d-transparencia-algoritmica), and its verdict can be Covered — the signed bundle satisfies it (the accessible record of the algorithm’s parameters, rules, and instructions). What rank 0 denies it is the presumption, not the verdict: a Covered on ET 64.4.d is a national obligation met, never a presumption of AI Act conformity. The practical consequence: the live obligation of this overlay is anchored on national labour law + RGPD (what they actually require now), not on the deferred AI Act high-risk gate date — but “ET 64.4.d Covered” must never be read as “the AI Act is presumed met”.
Where to go next
Section titled “Where to go next”You have now met your third overlay, the distinction between a certifiable safety defect and a declared data/fairness limit, a legitimate approval with the advisory residual on record, the live anchor in national law + RGPD, Article 10 in depth, and Aitor’s public-key-only verification. To go deeper on the mechanisms this level touched: