A test that was not run is not a test that was passed. 25 of the 29 cells below have been run. The rest are marked not applicable or blocked, and neither claims a result. Nothing here is scored out of anything.
A null result is a result, and it stays on the page. PACE’s first forward criterion read came back null. It is published as null rather than removed, because a validation page that only carries wins is marketing.
These are the technical headlines, not a summary of them. Each line is the figure the test itself produced. The full method for each model is published: PACE, C-RISK, C-FIN.
PACEProtected-area conservation effectiveness
PACE rates how effectively each protected area is actually being managed. It scores ten dimensions, from biodiversity and ecosystem services through security, governance, tourism, community development, leadership and budget, each on a one-to-five scale, across 200 areas in 45 countries. A higher score means a better-run area.
Internal structureRun
Do the dimensions measure one underlying thing, or several unrelated things wearing one name?
PCA PC1 61.9% (primary_7dim, n=200)
Method: PCA PC1-variance (coverage-aware specs), Cronbach alpha, inter-dim Spearman, VIF, corrected item-total Covers 200 of 200. Cadence: recurring.
ReliabilityRun and published
Score the same subject twice, blind, and compare. Does the model give the same answer?
Test-retest rho 0.864 (n=20), clears the 0.85 strong threshold
Method: Test-retest reliability: blind re-draft of the same PA by the same model, composite rank correlation. Cadence: recurring.
RobustnessRun
Move the weights and re-rank. Does the ordering survive choices that were judgment calls?
Weight sensitivity min Spearman 0.9959 across 18 sets (+/-5pp); band movers mean 13.8 max 39 of 199 (most sensitive biodiversity)
Method: Robustness: equal-weight 1/9 dimensions, each perturbed +/-5pp with the rest renormalised, PACE composite recomputed for every PA, Spearman vs baseline; band movers counted against the published score bands. Ingested from the published per-PA sensitivity artifact. Cadence: recurring.
External comparisonRun and published
Set the score against independent datasets measuring something related. Does it agree where it should?
Benchmarked vs PAVIS, UCDP, KBA, Hansen, FIRMS; KBA +0.44, FIRMS -0.20, Hansen null
Method: Convergent validation against independent third-party datasets across the relevant dimensions. Cadence: recurring.
BiasRun
Does the score systematically favour or penalise a group for something other than what it claims to measure?
Francophone composite runs -0.4154 below anglophone with GDP and area controlled (p <0.001, n=193); disclosed, not corrected. Raw group means anglophone 2.869, francophone 2.411
Method: Bias: PACE composite against working-language group, with a multivariate OLS controlling for log GDP per capita and log area (anglophone reference). Ingested from the published per-PA structural-bias artifact. Cadence: recurring.
StressRun
Push the inputs to an adverse case. How far can the published number move?
Adverse shock worst-case drop 1.0 (mean 0.918); degradation max swing 0.271 on dropping biodiversity; no dropped dimension outscores its scale maximum
Method: Stress: adverse 1-point input shock (floored at scale min) and single-input data degradation, recomputed through the model combination rule. Also tests whether dropping a dimension can outscore setting it to its scale maximum, i.e. whether the model rewards non-disclosure over perfect evidence. Cadence: recurring.
CriterionRun
Does the score predict the outcome it claims to predict?
First forward read is null: PACE at t versus post-snapshot savanna NDVI, rho -0.00 over an 8-week window (n=181). Harness runs monthly; a short-horizon read, not yet a verdict.
Method: PACE score at snapshot t versus the vegetation (NDVI) trajectory measured strictly after t, the savanna-visible outcome that replaces Hansen tree-cover loss. Spearman with a bootstrap 95% CI, a forest/savanna biome re-cut of the Hansen null, and a data-latency gate on the C-RISK arm. Cadence: recurring (monthly; gains power as the satellite series lengthens).
C-RISKCountry risk for conservation capital
C-RISK scores country risk for conservation capital across the 54 African countries. It works across five slices, fiscal capacity, the deal environment, the conservation asset, delivery capacity and traction, each on a one-to-five scale. Higher is safer: a high composite means a country better placed to take a new dollar of conservation funding, and a low one means that dollar runs more risk to get the same work done. Two gates, active armed conflict and sovereign distress, cap the readiness call regardless of the number, and governance is read as context rather than scored into the headline.
Internal structureRun
Do the dimensions measure one underlying thing, or several unrelated things wearing one name?
PCA PC1 44.2% across the five readiness slices (n=44 complete cases)
Method: PCA PC1-variance (coverage-aware specs), Cronbach alpha, inter-dim Spearman, VIF, corrected item-total Covers 44 of 54, 10 dropped. Cadence: recurring.
ReliabilityRun
Score the same subject twice, blind, and compare. Does the model give the same answer?
Rated-input arm min Spearman 0.9899; quantitative measurement arm 0.9998 (n=54)
Method: Split reliability: +/-0.5 rater jitter only on the three rated deal_env inputs; separate 2dp measurement-precision perturbation on quantitative slices; seeded Monte Carlo Cadence: recurring.
RobustnessRun
Move the weights and re-rank. Does the ordering survive choices that were judgment calls?
Slice-weight sensitivity min Spearman 0.9764 across 10 sets (+/-5pp); public-call movers mean 0.9 max 2 of 54 (most sensitive delivery)
Method: Robustness: each of the five readiness slice weights perturbed +/-5pp with the rest renormalised; score renormalised over slices present. Reports Spearman rank stability and movers across the four post-gate public readiness calls. Cadence: recurring.
External comparisonRun
Set the score against independent datasets measuring something related. Does it agree where it should?
Pre-registered WGI decorrelation gate: Pearson 0.601 (the published +0.86 -> +0.60 figure), Spearman 0.515 (n=54); secondary EPI 2024 overall arm 0.149 (n=51)
Method: External: lead with the pre-registered readiness-vs-WGI decorrelation gate; retain overall Yale EPI 2024 as a secondary exploratory continuity arm. Both higher is better. Cadence: recurring.
BiasRun
Does the score systematically favour or penalise a group for something other than what it claims to measure?
Spearman 0.336 readiness vs income rank; partial rho 0.135 controlling fiscal; Kruskal-Wallis H 7.15, p 0.0671 across 4 income classes (n=54)
Method: Bias: readiness vs World Bank income rank, plus partial Spearman controlling for the 0.30 fiscal slice; Kruskal-Wallis across classes as support. Cadence: recurring.
StressRun
Push the inputs to an adverse case. How far can the published number move?
Slice-drop max swing 0.845 on fiscal; public-call movers max 15; disabling gates moves 1 active-conflict and 0 sovereign-distress calls; 9 countries score without asset/delivery
Method: Stress: five readiness slices dropped individually with renormalisation; gate caps disabled individually; live partial coverage characterised separately. Cadence: recurring.
CriterionBlocked
Does the score predict the outcome it claims to predict?
Cannot be run yet: it waits on data that does not exist. No result is claimed.
Cadence: recurring (unblocks on the UCDP 2026 release, or a current-year conflict feed).
C-CARBONCarbon project ratings
C-CARBON rates carbon projects for quality and risk. It scores five dimensions, including methodology, financial viability and permanence, and scores each of 269 projects on a 1 to 5 scale. A higher score means a more credible project.
Internal structureRun
Do the dimensions measure one underlying thing, or several unrelated things wearing one name?
PCA PC1 40.9% (all dims)
Method: PCA PC1-variance (coverage-aware specs), Cronbach alpha, inter-dim Spearman, VIF, corrected item-total Covers 274 of 274. Cadence: recurring.
ReliabilityRun
Score the same subject twice, blind, and compare. Does the model give the same answer?
deterministic, re-run dev 0.0 (n=273)
Method: deterministic re-derivation: parity-checked mirror re-run, composite identity Cadence: recurring.
RobustnessRun
Move the weights and re-rank. Does the ordering survive choices that were judgment calls?
Weight sensitivity min Spearman 0.9916 across 10 sets (+/-5pp); band movers mean 12.5 max 28 of 273 (most sensitive community_architecture)
Method: Robustness: each composite weight perturbed +/-5pp with the rest renormalised, composite re-derived through the C-CARBON mirror, rank correlation vs baseline; band movers counted against the published score bands. Cadence: recurring.
BiasRun
Does the score systematically favour or penalise a group for something other than what it claims to measure?
Portfolio is effectively single-registry; registry confound not testable (n Verra 272, n GS 0, 1 non-Verra excluded). Verra mean 2.201, no Gold Standard comparison group
Method: Bias: composite by registry standard (Verra vs Gold Standard), group means and Mann-Whitney; other or unmatched registries reported as an excluded count. Cadence: recurring.
StressRun
Push the inputs to an adverse case. How far can the published number move?
Adverse shock worst-case drop 1.0 (mean 0.713); degradation max swing 0.6 on dropping sovereign; no dropped dimension outscores its scale maximum
Method: Stress: adverse 1-point input shock (floored at scale min) and single-input data degradation, recomputed through the model combination rule. Also tests whether dropping a dimension can outscore setting it to its scale maximum, i.e. whether the model rewards non-disclosure over perfect evidence. Cadence: recurring.
CriterionRun
Does the score predict the outcome it claims to predict?
Concurrent known-groups signal: withdrawn/terminal projects rate below live ones, and it holds on the status-independent part of the score (rank-biserial -0.31, p=0.02, n=248 vs 21). Not yet a forward test: only 3 in-window withdrawals so far.
Method: C-CARBON rating at t versus the withdrawn/terminal cohort (fetch_verra in_cohort=false, 21 of 269, incl. Kariba). Mann-Whitney U of the live cohort's ratings against the terminal cohort's, with a rank-biserial effect and a bootstrap 95% CI, read on the full composite and on a status-independent composite (methodology + sovereign + community), because 0.45 of the composite weight is mechanically coupled to lifecycle status. A forward t->t+1 arm counts active->terminal transitions and waits for power. Cadence: recurring (monthly; the forward arm gains power as snapshots accrue withdrawals).
Country scoreC-RISK readiness rescaled 20 to 100
The country score is one number from 20 to 100 for each of the 54 African countries, ranking them by how ready each one is for long-term conservation funding. It is the C-RISK readiness score rescaled, readiness from 1 to 5 multiplied by 20, so the two always move together and the attainable floor is 20 rather than 0. Two gates, active conflict and sovereign distress, can cap the tier a country lands in without changing its number. The formula is published so anyone can re-run it.
Internal structureNot applicable
Do the dimensions measure one underlying thing, or several unrelated things wearing one name?
This family does not apply to this model. It was not run, and no result is claimed.
ReliabilityRun
Score the same subject twice, blind, and compare. Does the model give the same answer?
Arithmetic alias of c-risk. Rated-input arm min Spearman 0.9899; quantitative measurement arm 0.9998 (n=54)
Method: Arithmetic alias of c-risk reliability: CCS = 20 × c-risk readiness, so rank reliability is identical; no independent validation arm. Split reliability: +/-0.5 rater jitter only on the three rated deal_env inputs; separate 2dp measurement-precision perturbation on quantitative slices; seeded Monte Carlo Cadence: recurring.
RobustnessNot applicable
Move the weights and re-rank. Does the ordering survive choices that were judgment calls?
This family does not apply to this model. It was not run, and no result is claimed.
External comparisonRun
Set the score against independent datasets measuring something related. Does it agree where it should?
Arithmetic alias of c-risk. Pre-registered WGI decorrelation gate: Pearson 0.601 (the published +0.86 -> +0.60 figure), Spearman 0.515 (n=54); secondary EPI 2024 overall arm 0.149 (n=51)
Method: Arithmetic alias of c-risk external validity: CCS = 20 × c-risk readiness, so Pearson and Spearman are identical; no independent validation arm. External: lead with the pre-registered readiness-vs-WGI decorrelation gate; retain overall Yale EPI 2024 as a secondary exploratory continuity arm. Both higher is better. Cadence: recurring.
BiasRun
Does the score systematically favour or penalise a group for something other than what it claims to measure?
Spearman 0.246 CCS vs coverage tier (higher-coverage countries score higher); per-tier means imputed 57.9, thin 50.7, moderate 58.8, rich 60.8 (n=54, 9 imputed)
Method: CCS bias confound is DATA-COVERAGE TIER, not income. CCS = 20 x c-risk readiness, so this equals a readiness-vs-coverage-tier measurement; it is reported here because no other C-VAL cell tests coverage tier (pace tests working language, c-risk tests income, c-carbon tests registry standard). It is an alias in mechanism, not a duplicate in question. Bias: CCS vs data-coverage tier (rated-PA count band), Spearman on the tier with per-tier mean CCS as support; imputed countries are tier 0 and flagged. Cadence: recurring.
StressNot applicable
Push the inputs to an adverse case. How far can the published number move?
This family does not apply to this model. It was not run, and no result is claimed.
CriterionRun
Does the score predict the outcome it claims to predict?
First country-level read is directionally positive but not yet significant: CCS vs post-t NDVI rho +0.19 over an ~8-week window (n=41). Replicated at t=2026-06-01: rho +0.35 (n=39, p=0.03). A short-horizon first read that gains power monthly, not a verdict.
Method: CCS at t versus the country-level aggregation of the PACE forward outcome: mean post-t NDVI deviation from seasonal baseline across each country's PAs. Spearman with a bootstrap 95% CI, primary at t=2026-05-01, replicated at t=2026-06-01. NOTE: reads CCS as published at t from the snapshots, which predate the v1.7 reframe, so this measures the retired four-component CCS. The current CCS (20 x C-RISK readiness) has no forward history yet and inherits this arm as it accrues. Cadence: recurring (monthly; inherits the PACE forward arm, gains power as the series lengthens).
PropagationRun
Does an error in an input carry through to the published number, or get absorbed?
42 of 54 tier-stable, 12 fragile near band edges; weight-set min Spearman 0.9772 (benchmark 0.989)
Method: Two-arm uncertainty propagation through live CCS=readiness x20: readiness via 10 five-slice weight perturbations, PACE sampling uncertainty via the 0.20 asset slice; band-flip probability against the 65/50/35 tiers. Cadence: recurring.
RedundancyRun
Does this score say anything the model it derives from does not already say?
PACE scoring exposure is 0.20 through readiness.asset; the legacy 0.25 PACE arm contributes 0 because it is display-only in ccs.js v1.7. As a display-signal diagnostic, readiness holds 64.48% of weighted spread to PACE's 35.52% under the retired 0.50/0.25 weights (SD ratio 0.908x); the country trimmed_max is the most C-RISK-redundant PACE input (r^2 0.277), p75 the least (r^2 0.159); 0 unresolved PA joins
Method: Embedded-input redundancy: verify the live CCS formula, quantify PACE exposure through readiness.asset, and retain the old RISK/PACE weighted-spread comparison only as a clearly non-scoring display-signal diagnostic; candidate comparison of PACE statistics (mean / within-SD / range / best / p75 / trimmed-max) by how much each restates C-RISK on the >= 3-PA countries. Also asserts the canonical iso3 PA->country join is complete (Cote d'Ivoire / DRC sentinel). Cadence: recurring.
That the scores are correct. Validation tests whether a model behaves like a measure: whether it is internally coherent, repeatable, stable under different reasonable choices, and free of a bias it does not intend. None of that makes a score true.
That the samples are large. Several of these tests run on small samples and their intervals are wide. Where a figure has an interval, the interval is the honest reading of it, not the point estimate.
That the coverage is even. Families differ by model, because the tests that apply differ by model. A model with fewer rows has not been examined less carefully; it has been given the tests that mean something for it.
Canopy is a nonprofit and everything here is free to use. The one condition is attribution.