10Canopy Proprietary

PACE Methodology

Protected Area Conservation Effectiveness
PACE rates how well each major African protected area is being run. Each park is scored on ten measures across four areas: ecology (is wildlife thriving?), operations (is the park well-managed and secure?), finance (does it have stable funding?), and community (are local people benefiting?). The full methodology, including what data we use and where it comes from, is below.
Go deeper
This page sets out how PACE works and what it is built to withstand. The exact source list, dimension weightings, scoring thresholds, and prompts sit behind it. We walk funders and partners doing serious diligence through them directly, one to one.
Request a methodology walkthrough →

Chapter 1: The model

Data sources flowing into the ten PACE dimensions
DATA SOURCES Registries and ecology Species, habitat and forest datasets Operator and authority filings Reports, filings and audited accounts Qualitative evidence News · gov statements · community reports TEN DIMENSIONS, TWO TIERS TIER 1 · RUBRIC-CLEAR · 6 OF 10 TIER 2 · SYNTHESIS-REQUIRED · 4 OF 10 Biodiversity protection Ecosystem services (carbon and water) Ecotourism infrastructure Community development Budget Fundraising performance Security and anti-poaching synthesis required Ecotourism revenues disclosure-dependent Governance and accountability qualitative oversight Political and community leadership explicit editorial judgement COMPOSITE Weighted mean of ten dimensions, scored 1.0 to 5.0 SCORE BAND 4.6 4.2 3.8 3.4 3.0 2.5 <2.5
Sources feed ten dimensions across two data tiers, combined into a weighted composite score from 1 to 5.
1.1 What PACE measures

Each protected area is scored 1 to 5 across ten dimensions covering ecology, operations, finance, and community outcomes. Scores combine into a single composite score from 1 to 5. PACE answers a specific question: is this place being run well enough to merit long-horizon conservation capital. Month to month, a published score is anchored to the prior month and moves only on new dated evidence, with every change labelled; section 4.4 sets out the mechanism.

Coverage is 200 of Africa's most operationally significant PAs across African Parks managed sites, Peace Parks landscapes, SANParks flagships, Kenya Wildlife Service parks, partnership-run reserves, and major state-managed icons. Coverage expands only as data quality allows defensible scoring.

1.2 How inputs are sourced

PACE has no automated numeric feeds. There is no API pulling species counts into the rating file, no scheduled ingestion of forest or carbon data, no financial-filing line-item parser, no ranger-density data link. Every score across all ten dimensions is generated by a structured large-language-model pipeline that consults the relevant source material per dimension, reads the underlying evidence, and applies a fixed rubric to return a 1-to-5 score with a source trail.

The Tier 1 and Tier 2 split that follows is about source-material clarity, not process. Tier 1 dimensions have source material that maps directly to the rubric: disclosed numbers, published trends, registry entries. Tier 2 dimensions require the drafter to weigh ambiguous or qualitative evidence. The same drafter runs both. The same source trail is published for both.

The defensibility of this approach rests on three things. First, a bounded scope of 200 auditable PAs where any reader can interrogate any score against the published source trail. Second, a fixed published rubric per dimension that constrains the drafter to a transparent scoring logic. Third, immutable monthly snapshots that build a panel record over time, so changes to scores can be tracked, attributed, and reviewed.

1.3 Two-tier scoring

PACE splits its ten dimensions into two tiers based on data availability. This reflects what is measurable from public sources and what requires synthesis. Both tiers run through the same drafter: the split records how directly source material maps to the rubric, not a difference in review process.

Tier 1 - Rubric-clear
6 dimensions where source material maps directly to the rubric
Biodiversity protection, ecosystem services, ecotourism infrastructure, community development, budget, and fundraising performance. Source material here generally maps more directly to the rubric anchors: disclosed numbers, published trend data, registry entries. That is not the same as unambiguous - species counts move with survey effort, and community spending is not by itself community benefit - but the drafter has clear evidence to grade against: operator and authority reporting, ecological and forest datasets, financial filings, and park-authority disclosures.
Tier 2 - Synthesis-required
4 dimensions requiring synthesis and review
Security and anti-poaching, ecotourism revenues, governance and accountability, and political and community leadership. Data here is either uneven across operators or inherently qualitative. Canopy drafts scores with full source trail using LLM-assisted synthesis. The editorial review layer that gates these scores is built but currently unstaffed - see Editorial review section below.
1.4 The ten dimensions
Dimension Tier Primary data sources
Biodiversity protection T1 Species-status registries, operator wildlife and census records, and population-trend statements in operator reporting
Ecosystem services (carbon and water) T1 Satellite forest-cover and carbon datasets, water-catchment studies, and carbon-registry records where the PA hosts a project
Security and anti-poaching T2 Operator anti-poaching reporting, wildlife-crime statistics, and news coverage of major incidents
Ecotourism infrastructure T1 Accommodation and access inventories, park-authority listings, travel-platform data, and road and airstrip accessibility
Ecotourism revenues T2 Park authority financial disclosures where available, operator annual report income statements, state tourism ministry data
Community development T1 Annual report disclosures on schools, clinics, employment, and skills programmes. Operator community investment data. Independent evaluations where available
Political and community leadership T2 Government and community statements, conflict and dispute history, leadership continuity, press coverage of governance
Governance and accountability T2 Oversight and institutional quality of the managing authority, accountability and transparency, audit findings, adherence to rules and mandates. Qualitative; drawn from published reviews, governance reporting, and press coverage
Budget T1 Operator audited accounts, parastatal and authority annual reports, and sub-national budgets where available. Several public-finance channels (national appropriations, donor and multilateral project finance, trust funds, concession and carbon revenues) are not yet incorporated and are a roadmap item. Where per-park budgets are not published, budget is scored on the managing operator's financial strength rather than left null; genuine weakness is floored to a low score with the reason recorded.
Fundraising performance T1 Audited income growth, grants received, announced multi-year funding commitments, and donor diversification
1.5 Scoring logic

Each dimension is rated 1 to 5 against a published rubric, where 5 represents best in class and 1 represents severely deficient. The composite score is a weighted mean of the ten dimensions: conservation outcomes carry the most weight, the institutions that produce them next, financial durability and community benefit next, and tourism least, since it is a means rather than an end. The exact weights are coarse bands set by judgement, not fitted to results; weight-sensitivity testing shows the ranking is near-invariant to the precise values (minimum Spearman rho 0.9968 across all single-dimension perturbations). Where a dimension cannot be scored due to insufficient data, it is marked "insufficient data" and excluded from the composite rather than scored conservatively.

Weighting bands

The ten dimensions sit in four weighting bands, set by judgement to reflect what each measures. Conservation outcomes carry the most weight, the institutions that produce them next, financial durability and community benefit next, and tourism least, since it is a means rather than an end. The bands sum to 1.0 and the composite renormalises over the dimensions actually scored for a given PA. (These weighting bands are distinct from the Tier 1 and Tier 2 data-availability split described above.)

BandDimensionsRelative weight
OutcomeBiodiversity protection, ecosystem servicesHighest
InstitutionalSecurity and anti-poaching, governance and accountability, political and community leadershipHigh
Sustainability and communityBudget, fundraising performance, community developmentModerate
MeansEcotourism infrastructure, ecotourism revenuesLowest

These are coarse bands, not values fitted to results. Weight-sensitivity testing confirms the ranking is near-invariant to the precise figures: across every single-dimension perturbation of plus or minus 5 percentage points, with the residual redistributed proportionally, the minimum Spearman rank correlation against the baseline is 0.9968 (mean 0.9986), with at most 12 PAs crossing a score band.

Operator-level conditioning on the financial dimensions

Three dimensions are financial: budget, fundraising performance, and ecotourism revenues. Per-park budgets are almost never published, so scoring them on park-level disclosure alone would floor most parks to the same low value and tell the reader little. Instead these three are conditioned on the financial strength of the managing operator across a reference set of the major managing operators, so a park run by a well-capitalised, well-governed operator is read differently from one run by a fragile one. The same multi-sample draft-and-critique method runs on these dimensions as on the other seven.

Where the evidence genuinely shows weak or absent financing, the dimension is floored to a low score carrying an explicit recorded reason rather than a null, distinguishing a channel that is simply silent from funding that is genuinely thin. On budget and fundraising, non-disclosure by an operator with an independently-evidenced durable funding base is read as adequate at portfolio level and floored at 3 rather than 2, capped there until park-level figures corroborate more; genuine underfunding still scores 2 or below. Most floored scores across the financial dimensions reflect non-disclosure rather than confirmed absence, with a portion sitting at 3 on operator evidence (22 July 2026 snapshot, rubric v1.1). This keeps the financial dimensions honest about what is unknown without treating an unknown as a failure.

Of the 200 rated PAs, 199 are scored under this current ten-dimension tier-weighted rubric. One, boucle-du-baoule (Mali), retains a prior-version nine-dimension score pending re-scoring and is disclosed as the single exception.

Score interpretation
4.60 - 5.00exemplary across most dimensions, a reference case for the sector
4.20 - 4.59strong performance, well-run operation with minor gaps
3.80 - 4.19above average, credible operation
3.40 - 3.79solid but with identifiable weaknesses
3.00 - 3.39mixed performance, meaningful operational gaps
2.50 - 2.99significant concerns across multiple dimensions
Below 2.50severely underperforming on core conservation outcomes
1.6 How missing data is handled

A recurring question is what happens to a protected area that discloses little or nothing: is it rewarded, punished, or left untouched. The short answer is that silence never produces a high score. Missing data is resolved in one of two ways depending on the dimension, and the guiding principle is the one stated throughout this page, PACE rewards disclosure. The two mechanisms below are the whole of the policy. There is no third path and no imputation to a midpoint.

Mechanism 1: exclude and renormalise (the seven non-financial dimensions)

Biodiversity protection, ecosystem services, security and anti-poaching, governance and accountability, ecotourism infrastructure, community development, and political and community leadership are scored only where the published evidence lets the drafter place the dimension against the 1–3–5 anchors. Where there is genuinely no basis to do so, the dimension is set to a null "insufficient data" flag: it is not scored, not imputed to a midpoint, and dropped from the composite. The weighted mean then renormalises over the dimensions actually scored, so the weights again sum to 1.0. A null is therefore neither rewarded nor punished, it is excluded, and the composite is transparent about resting on fewer inputs.

In practice this is rare. Across the 200-PA cohort, only two of the roughly 2,000 dimension scores are null today, both on a single park, Abumonbazi in the Democratic Republic of the Congo, where biodiversity and ecosystem services could not be defensibly scored, whose composite is the renormalised mean of its remaining eight dimensions. The drafter can almost always locate a park somewhere against the anchors from public evidence, so a true null is the exception, not a routine consequence of thin disclosure.

Does dropping a dimension keep the scores comparable?

Renormalising raises a fair question: if one park is averaged over ten dimensions and another over eight, are the two composites still a like-for-like comparison. The honest answer is that for 199 of the 200 PAs it is exactly like-for-like, and for the one park with a genuine gap it is very nearly so, with a caveat we state openly rather than bury.

Two structural facts keep the cohort comparable. First, almost every park is scored on all ten dimensions, so their composites are built from the same ingredients under the same weights, the same recipe applied to every park. Second, the three financial dimensions can never be blank, because non-disclosure there is floored rather than dropped (to a 2, or a 3 where a strong operator backs it), as Mechanism 2 below sets out. So the only dimensions that can ever fall out of a composite are the seven non-financial ones, and in practice that has happened just twice, at a single park, out of roughly 2,000 dimension scores.

What renormalising assumes is worth stating plainly. When a dimension is dropped and the remaining weights are rescaled to sum back to 1.0, the composite is implicitly treating the dimensions that were scored as representative of the park as a whole, in effect, that the park is about as strong on the missing dimension as on the ones we can see. That is a deliberately neutral assumption, not a measured fact. It neither invents a number nor assumes the worst; it simply declines to guess. A fully-scored park needs no such assumption, so a park scored on eight dimensions is not a perfect like-for-like with one scored on ten: its number rests on a smaller, slightly different set of ingredients.

For the one affected park, Abumonbazi, that assumption carries a specific edge worth naming. The two dimensions it lacks, biodiversity and ecosystem services, are the outcome band: they carry the highest weight of any single dimension in the model, the largest combined share of any band. With those two removed, Abumonbazi's composite is a renormalised average of its remaining eight dimensions, its security, governance, tourism infrastructure, tourism revenue, community, budget, fundraising, and leadership, and does not include the ecological outcomes that ordinarily carry the most weight. Its score is therefore answering a subtly narrower question than a fully-scored park's, and is best read with that in mind. This is a limitation of that one record, disclosed by name, not a property of the cohort.

Renormalising is chosen because every alternative is worse. Imputing a midpoint 3 would manufacture evidence that does not exist. Scoring the gap as a zero or another low value would punish a park for a reporting or research gap rather than for anything it has actually done, and would quietly rank a well-run but under-documented park below a genuinely worse one. Averaging over what is truly known invents nothing and penalises nothing; the price is the small, bounded loss of comparability described above, which is why it is confined to genuine nulls and disclosed at the level of the individual park rather than applied silently across the cohort.

Mechanism 2: a floored, explained low score (the three financial dimensions)

Budget, fundraising performance, and ecotourism revenues are treated differently. On these three, non-disclosure is not permitted to be null. It carries a floor of 2, below adequate, "significant concerns", with an explicit recorded reason that distinguishes whether the low score reflects funding that is genuinely thin or a channel that is simply silent. On budget and fundraising the floor rises to 3, adequate at portfolio level, where the managing operator has an independently-evidenced durable multi-donor funding base: a well-funded operator that reports at group rather than park level is read as adequate-but-opaque, not as under-resourced. It never rises above 3 on operator strength alone, a 4 or 5 requires park-level corroboration, and any park whose own evidence shows genuine underfunding stays at 2 regardless of operator. Ecotourism revenues stay park-specific and are not lifted by operator strength. Money that is not reported is therefore still scored, never waved through, and the reason is always attached so the reader can see why the score sits where it does.

What happens to each dimension when data is missing
Dimension If the underlying data is missing Effect on the composite
Biodiversity protectionNull: "insufficient data"Dropped; composite renormalises over the dimensions scored
Ecosystem servicesNull: "insufficient data"Dropped; composite renormalises over the dimensions scored
Security and anti-poachingNull: "insufficient data"Dropped; composite renormalises over the dimensions scored
Governance and accountabilityNull: "insufficient data" (see note below)Dropped; composite renormalises over the dimensions scored
Ecotourism infrastructureNull: "insufficient data"Dropped; composite renormalises over the dimensions scored
Community developmentNull: "insufficient data"Dropped; composite renormalises over the dimensions scored
Political and community leadershipNull: "insufficient data"Dropped; composite renormalises over the dimensions scored
BudgetFloored to a low 2, reason recordedCounts as a low 2; never null
Fundraising performanceFloored to a low 2, reason recordedCounts as a low 2; never null
Ecotourism revenuesFloored to a low 2, reason recordedCounts as a low 2; never null

Note on governance. A park with no delegation agreement is not treated as missing data. State-direct parks, which have no delegation deal by design, are scored on the managing agency's own institutional quality, independence, and accountability rather than auto-penalised. The drafter only reaches for a low non-disclosure score where a delegation arrangement exists but its terms are not made public.

Why the financial dimensions are floored rather than excluded

Per-park budgets are almost never published. If the three financial dimensions were scored on park-level disclosure alone, most parks would floor to the same low value and the score would tell the reader nothing. So budget and fundraising are additionally conditioned on the financial strength of the managing operator across a reference set of the major managing operators: a park run by a well-capitalised, well-governed operator is read differently from one run by a fragile one, and sits at 3, adequate at portfolio level, on operator evidence even without park-level disclosure, capped there until park-level figures corroborate more. The capped-low 2 bites where neither park disclosure nor operator strength supports a higher score, and it holds at 2 where a park's own evidence shows genuine underfunding despite a strong operator. The breakdown, on the 22 July 2026 snapshot (rubric v1.1):

Financial dimension Scored at capped-low 2 Undisclosed Genuinely thin
Budget14816914
Fundraising performance1091145
Ecotourism revenues1539381

Across the three financial dimensions that is 376 undisclosed and 100 genuinely thin, with no missing values. Of the 376 undisclosed cells, 49 now sit at 3 where an evidenced operator funding base supports portfolio-level adequacy and 326 remain at the capped-low 2. Every low financial score is explained rather than blank, and 181 of the 200 PAs still carry at least one financial dimension at the capped-low 2, a direct measure of how thin public conservation finance disclosure is.

So: rewarded or punished?

Neither outcome of silence is a reward. On the seven non-financial dimensions, a genuine data gap is excluded and the composite is built from what is known, neutral, and vanishingly rare in practice. On the three financial dimensions, non-disclosure is scored as a low 2, a mild penalty, always tagged with its reason. A further guard runs across every dimension: where the case for a top mark of 5 rests only on operator-stated or otherwise uncorroborated figures, the ceiling is 4, so a park cannot earn the highest score on its own unverified claims. The asymmetry is deliberate and disclosed. Transparent operators score well because the evidence to score them exists, while opaque parks either score low on finance or carry an insufficient-data flag on ecology. Nowhere does non-reporting produce a high score.

1.7 How PACE handles counterintuitive signals

Some data points mean the opposite of what they appear to at first glance. The clearest example is forest cover. Deforestation is plainly bad, so the reflexive reading is that a rise in forest cover must be good. But a gain can be the establishment of a monoculture plantation grown for eventual logging, a timber crop that may even have replaced natural habitat. Read naively, "forest went up" would reward exactly the wrong thing. How PACE deals with this case, and the broader class it belongs to, follows.

Why the mechanical version of the trap never fires

The "more trees, therefore a better score" reflex is only possible in a system that ingests a metric and maps it directly to a score. PACE is not that system. As set out in section 1.2, there are no automated numeric feeds: no forest-gain figure, no carbon-flux number, and no visitor count feeds any dimension mechanically. Every ecosystem and biodiversity score is produced by the structured large-language-model pipeline reading the underlying source material and judging it against a rubric written in terms of conservation quality rather than raw quantity. Because nothing maps a forest-area delta to a score automatically, the naive reflex never gets the chance to fire on its own.

Why the rubric itself resists it

The anchors reward habitat quality and verified service delivery, not area of tree cover. The ecosystem-services top score requires independently verified, quantified delivery of major services rather than raw tree-cover area; the biodiversity top score requires corroborated stable or recovering populations of key species and intact or expanding priority habitat. A plantation grown for logging is neither priority habitat, forest retention, nor a durable carbon service, it is a crop. It cannot satisfy the anchor on its face. Worse, where a gain reflects natural habitat converted to plantation, that reads as degradation under the biodiversity anchor for significant habitat loss, not as a gain. The rubric never asks "how many trees"; it asks "what kind of habitat, delivering what verified service, for how long", and a timber crop answers those questions badly.

The general pattern

Forest gain is one instance of a wider class: a surface signal whose obvious reading is wrong. PACE meets each the same way, by scoring against a quality-based anchor and a corroboration standard rather than the surface number.

Surface signal The naive reading How PACE reads it
Forest cover has increased "More trees, so conservation is working." Rewards intact or expanding priority habitat and verified, quantified service delivery, not canopy area. A plantation grown for logging is a crop, not habitat; natural habitat converted to plantation reads as degradation.
No poaching incidents on record "Zero poaching, so security is excellent." The security anchor requires a corroborated low poaching rate backed by a resourced, effective ranger force. An empty record can mean no monitoring rather than no poaching; absence of data is not read as success.
A large budget or revenue figure "Well funded, top marks." Where the figure is only operator-stated, the evidence cap holds the ceiling at 4; a 5 needs independent corroboration. A one-off grant is weighed differently from a durable, diversified multi-year base.
Extensive new tourism infrastructure "Lots of lodges, so the park is thriving." Tourism is scored as a means, not an end, and carries the lowest weight. Displacement or community conflict from rapid commercialisation surfaces in the community and leadership dimensions, which can pull the composite the other way.
The park has a formal effectiveness assessment "It has been assessed, so it must be good." Whether a park has ever been assessed is treated as coverage, not score (see 3.4). Rewarding "has been studied" would introduce a survey-effort bias, since globally only about one PA in ten has any assessment on record.
The catch mechanisms when a first read is naive

A single drafter can still misread a source, so three guards sit behind every score. First, multi-sample critique: each dimension is independently drafted and then reviewed by several adversarial critics, with more critics escalated on a split or a boundary. One critic's explicit job is to challenge a too-generous read, "this gain is a plantation, not recovery" is exactly the correction it exists to make, and the great majority of critique-driven score changes push scores down, so the pass functions as a brake on over-scoring rather than a source of noise. Second, the evidence cap: a 5 built only on operator-stated or otherwise uncorroborated figures is held to 4, so an unverified "forest area increased" claim cannot drive a top score. Third, source-quality weighting: the drafter is instructed to privilege primary and independently audited sources over marketing pages and single news stories, so an operator publicising "reforestation" without independent backing is discounted.

The honest limits
This is a disclosed weak spot, not a solved problem. Ecosystem services carries the lowest test-retest reproducibility of any dimension measured, Spearman rho 0.47 against roughly 0.86 for the composite, and rubric tightening on it is the highest-priority fix on the PACE roadmap. Both figures rest on a single 20-park blind re-draft, which is a small sample, and they are now published with intervals rather than as bare point estimates: the composite runs 0.61 to 0.97 and ecosystem services runs from about zero to 0.82 (percentile bootstrap over parks). Every dimension interval overlaps every other, so 0.47 is the lowest reading and a direction to work on, not an established rank. Two further limits of that run are worth stating: it predates the tenth dimension, so governance has no reliability figure at all, and budget and ecotourism revenues had too few scored pairs in the sample (5 and 2) to support a correlation. The external check against the Hansen Global Forest Change dataset returned a null (rho +0.10, n=192), disclosed rather than withheld; the stated reason is instructive here, that a buffer taken from the park centroid samples regional landscape dynamics rather than inside-PA conditions, so raw forest signals near a PA are known to be an unreliable proxy. Critically, PACE runs no automated natural-forest-versus-plantation classifier: Canopy's satellite forest and vegetation pipelines feed the C-INTEL anomaly and news layer, not PACE scores. So the plantation nuance is caught only where the source evidence the drafter reads distinguishes it. Good sources, caught; thin sources, potentially missed. The design reduces the risk of the naive reading; it does not eliminate it.
The principle

PACE handles counterintuitive signals the way a careful analyst does, interpretive judgement against quality-based anchors, stress-tested by adversarial critics, an evidence cap, and source-quality weighting, rather than through mechanical metric-to-score rules. That is a genuine advantage for exactly the "the obvious reading is wrong" case: a rule-based system rewards the surface number, while a reasoning system can ask what kind of forest, delivering what service, for whom. The cost is that the score is only as good as the evidence and the reasoning behind it, which is why the weak dimension is quantified, the failing external test is published, and the residual risk is disclosed here rather than hidden behind a claim of objectivity.

1.8 Is PACE difficulty-adjusted?

Operators run protected areas in wildly different settings. Some manage collapsed parks in conflict-torn or fragile states, the Congo, the Central African Republic, post-conflict Mozambique. Others run well-resourced flagships in stable, wealthy countries. A fair question is whether PACE accounts for that gap in the degree of difficulty. The answer is deliberately split, and conflating the two halves is the usual mistake: the park score is not difficulty-adjusted, but the operator comparison is. The two do different jobs and are handled differently on purpose.

The park score is absolute, not handicapped

A park in the DRC is scored against exactly the same rubric as a park in South Africa. There is no degree-of-difficulty bonus baked into the 1–5 dimension scores. This is intentional. PACE answers a specific question, is this place being run well enough to merit long-horizon conservation capital, and a funder deciding where money goes needs the actual state of the park, not a difficulty-graded version of it. A well-intentioned but overwhelmed park in a collapsed state should read as struggling, because it is struggling, however heroic the effort behind it. Grading on a curve would tell the funder something false about the place.

The consequence is visible in the raw numbers. Across the cohort, African Parks-managed sites average 3.06 on the composite and state-run parks 2.38, but, as section 3.1 sets out, part of that gap reflects the parks each operator runs, not the management itself. African Parks tends to take over collapsed PAs in difficult countries; state-run sets include easier wins in stable ones. On the raw score an operator in a hard place is not handed a leg up. If anything, the difficulty drags the absolute score down, and PACE lets it, because the score is a statement about the park rather than a report card graded against expectations.

The operator comparison strips difficulty out

The place difficulty is handled is the operator-effect analysis in section 3.1, which asks the different question of which operator actually adds value. That analysis does not compare raw averages. It removes the difficulty of the parks each operator runs in two ways. First, country fixed effects: each park is compared only against other parks in the same country, so a DRC park is measured against other DRC parks and never against a South African one. This removes the entire "the Congo is harder than South Africa" gap by construction, because the model never compares across it. Second, park size held constant via log(area_km2), which on this run carries a coefficient of -0.04 per log unit, close to zero. Size was materially negative on earlier runs (-0.27 in June 2026) and the control stays in the specification regardless: dropping a control because it went quiet on one rescore would be fitting the specification to the data.

What survives those controls is an adjusted association: the score gap that remains once country and park size are held constant, which is not the same as the effect of changing the operator. That gap is large. On the 3 September 2026 estimation run, African Parks scores about 0.84 points higher than state management in the same country, 95% CI +0.67 to +1.01, clustered on 13 countries, and government–NGO partnerships come in at +0.53. The estimate is rebuilt from the PACE scores on the monthly score cycle, so it moves when they do; these figures are from the scores published for 2026-09-01. The difficulty-adjusted African Parks effect is in fact larger than the naive raw gap of 0.68 implies, because once you stop crediting them for easy parks and start comparing them only against peers running equally hard parks, their turnaround work in difficult countries counts in their favour rather than against them. So an operator in the Congo earns no bonus on the park score for the difficulty, but is given full credit for it in the analysis that asks whether they are good at the job.

Where institutional context does enter a dimension

One dimension carries a related, narrower adjustment. On governance and accountability, the rubric instructs that state-direct parks, those with no delegation agreement to an NGO by design, are scored on the managing agency's own institutional quality, independence, and accountability rather than being auto-penalised for lacking a delegation deal. This is a fairness adjustment for institutional context, not for country difficulty as such: a park is not marked down merely for sitting inside a state system, only for genuine weakness in the institution actually running it.

The wealth check

If PACE were quietly rewarding easy, wealthy countries, it would show up as a correlation between the composite and national income. The structural bias audit tested exactly that and found a null: the composite's correlation with log GDP per capita is Spearman rho +0.06, essentially zero. Richer or easier countries do not systematically float to the top of the raw scores. The one real structural bias the same audit surfaced and disclosed is a francophone–anglophone scoring gap of −0.42, but that is a language and source-availability artefact, less English-language material to score against, not a difficulty effect, and its roadmap mitigation is a French-prompt blind re-draft of the francophone sample.

The honest limit
The operator comparison controls for difficulty; it does not fully model selection. Which operators take which parks, and when, is not modelled, and reverse causation cannot be ruled out. African Parks may look strong partly because it selects turnarounds it believes it can fix. The +0.84 is the effect that remains after country and size controls, not a clean causal claim that swapping the operator would move a given park by that amount. The effect is also an average: individual sites vary widely within every operator class.
Summary
Question Difficulty-adjusted?
The park's PACE score (what a funder screens on) No. Absolute, same rubric everywhere. A hard place reads as hard.
The operator comparison (who adds value, section 3.1) Yes. Country fixed effects plus size controls compare like with like.
Governance dimension Partially. Judged on the managing agency's own quality; state-run parks not auto-penalised.

The design philosophy is straightforward: tell the funder the truth about the park, undiscounted, but when judging the manager, compare them only against peers running equally hard parks. An operator in the Congo is not rewarded with a higher park score for the difficulty, but is given full credit for it in the analysis that asks whether they run parks well.

Chapter 2: Stress-testing and validation

2.1 Where the validation lives

PACE is validated across six families: reliability, internal structure, robustness, stress, external benchmarking, and bias. So that every Canopy scoring model is held to one consistent standard, the full method and results for each are reported in the C-VAL validation matrix rather than duplicated here.

In summary, PACE clears its reproducibility, design-robustness and construct-validity checks; the one finding disclosed rather than passed clean is the francophone-anglophone scoring gap surfaced by the bias audit. The matched datasets, analysis scripts and method documentation remain published for review and replication at validation/pavis, validation/gfw, validation/ucdp, validation/kba, and validation/firms. The KBA match CSV contains PACE-side fields plus a derived per-PA proximity count only; raw KBA data is not redistributed in line with WDKBA Terms of Service Section 4.

The validation roadmap, rubric tightening, biome-stratified re-runs, PAME convergence, polygon-enabled re-runs and operator-type stratification, is tracked alongside the results in the C-VAL matrix.

Chapter 3: Findings and disclosures

3.1 Operator effect on PACE scores

Management matters. African Parks-managed sites score 0.84 points higher on the 1-5 PACE composite than state-managed sites in the same country, controlling for park size - the largest effect of any operator class with a substantial sample. Government-NGO partnerships add +0.53. Estimated 3 September 2026 on the scores published for 2026-09-01.

Operator effect on PACE composite score Effect vs state management in the same country, with park size held constant. Ten-dimension composite, n=185 across 31 countries, scored 2026-09-01. +0.5 +1.0 +1.5 0 (state baseline) African Parks n=24, 13 countries +0.84 p<0.001 95% CI +0.67 to +1.01 Peace Parks Foundation n=4, 2 countries +0.60 not calibrated 95% CI +0.48 to +0.72 Other NGO-led n=4, 2 countries +0.54 not calibrated 95% CI -0.45 to +1.54 Government-NGO partnership n=30, 16 countries +0.53 p<0.001 95% CI +0.32 to +0.73 SANParks n=8, 1 country +0.39 not calibrated 95% CI -0.46 to +1.24 Kenya Wildlife Service n=8, 1 country +0.30 not calibrated 95% CI -0.22 to +0.82 OLS: composite ~ operator type + country fixed effects + log(area). R-squared 0.60, adjusted 0.49. Clustered standard errors for multi-country operators; within-country standard errors where an operator sits in one country. Solid bars are published as significant at p<0.05. Hatched bars are not: either the interval crosses the state baseline, or it is marked NOT CALIBRATED because the operator spans fewer than 5 countries, at which point the cluster-robust interval is too tight to trust and the estimator cannot support a significance claim either way. Other NGO-led: a conservation NGO other than African Parks or Peace Parks running the park (e.g. Ol Pejeta, Makira). Government-NGO partnership: a state authority co-managing with a conservation NGO (e.g. North Luangwa with FZS, Okapi with WCS).

The naive comparison favours the standout operators by a wide margin: African Parks averages 3.06 on the composite, state-run parks 2.38. But part of that gap reflects the parks each operator runs, not the management itself. African Parks tends to take over collapsed PAs in difficult countries (CAR, DRC, post-conflict Mozambique). State-run parks include easier wins in stable countries. To strip that out, each park is compared only against other parks in the same country, with park size held constant. What's left is an adjusted association: the score gap that remains after those controls, which is not the same as the effect of changing the operator.

After those controls, African Parks sites score 0.84 points higher than state-managed sites in the same country, p<0.001 across 13 countries, and it is the strongest effect among operator classes with a broad footprint. Government-NGO partnership sites - a state authority co-managing with a conservation NGO - show a +0.53 lift, also significant across 16 countries. Those two are the only classes whose parks span enough countries for a cluster-robust interval to mean anything. The rest sit in fewer than 5: Peace Parks Foundation is +0.60 across 2 countries; Other NGO-led is +0.54 across 2 countries; SANParks is +0.39 across 1 country; Kenya Wildlife Service is +0.30 across 1 country. Their point estimates are reported, their intervals are not calibrated at that cluster count, and none of them is published as significant in either direction.

What this does not show: the PACE composite is built from Canopy's own ratings across ten dimensions, not external outcomes such as animal counts or deforestation rates. Selection is not modelled - which operators take which parks and when - and reverse causation cannot be ruled out. The effect is also an average; individual sites vary widely within every operator class.

Method and scope

DesignCross-sectional. One snapshot of the PACE dataset; no within-PA change over time.
ModelOLS regression of composite score on operator type, country fixed effects, and log(area_km2).
Sample185 PACE-rated PAs across 31 countries. Of the 200-PA cohort, 15 are excluded: 14 as the only rated park in their country (country fixed effect unidentified) and 1 carrying no operator label. Reference category: state management, n=106.
Standard errorsClustered at country level for operators spanning multiple countries. An operator confined to a single country cannot support country-clustered inference at all, since the cluster-robust estimator is undefined with one cluster, so it is tested with within-country standard errors. Beyond that, any operator whose parks span fewer than 5 countries is flagged as not calibrated and is not published as significant in either direction: the cluster-robust interval is too tight to trust at that cluster count. An earlier version drew this line at one cluster, which was not enough. Peace Parks has since moved from one country to 2, and under the old rule it would have been silently promoted to "clustered" and printed with a deceptively tight interval. On this run 2 of the 6 charted classes clear the threshold: African Parks and Government-NGO partnership.
Operator bucketsOperator field cleaned to merge "government" with "state" and classify by operator string. 7 operator classes are estimated against the state baseline. The chart shows the 6 with more than one park; a single private reserve (n=1) is in the estimation sample but not charted, since one park is not an interpretable class. Classes whose confidence interval crosses zero, and classes whose interval is not calibrated, are both drawn hatched and neither is published as significant.
What it cannot saySelection is unmodelled - which operators take which parks, and when. Reverse causation cannot be ruled out. Composite is built from Canopy's ten-dimension ratings, not from external outcomes such as wildlife counts or deforestation rates.
BaselineRun on the current ten-dimension composite over the 200-PA cohort. Supersedes the April 2026 nine-dimension run, on which the African Parks effect was +0.56. Estimates move with the monthly rescore: the 8 June 2026 ten-dimension run put African Parks at +0.95 and partnership at +0.55, and three rubric and cap revisions later this run puts them at +0.84 and +0.53. The order at the top is unchanged.
FitR-squared 0.60, adjusted 0.49. Full results held in the validation record.
DataThe published PACE dataset
PublishedFirst published 30 April 2026 on the nine-dimension composite. Re-estimated 3 September 2026 against the PACE scores published for 2026-09-01 (pace-v1.1), and rebuilt from pipeline/build_operator_effect.py on the monthly score cycle, so the figures above and the scores they describe move together.
3.2 Operator scores: aggregation and reliability adjustment

Each operator's score is the mean of the PACE composites of the protected areas it manages, rescaled to 0-100. Operators with fewer than three attributed parks are listed separately and not ranked: a one or two-park average is too thin a sample to place on a league table.

Among ranked operators, the headline order uses a reliability-adjusted mean rather than the raw mean. Each operator's mean is shrunk toward the all-operator average by an amount that depends on how many parks it covers, so a three-park portfolio is pulled toward the middle far more than a twenty-park one. The shrinkage strength k is the empirical-Bayes estimate, k = within-operator variance divided by between-operator variance, recomputed from the data each month (currently about 0.86) rather than hand-set. This stops a small portfolio topping the table on a thin, noisy sample. The raw mean, median and park count are kept and shown next to the adjusted score, so the adjustment is visible rather than hidden.

The order is not sensitive to the exact value of k. The largest portfolio leads for any k above roughly 0.12, and the empirical-Bayes value sits well past that point, so the result does not hinge on a chosen constant. The full k-sweep is recorded as a robustness artifact at validation/operator_shrinkage_robustness_2026-06-08, reported alongside the other model robustness tests in the C-VAL matrix.

Operator view and network view. A handful of operators run a park through a dedicated legal vehicle created for that one site, the Gonarezhou Conservation Trust at Gonarezhou, the Nouabale-Ndoki Foundation at Nouabale-Ndoki. By default each park is credited to the entity that manages it, so these vehicles are ranked in their own right and their parent NGO is credited only with the parks it manages directly. The network view is the second lens: it re-credits each dedicated vehicle to its parent NGO, so an operator whose estate is split across such vehicles, the Frankfurt Zoological Society and the Wildlife Conservation Society are the two current cases, is ranked at its full delegated footprint rather than being held below the three-park threshold by an accident of legal structure. It is a strict re-partition of the same parks: every park still belongs to exactly one network, so nothing is double-counted, and the reliability adjustment above is recomputed within the network partition. Only delegated or co-management relationships roll up. A funding or technical-partnership tie to a government-run park does not, so an NGO never inherits parks that a state authority manages.

3.3 Caveats
All inputs are LLM-sourced. PACE does not ingest numeric data feeds. Every score across all ten dimensions is generated by a structured large-language-model pipeline reading published source material and applying the rubric. Defensibility rests on bounded scope, published source trails per dimension, a fixed published rubric, and immutable monthly snapshots, not on automated numeric ingestion.
Data asymmetry is real. Transparent operators will score well because the data to score them exists. Opaque private reserves and state-run parks that publish nothing will either score poorly on Tier 1 dimensions or carry "insufficient data" flags. This is by design. PACE rewards disclosure.
Revenues are the hardest number. National park authorities sometimes publish ecotourism revenue (SANParks, Kenya Wildlife Service). Private reserves and partnership operators rarely do. Where data is absent, PACE flags this rather than estimates.
Leadership scoring is explicit editorial judgement. Unlike every other dimension, leadership cannot be reduced to a data point. PACE scores this dimension based on synthesised evidence across government statements, community reports, dispute history, and continuity. Every leadership score carries a source trail and Canopy stands behind the assessment.
Coverage before opinion. Coverage is 200 PAs where the full ten-dimension score can be defended, expanded from 38 at launch. Expansion is deliberate. PACE will not inflate coverage at the cost of rating quality.
PACE is not a substitute for due diligence. Ratings are a public reference for comparison and early screening. They are not investment advice, partnership recommendations, or a licence to skip direct operator assessment.
3.4 Independent assessment cross-check (GD-PAME)

Each protected-area page carries an Independent assessment panel recording whether the park has ever had a formal, third-party management-effectiveness assessment, by METT (the Management Effectiveness Tracking Tool) or one of roughly sixty other methods, and how recently. The source is UNEP-WCMC's Global Database on Protected Area Management Effectiveness (GD-PAME). Of the 200-PA cohort, 156 have at least one assessment on record (106 including METT specifically); the join is reproduced in the published cross-reference dataset.

This is a coverage signal and a cross-check, not a scoring input. Two reasons it is deliberately kept out of the PACE composite. First, the public GD-PAME records only that a park was assessed, method, year, and a source link; it does not publish the numeric METT/PAME scores, which are confidential. So there is no score here to fold in. Second, the fact of assessment is itself unevenly distributed (globally only about one PA in ten has ever been assessed), so treating "has an assessment" as a positive would reward parks for having been studied rather than for being well run, a survey-effort bias. The panel therefore sits alongside the PACE score as an independent reference: where an external assessment exists it corroborates coverage, and a park with no assessment on record is flagged as such rather than penalised.

3.5 Live satellite monitoring and confidence tiering

Every protected-area page carries a Current flags board, at the head of its Satellite feeds section, showing the current status of eight independent detectors on a range of cadences: fire (near-daily, FIRMS-derived) and forest-loss alerts (near-daily); vegetation health via NDVI and surface-water extent via NDWI (a slower satellite cadence); night-time lights and burned area (monthly); and buffer land-use change and biomass (multi-year trend measures). Each is flagged against its own recent-history or same-season baseline. The panel also carries one static measured layer that is not a live detector: a mining-footprint overlap, flagging where a mapped mining footprint from the Maus et al. (2022) global mining polygons (Sentinel-2, 2019 imagery) intersects the park. It is a one-time 2019 snapshot rather than a monitored feed, so a mine opened after 2019 will not appear and a quiet reading means no mapped 2019-vintage overlap, not no mining risk. The 162 parks with a real boundary polygon get a true polygon intersection; the 38 without one fall back to a 10km point-buffer proxy, shown at low confidence. Like the live feeds, it is non-scoring. Two further measured reads now sit alongside these detectors on each park page, both context only and non-scoring: the vegetation and water flags carry a fused rainfall-and-soil-moisture drought cross-check (CHIRPS rainfall as the drought input and NASA SMAP L4 root-zone soil moisture as the drought state, resolved into a single verdict rather than two votes and falling back to rainfall alone where a soil reading is missing), and a GEDI forest-structure read pools NASA's spaceborne-LiDAR footprints into one canopy-height-and-cover fingerprint per park (whole-mission, not change detection). Both are documented in full elsewhere, the drought cross-check in the C-INTEL methodology and the structure read in the C-CARBON methodology (Measured forest structure). The main PACE list and map carry a compact version of the same signal: a small dot next to each park's name, red if at least one of the four faster feeds (fire, forest, vegetation, water) is currently flagged and green if they are all quiet, so a reader scanning the full 200-park cohort can see live conditions without opening every page. The dot deliberately tracks only those faster feeds, which reflect current on-the-ground conditions; the monthly and multi-year feeds carry their full detail on each park's own page.

The monitoring panel is a live-data display, not a scoring input. As stated in sections 1.2 and 1.7, no automated numeric feed maps to a PACE dimension score: a fire or forest-loss flag on this panel does not, by itself, move a park's composite. It is a real-time cross-check the reader can weigh alongside a score that reflects the drafter's read of published sources as of its last critique pass. This is a deliberate division of labour: PACE answers "how is this park run, on the evidence," and the satellite panel answers "what is happening on the ground right now," and conflating the two would either make the composite jump on noisy short-term anomalies or smuggle an ungrounded automated judgement into a rubric-anchored score. There is one deliberate, narrow, human-gated exception, set out in section 3.6: a verified inside-boundary loss can, once an editor confirms it, hold one of the three physically observable dimensions down. That exception is one-directional and never automatic, so the principle still holds, the feed never moves a score on its own, a person does, on the measured evidence.

Each park also carries a confidence tier (High, Medium, or Low), shown next to the composite score. This is not a new judgement call: it is the existing critique-pipeline state, made visible. High requires the park to have passed critique, for that critique to have raised no unresolved parse failure, and for at least three sources to be attached to the record; Medium is a park that has been critiqued but falls short of that bar; Low is a park that has not yet been through the critique pass. As of this writing the 200-park cohort is 18 High and 182 Medium, with zero Low, since every published record has been critiqued at least once; Low exists in the schema for future coverage expansion into thinner-sourced parks, so the tier does not silently disappear as coverage grows.

Neither of these two additions, the monitoring panel or the confidence tier, changes a single PACE score. Both are additive, non-scoring fields, computed from data (satellite feeds) or process state (critique/sourcing fields) Canopy already had; the confidence tier reuses the evidence-cap logic already implicit in the critique pipeline rather than inventing a new signal, and the satellite panel reuses the monitoring dataset already wired into each PA's detail page. The one place where a satellite measurement does move a score, the human-gated measured cap, is a separate mechanism and is set out in section 3.6.
3.6 Measured caps

Sections 1.2 and 1.7 state that PACE has no automated numeric feed that maps to a dimension score, and section 3.5 describes the satellite panel as a live-data display that sits alongside a score without moving it. Measured caps are the one deliberate, narrow exception to that rule: the only place where a satellite measurement can change a published PACE score. The exception is tightly bounded. It applies to three dimensions only, it can only lower a score, and it never fires automatically. A person confirms every cap.

For the three physically observable dimensions, biodiversity, ecosystem services, and security, PACE's rubric judgement is checked against measured satellite evidence. Where an automated feed detects a verified loss inside a park's boundary (canopy disturbance from the GFW DIST-ALERT layer, a vegetation or open-water drop not explained by rainfall, an anomalous burn, new night-time lights, or an intersecting mapped mining footprint), the affected dimension is flagged for review and cannot score above adequate, a 3 on the 1–5 scale, once an editor confirms the loss. A dimension already scored 3 or below is never changed, and the seven dimensions that are money, institutions, and documents are never touched.

The check is one-directional. Measured loss can only hold a score down, never raise one, and it never credits satellite "gain," because remote sensing cannot tell recovering habitat from a timber plantation (the counterintuitive-signal problem of section 1.7). It is also inside-boundary only: a cap fires on loss measured within the park, never on external buffer pressure, so the score stays absolute and is not discounted for a park's neighbourhood, consistent with the absolute, not difficulty-adjusted, principle of section 1.8.

A flag is a trigger to look, not an automatic change. This is what keeps measured caps consistent with the no-automated-feed rule: the feed never moves a score, an editor does, on the measured evidence, in the same kind of human review used for the C-INTEL weekly digest. Before any cap applies, an editor confirms the loss is real and management-relevant, not a permitted community offset or a natural burn. Three safeguards sit under that judgement. Drought-driven vegetation and open-water drops are filtered out automatically against a rainfall (CHIRPS) record before they ever reach review. Every cap carries its sensor, date, and measured extent, so a third party can reproduce it from public imagery. And a stale or broken feed cannot trigger a cap: a feed that has gone quiet because it failed reads the same as a feed that is genuinely clear, so any flag whose source feed is marked stale is held back rather than applied. This is the silent-sentinel guard. Above-ground biomass, whose multi-year trend is too noisy and confound-prone to veto a score, is carried as scorecard context only and never caps.

A confirmed cap becomes the dimension's single final grade: it overwrites the score, the composite recomputes with the existing weighting, and the reason is recorded in the dimension's note, shown beside the grade on the park page. There is one score per park and no second number. At the 1 August 2026 refresh, the eligible triggers are acute inside-boundary loss plus the static mining-footprint overlap; of the 200 parks, 41 are flagged for an editor to review, and because the cap is one-directional, the caps that actually change a grade are the divergence cases, at launch all on biodiversity, where the rubric scored a physical dimension above adequate while the satellite measured a real inside-boundary loss. Security caps no parks at launch, because the mining-flagged parks already score security at or below adequate; that is expected. Measured caps are deliberately conservative: for the answer to "is PACE just an LLM," what matters is that a reproducible measurement can overrule the narrative where the two disagree, and that every such cap is auditable, not that the count is large.

What measured caps do not do. They do not re-score a park, do not raise any score, do not touch the seven non-physical dimensions, and do not run without a human. They are not a mechanical metric-to-score map, the trap section 1.7 exists to avoid: the rubric still produces the score, and a confirmed measurement can only pull it down where the ground contradicts the narrative. A review queue surfaces candidate divergences and a separate confirmed-cap step applies only those an editor has signed off, with every confirmation logged. Sustained multi-cycle decline trends are being developed as an additional review signal and are not part of the launch cap set.
3.7 Land-use conversion detector

The Conversion card on each protected-area page answers one question: is the land inside the park being cleared or turned over to crops? It is annual, because a fire is an event and conversion is a multi-year trend. Standard satellite land-cover maps read dry grass, fire scars and seasonally bare ground as new farmland, so a cropland flag on its own is treated as a suspicion, not a finding: every flag is checked against a second, independent reading of the same ground, which looks for the bare, tilled soil that real cropland leaves behind and dry grass does not. When the evidence is thin the card abstains and shows nothing, rather than a false "no conversion."

Context only, never a score. Unlike the measured caps of section 3.6, the conversion detector is never a PACE input: it raises no alert and cannot move or cap any dimension score, in either direction. It is a slow, structural cross-check to weigh alongside a score. The method of record, with the exact detectors, thresholds, cross-checks and validation results, is the Land-use conversion entry in the Keystone Watch monitoring methodology on Canopy's partner portal. This page does not restate its figures, so the two cannot drift.

Chapter 4: Governance and change control

4.1 Editorial review and correction

PACE draft scores are generated by a structured large-language-model pipeline consulting the relevant source material per dimension and producing dimension scores, summary copy, and a sources-consulted trail per PA. The pipeline writes drafts into a separate file; nothing reaches the live site without an explicit promote step. An editorial review layer is specified between drafter and promote, with particular focus on Tier 2 dimensions where synthesis is required.

That layer is not staffed. Drafts are promoted to live data after automated parse validation and the automated critique-and-revise pass; no person signs off on a score between drafter and publication. The architecture is in place and staffing it is a function of platform stage. Read the process above as what PACE is specified to run, not as what gates a published score today.

Every score carries a source trail and a scoring date. Corrections are welcomed - operators, funders, or researchers spotting a factual error should email us.

4.2 Update cadence

PACE scores are refreshed when underlying data changes. Annual report releases, major news events, leadership transitions, and registry updates all trigger review of affected dimensions. A scoring date is displayed on each PA record. No PA score is published older than 18 months without re-review.

4.3 Version history

v1.31 - effective 3 September 2026. The operator effect is now rebuilt from the scores instead of being carried forward by hand. Sections 1.8 and 3.1 published an 8 June 2026 estimate: African Parks +0.95, government-NGO partnership +0.55, R-squared 0.65. There was no builder behind those figures. The chart was hand-authored inline SVG whose bar widths and whisker coordinates had only ever been rewritten by a one-shot patch script, so the estimate sat still through three months of monthly rescores (the v1.26 rubric change, the v1.26 measured caps and the v1.28 anchoring pass). Re-estimated on the same model against the scores published for 1 September 2026, African Parks is +0.84 (95% CI +0.67 to +1.01, clustered on 13 countries), partnership +0.53, R-squared 0.60, and the size control has fallen from about -0.27 to -0.04. Every qualitative conclusion is unchanged. Three further corrections travel with the rebuild. (1) The few-cluster rule from the same September run is now published: the 9 June 2026 correction switched single-country operators to within-country standard errors, but cluster-robust intervals are unreliable well above one cluster, and Peace Parks has since crossed from one country to two, which under the old rule alone would have silently promoted it to "clustered" and printed a deceptively tight interval. Any operator spanning fewer than five countries is now marked NOT CALIBRATED on the chart and is not published as significant in either direction; only African Parks and partnership clear it. (2) The chart states each class's country count and its full 95% interval, both of which it previously omitted. (3) Two layout defects are fixed: the SANParks and Kenya Wildlife Service intervals struck through their own labels, and the first legend line was cut off mid-sentence at the chart edge. The figures, the chart and the run date on all three surfaces are regenerated by pipeline/build_operator_effect.py on the monthly score cycle and guarded by tests/test_operator_effect_sync.py. Published figures change; the model, the weighting and the park scores do not.

v1.30 - effective 1 September 2026. Four claims tightened, no score changes. (1) The operator comparison in sections 1.8 and 3.1 described the residual after country and size controls as the part of the score "genuinely attributable to who is running the place". That overstates what the specification identifies: selection is unmodelled and reverse causation is not ruled out, both already disclosed in the same sections, so the residual is an adjusted association and is now described as one. The African Parks 95 percent confidence interval (+0.75 to +1.14, clustered on 13 countries) is now published alongside the point estimate rather than the p-value alone, and both now carry the 8 June 2026 run date: the operator regression is re-run on demand rather than on the monthly score cycle, so it lags the current scores and a reader could not previously tell. (2) Tier 2 renamed from "Editorial-assisted" to "Synthesis-required". v1.4 reframed the tier split as rubric clarity rather than process, and both tiers have run through the same drafter since; the old label survived that reframe and implied a human editorial step that the tier does not carry. (3) The Tier 1 description claimed its source material "is unambiguous". Replaced: Tier 1 source material maps more directly to the rubric anchors, which is not the same as unambiguous, and the two standing counterexamples (species counts moving with survey effort, community spending not being community benefit) are now stated in the tier description. (4) Section 4.1 asserted an editorial review layer and then retracted it in the next sentence under an "As of April 2026" stamp that had gone five months stale. Restated once, in the present tense, with no date to rot: the layer is specified, it is not staffed, and no person signs off on a score between drafter and publication. (5) The test-retest reliability figures are now published with bootstrap intervals and with the limits of the run that produced them. The composite 0.86 and the ecosystem services 0.47 were stated as bare point estimates; they come from one 20-park blind re-draft computed 7 May 2026, and the intervals are wide enough (composite 0.61 to 0.97, ecosystem services about zero to 0.82) that the dimension ordering they were used to assert is not separated. That run also predates the tenth dimension, so governance has never been reliability-tested, and two dimensions had too few scored pairs in it to support a correlation. The lite summary on the PACE page carries the same correction. Wording and disclosure only; no methodology, weighting, or score changes.

v1.29 - effective 13 August 2026. Section 3.5 extended to record two further measured reads now shown on each PA page alongside the eight live detectors and the static mining overlap, both context only and non-scoring: a fused rainfall-and-soil-moisture drought cross-check on the vegetation and water flags (CHIRPS rainfall as the drought input, NASA SMAP L4 root-zone soil moisture as the drought state; falls back to rainfall alone where a soil reading is missing), and a GEDI forest-structure read (NASA spaceborne LiDAR pooled into one canopy-height-and-cover fingerprint per park, whole-mission rather than change detection). Neither is new data; both are already documented in full, the cross-check in the C-INTEL methodology and the structure read in the C-CARBON methodology (Measured forest structure). This entry only records their disclosure on the PA satellite panel. Presentation and disclosure only; no methodology, weighting, or score changes.

v1.28 - effective 11 August 2026. Month-to-month re-scoring is now anchored to the prior published snapshot. At each re-score the drafter receives the previous month's composite and per-dimension scores as an explicit baseline and reproduces them by default; a dimension moves only when tied to specific new dated evidence, and a single-scale-point move with no cited evidence is held at the prior value. The one-point deadband is calibrated on the first full month-over-month diff, 1 July to 1 August 2026, where 27 percent of all dimension scores had moved by exactly one point, near-symmetrically up and down, the signature of re-rating variance rather than events. Every promoted score now carries a per-dimension delta record diffed against the prior snapshot, each change stamped with its evidence-tied reason and each held move recorded as held. This changes how future re-scores are produced and disclosed; it does not re-baseline the current published scores. Section 4.4 updated.

v1.27 - effective 8 August 2026. New section 3.7 (Land-use conversion detector) added. Documents the interior land-use conversion detector behind the Conversion card on each PA page: an interior Dynamic World crops-and-built screen, a Sentinel-2 spectral-unmixing de-confuser (photosynthetic / non-photosynthetic vegetation / bare soil) with a low-bare grass guard, a multi-year persistence gate, and a lower-confidence Sentinel-1 radar fallback. Restates that this feed is context and corroboration only: never a PACE input, it raises no alert and caps no score, and it abstains (insufficient data, the card is not shown) rather than reading a false zero. Summarises the Keystone Watch land-use methodology (version 1.8, 8 August 2026), the maintained method of record. Presentation and disclosure only; no methodology, weighting, or score changes.

v1.26 - effective 22 July 2026. Rubric raised to pace-rubric-v1.1: an operator-aware non-disclosure floor on the budget and fundraising dimensions (issue #54). Previously all non-disclosure on the three financial dimensions floored to a flat capped-low 2, which scored a well-funded operator that reports at group rather than park level identically to a genuinely cash-strapped one. Under v1.1, park-level non-disclosure by an operator with an independently-evidenced durable multi-donor funding base (audited group or consolidated accounts, public financial filings, a documented multi-year donor base or endowment) floors at 3, adequate at portfolio level, capped there without park-level corroboration; genuine, evidenced underfunding still scores 2 or below, and ecotourism revenues stay park-specific and are not lifted by operator strength. Applied to the current cohort this lifts 37 budget and fundraising cells across 29 parks from 2 to 3, with 26 non-disclosure cells reviewed and held at 2 where the park's own evidence shows genuine strain. The ranking is near-invariant: Spearman rho 0.9965 against the v1.0 composite, and six parks move up one score band. Unlike the recent presentation-only releases, this is a deliberate scoring change with a full re-baseline of the two affected dimensions across all PAs.

v1.26 - effective 1 August 2026. New section 3.6 (Measured caps) added, and section 3.5 reworded to carve out its one deliberate exception. For the first time a satellite measurement can move a published PACE score. A verified inside-boundary loss on one of the three physically observable dimensions (biodiversity, ecosystem services, security), once an editor confirms it in a C-INTEL-style human review, caps that dimension at adequate (3) and the composite recomputes. The cap is one-directional (it can only lower a score already above the ceiling, never raises one and never credits satellite gain), inside-boundary only, drought-filtered against rainfall data, and guarded so a stale or broken feed cannot cap (the silent-sentinel guard); biomass is context only. It stays consistent with sections 1.2 and 1.7 because it is human-gated: no automated feed moves a score, a person does, on measured evidence. Landed at the 1 August refresh boundary so all scores move together. Unlike the additive 3.5 additions, this release changes scores: the confirmed divergence parks show a capped dimension grade and a lower composite. Sustained multi-cycle trend caps are deferred to a later release.

v1.25 - effective 21 July 2026. Section 3.5 extended to document a ninth entry on the per-PA Satellite monitoring panel: a static mining-footprint overlap (Maus et al. 2022 global mining polygons, Sentinel-2 2019 imagery; PANGAEA, CC-BY-SA-4.0), flagging where a mapped mining footprint intersects the park. Distinct from the eight live detectors, it is a one-time 2019 snapshot, not a monitored feed; 162 parks use a true boundary-polygon intersection and 38 without a boundary use a 10km point-buffer proxy shown at low confidence. 12 of 200 parks carry a nonzero overlap. Explicitly non-scoring, disclosed the same way as the other satellite additions. Presentation and disclosure only; no methodology, weighting, or score changes.

v1.24 - effective 19 July 2026. New section 3.5 (Live satellite monitoring and confidence tiering) added. Documents, for the first time, the existing per-PA Satellite monitoring panel (eight feeds: fire, forest-loss, vegetation, water, night-time lights, burned area, buffer land use, biomass) and extends it with a compact live-signal indicator on the main PACE list and map, tracking the four faster feeds (fire, forest, vegetation, water). Also introduces a confidence tier (High/Medium/Low) per PA, computed from existing critique-pipeline signals, current cohort 18 High / 182 Medium / 0 Low. Both additions are explicitly non-scoring, disclosed the same way as the GD-PAME cross-check in section 3.4. Presentation and disclosure only; no methodology, weighting, or score changes.

v1.23 - effective 2 July 2026. New section 1.8 (Is PACE difficulty-adjusted?) added, setting out the deliberate split: the absolute park score carries no degree-of-difficulty bonus (a funder needs the true state of the park), while the section 3.1 operator comparison strips difficulty out via country fixed effects and a size control, yielding the +0.95 African Parks and +0.55 partnership effects. Also documents the governance institutional-context adjustment, the GDP-per-capita null (rho +0.06) confirming easy/wealthy countries are not rewarded, and the unmodelled-selection limit. Presentation and disclosure only; no methodology, weighting, or score changes.

v1.22 - effective 2 July 2026. New section 1.7 (How PACE handles counterintuitive signals) added, using forest-cover gain (which can be a logging plantation rather than habitat recovery) as the worked case. Sets out why the mechanical metric-to-score trap never fires given PACE has no numeric feeds, why the quality-based rubric anchors resist it, a table of five counterintuitive signals against how PACE reads each, the three catch mechanisms (multisample critique, evidence cap, source-quality weighting), and the honest limits (ecosystem_services test-retest rho 0.47, the Hansen null, and the absence of any automated plantation-versus-natural classifier in PACE). Presentation and disclosure only; no methodology, weighting, or score changes.

v1.21 - effective 2 July 2026. New section 1.6 (How missing data is handled) added, setting out the full missing-data policy per dimension: the seven non-financial dimensions drop to a null insufficient-data flag and renormalise out of the composite, while the three financial dimensions floor non-disclosure to a capped-low 2 carrying a recorded reason. Worked example disclosed (Abumonbazi, DRC, the only park with null dimensions: 2 of ~2,000 dimension scores). The section also treats cross-park comparability directly, setting out what renormalisation assumes and its specific effect on Abumonbazi, whose two null dimensions are the heavily-weighted outcome pair, and why renormalising is preferred to midpoint imputation or a low-score penalty. Same release, the live non-disclosure counts in section 1.5 were refreshed to the 1 July 2026 snapshot, correcting stale figures carried since v1.19. Presentation and disclosure only; no methodology, weighting, or score changes.

v1.20 - effective 6 June 2026. Coverage expanded from 198 to 200 PAs (162 RWF keystones plus 38 Canopy editorial additions). Two African Parks reserves added: Mangochi Forest Reserve (Malawi), managed as a single unit with Liwonde, and Siniaka-Minia (Chad), part of the Greater Zakouma Ecosystem; both scored under the current ten-dimension rubric. Kundelungu National Park (DRC) re-credited from state to African Parks following the October 2025 ICCN co-management mandate. African Parks operator coverage is now 24 of its 24 managed parks. No methodology or weighting change, and no change to existing scores.

v1.19 - effective 4 June 2026. Governance and accountability added as a tenth dimension (Tier 2 data-availability, institutional weighting band), carried by 197 of 198 PAs. The composite moved from an equal-weight mean of nine dimensions to a tier-weighted mean of ten, with the outcome band carrying the most weight and tourism the least, renormalised over the dimensions actually scored. All ten dimensions are now multisample-voted (independently drafted then reviewed by several adversarial critics). The three financial dimensions additionally use operator-level conditioning, with non-disclosure floored to a low score carrying a recorded reason in place of nulls; these dimensions were re-scored under this method across 88 operator-matched parks. Weight-sensitivity was re-run on the tiered ten-dimension scheme: minimum Spearman rho 0.9968, mean 0.9986, at most 12 PAs crossing a band (n=197, boucle-du-baoule excluded). One PA, boucle-du-baoule (Mali), retains a prior-version nine-dimension score pending re-scoring. Scores changed across the cohort from the new weighting and the financial re-score; coverage unchanged at 198.

v1.18 - effective 21 May 2026. Subsection numbering added. Section headings now carry chapter.section numbers (1.1, 2.3, and so on) and the five external-validation benchmarks are numbered 2.6.1 to 2.6.5. Repeated Method, Result, and Verdict subtitles are left unnumbered. Presentation only; no content, figure, or number changed.

v1.17 - effective 21 May 2026. Four numbered chapter headings added for navigability (The model; Stress-testing and validation; Findings and disclosures; Governance and change control), each with an anchor id for direct linking. Presentation only; no section, content, figure, or number changed.

v1.16 - effective 21 May 2026. Statistical-tests and validation sections reorganised into two clusters mirroring the Canopy maths primer Chapter 6 structure. New top-frame section (How PACE is stress-tested) added, carrying a two-family summary card across the four robustness tests and five external benchmarks. Robustness cluster reordered to test-retest, weight sensitivity, subscore PCA, structural bias; external-validation subsections reordered to PAVIS, UCDP, KBA, Hansen, FIRMS to group the significant positives ahead of the null and the weak diagnostic. The standalone external summary card was folded into the top frame, and each robustness test now closes with a Verdict against an explicit threshold. Presentation only; no methodology, score, result, or figure changed.

v1.15 - effective 13 May 2026. Weight sensitivity section added after Scoring logic. For each of the nine dimensions, the equal-weighting baseline (each dimension weighted equally) is perturbed by +/-5pp with remaining dimensions renormalised; the composite is recomputed for all 199 PAs and Spearman rank correlation against baseline is measured. Minimum rho 0.9959 across all 18 perturbations (worst case biodiversity -5pp), mean rho 0.9977. Mean grade-band movement 14 PAs per perturbation; maximum 39 PAs (also biodiversity -5pp). Clears the CCS sensitivity benchmark (minimum rho >= 0.989). Equal weighting is now defended directly rather than indirectly via PCA. No score changes.

v1.11 - effective 7 May 2026. Fifth external validation added: NASA FIRMS active fire detections (VIIRS S-NPP, 2020-2024) versus PACE ecosystem_services. Headline rho -0.20 across 193 PAs (p = 0.006); 3.35 million fire detections within 50 km buffers, 98% of PAs exposed. Buffer sensitivity flat across 25-100 km. Unexpected secondary finding: same fire metric correlates more strongly with PACE biodiversity (rho -0.28, p < 0.0001) than with ecosystem_services. Combined with the GFW null and test-retest rho 0.47, this is now a third independent signal that ecosystem_services rubric tightening is the highest-priority methodology fix on the PACE roadmap. Summary card extended to five rows. No score changes.

v1.14 - effective 7 May 2026. Subscore PCA section added. Primary 7-dim PCA (n=170, drops budget and ecotourism_revenue as structurally sparse) and robustness 5-dim PCA (n=193, drops also fundraising and community_development) both run jointly. PC1 explains 56.2% (primary) and 57.9% (robustness); cross-spec difference 1.7pp. PC1+PC2 cumulative 70-76%. PC2 has interpretable structure as ecology-versus-operations axis (biodiversity and ecosystem_services on one pole, leadership and security on the other). Highest pairwise correlation leadership x security r = +0.69, below collinearity threshold. Multi-dim structure justified. No score changes. Concurrent: structural bias audit (v1.13) silently dropped 4 of 9 dimensions from per-dimension univariate output due to a dim-name mismatch; the composite-level OLS and the 5 dims that did report were correct, but per-dim coverage in the v1.13 published numbers is partial. Script fixed in v1.14; re-run produced identical headline.

v1.13 - effective 7 May 2026. Structural bias audit added as new methodology section between External validation and Test-retest reliability. Tests PACE composite and 9 dimensions against four structural regressors: log GDP per capita (WDI 2023), working language (4-way), log PA area, establishment year. GDP per capita null (rho +0.06), establishment year null (rho +0.05), PA size reverse direction (rho -0.27, large PAs score lower - methodologically defensible). Francophone-anglophone gap of -0.42 (p < 0.0001) disclosed; roadmap mitigation is French-prompt blind re-draft of francophone sample. Multivariate R2 = 0.16 - structural variables explain one-sixth of composite variance. No score changes.

v1.12 - effective 7 May 2026. UCDP buffer-sensitivity sweep extended from 4 to 5 buffers (10, 25, 50, 75, 100 km), matching KBA. Per-buffer Spearman rho now published explicitly: -0.18 at 10 km, -0.33 at 25 km, -0.37 at 50 km (peak), -0.34 at 75 km, -0.30 at 100 km. Inverted-U shape with peak at 50 km mirrors the KBA buffer test; defends choice of buffer against cherry-picking. Headline rho unchanged. No score changes.

v1.10 - effective 7 May 2026. External validation section reorganised. New summary card at the top shows all four benchmarks (PAVIS, Hansen GFW, UCDP, KBA) at a glance; each test now sits under a labelled sub-section with its dataset, method, result, scatter, and caveats. New "Validation roadmap" sub-section consolidates outstanding work. Presentation only; no methodology, score, or result changes.

v1.9 - effective 7 May 2026. Fourth external validation added: Spearman rho +0.44 across 193 PAs (p < 0.0001) on log count of Key Biodiversity Areas within 50 km buffer versus PACE biodiversity score. Benchmarked against the World Database of Key Biodiversity Areas (BirdLife International / KBA Partnership, March 2026 release, 2,111 African KBAs). 75% of PACE PAs have at least one KBA within their 50 km buffer; 329 KBA-PA proximity pairs captured. Buffer sensitivity sweep at 10/25/50/75/100 km confirms 50 km as the signal peak, mirroring the UCDP buffer-test shape. WDKBA used under non-commercial pre-approval from the KBA Secretariat; compliance trail at docs/kba_compliance.md. No score changes.

v1.8 - effective 7 May 2026. Fourth validation added: an internal test-retest reliability assessment. Composite Spearman rho is 0.864 across 20 blind-redrafted PAs (drafter plus critique-and-revise) and 0.883 on the drafter alone. Two-stage attribution is published: 18 of 20 critique-driven score moves were downward, indicating the critique pass functions as calibration of a +0.25 drafter generosity bias rather than as additional noise. ecosystem_services flagged as a v1.8+ rubric tightening target (rho 0.47). Three reproducibility tiers documented per dimension. Sample, retest scores, and analysis scripts published in validation/test_retest/. No score changes.

v1.7 - effective 7 May 2026. Visual added: scatter charts for PAVIS and UCDP validations now embedded directly in the External validation section. No score changes, no methodological changes.

v1.6 - effective 6 May 2026. Third external validation added, against UCDP Georeferenced Event Dataset v25.1 (Uppsala Conflict Data Program, CC BY 4.0). Result: Spearman rho -0.37 across 193 PAs (p < 0.0001) on log fatalities versus PACE security score. Direction is correct (higher security score correlates with fewer nearby fatal conflict events). Three benchmarks attempted to date: PAVIS positive, GFW null, UCDP positive. No score changes.

v1.5 - effective 6 May 2026. Second external validation attempted against the Hansen Global Forest Change dataset; null result (Spearman rho +0.10 on ecosystem_services, n = 192). Disclosed as a methodology-limited null in the External validation section rather than withheld. Most likely explanation: the buffer-from-centroid approach samples regional landscape dynamics rather than inside-PA conditions, an issue that would resolve with protected-area boundary-polygon access. No score changes.

v1.4 - effective 6 May 2026. Two changes, no score changes. (1) Methodology page rewritten to make explicit that all PACE inputs are LLM-sourced, with no automated numeric data feeds. The Tier 1 to Tier 2 distinction was reframed as rubric-clarity rather than process; both tiers run through the same drafter, the difference is whether source material maps directly to the rubric or requires synthesis. (2) External validation section added: across 105 of 193 PACE-rated PAs with public visitor data, the ecotourism_infrastructure score correlates with log of mean recent annual visitors at Spearman rho 0.68 (p < 0.0001), benchmarked against the PAVIS v1.0 dataset (Buschke et al. 2024).

v1.3 - effective 5 May 2026. Scope rule clarified after audit. PACE was not in fact a strict subset of the RWF 176-keystone list as previously documented; v1.2 quietly included 15 non-keystone additions for major-operator coverage (3 African Parks, 3 Kenya Wildlife Service, 2 SANParks, 1 Peace Parks, 1 WCS, 5 state-run majors). Going forward: PACE backbone is the RWF 176-keystone list, plus a flagged set of operator-coverage additions where excluding them would create inconsistent operator representation. Every PA record now carries a keystone_member: bool field making membership explicit. Same release: five individual operator_type corrections (W Benin, Gambella, Boma, Manovo-Gounda, Luengue-Luiana) and a 21-PA government to state/kws merge aligning data to the documented operator-type taxonomy. No score changes.

v1.2 - effective 26 April 2026. Coverage expanded to 192 PAs (from 128 in v1.1) through two further depth passes. Batch 3 added 30 second-tier keystones with the first deployment of the critique-and-revise quality pass. Batch 4 added 36 country-residual keystones, also with critique enabled. Operator-mix continued moving away from African Parks dominance, now approximately 12% of cohort. Critique-and-revise infrastructure is now standard on every batch. One RWF keystone (Boumba-Djombi) excluded after editorial review found no verifiable standalone PA by the listed name; logged in an internal exclusions record. Luasi was initially excluded on the same grounds and subsequently promoted to the live PACE record on 3 May 2026; public protected-area registry verification confirmed 4 May 2026 that the canonical name is Lwafi-Nkamba Game Reserve (3,389 km2 combining the historic Lwafi GR with the adjoining Nkamba area), and the record was renamed accordingly.

v1.1 - effective 25 April 2026. Coverage expanded to 128 PAs (from 38 at launch) through staged additions: a country-coverage pass to widen pan-African breadth, then two depth passes adding 63 second-tier keystones predominantly under state management. Operator-mix rebalanced as a result: African Parks share fell from 44% to approximately 17% of cohort. PACE remains a strict subset of the curated Keystone Protected Areas list.

v1.0 - launched April 2026. Initial release with 38 PAs covering African Parks managed sites, Peace Parks landscapes, SANParks flagships, Kenya Wildlife Service parks, and major partnership-run reserves. Nine-dimension composite scoring framework established.

4.4 Time-series capture

PACE began as a snapshot tool. Each PA was scored, the rating was published, and the rating evolved silently in place as new evidence arrived. Useful for the moment, but no history was kept.

From 1 May 2026, every PACE record carries provenance fields recording which rubric was applied, when the score was generated, and which underlying source-data versions were used. The canonical published dataset continues to hold the current rating for each PA. Each month, an immutable snapshot of that dataset is preserved with cryptographic hashes and a manifest. Historical state can be reconstructed exactly.

The first snapshot was captured on 1 May 2026. Snapshots cannot be reconstructed retroactively, so this is the start of Canopy's panel view.

Between snapshots, each monthly re-score is anchored to the previous month's published score. The drafter is given the prior composite and every per-dimension score as an explicit baseline and reproduces them by default. A dimension moves only when the change is tied to specific new dated evidence, a new filing, dated news, or a satellite trigger, that the prior score did not reflect; a re-reading of the same evidence is not a reason to move a score. A single-scale-point move that cites no new evidence is held at the prior value, because that is the size of the noise: on the first full month-over-month diff, between the 1 July and 1 August 2026 snapshots, 27 percent of all dimension scores had shifted by exactly one point, almost perfectly balanced between up and down, the signature of run-to-run re-rating rather than change on the ground. Every surviving change is stored with a short reason tied to its evidence, and every held move is recorded as held, so a reader can separate real movement from noise.

This stabilises the published rating against re-rating variance and labels every change, but it still tracks change month to month rather than re-scoring the same PA across years. Systematic multi-year re-scoring, which would let the panel resolve causal questions in the cross-sectional findings, remains on the broader roadmap.

4.5 Disclaimer

PACE scores are informational only. They are not conservation partnership recommendations, grant-making advice, or a solicitation for any financial transaction. Canopy makes no warranty as to the accuracy or completeness of the underlying data. All philanthropic or investment decisions should be made on the basis of independent due diligence.

4.6 Coverage: the 162 and the 38

PACE rates 200 of Africa's most operationally significant protected areas under a single ten-dimension scoring rubric. That cohort is built in two layers. The first is the 162 Keystone protected areas - the canonical list behind the Africa Keystone Partnership, launched at the UN General Assembly in September 2025 by the Presidents of Botswana, Mozambique, Namibia, and South Africa, and backed by the Rob Walton Foundation, African Parks, WCS, and the Frankfurt Zoological Society. This is the spine of our coverage. The second is a set of 38 Canopy additions, protected areas we added on top of the Keystone list to make the cohort operationally complete and continentally representative rather than a re-presentation of a single partnership's roster. Membership is explicit in the data: every PA record carries a keystone_member field.

The 38 exist for three reasons. First, to complete every major operator's flagship estate. A credible operator comparison requires the whole operator, not a sample. The additions bring African Parks to full coverage at 24 of its 24 managed parks, and round out the SANParks flagship set (Addo Elephant, Table Mountain, Mapungubwe), the Kenya Wildlife Service icons (Amboseli, Lake Nakuru, Meru), and leading NGO- and partnership-run reserves (Ol Pejeta, Gola Rainforest, Makira). Without them, head-to-head operator scoring would carry blind spots.

Second, to include the marquee state-run icons any serious index must carry. Twenty of the 38 are major state-managed protected areas, many of them World Heritage Sites, that the Keystone list did not include: Simien Mountains (Ethiopia), Masoala (Madagascar), Banc d'Arguin (Mauritania), Aldabra Atoll (Seychelles), Ichkeul (Tunisia), and Tassili n'Ajjer (Algeria) among them. These are reference points the conservation field expects to see scored.

Third, to extend continental reach. The 38 bring PACE into 12 countries the Keystone 162 did not reach, including Madagascar, Morocco, Tunisia, Egypt, Mauritania, the Seychelles, Sierra Leone, Guinea-Bissau, and Burundi. The full 200-PA cohort now spans 31 countries; the 38 additions alone touch 22.

Every one of the 38 is scored under the same ten-dimension, tier-weighted rubric as the 162, with the same multi-sample voting (a draft reviewed by several adversarial critics), the same operator-level financial conditioning, and the same source discipline. There is no separate or softer methodology for the additions. One disclosed exception applies across the whole cohort: Boucle du Baoule (Mali), one of the 38, still carries a legacy nine-dimension score pending re-scoring.

The framing is straightforward. The Keystone 162 anchors PACE in an established, head-of-state-backed conservation initiative. The 38 additions demonstrate what makes PACE valuable: the ability to define and score the right universe on its own terms - complete operator coverage, the icons the field expects, and genuine continental breadth - under one consistent, defensible standard.