If you evaluate resilience instruments seriously, you eventually arrive at the obvious question: which one is the standard? Medicine has reference tests. Psychology has its established batteries. Surely resilience, after decades of research, has settled on a gold standard against which the rest are judged.
It has not, and that is not an opinion. It is the finding of the two most rigorous reviews of the field, run fourteen years apart, and knowing what they actually found will make you a sharper judge of every instrument you consider, including ours.
Is there a gold standard for measuring resilience? No. A 2011 methodological review of 19 resilience measures found that among the 15 distinct ones, none met the bar for a gold standard [1]. A 2025 systematic review of 24 scales developed since 2013 found only 11 (45.8%) met a two-thirds psychometric quality threshold, with stability the weakest property: just 4 of 24 scored any points on it [2]. Coaches evaluating instruments should ask for property-level evidence rather than the word “validated”.
Table Of Contents:
- What Windle found in 2011
- What changed by 2025
- The stability problem, precisely
- The six properties: the field, and the PRI, honestly
- What this means when you choose an instrument
- FAQ
- References
What Windle found in 2011
Windle, Bennett and Noyes reviewed 19 resilience measures against reliability, validity and practical use [1]. Four of the 19 were refinements of an original scale rather than distinct instruments, and among the 15 distinct measures, their conclusion was blunt: not one met the bar for a gold standard. Some scored well on individual properties; none put the whole picture together. For a young measurement field, that verdict was understandable. The reasonable expectation was that the next generation of instruments would close the gap.
What changed by 2025
Fourteen years later, Huerzeler, Boss and Thoma tested that expectation. Their 2025 systematic review in The Journal of Positive Psychology examined the 24 resilience scales developed between 2013 and 2024, the generation built after Windle’s verdict, and scored each against a structured set of psychometric quality criteria [2]. The headline: only 11 of the 24, 45.8%, reached a two-thirds quality threshold. The field improved; it did not resolve.
The per-property breakdown is more instructive than the headline, because it shows a pattern rather than a shortfall. The field scored well where evidence is cheapest to produce: 21 of 24 scales scored maximum on criterion validity, 19 on content validity, 18 on reliability. It scored poorly exactly where evidence is expensive: 8 of 24 on replicability, 7 on construct validity, and, weakest by a distance, 4 of 24 scored any points at all on stability. Read as a shape, the finding is uncomfortable and useful: resilience instruments are routinely good at the properties a single study can establish, and routinely unexamined on the properties that take sustained, costly work.
The stability problem, precisely
Stability deserves a precise definition, because it is the property most often blurred. In the 2025 review, stability means test-retest and parallel-forms reliability: if the same person, genuinely unchanged, takes the instrument twice, do the scores agree (the applied threshold being r ≥ 0.75 [3])? It does not mean responsiveness, the ability to detect real change after an intervention. Those are different properties answering different questions, and the review does not score responsiveness at all. An instrument can be stable but insensitive, responsive but noisy, or, ideally, both. Conflating them flatters everyone and informs no one, which is why this piece keeps them apart, including when assessing our own instrument below. Explainers on convergent validity and on Cronbach’s alpha go deeper on the terms used here.
The six properties: the field, and the PRI, honestly
The fair test of any instrument, ours included, is the same six properties the field was scored on. The table below puts the 2025 aggregate findings beside the PRI’s own published figures. Every PRI figure here is public on our Evidence page; nothing is claimed that cannot be checked.
| Property (per the 2025 review [2]) | What the field scored (of 24 scales, 2013 to 2024) | Where the PRI stands (public figures only) |
|---|---|---|
| Criterion validity | 21 of 24 scored maximum | Convergent validity published: r = 0.72 with the Sense of Coherence scale (SOC-13); r = 0.67 with positive and −0.55 with negative affect (PANAS) |
| Content validity | 19 of 24 scored maximum | 260 candidate items developed from a research-grounded model of resilience, refined to the 64 scored items through testing |
| Reliability | 18 of 24 scored maximum | α = 0.94 overall; 0.76 to 0.85 across the six domains |
| Replicability | 8 of 24 scored maximum | Methodology documented in the PRI Technical Report [4], summarised publicly on the Evidence page |
| Construct validity | 7 of 24 scored maximum | In the validation data, multi-factor models clearly outperformed single-factor models |
| Stability (test-retest and parallel-forms reliability) | 4 of 24 scored any points at all; the field’s weakest property | Formal test-retest work not yet complete, described in the validation documentation as upcoming, and stated openly on our Evidence page. The PRI has demonstrated responsiveness (driver-level scores moving with real coaching and training interventions), which is a distinct property, not a substitute for stability, and one the 2025 review does not score |
The stability row deserves the plain-language version. The PRI does not yet have formal test-retest figures; our validation documentation describes that work as upcoming, and our Evidence page says so openly. What the PRI does have is demonstrated responsiveness: driver-level scores that move when clients do genuine coaching and training work, which is a valuable property for a development instrument and a distinct one. We state clearly that responsiveness is not a substitute for stability. The honest claim, and we think the stronger one, is this: on the property where 20 of 24 recent scales simply are not measured at all, the PRI is unusually transparent about exactly where its evidence ends and its ongoing work begins.

What this means when you choose an instrument
The practical lesson of both reviews is not that resilience measurement is hopeless; it is that the word “validated” has been doing too much work. In a field where fewer than half the recent instruments clear a two-thirds quality bar, validation is not a badge, it is a set of specific properties, each with evidence or without it. So evaluate the way the reviewers do: ask any instrument, ours included, for its reliability figure, its convergent validity correlations, its construct evidence, and its honest position on stability. An instrument that answers property by property respects your judgement. One that answers with an adjective is hoping you will not ask twice.
If that is the standard you want behind your own practice, the PRI Certification Training trains you to work with an instrument that publishes its evidence, and to read any instrument’s evidence with the same eyes.
FAQ
Is there a gold standard resilience measure?
No, and there never has been. The 2011 methodological review found no gold standard among 15 distinct measures, and the 2025 systematic review of 24 newer scales found only 11 met a two-thirds quality threshold. The practical consequence for you: treat “validated” as the start of your questions, not the end of them, and ask for property-level evidence.
What does stability mean in psychometrics?
Test-retest and parallel-forms reliability: whether an unchanged person scores the same twice, with r ≥ 0.75 the commonly applied bar [3]. Here is the distinction people blur: stability is not responsiveness, the ability to detect real change after an intervention. They are different properties, and an honest instrument tells you its position on each rather than letting one stand in for the other.
Why do so few resilience scales measure stability?
Because it is expensive in exactly the way most validation studies are not. Testing stability means bringing the same people back after a controlled interval and being able to argue they have not truly changed in between, which for a construct as dynamic as resilience is genuinely hard to design. That difficulty explains the 2025 finding, only 4 of 24 scales scored any points on it, but it does not excuse silence: the honest move is to state the gap, which is what we do on our Evidence page.
How does the PRI perform against these criteria?
Property by property, and publicly: reliability of α = 0.94 overall and 0.76 to 0.85 by domain, convergent validity of r = 0.72 with the SOC-13 and r = 0.67 and −0.55 with the PANAS, construct evidence that multi-factor models clearly outperformed single-factor models, and an openly stated position on stability: formal test-retest work is upcoming, while demonstrated responsiveness to real interventions, a distinct property, is already in evidence. The full figures are on the Evidence page.
References
[1] Windle, G., Bennett, K. M., & Noyes, J. (2011). A methodological review of resilience measurement scales. Health and Quality of Life Outcomes, 9(1), 8.
[2] Huerzeler, H. E., Boss, P., & Thoma, M. V. (2025). Addressing the heterogeneity of resilience scales: A systematic review and development of a unified resilience construct within a standardized resilience framework. The Journal of Positive Psychology, 1–20. https://doi.org/10.1080/17439760.2025.2574049
[3] Portney, L. G., & Watkins, M. P. (2015). Foundations of clinical research: Applications to practice (3rd ed.). F. A. Davis.
[4] Sinclair, N., Hafner, G., & Sinclair, P. D. (2022). Personal Resilience Indicator: Validation summary and psychometric properties (PRI Technical Report). Mind Matters. Available on the Evidence page.
