The Ruler and the Thermometer: Stability vs Responsiveness in Resilience Measurement

BY: Nadine SinclairAugust 12, 2026
  • Home
  • »
  • Gain Insights
  • »
  • The Ruler and the Thermometer: Stability vs Responsiveness in Resilience Measurement

Somewhere in your evaluation of assessment tools, you will have met two claims that sound like the same virtue. One spec sheet says an instrument is stable: test it twice, get the same answer. Another says an instrument is sensitive to change: run an intervention, watch the scores move. Both get filed under “reliable,” both sound like quality, and most buyers never notice that they are pulling in opposite directions.

The confusion is not your fault; it is baked into how instruments are marketed. But the distinction decides what an assessment can honestly do for your practice, so it is worth fifteen minutes to own it permanently. Two household objects will do most of the work.

What is the difference between stability and responsiveness? Stability (test-retest and parallel-forms reliability) is whether an unchanged person gets the same score twice, with r ≥ 0.75 the commonly applied bar [1]. Responsiveness is whether an instrument detects real change when change actually happens. The international COSMIN consensus treats them as separate measurement domains [2]. An instrument can have either without the other, and which one matters depends on whether you are describing people or developing them.

Table Of Contents:

The Ruler

Measure an adult’s height today, then again next month. If the ruler reads differently, you do not conclude the person grew; you conclude the ruler is faulty, because adult height barely moves. That is stability, and test-retest reliability is exactly the right exam for it: same person, genuinely unchanged, measured twice, and the scores should agree. Instruments that measure trait-like constructs, the stable architecture of who someone is, are rulers. For them, stability is the headline virtue, and an unstable trait measure is simply broken.

The Thermometer

Now read a thermometer at seven in the morning and again at seven in the evening. The numbers differ, and nobody calls the thermometer unreliable, because the temperature actually changed. Faulting a thermometer for moving is faulting it for working. That is responsiveness: the capacity to detect real change when real change occurs. Instruments built for development work, where the entire point is that the construct should move under intervention, are thermometers, and for them responsiveness is not a nice-to-have; it is the job description.

Here is the part almost nobody tells practitioners: the two properties belong to formally separate domains in the international consensus taxonomy of measurement properties (COSMIN) [2], alongside validity. And the same consensus work notes why the distinction is so often missing from psychology-adjacent instruments: the psychology tradition’s own frameworks historically excluded responsiveness altogether, because psychological measures were mostly built to discriminate between people at one point in time, not to evaluate change across time. An entire branch of psychometrics, latent state-trait theory, exists precisely to model how much of any score reflects stable trait and how much reflects current state. The distinction is old, established, and routinely ignored in marketing copy.

The Stopped Clock

One more object completes the set. A stopped clock delivers perfect test-retest reliability: read it twice, get identical answers, every time. It is also useless, because its stability is achieved through total insensitivity. This is the failure mode the ruler-lovers forget: an instrument can score beautifully on stability by being incapable of registering anything, including the change your client just worked six months to achieve. Stability without responsiveness describes; it cannot evaluate. If your practice is development, a stopped clock with excellent psychometrics is still a stopped clock.

Traits, states, and the interval problem

The two metaphors map directly onto the trait-and-state distinction that runs through resilience science (see our companion piece, state vs trait resilience). Trait constructs are ruler territory: stable over months and years, properly tested with a straightforward retest. State constructs are thermometer territory: genuinely dynamic, responsive to sleep, stress, intervention and circumstance. And that creates a real design problem for testing stability in state-sensitive instruments, because the standard test-retest convention quietly assumes the construct barely moves across the interval. Retest too soon and people remember their answers; wait the conventional few weeks and a state construct may have genuinely changed, so real movement gets misread as instrument error.

One fence around this point, because it matters: being a thermometer is a design constraint on how stability must be tested, not an excuse for never testing it. Well-designed studies handle state-sensitive constructs with more careful interval choices and modelling, and a serious instrument owes the field that work. The honest position is never “stability does not apply to us”; it is “stability testing for a state measure takes more care, and here is where ours stands.”

Where resilience measurement stands

Now the field data. In the 2025 systematic review of the 24 resilience scales developed between 2013 and 2024 [3], stability was the weakest property by a distance: only 4 of 24 scales scored any points on it at all. And responsiveness was not scored in the review, consistent with the psychology tradition described above. Put plainly: on the ruler question, five of every six recent resilience scales are simply unmeasured, and on the thermometer question, the field’s standard evaluation does not even ask. For a construct that practitioners overwhelmingly want to develop rather than merely describe, that second silence is arguably the more remarkable one. The full field picture, including the 2011 review that found no gold standard [4], is in our companion piece: is there a gold standard for measuring resilience?

What the PRI does with this

The Personal Resilience Indicator sits deliberately on both sides of the map, which is exactly why this distinction matters to how it is tested. Some of its twelve drivers are trait-leaning, Intuition and Connection sit closer to the ruler. Others are state-sensitive by design, Sleep, Motivation and Emotional Agility can genuinely shift within weeks, which is thermometer territory, and the point of a development instrument. Its position on the two properties, stated the same way here as on our Evidence page: on responsiveness, the PRI has demonstrated evidence, driver-level scores that move with real coaching and training interventions, which is the property a development instrument exists for. On stability, formal test-retest work is not yet complete; our validation documentation describes it as upcoming, and because several drivers are state-sensitive, that study has to be designed with the interval problem above taken seriously rather than run as a naive two-week retest. We publish the figures when they are sound, and until then we say so, which, in a field where 20 of 24 recent scales are not measured on stability at all, is the position we would want any instrument to take.

What to ask any instrument

  • Is the construct you measure a trait, a state, or a mix? The answer determines which property matters most.
  • What is your test-retest evidence, over what interval, and why that interval? A bare correlation without an interval rationale is half an answer.
  • What is your responsiveness evidence? If the instrument is sold for development work, this is the property the sale rests on.
  • Where the evidence does not yet exist, does the instrument say so? Silence on a property is a finding about the vendor, not the property.

If you want the skills to evaluate and use instruments at this level, property by property, the PRI Certification Training teaches exactly that, on an instrument that answers these four questions in public.

FAQ

What is the difference between test-retest reliability and responsiveness?

They answer opposite questions. Test-retest reliability (stability) asks: if nothing has truly changed, do two measurements agree? Responsiveness asks: if something has truly changed, does the instrument notice? The international COSMIN consensus treats them as separate measurement domains, and an instrument can have either without the other. A stopped clock has perfect stability and zero responsiveness, which is all you need to remember.

Can an instrument be reliable but unable to measure change?

Yes, easily, and it is common. High internal consistency and strong test-retest figures tell you an instrument is stable and coherent; they say nothing about whether it can register the change your client just achieved. If you use assessments to evaluate development rather than describe people, ask for responsiveness evidence specifically, because “reliable” on a spec sheet does not include it.

Why do resilience scales rarely report responsiveness?

Tradition, mostly. The consensus literature notes that psychology’s own measurement frameworks historically excluded responsiveness, because psychological instruments were built to discriminate between people at a point in time rather than to evaluate change. Resilience scales inherited that habit, and the 2025 systematic review of the field did not score responsiveness at all. The awkward result: a field full of instruments sold for development, evaluated by standards built for description.

How does the PRI handle stability and responsiveness?

Openly, and differently for each. On responsiveness, the PRI has demonstrated evidence: driver-level scores move with genuine coaching and training interventions, which is what a development instrument is for. On stability, formal test-retest work is upcoming and stated as such on our Evidence page, partly because several PRI drivers are state-sensitive by design, so the study must be built around the interval problem rather than run naively. The two properties are kept distinct, in our claims and in our plans.

References

[1] Portney, L. G., & Watkins, M. P. (2015). Foundations of clinical research: Applications to practice (3rd ed.). F. A. Davis.

[2] Mokkink, L. B., Terwee, C. B., Patrick, D. L., Alonso, J., Stratford, P. W., Knol, D. L., Bouter, L. M., & de Vet, H. C. W. (2010). The COSMIN study reached international consensus on taxonomy, terminology, and definitions of measurement properties for health-related patient-reported outcomes. Journal of Clinical Epidemiology, 63(7), 737-745.

[3] Huerzeler, H. E., Boss, P., & Thoma, M. V. (2025). Addressing the heterogeneity of resilience scales: A systematic review and development of a unified resilience construct within a standardized resilience framework. The Journal of Positive Psychology, 1–20. https://doi.org/10.1080/17439760.2025.2574049

[4] Windle, G., Bennett, K. M., & Noyes, J. (2011). A methodological review of resilience measurement scales. Health and Quality of Life Outcomes, 9(1), 8.

[5] Sinclair, N., Hafner, G., & Sinclair, P. D. (2022). Personal Resilience Indicator: Validation summary and psychometric properties (PRI Technical Report). Mind Matters. Available on the Evidence page.

{"email":"Email address invalid","url":"Website address invalid","required":"Required field missing"}

Author Profile

Nadine Sinclair 

Dr. Nadine Sinclair is a co-developer of the Personal Resilience Indicator and co-founder and managing director of Mind Matters. A scientist by training, she conducted her doctoral research at the Max Planck Institute for Biophysical Chemistry and brought that research discipline to the PRI's development and independent validation. Before founding Mind Matters, she spent 18 years as a management consultant, at McKinsey and independently, advising many of the world's leading companies, with more than 30,000 hours of hands-on client work. Today she works with coaches, teams and organisations that want to measure resilience rather than guess at it.

Related Posts