Aug 15, 2026

Before Asking If a Scale Is Good, Ask If the Thing Is Measurable

Human Elements System · 3 min readRead on Substack ↗

Psychometrics has spent a century asking: is this instrument any good? Reliability, validity, replication, norms. Whole shelves of methodology. Fair enough.

But there’s a question underneath that one. A question that stays hidden because everyone assumes the answer is yes.

Is the thing you’re trying to measure actually measurable yet?

Not “is your scale good.” Not “did you validate it in three samples.” The question is one level up: does the construct itself sit at a point in its scientific life where measurement is even the right verb?

Some constructs are ready. Working memory has decades of instruments behind it. Cronbach’s alphas north of .80, cross-cultural validation, normative datasets. If someone hands you a working-memory task, you know roughly what to trust.

Other constructs are not ready. They appear in papers, sometimes in whole review articles, but no dedicated instrument exists. Or the instrument that exists measures something adjacent and gets treated as if it measures the thing. Or the reliability is low and everyone quietly ignores it.

And some constructs sit in between. They have partial measurement. Or one good instrument but only one, without independent replication. Or good psychometrics on the wrong sample.

This is what the MML framework tries to make visible.

MML — Measurement Maturity Levels — is a construct-level classification system. Five levels, from MML-1 (Established: multiple validated instruments, cross-cultural evidence, normative data) down to MML-5 (Novel: not yet present in peer-reviewed literature under this or an equivalent name). Each level is assigned by walking a decision tree. Does a dedicated instrument exist? Is reliability above threshold? Does convergent validity hold? And so on.

The point isn’t to grade constructs on how “good” they are. The point is to make the measurement landscape explicit, so that when someone claims to have measured X, we know whether X is at the point where such claims are stable, or whether we’re still building the road.

I applied MML to a dataset of 343 constructs drawn from the Human Elements System, the framework I’ve been writing about here. The results surprised me. 60% sit at MML-1 or MML-2, established or adequate measurement. Another 22% are Emerging (MML-3). And 18% are what MML calls Pre-measurement (MML-4), recognized concepts, real research, but no dedicated instrument yet. Zero at MML-5, because HES was built by anchoring to existing literature.

The 18% is the interesting part. Those aren’t fringe or speculative constructs. Many are things psychology talks about all the time. Dread. Relief. Schadenfreude. Complex emotional states that appear in every serious taxonomy of affect but don’t have their own validated scale. We measure them by proxy or don’t measure them at all.

MML just makes this visible instead of leaving it implicit.

The paper is now on PsyArXiv:

https://doi.org/10.31234/osf.io/ujctf_v1

It’s short (~7,000 words) and includes the full 343-construct classification as a supplementary table. If you work with psychological measurement, or if you’ve ever wondered why some constructs feel harder to pin down than others, the paper might be worth a look.

The idea has been quietly sitting inside HES from the beginning. Publishing it separately makes it usable outside HES too. Any taxonomy, any construct set, any research program can run its own readiness audit.

That’s the small unlock.