In elementary school I wanted to be an astronaut. Or an astronomer. I was drawn to the big questions — what does it mean for space to be infinite? What was there before the Big Bang? The kind of questions that make you feel very small in a satisfying way.

But somewhere along the way, the questions shifted. I stopped asking “what is space?” and started asking “why can’t I imagine infinity as a whole in my head?” Which led to: what is a mind? What does it mean to be conscious? Is it possible that I’m the only conscious thing in the universe and the rest of you only exist in my imagination?

By high school I was deep into Carl Jung and Joseph Campbell. The inner world had become more interesting than the outer one — though I still thought robots and spacecraft were pretty great.

In college I majored in psychology with an emphasis on research. I loved measuring things. I took everything measurement-related I could find: three semesters of applied statistics, psychophysiological recording, clinical neuropsychology, design of experiments. When I finished, I had a rigorous set of tools for measuring things that were hard to define and impossible to directly observe.

I just didn’t know what to do with them.


This was the mid-90s. The internet was the big new thing. I found my way into software quality assurance — it’s kind of like psychology, if you squint. Evaluation of behavior in a simplified form. I had a paycheck now.

I also had a wealth of knowledge burning a hole in the back of my head that I couldn’t use. Every time I tried to apply what I knew — construct validity for test design, reliability measurement instead of pass/fail counts, controlled evaluation methodology — the answer was always one of two things: this is the wrong tool for a deterministic system, or this is overkill. Management was never interested in a deep dive.

My first QA mentor could have predicted that. He had done verification work on software for the NASA Space Shuttle program. He was also a psychology major. When I told him I wanted to apply measurement science to software quality, he didn’t argue with the idea. He argued with the timing. Software companies, he said, were too busy building software bound to fail to invest in a process that would stop this. His advice: don’t go into process or quality engineering.

I went into quality engineering anyway.


The clearest illustration of why he was right came years later, during a death march — one of those development cycles where schedule slips eat the testing time until there’s none left. A member of the testing team pointed out the obvious: we were about to ship a product that hadn’t been tested. The QA manager’s response was matter-of-fact.

“We’re not building rocket ships here.”

He wasn’t wrong. We weren’t.

But I remember thinking: if shipping untested software is acceptable when the schedule is tight, what exactly is the purpose of this work? There had to be more to quality engineering than a checkbox. Something with a scientific basis for what “good enough” actually meant.

My mentor had literally worked on rocket ships. He knew what measurement rigor looked like when the stakes were high enough to demand it. His prediction wasn’t that rigorous measurement was impossible. It was that software companies weren’t ready to pay for it.

I spent the next two decades proving him right — climbing from junior manual tester to test automation framework architect, performance engineer, SDET. All the while being reminded, with some regularity, that my psychology degree was not a computer science degree.


Then the AI revolution happened.

AI systems have a layer that traditional software does not. They acquire structured tendencies from training data — implicit assumptions about language, evidence, and social categories — that function like psychological constructs. You cannot read these tendencies from the code. You cannot unit-test your way to them. You can only infer them from behavior.

Machine learning engineers are finding this layer genuinely difficult to evaluate, because the tools they learned — built for deterministic systems with verifiable rules — cannot reach it. So far, the solution has eluded them.

It hasn’t eluded me.

This is the exact type of system that psychometrics was invented to evaluate, decades before the electronic computer existed. The field spent 70 years developing rigorous methods for measuring constructs you cannot directly observe: construct validity, inter-rater reliability, measurement invariance, differential item functioning. The AI industry is now rediscovering every one of these problems, often by running into them in production, and calling them by different names.

Now — I know what you’re thinking. Your crazy ex from college may have majored in psychology. Fine. But at its core, psychology is a measurement science. The evaluation frameworks the AI industry needs already exist. They just live in a field most ML engineers have never heard of.


My mentor told me not to go into quality engineering. I’m glad I didn’t take his advice. But he was right that the industry wasn’t ready.

I founded Noometric because now it is.

I’m writing here about the intersection of psychometric methodology and AI evaluation — what 80 years of measurement science has to say about the problems AI teams are actually facing. If that’s your problem too, please subscribe.