~/tools/llm-scales

If you are a human–AI interaction researcher like me, you have no doubt had the pleasure of trying to figure out which scales are out there, which ones were actually validated, and whether that validation ever made it past a preprint. So I set out to build a list.

Specifically, I used Claude Opus 5.5 to run a multi-agent deep dive: 26 AI agents searched the literature over about 10 hours, and every scale was then checked against its source paper by a separate agent. See the model, agents, time and prompts used →

–validated scales
–published
–preprints
–languages

scales how scales are chosen ↓

$ grep
construct
evidence what's flagged?
status

loading…

    no scales match these filters

    download json ↓

    about

    what counts
    A scale is listed when its items target LLMs, generative AI or AI chatbots (ChatGPT, Copilot, Gemini, Claude, AI companions) and a paper or preprint whose main aim is developing or validating it reports (a) the factor structure, (b) reliability and (c) at least one form of validity evidence. Preprints are included and labelled.
    not included
    General “AI” scales that don’t target generative AI (e.g., attitudes toward AI in general, AI anxiety, AI literacy), robot and automation scales, and adoption models that adapt items without a validation study.
    flagged
    Scales that target generative AI and have a development paper, but that an independent check could not confirm meet the bar above. They are hidden by default; use the evidence filter to show them. Partial evidence means the paper was read and something is missing, usually validity evidence beyond factor analysis. Unconfirmed means the full text could not be accessed, so the scale may well qualify. Each card says why it was flagged.
    translations
    Validated translations and cross-cultural adaptations are listed inside the original scale’s entry.
    method
    Built from an AI-assisted literature sweep in which every entry was then checked against its source paper, followed by manual spot checks (full method and prompts). Mistakes are still possible: always read the original paper before using a scale. Last updated –.
    missing one?
    Suggest a scale on GitHub or by email. Corrections are welcome too.