~/tools/llm-scales/method
How the LLM Scale Finder list was made, so you can judge it, and reuse or improve the approach.
26AI agents
~10 hwall-clock time
1,751searches & page reads
174 → 118candidates → listed
the run
- model
- Claude Opus 5.5 (
claude-opus-5-5, Anthropic) for every agent, run through the Workflow tool in Claude Code. - when
- 28 September 2026, 8:47 pm → 29 September 2026, 7:10 am (Toronto). That includes two pauses when the account hit its usage limit; the run resumed automatically each time.
- scale
- 26 agents, about 4.5 million tokens, and 1,751 tool calls (web searches and reads of publisher pages, PubMed Central, preprint servers and OpenAlex).
- code
- The complete workflow script, including every prompt and the output schema: llm-scales-workflow.js. The resulting data: llm-scales.json.
steps
- Sweep (7 agents in parallel). Each agent searched one cluster of constructs and returned every candidate scale with its citation, subscales, psychometrics and a note on where it saw the evidence. Result: 158 unique candidates, plus near-misses set aside.
- Gap-fill (1 agent). A “completeness critic” got the full list and searched for what was missing: non-English journals, specific professions, code assistants, children and older adults, conference papers, and recent preprints. It added 16 candidates.
- Verify (18 agents). The 174 candidates were split into batches of 10. For each batch, a separate, skeptical agent opened every source and checked the citation, the inclusion criteria, subscale names, item counts, reliability numbers and publication status, then corrected the record. Result: 128 kept, 41 borderline, 5 dropped, 1 not checked.
- Clean-up (by Claude in the main session, then me). Validated translations were nested under their original scales (9 entries). Population and language labels were standardised. All 143 DOIs were checked to resolve at doi.org, and 12 entries (about 10%) were spot-checked against their papers. Of the 41 borderline scales, 35 are listed as flagged (hidden by default, each with the reason). The other 6 were left out: 5 whose items may not target generative AI, and 1 translation that was already listed.
my request
The instruction I gave Claude (verbatim, typos included), followed by the scope decisions I made when it asked:
I want you to /plan a new tool. Essentially, I want you to do a deep deep deep dive on Human-AI interaction work and get me a list of all psychometrically validatd survey. The point of this tool is to be able to search so that students/researchers have an easier time to know what is out there, that is validated, what is a pre-print or published, and links to the paper and summaries of the subscales/point of the survey.
- Validation bar: a dedicated development/validation paper or preprint reporting factor structure, reliability and at least one form of validity evidence.
- Scope: modern LLMs specifically; general “AI” scales excluded.
- Translations: nested under the original scale.
- Research approach: a multi-agent workflow with independent verification.
prompts
Every agent received the shared instructions, followed by its role-specific prompt. They are reproduced verbatim; bracketed text marks where the script inserted content.
shared instructions (all agents)
TASK CONTEXT: We are building a public, searchable catalogue of PSYCHOMETRICALLY VALIDATED survey scales about people's interaction with MODERN LLMs / GENERATIVE AI, for students and researchers. INCLUDE a scale only if ALL hold: 1. Its items target LLMs, generative AI, or AI chatbots (ChatGPT, Copilot, Gemini, Claude, Bard, Ernie, AI companions like Replika/Character.AI, "generative AI tools"). A general-AI scale re-worded for ChatGPT/GenAI and fully re-validated counts (as its own entry). 2. It has a dedicated development/validation paper or preprint (main aim = develop or validate the scale), reporting (a) factor structure (EFA, CFA, ESEM, IRT/Rasch, network), (b) reliability (alpha, omega, CR, test-retest), and (c) at least one validity evidence (convergent/discriminant, criterion, known-groups, measurement invariance, nomological). 3. Published (journal or peer-reviewed proceedings) OR preprint (PsyArXiv, OSF, arXiv, SSRN, Research Square, medRxiv, etc.). EXCLUDE: general "AI" scales not targeted at GenAI/LLMs (GAAIS, ATAI, AI anxiety scale, MAILS, AI literacy scales about AI broadly), robots, automation/decision-aid trust, TAM/UTAUT/structural models that merely adapt items without a validation study, theses/dissertations, papers that only USE a scale. TRANSLATIONS/ADAPTATIONS: if a validated translation/cross-cultural adaptation of a qualifying scale exists, list it in that scale's "adaptations" array. If you only find the adaptation, still report it as a record and set "adaptation_of" to the original scale's name + DOI. SEARCH HOW: First load web tools: call ToolSearch with query "select:WebSearch,WebFetch". Search Google Scholar-style queries, PsyArXiv, OSF, arXiv, SSRN, ERIC, Research Square, and publisher sites (Elsevier, Springer, Taylor & Francis, Frontiers, MDPI, Wiley, SAGE, APA). Use many query variants (e.g. "ChatGPT scale development validation", "generative AI" scale psychometric, "chatbot" questionnaire CFA reliability, Turkish/Chinese/Spanish/German adaptation). Snowball: check reference lists of reviews of GenAI measurement and "cited by" of known scales. Be exhaustive for your cluster — aim to find every qualifying scale, 2023-2026 especially. HONESTY: Only report what you actually saw in a source (full text preferred; abstract minimum). Never guess DOIs — fetch https://doi.org/<doi> or the publisher page to confirm. If a field is unknown write null (or empty array), do not invent. In evidence_notes say what you saw and where (full text vs abstract only). FIELD RULES: - id: kebab-case slug of acronym or short name (e.g. "chatgpt-dependence-scale"). - summary: 1-2 plain-language sentences for a student: what it measures and when you'd use it. - constructs: 1-3 values ONLY from: attitudes, acceptance-use, trust, reliance, credibility, dependence, literacy, self-efficacy, anxiety, ethics-concerns, privacy, academic-integrity, learning, anthropomorphism, social-presence, relationships, workplace, health, creativity, other. - subscales: name, item count, one-line description of what that subscale captures. Unidimensional → one entry named "(unidimensional)". - psychometrics.reliability: short string with numbers, e.g. "α .84–.92; ω .86". - citation: APA 7. - status: "published" or "preprint" (if a preprint later appeared in a journal, status=published and keep preprint_url). - items_available: true only if full item wording is given in the paper/appendix/OSF.
sweep agents (7)
[shared instructions above] YOUR CLUSTER: [cluster description] Find every qualifying scale in this cluster (scales that also fit other clusters are fine to include). Put near-misses you are unsure about in "borderline" with the reason.
The seven cluster descriptions:
- attitudes: Attitudes toward ChatGPT/GenAI, perceptions, acceptance, adoption, use intention, usage frequency/behaviour, satisfaction, perceived usefulness scales that were properly validated.
- trust: Trust in LLMs/chatbots, reliance, over-reliance, verification behaviour, perceived credibility/accuracy of GenAI output, scepticism, calibration.
- dependence: ChatGPT/GenAI dependence, addiction, problematic or compulsive use, overuse, AI chatbot dependency, withdrawal, fear of missing out on AI.
- literacy: GenAI literacy, ChatGPT competence, prompt engineering skill, GenAI self-efficacy, digital competence with GenAI, critical evaluation skills, readiness to use GenAI.
- anxiety: GenAI/ChatGPT anxiety, technostress, job/replacement threat from GenAI, ethical concerns, perceived risks, privacy concerns, misinformation worry, moral judgments of GenAI use.
- education: Education-specific: student and teacher GenAI/ChatGPT use scales, academic integrity/cheating/plagiarism with GenAI, GenAI in learning/writing/assessment, teacher acceptance, learning outcomes, AI-assisted learning motivation.
- social: Anthropomorphism/mind perception of LLM chatbots, social presence, parasocial/companion relationships (Replika, Character.AI), emotional attachment, loneliness with AI companions, plus workplace GenAI scales (employee use, AI-augmented work, creativity with GenAI) and health contexts (patients/clinicians using ChatGPT).
completeness critic (1)
[shared instructions above] YOU ARE THE COMPLETENESS CRITIC. Seven cluster sweeps already found these scales: [list of every scale found so far] Find qualifying scales that are MISSING from this list. Try angles the sweeps may have missed: non-English origin scales (Chinese, Turkish, Korean, Arabic, Spanish, Portuguese, German, Indonesian, Persian journals), specific professions (nurses, physicians, lawyers, programmers, journalists), code assistants (GitHub Copilot), AI image generators if paired with text GenAI, older-adult samples, children/adolescents, conference proceedings (CHI, CSCW), recent 2025-2026 preprints, systematic reviews of GenAI instruments and their tables. Also list validated translations/adaptations of the known scales (report them as records with adaptation_of set). Return ONLY new items, not ones already listed.
verifiers (18)
[shared instructions above] YOU ARE AN INDEPENDENT, SKEPTICAL VERIFIER. Other agents proposed the records below. For EACH record, open the actual source (fetch the DOI/URL; find full text via publisher, PMC, ResearchGate, preprint server) and check: 1. The DOI/URL resolves to this paper; citation (authors, year, title, venue) is correct. 2. It meets ALL inclusion criteria (GenAI/LLM-targeted items; dedicated validation; factor structure + reliability + validity evidence). If it fails, verdict "drop" with reason. If you truly cannot tell, "borderline". 3. Subscale names, item counts, total items, response format, reliability numbers, samples match the paper. FIX any errors in "record". 4. Status is current: if a preprint now has a journal version, set status "published", update citation/doi/venue, keep preprint_url. 5. Adaptations listed are real (check each quickly); remove any you cannot confirm. 6. Summary is accurate and plain-language; constructs fit. If the record has adaptation_of set, confirm the original and keep adaptation_of (we will nest it). Return one result per input record (same id), with the corrected full record. In corrections, list what you changed (or "none"). In record.evidence_notes, state what you verified and whether from full text or abstract. RECORDS: [a batch of about 10 proposed records]
limits
- AI agents can misread papers. Verification and spot checks reduce errors but don't remove them; always read the original paper before using a scale.
- Some papers were paywalled, so a few records rest on abstracts rather than full text.
- Coverage reflects what was findable online as of 29 September 2026. Suggest a missing scale.