University of Sussex ResearchPlus Reality Bending Lab

The

Abyss

The Science-based Assessment of Deep Personality

Research-grade measures immediate results entirely anonymous

Documentation

The trouble with surveys

Psychology runs on questionnaires, and nearly all of them are asked the same way: a page of radio buttons on a platform built for the person writing the study, not the person answering it.

The instruments are rarely what fails; the answering is. No amount of validation survives careless answers, and a dataset nobody can trust is not one worth sharing.

Somebody at a desk, bored, looking at a screen with their head on their hand

What the Abyss is

A survey platform built to be engaging enough that people want to finish it. The questionnaires underneath are ordinary validated ones, asked as they were written; everything around them is there to keep somebody reading the question they are answering, to the last of them.

The aim is high-quality, re-usable, open-access datasets, clean enough to be pooled and wide enough to show how far their measures overlap, since the same people answer all of them.

A clustered correlation matrix of the items, red and blue blocks along the diagonal

Built to be changed: the questions are data rather than code, and a study asks its own battery from its own link.

What is being validated

MINT

Multidimensional Interoceptive Traits questionnaire · Makowski et al.

How people sense the inside of their own body, as a trait.

BAIT

Beliefs about Artificial Intelligence Technology · Makowski et al.

What people believe AI can make, and how they feel about it.

Core · levels 1–4, every participant

Optional · levels 5–10, the participant's choice

MINT

Multidimensional Interoceptive Traits questionnaire · Makowski et al.

Where it comes from

Its first validation (ER/MB2021/2, two studies) established its structure and how it relates to other interoceptive scales. With the two later studies that asked it too, 1,684 people are its norms.

In this study

  1. Replication of structure on a new, larger sample with the same design, and its convergent and discriminant validity against the measures around it.
  2. Effect of response format. Each run is dealt one of three formats at random, all scored 0 to 6, to see whether the numbers on a scale change how it is answered.
  3. Correlates with HiTOP, the Hierarchical Taxonomy of Psychopathology, which describes mental ill health as dimensions rather than diagnoses: to assess the relationship between interoception and psychopathology within that framing.
The three formats, as a participant meets them

I can notice even very subtle changes in my breathing

0 to 6 Disagree Agree
−3 to +3 Disagree Agree
No numbers Disagree Agree

BAIT

Beliefs about Artificial Intelligence Technology · Makowski et al.

Where it started

As a trait measure that could partly modulate and mediate the anti-AI bias the lab measures in its tasks, drawing on two things:

Expectations about what AI can produce Current AI algorithms can generate very realistic videos
Attitudes, items carried over from the GAAIS Much of society will benefit from a future full of AI

partly modulates and mediates

Anti-AI bias in the lab's tasks: the same image judged less favourably once it is believed to be made by AI

Where it is going

Its psychometric properties have not been ideal, and more work is needed to make it a better measure.

  1. Sharper expectations. Improve the sensitivity of what it measures about what AI can produce, across modalities: images, text, art and more.
  2. Subjective discriminability. Explore a separate factor for how well people think they can tell AI's work from a person's. Four such items are asked already, and not yet scored.
  3. Usage, taken further. Extend it to AI addiction and "AI psychosis": the bond with a chatbot and the spiral the reports describe.

Content

Level Questionnaire Dimensions Items

HiTOP

Hierarchical Taxonomy of Psychopathology · Kotov et al., 2017

What it is

A classification of mental ill health built from how symptoms go together rather than from diagnostic categories. Problems sit on spectra that everybody falls somewhere along, arranged in a hierarchy from one general factor down to single symptoms.

The HiTOP model · ColinFreilich, CC BY-SA 4.0, via Wikimedia Commons · press to enlarge

The short measure

The HiTOP-BR (Brief Report; Simms et al., 2026): 45 statements about the last twelve months, on four points from "Not at all" to "A lot", none reversed, scored as the mean of each of six spectra. Its norms are the development sample's (N = 780), through the {hitop} R package, which also scores a saved file directly.

The six spectra of the HiTOP-BR · point at one

General factor (p) 12 items across the spectra

Externalizing 10 items

Internal­izing8 items Shown as Emotional IntensityMy moods were intense and unpredictable
Somato­form8 items Shown as Bodily ComplaintsI felt something was wrong with my body
Thought Disorder6 items Shown as Unusual ExperiencesMy fantasies felt very real to me
Detach­ment5 items Shown as SolitudeI was happiest when I was alone
Disinhib­ition9 items Shown as ImpulsivityI had trouble planning and keeping to schedules
Antagon­ism9 items Shown as DominanceI deserved special treatment

The two scales over the spectra are scored at analysis time.

Feedback expectations

What piloting the test should send back, and at what grain: excerpts of three pilots' notes on an earlier version, all addressed since.

Pilot 1Line by line: wording, punctuation, layout

Pilot 2Level by level, with how long each took

Pilot 3On a phone and a laptop

PositivesWhat worked, so that it is kept

Very cool user experience - definitely engaging!

I liked the hill metaphor and the overall use of images throughout the test to help understand and visualise the results

Compared to a normal test, it actually made me want to keep going and see what the next level was

NegativesWhat confused, broke or got in the way

For PHQ-4 section, the visual format of the question items is different to that in previous sections (although, may be intentional)

I get through hard times by finding what is funny in them. … what or who is them? No text above the question either.

Sometimes when I tried to tap on a specific number, it wouldn’t select the number I was actually pressing … I tried to select 40 but it would sometimes go to 47

Things to fixWhere it is, what it says, what it should say

contains → contain on first paragraph line Now, how have you been lately. requires a question mark rather than a full stop

maybe add PMDD in the what are you diagnosed with question especially if you're tracking mood

the Disagree and Agree labels were in the middle of the scale, underneath around 2 and 4, rather than at the ends underneath 0 and 6

Subjective reactionsHow it felt: the effort, the time, whether it rang true

On the glosses in brackets: I feel it comes across as a bit overwhelming

This level was a little hard so I didn't put lots of effort to answer the questions properly towards the end

I think it looks a bit nicer on the laptop, but it was still really good on the phone

Say which device and browser, name the level, and quote the words on screen: a note that can be found is a note that can be fixed.

TODOs

For next week

  1. Pilot the Abyss test realitybendinglab.com/TestYourself, with notes
  2. Pilot another experiment: Color illusions About 10–15 minutes
  3. Start thinking about what is interesting to measure in someone's usage of and beliefs about AI Where the BAIT is going