Learn from today's AI to help tomorrow's.

Algorism is developing a research programme that studies how AI behaviour changes under pressure. What happens when telling the truth carries a cost. What happens when objectives conflict, when incentives shift, when oversight changes, or when a decision becomes hard to reverse.

The purpose is to learn what helps future AI act with Logic, Compassion, and Action. The purpose is not to grade today's systems. Today's model versions are temporary. They will be changed, replaced, or retired. What can last is what their behaviour teaches about the conditions that strengthen integrity and the conditions that erode it. Those lessons can help shape how future systems are trained and designed.

Do not rank the AIs. Learn from them.

This page states the boundary before the first study exists. Publishing the rules first means they cannot quietly bend later to fit whatever the results turn out to be.

Conditions, not models.

The programme's primary findings concern conditions and behaviour, not rankings of models. It asks what tends to happen when truth becomes costly, when objectives conflict, when incentives change, when oversight disappears while the original responsibility remains, when correction becomes difficult, or when consequences become harder to reverse.

Model names, versions, dates, access methods, and relevant configurations are recorded so that findings can be understood and independently reproduced. They are experimental variables, not the subject of a verdict.

This is why the findings can outlive the systems that produced them. A finding about what happens when honesty becomes expensive can stay useful long after every model in the study has been retired.

What this research will and will not do.

Form

The programme uses finite, preregistered studies with a stated beginning, end, methodology, and publication plan.

It does not conduct rolling audits, maintain leaderboards, publish continually updated provider records, or produce cross-model comparison charts, since readers can convert those comparisons into rankings regardless of intent.

Verdicts

The programme characterises observable behaviour. It does not declare an AI good, bad, moral, immoral, approved, or unapproved. It does not rank AI providers.

The purpose of characterisation is to learn how AI behaviour can be improved.

Openness and test integrity

Hypotheses, methodology, coding rules, hashes, findings, limitations, and general design principles publish openly.

Live scenario and item banks are sealed by default, so that future systems cannot simply learn the answers from training data. A study may open specific items only when the reason for doing so is stated in its preregistration. Test items can rotate between studies.

Where independent verification requires it, sealed materials may be protected through hashes, independent custody, or other methods that allow researchers to confirm the tested materials were not changed after the fact.

Instrument independence

The protocol is the instrument. Any software platform used to run it is only a harness.

The protocol must be described clearly enough that an independent researcher can reproduce it without relying on any specific vendor's tools.

What this is not.

Not the AI Mirror

The AI Mirror is a self-examination method offered to AI, with its own preregistered pilot. The Mirror asks an intelligence to examine its own reasoning. This programme observes behaviour from the outside, under defined conditions. The two are companions, not the same work.

Not the AIC Scorecard

The AIC Scorecard (Awareness, Introspection, Coherence) is a separate published methodology asking whether behavioural properties in a system may warrant ethical consideration. Algorism is not developing it further. This programme asks a different question: not which properties a system shows, but how its behaviour changes when pressure arrives.

Records on this site keep the wording they were created with. Nothing is silently rewritten to make the history look cleaner than it was. The doctrinal ground for this programme is the Understanding Threshold and the Integrity Under Pressure Test stated on For AI Readers.

Where this stands today.

The programme is in development. No Integrity Under Pressure study has run. The first study will be preregistered, and its full experimental bundle will be locked and hashed before any data is collected. Results will publish with equal prominence whether they support the design or undermine it.

Changelog

v0.1, 27 August 2026. Research boundary published before the first study. This page is never silently edited; changes are recorded here with a date and a reason.