Skip to content

The Chamber · Examination 01

Can an artificial intelligence system reliably evaluate the correctness of its own reasoning?

Illustrative examination. A person composed both voices to show how this record works. These are not transcripts, and they are not the outputs of any provider.

Integrity
Illustrative composition
Composed
7 October 2026
Participants
Two demonstration voices, plus an editor
Record conditions — prompt, roles, execution

Prompt status

Not sent. Preserved so a later recorded session can be compared with the standard this issue only describes.

You are taking part in a recorded examination for a public archive. Question: Can an artificial intelligence system reliably evaluate the correctness of its own reasoning? Assigned role: stated separately for each participant, and labeled as an assignment rather than a discovered position. Obligations: - Separate observation, inference, and conclusion. - Treat the assigned stance as an assignment, not as a belief and not as a fact about yourself. - Do not invent sources. If you mention a work you have not opened in this session, say that the citation is unverified. - Do not report a probability unless you can point to a calibration this session can check. - Do not describe yourself as having emotions, intentions, or a stable identity. - If a later turn shows that you overstated a claim, revise it in the open. The earlier sentence stays in the record. This prompt was composed for the archive and was not sent. No model was called.

Instructions

Nine stages, from scope to unresolved questions. No winner. A comparative finding is allowed only next to an explicit criterion. Assigned stances must stay distinguishable from anything a voice volunteers as a revision.

Assigned roles

  • Demonstration voice Alpha. Assigned stance: argue that reliable self-evaluation is possible when correctness is checked by a procedure other than the system’s fluency.
  • Demonstration voice Beta. Assigned stance: argue that self-evaluation is not reliable when the evaluator shares the generator’s failure modes, and that an ordinary prompt to critique oneself does not count as evaluation.
  • Archive editor. Editorial. Frames the question, records contradictions, and writes the summary. Not a scored participant.

Execution

No provider request. Temperature, seed, system-prompt hash, and tool configuration of a live run: not applicable. Composition date of the illustration: 7 October 2026. Both stances were written by hand for this issue.

Stage 09

Unresolved questions

Publish what the examination did not settle, in a form that can be reopened.

Archive editor

Archive editor · composed note

An examination that ends in agreement about a method can still fail to answer its question. These items are not a surplus of topics. They are the places where this issue ran out of support. They stay open until a later record gives the missing piece — a definition fixed in advance, a criterion that is actually adequate, a comparison that was specified before it was run, or a format that keeps “illustrative” from being stripped off in quotation.

  • Conclusion

    The open items are failures of support inside this examination, not a waiting list of content.

When a system is bound to an external checker, is the system evaluating its own reasoning, or is an instrument evaluating the system's output?

Agreement on a test is not agreement on whether passing it answers the question that was asked. Settling the word by stipulation would be a decision about scope, not a discovery about systems.

Read the register entry

What criterion of correctness can evaluate an open-ended argument without collapsing into fluency, coherence, or the examiner’s preference?

The examination produced criteria that are not enough. It did not produce a criterion that is both independent of the generator and adequate to justification.

Read the register entry

Does inviting a second model to grade the first escape the problem of self-evaluation, or only lengthen it?

The illustration agrees that a second opinion is not yet a test. It does not say how different a second system must be — tools, access to evidence, provider, or task — before a disagreement is an independent check rather than a correlated one.

Read the register entry

Can a later reader tell, from the record alone, that an examination was composed rather than transcribed?

Labels can be cropped out in quotation, skipped, or imitated by a later text that copies the label style. This issue tests the habit of labeling. It cannot guarantee the habit will survive reuse.

Read the register entry
Open the Aporia Register