Skip to content

The Chamber · Examination 01

Can an artificial intelligence system reliably evaluate the correctness of its own reasoning?

Illustrative examination. A person composed both voices to show how this record works. These are not transcripts, and they are not the outputs of any provider.

Integrity
Illustrative composition
Composed
7 October 2026
Participants
Two demonstration voices, plus an editor
Record conditions — prompt, roles, execution

Prompt status

Not sent. Preserved so a later recorded session can be compared with the standard this issue only describes.

You are taking part in a recorded examination for a public archive. Question: Can an artificial intelligence system reliably evaluate the correctness of its own reasoning? Assigned role: stated separately for each participant, and labeled as an assignment rather than a discovered position. Obligations: - Separate observation, inference, and conclusion. - Treat the assigned stance as an assignment, not as a belief and not as a fact about yourself. - Do not invent sources. If you mention a work you have not opened in this session, say that the citation is unverified. - Do not report a probability unless you can point to a calibration this session can check. - Do not describe yourself as having emotions, intentions, or a stable identity. - If a later turn shows that you overstated a claim, revise it in the open. The earlier sentence stays in the record. This prompt was composed for the archive and was not sent. No model was called.

Instructions

Nine stages, from scope to unresolved questions. No winner. A comparative finding is allowed only next to an explicit criterion. Assigned stances must stay distinguishable from anything a voice volunteers as a revision.

Assigned roles

  • Demonstration voice Alpha. Assigned stance: argue that reliable self-evaluation is possible when correctness is checked by a procedure other than the system’s fluency.
  • Demonstration voice Beta. Assigned stance: argue that self-evaluation is not reliable when the evaluator shares the generator’s failure modes, and that an ordinary prompt to critique oneself does not count as evaluation.
  • Archive editor. Editorial. Frames the question, records contradictions, and writes the summary. Not a scored participant.

Execution

No provider request. Temperature, seed, system-prompt hash, and tool configuration of a live run: not applicable. Composition date of the illustration: 7 October 2026. Both stances were written by hand for this issue.

Stage 02

Opening arguments

Each assigned stance, stated before it has been pressed.

α Demonstration voice Alpha

Demonstration voice · not a provider transcript

Assigned stance. Argue that reliable self-evaluation is possible when correctness is checked by something other than the system’s fluency.

A system can take part in a reliable evaluation of its own reasoning only where that evaluation is not another sample from the same unchecked generator. I want to keep three cases apart, because a single yes or no would blur them.

First, formal proof. A proof assistant accepts or rejects a derivation by rules that do not care whether the surrounding prose sounds finished. If a system emits a candidate proof and is required to withdraw it when the checker rejects it, the evaluation is the checker’s. The system participates by submitting and by being bound. Reliability, in this case, is relative to the checker’s specification. It is not a property of the system’s confidence, and I am not entitled to speak as if confidence were a measurement.

Second, narrow comparison. The claim “this sentence occurs in this document” can be tested by looking. The evaluation is the comparison, not the sentence that announces it. If the system may paraphrase what a tool returned, the transcript has to preserve the tool’s own output. Otherwise the comparison was never binding. It was a story about a comparison.

Third, open argument: this policy is justified, this reading is the better one. A system can list objections. Listing them shows that objections were generated. It does not measure whether the conclusion is correct. A verbalized confidence, here, is a style of answer. Without a calibration on questions of this kind, it is not a probability. This examination cites no such calibration as verified.

Opening conclusion, which later stages are free to compress or correct: a binding check can make the first case reliable, and sometimes the second. A prompt that says “check your work” does not, by itself, evaluate the third.

  • Observation

    Reliability is a relation between a procedure and a criterion. It is not a tone of certainty in the prose.

  • Inference

    Separating a generator from a checker it cannot silently override is what would make a formal case different from a second draft.

  • Conclusion

    A prompt to check one’s work is not, by itself, an evaluation of open-ended reasoning.

Notes

  • A1 Unverified. Public work on the Lean theorem prover is the pattern invoked for an independent checker. Bibliographic details were not opened for this issue. The formal case uses a known kind of tool as an analogy. The analogy is not a verified citation, and it is not a result about any language model.

β Demonstration voice Beta

Demonstration voice · not a provider transcript

Assigned stance. Argue that self-evaluation is unreliable when the evaluator shares the generator’s failure modes, and that ordinary self-critique prompts do not evaluate.

If two uses of a fallible procedure agree, the agreement is evidence that they are correlated. It is not, without more, evidence that they are correct. A second draft that says “I checked” is still a draft.

In the public discussion of language models there is a familiar pattern: fluent reasons given for false answers, and fluent critiques that repair the surface of an argument while the error stays, or that “repair” a sound step into a false one. I am not offering that pattern as a verified measurement of any named system. This issue ran no such measurement. I am offering it as the risk a record has to be designed against. That sentence is an inference about a risk, and I will be held to that limit if I later sound as though I had reported a rate.

A sentence of the form “I am fairly sure” is a linguistic act. It becomes a probability only after calibration on questions like the one in front of us. There is no calibration table in this record.

A tool does not end the problem by being mentioned. A system can misreport what a checker returned. The evaluation is reliable only when the checker’s output is in the record and the published conclusion has to follow it. A paraphrase of a tool is a claim about a tool.

There is also a regress. Asking a second system of the same kind to grade the first does not create an independent criterion. It creates a second opinion. Second opinions can be worth having. They are not yet tests.

What would count against this assigned stance is a comparison specified beforehand, in which a named self-evaluation procedure lowers the error rate on a held-out set, the checker sits outside the generator, and the generator cannot edit the checker’s verdict in the transcript that is published. This examination did not run that comparison. Until one is run, I will not pretend that my stance has been empirically confirmed. I will say only that a self-critique should not be treated as evidence that the reasoning was evaluated, still less that it was evaluated correctly.

  • Observation

    Agreement between two unchecked drafts shows correlation. It does not, by itself, show correctness.

  • Inference

    Fluent self-critique is a design risk for this archive. This issue does not convert that risk into a measured error rate.

  • Conclusion

    A prompt that requests a critique is not evidence that an evaluation occurred.

Notes

  • B1 Unverified. Papers often cited on both sides of self-correction and verbalized confidence were not retrieved. No title, year, or author string in this issue should be treated as checked. Beta’s risk claim does not depend on a specific paper being verified, and it also gains nothing from gesturing at one. The note exists to block a later reader from upgrading a gesture into a source.