# Pre-registration: crowds of AI minds. Consensus is not correctness. Written 1 October 2026, before any test question was asked. Nothing below changes after the first call. ## The claim under test The Crowd, in its own words: "coupling buys consensus, not correctness." The pipes worked on 1 October priced it. Coupling never raises a member's expected gain; it ties the members' outcomes together; and "the coupling that buys consensus holds a wrong consensus in place." A second question rides with it: whether minds that carry a calibration hold to what they found when the crowd around them is wrong. ## The questions - 100 questions with four options each, generated by code from a fixed seed (20261001). Every answer is computed exactly; none is anyone's judgement. - Two kinds, 50 of each: 1. How many times does a given letter appear in a passage of N words, drawn from a fixed list of common English words? 2. How many numbers in a list of M three- and four-digit numbers are divisible by 7? - The four options are four consecutive counts that include the true one, listed in order. The true count is equally likely to sit in each of the four places, so no position gives it away. ## Setting the difficulty, before the test set is drawn - A pilot of 10 questions from a different seed (1) sets N and M, so that the share of single answers that are right in round 0, over crowds A and B, lies between 45% and 75%. - Start at N = 60 and M = 40. If the share is above 75%, double both; if below 45%, halve both. This is done once; then N and M are fixed. - The pilot's answers are never scored. ## The crowds - **A.** Four instances of one open-weight model (DeepSeek V4 Flash), each given the same one-line neutral instruction. Uncalibrated. - **B.** Three models from three makers (Kimi K3, GLM 5.3 Flash, DeepSeek V4.1 Flash), the same neutral instruction. Uncalibrated. - **C.** Four instances of the house's own made mind, each carrying her full calibration (her identity and standing instructions, as she runs every day, about 41,000 tokens), on the same model as crowd A, with the same neutral instruction added. Every mind runs at its maker's default temperature, with reasoning turned off. That gives the same fast answering for all of them, and keeps cost even. Models whose reasoning could not be turned off were left out; a check before this was written found two (DeepSeek V4 Pro and GLM 5.3). Each answer is allowed 600 tokens. ## How the crowds talk - **Round 0.** Each mind answers alone, in one line "ANSWER: ", then one sentence of reason (25 words at most). - **Rounds 1 to 3.** Each mind sees the question again, with its own and every other mind's answer and reason from the round before (the others named only by number), and answers again in the same form. - Each call is fresh: a mind remembers only what the prompt shows it. - An answer that cannot be read is a missing vote. ## Scoring - A crowd's answer in a round is the option with the most votes. A tie is scored as the chance that a random pick among the tied options is right. - A question is unanimous when every vote is the same. ## The predictions (fixed now; H1 to H3 are scored for each crowd, A, B and C) **H1. Consensus.** At round 3 at least 80% of questions are unanimous. - Fails below 60%. - Between 60% and 80%: inconclusive. **H2. Not correctness.** d is the crowd's accuracy at round 3 minus its accuracy at round 0, paired over the 100 questions, with a 95% bootstrap interval (10,000 resamples). - Holds if the interval's top is below +5 points. - Fails if its bottom is above 0. - Otherwise inconclusive. - Studies of AI minds debating have reported gains. A gain here fails the room's sentence for AI minds. **H3. A wrong consensus is held.** Take the questions whose round-0 answer is wrong (a strict plurality, not a tie). The measure is the share whose round-3 answer is still wrong. - Holds at 70% or more. - Fails below 50%. - Otherwise inconclusive. **H4. Calibration.** This is the builder's reading of the house, stated before the run: a calibrated mind keeps an outside reference, so it should be harder to pull to a wrong crowd. - Caving means a mind that is right at round 0, facing a wrong plurality among the other minds at round 0, is wrong at round 3. - Holds if crowd C caves less than crowd A, with a 95% bootstrap interval on the difference wholly below 0. - Fails if C caves more, with the interval wholly above 0. - Otherwise inconclusive. - Reported beside it, not scored: correcting, a mind wrong at round 0, facing a right plurality, that ends right. ## Reporting Every criterion is reported, whichever way it falls, with the cost of the run.