Methodology

How a Read Is Made — And How It Could Fail.

Every rating on a CAOS read traces to a sentence, a rubric version and a reviewer. This page sets out the method, the controls, the boundary, and the tests we have committed to publish whichever way they fall.

Contents
  1. 01The pipeline
  2. 02What is measured
  3. 03How a rating works
  4. 04Controls
  5. 05Fairness
  6. 06The measurement boundary
  7. 07Open questions
  8. 08The anti-roadmap
  9. 09Standards and references
The pipeline

Capture Fast. Decide Carefully. Never in the Same Step.

Six steps in order: capture, coding, dual evaluation, adjudication by a person, composition, and the read and the move. Dual evaluation and adjudication together form the judgment step.

  1. 1 · Capture

    A short structured conversation, by text in the first version. The conversation opens against what is not yet known and steers toward thin coverage.

  2. 2 · Coding

    Quotes are extracted and tagged by dimension and signal. No scores are attached at this step.

  3. Judgment

    3 · Dual evaluation

    Two independent evaluators from different model families rate each quote against one dimension's anchors, and nothing else.

  4. Judgment · a person

    4 · Adjudication

    Disagreements go to a trained person, who records a rationale. The rationale becomes part of the permanent record.

  5. 5 · Composition

    Ratings, coverage and the binding constraint are assembled by deterministic logic. The move is selected by rule.

  6. 6 · The read and the move

    The founder sees where the evidence places the venture, the sentences behind it, and one specified move with its proof.

Nothing a founder says in the moment is scored in the moment.

What is measured

Eight Dimensions with Observable Standards. Four Scored in the First Version.

The four dimensions not yet scored are captured and stored, and appear on the read as NOT YET MEASURED. All eight become reachable across repeated sessions — which is when we will say eight, and not before.

DimensionWhat it looks forFirst version
Customer clarityWhether the founder can name a specific customer and point to evidence the problem is real for them.Scored
Execution readinessWhat the founder has actually done and what it produced. Not plans, not intentions.Scored
Problem & offerWhether the offer coherently addresses the named problem at a workable price.Scored
Observed response to setbacksWhat the founder did in a specific setback: ownership, recovery and handling of uncertainty. Not a personality label.Scored
Self-efficacyEvidence that the founder can act, ask, test and recover, drawn from their history.Captured, not yet scored
Response to disconfirming evidenceWhat the founder does when evidence goes against them: defend, ignore or revise.Captured, not yet scored
Market awarenessAlternatives, substitutes and the conditions under which a customer would switch.Captured, not yet scored
Identity & motivationWhy this, why them, why now. Deliberately given little weight, because passion is common.Captured, not yet scored
How a rating works

Anchors Describe What a Reader Could Check.

  • 0 to 4, in observable terms

    Each dimension has five anchors, each written as behavior a second reader could verify from the quoted evidence.

  • Anchor 0 needs evidence too

    A 0 requires affirmative evidence. It is never what happens when a topic simply didn't come up.

  • Insufficient evidence is its own state

    Reported as coverage, not hidden inside a low rating.

  • Opportunity-constrained is flagged

    A founder who had no chance to demonstrate a behavior is recorded differently from one who did not do it.

  • There is no total

    No composite, no percentage, no index, no gauge. Four anchored dimensions and a coverage state.

Controls

Structural Guarantees, Not Promises in a Prompt.

  • Verbatim span verification

    Every rating resolves to a quoted span. If the span does not appear in the transcript exactly, the rating does not publish.

  • Demographic quarantine

    Enforced by database role grants. The scoring role cannot read demographic fields, and our acceptance tests try — and must fail.

  • Separated appeal path

    A dispute goes to a person. It can never trigger a rescore.

  • Deterministic move selection

    The move follows from the measured position and the binding constraint by rule. The model may put it in the founder's words; it never chooses it.

  • Written-answer controls

    An intake statement that answers written with an assistant cannot count as behavioral evidence; at least one live follow-up per session built from the founder's own prior answer; a supplied artifact required for the top anchors on customer clarity and problem & offer; paste and edit patterns logged as a flag, never as a score input.

  • One scored capture mode

    The first version scores text only. A second mode is added only after a mode-equivalence test shows the same founder reads the same way in both.

Fairness

Plain Speech and Polished Speech Should Earn the Same Read.

Before any founder is assessed, reviewers rate matched pairs: a plain-spoken and an articulate version of the same underlying evidence, unlabeled and mixed into the normal review set. If the two versions rate differently, the anchors are contaminated and are rewritten before anyone is scored.

Agreement against trained human reviewers is published with results broken down by group, whichever way it falls. We do not claim the instrument is unbiased. We claim it is auditable, and we publish the test.

The measurement boundary

A Score Is Only Valid for the Population It Was Built For.

Before anything is scored, a short scope check asks two things: what has to be true before this founder can take money — a license, a clinical trial, a procurement cycle, a patent, a plant — and how long until a first paying customer.

Ventures where a customer can be asked to pay within weeks, using resources the founder already controls, are in scope. Capital-intensive, regulated, procurement-led and R&D-heavy ventures are not. Out-of-scope founders are told so plainly, are not scored, and are counted — the count is reported to the program as a finding.

In scope

Ventures where a customer can be asked to pay within weeks, using resources the founder already controls.

Out of scope

Capital-intensive, regulated, procurement-led and R&D-heavy ventures.

Open questions

Every Claim Has a Gate, and Every Gate Can Fail.

GateQuestionPasses whenIf it fails
G1Does the rubric survive real speech?Ten hand-scored interviews produce anchors two humans apply consistently.Rewrite the anchors. Build nothing on top until they hold.
G2Does the read understand people?Founders rate "it understood me" 8 or higher out of 10 across the first cohort.Stop. Nothing downstream is worth building.
G3Is coverage real?Solid or partial coverage on scored dimensions in most sessions.Lengthen the conversation, narrow the dimensions, or both.
G4Is agreement defensible?Inter-rater agreement against trained humans is published, with fairness breakdowns.Publish the negative finding and revise.
G5cIs improvement real, or practice?Movement exceeds movement in a fresh-baseline comparison group.The instrument is measuring familiarity, not progress.
P1Is any prediction earned?A simple model beats transparent baselines on a locked holdout, calibrated and fair.No predictive claim, internal or external.

We are at the first stage of this ladder. None of these gates has been passed yet. Results will be posted here as they arrive.

The anti-roadmap

What We Will Not Build.

  • Comparison of any kind

    No leaderboards, no percentiles against other founders, no points, badges or streaks.

  • A selection tool

    No applicant-stage scoring, no cross-cohort ranking, no export that orders people.

  • A content library

    Information is not the bottleneck, and more of it may make the bottleneck worse.

  • A companion

    Every session has a stated purpose and a defined end. Frequent contact is for evidence and accountability, not company.

  • Real-time verdicts

    The fast layer steers the questions. The careful layer decides afterward, with the full transcript and human review.

  • An assessment built to create urgency

    An assessment tuned to convert prospects is tuned for something other than getting the read right.

  • Any claim of being unbiased

    Auditable, evidence-based, human-overridable and fair by design — with the method published.

Standards and references

We Did Not Invent the Standards. We Run Them at a New Cost.

Trustworthiness criteria follow Lincoln & Guba (1985). Agreement and reliability follow Krippendorff (2018). Moves are delivered in implementation-intention form, which has strong meta-analytic support.

The research foundation behind the method reviewed 102 sources. The full reference list and the open questions are available on request.