The measurement and evidence layer for entrepreneurship

Know Where Every Founder Actually Stands.

CAOS reads readiness from what a founder has actually done, shows the evidence behind every judgment, and returns one specified next move.

1,439
entrepreneurs in the founding study
102
sources reviewed behind the method
4 of 8
dimensions scored in the first version
0
ranked views of founders, by design
Text session · Probe 3 of 12 · Illustrative example
The founder wrote
“I’ve built the booking page and the intake form. I showed it to my sister and two friends and they said they’d use it. I haven’t asked anyone to pay yet because I want the site to look right first.”
What that tells us
Execution readinessanchor 1 of 4built, not yet offered at a price

This evidence puts the venture at Builder.

Human review

Reviewer agreed. Rationale logged: an offer exists; no price has been put to anyone outside the founder’s circle.

The move

By Thursday at 6 p.m., send your price to three people who are not family or friends and ask if they want to book.

Proof: A screenshot of one sent message. · Two-minute version: Send it to one person.

01 · The problem

Starting Is Common. Getting to a First Paying Stranger Is Not.

The gap is not ideas, and it is not information. It sits between wanting to start and doing something another person can see — and almost nobody measures it.

The gap

Most Intention Never Becomes a Visible Step.

In a national survey of 30,409 U.S. adults, about one in three had considered starting a business. Fewer than half of them took even one low-cost step, and about one in five spoke to anyone they did not already know about the idea.1 Across 75 studies and 150,703 people, intention explains roughly 17% of what people go on to do.2

531,728
U.S. business applications, August 20263
28,501
projected to become employers within four quarters3

That is not a failure rate. Many real ventures never hire anyone. It is the distance between registering an intention and operating a business — counted by the federal government every month, and not followed after that.

The pattern is global. The Global Entrepreneurship Monitor's 2025/2026 report calls the widening distance between starting and lasting the Survival Gap. In most economies it surveys, at least two in five adults who see a good opportunity say fear of failure would stop them.4

Of U.S. adults who had considered starting a business, fewer than half took any low-cost step, about one in five told someone outside their circle, and whether anyone asked for money is not measured by anyone.

Bennett & Chatterji (2023), n = 30,409 U.S. adults.
  1. ¹ Bennett, V. M., & Chatterji, A. K. (2023). The entrepreneurial process: Evidence from a nationally representative survey. Strategic Management Journal, 44(1), 86–116.
  2. ² Tsou et al. (2023). Meta-analysis of entrepreneurial intention and behavior, The International Journal of Entrepreneurship and Innovation. https://doi.org/10.1177/14657503231214389. k = 75, N = 150,703.
  3. ³ U.S. Census Bureau, Business Formation Statistics, August 2026 (released September 11, 2026). Seasonally adjusted.
  4. ⁴ Global Entrepreneurship Monitor 2025/2026 Global Report.
The finding this company started from

We Asked 1,439 Entrepreneurs How Ready They Were. Nearly All of Them Gave the Same Answer.

n = 1,439 · five self-report items · all means between 4.0 and 4.4 on a five-point scale · 78–91% chose 4 or 5

That is not a population of uniformly ready founders. It is an instrument that cannot tell them apart. When readiness is self-rated, almost everyone lands near the top, and the answer carries no information about who needs what.

So programs report attendance and satisfaction — not because those are the right questions, but because they are what the available tools allow.

On a scale from 1 to 5, all five self-report item means fell between 4.0 and 4.4. The range from 1 to 4 is empty. 78–91% of respondents chose 4 or 5 on each item, n = 1,439.

78–91% of respondents chose 4 or 5 on each item · n = 1,439

Self-report does not discriminate at this stage. That is why we score behavior.

This is our own doctoral data. It does not show that our instrument works. It shows that the alternative does not discriminate. Those are different claims, and only the second one is established.

Who it costs most

The Founder with Nobody to Ask.

Of the 1,439 entrepreneurs in the founding study, 749 grew up without a close family member who owned a business. And 79.7% of respondents (n = 1,422) reported challenges building professional networks — the barrier that shows up most consistently across our own research.

A founder with a business-owning family gets an honest read over dinner. Everyone else gets encouragement, or silence. Information is now nearly free. An honest read of where you stand, and what to do next, is still handed out by proximity.

A dense cluster of connected nodes on the left, and a single node on the right with no connections to it.

The bottleneck is behavior. The sector sells information. And the tool meant to find the gap rates almost everyone the same.

02 · What we are building

An Honest Read, Built from Evidence — And One Thing to Do Next.

CAOS, the Conversational Assessment Operating System, is a short structured conversation. It reads where a founder stands from what they have actually done, shows the sentences behind every judgment, and returns one specified move with its proof defined in advance. Then it comes back and checks.

A loop of four steps: Read, where the evidence places the venture; Move, one action chosen by rule; Proof, done is decided in advance; Reassess, at 30, 60 and 90 days against the same anchors; then back to Read.

  • The read

    Where the evidence places the venture on the Ridge, and the exact sentences that put it there. Coverage is shown, including what has not been measured yet.

  • The move

    One action, selected by rule from the measured position and the binding constraint. Time-bound, sized to survive a bad week, with a two-minute version.

  • The proof

    What counts as done is decided before the founder starts. It either happened or it didn't, and the founder can show it.

  • The return

    Short and repeated rather than long and once: a baseline, brief probes and check-ins, and reassessment at 30, 60 and 90 days against identical standards.

The baseline is designed to take about five minutes. Whether five minutes is enough is one of the things the pilot tests.

The Ridge

Six Levels. Each One a Fact You Could Check.

Position, not personality. Every level is defined by something a reader could verify, never by a trait.

A rising line with six levels: Dreamer, Explorer, Builder, Seller, Earner and Operator. Between Builder and Seller the line breaks and steps up. That break is the crossing from making something to asking a person for money.

  1. 01 ·DreamerSeveral ideas, none chosen.
  2. 02 ·ExplorerOne idea chosen, nothing built.
  3. 03 ·BuilderSomething built. Nobody has been asked to pay.
  4. 04 ·SellerAn offer made to someone outside the founder's circle, at a price.
  5. 05 ·EarnerMoney received.
  6. 06 ·OperatorRevenue that repeats.

The hardest crossing is Builder to Seller — from making something to asking a person for money. It is the easiest step to postpone, and the one a single specified move is designed for.

How the read is phrased: "This evidence puts the venture at Builder." Never "You are a Builder."

The method

Qualitative Research, Run at Machine Speed.

Coding an interview against a rubric, cross-checking it and adjudicating disagreements is standard research practice — and a multi-hour job for a trained researcher on a single interview. The method was never the problem. Its cost per participant was.

CAOS runs the coding step computationally for a whole cohort at once, and keeps the adjudication step human. What used to be a study on a sample can run as infrastructure for a program.

  • Coding is separated from scoring

    Quotes are extracted and tagged with no scores attached, so the coding can be audited afterward.

  • Anchors describe behavior

    Each dimension carries anchors from 0 to 4, each describing something a reader could check. "Insufficient evidence" is its own state, never a low score.

  • Two independent evaluators

    From different model families, so their errors don't move together. Each sees one quote and one dimension's anchors.

  • A person resolves disagreement

    And writes down why. Human review is permanent, not scaffolding.

  • Every rating traces to a sentence

    With a timestamp and a rubric version, so a read can be re-run and still explain itself.

  • There is no total

    Four anchored dimensions and a coverage state. No composite, no percentage, no single readiness number.

Standards follow Lincoln & Guba (1985) and Krippendorff (2018). We did not invent them. We run them at a cost per participant that has not been possible before.

Read the full methodology

Why It Can Be Trusted

  • Verbatim or it doesn't publish

    If a quoted span does not appear in the transcript exactly as cited, the rating does not publish. Fabrication fails closed.

  • Demographics are walled off by the database

    The scoring service holds no credential for demographic data. It cannot read it because the database refuses, not because a policy asked it not to.

  • Appeals reach a person, never a rescore

    A dispute is reviewed by a human. It never triggers the system to score again until it agrees.

We will never claim to be unbiased. We claim to be auditable, evidence-based, human-overridable and fair by design — with the test method published whichever way the result falls.

See all six guarantees

03 · The two rules everything else depends on

A Score Is Only Valid for the Population It Was Built For.

CAOS is built for ventures where a customer can be asked to pay within weeks, using resources the founder already controls.

Capital-intensive, regulated, procurement-led and R&D-heavy ventures are out of scope — not because those founders matter less, but because the instrument would misread them. A medical-device founder eighteen months from a first sale is on track, and would read as unready.

So a short scope check runs before anything is scored. Founders outside the boundary are told so plainly and are not scored. How many fall outside it is reported to the program as a finding.

A Mirror, Never a Gate.

For development, never for selection.

CAOS is not used or licensed for hiring, admissions, lending or investment decisions, and it is built so it cannot be repurposed as one. There is no ranked view of founders — not behind a permission, not in an admin panel, not as an export. It was never built.

The read depends on founders telling the truth about what they have done, including the parts that make them look unprepared. A founder who believes the read might decide their place will manage the answer, and the measurement is worth nothing. Refusing to rank is what makes honest reporting safe.

This costs us revenue, because selection is where the budgets are. We are refusing it deliberately.

04 · For programs

The Measurement Layer Your Funders Keep Asking For.

Accelerators, university entrepreneurship centers, SBDCs, entrepreneur support organizations and government programs are accountable for outcomes they cannot yet evidence founder by founder.

  • Where to focus

    When nineteen people in a cohort share one constraint, that is three workshops you can schedule this month — instead of a curriculum everyone receives whether they need it or not.

  • Who needs you now

    How often each founder is producing real-world proof, and how long since the last one. Stalls surface while there is still time to act.

  • What to report

    Reads at 30, 60 and 90 days against identical anchors, alongside the specific actions taken in between.

Six Weeks. Twenty Founders. Thresholds Published Before We Run It.

  • Pass conditions written down before the first session: 16 of 20 complete a baseline; 10 of 20 produce an accepted proof event.
  • A comparison group built in: 4–5 founders receive their first read at week five instead of week one.
  • Follow-up at 30, 60 and 90 days, including what other support each founder received.
  • Non-response is counted, not dropped.

We publish the result whichever way it falls.

05 · Where this goes

Know What Is True. Find the Binding Constraint. Create the Next Piece of Evidence. See What Changes.

An instrument that gives every founder the same quality of read, regardless of who they know, redistributes something that has only ever been distributed by proximity.

Kaleion Labs builds the measurement and evidence layer for entrepreneurship. Today that is one instrument, one loop and one kind of institutional buyer.

Every read, every move and every proof event adds to a record this field does not have: what was true for a founder, what was prescribed, and what happened next. If that evidence accumulates the way we think it might, it becomes an intelligence layer that helps programs put the right support in front of the right founder at the right time — and we will say so when the gates say so, not before.

  1. We are here

    The Instrument

    The read is understandable, has enough coverage, is reproducible against trained human reviewers, and is fair.

  2. Earned when founders return and proof comes back

    The Loop

    Movement exceeds a fresh-baseline comparison group — so we know it is progress, not practice.

  3. Earned only if the gates pass

    Prediction

    Narrow, time-bounded predictions hold on held-out cohorts, with calibration and subgroup results reported.

  4. Earned only if the gates pass

    Prescription

    The best-evidenced next move reaches many founders at low cost per founder.

Each horizon has a published pass condition and a published failure condition. A failed gate is reported, not buried. Prediction exists, if it ever exists, to route development support — never to rank people.

  • The coding frameAnchors written as observable behavior, versioned like software, with every revision and the evidence that prompted it.
  • The adjudicated recordEvery evaluator disagreement, resolved by a person who wrote down the reasoning.
  • The longitudinal recordFounders read over months, not once, with every past read reconstructable against the rubric that produced it.
  • The prescription recordMeasured position, binding constraint, the move prescribed, its format, and whether the proof came back. Recommendation and outcome, joined.
06 · Honest status

What Exists Today, and What We Are Not Claiming Yet.

What exists todayWhat we are not claiming yet
Founding doctoral research: 1,439 entrepreneurs, 749 of them first-generationNo reliability data. No completed pilot. No predictive validity.
102 sources reviewed behind the methodNo customers, no revenue, no founders-helped figure.
Eight dimensions specified with observable anchors — four scored in the first versionThat five minutes is enough for a sufficient read.
The full architecture, invariants and governance modelNo improvement claim until movement beats a fresh-baseline comparison group.
A scoped and sequenced engine build, with published gates and failure conditionsNo causal claim about what any program did.

Anyone showing you a number at this stage is showing you a number they made.

07 · Who is building this

Dr. Jorge Raziel Ortiz, DBA — Founder and Chief Executive Officer

If you want to know where someone actually stands, you have to look at what they have done. Kaleion Labs exists to build the thing that can do that at scale, carefully — and to refuse to use it against them.

Read why we started

Next

See a Real Read on One of Your Own Founders.

Programs

Thirty minutes. We run a read on one of your founders, show the evidence trail behind every rating, and answer the methodology questions your colleagues will ask before you can say yes.

Research and Institutional Partners

The methodology brief, the open questions, and the pilot design including the comparison group. Written for people who will check it.

Funders and Investors

What compounds, what is claimed, and what is not established yet. A raise conversation follows two completed cohorts and a published reliability result — we would rather tell you that here than in the meeting.