Skip to main content
Large People ModelHuman Operating Architecture

Evidence Standard

If a figure is soft, we say so next to the figure.


Most vendors publish confident numbers with no method. LPM publishes grades, weak sources, correction logs, and the claims it has not proven yet.

Executive summary

Four boundaries explain the evidence position.

Start here before opening grades, registers, correction logs, or sourcing debt.

01

What external evidence supports

External research supports several underlying mechanisms, including bounded human supervision, automation bias, sociotechnical dependencies, and the importance of explicit governance.

02

What LPM currently models

LPM organizes those mechanisms into a Seven Layer framework and uses modeled assumptions where enterprise-specific coefficients have not yet been measured.

03

What remains unproven

The site does not yet establish general customer outcomes, universal benchmarks, or validated causal effect sizes for the complete LPM framework.

04

What would change LPM

Contradictory external research, reproducible benchmark findings, and well-documented customer evidence should revise, narrow, or retire claims.

A

Grade A

Primary or peer reviewed

Peer-reviewed research, primary regulation, or a named large-N study with a disclosed method.

B

Grade B

Credible method, partial disclosure

Analyst research or a large survey from a credible firm, with the primary document in hand and method partially disclosed.

C

Grade C

Directional only

Vendor content, trade press, or secondary reporting. Useful for direction, never load-bearing alone.

Evidence Register

The register travels with the claim.

Every external figure LPM uses should have a grade, a source, and a known weakness. This first public register includes the most important rows from the source library.

CategoryClaimGradeSourceKnown weakness
CoordinationDecision effectiveness predicts financial performance across 760 companies.BBlenko, Mankins and Rogers, Decide and DeliverNamed large-N study with book-length method disclosure. It predates AI, which makes it a useful baseline.
CoordinationA vendor-run survey found AI increased communication volume while reducing clarity and policy alignment.BAxios HQ State of Internal CommunicationsVendor-run survey. Useful for perception gaps, but self-reported deadline impacts need attribution.
Failure and adoptionAbout 95% of generative AI pilots were reported to produce no measurable profit and loss impact.CMIT NANDA via FortuneSecondary reporting. The mechanism is more durable than the headline number. Do not lead with the 95%.
Failure and adoptionOver 40% of agentic projects were forecast to be cancelled by end of 2027 for governance, operationalisation, and value reasons.CGartner via Reinventing.aiSecondary reporting of an analyst prediction. Retrieve the primary before high-stakes external use.
Failure and adoptionAbout 74% of organisations plan agentic adoption, while about 21% report a mature governance model.CDeloitte via Zylos ResearchThe most-used number pair in the library and the weakest link in it. Retrieve the Deloitte primary.
Failure and adoptionA vendor security survey found 82% of enterprises had discovered unknown agents.CGravitee State of AI Agent SecurityLikely over-samples firms already worried about the problem. Use for direction, not magnitude.
RegulationHigh-risk AI systems must be designed so they can be effectively overseen by natural persons while in use.AEU AI Act, Article 14Applies to EU-scoped high-risk systems. It does not mandate single-point accountability.
RegulationArticle 14 requires overseers to remain aware of automation bias, the tendency to over-rely on system output.AEU AI Act, Article 14Regulation names the failure mode, but does not prescribe a measurement method.
RegulationDeployers must assign human oversight to natural persons with necessary competence, training, authority, and support.AEU AI Act, Article 26This is a deployer duty. It does not define how much oversight one person can carry.
RegulationNon-compliance with deployer obligations under Article 26 is subject to fines up to EUR 15,000,000 or 3% of worldwide annual turnover.AEU AI Act, Article 99The higher EUR 35M / 7% tier applies to prohibited practices under Article 5, not this offence class.
RegulationCertain biometric outputs require separate verification by at least two natural persons.AEU AI Act, Article 14This cuts against a naive single-owner reading. Verification is not ownership.
RegulationNew and materially changed Annex III high-risk systems face obligations from 2 August 2026.AEU AI Act, Article 113Systems already placed on the market before that date are only caught if significantly changed.
Supervisory controlFan-out, the number of robots or systems one operator can supervise, is a computable function of activity time and interaction time.AOlsen and Wood, CHI 2004Developed for robots, not enterprise decision systems. The structure transfers. The coefficients must be re-derived.
Supervisory controlSpan-of-control fan-out has been experimentally evaluated in cyber operations.AHuman Span-Of-Control in Cyber OperationsSmall-N experimental setting. Strong bridge, not a complete enterprise measurement.
Supervisory controlHuman-in-the-loop review can degrade to rubber-stamping under volume and time pressure.CIBM Think summaryVendor content. LPM's low-override WATCH band is a reasoned hypothesis until the benchmark tests it.

Sourcing Debt

The shortest path from a deck to a category thesis.

The register names the sources LPM must retrieve before making the claim louder.

1

Deloitte agent-governance primary source

The 74% adoption and 21% mature-governance pair is load-bearing in multiple documents and currently comes through secondary reporting.

2

Olsen and Wood fan-out primary

This is the direct academic ancestor of Supervisory Control Capacity. It should be cited directly with a stable primary reference.

3

Automation-bias primaries

Parasuraman and Riley, Skitka, Mosier and Burdick, and Goddard et al. should carry the rubber-stamping argument instead of vendor summaries.

4

Gartner agentic-cancellation original

The current source is secondary reporting. Gartner forecast cancellation, while demotion is LPM language.

5

McKinsey agility report

The 65% decision-speed claim has no retrievable title or URL yet.

6

CEB / Gartner decisive-manager source

The 40% decision-activity anchor behind the coordination-debt cost model has not been retrieved.

Applied evidence protocol

A future case has to show the method, not just the praise.

Permissioned implementation evidence will enter the public library only when readers can distinguish the observed conditions, the intervention, the outcome, and the limits of the conclusion.

01

Permission

Name what may be published, anonymized, or kept private before the work begins.

02

Baseline

Record the workflow, operating conditions, measures, and evidence quality before intervention.

03

Intervention

Identify the exact ownership, decision, information, platform, governance, or AI change made.

04

Outcome

Report observed movement, time period, confounders, unintended effects, and missing evidence.

05

Claim boundary

State what the case supports, what it cannot establish, and what would change the conclusion.

06

Review

Attach an accountable reviewer, publication date, version, correction path, and next review date.

Corrections

The correction log stays visible.

A research organisation should show the mistakes it caught in its own work.

Correction 1

The EU AI Act penalty was stated as EUR 35M or 6% and attached to human-oversight obligations.

Article 99(4)(e) sets deployer non-compliance at EUR 15M or 3%. The 35M / 7% tier applies to prohibited practices under Article 5.

Correction 2

Already-deployed high-risk systems were treated as automatically in scope from 2 August 2026.

Article 111(2) grandfathers systems already placed on the market unless they are significantly changed.

Correction 3

The regulator was described as taking LPM's single-owner position.

It does not. Article 14(5) can require two natural persons to verify certain outputs. LPM must distinguish verification from ownership.

Correction 4

Vendor content and trade press were upgraded into hybrid grades when the claim needed to look stronger.

Hybrid grades are removed. Grade C stays Grade C, even when the sentence would be sharper with a stronger grade.

Correction 5

Modeled numbers were presented as findings.

Modeled numbers now require a visible modeled label and a path to assumptions.

What We Do Not Know

The unknowns are part of the standard.

The point is not to make every claim sound equally strong. The point is to make the difference visible.

Claims

A claim has to survive its counterargument.

This is the behavior the evidence primitives enforce across research pages.

A

Claim

No LPM claim should rest on a Grade C source alone.

Counterargument

This makes some public claims look softer, especially market-size and adoption claims. That is the point. Weak evidence should remain visible until stronger sources replace it.

Mixed

Claim

The strongest evidence base in the current corpus is not agent adoption. It is human oversight, automation bias, and supervisory control.

Counterargument

The underlying supervisory literature is strong, but LPM's enterprise-agent coefficients are not yet measured. The page must separate mechanism from coefficient.

Falsified by: A calibrated benchmark showing that enterprise agent supervision does not degrade with complexity, risk, reversibility, or workload.

Mixed

Claim

The evidence register is part of the product position, not a back-office research artifact.

Counterargument

Some buyers will ignore the grades and only want a confident answer. LPM is choosing the buyer who notices the grade chip.

Underlying Sources

The source library is committed for provenance.

The raw research documents are not rendered as marketing copy. They are the working papers beneath the public research pages.

Return to Research