01
What external evidence supports
External research supports several underlying mechanisms, including bounded human supervision, automation bias, sociotechnical dependencies, and the importance of explicit governance.
Evidence Standard
Most vendors publish confident numbers with no method. LPM publishes grades, weak sources, correction logs, and the claims it has not proven yet.
Executive summary
Start here before opening grades, registers, correction logs, or sourcing debt.
01
External research supports several underlying mechanisms, including bounded human supervision, automation bias, sociotechnical dependencies, and the importance of explicit governance.
02
LPM organizes those mechanisms into a Seven Layer framework and uses modeled assumptions where enterprise-specific coefficients have not yet been measured.
03
The site does not yet establish general customer outcomes, universal benchmarks, or validated causal effect sizes for the complete LPM framework.
04
Contradictory external research, reproducible benchmark findings, and well-documented customer evidence should revise, narrow, or retire claims.
Primary or peer reviewed
Peer-reviewed research, primary regulation, or a named large-N study with a disclosed method.
Credible method, partial disclosure
Analyst research or a large survey from a credible firm, with the primary document in hand and method partially disclosed.
Directional only
Vendor content, trade press, or secondary reporting. Useful for direction, never load-bearing alone.
Evidence Register
Every external figure LPM uses should have a grade, a source, and a known weakness. This first public register includes the most important rows from the source library.
| Category | Claim | Grade | Source | Known weakness |
|---|---|---|---|---|
| Coordination | Decision effectiveness predicts financial performance across 760 companies. | B | Blenko, Mankins and Rogers, Decide and Deliver | Named large-N study with book-length method disclosure. It predates AI, which makes it a useful baseline. |
| Coordination | A vendor-run survey found AI increased communication volume while reducing clarity and policy alignment. | B | Axios HQ State of Internal Communications | Vendor-run survey. Useful for perception gaps, but self-reported deadline impacts need attribution. |
| Failure and adoption | About 95% of generative AI pilots were reported to produce no measurable profit and loss impact. | C | MIT NANDA via Fortune | Secondary reporting. The mechanism is more durable than the headline number. Do not lead with the 95%. |
| Failure and adoption | Over 40% of agentic projects were forecast to be cancelled by end of 2027 for governance, operationalisation, and value reasons. | C | Gartner via Reinventing.ai | Secondary reporting of an analyst prediction. Retrieve the primary before high-stakes external use. |
| Failure and adoption | About 74% of organisations plan agentic adoption, while about 21% report a mature governance model. | C | Deloitte via Zylos Research | The most-used number pair in the library and the weakest link in it. Retrieve the Deloitte primary. |
| Failure and adoption | A vendor security survey found 82% of enterprises had discovered unknown agents. | C | Gravitee State of AI Agent Security | Likely over-samples firms already worried about the problem. Use for direction, not magnitude. |
| Regulation | High-risk AI systems must be designed so they can be effectively overseen by natural persons while in use. | A | EU AI Act, Article 14 | Applies to EU-scoped high-risk systems. It does not mandate single-point accountability. |
| Regulation | Article 14 requires overseers to remain aware of automation bias, the tendency to over-rely on system output. | A | EU AI Act, Article 14 | Regulation names the failure mode, but does not prescribe a measurement method. |
| Regulation | Deployers must assign human oversight to natural persons with necessary competence, training, authority, and support. | A | EU AI Act, Article 26 | This is a deployer duty. It does not define how much oversight one person can carry. |
| Regulation | Non-compliance with deployer obligations under Article 26 is subject to fines up to EUR 15,000,000 or 3% of worldwide annual turnover. | A | EU AI Act, Article 99 | The higher EUR 35M / 7% tier applies to prohibited practices under Article 5, not this offence class. |
| Regulation | Certain biometric outputs require separate verification by at least two natural persons. | A | EU AI Act, Article 14 | This cuts against a naive single-owner reading. Verification is not ownership. |
| Regulation | New and materially changed Annex III high-risk systems face obligations from 2 August 2026. | A | EU AI Act, Article 113 | Systems already placed on the market before that date are only caught if significantly changed. |
| Supervisory control | Fan-out, the number of robots or systems one operator can supervise, is a computable function of activity time and interaction time. | A | Olsen and Wood, CHI 2004 | Developed for robots, not enterprise decision systems. The structure transfers. The coefficients must be re-derived. |
| Supervisory control | Span-of-control fan-out has been experimentally evaluated in cyber operations. | A | Human Span-Of-Control in Cyber Operations | Small-N experimental setting. Strong bridge, not a complete enterprise measurement. |
| Supervisory control | Human-in-the-loop review can degrade to rubber-stamping under volume and time pressure. | C | IBM Think summary | Vendor content. LPM's low-override WATCH band is a reasoned hypothesis until the benchmark tests it. |
Sourcing Debt
The register names the sources LPM must retrieve before making the claim louder.
The 74% adoption and 21% mature-governance pair is load-bearing in multiple documents and currently comes through secondary reporting.
This is the direct academic ancestor of Supervisory Control Capacity. It should be cited directly with a stable primary reference.
Parasuraman and Riley, Skitka, Mosier and Burdick, and Goddard et al. should carry the rubber-stamping argument instead of vendor summaries.
The current source is secondary reporting. Gartner forecast cancellation, while demotion is LPM language.
The 65% decision-speed claim has no retrievable title or URL yet.
The 40% decision-activity anchor behind the coordination-debt cost model has not been retrieved.
Applied evidence protocol
Permissioned implementation evidence will enter the public library only when readers can distinguish the observed conditions, the intervention, the outcome, and the limits of the conclusion.
Name what may be published, anonymized, or kept private before the work begins.
Record the workflow, operating conditions, measures, and evidence quality before intervention.
Identify the exact ownership, decision, information, platform, governance, or AI change made.
Report observed movement, time period, confounders, unintended effects, and missing evidence.
State what the case supports, what it cannot establish, and what would change the conclusion.
Attach an accountable reviewer, publication date, version, correction path, and next review date.
Corrections
A research organisation should show the mistakes it caught in its own work.
Correction 1
Article 99(4)(e) sets deployer non-compliance at EUR 15M or 3%. The 35M / 7% tier applies to prohibited practices under Article 5.
Correction 2
Article 111(2) grandfathers systems already placed on the market unless they are significantly changed.
Correction 3
It does not. Article 14(5) can require two natural persons to verify certain outputs. LPM must distinguish verification from ownership.
Correction 4
Hybrid grades are removed. Grade C stays Grade C, even when the sentence would be sharper with a stronger grade.
Correction 5
Modeled numbers now require a visible modeled label and a path to assumptions.
What We Do Not Know
The point is not to make every claim sound equally strong. The point is to make the difference visible.
Claims
This is the behavior the evidence primitives enforce across research pages.
Claim
Counterargument
This makes some public claims look softer, especially market-size and adoption claims. That is the point. Weak evidence should remain visible until stronger sources replace it.
Claim
Counterargument
The underlying supervisory literature is strong, but LPM's enterprise-agent coefficients are not yet measured. The page must separate mechanism from coefficient.
Falsified by: A calibrated benchmark showing that enterprise agent supervision does not degrade with complexity, risk, reversibility, or workload.
Claim
Counterargument
Some buyers will ignore the grades and only want a confident answer. LPM is choosing the buyer who notices the grade chip.
Underlying Sources
The raw research documents are not rendered as marketing copy. They are the working papers beneath the public research pages.