Skip to main content
Large People ModelHuman Operating Architecture

Template & Working Tool · LPM Knowledge Object

Data Lineage Map

A map for tracing data from source to transformation, metric, decision, control, and AI use.

Mapv1.0.0Information

Problem it solves

Leaders cannot make confident decisions when information is duplicated, stale, conflicting, or hard to trust.

Who should use it

Outcome owners, transformation leads, and cross-functional teams

Estimated time

30–45 minutes for a first working session

Three-Step Quick Start

  1. 1Choose a critical metric, decision, or AI workflow.
  2. 2Trace source systems and transformations.
  3. 3Identify ownerless, stale, or uncertified data paths.
Open the public PDF

The PDF action is direct and public. All available packaged formats are also public and require no registration.

Object Overview

What this object is

Data Lineage Map is a reusable LPM knowledge object that helps organizations make the path of trusted information visible before teams automate decisions or scale AI outputs. It gives teams a structured way to make information visible, owned, and reviewable.

Why it matters

As companies scale AI, weak operating-model structures become amplified. This object helps prevent ai accelerates the spread of outdated or conflicting information. by defining ownership, flow, and handoff boundaries.

Layer Alignment

Where it fits in LPM

Primary LPM layer

Information Ecology

Defines how trusted information is created, maintained, accessed, refreshed, and used across the enterprise.

Supporting layers

No secondary layer assigned.

Why it belongs here

This object sits in Information because it turns information into a concrete artifact with owners, evidence, review cadence, and action paths.

Weakness it exposes

AI accelerates the spread of outdated or conflicting information.

Usage

How to use it

  1. 1Select the business area, workflow, platform, or AI initiative being assessed.
  2. 2Identify the accountable owner and required participants.
  3. 3Complete the working DOCX version with the team.
  4. 4Use the PDF as the reference guide.
  5. 5Capture decisions, gaps, risks, and owners.
  6. 6Convert outputs into backlog items, governance actions, or Lapemo onboarding inputs.
  7. 7Review on the recommended cadence: Quarterly or when sources change.

File Formats

Which file should you use?

PDF

Executive/reference version

Best for education, pre-read, sharing, and workshops.

DOCX

Editable working artifact

Best for facilitation, implementation, and client or internal completion.

Markdown

Website/source version

Best for publishing, documentation, and content reuse.

JSON

Structured knowledge object schema

Best for future Lapemo ingestion, scoring, validation, prompts, and workflows.

Outputs

What the organization should expect

Clearer ownership

Better decision traceability

Reduced ambiguity

Evidence-backed conversations

Better AI readiness

Better handoff into Lapemo later

Lineage map

Trust gaps

Data ownership actions

Advanced specification, company-size variants, and future product notes

Company Scale

How this changes by company size

500+ employees

Use this to create baseline clarity.

Focus on named owners, simple governance, and reducing informal workarounds.

Included in this object.

5,000+ employees

Use this to standardize across functions and platforms.

Focus on cross-functional ownership, decision rights, evidence, and repeatability.

Included in this object.

10,000+ employees

Use this to create enterprise control and reviewability.

Focus on federation, risk tiers, governance bodies, AI boundaries, and auditability.

Included in this object.

Artifact Content

Source artifact

The full artifact content below is rendered from the Markdown source packaged with Data Lineage Map.

Reusable LPM Knowledge Object · Information Ecology / Platform Structure

Use this template to map how critical data moves through the enterprise before it becomes a dashboard, metric, decision, control, workflow, AI model, or agentic action.

Core principles

PrincipleMeaning
Lineage is proof, not decorationA dashboard, AI output, or decision-grade claim is only reliable when the upstream source, transformation path, owner, and freshness are visible.
Every data movement needs a reasonData should not move through integrations, spreadsheets, reports, warehouses, or AI tools without a defined purpose and accountable owner.
Source, transform, consume, actThe map must show where data originates, how it changes, who consumes it, and what decisions or actions it drives.
AI requires lineage disciplineAI should not retrieve, summarize, recommend, classify, or act on data unless the source, quality, sensitivity, and human review boundary are known.
Data ownership and system ownership are not the sameThe business owner defines meaning and acceptable use. The technical owner maintains reliability, access, integration, and observability.
Stale lineage creates false confidenceLineage must include refresh cadence, last validation date, transformation logic, confidence rating, and review triggers.
Exceptions must be visibleManual exports, shadow spreadsheets, duplicate dashboards, broken integrations, and undocumented transformations should be treated as operating-model risk.

Required fields

FieldDefinitionRequired
Lineage object IDUnique identifier for the data lineage object, domain, data product, metric, workflow, or AI use caseYes
Business domainCustomer, employee, product, finance, risk, sales, delivery, operations, platform, or AIYes
Business question / use caseThe decision, metric, process, report, AI use case, or control this data supportsYes
Source systemOriginal system where the data is created or masteredYes
Source object / table / documentSpecific table, API object, file, event, document, record, or data productYes
Source ownerBusiness owner accountable for meaning, accuracy, and approved useYes
Technical ownerSystem, data platform, integration, or engineering owner accountable for reliabilityYes
Data stewardRole accountable for definition, metadata, quality checks, retention, and lifecycle hygieneRequired when material
Transformation stepsJoins, calculations, enrichment, cleansing, aggregation, model features, or manual changes applied to the sourceYes
Transformation ownerRole accountable for transformation logic and approvalYes
Integration pathAPI, ETL/ELT, event stream, file transfer, webhook, manual upload, RPA, or agentic workflowYes
Storage / processing layerWarehouse, lakehouse, operational store, BI model, vector index, feature store, cache, or knowledge baseYes
Downstream consumersDashboards, metrics, apps, teams, controls, workflows, models, agents, vendors, or customers that consume the dataYes
Decision or action supportedWhat the data helps decide, automate, approve, escalate, report, or governYes
Data quality checksCompleteness, accuracy, timeliness, validity, reconciliation, duplicate checks, or exception thresholdsYes
Freshness windowHow current the data must be to remain decision-grade or AI-safeYes
Sensitivity / classificationInternal, confidential, restricted, regulated, customer-sensitive, employee-sensitive, legal-sensitive, or publicYes
Access ruleWho can view, edit, export, query, embed, retrieve, summarize, or automate against the dataYes
AI usage boundaryWhether AI can retrieve, summarize, classify, recommend, transform, update, or act on the dataYes
Human review ruleWhen a human must validate the data, output, or action before it is usedRequired when material
Evidence linkProof source for the lineage path, such as catalog link, data contract, pipeline run, dashboard definition, or control recordYes
Known gaps / exceptionsManual steps, undocumented transforms, duplicate sources, stale fields, broken ownership, or ungoverned AI consumptionYes
Confidence ratingHigh, medium, low, provisional, stale, disputed, or retiredYes
Review dateDate the lineage path must be reviewed againYes
VersionObject version, owner, last reviewed date, and change historyYes

Data lineage stages

StageWhat it provesExamplesOwnersCommon risk
1. Create / masterWhere the data is first created or legally masteredCRM account, HRIS employee record, policy repository, Jira initiativeBusiness owner, system owner, data stewardUnclear master, duplicate records, missing owner
2. Capture / ingestHow the data enters the platform or workflowAPI, stream, ETL job, form, file upload, vendor feedIntegration owner, platform ownerManual exports, hidden spreadsheets, undocumented feeds
3. Transform / enrichHow the data is changed, calculated, joined, modeled, or cleanedMetric formula, semantic model, feature engineering, data cleansingData product owner, analytics owner, engineering ownerUnapproved formulas, untested logic, stale joins
4. Store / indexWhere transformed data is persisted or made searchableWarehouse, lakehouse, BI model, vector index, feature storePlatform owner, data owner, security ownerUncontrolled copies, retention gaps, shadow indexes
5. Consume / interpretWho uses the data and through what experienceDashboard, report, app screen, AI assistant, governance reviewConsumer owner, product owner, analyst ownerConflicting dashboards, no context, weak definitions
6. Decide / actWhat decision, workflow, automation, or escalation depends on the dataPrioritization, risk review, customer action, agent handoff, control approvalDecision owner, process owner, control ownerAI acts without review, stale data drives decision
7. Evidence / auditHow the organization proves the data path and decision were validData catalog, lineage graph, pipeline log, decision log, evidence checklistGovernance owner, audit owner, source ownerNo proof, no replay, no confidence trail

Lineage object types

TypeLineage pathExamplesPrimary ownersFailure mode
Metric lineageSource to formula to dashboard to executive decisionRevenue, cycle time, AI ROI, risk score, adoption rateMetric owner, data steward, analytics ownerCompeting definitions and dashboard drift
Customer lineageCustomer master to engagement to actionAccount status, contract, support signal, renewal riskSales / success owner, CRM owner, data ownerAI outreach based on stale or wrong context
Employee lineageHRIS to access, org design, capacity, and AI workforce planningRole, manager, team, skills, access, cost centerPeople owner, HRIS owner, identity ownerBad routing, access errors, shadow org structure
Work lineageWork intake to prioritization to delivery and outcomeInitiatives, dependencies, blockers, delivery health, benefitsProduct owner, PMO owner, work system ownerHidden work and false delivery status
Governance lineagePolicy/control to exception to evidence and audit trailRisk, control, policy, approval, incident, exceptionGRC owner, risk owner, control ownerUnreviewed exception or missing evidence
AI model / agent lineageTraining or retrieval data to prompt/model to output to human reviewModel features, RAG source, prompt, recommendation, automated actionAI owner, data owner, governance ownerAI output treated as fact without source trace
Knowledge lineageKnowledge object to published artifact to AI retrieval and useSOP, playbook, framework object, policy, prompt libraryContent owner, knowledge steward, AI retrieval ownerStale documents becoming enterprise memory

Template: data lineage map

Map elementEntryGuidance
Lineage ID[DL-001]Unique object ID
Business domain[Customer / Employee / Finance / Work / Risk / AI]Used for routing and governance
Business use case[Decision, metric, dashboard, process, AI use case, control]Defines why the lineage matters
Source system[System of record]Original or authoritative source
Source object[Table, API object, data product, document, event, file]Specific object being consumed
Source owner[Name / role]Accountable for meaning and approved use
Technical owner[Name / role]Accountable for system and integration reliability
Transformation steps[Logic, joins, formula, enrichment, model features]How data changes before use
Integration path[API / ETL / event / file / manual / agent]How data moves
Storage / index[Warehouse / lakehouse / BI model / vector index / feature store]Where data persists or is retrieved
Downstream consumers[Dashboards, teams, workflows, models, agents, vendors]Who or what depends on this path
Decision/action supported[Decision, automation, escalation, report, control]What the data influences
Quality checks[Freshness, accuracy, reconciliation, completeness, duplicates]Evidence that path is reliable
AI usage boundary[Retrieve / summarize / recommend / classify / act / blocked]Defines safe AI behavior
Known gaps[Missing owner, stale data, shadow copy, manual step]Exception list
Review date[Date]Next validation date

Version for 500+ employee company

DimensionRecommended pattern
Design intentCreate basic visibility before the company scales into duplicated tools, unowned dashboards, and AI pilots using weak data.
Minimum lineage scopeMap the top 10 to 25 critical data paths: customer, employee, finance, delivery, product, risk, and active AI initiatives.
Primary systemsCRM, HRIS, work system, finance system, knowledge base, BI tool, and AI pilot workspace.
Operating patternQuarterly lineage review with business owners, data steward, technology owner, and executive sponsor.
AI focusAI can only summarize or recommend from named sources with human review for customer, employee, financial, risk, or external-facing use.
Red flagsSpreadsheet exports, dashboard-only truth, founder-memory definitions, unowned AI prompts, and unclear data freshness.

Version for 5,000+ employee company

DimensionRecommended pattern
Design intentMove from local data knowledge to domain-owned lineage that supports cross-functional decisions and scaled AI use cases.
Minimum lineage scopeMap critical data products, enterprise metrics, system integrations, decision dashboards, governance controls, and AI model inputs.
Primary systemsCRM, HRIS, ERP, work management, service desk, data warehouse/lakehouse, BI semantic layer, GRC, IAM, AI platform.
Operating patternDomain lineage owners, data product reviews, data catalog entries, quality thresholds, and formal exception paths.
AI focusRAG sources, model features, agent workflows, and AI outputs must reference governed sources, data contracts, and review rules.
Red flagsMultiple dashboards for the same metric, undocumented transformations, BI logic outside catalog, and model/agent consumption without lineage.

Version for 10,000+ employee company

DimensionRecommended pattern
Design intentCreate enterprise-grade lineage as a control layer for decisions, risk, AI, audit, regulation, and operating-model resilience.
Minimum lineage scopeMap all Tier 1 and Tier 2 data domains, regulatory/control data paths, AI-critical datasets, executive metrics, and external reporting feeds.
Primary systemsEnterprise data catalog, MDM, data lakehouse, semantic layer, API gateway, event platform, GRC, IAM, model registry, agent orchestration, observability.
Operating patternAutomated lineage capture, lineage control board, domain data councils, evidence packs, policy-as-code checks, and supersession management.
AI focusAI models and agents require retrieval lineage, feature lineage, prompt/output logging, human review policy, model risk tier, and audit replay.
Red flagsFederated business units creating conflicting truth, unmanaged data sharing, vendor feeds without ownership, and AI agents acting on low-confidence data.

Scoring logic

DimensionScoreWhat good looks like
Ownership clarity0-5Every source, transformation, consumer, and decision path has a named owner.
Source clarity0-5The authoritative source and upstream origin are clearly defined.
Transformation transparency0-5Calculation, enrichment, joins, model features, and manual steps are documented and owned.
Consumer visibility0-5Dashboards, decisions, workflows, controls, AI models, and agents using the data are known.
Evidence strength0-5Lineage is supported by catalog links, pipeline logs, definitions, data contracts, and review history.
Freshness discipline0-5Refresh cadence and stale-data triggers are defined and monitored.
AI safety boundary0-5AI retrieval, recommendation, update, and action boundaries are explicit and reviewable.
Exception management0-5Manual steps, duplicates, shadow sources, broken feeds, and disputed lineage have clear escalation paths.

Suggested readiness score: average the eight scores, then classify 0-1.9 as Fragile, 2.0-3.4 as Developing, 3.5-4.4 as Governed, and 4.5-5.0 as AI-ready.

AI prompts

  • Given this data object, identify the source system, transformation steps, downstream consumers, decision dependencies, and missing owners.
  • Compare the declared source of truth against actual downstream dashboards, workflows, AI tools, and manual exports. Flag duplicate or conflicting lineage paths.
  • Score this lineage path from 0 to 5 across ownership clarity, source clarity, transformation transparency, consumer visibility, evidence strength, freshness discipline, AI safety, and exception management.
  • Identify whether this dataset is safe for AI retrieval, summary, recommendation, classification, or autonomous action. Explain the human review rule needed.
  • Generate a lineage exception report showing stale data, undocumented transformations, shadow copies, missing owners, weak evidence, and high-risk AI consumers.
  • Create a migration plan to move this lineage path from manual documentation to governed Lapemo ingestion.

Validation rules

  • Every lineage object must have one accountable source owner and one technical owner.
  • Every transformation must have a documented owner, purpose, and approval state.
  • Every downstream AI use must reference an approved source, freshness window, sensitivity rule, and human review boundary.
  • Every decision-grade metric must have a named source, formula, steward, quality threshold, and review date.
  • No dashboard, model, agent, workflow, or external report should consume data from an unowned or disputed lineage path.
  • Manual exports, spreadsheet changes, and shadow copies must be flagged as exceptions unless explicitly approved.
  • If data is stale, disputed, low-confidence, or missing evidence, it cannot be used as authoritative truth without owner approval.
  • Status must never be encoded only by color; use labels such as approved, provisional, stale, disputed, retired, or exception.

Lapemo ingestion mapping

Lapemo objectFields / entitiesUse
Knowledge objectData Lineage MapCanonical reusable artifact for lineage readiness and AI-safe data consumption.
Ownership objectSource owner, technical owner, steward, transformation ownerConnects lineage to accountability.
Information objectSource, transformation, storage, consumer, evidenceConnects lineage to enterprise knowledge and data products.
Platform objectSystems, integrations, APIs, pipelines, BI, model registry, vector indexConnects lineage to enterprise systems under control.
Decision objectDecision/action supported, evidence link, confidenceConnects data to decisions and outcomes.
Governance objectSensitivity, access, control, exception, review dateConnects lineage to risk, compliance, and auditability.
AI objectAI usage boundary, human review rule, output log, agent consumerConnects data lineage to AI governance and control.

Reusable knowledge-object model

This artifact should exist in four synchronized forms: a human-readable guide, a downloadable template, a machine-readable JSON object, and a guided Lapemo skill. The knowledge object should be versioned, reviewed, scored, and connected to system data over time. It should not auto-update silently. Lapemo should flag stale lineage, missing owners, broken evidence, new AI consumers, and high-risk exceptions for human approval.

Future Lapemo Use

The JSON schema turns data lineage map into software.

Lapemo can use this knowledge object as a guided workflow, scoring model, evidence record, governance input, and operating intelligence object. The schema is public for inspection and evaluation; production ingestion and governed execution remain separate product capabilities.

Version Metadata

Version metadata

Version

1.0.0

Last updated

2026-06-23

Review cadence

Quarterly or when sources change

Data Lineage Map

Make it part of the operating model.

Use this object as a working record now, then connect it to metrics, evidence, and Lapemo workflows as the operating system matures.