X. DATA
Nexus data governance for data commons, rights, metadata, lifecycle controls, secure rooms, compute-to-data, lawful handoff context, and non-executing data boundaries.
This section defines data in Nexus as governed public-good infrastructure.
It explains how data is classified, accessed, transformed, reviewed, published, restricted, corrected, archived, and handed off. It also sets the core boundary rule: data may support learning, evidence, intelligence, and coordination, but it does not become public authority action, procurement, finance, insurance, consent, deployment authorization, or execution by implication.
10.1 Data Commons Doctrine
10.1.1 Data, Innovation, Commons, and Evidence Governance Layer
Decentralized Innovation Commons (DICE) is the data, innovation, commons, and evidence governance layer of Nexus Ecosystem. It governs how data, metadata, datasets, data products, data rooms, clean rooms, secure rooms, compute-to-data workflows, public-safe summaries, AI-use inputs, Observatory signals, DRI indicators, GRIx mappings, Studio workflows, Reports, Academy objects, Campaign objects, Grid and TRL records, National Portfolio objects, Nexus Universe outputs, Marketplace listings, Registry records, and handoff packages are created, classified, accessed, used, reviewed, transformed, published, restricted, corrected, archived, or handed off.
DICE exists because data is not neutral raw material. Data may carry rights, omissions, power, bias, uncertainty, privacy risk, cybersecurity risk, public authority sensitivity, geospatial sensitivity, protected knowledge, community context, Indigenous protocol obligations where applicable, commercial restrictions, sovereign data controls, cross-border transfer limits, AI-use restrictions, and downstream reliance risk. Nexus therefore treats data as governed infrastructure, not as an extractive input.
DICE supports innovation by making data usable under disciplined conditions. It supports commons by making public-good data and metadata discoverable where safe. It supports evidence by requiring provenance, lineage, source context, quality, uncertainty, review, public-safe transformation, and correction. It supports lawful handoff by ensuring data conditions travel with the object and are not stripped away when an output becomes useful to downstream actors.
DICE may operate across open, controlled, restricted, secure-room-only, data-room-only, clean-room-only, compute-to-data-only, National Node-only, protected-knowledge-controlled, public-authority-learning-only, handoff-recipient-only, archive-only, sealed, or non-continuing data states. Its purpose is not to make all data open. Its purpose is to make data governance precise enough that openness, control, restriction, learning, intelligence, publication, and handoff can each occur without collapsing into one another.
DICE does not create data ownership, unrestricted access, AI-training rights, publication rights, public authority approval, procurement status, financeability, insurability, certification, consent, deployment authorization, or execution authority by default. It creates the governance layer that allows data to move responsibly where movement is lawful, safe, recorded, and bounded.
10.1.2 Data Commons Defined
A Data Commons is a governed public-good environment in which data objects, metadata objects, schemas, ontologies, data dictionaries, data products, data-quality records, lineage records, access rules, public-safe summaries, DICE records, GRIx mappings, DRI indicators, Observatory inputs, Studio-ready datasets, Reports-ready datasets, Academy datasets, National Portfolio datasets, Nexus Universe datasets, and handoff-relevant data packages may be organized for shared learning, reuse, review, and public-good production.
A Data Commons is not an unrestricted data lake. It is not a mass-upload repository, surveillance layer, commercial data marketplace, unbounded AI-training corpus, procurement platform, public authority database, consent system, or execution system. Its public-good value depends on governance: identity, metadata, rights, access class, sensitivity class, data-use label, AI-use label, jurisdictional context, review status, public-safe status, support class, correction pathway, and archive rule.
A Data Commons may include:
open data objects, where lawful and public-safe release is permitted;
controlled data objects, where access is limited to defined participants, reviewers, learners, rooms, institutions, or pathways;
restricted data objects, where privacy, cyber, public authority, protected knowledge, geospatial, legal, commercial, sovereign, or safeguard concerns require stronger controls;
metadata-first objects, where metadata can be made discoverable even where source data remains restricted;
public-safe data products, where raw or sensitive data has been transformed into safe summaries, aggregates, indicators, or learning objects;
compute-to-data objects, where permitted computation occurs without raw data export;
handoff-context data objects, where data conditions are transferred to competent recipients without transferring unrestricted rights.
A Data Commons should be commons-based in governance and public-good purpose, not commons-washed in language. Data enters the commons only within lawful, recorded, rights-respecting, consent-aware where applicable, privacy-preserving, secure, public-safe, and correctionable conditions. The commons is a disciplined sharing architecture, not a default claim that all data belongs to everyone.
10.1.3 Public-Safe Data Defined
Public-safe data is data or a data-derived object that has been reviewed, transformed, classified, and released in a manner suitable for public or broader controlled communication without creating unacceptable risk of harm, privacy exposure, protected knowledge exposure, public authority confusion, cybersecurity risk, geospatial risk, community harm, Indigenous protocol breach where applicable, public panic, false reassurance, finance or insurance overclaim, procurement implication, consent overclaim, or execution overclaim.
Public-safe data may include aggregated indicators, de-identified summaries, generalized geospatial layers, masked records, public-safe dashboards, public-safe DRI summaries, public-safe Observatory outputs, public-safe National Portfolio summaries, public-safe Reports tables, public-safe Academy datasets, public-safe Campaign metrics, public-safe Nexus Universe outputs, or public-safe Marketplace descriptions.
Public-safe data should identify:
source data class and source restrictions;
transformation method;
aggregation, masking, redaction, generalization, or de-identification steps;
public-safe reviewer or review pathway;
residual risk;
uncertainty and limitation language;
permitted audience;
prohibited use;
update and correction pathway;
archive treatment.
Public-safe data is not necessarily open data. Some public-safe outputs may be suitable for public communication; others may be suitable only for controlled participants, public authority learning rooms, media-safe rooms, Academy pathways, National Nodes, or handoff recipients. Public-safe status is also not permanent. New information, linkage risk, geospatial sensitivity, AI re-identification risk, legal change, community concern, Indigenous protocol concern where applicable, or public-safe reporting issue may require correction, restriction, withdrawal, recall, or archive.
Public-safe data informs responsibly. It does not become public warning, public authority action, certification, procurement status, financeability, insurability, consent, deployment authorization, or execution authority.
10.1.4 Controlled Data Defined
Controlled data is data that may be used within Nexus only under defined access, role, purpose, security, privacy, public-safe, data-use, AI-use, jurisdictional, safeguard, and correction conditions. Controlled data is not fully open, but it is not necessarily inaccessible. It exists between unrestricted public release and strict non-use.
Controlled data may include:
participant-limited datasets;
National Node datasets;
public authority learning datasets;
data-room materials;
secure-room materials;
clean-room materials;
compute-to-data datasets;
geospatial-sensitive datasets;
infrastructure-sensitive datasets;
cyber-sensitive datasets;
community-sensitive datasets;
health-sensitive or youth-sensitive datasets;
Indigenous protocol-sensitive datasets where applicable;
protected knowledge datasets;
provider-contributed or sponsor-supported restricted datasets;
handoff-recipient-only datasets.
A Controlled Data Record should identify the data steward, access class, approved users, approved purpose, prohibited purpose, retention rule, output review rule, download rule, export rule, cross-border transfer rule, AI-use rule, publication rule, public-safe transformation rule, correction pathway, incident pathway, and archive rule.
Controlled data should not be copied into open repositories, used for AI training, displayed in public dashboards, exported across borders, included in Reports, listed in Marketplace, used in Campaigns, included in Nexus Universe materials, or handed off unless the applicable controls permit that specific use. Access to controlled data is not unrestricted permission. Use is always purpose-bound and record-bound.
Controlled data allows Nexus to learn from sensitive information without pretending sensitivity has disappeared.
10.1.5 Metadata as Public-Good Bridge
Metadata is the public-good bridge between data that may be openly known, data that may be controlled, and data that must remain restricted. It allows the ecosystem to understand that data exists, what it is about, who stewards it, what rights and restrictions apply, what quality and uncertainty conditions exist, what review has occurred, what public-safe outputs may be available, and how the data may be requested, studied, transformed, or handed off without exposing the underlying data improperly.
Metadata may describe:
dataset identity and title;
steward and source pathway;
subject matter and scope;
geography and time period;
data class and sensitivity class;
rights and license conditions;
access and release class;
data-use label;
AI-use label;
quality, lineage, provenance, uncertainty, and limitations;
update cadence and support status;
public-safe summary availability;
DICE, GRIx, DRI, Observatory, Studio, Reports, Academy, Marketplace, Registry, Grid, National Portfolio, Nexus Universe, or handoff relationships;
correction and archive pathway.
Metadata can be more open than source data where safe, but metadata can also be sensitive. The fact that a dataset exists, its location, its steward, its subject, its geography, or its relationship to a public authority, community, Indigenous institution where applicable, critical infrastructure, health system, protected site, or cyber system may itself require control.
Metadata should never be treated as decorative. It is the bridge that lets Nexus build commons without forcing unsafe disclosure. It makes data discoverable, governable, and correctable while preserving lawful boundaries.
10.1.6 Data as Evidence-Bearing Object
Data as an evidence-bearing object means that data within Nexus must carry provenance, lineage, method context, quality context, uncertainty, limitations, sensitivity, review status, and correction pathway. Data is not evidence merely because it exists. It becomes evidence-bearing when its source, collection method, transformation history, scope, gaps, assumptions, rights, restrictions, and intended use are recorded.
A data object used as evidence should identify:
source and provenance;
collection or generation method;
lineage and transformation history;
data quality, missingness, bias, uncertainty, and confidence;
spatial, temporal, sectoral, and population scope;
rights, permissions, license, and consent conditions where applicable;
privacy, cyber, geospatial, protected knowledge, public authority, community, and Indigenous protocol sensitivity where applicable;
data-use and AI-use labels;
review status;
public-safe transformation status;
correction and archive rules.
Evidence-bearing data may support Reports, DRI indicators, GRIx mappings, Observatory outputs, Studio workflows, dashboards, digital twins, simulations, Academy materials, Campaign summaries, Grid and TRL records, National Portfolio objects, Nexus Universe outputs, Marketplace listings, Registry records, and handoff packages. Its evidentiary meaning remains bounded by purpose and scope.
Data evidence is not universal truth. A dataset may be sufficient for exploratory learning but insufficient for public-safe reporting. It may be sufficient for Studio demonstration but insufficient for handoff. It may be sufficient for handoff context but insufficient for deployment. Evidence-bearing status supports interpretation; it does not create certification, public authority action, procurement, finance, insurance, consent, deployment, or execution.
10.1.7 Data as Learning Object
Data as a learning object means that datasets, metadata, synthetic datasets, public-safe datasets, example datasets, data dictionaries, notebooks, dashboards, quality records, DICE records, GRIx mappings, DRI indicators, Observatory outputs, and Studio workflows may be used to build capability through Nexus Academy, Risk Academy, WILPs, Foundry contributor learning, Studio learning, public authority learning, data and AI literacy, public-safe reporting learning, finance-readiness literacy, insurance-readiness literacy, and handoff literacy.
A data learning object should identify:
learning purpose;
learner class;
data class and sensitivity;
whether data is real, synthetic, anonymized, aggregated, masked, public-safe, controlled, restricted, or compute-to-data-only;
permitted exercises;
prohibited uses;
AI-use restrictions;
public-safe output rules;
accessibility and localization status;
correction and archive pathway.
Learning use must not be treated as unrestricted data use. A learner may access a dataset for a controlled exercise without acquiring the right to publish it, train AI on it, download it, transfer it, combine it with other data, use it commercially, submit it to a public authority, include it in procurement materials, or hand it off. Where sensitive data is not needed for learning, synthetic or public-safe substitute data should be preferred.
Data learning objects build capability. They do not license professionals, certify competence externally, grant data rights, authorize AI training, create public authority action, create financeability, create insurance approval, grant consent, deploy systems, or execute projects.
10.1.8 Data as Intelligence Object
Data as an intelligence object means that data may be organized, analyzed, transformed, and interpreted to support disaster-risk intelligence, systems-risk awareness, Observatory signals, DRI indicators, GRIx mappings, degraded-mode awareness, WFEH-B analysis, digital twin needs, Studio scenarios, public authority learning, National Portfolio preparation, Nexus Universe outputs, Reports, Campaigns, Grid and TRL context, finance-readiness questions, insurance-readiness questions, and lawful handoff dependencies.
An intelligence data object should identify:
intelligence question;
source data and metadata;
method and transformation logic;
uncertainty, confidence, limitations, and gaps;
public-safe status;
sensitivity and access class;
interpretation limits;
public authority boundary language;
finance and insurance boundary language where relevant;
correction, update, and archive pathway.
Data intelligence must distinguish signal from decision. A risk signal is not a public warning. An indicator is not an official classification. A hotspot is not a public authority designation. A dashboard is not emergency command. A finance-relevant indicator is not investment advice. An insurance-relevant signal is not underwriting. A National Portfolio intelligence object is not government approval. A handoff-relevant data object is not execution authority.
Data as intelligence is powerful because it supports better questions and better preparedness. It remains legitimate only when interpretation limits, public-safe rules, and no-conversion boundaries travel with the object.
10.1.9 Data Without Unrestricted Rights by Default
Nexus operates under the doctrine of data without unrestricted rights by default. No data object is presumed to be freely accessible, reusable, publishable, transferable, downloadable, commercializable, AI-trainable, handoff-transferable, or deployable merely because it appears within Nexus, supports a public-good purpose, is referenced in metadata, is used in a dashboard, is summarized in a Report, is stored in a repository, is discussed in a room, is included in Nexus Universe, is listed in Marketplace, or is recorded in Registry.
Data rights must be specific, recorded, and attached to the data object. A data object should distinguish:
right to know metadata;
right to view;
right to query;
right to compute-to-data;
right to download;
right to transform;
right to publish;
right to combine;
right to train AI;
right to share with participants;
right to transfer across borders;
right to hand off to a recipient;
right to commercialize or use in enterprise context;
right to archive;
duty to delete, seal, restrict, or recall.
Public-good purpose does not override privacy, data protection, intellectual property, contract, consent, Indigenous protocol where applicable, protected knowledge, public authority sensitivity, cyber sensitivity, health sensitivity, youth protection, sovereign data rules, cross-border transfer limits, or other legal and ethical controls.
The default rule is controlled meaning: data can be used only within the rights, access, purpose, and pathway recorded for it. Anything beyond that requires separate permission or authority.
10.1.10 Data Without AI-Use Permission by Default
Nexus operates under the doctrine of data without AI-use permission by default. Data access does not imply permission to use the data for AI training, fine-tuning, embedding, retrieval augmentation, automated classification, generative output, agentic workflow, model evaluation, synthetic data generation, public-facing AI tool operation, or downstream AI product development.
AI-use permission must be separately recorded through an AI-Use Label or equivalent data-use instrument. Where AI use is permitted, the record should identify:
permitted AI-use class;
prohibited AI-use class;
model or system class where relevant;
whether training, fine-tuning, embedding, retrieval, summarization, classification, generation, simulation, or agentic use is allowed;
whether outputs may be published, taught, listed, reported, demonstrated, or handed off;
whether human review is required;
whether secure-room, data-room, clean-room, or compute-to-data conditions apply;
whether protected knowledge, personal data, health data, youth data, community data, Indigenous protocol-sensitive data where applicable, cyber-sensitive data, geospatial-sensitive data, or public authority-sensitive data is excluded;
correction, deletion, model-output recall, and archive rules.
The absence of an AI-use permission means no AI use beyond ordinary non-AI handling unless a competent steward records otherwise. Public availability does not mean AI-training permission. A dataset used for a report does not become a model-training corpus. A data-room review does not authorize extraction into AI tools. A community contribution does not authorize AI training. Indigenous participation where applicable does not authorize protected knowledge ingestion. Handoff context does not transfer AI-use rights unless expressly recorded.
The final DICE and Data Commons Doctrine rule is: data is governed evidence, learning, and intelligence infrastructure; commons means lawful, recorded, rights-respecting, public-safe sharing, not unrestricted extraction; metadata bridges openness and control; data rights and AI-use permissions must be explicit; no data object becomes open, publishable, AI-trainable, financeable, insurable, public-authority-approved, consented, deployable, or executable by implication.
10.2 Data Object Classes
10.2.1 Raw Data
Raw data is data in its original or near-original collected, received, captured, sensed, submitted, extracted, observed, logged, generated, or imported form before substantial cleaning, transformation, aggregation, normalization, public-safe conversion, modeling, or publication. Raw data may come from sensors, surveys, administrative records, public authority sources, research sources, community inputs, Indigenous knowledge contexts where applicable, geospatial sources, Earth observation, cyber telemetry, infrastructure logs, health systems, WFEH-B systems, provider systems, Studio workflows, Observatory nodes, DRI pipelines, Campaign inputs, National Portfolio processes, Nexus Universe rooms, or handoff-recipient materials.
Raw data carries the highest governance burden because it may include personal information, sensitive attributes, location data, protected knowledge, confidential institutional information, cyber-sensitive details, infrastructure-sensitive details, public authority-sensitive records, commercial restrictions, data-quality defects, missingness, bias, uncertainty, or context that can be lost when separated from its source. Raw data should therefore be treated as controlled or restricted by default unless a competent steward records that it is lawful, safe, and appropriate for open use.
A Raw Data Record should identify:
source and steward, including who provided, collected, generated, or controls the data;
collection or generation method, including sensor, survey, administrative process, observation, model output, system log, public record, research process, community process, or other source pathway;
rights and lawful basis, including license, permission, consent where applicable, contract, public authority basis, research basis, or restriction;
sensitivity and access class, including privacy, cyber, health, youth, geospatial, infrastructure, community, Indigenous protocol where applicable, protected knowledge, public authority, sovereign, legal, or security sensitivity;
data-use and AI-use labels, including whether viewing, querying, transformation, publication, AI training, embedding, summarization, classification, simulation, handoff, or cross-border transfer is permitted;
quality and limitation notes, including completeness, missingness, uncertainty, bias, timeliness, spatial resolution, temporal resolution, and known defects;
correction, restriction, deletion, sealing, recall, and archive rules.
Raw data should not be placed in open repositories, used in public dashboards, included in public Reports, used for AI training, displayed at Nexus Universe, listed in Marketplace, transferred to handoff recipients, or reused in Academy materials unless the relevant rights, reviews, transformations, public-safe controls, and access conditions allow that specific use. Raw data is evidence potential, not public-good release by default.
10.2.2 Processed Data
Processed data is data that has been cleaned, normalized, transformed, validated, enriched, joined, filtered, structured, geocoded, labeled, encoded, classified, deduplicated, quality-checked, or otherwise modified from raw form for a defined Nexus purpose. Processed data may support DICE objects, DRI indicators, GRIx mappings, Observatory outputs, Studio workflows, dashboards, digital twins, simulations, Reports, Academy modules, Campaign summaries, Grid or TRL records, Marketplace listings, Registry records, National Portfolio objects, Nexus Universe outputs, or handoff packages.
Processed data should preserve lineage. Transformation must not erase the source, rights, sensitivity, uncertainty, missingness, bias, or limitations of the raw data. Processing may improve usability, but it may also introduce assumptions, errors, exclusions, or interpretation changes. Nexus therefore treats processed data as a new governed data object connected to its source objects and transformation method.
A Processed Data Record should identify:
source data objects and versions;
processing method, including cleaning, normalization, transformation, linkage, aggregation, enrichment, geocoding, labeling, classification, or feature generation;
processing environment, including open, controlled, secure-room, data-room, clean-room, compute-to-data, National Node, Studio, or handoff-recipient environment;
quality changes, including errors corrected, records excluded, fields changed, assumptions applied, uncertainty introduced, and known limitations;
rights and restrictions carried forward, including license, consent limits, data-use limits, AI-use limits, privacy restrictions, public authority restrictions, protected knowledge restrictions, and jurisdictional conditions;
review status, including data review, method review, public-safe review, AI review where applicable, and safeguard review where relevant;
correction and archive relationship, including source correction propagation and downstream object correction.
Processed data is not automatically safer than raw data. De-identification may be incomplete. Aggregation may still reveal sensitive locations. Normalization may hide bias. Labeling may encode contested assumptions. Processing improves structure; it does not create unrestricted rights, public-safe status, certification, public authority approval, consent, deployment authorization, or execution authority.
10.2.3 Public-Safe Data
Public-safe data is data or data-derived material that has passed a public-safe transformation and review pathway for a defined audience and release class. Public-safe data may be produced through aggregation, masking, redaction, de-identification, synthetic substitution, geospatial generalization, uncertainty labeling, limitation language, sensitive-field removal, protected knowledge exclusion, public authority boundary review, and no-conversion language.
Public-safe data may include public-facing tables, maps, indicators, dashboards, summaries, charts, teaching datasets, public-safe DRI extracts, public-safe Observatory summaries, public-safe National Portfolio summaries, Campaign metrics, Nexus Universe public materials, Marketplace descriptions, Reports figures, and handoff-facing summaries. Its purpose is to communicate data-derived knowledge without exposing sensitive source data or creating unsafe reliance.
A Public-Safe Data Record should identify:
source data and source access class;
public-safe transformation method;
removed, masked, generalized, or restricted fields;
audience and release class;
residual risk, including re-identification, geospatial exposure, community harm, public authority confusion, finance or insurance overclaim, and consent overclaim;
required limitation and uncertainty language;
review pathway, including public-safe, data, privacy, cyber, safeguard, community, Indigenous protocol where applicable, and public authority boundary review where relevant;
correction, withdrawal, recall, and archive pathway.
Public-safe data is not automatically open data. It may be public-safe only for a specific summary, chart, room, classroom, report, dashboard, or audience. Public-safe status may change if new linkage risks arise, source data is corrected, public authority context changes, community concerns arise, Indigenous protocol concerns arise where applicable, or downstream misuse appears.
Public-safe data informs responsibly. It does not become public warning, official determination, procurement recommendation, finance signal, insurance rating, consent record, deployment authorization, or execution instruction.
10.2.4 Aggregated Data
Aggregated data is data combined or summarized across records, persons, places, institutions, systems, time periods, sectors, hazards, indicators, or categories so that individual-level or source-level detail is reduced. Aggregation may support public-safe reporting, dashboards, DRI indicators, Observatory outputs, National Portfolio summaries, Reports, Academy materials, Campaign metrics, Grid context, Nexus Universe outputs, and handoff context.
Aggregation may reduce sensitivity but does not eliminate risk by default. Aggregated data can still expose small populations, vulnerable groups, sensitive locations, Indigenous lands or knowledge contexts where applicable, health patterns, infrastructure vulnerabilities, cyber posture, public authority capacity, commercial information, or community-sensitive conditions. Aggregation must therefore be governed by threshold rules, suppression rules, geospatial generalization, uncertainty labeling, and public-safe review.
An Aggregated Data Record should identify:
source data objects;
aggregation method, including grouping variables, thresholds, time period, geography, categories, weighting, and suppression rules;
minimum cell-size or disclosure-control rules where relevant;
geospatial treatment, including masking, displacement, generalization, or restricted-resolution display;
bias and missingness implications;
public-safe status and audience;
data-use and AI-use restrictions;
correction propagation from source data and to downstream objects.
Aggregation is not anonymization by default. It is not consent. It is not permission to publish. It is not permission to train AI. Aggregated data must still carry its rights, limitations, access class, public-safe status, and correction pathway.
10.2.5 Synthetic Data
Synthetic data is artificially generated data designed to resemble, simulate, teach, test, benchmark, demonstrate, or approximate real-world data patterns without directly exposing raw source records. Synthetic data may support Academy learning, software testing, dashboard demonstrations, Studio workflows, model evaluation, API examples, DICE exercises, GRIx mapping, DRI training, public-safe reporting demonstrations, Nexus Universe demonstrations, and handoff preparation.
Synthetic data can be useful because it may reduce exposure of sensitive source data. It does not eliminate risk by default. Synthetic data may still leak patterns from source data, reproduce bias, create false realism, mislead users, imply representativeness, or be misused as real evidence. It must be labeled clearly and governed as a data object.
A Synthetic Data Record should identify:
generation method, including rule-based generation, simulation, model-generated data, anonymized derivative generation, or manually constructed examples;
source relationship, including whether real data informed the structure and what restrictions carry forward;
intended use, including testing, teaching, demonstration, benchmarking, public-safe illustration, or handoff training;
prohibited use, including real-world inference, public authority decision, finance decision, insurance decision, procurement decision, operational deployment, or evidence claim beyond scope;
bias, realism, and limitation notes;
AI-use permissions and restrictions;
review, correction, and archive pathway.
Synthetic data is not automatically public-safe, unrestricted, or evidence-bearing. It may be a learning object, test object, or demonstration object. It must not be represented as real observation unless explicitly and truthfully described as such. Synthetic data helps protect source data; it does not create authority.
10.2.6 Metadata-Only Records
Metadata-only records are records that describe data objects without exposing the underlying data. They may identify title, steward, source pathway, subject matter, geography, time period, data class, access class, sensitivity class, data-use label, AI-use label, quality summary, public-safe summary availability, update cadence, review status, Registry status, Marketplace discoverability, National Portfolio relationship, Nexus Universe relationship, and handoff relevance.
Metadata-only records are essential where source data cannot be opened but its existence, governance, and possible public-good relevance should be known. They allow discovery, review routing, data access requests, public-safe reporting planning, Studio planning, Academy planning, Foundry planning, Grid routing, Registry status truth, Marketplace controlled discovery, and handoff dependency mapping without exposing restricted data.
A Metadata-Only Record should identify:
what may be known publicly or by controlled users;
what source data remains restricted;
who controls access;
what access request pathway applies;
what public-safe derivative exists, if any;
what AI-use restrictions apply;
what jurisdictional, privacy, cyber, community, Indigenous protocol where applicable, protected knowledge, or public authority restrictions apply;
correction and archive pathway.
Metadata itself may be sensitive. The existence of a dataset, its location, steward, geography, public authority relationship, infrastructure relationship, or community relationship may require control. Metadata-only does not mean harmless. It means source data is not exposed.
10.2.7 Data Dictionaries
Data dictionaries are structured metadata objects that define fields, variables, columns, codes, units, formats, permissible values, missing-value meanings, sensitivity labels, data types, quality notes, lineage notes, and interpretation rules for a data object or data object family.
Data dictionaries make data usable, reviewable, interoperable, teachable, and correctable. They support DICE, GRIx, DRI, Observatory, Studio, Reports, Academy, Marketplace, Registry, Grid, National Portfolios, Nexus Universe, and handoff packages by explaining what data fields mean and how they may be safely interpreted.
A Data Dictionary Record should identify:
dataset or object family covered;
field names and definitions;
data types, units, formats, and allowed values;
missingness codes and quality notes;
sensitivity labels by field;
data-use and AI-use restrictions by field where relevant;
controlled vocabulary or ontology links;
localization or translation status;
review status and steward;
correction and archive pathway.
A data dictionary does not make the underlying data open, accurate, complete, public-safe, AI-trainable, or authorized for handoff. It is an interpretive object. If the dictionary is wrong, every downstream use may be wrong. Data dictionaries therefore require versioning and correction discipline.
10.2.8 Codebooks
Codebooks are data interpretation objects that define codes, categories, labels, survey responses, qualitative codes, classification rules, ontology mappings, risk categories, DRI categories, GRIx categories, public authority learning categories, safeguard categories, consent boundary categories, or other coded values used in data objects.
Codebooks may support survey data, administrative data, qualitative research, community inputs, public authority learning records, DRI datasets, GRIx mappings, Observatory datasets, Campaign data, Academy assessments, National Portfolio objects, Nexus Universe room records, and handoff packages.
A Codebook Record should identify:
coded dataset or object;
code values and meanings;
classification method;
coder or steward;
inter-coder or review process where applicable;
controlled vocabulary or ontology relationship;
cultural, linguistic, legal, community, or Indigenous protocol context where applicable;
sensitivity and public-safe implications;
bias, ambiguity, and limitation notes;
correction, recoding, supersession, and archive pathway.
Codebooks are powerful because they shape what data appears to mean. They can also encode bias, erase local context, flatten protected knowledge, or create false categories. A codebook must not be treated as neutral simply because it is structured.
A codebook does not create official classification, public authority determination, certification, finance or insurance rating, procurement category, consent status, deployment authorization, or execution authority. It records meaning within scope.
10.2.9 Schemas
Schemas are structural data objects that define how data, metadata, records, APIs, Registry entries, Marketplace listings, Studio workflows, Grid records, TRL records, National Portfolio objects, Nexus Universe outputs, proof receipts, ledgers, registers, and handoff packages must be represented.
Schemas may specify fields, data types, required values, validation rules, controlled vocabularies, ontology relationships, sensitivity labels, access labels, public-safe labels, data-use labels, AI-use labels, version fields, correction fields, archive fields, localization fields, and interoperability rules.
A Schema Record should identify:
schema purpose and object class;
version and steward;
required and optional fields;
validation rules;
controlled vocabulary and ontology dependencies;
sensitivity and access fields;
public-safe and no-conversion fields;
data-use and AI-use fields;
interoperability and API relationships;
correction, migration, deprecation, and archive pathway.
A schema structures data; it does not validate the truth of the data entered. Schema conformance is not evidence sufficiency, public-safe approval, certification, procurement status, financeability, insurability, public authority action, consent, deployment authorization, or execution authority.
Schema discipline makes object movement possible. Review discipline makes object meaning trustworthy.
10.2.10 Data Pipelines
Data pipelines are governed workflow objects that move, transform, validate, enrich, aggregate, anonymize, classify, label, route, publish, restrict, or hand off data across systems, rooms, repositories, dashboards, Studio workflows, Observatory nodes, DRI workflows, GRIx mappings, Reports, Academy pathways, Marketplace listings, Registry records, Grid records, National Portfolios, Nexus Universe outputs, or lawful recipients.
A data pipeline may include ingestion, validation, transformation, metadata generation, quality checks, DICE controls, AI-assisted classification, feature generation, aggregation, public-safe transformation, secure-room processing, clean-room computation, compute-to-data execution, output review, Registry update, Marketplace update, or handoff packaging.
A Data Pipeline Record should identify:
pipeline purpose and source systems;
input data objects and access classes;
processing steps and transformation methods;
runtime environment, including open, controlled, secure room, data room, clean room, compute-to-data, Studio, National Node, or handoff-recipient environment;
permissions and controls, including authentication, authorization, logging, no-download, no-write-back, output review, and cross-border transfer controls;
AI-use, including whether AI assists classification, transformation, summarization, generation, or review;
outputs and release classes;
quality checks, failure handling, incident pathway, correction pathway, rollback, archive, and non-continuation rules.
Data pipelines can create hidden authority if they automatically publish, update dashboards, trigger reports, route status, or package handoff materials. Pipeline outputs should require review gates where public-safe, public authority, finance, insurance, procurement, consent, or handoff implications arise. Automation is not authority.
10.2.11 Feature Sets
Feature sets are structured variables or derived attributes prepared for modeling, analysis, dashboards, DRI indicators, AI workflows, digital twins, simulations, risk scoring, Observatory outputs, Grid reviews, or handoff-context evaluation. Feature sets may be generated from raw data, processed data, aggregated data, synthetic data, geospatial data, sensor data, metadata, or qualitative coding.
Feature sets are high-risk because they transform data into model-ready or decision-adjacent form. A feature may encode assumptions, proxies, bias, sensitivity, protected attributes, community context, Indigenous protocol-sensitive knowledge where applicable, or public authority implications. Feature engineering must therefore be governed as a substantive act, not a technical detail.
A Feature Set Record should identify:
source data objects;
features included and excluded;
feature definitions and calculation methods;
sensitivity, proxy, bias, and fairness considerations;
data-use and AI-use permissions;
model, dashboard, DRI, Studio, Grid, or handoff relationship;
quality, missingness, uncertainty, and limitation notes;
review status, including data, AI, public-safe, safeguard, and public authority boundary review where relevant;
correction, deprecation, withdrawal, recall, and archive pathway.
A feature set is not a model certification. It is not a risk rating, public authority classification, finance score, insurance score, procurement priority, consent record, deployment authorization, or execution instruction. It is a governed input that must remain interpretable and correctable.
10.2.12 Benchmark Datasets
Benchmark datasets are data objects used to evaluate, compare, test, calibrate, demonstrate, or review software, models, AI workflows, dashboards, DRI indicators, GRIx mappings, data pipelines, Studio workflows, digital twins, simulations, or public-good technical baselines.
Benchmark datasets may be open, controlled, synthetic, restricted, secure-room-only, data-room-only, clean-room-only, National Node-specific, public-safe, or handoff-recipient-only. Their value depends on clear scope and limitation. A benchmark that is narrow, biased, outdated, synthetic, localized, incomplete, or context-specific should not be represented as universal.
A Benchmark Dataset Record should identify:
benchmark purpose;
source and construction method;
coverage and exclusions;
ground truth or reference standard where applicable;
metrics supported and metrics not supported;
bias, missingness, uncertainty, and limitation notes;
data-use and AI-use permissions, including whether model training, fine-tuning, evaluation, publication, or redistribution is permitted;
access class and sensitivity class;
review status;
correction, versioning, supersession, withdrawal, recall, and archive rules.
Benchmark performance is not certification. A model performing well on a benchmark is not safe for deployment. A dashboard passing benchmark tests is not public authority-ready. A software tool matching benchmark expectations is not procurement-ready. Benchmarks are evidence inputs, not authority.
10.2.13 DRI Datasets
DRI datasets are data objects used to support disaster-risk intelligence, systemic-risk awareness, hazard, exposure, vulnerability, capacity, resilience, degraded-mode, cascade, WFEH-B, infrastructure, cyber-physical, climate, nature, protection-gap, public authority learning, finance-readiness, insurance-readiness, National Portfolio, Nexus Universe, Studio, Reports, Grid, or handoff-context workflows.
DRI datasets may include hazard data, exposure data, vulnerability data, resilience data, infrastructure data, WFEH-B data, geospatial data, sensor data, Earth observation data, public service data, logistics data, climate and weather data, biodiversity data, cyber-physical telemetry, public authority learning inputs, public-safe community context, and finance or insurance-relevant dependency data.
A DRI Dataset Record should identify:
risk-intelligence purpose;
hazard, exposure, vulnerability, capacity, resilience, or dependency category;
source, provenance, and lineage;
geography and time period;
quality, uncertainty, confidence, missingness, and bias;
public-safe status;
data-use and AI-use labels;
geospatial, infrastructure, privacy, cyber, community, Indigenous protocol where applicable, protected knowledge, public authority, finance, or insurance sensitivity;
DRI indicator relationship and GRIx mapping relationship;
Studio, Observatory, Reports, Grid, National Portfolio, Nexus Universe, and handoff relationships;
correction and archive pathway.
A DRI dataset is not a public warning. It is not official hazard designation, public authority classification, insurance rating, investment signal, procurement priority, consent record, deployment authorization, or execution instruction. It supports risk intelligence within recorded limits.
10.2.14 Observatory Datasets
Observatory datasets are data objects used by Nexus Observatory to collect, organize, interpret, route, summarize, or display signals related to systems risk, WFEH-B conditions, climate and nature stress, infrastructure resilience, digital infrastructure, cyber-physical systems, sensor and edge signals, Earth observation, geospatial layers, DRI indicators, GRIx mappings, degraded-mode awareness, digital twin needs, National Portfolios, Nexus Universe preparation, Studio workflows, Reports, Campaigns, Grid, and handoff context.
Observatory datasets may come from sensors, remote sensing, field observations, public datasets, partner datasets, National Node datasets, public authority learning sources, community inputs where appropriate, providers, research bodies, secure rooms, data rooms, clean rooms, compute-to-data workflows, or derived public-safe outputs.
An Observatory Dataset Record should identify:
signal purpose and signal class;
node, hub, regional cluster, or National Dense Nexus Core relationship;
source and steward;
collection method and update cadence;
geography, resolution, and temporal coverage;
sensitivity, including geospatial, infrastructure, cyber, privacy, community, Indigenous protocol where applicable, protected knowledge, ecological, public authority, or sovereign sensitivity;
public-safe transformation rules;
data-use and AI-use labels;
Studio, DRI, GRIx, Reports, Grid, National Portfolio, Nexus Universe, and handoff relationships;
correction, withdrawal, recall, and archive pathway.
Observatory datasets must not turn observation into surveillance. Sensor visibility, map layers, or signal records do not create public warning authority, emergency command, public authority decision, procurement priority, financeability, insurability, consent, deployment authorization, or execution authority.
10.2.15 National Portfolio Datasets
National Portfolio datasets are data objects that support country-level Nexus public-good memory, national systems-risk understanding, National Challenge Briefs, Evidence Need Records, DICE records, GRIx localization, DRI localization, Observatory needs, Studio workflow candidates, Academy pathways, Campaign candidates, Foundry builds, Grid and TRL context, public authority learning, safeguard records, finance-readiness questions, insurance-readiness questions, donor-readiness questions, public finance learning, Nexus Universe preparation, lawful handoff dependencies, correction, and archive.
National Portfolio datasets may be open, controlled, restricted, National Node-only, public-authority-learning-only, secure-room-only, data-room-only, clean-room-only, protected-knowledge-controlled, public-safe-summary-only, handoff-recipient-only, archive-only, or non-continuing. They should be governed according to national ownership, national data sovereignty, public authority boundaries, language and localization needs, community safeguards, Indigenous protocols where applicable, and national lawful handoff pathways.
A National Portfolio Dataset Record should identify:
country and National Node relationship;
National Nexus Consortium, National Council, Working Group, or Competence Cell relationship where applicable;
national purpose, including risk mapping, evidence need, public authority learning, Academy, Foundry, Campaign, Studio, Nexus Universe, Grid, or handoff context;
source, steward, provenance, and jurisdictional context;
data sovereignty, privacy, cyber, public authority, community, Indigenous protocol where applicable, protected knowledge, and cross-border transfer conditions;
public-safe status and localization status;
data-use and AI-use labels;
finance-readiness, insurance-readiness, donor-readiness, or public finance learning relationship where applicable;
correction, recall, archive, and national continuation pathway.
A National Portfolio dataset is not government approval, public authority action, procurement status, financeability, insurability, consent, deployment authorization, or execution authority. It is country-level public-good memory and preparation infrastructure. It helps a nation see, learn, prepare, and route context without bypassing competent national actors.
The final Data Object Classes rule is: raw data requires the strongest caution; processed data requires lineage; public-safe data requires transformation and review; aggregated data still carries risk; synthetic data must not pretend to be real; metadata-only records enable safe discovery; dictionaries, codebooks, and schemas preserve meaning; pipelines move data under controls; feature sets govern model-ready inputs; benchmarks support evaluation without certification; DRI and Observatory datasets support intelligence without warning authority; National Portfolio datasets preserve country-level memory without national approval or execution by implication.
10.3 Data Lifecycle
10.3.1 Data Signal
A data signal is the earliest indication that a data object, data need, data gap, data source, data risk, metadata need, public-safe data opportunity, DICE object, DRI dataset, Observatory dataset, National Portfolio dataset, Studio data workflow, Reports data dependency, Academy data object, Campaign data object, Grid or TRL data dependency, Nexus Universe data object, or handoff-related data package may require Nexus attention.
A data signal may arise from Observatory nodes, DRI workflows, GRIx mappings, public authority learning rooms, National Nodes, National Portfolios, Working Groups, Competence Cells, Nexus Labs, Nexus Foundry, Nexus Studio, Nexus Academy, Risk Academy, Nexus Campaigns, Nexus Reports, Nexus Marketplace, Nexus Registry, Nexus Grid, Nexus Universe, community safeguard rooms, protected knowledge rooms, secure rooms, data rooms, clean rooms, providers, sponsors, universities, public authorities acting separately, capital readers, insurers, donors, or lawful handoff recipients.
A data signal should identify, at minimum:
the apparent data subject or need, including risk, system, technology, geography, community, institution, workflow, model, report, dashboard, or handoff dependency;
the apparent source pathway, including whether the signal comes from observation, public source, private source, public authority learning context, community context, provider contribution, sponsor-supported pathway, research pathway, Studio workflow, National Portfolio pathway, Nexus Universe cycle, or handoff discussion;
the possible data class, including raw, processed, public-safe, aggregated, synthetic, metadata-only, DRI, Observatory, National Portfolio, benchmark, feature set, data dictionary, codebook, schema, or data pipeline object;
the immediate risk indicators, including privacy, cyber, geospatial, infrastructure, public authority, health, youth, community, Indigenous protocol where applicable, protected knowledge, sovereign, legal, finance, insurance, procurement, consent, public-safe, or AI-use concerns;
the next routing need, including intake, restriction, secure-room handling, data-room handling, public-safe review, rights review, DICE routing, National Node routing, or correction routing.
A data signal is not data approval. It is not permission to collect, access, process, publish, train AI, list, register, transfer, hand off, deploy, or execute. It is a trigger for disciplined intake and classification.
10.3.2 Data Intake
Data intake is the lifecycle state in which a data signal, dataset, metadata record, data product, data dictionary, codebook, schema, data pipeline, feature set, benchmark dataset, DRI dataset, Observatory dataset, National Portfolio dataset, or handoff data object enters a governed Nexus pathway for preliminary handling.
Data intake should occur before data is placed into repositories, used in software, displayed in dashboards, processed through AI tools, routed to Studio, converted into Reports, used in Academy materials, listed in Marketplace, recorded in Registry, classified through Grid or TRL, presented at Nexus Universe, included in National Portfolios, or transferred through handoff packages.
A Data Intake Record should identify:
data object identity, including provisional identifier, title, source, steward, and submitting pathway;
intake pathway, including DICE, Observatory, DRI, GRIx, Studio, Labs, Foundry, Academy, Reports, Campaigns, Grid, Marketplace, Registry, National Node, Nexus Universe, public authority learning room, secure room, data room, clean room, or handoff pathway;
data type and format, including file type, database, API, sensor feed, geospatial layer, model output, notebook output, dashboard input, report table, or metadata-only record;
initial access class, including open, controlled, restricted, secure-room-only, data-room-only, clean-room-only, compute-to-data-only, National Node-only, protected-knowledge-controlled, handoff-recipient-only, archive-only, sealed, or non-public;
known rights and restrictions, including license, contract, consent, public authority basis, provider terms, sponsor terms, community terms, Indigenous protocol where applicable, protected knowledge conditions, privacy, cyber, data sovereignty, cross-border transfer, and AI-use limits;
immediate handling instructions, including quarantine, no-download, no-AI-use, no-publication, no-sharing, no-cross-border-transfer, no-linkage, no-dashboard-display, no-Marketplace-listing, no-handoff, or secure-room routing.
Data intake does not authorize use. It creates a controlled record so that use can be assessed. Data that cannot be safely classified at intake should default to restriction until source, rights, sensitivity, data-use, AI-use, and public-safe conditions are reviewed.
10.3.3 Source Review
Source review examines where data came from, how it was collected or generated, who provided it, what authority or permission applies, what context may be missing, what biases may exist, and what source-related restrictions or risks must travel with the data.
Source review should determine whether the source is:
publicly available, and if so, under what license, terms, scraping limits, attribution duties, reuse limits, AI-use limits, and public-safe conditions;
institutionally provided, and if so, under what agreement, role, confidentiality condition, data-sharing condition, or support record;
public authority-related, and if so, whether the source relates to public authority learning, official records, public law restrictions, public finance, emergency authority, regulatory procedure, procurement, or public communications;
community-sourced, and if so, whether participation, consent boundaries, non-extraction controls, dignity safeguards, and public-safe limits are recorded;
Indigenous protocol-sensitive where applicable, and if so, whether governance, custodianship, rights, protected knowledge, data sovereignty, consent, and access restrictions are properly recorded;
provider-contributed, and if so, whether provider contribution is recorded without provider validation;
sponsor-supported, and if so, whether sponsor support is recorded without sponsor control;
machine-generated or model-generated, and if so, whether model source, synthetic status, uncertainty, limitations, and AI-use restrictions are recorded;
sensor, edge, geospatial, Earth observation, cyber, or infrastructure-derived, and if so, whether technical provenance, collection conditions, resolution, sensitivity, and public-safe limits are recorded.
A Source Review Record should state what is known, what is unknown, what must be verified, what restrictions apply, and whether the data may proceed to rights review, sensitivity review, DICE routing, secure-room handling, public-safe transformation, or archive.
Source review is not source endorsement. It establishes provenance and risk context. It does not create data rights, AI-use permission, public authority approval, consent, publication rights, deployment authorization, or execution authority.
10.3.4 Rights Review
Rights review determines what legal, contractual, institutional, community, Indigenous protocol where applicable, privacy, public authority, license, intellectual property, database, data protection, confidentiality, export, cross-border transfer, and use rights apply to a data object.
Rights review should distinguish separate rights rather than treating data access as a single permission. A Rights Review Record should determine whether the data may be:
viewed;
stored;
indexed;
queried;
cleaned or transformed;
linked with other data;
aggregated;
de-identified or masked;
used in dashboards;
used in Reports;
used in Academy or Campaign materials;
used in Studio;
used in DRI, GRIx, Observatory, Grid, Marketplace, or Registry workflows;
used for AI retrieval, classification, summarization, embedding, training, fine-tuning, simulation, or agentic workflows;
transferred across borders;
shared with participants;
published;
licensed onward;
handed off to a recipient;
archived, sealed, deleted, or recalled.
Rights review should identify the source of each right, any conditions, any prohibited uses, any expiration, any withdrawal rights, any attribution duties, any audit duties, any consent requirements, and any data subject, community, Indigenous, institutional, or public authority rights that must be respected.
Rights review does not convert controlled data into open data. It prevents misuse by making permissions explicit. Where rights are unclear, the most restrictive reasonable handling should apply until clarification is recorded.
10.3.5 Sensitivity Review
Sensitivity review classifies the risks that may arise from data access, use, disclosure, linkage, aggregation, publication, AI use, geospatial display, public-safe reporting, Marketplace discovery, Registry display, Studio workflow, National Portfolio inclusion, Nexus Universe presentation, handoff, correction, or archive.
Sensitivity review should identify whether the data is or may be:
personal-data-sensitive;
health-sensitive;
youth-sensitive;
vulnerable-population-sensitive;
community-sensitive;
Indigenous protocol-sensitive where applicable;
protected-knowledge-sensitive;
cultural-heritage-sensitive;
ecological-sensitive;
geospatial-sensitive;
infrastructure-sensitive;
cyber-sensitive;
public authority-sensitive;
sovereign-sensitive;
finance-sensitive;
insurance-sensitive;
procurement-sensitive;
legal-sensitive;
security-sensitive;
dual-use-sensitive;
export-control-sensitive where applicable;
archive-sensitive;
sealed or non-public.
Sensitivity review should consider both direct exposure and inference risk. Data may appear harmless alone but become sensitive when linked with other datasets, mapped at high resolution, processed by AI, included in dashboards, used in public reporting, routed to capital or insurance rooms, included in handoff materials, or displayed at Nexus Universe.
A Sensitivity Review Record should identify sensitivity class, access class, public-safe restrictions, output review requirements, secure-room or data-room needs, AI-use restrictions, cross-border transfer restrictions, correction triggers, incident pathway, and archive rule.
Sensitivity classification does not prohibit all use. It determines lawful and safe conditions of use. Where ambiguity exists, the more restrictive control should govern until review supports a narrower classification.
10.3.6 Data-Use Labeling
Data-use labeling assigns a recorded label to the data object describing what uses are permitted, restricted, prohibited, conditional, review-required, room-required, public-safe-transformation-required, handoff-limited, archive-only, or non-continuing.
Data-use labels should be specific enough to prevent broad inference. A data-use label may state that data is:
open for public use;
public-safe summary only;
controlled participant access;
internal Nexus use only;
National Node-only;
public authority learning only;
secure-room-only;
data-room-only;
clean-room-only;
compute-to-data-only;
dashboard-display permitted only in public-safe aggregate;
Reports use permitted only after public-safe review;
Academy use permitted only with synthetic or masked version;
Marketplace metadata-only;
Registry metadata-only;
handoff-recipient-only;
no download;
no cross-border transfer;
no publication;
no linkage;
archive-only;
deletion or sealing required;
non-continuing.
A Data-Use Label should travel with the data and all derived objects unless a lawful and recorded transformation changes the label. The label should appear in metadata, repositories, DICE records, Studio records, Reports records, Marketplace records, Registry records, Grid records, National Portfolio records, Nexus Universe records, and handoff packages where relevant.
Data-use labeling does not create permission beyond its terms. It constrains interpretation and prevents data from being treated as generally available merely because it has entered the ecosystem.
10.3.7 AI-Use Labeling
AI-use labeling assigns a recorded label to the data object describing whether and how data may be used with AI systems. It is separate from general data-use labeling because data that may be viewed or analyzed by humans may still be prohibited from AI training, embedding, retrieval augmentation, automated classification, summarization, simulation, generation, model evaluation, agentic workflow, or downstream AI product development.
AI-use labels may include:
no AI use permitted;
AI-assisted metadata generation permitted;
AI-assisted summarization permitted after review;
AI-assisted classification permitted under controlled conditions;
AI-assisted translation permitted with human review;
retrieval use permitted without retention;
embedding permitted in controlled environment;
model evaluation permitted;
simulation use permitted;
synthetic data generation permitted;
training prohibited;
fine-tuning prohibited;
agentic use prohibited;
public AI tools prohibited;
secure-room-only AI use;
data-room-only AI use;
clean-room-only AI use;
compute-to-data-only AI use;
protected knowledge excluded from AI use;
personal data excluded from AI use;
public-safe outputs only;
AI-use under suspension or correction.
An AI-Use Label should identify permitted model classes, prohibited model classes, approved environments, human review requirements, output review requirements, retention rules, logging rules, prompt-injection controls, data leakage controls, protected knowledge controls, deletion or recall obligations, and downstream restriction propagation.
No AI-use permission should be inferred from data access. The default is no AI use unless an AI-use label permits the specific use. This rule protects data subjects, communities, Indigenous institutions where applicable, protected knowledge, public authorities, data stewards, and Nexus from unauthorized AI extraction.
10.3.8 Lineage Capture
Lineage capture records the history of a data object from source through processing, transformation, review, public-safe conversion, repository routing, DICE routing, Registry status, Marketplace listing, Studio workflow, Report, Academy object, Campaign object, Grid input, National Portfolio object, Nexus Universe output, handoff package, correction, and archive.
Lineage capture should identify:
source data object or source pathway;
source version;
intake record;
rights review record;
sensitivity review record;
data-use label;
AI-use label;
processing steps;
transformation methods;
aggregation, masking, de-identification, geospatial generalization, or synthetic generation steps;
tools, software, models, notebooks, APIs, pipelines, or AI systems used;
reviewers and review dates;
derived objects;
downstream objects and dependencies;
corrections, withdrawals, recalls, supersessions, and archive status.
Lineage is essential because data meaning changes as data moves. A dashboard may depend on processed data, which depends on raw data, which depends on a source license, which depends on consent, which may later be withdrawn or restricted. Without lineage, correction cannot propagate and public-safe status cannot be trusted.
Lineage capture does not validate the data. It records the path by which data became an object. That path supports evidence, correction, and accountability without creating authority by implication.
10.3.9 Quality Assessment
Quality assessment evaluates whether a data object is fit for a defined Nexus purpose. Data quality is purpose-relative. A dataset may be sufficient for exploratory learning but insufficient for public-safe reporting; sufficient for a Studio demonstration but insufficient for handoff; sufficient for a DRI signal but insufficient for public authority action; sufficient for a report table but insufficient for model training.
Quality assessment should consider:
completeness;
accuracy;
consistency;
timeliness;
provenance;
representativeness;
missingness;
bias;
uncertainty;
spatial resolution;
temporal resolution;
measurement error;
collection method limitations;
transformation error;
interoperability;
documentation quality;
metadata completeness;
public-safe adequacy;
review sufficiency;
correction readiness.
A Quality Assessment Record should identify the data object, purpose assessed, methods used, findings, limitations, quality class where used, confidence, unresolved issues, recommended use, prohibited use, required corrections, review needs, and archive implications.
Quality assessment is not certification. It does not mean the data is true for all purposes, legally usable for all purposes, AI-trainable, public-safe, finance-relevant, insurance-relevant, public-authority-ready, consented, deployment-ready, or executable. It indicates whether the data is adequate for a recorded Nexus use and what limitations apply.
10.3.10 Public-Safe Transformation
Public-safe transformation converts raw, processed, controlled, restricted, geospatial, sensitive, community, Indigenous protocol-sensitive where applicable, public authority-sensitive, cyber-sensitive, infrastructure-sensitive, finance-sensitive, insurance-sensitive, or otherwise risky data into a form suitable for a defined public-safe or controlled communication purpose.
Public-safe transformation may include:
aggregation;
masking;
redaction;
de-identification;
suppression;
geospatial generalization;
spatial displacement;
time-window generalization;
field removal;
uncertainty and limitation labeling;
synthetic substitution;
protected knowledge exclusion;
public authority boundary language;
finance and insurance boundary language;
consent boundary language;
no-warning, no-procurement, no-certification, no-deployment, and no-execution notices.
A Public-Safe Transformation Record should identify source data, transformation method, target audience, release class, residual risk, reviewer class, fields removed or modified, limitations, permitted outputs, prohibited outputs, public-safe status, correction pathway, withdrawal pathway, recall pathway, and archive rule.
Public-safe transformation does not make all underlying data open. It creates a specific transformed object for a defined use. If the source data is corrected, rights change, linkage risk changes, public authority context changes, or safeguard concerns arise, the transformed output may require correction, restriction, withdrawal, recall, or archive.
10.3.11 Repository Routing
Repository routing determines where a data object, metadata-only record, data dictionary, codebook, schema, processed dataset, public-safe dataset, benchmark dataset, DRI dataset, Observatory dataset, National Portfolio dataset, or handoff data package should be stored, mirrored, accessed, versioned, corrected, and archived.
Repository routing may direct data to:
open public repository;
controlled Nexus repository;
restricted repository;
National Node repository;
sovereign data repository;
secure-room repository;
data-room repository;
clean-room environment;
compute-to-data environment;
protected knowledge repository;
public authority learning repository;
Studio repository;
Reports repository;
Academy repository;
Registry metadata repository;
Marketplace metadata record;
handoff-recipient repository;
archive repository;
sealed or non-public storage;
deletion or non-continuation pathway.
Repository routing should follow rights review, sensitivity review, data-use labeling, AI-use labeling, and jurisdictional context. Repository convenience must not override data sovereignty, privacy, cyber, protected knowledge, community, Indigenous protocol where applicable, public authority, legal, or public-safe controls.
A Repository Routing Record should identify selected repository, rationale, access class, data-use limits, AI-use limits, cross-border conditions, mirroring conditions, retention conditions, correction and recall pathway, and archive rule.
Repository placement is not permission. A dataset in a repository remains governed by its labels and records.
10.3.12 DICE Routing
DICE routing determines how a data object moves through the data, innovation, commons, and evidence governance layer. It connects data lifecycle state to the correct DICE pathway: rights review, sensitivity review, metadata creation, public-safe transformation, data commons inclusion, Studio routing, Reports routing, Academy routing, Observatory routing, DRI routing, GRIx routing, Grid routing, Marketplace routing, Registry routing, National Portfolio routing, Nexus Universe routing, handoff routing, correction routing, or archive routing.
DICE routing should identify:
whether the object may enter a Data Commons and under what access class;
whether only metadata may be discoverable;
whether public-safe transformation is required;
whether secure-room, data-room, clean-room, or compute-to-data handling is required;
whether the object may be used for AI and under what label;
whether the object may support Reports, Academy, Studio, Campaigns, Foundry, Grid, Registry, Marketplace, National Portfolios, Nexus Universe, or handoff packages;
whether the object must be restricted, sealed, deleted, archived, or marked non-continuing;
what correction propagation rules apply.
DICE routing is the governance bridge between data availability and data use. It prevents data from moving because it is technically easy rather than because it is lawful, safe, and appropriate.
DICE routing does not approve external action. It routes data through Nexus public-good pathways while preserving rights, restrictions, and no-conversion boundaries.
10.3.13 Registry Record
A Registry Record for data records the data object’s identity, lifecycle state, access class, data-use label, AI-use label, rights status, sensitivity class, public-safe status, source review, rights review, quality assessment, lineage, support status, repository routing, Marketplace relationship, Studio relationship, Reports relationship, Grid relationship, National Portfolio relationship, Nexus Universe relationship, handoff relationship, correction history, and archive state.
A Data Registry Record should identify:
data object identifier;
title and description;
source pathway;
steward;
data class;
version;
jurisdictional context;
access class;
data-use label;
AI-use label;
rights review status;
sensitivity review status;
quality assessment status;
public-safe transformation status;
repository location or metadata-only location;
support and update status;
linked objects;
correction, withdrawal, recall, supersession, archive, and non-continuation status.
Registry recording creates status truth. It does not make the data open, accurate for all purposes, AI-trainable, certified, public-authority-approved, financeable, insurable, consented, deployable, or executable. It states what the data’s current Nexus status is and what boundaries apply.
10.3.14 Marketplace Listing
Marketplace listing for data allows a data object, metadata-only record, public-safe dataset, data dictionary, codebook, schema, benchmark dataset, DRI dataset, Observatory dataset, National Portfolio dataset, or handoff data context to become discoverable under Marketplace governance.
A data Marketplace listing should identify:
listed object identity and version;
Registry status;
data class;
steward;
access class;
data-use label;
AI-use label;
sensitivity class;
license or rights summary;
public-safe status;
quality assessment summary;
update and support status;
allowed discovery pathway;
access request pathway where applicable;
prohibited uses;
correction and delisting pathway.
Marketplace listing may be public, controlled, metadata-only, National Node-only, public-authority-learning-only, secure-room-aware, data-room-aware, handoff-awareness-only, or archive-listed. A listing should not expose sensitive metadata where metadata itself is restricted.
Marketplace listing is discovery, not permission. A data listing is not a license to download, reuse, train AI, publish, combine, commercialize, hand off, or deploy. It is not procurement, certification, finance, insurance, public authority action, consent, or execution.
10.3.15 Correction
Correction is the data lifecycle process for addressing errors, rights issues, sensitivity issues, quality defects, lineage gaps, metadata errors, data-use label errors, AI-use label errors, public-safe transformation errors, repository routing errors, Registry errors, Marketplace listing errors, Studio workflow errors, Reports errors, Academy errors, Campaign errors, Grid errors, National Portfolio errors, Nexus Universe errors, or handoff package errors affecting a data object.
Data correction may include:
metadata correction;
source correction;
rights correction;
sensitivity reclassification;
data-use label correction;
AI-use label correction;
quality correction;
lineage correction;
transformation correction;
public-safe output correction;