The nation decides what its models are allowed to learn from — source by source, license by license. Inclusion is a governed decision, not a scraping accident.
Every source that reaches training passes through an explicit gate the owner defines and can inspect.
Each dataset is registered with its origin, custodian, legal basis, and intended use before ingestion. Unregistered data has no path into a training run.
Admission rules — licensing, classification, jurisdiction, consent status — are expressed as machine-checked policy. A source that fails any rule is quarantined rather than silently included.
Categories the owner marks as sensitive require a named approver's signature. The approval, approver, and rationale are recorded on the tamper-evident ledger.
There is no ambient crawler pulling whatever it finds. The corpus grows only through deliberate, attributed additions the owner authorized.
What the model should know, forget, and never touch is a sovereign editorial choice made explicit in the pipeline.
Owners define allowlists and denylists at the source, domain, and document level. Financial regulation, official records, and national-language material can be prioritized deliberately.
Sampling weights let the owner over- or under-represent domains — for example emphasizing in-nation financial and legal text over generic web content — with the mixture recorded per run.
State secrets, personal data, and restricted material are handled under owner-set rules: excluded, redacted, or confined to isolated enclaves that never reach a general model.
Curation can target the nation's official languages and legal corpus so the model reasons in the terms its institutions actually use.
The right to train on a piece of data is tracked at the granularity of that data.
Licenses, consent scope, and usage restrictions are stored alongside each source and enforced downstream. A restriction on one source never has to be inferred from a broad policy.
The system can answer whether a given document was permitted for training at the time a given model was built, using the consent state of record for that moment.
When a legal basis is revoked, the affected documents are flagged across the corpus and their influence is scheduled for removal from future and, where required, existing models.
Personal data can be detected, tokenized, or excluded per policy, and its presence is recorded so the owner always knows what personal information a model may have seen.
Corpus decisions are actions taken by named people under recorded authority, not opaque configuration.
Different institutions and officers hold different rights over admission, weighting, and retirement. Authority maps to real organizational mandate, not shared credentials.
Adding a source, changing a weight, or excluding a domain is a signed event on the hash-chained ledger, reconstructable long after the fact.
Owners can preview how a corpus change would alter the training mixture before committing, so editorial decisions are made with their downstream effect visible.
An auditor can be granted a read-only view of what the corpus contains and why, without the ability to alter it, supporting external review of what the nation's models learned from.
Talk to us about owner-controlled corpus in a sovereign deployment.