Four-domain samples
← Back to the four-domain report
Every window that the four-domain comparison trained on or measured with: 1,200 training windows, 32 held-out windows and the 64-window never-trained control. Token ids are the executed sample. The text is decoded from them for reading and is not normalized, so a window can start or end inside a word or a multi-byte character. Index values are zero-based. The rows are in the order the Learner 1.0 conditions trained them; LoRA trained the same 300 windows of each domain in a seeded shuffled order. The Yoruba, Amharic, news and control text comes from third-party sources and keeps its source terms.
Token ids ()
Downloads
| file | set | split | rows |
|---|
windows_index.json holds the sources, revisions, counts and a checksum for every window. README.md describes the fields.