Analytics · 計量
Measurements over the selected text layer; the modern layer is the machine re-transcription still under editorial review. The stylometric distances below are exploratory: they register differences in habitual small-word usage between books, which may reflect genre, source text, or the several Ainu writers Batchelor names in his preface — the numbers alone cannot say which.
Most frequent words
Top 30 word forms of the 1897 spelling, whole corpus.
Dispersion
Where selected words occur across the volume's 10,546 verses, in canonical order. Proper names first, then the commonest content words.
Keywords per book (log-likelihood)
Words significantly overused in one book against the rest of the corpus (Dunning G², p < .001), with %DIFF effect size.
| word | count | G² | %DIFF |
|---|---|---|---|
| itak | 634 | 147.0 | +76% |
| sangere | 39 | 128.8 | +3,245% |
| ene | 305 | 113.6 | +109% |
| wa | 830 | 78.6 | +41% |
| hi | 304 | 76.1 | +80% |
| koikara | 81 | 70.8 | +241% |
| oman | 82 | 51.5 | +172% |
| araki | 84 | 51.2 | +167% |
| koye | 11 | 49.7 | +37,863,141,852% |
| orota | 142 | 48.7 | +102% |
Function-word fingerprint
Per-book usage rate of the 25 commonest word forms, z-scored across books (blue = below corpus norm, red = above). The per-book profile of small words is the classic authorship signal.
| mat | mrk | luk | jhn | act | rom | 1co | 2co | gal | eph | php | col | 1th | 2th | 1ti | 2ti | tit | phm | heb | jas | 1pe | 2pe | 1jn | 2jn | 3jn | jud | rev | psa | jon | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ne | |||||||||||||||||||||||||||||
| utara | |||||||||||||||||||||||||||||
| ruwe | |||||||||||||||||||||||||||||
| gusu | |||||||||||||||||||||||||||||
| anak | |||||||||||||||||||||||||||||
| an | |||||||||||||||||||||||||||||
| wa | |||||||||||||||||||||||||||||
| nei | |||||||||||||||||||||||||||||
| no | |||||||||||||||||||||||||||||
| koro | |||||||||||||||||||||||||||||
| orowa | |||||||||||||||||||||||||||||
| ku | |||||||||||||||||||||||||||||
| itak | |||||||||||||||||||||||||||||
| shinuma | |||||||||||||||||||||||||||||
| otta | |||||||||||||||||||||||||||||
| nisa | |||||||||||||||||||||||||||||
| na | |||||||||||||||||||||||||||||
| echi | |||||||||||||||||||||||||||||
| guru | |||||||||||||||||||||||||||||
| ambe | |||||||||||||||||||||||||||||
| kuni | |||||||||||||||||||||||||||||
| kamui | |||||||||||||||||||||||||||||
| ki | |||||||||||||||||||||||||||||
| e | |||||||||||||||||||||||||||||
| kusu |
Verse-length ratio against the parallels
Mean Ainu tokens per Greek token (and per English RV token where that layer exists), verse-aligned. Values above 1 mean the Ainu rendering runs longer than its source.
Distance between books (Burrows delta)
Deeper color = greater stylometric distance (1897 spelling). Features:
| mat | mrk | luk | jhn | act | rom | 1co | 2co | gal | eph | php | col | 1th | 2th | 1ti | 2ti | tit | phm | heb | jas | 1pe | 2pe | 1jn | 2jn | 3jn | jud | rev | psa | jon | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| mat | |||||||||||||||||||||||||||||
| mrk | |||||||||||||||||||||||||||||
| luk | |||||||||||||||||||||||||||||
| jhn | |||||||||||||||||||||||||||||
| act | |||||||||||||||||||||||||||||
| rom | |||||||||||||||||||||||||||||
| 1co | |||||||||||||||||||||||||||||
| 2co | |||||||||||||||||||||||||||||
| gal | |||||||||||||||||||||||||||||
| eph | |||||||||||||||||||||||||||||
| php | |||||||||||||||||||||||||||||
| col | |||||||||||||||||||||||||||||
| 1th | |||||||||||||||||||||||||||||
| 2th | |||||||||||||||||||||||||||||
| 1ti | |||||||||||||||||||||||||||||
| 2ti | |||||||||||||||||||||||||||||
| tit | |||||||||||||||||||||||||||||
| phm | |||||||||||||||||||||||||||||
| heb | |||||||||||||||||||||||||||||
| jas | |||||||||||||||||||||||||||||
| 1pe | |||||||||||||||||||||||||||||
| 2pe | |||||||||||||||||||||||||||||
| 1jn | |||||||||||||||||||||||||||||
| 2jn | |||||||||||||||||||||||||||||
| 3jn | |||||||||||||||||||||||||||||
| jud | |||||||||||||||||||||||||||||
| rev | |||||||||||||||||||||||||||||
| psa | |||||||||||||||||||||||||||||
| jon |
Stylometric map of the books
Classical MDS of the word-delta matrix — books that sit close use the corpus's frequent words at similar rates. Colors mark the traditional canon divisions, drawn on for orientation only.
Topics (NMF over chapters)
Five components of the chapter × content-word tf-idf matrix. A topic is a weighted word list; a book's bar shows how its chapters load on each. Components register vocabulary domains, which here track genre.
Collocations (PMI)
Strongest two-word bonds of the 1897 spelling, bigram count ≥ 25.
| collocation | count | PMI |
|---|---|---|
| yaku etaye | 25 | 11.70 |
| sen nin | 32 | 11.58 |
| yaisambepokash humi | 25 | 11.29 |
| utasa chikuni | 68 | 10.82 |
| kambi sosh | 54 | 10.22 |
| okarituye buri | 71 | 10.11 |
| shimon tek | 47 | 10.05 |
| tokap chup | 29 | 9.88 |
| ikashima wan | 61 | 9.86 |
| kora kenru | 264 | 9.84 |
| yange set | 30 | 9.81 |
| ashikne hotne | 26 | 9.65 |
| otek sak | 54 | 9.55 |
| buri akire | 33 | 9.53 |
| ko oripak | 66 | 9.52 |
| tun ikashima | 29 | 9.51 |
| hitsuji chikoikip | 42 | 9.43 |
| tokap to | 38 | 9.41 |
| shimge toho | 27 | 9.41 |
| upakashnu chisei | 25 | 9.35 |
Rank–frequency (Zipf)
Word frequency against frequency rank, both axes logarithmic.
Per-book measurements
| book | chapters | verses | tokens | types | hapax | Yule K | Honoré R | TTR | tokens / verse | reviewed |
|---|---|---|---|---|---|---|---|---|---|---|
| Matthewマタイ傳福音書 | 28 | 1,099 | 29,052 | 1,882 | 779 | 148.8 | 1,753 | 0.065 | 26.4 | 0 |
| Markマルコ傳福音書 | 16 | 691 | 17,854 | 1,410 | 603 | 138 | 1,711 | 0.079 | 25.8 | 0 |
| Lukeルカ傳福音書 | 24 | 1,169 | 31,354 | 1,930 | 828 | 150.1 | 1,813 | 0.062 | 26.8 | 0 |
| Johnヨハネ傳福音書 | 21 | 898 | 24,803 | 1,170 | 432 | 187.4 | 1,604 | 0.047 | 27.6 | 0 |
| Acts使徒行傳 | 28 | 1,027 | 29,363 | 1,921 | 840 | 139.2 | 1,828 | 0.065 | 28.6 | 0 |
| Romansロマ書 | 16 | 447 | 11,814 | 1,063 | 497 | 201 | 1,761 | 0.090 | 26.4 | 0 |
| 1 Corinthiansコリント前書 | 16 | 452 | 11,869 | 1,050 | 467 | 200.8 | 1,690 | 0.088 | 26.3 | 0 |
| 2 Corinthiansコリント後書 | 13 | 268 | 8,165 | 880 | 436 | 219.4 | 1,785 | 0.108 | 30.5 | 0 |
| Galatiansガラテヤ書 | 6 | 155 | 4,097 | 564 | 267 | 210.7 | 1,580 | 0.138 | 26.4 | 0 |
| Ephesiansエペソ書 | 6 | 161 | 3,679 | 598 | 318 | 173.2 | 1,754 | 0.163 | 22.9 | 0 |
| Philippiansピリピ書 | 4 | 108 | 2,868 | 446 | 216 | 180.4 | 1,544 | 0.155 | 26.6 | 0 |
| Colossiansコロサイ書 | 4 | 99 | 2,568 | 494 | 257 | 159.6 | 1,636 | 0.192 | 25.9 | 0 |
| 1 Thessaloniansテサロニケ前書 | 5 | 92 | 2,516 | 423 | 205 | 200 | 1,519 | 0.168 | 27.3 | 0 |
| 2 Thessaloniansテサロニケ後書 | 3 | 50 | 1,300 | 296 | 154 | 188.7 | 1,495 | 0.228 | 26 | 0 |
| 1 Timothyテモテ前書 | 6 | 119 | 3,026 | 580 | 311 | 170.1 | 1,728 | 0.192 | 25.4 | 0 |
| 2 Timothyテモテ後書 | 4 | 86 | 2,172 | 475 | 265 | 161.6 | 1,738 | 0.219 | 25.3 | 0 |
| Titusテトス書 | 3 | 48 | 1,222 | 328 | 187 | 171.3 | 1,654 | 0.268 | 25.5 | 0 |
| Philemonピレモン書 | 1 | 25 | 628 | 182 | 97 | 230.9 | 1,379 | 0.290 | 25.1 | 0 |
| Hebrewsヘブル書 | 13 | 313 | 8,718 | 1,032 | 492 | 179.9 | 1,734 | 0.118 | 27.9 | 0 |
| Jamesヤコブ書 | 5 | 113 | 2,849 | 582 | 311 | 180.3 | 1,708 | 0.204 | 25.2 | 0 |
| 1 Peterペテロ前書 | 5 | 110 | 2,974 | 578 | 325 | 180.2 | 1,827 | 0.194 | 27 | 0 |
| 2 Peterペテロ後書 | 3 | 64 | 1,922 | 438 | 230 | 155.2 | 1,592 | 0.228 | 30 | 0 |
| 1 Johnヨハネ第一書 | 5 | 110 | 3,213 | 330 | 135 | 280.4 | 1,367 | 0.103 | 29.2 | 0 |
| 2 Johnヨハネ第二書 | 1 | 14 | 376 | 136 | 73 | 197.8 | 1,280 | 0.362 | 26.9 | 0 |
| 3 Johnヨハネ第三書 | 1 | 15 | 382 | 136 | 72 | 214.9 | 1,263 | 0.356 | 25.5 | 0 |
| Judeユダ書 | 1 | 26 | 756 | 264 | 162 | 160.5 | 1,715 | 0.349 | 29.1 | 0 |
| Revelationヨハネ默示錄 | 22 | 425 | 14,623 | 1,202 | 470 | 148.8 | 1,575 | 0.082 | 34.4 | 0 |
| Psalms詩篇 | 150 | 2,725 | 52,518 | 2,597 | 1,139 | 184.5 | 1,936 | 0.049 | 19.3 | 0 |
| Jonahヨナ書 | 4 | 51 | 1,548 | 318 | 147 | 148.4 | 1,366 | 0.205 | 30.4 | 0 |
Burrows delta over chapter chunks of the 1897 text: 150 most frequent words (and, separately, 300 most frequent character 4-grams), relative frequencies z-scored across chunks, mean absolute z-difference between book centroids; 2D map by classical MDS of the word-delta matrix.
Method & prior art
- Burrows delta: J. Burrows, Literary and Linguistic Computing 17(3), 2002. Function-word fingerprints per book follow A. Kenny, A Stylometric Study of the New Testament, 1986.
- Keywords: log-likelihood G² per T. Dunning, Computational Linguistics 19(1), 1993, with %DIFF effect size per Gabrielatos & Marchi 2015. Applied here to a biblical corpus without a direct published precedent.
- Vocabulary richness: Yule's K (G. U. Yule, The Statistical Study of Literary Vocabulary, 1944) and Honoré's R (A. Honoré 1979), as used in the authorship literature on the Greek NT. Word-level tokens of the 1897 spelling; Ainu polysynthesis makes tokenization choices decisive, so these are comparable within this corpus only.
- Dispersion plots continue the concordance tradition in its modern word-offset form.
- Verse-length ratios use the verse alignment that also underlies the massively parallel Bible corpora (Christodoulopoulos & Steedman 2015; Mayer & Cysouw 2014).
- The NMF topic view has no published Bible-specific precedent and is exploratory.