アイヌ語聖書 ジョン・バチェラー訳 新約聖書(1897年)・詩編・ヨナ書

Analytics · 計量

Measurements over the selected text layer; the modern layer is the machine re-transcription still under editorial review. The stylometric distances below are exploratory: they register differences in habitual small-word usage between books, which may reflect genre, source text, or the several Ainu writers Batchelor names in his preface — the numbers alone cannot say which.

278,229
tokens (1897 spelling)
5,938
distinct word forms
10,960
verses
29
books
137
stylometric chunks (≥1,500 tokens)

Most frequent words

Top 30 word forms of the 1897 spelling, whole corpus.

Dispersion

Where selected words occur across the volume's 10,546 verses, in canonical order. Proper names first, then the commonest content words.

matmrklukjhnactrom1corevpsayesukiristokamuinukaratankonyaikotakarapirikashuishinetap

Keywords per book (log-likelihood)

Words significantly overused in one book against the rest of the corpus (Dunning G², p < .001), with %DIFF effect size.

wordcount%DIFF
itak634147.0+76%
sangere39128.8+3,245%
ene305113.6+109%
wa83078.6+41%
hi30476.1+80%
koikara8170.8+241%
oman8251.5+172%
araki8451.2+167%
koye1149.7+37,863,141,852%
orota14248.7+102%

Function-word fingerprint

Per-book usage rate of the 25 commonest word forms, z-scored across books (blue = below corpus norm, red = above). The per-book profile of small words is the classic authorship signal.

matmrklukjhnactrom1co2cogalephphpcol1th2th1ti2tititphmhebjas1pe2pe1jn2jn3jnjudrevpsajon
ne
utara
ruwe
gusu
anak
an
wa
nei
no
koro
orowa
ku
itak
shinuma
otta
nisa
na
echi
guru
ambe
kuni
kamui
ki
e
kusu

Verse-length ratio against the parallels

Mean Ainu tokens per Greek token (and per English RV token where that layer exists), verse-aligned. Values above 1 mean the Ainu rendering runs longer than its source.

vs Greek (WH) vs English (RV 1885) | = ratio 1.0

Distance between books (Burrows delta)

Deeper color = greater stylometric distance (1897 spelling). Features:

matmrklukjhnactrom1co2cogalephphpcol1th2th1ti2tititphmhebjas1pe2pe1jn2jn3jnjudrevpsajon
mat
mrk
luk
jhn
act
rom
1co
2co
gal
eph
php
col
1th
2th
1ti
2ti
tit
phm
heb
jas
1pe
2pe
1jn
2jn
3jn
jud
rev
psa
jon

Stylometric map of the books

Classical MDS of the word-delta matrix — books that sit close use the corpus's frequent words at similar rates. Colors mark the traditional canon divisions, drawn on for orientation only.

matmrklukjhnactrom1co2cogalephphpcol1th2th1ti2tititphmhebjas1pe2pe1jn2jn3jnjudrevpsajon
Gospels · ActsPauline epistlesGeneral epistlesRevelationSeparate volume

Topics (NMF over chapters)

Five components of the chapter × content-word tf-idf matrix. A topic is a weighted word list; a book's bar shows how its chapters load on each. Components register vocabulary domains, which here track genre.

topic 1
ramye
yah
humi
kando
ushiketa
yanro
koto
rappa
kotoro
so
topic 2
aa
shinotcha
kashiobiuki
oupeka
kara
iyohaichish
wen
dabid
sera
goro
topic 3
sangere
yosep
babironia
ikiri
atura
machi
dabid
yuda
maria
yakob
topic 4
yehoba
yona
hauturumbe
atui
chikoikip
arawan
amset
shine
kata
hawe
topic 5
kiristo
yaikota
tan
iriwak
pirika
tuitak
rai
eishokor
eraman
tap

Collocations (PMI)

Strongest two-word bonds of the 1897 spelling, bigram count ≥ 25.

collocationcountPMI
yaku etaye2511.70
sen nin3211.58
yaisambepokash humi2511.29
utasa chikuni6810.82
kambi sosh5410.22
okarituye buri7110.11
shimon tek4710.05
tokap chup299.88
ikashima wan619.86
kora kenru2649.84
yange set309.81
ashikne hotne269.65
otek sak549.55
buri akire339.53
ko oripak669.52
tun ikashima299.51
hitsuji chikoikip429.43
tokap to389.41
shimge toho279.41
upakashnu chisei259.35

Rank–frequency (Zipf)

Word frequency against frequency rank, both axes logarithmic.

rank (log)frequency (log)

Per-book measurements

bookchaptersversestokenstypeshapaxYule KHonoré RTTRtokens / versereviewed
Matthewマタイ傳福音書281,09929,0521,882779148.81,7530.06526.40
Markマルコ傳福音書1669117,8541,4106031381,7110.07925.80
Lukeルカ傳福音書241,16931,3541,930828150.11,8130.06226.80
Johnヨハネ傳福音書2189824,8031,170432187.41,6040.04727.60
Acts使徒行傳281,02729,3631,921840139.21,8280.06528.60
Romansロマ書1644711,8141,0634972011,7610.09026.40
1 Corinthiansコリント前書1645211,8691,050467200.81,6900.08826.30
2 Corinthiansコリント後書132688,165880436219.41,7850.10830.50
Galatiansガラテヤ書61554,097564267210.71,5800.13826.40
Ephesiansエペソ書61613,679598318173.21,7540.16322.90
Philippiansピリピ書41082,868446216180.41,5440.15526.60
Colossiansコロサイ書4992,568494257159.61,6360.19225.90
1 Thessaloniansテサロニケ前書5922,5164232052001,5190.16827.30
2 Thessaloniansテサロニケ後書3501,300296154188.71,4950.228260
1 Timothyテモテ前書61193,026580311170.11,7280.19225.40
2 Timothyテモテ後書4862,172475265161.61,7380.21925.30
Titusテトス書3481,222328187171.31,6540.26825.50
Philemonピレモン書12562818297230.91,3790.29025.10
Hebrewsヘブル書133138,7181,032492179.91,7340.11827.90
Jamesヤコブ書51132,849582311180.31,7080.20425.20
1 Peterペテロ前書51102,974578325180.21,8270.194270
2 Peterペテロ後書3641,922438230155.21,5920.228300
1 Johnヨハネ第一書51103,213330135280.41,3670.10329.20
2 Johnヨハネ第二書11437613673197.81,2800.36226.90
3 Johnヨハネ第三書11538213672214.91,2630.35625.50
Judeユダ書126756264162160.51,7150.34929.10
Revelationヨハネ默示錄2242514,6231,202470148.81,5750.08234.40
Psalms詩篇1502,72552,5182,5971,139184.51,9360.04919.30
Jonahヨナ書4511,548318147148.41,3660.20530.40

Burrows delta over chapter chunks of the 1897 text: 150 most frequent words (and, separately, 300 most frequent character 4-grams), relative frequencies z-scored across chunks, mean absolute z-difference between book centroids; 2D map by classical MDS of the word-delta matrix.

Method & prior art

  • Burrows delta: J. Burrows, Literary and Linguistic Computing 17(3), 2002. Function-word fingerprints per book follow A. Kenny, A Stylometric Study of the New Testament, 1986.
  • Keywords: log-likelihood G² per T. Dunning, Computational Linguistics 19(1), 1993, with %DIFF effect size per Gabrielatos & Marchi 2015. Applied here to a biblical corpus without a direct published precedent.
  • Vocabulary richness: Yule's K (G. U. Yule, The Statistical Study of Literary Vocabulary, 1944) and Honoré's R (A. Honoré 1979), as used in the authorship literature on the Greek NT. Word-level tokens of the 1897 spelling; Ainu polysynthesis makes tokenization choices decisive, so these are comparable within this corpus only.
  • Dispersion plots continue the concordance tradition in its modern word-offset form.
  • Verse-length ratios use the verse alignment that also underlies the massively parallel Bible corpora (Christodoulopoulos & Steedman 2015; Mayer & Cysouw 2014).
  • The NMF topic view has no published Bible-specific precedent and is exploratory.
バチェラー訳 1897年 · 原文:パブリックドメイン · 編集層: CC BY 4.0 · 画像:Internet Archive bible.aynu.org