Paper deep dive
Automatic Speech Recognition for Documenting Endangered Languages: Case Study of Ikema Miyakoan
Chihiro Taguchi, Yukinori Takubo, David Chiang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/31/2026, 1:40:38 AM
Summary
This paper presents the development of an automatic speech recognition (ASR) system for Ikema, a severely endangered Ryukyuan language. The authors constructed a 6.33-hour speech corpus, trained a Wav2Vec2-based ASR model achieving a 14.8% character error rate, and demonstrated that integrating this model into the ELAN annotation software significantly reduces transcription time and cognitive load for researchers.
Entities (5)
Relation Signals (3)
ASR system for Ikema â integratedinto â ELAN
confidence 100% · The integration of the ASR model into the linguistic annotation tool (ELAN)
Chihiro Taguchi â developed â ASR system for Ikema
confidence 95% · We present an ongoing effort to develop an ASR system for Ikema
Wav2Vec2 â usedfor â ASR system for Ikema
confidence 95% · we fine-tune pretrained Wav2Vec2 models... for Ikema
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Language endangerment poses a major challenge to linguistic diversity worldwide, and technological advances have opened new avenues for documentation and revitalization. Among these, automatic speech recognition (ASR) has shown increasing potential to assist in the transcription of endangered language data. This study focuses on Ikema, a severely endangered Ryukyuan language spoken in Okinawa, Japan, with approximately 1,300 remaining speakers, most of whom are over 60 years old. We present an ongoing effort to develop an ASR system for Ikema based on field recordings. Specifically, we (1) construct a {\totaldatasethours}-hour speech corpus from field recordings, (2) train an ASR model that achieves a character error rate as low as 15\%, and (3) evaluate the impact of ASR assistance on the efficiency of speech transcription. Our results demonstrate that ASR integration can substantially reduce transcription time and cognitive load, offering a practical pathway toward scalable, technology-supported documentation of endangered languages.
Tags
Links
- Source: https://arxiv.org/abs/2603.26248v1
- Canonical: https://arxiv.org/abs/2603.26248v1
Trouble viewing inline? Open PDF directly â
Full Text
32,399 characters extracted from source content.
Expand or collapse full text
Automatic Speech Recognition for Documenting Endangered Languages: Case Study of Ikema Miyakoan Chihiro Taguchi 1 , Yukinori Takubo 2 , David Chiang 1 1 University of Notre Dame, 2 National Institute for Japanese Language and Linguistics 1 Notre Dame, IN, USA, 2 Tokyo, Japan ctaguchi@nd.edu, ytakubo@ninjal.ac.jp, dchiang@nd.edu Abstract Language endangerment poses a major challenge to linguistic diversity worldwide, and technological advances have opened new avenues for documentation and revitalization. Among these, automatic speech recognition (ASR) has shown increasing potential to assist in the transcription of endangered language data. This study focuses on Ikema, a severely endangered Ryukyuan language spoken in Okinawa, Japan, with approximately 1,300 remaining speakers, most of whom are over 60 years old. We present an ongoing effort to develop an ASR system for Ikema based on field recordings. Specifically, we (1) construct a 6.33-hour speech corpus from field recordings, (2) train an ASR model that achieves a character error rate as low as 15%, and (3) evaluate the impact of ASR assistance on the efficiency of speech transcription. Our results demonstrate that ASR integration can substantially reduce transcription time and cognitive load, offering a practical pathway toward scalable, technology-supported documentation of endangered languages. Keywords: language documentation, automatic speech recognition, language resources 1. Introduction Language endangerment is a pressing global issue, with thousands of languages at risk of disappearing within the coming decades. Recent advancements in language technologies have opened up new op- portunities for language documentation, offering computational tools that can assist researchers in preserving linguistic data more efficiently. In par- ticular, automatic speech recognition (ASR) has emerged as a promising technique to support the transcription of spoken data from endangered lan- guages. Ikema Miyakoan (Ikema henceforth), a highly endangered Ryukyuan language spoken on the Miyako Islands in Japan, represents one such case. With an estimated speaker population of approxi- mately 1,300 individuals, most of whom are over 60 years old, the language faces severe risk of extinction (Nakama et al., 2025). In this context, leveraging speech technologies for Ikema is both urgent and valuable for documentation and revital- ization efforts. This paper reports on our ongoing work to de- velop a speech dataset and an ASR system for Ikema as part of a broader language documenta- tion project. We describe the construction of a 6.33- hour speech dataset collected from field recordings, pronounced dictionary entries, and audio books. Then, we report the training of an ASR model that achieves a character error rate (CER) as low as 14.8%. The trained model is then integrated into the linguistic annotation software, and we evalu- ate how much ASR can facilitate the transcription process. Our findings indicate that ASR-assisted transcription not only speeds up annotation but also reduces the cognitive burden on annotators, high- lighting the practical benefits of integrating ASR into endangered language documentation workflows. The key contributions of this study are: âąThe construction of the first speech dataset for Ikema; âąThe development of the first ASR model for Ikema with a CER as low as 14.8%; âąThe integration of the ASR model into the lin- guistic annotation tool (ELAN); âąPresenting the positive evidence for the poten- tial of ASR-assisted transcription in language documentation. The dataset, the model, and the experimental code developed in this study are publicly released. 1 2. Related Work 2.1. Language Ikema is an endangered Japonic language spoken in the Miyako Islands of Okinawa, Japan. Its linguis- tic classification is illustrated in Figure 1. The lan- guage is spoken in three villages: Ikema Island, the Nishihara village on Miyako Island, and the Sara- hama village on Irabu Island, as shown in Figure 2. A recent study predicts that Ikema will become âcrit- ically endangeredâ according to UNESCOâs criteria within 30 years, when the youngest active speakers 1 https://github.com/ctaguchi/ikema_asr arXiv:2603.26248v1 [cs.CL] 27 Mar 2026 reach 90 years of age (Nakama et al., 2025). Never- theless, Ikema is comparatively better studied and documented than other Miyakoan varieties, with an existing descriptive grammar (Hayashi, 2013) and dictionary (Nakama et al., 2025). Yamada et al. (2020) report that Ikema is largely unintelligible to speakers of Japonic languages outside the South- ern Ryukyuan subgroup. Although Ikema currently lacks an official orthog- raphy, several writing systems have been proposed to document the language. In linguistic literature on the Ikema grammar, phonemic transcription using either the International Phonetic Alphabet (IPA) or romanization has been commonly employed. In contrast, the speaker community more frequently uses hiragana or katakana (collectively referred to as kana) which are syllabic writing systems used in Japanese (Takubo, 2021a). The existing dictionary (Nakama et al., 2025) defines a phonemic kana orthography for Ikema, incorporating several ex- tensions to represent phonemes absent in Modern Standard Japanese, such as a syllabic devoiced nasal consonant. The phonemic kana orthography can be deterministically converted into romanized script. Example (1) shows a comparison of the three writing systems: (1a) for kana, (1b) for phone- mic romanization (romaji henceforth), and (1c) for phonemic IPA. For reference, (1d) 2 provides the gloss and the translation of the example. (1) a. ăŁă ăăăŒ ăïŸăŹ ăȘăăă©ă ă»ă ăŒăăŒ b. vvaa N Ì nu nauyudu huutaa c. vvaa n Ì nu naujudu huutaa d. vva-a 2sg-top n Ì nu yesterday nau-ju-du what-acc-foc huutaa do.prog.pst.q âWhat were you doing yesterday?â 2.2. Speech recognition Speech recognition technologies have advanced drastically with the development of deep learning, and their applications have increasingly extended to low-resource and endangered languages. Such efforts have led to the development of ASR systems for Nasal (Billings and McDonnell, 2025), Newar and Dzardzongke (OâNeill et al., 2023), Sardinian (Chizzoni and Vietti, 2024), Mvskoke (Mainzinger and Levow, 2024), Kichwa (Taguchi et al., 2024), Japhug (Guillaume et al., 2022), and Khinalug (Li et al., 2024), among others. Notably, several of 2 2sg: 2nd person singular pronoun; acc: accusative case; foc: focus marker; prog: progressive aspect; pst: past tense; q: interrogative; top: topic marker Dataset#Samples Duration Avg. duration (sec)(sec) Field11126 160581.44 Dictionary568053790.95 Audiobooks28513484.73 Total17091 227851.33 Table 1: Statistics of the datasets. these studies emphasize the use of ASR systems to accelerate the transcription process, which is an essential yet labor-intensive stage of language documentation. Transcription requires significant time investment and specialized linguistic literacy and is often referred to as the âtranscription bot- tleneckâ in language documentation (Seifart et al., 2018). Since documenting endangered languages is inherently time-sensitive, mitigating this bottle- neck through ASR support represents a promising direction. Indeed, case studies on Yongning Na (Michaud et al., 2018) and Seneca (Jimerson and Prudâhommeaux, 2018) have demonstrated the util- ity of ASR integration in transcription workflows, where ASR outputs serve as drafts for human post- editing. While these studies confirm the usefulness of the approach, quantitative evidence on the extent to which ASR accelerates transcription in practice remains limited. 3. Dataset The dataset constructed in this study is com- posed of three sources: video recordings collected through nearly twenty years of fieldwork (hereafter, âFieldâ data), pronounced entries from the Ikema dictionary (Nakama et al., 2025) (hereafter, âDic- tionaryâ data), and audiobooks. A large portion of the Field data consists of semi-spontaneous monologues by a single male speaker. Most of these monologues are planned speech, in which the speaker was either provided with a specific topic or has read a prepared script before recording. The Dictionary data contain scripted speech produced by the same male speaker. Most audio segments in the Dictionary data are shorter than one second, as each recording corresponds to a single dictionary entry (i.e., a word). The audiobook data consist of stories read aloud by a female speaker. See Table 1 for the detailed statistics of the data. The video recordings were first converted to WAV files, after which the audio was segmented and an- notated with transcriptions. Transcriptions were provided in kana in a phonetically faithful manner; that is, utterances were transcribed as they were pronounced, including fillers, disfluencies, and rep- etitions. After the kana transcription was completed, it was transliterated into romaji using deterministic Japonic JapaneseRyukyuan Southern Ryukyuan MacroâYaeyamaMiyako MiyakoTarama Ì OgamiIrabuIkema Northern Ryukyuan AmamiOkinawa Figure 1: A classification of Japonic languages and Ikemaâs position thereof. The classification is largely based on Shimoji (2008) and Pellard (2015). Figure 2: Map of the Miyako Islands (bottom) in Japan (top). The Ikema-speaking villages are marked with a red circle. mapping rules. In the dataset, word boundaries are indicated by a half-width space (U+0009). The transcriptions include several types of tags that provide detailed linguistic information, such as code-switched segments in Japanese (<ja>...</ja>), disfluencies (<dis>...</dis>), songs (<song>...</song>), and personal names (<name>...</name>). In addition, utterances that an- notators were unable to understand are annotated with the tag<unsure>...</unsure>. This design allows users who wish to train models on clean transcriptions without disfluencies or repetitions to automatically remove them from the dataset, and also enables the easy exclusion of samples contain- ing personally identifiable information. Given the nature of spontaneous speech, sentence bound- aries are often difficult to determine. Therefore, segment boundaries were determined based on pauses between speech regions. In addition, the annotated data include a tier with longer segments that roughly correspond to sentence-level units, as shown in Figure 3. 4. Experiments 4.1. Setup In our experiments, we train automatic speech recognition (ASR) models on the newly devel- oped Ikema speech dataset. Specifically, we fine- tune pretrained Wav2Vec2 models (Baevski et al., 2020) using a Connectionist Temporal Classifica- tion (CTC) decoder layer (Graves et al., 2006). Wav2Vec2 is a self-supervised model that learns speech representations from large amounts of unla- beled audio. Among its multilingual variants, XLS-R (Babu et al., 2021) and MMS (Pratap et al., 2023) have been trained on 128 and 1,406 languages, respectively, enabling robust cross-lingual transfer. For Ikema, we fine-tune these pretrained models by adding a final CTC layer that maps frame-level features to textual transcriptions. The CTC layer produces a token prediction for each frame, and Figure 3: Segmentation and annotation in ELAN. The top transcription tier contains pause-based segments, while the bottom tier contains longer segments that concatenate multiple pause-based segments to approximate sentence-level units. identical consecutive tokens are collapsed into a single symbol to form the decoded output, typically at the character or grapheme level. Wav2Vec2- based models have been shown to perform ef- fectively even in extremely low-resource settings (Liang and Levow, 2025), making them well-suited for our experiments. While automatic speech recognition (ASR) mod- els with autoregressive decoders, such as Whis- per (Radford et al., 2022), achieve strong perfor- mance for high-resource languages, they require large amounts of data to train the decoder to model subword dependencies effectively. In contrast, a model with a standard CTC decoder does not ex- plicitly learn dependencies across tokens, allowing it to directly learn the alignment between the input speech and the corresponding output characters at each time frame. Furthermore, CTCâs per-frame outputs are particularly suitable for our goal of pro- ducing phonetically faithful transcriptions, as the transcribed data may contain substantial disfluen- cies and fillers that should not be omitted in lan- guage documentation. In our experiments, we compare the performance ofwav2vec2-xls-r-300m(300M parameters, 128 languages),wav2vec2-xls-r-1b(1B parameters, 128 languages), andmms-1b(1B parameters, 1,406 languages). The tags in the transcription are re- moved in the preprocessing. Of the dataset, 80% is used as the training data, 10% is reserved for the validation data, and 10% for the test data. We also train models in two writing systems: kana and romaji. The token vocabulary is constructed based on syllables (or more precisely, morae) for kana and phonemes for romaji; for example, âăăâ /Ăœa/ is counted as one token in the kana model, and âzyâ /Ăœ/ as one token in the romaji model. We report character error rate (CER) and word error rate (WER), defined as the percentage of incorrect characters or words out of 100. It is im- portant to note, however, that WER is less infor- kanaromaji ModelCER WER CER WER xls-r-300m 20.94 62.56 14.80 64.99 xls-r-1b24.07 66.61 20.33 76.73 mms-1b21.86 63.22 18.98 73.87 Table 2: ASR performances of the trained models. mative for Ikema, where word boundaries are not well defined and variations appear even in the gold transcriptions. The learning rate is set to 0.0003. Preliminary experiments showed higher learning rates exhibited to slower and unstable learning curves. The batch size was set to 16 withwav2vec2-xls-r-300mand 4 withwav2vec2-xls-r-1bandmms-1b. The model is trained for 50 epochs. All the models are trained on a single A10 GPU with 16GB RAM. 4.2. Results Table 2 shows the performance of the trained ASR models under different conditions. Compared to the kana transcription, the romaji transcription con- sistently yielded lower CERs. This is likely because the kana writing system encodes more phological information than romaji. This finding aligns with the general tendency of Wav2Vec2 models to under- perform when transcribing more complex writing systems (Taguchi and Chiang, 2024). The best performing model was the fine-tuned model based onxls-r-300min both kana and ro- maji transcriptions. On the other hand, the larger model (xls-r-1b) performed the worst among the three models. These results suggest that a larger model size does not necessarily promise a suc- cessful training. The lightest model (xls-r-300m) was not only effective in its performance but also advantageous in the training time and the total train- ing steps because of its increased batch size as Figure 4: The curves of the CERs on the validation data during the training. illustrated in Figure 4. In the experiments, the train- ing withxls-r-300mcompleted in approximately 10 hours, while the other 1B models took more than 50 hours to complete. In addition, the lightweight model takes up to approximately 1.2GB as opposed to the 1B modelsâ 3.8GB, and is beneficial for run- ning locally in the integrated annotation tool. 5. Is ASR-assisted transcription helpful? Whether ASR can truly benefit annotators in the transcription process has been a point of debate. Although many studies have argued for the poten- tial advantages of ASR in language documenta- tion, Prudâhommeaux et al. (2021) reported that members of some speaker communities preferred unassisted transcription without ASR support. To examine this question in the case of Ikema, we con- ducted a human evaluation to assess the extent to which ASR can actually accelerate the transcription process. 5.1. Setup For this experiment,the fine-tuned wav2vec2-xlr-s-300mmodel trained on the kana-script is integrated into ELAN, linguistic annotation software (Wittenburg et al., 2006), as an extension to transcribe Ikema automatically on the user interface. In this experiment, we choose an unannotated 5-minute recording, segment the spoken parts, and asked two annotators to annotate the same segments on ELAN under the following conditions. The total number of segments is 121, and their total duration is 217.64 seconds. Annotator A is a non-native expert who has worked on the documentation of Ikema for 19 years, and Annotator B is a non-native learner with less than one year of experience. Both are native speakers of Japanese and familiar with the kana orthography of Ikema. Annotator A is given an EAF file (an extended XML for ELAN containing the annotation text) in which the first half of the segments (61 w/o ASR w/ ASR Speedup Annotator A11:369:43 +19.38% Annotator B30:4925:00 +23.27% Table 3: A comparison of annotation speed, with and without ASR output. The speedup is computed by dividing the time without ASR by the time with ASR. samples, 110.12 seconds) are blank and the second half (60 samples, 107.52 second) are filled in with the modelâs prediction output. Conversely, Annotator B is provided with an EAF file where the first half contains the modelâs transcription and the second half is left blank. Both annotators are asked to transcribe the audio manually, and their annotation times are recorded separately for each half. Note that they are both allowed to have listened to the recording before and know what is talked about. 5.2. Results As Table 3 demonstrates, both annotators reported a positive speedup in ASR-assisted transcription. The annotators also noted that they felt it some- what easier to annotate the data with the ASR draft despite its errors. Annotator B noted that segments with a frequently occurring filler and a fixed expres- sion were often transcribed correctly, decreasing the need of manual correction with clicking, typing, and re-listening. The results confirmed the positive effect of ASR outputs in language documentation even when the modelâs accuracy is not ideal. As for the inter-annotator agreement between the two annotators, there was a CER of 17.32% and a WER of 46.24%. The CER of the ASR output was 29.61% with Annotator A, and 25.25% with Annotator B. However, when minor orthographic ambiguities are ignored, the CER of the ASR output decreases to 20.34%. The high error rates reflect the non-standardized aspects of the languageâs writing system, such as post-lexical morphophono- logical changes (Takubo, 2021b) and the position of whitespaces. Example (2) shows an instance of such cases. 3 The words umatti âcoming hereâ and Nmeera âummâ (a filler) are often pronounced as if they are single words. For this reason, An- notator A (2a) transcribed them as single words, while Annotator B (2b) inserted spaces based on their underlying analyzed forms. In addition, while Annotator A often omitted disfluencies and repeti- tions, Annotator B transcribed the speech faithfully including the speech errors. For reference, (2c) shows the actual output by 3 cond: conditional conjunctive suffix; dsc: discourse sentence-final particle; inf: infinitive form. the recognizer. While the high error rates may sug- gest low quality, a closer look at the output reveals that the modelâs prediction reflects how the sample sounds. In fact, the suffix -tigaa often realizes as phonetically shortened forms -tiyaa, -tiya, or -taa, the filler Nme is also pronounced as Nmya, and the sentence-final particle ira can also realize as ra. This phonetic faithfulness is characteristic of the CTC decoder outputting frame by frame without a language model. While the model still mispre- dicts token boundaries, the errors observed in the outputs are often in fact free variations. This obser- vation is coherent with the discussion provided in the previous study on ASR for Japhug (Guillaume et al., 2022). (2)a.ăăŸăŁăŠă ăăăȘăŒăăŠăăăŒ ăăăŒă umatti uma+tti here+come.inf ugunaaitigaa ugunaai-tigaa gather-cond Nmeera Nme+ira well+dsc âIf (they) come here and get together, wellâ b.ăăŸ ăŁăŠă ăăăȘăŒăăŠăăăŒ ăă ăă uma uma here tti tti come.inf ugunaaitigaa ugunaai-tigaa gather-cond Nme Nme well ira ira dsc c.ăăŸ ăŠăăăăȘăăăŒăăżăă uma tiugunaitaaNmyara 6. Conclusion This study presented an ongoing effort to develop an automatic speech recognition (ASR) system for Ikema Miyakoan, an endangered Ryukyuan lan- guage spoken in Okinawa, Japan. We constructed a 6.33-hour speech corpus from recordings through collaborative fieldwork with the speaker community. Based on this dataset, we trained an ASR model achieving a CER as low as 14.8%. Then, we inte- grated the model into ELAN as an ASR extension and evaluated its potential to mitigate the transcrip- tion bottleneck in language documentation. The results indicate that the integration of ASR into tran- scription workflows can substantially accelerate an- notation and reduce the cognitive burden on human transcribers. 7. Ethics statements This research is grounded in close collaboration with the Ikema-speaking community and adheres to ethical standards for language documentation and computational research on endangered languages. All speakers participated voluntarily with informed consent, and their privacy and data rights were re- spected throughout data collection and processing. The speech recordings are used strictly for research and documentation purposes, and any future public release of the dataset will follow community consul- tation and appropriate licensing to ensure respon- sible use. Potentially personally-identifiable names that appear in our datasets are annotated with the tags<name>...</name>so that such samples can be removed automatically. We recognize that the deployment of language technologies in small and vulnerable communi- ties entails complex ethical and social implications. While ASR systems can effectively accelerate tran- scription and support revitalization efforts, they may also raise concerns regarding data ownership, rep- resentation, and the potential misuse of linguistic data. To mitigate these risks, we emphasize trans- parent documentation of data ownership and licens- ing, in close collaboration with the community and through shared decision-making. Another relevant important consideration is that developing an ASR system for a traditionally unwrit- ten language may privilege a particular writing sys- tem that is not necessarily widely supported within the community. While the constructed dataset uses a kana (hiragana) orthography, some speakers in fact find this orthography difficult to read. This study, as well as the dataset and model developed in this work, does not intend to impose the use of the kana orthography on the community. More broadly, this work contributes to the eq- uitable development of language technologies by extending computational research beyond high- resource languages. By developing resources and methodologies for Ikema, we aim to promote lin- guistic diversity and inclusivity within NLP, demon- strating that advances in speech technology can serve both scientific and community-centered goals when guided by ethical collaboration and respect for speaker agency. 8. Acknowledgments The material was based on work supported in part by the US National Science Foundation un- der Grant Number BCS-2109709 and IIS-2137396. We also thank the Ikema Miyakoan native speakers who helped the authors collect the data and review the model output. 9. Appendix Table 4 lists the field recordings used in the Field dataset, as well as their basic statistics. 10. Bibliographical References Recording IDStyleDuration #Samples #Kana #Words I0482_usISpontaneous 111.0369 1105202 I0482_ngiSpontaneous 100.3672885201 I0482_pic_mucIusaSpontaneous85.9356683132 I0482_pic_buugiiSpontaneous72.9149534117 I0482_byuuigassaSpontaneous53.513545498 I0482_bippiiSpontaneous26.271522943 I0482_pic_barazanSpontaneous30.501632458 I0482_niguuSpontaneous86.8156675148 I0482_pic_takannaSpontaneous44.993441283 I0482_pic_takaSpontaneous 179.77108 1698345 I0412_suuniSpontaneous29.762124648 I0483_0414_satatinpuraSpontaneous44.302938277 I0482_pic_zzakugiiSpontaneous73.3748586122 I0482_gisIcISpontaneous 107.8772934193 I0482_kamiSpontaneous 197.85138 1521332 I0482_gazIhanagiiSpontaneous47.872942686 I0482_uganSpontaneous84.4761677129 I0483_0277_yasaiitameSpontaneous19.331611427 I0483_0276_waaSpontaneous16.761014330 I0483_0405_isIusISpontaneous 114.5177942203 I0483_0527_subaSpontaneous37.642234569 I0483_0412_waanimunSpontaneous34.992628361 I0483_0688_mamisuimai Spontaneous63.7940540110 I0483_0409_fukyagiSpontaneous66.8845568111 I0483_dakyauSpontaneous19.401512727 I0482_pc_taummabasISpontaneous 324.87185 3049586 I0483_0678_sagunaSpontaneous 118.4070 1093220 I0482_kkucIgiiSpontaneous65.8646534111 I0483_0222_avvansuSpontaneous23.821918039 I0488_ăăŒăăŒSpontaneous 100.7669832180 I0490_ăăŸăSpontaneous 111.1179958214 I0490_æ± é性æ©Spontaneous 167.68113 1485306 I0496_kaichosyokurekiConversation 912.26339 84071706 I0501_nakasonetuimyaaSpontaneous 305.9886 2495461 I0502_harimizuutakiSpontaneous 405.51151 3419660 I0503_NevskynohiSpontaneous 145.1660 1054204 I0506_aisatsutehon_VaSpontaneous19.1878519 I0506_aisatsutehon_VbSpontaneous13.6066219 I0506_aisatsutehon_VcSpontaneous15.88710921 I0506_aisatsutehon_VdSpontaneous 101.1525639131 I0506_aisatsutehon_VeSpontaneous 159.4947 1260247 I0506_aisatsutehon_VfSpontaneous79.7517612122 I0507_mingukaisetsu_Va Spontaneous 112.8167978215 I0507_mingukaisetsu_Vb Spontaneous 126.1839942191 I0508_nakamautaki_VSpontaneous 367.59102 2716507 I0509_zyaagama_VSpontaneous 254.0076 1918347 I0510_kyuukoominkan_Vb Spontaneous 116.3622787152 I0510_kyuukoominkan_Va Spontaneous 182.8539 1387263 Table 4: The detail of the spontaneous speech dataset (Field). The Audiobook data and the Dictionary data are not shown in this table. Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli. 2021. XLS-R: Self-supervised cross-lingual speech representation learning at scale. Alexei Baevski, Henry Zhou, Abdelrahman Mo- hamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. Blaine Billings and Bradley McDonnell. 2025. Con- necting automated speech recognition to tran- scription practices. In Proceedings of the Eighth Workshop on the Use of Computational Methods in the Study of Endangered Languages (Com- putEL), pages 128â132. Ilaria Chizzoni and Alessandro Vietti. 2024. To- wards an ASR system for documenting endan- gered languages: A preliminary study on Sar- dinian. In Proceedings of the 10th Italian Con- ference on Computational Linguistics (CLiC-it 2024), pages 214â220. Alex Graves, Santiago FernĂĄndez, Faustino Gomez, and JĂŒrgen Schmidhuber. 2006. Con- nectionist temporal classification: labelling un- segmented sequence data with recurrent neural networks. In Proceedings of the 23rd Interna- tional Conference on Machine Learning, pages 369â376. SĂ©verine Guillaume, Guillaume Wisniewski, CĂ©cile Macaire, Guillaume Jacques, Alexis Michaud, Benjamin Galliot, Maximin Coavoux, Solange Rossato, Minh-ChĂąu NguyĂȘn, and Maxime Fily. 2022. Fine-tuning pre-trained models for auto- matic speech recognition, experiments on a field- work corpus of Japhug (Trans-Himalayan family). In Proceedings of the Fifth Workshop on the Use of Computational Methods in the Study of En- dangered Languages, pages 170â178. Yuka Hayashi. 2013. Minami ryuukyuu miyakogo ikema hougen no bunpou. Ph.D. thesis, Kyoto University. Robbie Jimerson and Emily Prudâhommeaux. 2018. ASR for documenting acutely under-resourced indigenous languages. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018). Zhaolin Li, Monika Rind-Pawlowski, and Jan Niehues. 2024. Speech recognition corpus of the khinalug language for documenting endan- gered languages. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 15171â15180. Siyu Liang and Gina-Anne Levow. 2025. Breaking the transcription bottleneck: Fine-tuning ASR models for extremely low-resource fieldwork lan- guages. In Proceedings of the Fourth Workshop on NLP Applications to Field Linguistics, pages 26â37. Julia Mainzinger and Gina-Anne Levow. 2024. Fine- tuning ASR models for very low-resource lan- guages: A study on Mvskoke. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pages 76â82. Alexis Michaud, Oliver Adams, Trevor Cohn, Gra- ham Neubig, and SĂ©verine Guillaume. 2018. Inte- grating automatic transcription into the language documentation workflow: Experiments with Na data and the Persephone toolkit. Language Doc- umentation & Conservation, 12:393â429. Hiroyuki Nakama, Yukinori Takubo, Shoichi Iwasaki, Yosuke Igarashi, Daniel Wymark, and Natsuko Nakagawa. 2025. Dictionary of Ikema spoken in Nishihara village, a variety of Miyako Ryukyuan, 3rd edition. National Institute for Japanese Lan- guage and Linguistics. Alexander OâNeill, Marieke Meelen, Rolando Coto- Solano, Sonam Phuntsog, and Charles Ramble. 2023. Language preservation through ASR. In Cambridge Language Sciences Annual Sympo- sium. Thomas Pellard. 2015. The linguistic archeology of the Ryukyu Islands. In Patrick Heinrich, Shinsho Miyara, and Michinori Shimoji, editors, Hand- book of the Ryukyuan languages: History, struc- ture, and use, pages 13â37. De Gruyter Mouton, Berlin; Boston. Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, Alexei Baevski, Yossi Adi, Xiao- hui Zhang, Wei-Ning Hsu, Alexis Conneau, and Michael Auli. 2023. Scaling speech technology to 1,000+ languages. Emily Prudâhommeaux, Robbie Jimerson, Richard Hatcher, and Karin Michelson. 2021. Automatic speech recognition for supporting endangered language documentation. Language Documen- tation & Conservation, 15:491â513. Alec Radford, Jong Wook Kim, Tao Xu, Greg Brock- man, Christine McLeavey, and Ilya Sutskever. 2022. Robust speech recognition via large-scale weak supervision. Frank Seifart, Nicholas Evans, Harald Ham- marström, and Stephen C. Levinson. 2018. Lan- guage documentation twenty-five years on. Lan- guage, 94(4):e324âe345. Michinori Shimoji. 2008. A Grammar of Irabu, a Southern Ryukyuan Language. Ph.D. thesis, Australian National University. Chihiro Taguchi and David Chiang. 2024. Lan- guage complexity and speech recognition accu- racy: Orthographic complexity hurts, phonolog- ical complexity doesnât. In Proceedings of the 62nd Annual Meeting of the Association for Com- putational Linguistics (Volume 1: Long Papers), pages 15493â15503. Chihiro Taguchi, Jefferson Saransig, Dayana VelĂĄsquez, and David Chiang. 2024. Killkan: The automatic speech recognition dataset for kichwa with morphosyntactic information. In Proceed- ings of the 2024 Joint International Conference on Computational Linguistics, Language Re- sources and Evaluation (LREC-COLING 2024), pages 9753â9763. Yukinori Takubo. 2021a. Hougen wo kana de kaku: Ryuukyuu miyakogo ikemahougen wo rei ni. Ko- toba to Mozi. Yukinori Takubo. 2021b. Morphophonemics of Ikema Miyakoan. In Studies in Asian Histori- cal Linguistics, Philology and Beyond, chapter 5. Brill, Leiden, The Netherlands. Peter Wittenburg, Hennie Brugman, Albert Russel, Alex Klassmann, and Han Sloetjes. 2006. ELAN: a professional framework for multimodality re- search. In Proceedings of the Fifth International Conference on Language Resources and Evalu- ation (LREC). Masahiro Yamada, Yukinori Takubo, Shoichi Iwasaki, Celik Kenan, Soichiro Harada, Nobuko Kibe, Tyler Lau, Natsuko Nakagawa, Yuto Ni- inaga, Tomoyo Otsuki, Manami Sato, Rihito Shi- rata, Gijs van der Lubbe, and Akiko Yokoyama. 2020. Experimental study of inter-language and inter-generational intelligibility: Methodology and case studies of Ryukyuan languages. In Shoichi Iwasaki, Susan Strauss, Shin Fukuda, Sun-Ah Jun, Sung-Ock Sohn, and Kie Zuraw, editors, Japanese/Korean Linguistics, Vol. 26. CSLI Pub- lications, Stanford, CA.