ISSN: 2534-5192 (electronic) – 2681-8566 (print)| ◀ Yannis Haralambous | ▲ Proceedings | Saiya L. Karamali ▶ |
![]() ISBN: 978-2-487055-10-0 e-ISBN: 978-2-487055-11-7 ![]() | Audio LLM subtokens as encapsulated knowledge: the case of Persian subtoken graphemic representations in Whisper Behnoosh Namdarzadeh, Nicolas Ballier Abstract. Whisper is a widely used, open-access Large Language Model (LLM) trained using a multilingual paradigm. As such, it represents an important opportunity for researchers to study how multilingual LLMs function across languages. In this paper, we probe the grapholinguistic representations of Whisper for an under-resourced language with a Perso-Arabic script: Persian. DOI: https://doi.org/10.36824/2024-graf-namd
Ahmadi, Sina. 2020. “A Tokenization System for the Kurdish Language.” In Proceedings of the 7th Workshop on NLP for Similar Languages, Varieties and Dialects, 114–27.
Arab-Moghaddam, Narges, and Monique Sénéchal. 2001. “Orthographic and Phonological Processing Skills in Reading and Spelling in Persian/English Bilinguals.” International Journal of Behavioral Development 25:140–47. https://doi.org/10.1080/01650250042000320.
Ardila, R., M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber. 2020. “Common Voice: A Massively-Multilingual Speech Corpus.” In Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), 4211–15.
Asebriy, Zahra, Said Raghay, Omar Bencharef, and Younes Chihab. 2014. “Comparative Systems of Handwriting Arabic Character Recognition.” In 2014 Second World Conference on Complex Systems (WCCS), 90–93. https://doi.org/10.1109/ICoCS.2014.7060923.
Ballier, Nicolas, Léa Burin, Behnoosh Namdarzadeh, Sara B Ng, Richard Wright, and Jean-Baptiste Yunès. 2024. “Probing Whisper Predictions for French, English and Persian Transcriptions.” In Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024), edited by Mourad Abbas and Abed Alhakim Freihat, 129–38. Trento: Association for Computational Linguistics. https://aclanthology.org/2024.icnlsp-1.15.
Ballier, Nicolas, Adrien Meli, Maelle Amand, and Jean-Baptiste Yunès. 2023. “Using Whisper LLM for Automatic Phonetic Diagnosis of L2 Speech, a Case Study with French Learners of English.” In Proceedings of the 6th International Conference on Natural Language and Speech Processing (ICNLSP 2023), edited by Mourad Abbas and Abed Alhakim Freihat, 282–92. Online: Association for Computational Linguistics. https://aclanthology.org/2023.icnlsp-1.30.
Baluch, Bahman, and S. Choudhury. 2011. “Spelling Transparency and Its Impact on Short-Term Memory: Evidence from Persian and English.” Spelling Skills: Acquisition, Abilities, and Reading Connection, 93–106.
Bostrom, Kaj, and Greg Durrett. 2020. “Byte Pair Encoding Is Suboptimal for Language Model Pretraining.” In Findings of the Association for Computational Linguistics: EMNLP 2020, 4617–24.
Conneau, Alexis, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna. 2022. “FLEURS: Few-Shot Learning Evaluation of Universal Representations of Speech.”
———. 2023. “Fleurs: Few-Shot Learning Evaluation of Universal Representations of Speech.” In 2022 IEEE Spoken Language Technology Workshop (SLT), 798–805. IEEE.
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. “BERT: Pre-training of deep bidirectional transformers for language understanding.” In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171–86.
Downey, C.m., Terra Blevins, Nora Goldfine, and Shane Steinert-Threlkeld. 2023. “Embedding Structure Matters: Comparing Methods to Adapt Multilingual Vocabularies to New Languages.” In Proceedings of the 3rd Workshop on Multi-Lingual Representation Learning (MRL), edited by Duygu Ataman, 268–81. Singapore: Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.mrl-1.20.
Freihat, Abed Alhakim, and Mourad Abbas, eds. 2021. Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) Co-Located with ICNLSP 2021. Trento, Italy: Association for Computational Linguistics. https://aclanthology.org/2021.nsurl-1.0.
Gerganov, Georgi. 2003. “Whisper.cpp : A High-Performance Inference of OpenAI’s Whisper Automatic Speech Recognition (ASR) Model.” https://github.com/ggerganov/whisper.cpp.
Graham, Calbert, and Nathan Roll. 2024b. “Evaluating OpenAI’s Whisper ASR: Performance Analysis Across Diverse Accents and Speaker Traits.” JASA Express Letters 4 (2): 025206.
———. 2024a. “Evaluating OpenAI’s Whisper ASR: Performance analysis across diverse accents and speaker traits.” JASA Express Letters 4 (2): 025206. https://doi.org/10.1121/10.0024876.
Guerreiro, Nuno M., Duarte M. Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, and André F. T. Martins. 2023. “Hallucinations in Large Multilingual Translation Models.” Transactions of the Association for Computational Linguistics 11:1500–1517. https://doi.org/10.1162/tacl_a_00615.
Halbout, Dominique, and Muhammad-hossein Karimi. 2012. Le Persan. Assimil.
Han, Yujin, and Difan Zou. 2024. “Understanding and Mitigating Tokenization Bias in Language Models.” In Proceedings of the 41st International Conference on Machine Learning, edited by Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, 235:17480–504. Proceedings of Machine Learning Research. PMLR. https://proceedings.mlr.press/v235/han24g.html.
Hosseini, Fatemeh Sadat, Shima Kashef, Elham Shabaninia, and Hossein Nezamabadi-pour. 2021. “IDPL-PFOD: An Image Dataset of Printed Farsi Text for OCR Research.” In Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) Co-Located with ICNLSP 2021, edited by Abed Alhakim Freihat and Mourad Abbas, 22–31. Trento, Italy: Association for Computational Linguistics. https://aclanthology.org/2021.nsurl-1.4.
Jafari, Mohammad Mahdi, Somayyeh Behmanesh, Alireza Talebpour, and Ali Nadian Ghomsheh. 2021. “Improving Pre-Trained Language Model for Relation Extraction Using Syntactic Information in Persian.” In Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) Co-Located with ICNLSP 2021, edited by Abed Alhakim Freihat and Mourad Abbas, 38–44. Trento, Italy: Association for Computational Linguistics. https://aclanthology.org/2021.nsurl-1.6.
Karimi, Sarvnaz, Falk Scholer, and Andrew Turpin. 2007. “Collapsed Consonant and Vowel Models: New Approaches for English-Persian Transliteration and Back-Transliteration.” In Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics, edited by Annie Zaenen and Antal van den Bosch, 648–55. Prague, Czech Republic: Association for Computational Linguistics. https://aclanthology.org/P07-1082.
Karimi, Simin. 2005. A Minimalist Approach to Scrambling. Evidence from Persian. Berlin, New York: De Gruyter Mouton. https://doi.org/doi:10.1515/9783110199796.
Khalilia, Hadi, Abed Alhakim Freihat, and Fausto Giunchiglia. 2021. “The Dimensions of Lexical Semantic Resource Quality.” In Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) Co-Located with ICNLSP 2021, edited by Abed Alhakim Freihat and Mourad Abbas, 15–21. Trento, Italy: Association for Computational Linguistics. https://aclanthology.org/2021.nsurl-1.3.
Kudo, T. 2018. “Sentencepiece: A Simple and Language Independent Subword Tokenizer and Detokenizer for Neural Text Processing.”
Kummervold, Per E, Javier de la Rosa, Freddy Wetjen, Rolv-Arild Braaten, and Per Erik Solberg. 2024. “Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges.”
Mahootian, Shahrzad, and Lewis Gebhardt. 1997. Persian. Descriptive Grammars. London: Routledge.
Maleki, Jalal, and Lars Ahrenberg. 2008a. “Converting Romanized Persian to the Arabic Writing Systems.” In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08), edited by Nicoletta Calzolari, Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, and Daniel Tapias. Marrakech, Morocco: European Language Resources Association (ELRA). https://aclanthology.org/L08-1315/.
———. 2008b. “Converting Romanized Persian to the Arabic Writing Systems.” In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08). Marrakech, Morocco: European Language Resources Association (ELRA). http://www.lrec-conf.org/proceedings/lrec2008/pdf/738_paper.pdf.
Manova, Stela. 2023. “ChatGPT, n-Grams and the Power of Subword Units: The Future of Research in Morphology.” In Fourth International Workshop on Resources and Tools for Derivational Morphology, 5:1.
Manova, Stela, Harald Hammarström, Itamar Kastner, and Yining Nie. 2020. “What Is in a Morpheme? Theoretical, Experimental and Computational Approaches to the Relation of Meaning and Form in Morphology.” Word Structure 13 (1): 1–21.
Marszałek-Kowalewska, Katarzyna, ed. 2023. “Frontmatter.” In Persian Computational Linguistics and NLP, I–IV. Berlin, Boston: De Gruyter Mouton. https://doi.org/doi:10.1515/9783110619225-fm.
Mohebbi, Hosein, Grzegorz Chrupała, Willem Zuidema, and Afra Alishahi. 2023. “Homophone Disambiguation Reveals Patterns of Context Mixing in Speech Transformers.” In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, edited by Houda Bouamor, Juan Pino, and Kalika Bali, 8249–60. Singapore: Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.513.
Morris, Andrew Cameron, Viktoria Maier, and Phil D Green. 2004. “From WER and RIL to MER and WIL: improved evaluation measures for connected speech recognition.” In Interspeech, 2765–68.
Namdarzadeh, Behnoosh, Sadaf Mohseni, Lichao Zhu, Guillaume Wisniewski, and Ballier Nicolas. 2023. “Fine-Tuning MBART-50 with French and Farsi Data to Improve the Translation of Farsi Dislocations into English and French.” In Proceedings of Machine Translation Summit XIX, Vol. 2: Users Track, edited by Masaru Yamada and Felix do Carmo, 152–61. Macau SAR, China: Asia-Pacific Association for Machine Translation. https://aclanthology.org/2023.mtsummit-users.14/.
Oji, Romina, Nasrin Taghizadeh, and Heshaam Faili. 2021. “PerSpellData: An Exhaustive Parallel Spell Dataset for Persian.” In Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) Co-Located with ICNLSP 2021, edited by Abed Alhakim Freihat and Mourad Abbas, 8–14. Trento, Italy: Association for Computational Linguistics. https://aclanthology.org/2021.nsurl-1.2.
Patman, Chloe, and Eleanor Chodroff. 2024. “Speech Recognition in Adverse Conditions by Humans and Machines.” JASA Express Letters 4 (11).
Petrov, Aleksandar, Emanuele La Malfa, Philip Torr, and Adel Bibi. 2024. “Language Model Tokenizers Introduce Unfairness Between Languages.” Advances in Neural Information Processing Systems 36.
Radford, Alec, Jong Wook Kim, Tao Xu, Greg Brockman, Christine Mcleavey, and Ilya Sutskever. 2023. “Robust Speech Recognition via Large-Scale Weak Supervision.” In Proceedings of the 40th International Conference on Machine Learning, edited by Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, 202:28492–518. Proceedings of Machine Learning Research. PMLR. https://proceedings.mlr.press/v202/radford23a.html.
Rahbari, Noriyeh, and Monique Sénéchal. 2009. “Lexical and Nonlexical Processes in the Skilled Reading and Spelling of Persian.” Reading and Writing 22 (5): 511–30. https://doi.org/10.1007/s11145-008-9122-1.
Rastorgueva, Vera Sergeevna, and Steven P. Hill. 1964. A Short Sketch of the Grammar of Persian. Indiana University Research Center in Anthropology, Folklore and Linguistics Publication. Bloomington: Indiana University.
Salehi, Ali, and Cassandra L. Jacobs. 2024. “The Effect of Model Capacity and Script Diversity on Subword Tokenization for Sorani Kurdish.” In Proceedings of the 21st SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology, edited by Garrett Nicolai, Eleanor Chodroff, Frederic Mailhot, and Çağrı Çöltekin, 51–56. Mexico City, Mexico: Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.sigmorphon-1.6.
Sanabria, Ramon, Nikolay Bogoychev, Nina Markl, Andrea Carmantini, Ondrej Klejch, and Peter Bell. 2023. “The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR.” In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5. IEEE.
Sartakhti, Moein Salimi, Romina Etezadi, and Mehrnoush Shamsfard. 2021. “Improving Persian Relation Extraction Models by Data Augmentation.” In Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) Co-Located with ICNLSP 2021, edited by Abed Alhakim Freihat and Mourad Abbas, 32–37. Trento, Italy: Association for Computational Linguistics. https://aclanthology.org/2021.nsurl-1.5.
Sennrich, Rico, Barry Haddow, and Alexandra Birch. 2015. “Neural Machine Translation of Rare Words with Subword Units.” CoRR abs/1508.07909. http://arxiv.org/abs/1508.07909.
———. 2016. “Neural Machine Translation of Rare Words with Subword Units.” In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), edited by Katrin Erk and Noah A. Smith, 1715–25. Berlin, Germany: Association for Computational Linguistics. https://doi.org/10.18653/v1/P16-1162.
Shakeri, N., Z. Soleymani, T. Zarifian, and M. Kamali. 2015. “Investigating Phonological Awareness in Persian-Speaking Children With Phonological Disorders.” Middle East Journal of Rehabilitation and Health 2 (4): e32200. https://doi.org/10.17795/mejrh-32200.
Song, Xinying, Alex Salcianu, Yang Song, Dave Dopson, and Denny Zhou. 2021. “Fast Wordpiece Tokenization.” In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2089–2103.
Sproat, Richard, and Alexander Gutkin. 2021. “The Taxonomy of Writing Systems: How to Measure How Logographic a System Is.” Computational Linguistics 47 (3): 477–528.
Taghizadeh, Nasrin, Ali Ebrahimi, and Heshaam Faili. 2021. “NSURL-2021 Shared Task 1: Semantic Relation Extraction in Persian.” In Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) Co-Located with ICNLSP 2021, edited by Abed Alhakim Freihat and Mourad Abbas, 1–7. Trento, Italy: Association for Computational Linguistics. https://aclanthology.org/2021.nsurl-1.1.
Taguchi, Chihiro, and David Chiang. 2024. “Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn’t.”
Wisniewski, Guillaume, Lichao Zhu, Nicolas Ballier, and François Yvon. 2021. “Screening Gender Transfer in Neural Machine Translation.” In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, 311–21.
Yousef, Saeed, and Hayedeh Torabi. 2018. Persian. Routledge Comprehensive Grammars. Abingdon, Oxon: Routledge.
@INPROCEEDINGS{ahmadi2020tokenization,
AUTHOR = {Ahmadi, Sina},
TITLE = {A tokenization system for the {K}urdish language},
BOOKTITLE = {Proceedings of the 7th Workshop on NLP for Similar Languages, Varieties and Dialects},
YEAR = {2020},
PAGES = {114--127},
}
@ARTICLE{OrthoPersian2001,
AUTHOR = {Arab-Moghaddam, Narges and Sénéchal, Monique},
TITLE = {Orthographic and phonological processing skills in reading and spelling in {Persian/English} bilinguals},
JOURNAL = {International Journal of Behavioral Development},
YEAR = {2001},
VOLUME = {25},
PAGES = {140-147},
}
@INPROCEEDINGS{commonvoice:2020,
AUTHOR = {Ardila, R. and Branson, M. and Davis, K. and Henretty, M. and Kohler, M. and Meyer, J. and Morais, R. and Saunders, L. and Tyers, F. M. and Weber, G.},
TITLE = {Common Voice: A Massively-Multilingual Speech Corpus},
BOOKTITLE = {Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020)},
YEAR = {2020},
PAGES = {4211--4215},
}
@INPROCEEDINGS{handwritingArabiccharacterarticle,
AUTHOR = {Asebriy, Zahra and Raghay, Said and Bencharef, Omar and Chihab, Younes},
TITLE = {Comparative systems of handwriting Arabic character recognition},
BOOKTITLE = {2014 Second World Conference on Complex Systems (WCCS)},
YEAR = {2014},
VOLUME = {},
NUMBER = {},
PAGES = {90-93},
}
@INPROCEEDINGS{ballier-etal-2024-probing,
AUTHOR = {Ballier, Nicolas and Burin, L{\'e}a and Namdarzadeh, Behnoosh and Ng, Sara B and Wright, Richard and Yun{\`e}s, Jean-Baptiste},
EDITOR = {Abbas, Mourad and Freihat, Abed Alhakim},
TITLE = {Probing Whisper Predictions for {F}rench, {E}nglish and {P}ersian Transcriptions},
BOOKTITLE = {Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024)},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Trento},
YEAR = {2024},
PAGES = {129--138},
URL = {https://aclanthology.org/2024.icnlsp-1.15},
}
@INPROCEEDINGS{ballier-etal-2023-using,
AUTHOR = {Nicolas Ballier and Meli, Adrien and Amand, Maelle and Yun{\`e}s, Jean-Baptiste},
EDITOR = {Abbas, Mourad and Freihat, Abed Alhakim},
TITLE = {Using Whisper {LLM} for Automatic Phonetic Diagnosis of {L}2 Speech, a Case Study with {F}rench Learners of {E}nglish},
BOOKTITLE = {Proceedings of the 6th International Conference on Natural Language and Speech Processing (ICNLSP 2023)},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Online},
YEAR = {2023},
PAGES = {282--292},
URL = {https://aclanthology.org/2023.icnlsp-1.30},
}
@ARTICLE{Balucharticle,
AUTHOR = {Baluch, Bahman and Choudhury, S.},
TITLE = {Spelling transparency and its impact on short-term memory: Evidence from persian and english},
JOURNAL = {Spelling Skills: Acquisition, Abilities, and Reading Connection},
YEAR = {2011},
PAGES = {93-106},
}
@INPROCEEDINGS{bostrom2020byte,
AUTHOR = {Bostrom, Kaj and Durrett, Greg},
TITLE = {Byte Pair Encoding is Suboptimal for Language Model Pretraining},
BOOKTITLE = {Findings of the Association for Computational Linguistics: EMNLP 2020},
YEAR = {2020},
PAGES = {4617--4624},
}
@UNPUBLISHED{conneau2022fleursfewshotlearningevaluation,
AUTHOR = {Alexis Conneau and Min Ma and Simran Khanuja and Yu Zhang and Vera Axelrod and Siddharth Dalmia and Jason Riesa and Clara Rivera and Ankur Bapna},
TITLE = {FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech},
YEAR = {2022},
NOTE = {preprint arXiv:2205.12446},
}
@INPROCEEDINGS{conneau2023fleurs,
AUTHOR = {Conneau, Alexis and Ma, Min and Khanuja, Simran and Zhang, Yu and Axelrod, Vera and Dalmia, Siddharth and Riesa, Jason and Rivera, Clara and Bapna, Ankur},
TITLE = {Fleurs: Few-shot learning evaluation of universal representations of speech},
BOOKTITLE = {2022 IEEE Spoken Language Technology Workshop (SLT)},
YEAR = {2023},
PAGES = {798--805},
}
@INPROCEEDINGS{devlin2018bert,
AUTHOR = {Devlin, Jacob and Ming-Wei Chang and Kenton Lee and Kristina Toutanova},
TITLE = {{BERT: Pre-training of deep bidirectional transformers for language understanding}},
BOOKTITLE = {{Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies}},
YEAR = {2019},
PAGES = {4171--4186},
}
@INPROCEEDINGS{downey-etal-2023-embedding,
AUTHOR = {Downey, C.m. and Blevins, Terra and Goldfine, Nora and Steinert-Threlkeld, Shane},
EDITOR = {Ataman, Duygu},
TITLE = {Embedding Structure Matters: Comparing Methods to Adapt Multilingual Vocabularies to New Languages},
BOOKTITLE = {Proceedings of the 3rd Workshop on Multi-lingual Representation Learning (MRL)},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Singapore},
YEAR = {2023},
PAGES = {268--281},
}
@PROCEEDINGS{nsurl-2021-international-nlp,
EDITOR = {Freihat, Abed Alhakim and Abbas, Mourad},
TITLE = {Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) co-located with ICNLSP 2021},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Trento, Italy},
YEAR = {2021},
URL = {https://aclanthology.org/2021.nsurl-1.0},
}
@UNPUBLISHED{gerganov2023whispercpp,
AUTHOR = {Gerganov, Georgi},
TITLE = {whisper.cpp : A High-performance inference of {O}pen{AI}'s {W}hisper automatic speech recognition {(ASR)} model},
YEAR = {2003},
URL = {https://github.com/ggerganov/whisper.cpp},
}
@ARTICLE{graham2024evaluating,
AUTHOR = {Graham, Calbert and Roll, Nathan},
TITLE = {Evaluating OpenAI's Whisper ASR: Performance analysis across diverse accents and speaker traits},
JOURNAL = {JASA Express Letters},
PUBLISHER = {Acoustical Society of America},
YEAR = {2024},
VOLUME = {4},
NUMBER = {2},
PAGES = {025206},
}
@ARTICLE{GrahamWhisperforEnglish,
AUTHOR = {Graham, Calbert and Roll, Nathan},
TITLE = {{Evaluating OpenAI's Whisper ASR: Performance analysis across diverse accents and speaker traits}},
JOURNAL = {JASA Express Letters},
YEAR = {2024},
VOLUME = {4},
NUMBER = {2},
PAGES = {025206},
}
@ARTICLE{guerreiro-etal-2023-hallucinations,
AUTHOR = {Guerreiro, Nuno M. and Alves, Duarte M. and Waldendorf, Jonas and Haddow, Barry and Birch, Alexandra and Colombo, Pierre and Martins, Andr{\'e} F. T.},
TITLE = {Hallucinations in Large Multilingual Translation Models},
JOURNAL = {Transactions of the Association for Computational Linguistics},
PUBLISHER = {MIT Press},
ADDRESS = {Cambridge, MA},
YEAR = {2023},
VOLUME = {11},
PAGES = {1500--1517},
}
@BOOK{halbout2012persan,
AUTHOR = {Halbout, Dominique and Karimi, Muhammad-hossein},
TITLE = {Le persan},
PUBLISHER = {Assimil},
YEAR = {2012},
}
@INPROCEEDINGS{pmlr-v235-han24g,
AUTHOR = {Han, Yujin and Zou, Difan},
EDITOR = {Salakhutdinov, Ruslan and Kolter, Zico and Heller, Katherine and Weller, Adrian and Oliver, Nuria and Scarlett, Jonathan and Berkenkamp, Felix},
TITLE = {Understanding and Mitigating Tokenization Bias in Language Models},
BOOKTITLE = {Proceedings of the 41st International Conference on Machine Learning},
SERIES = {Proceedings of Machine Learning Research},
PUBLISHER = {PMLR},
YEAR = {2024},
VOLUME = {235},
PAGES = {17480--17504},
URL = {https://proceedings.mlr.press/v235/han24g.html},
}
@INPROCEEDINGS{of-printed-farsi-text-for-ocr-research-2021-fatemeh-sadat,
AUTHOR = {Hosseini, Fatemeh Sadat and Kashef, Shima and Shabaninia, Elham and Nezamabadi-pour, Hossein},
EDITOR = {Freihat, Abed Alhakim and Abbas, Mourad},
TITLE = {{IDPL}-{PFOD}: An Image Dataset of Printed {F}arsi Text for {OCR} Research},
BOOKTITLE = {Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) co-located with ICNLSP 2021},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Trento, Italy},
YEAR = {2021},
PAGES = {22--31},
URL = {https://aclanthology.org/2021.nsurl-1.4},
}
@INPROCEEDINGS{pre-trained-language-model-for-relation-extraction-using-syntactic-information-in-persian-2021-mohammad-mahdi,
AUTHOR = {Jafari, Mohammad Mahdi and Behmanesh, Somayyeh and Talebpour, Alireza and Ghomsheh, Ali Nadian},
EDITOR = {Freihat, Abed Alhakim and Abbas, Mourad},
TITLE = {Improving pre-trained Language Model for Relation Extraction Using Syntactic Information in {P}ersian},
BOOKTITLE = {Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) co-located with ICNLSP 2021},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Trento, Italy},
YEAR = {2021},
PAGES = {38--44},
URL = {https://aclanthology.org/2021.nsurl-1.6},
}
@INPROCEEDINGS{karimi-etal-2007-collapsed,
AUTHOR = {Karimi, Sarvnaz and Scholer, Falk and Turpin, Andrew},
EDITOR = {Zaenen, Annie and van den Bosch, Antal},
TITLE = {Collapsed Consonant and Vowel Models: New Approaches for {E}nglish-{P}ersian Transliteration and Back-Transliteration},
BOOKTITLE = {Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Prague, Czech Republic},
YEAR = {2007},
PAGES = {648--655},
URL = {https://aclanthology.org/P07-1082},
}
@BOOK{Karimi+2005,
AUTHOR = {Simin Karimi},
TITLE = {A Minimalist Approach to Scrambling. Evidence from Persian},
PUBLISHER = {De Gruyter Mouton},
ADDRESS = {Berlin, New York},
YEAR = {2005},
}
@INPROCEEDINGS{of-lexical-semantic-resource-quality-2021-hadi-khalilia,
AUTHOR = {Khalilia, Hadi and Freihat, Abed Alhakim and Giunchiglia, Fausto},
EDITOR = {Freihat, Abed Alhakim and Abbas, Mourad},
TITLE = {The Dimensions of Lexical Semantic Resource Quality},
BOOKTITLE = {Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) co-located with ICNLSP 2021},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Trento, Italy},
YEAR = {2021},
PAGES = {15--21},
URL = {https://aclanthology.org/2021.nsurl-1.3},
}
@UNPUBLISHED{kudo2018sentencepiece,
AUTHOR = {Kudo, T.},
TITLE = {Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing},
YEAR = {2018},
NOTE = {preprint arXiv:1808.06226},
}
@UNPUBLISHED{kummervold2024whispering,
AUTHOR = {Kummervold, Per E and de la Rosa, Javier and Wetjen, Freddy and Braaten, Rolv-Arild and Solberg, Per Erik},
TITLE = {Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges},
YEAR = {2024},
NOTE = {preprint arXiv:2402.01917},
}
@BOOK{PersianMahootian,
AUTHOR = {Mahootian, Shahrzad and Gebhardt, Lewis},
TITLE = {Persian},
SERIES = {Descriptive grammars},
PUBLISHER = {Routledge},
ADDRESS = {London},
YEAR = {1997},
NOTE = {Includes index},
}
@INPROCEEDINGS{inproceedingsJalalConvertingRomanizedPersiantoArabic,
AUTHOR = {Maleki, Jalal and Ahrenberg, Lars},
EDITOR = {Calzolari, Nicoletta and Choukri, Khalid and Maegaard, Bente and Mariani, Joseph and Odijk, Jan and Piperidis, Stelios and Tapias, Daniel},
TITLE = {Converting {R}omanized {P}ersian to the {A}rabic Writing Systems},
BOOKTITLE = {Proceedings of the Sixth International Conference on Language Resources and Evaluation ({LREC}'08)},
PUBLISHER = {European Language Resources Association (ELRA)},
ADDRESS = {Marrakech, Morocco},
YEAR = {2008},
URL = {https://aclanthology.org/L08-1315/},
}
@INPROCEEDINGS{maleki-ahrenberg-2008-converting,
AUTHOR = {Maleki, Jalal and Ahrenberg, Lars},
TITLE = {Converting {R}omanized {P}ersian to the {A}rabic Writing Systems},
BOOKTITLE = {Proceedings of the Sixth International Conference on Language Resources and Evaluation ({LREC}'08)},
PUBLISHER = {European Language Resources Association (ELRA)},
ADDRESS = {Marrakech, Morocco},
YEAR = {2008},
URL = {http://www.lrec-conf.org/proceedings/lrec2008/pdf/738_paper.pdf},
}
@INPROCEEDINGS{manova2023chatgpt,
AUTHOR = {Manova, Stela},
TITLE = {ChatGPT, \$n\$\yh-grams and the power of subword units: The future of research in morphology},
BOOKTITLE = {Fourth International Workshop on Resources and Tools for Derivational Morphology},
YEAR = {2023},
VOLUME = {5},
PAGES = {1},
}
@ARTICLE{manova2020morpheme,
AUTHOR = {Manova, Stela and Hammarstr{\"o}m, Harald and Kastner, Itamar and Nie, Yining},
TITLE = {What is in a morpheme? Theoretical, experimental and computational approaches to the relation of meaning and form in morphology},
JOURNAL = {Word Structure},
PUBLISHER = {Edinburgh University Press The Tun-Holyrood Road, 12 (2f) Jackson's Entry~…},
YEAR = {2020},
VOLUME = {13},
NUMBER = {1},
PAGES = {1--21},
}
@INBOOK{PersianNLP2023,
EDITOR = {Katarzyna Marszałek-Kowalewska},
TITLE = {Frontmatter},
BOOKTITLE = {Persian Computational Linguistics and NLP},
PUBLISHER = {De Gruyter Mouton},
ADDRESS = {Berlin, Boston},
YEAR = {2023},
PAGES = {I--IV},
}
@INPROCEEDINGS{mohebbi-etal-2023-homophone,
AUTHOR = {Mohebbi, Hosein and Chrupa{\l}a, Grzegorz and Zuidema, Willem and Alishahi, Afra},
EDITOR = {Bouamor, Houda and Pino, Juan and Bali, Kalika},
TITLE = {Homophone Disambiguation Reveals Patterns of Context Mixing in Speech Transformers},
BOOKTITLE = {Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Singapore},
YEAR = {2023},
PAGES = {8249--8260},
}
@INPROCEEDINGS{morris2004,
AUTHOR = {Morris, Andrew Cameron and Maier, Viktoria and Green, Phil D},
TITLE = {{From WER and RIL to MER and WIL: improved evaluation measures for connected speech recognition}},
BOOKTITLE = {Interspeech},
YEAR = {2004},
PAGES = {2765--2768},
}
@INPROCEEDINGS{namdarzadeh-etal-2023-fine,
AUTHOR = {Namdarzadeh, Behnoosh and Mohseni, Sadaf and Zhu, Lichao and Wisniewski, Guillaume and Nicolas, Ballier},
EDITOR = {Yamada, Masaru and do Carmo, Felix},
TITLE = {Fine-tuning {MBART}-50 with {F}rench and {F}arsi data to improve the translation of {F}arsi dislocations into {E}nglish and {F}rench},
BOOKTITLE = {Proceedings of Machine Translation Summit XIX, Vol. 2: Users Track},
PUBLISHER = {Asia-Pacific Association for Machine Translation},
ADDRESS = {Macau SAR, China},
YEAR = {2023},
PAGES = {152--161},
URL = {https://aclanthology.org/2023.mtsummit-users.14/},
}
@INPROCEEDINGS{persian-2021-romina-oji,
AUTHOR = {Oji, Romina and Taghizadeh, Nasrin and Faili, Heshaam},
EDITOR = {Freihat, Abed Alhakim and Abbas, Mourad},
TITLE = {PerSpellData: An Exhaustive Parallel Spell Dataset For {P}ersian},
BOOKTITLE = {Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) co-located with ICNLSP 2021},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Trento, Italy},
YEAR = {2021},
PAGES = {8--14},
URL = {https://aclanthology.org/2021.nsurl-1.2},
}
@ARTICLE{patman2024speech,
AUTHOR = {Patman, Chloe and Chodroff, Eleanor},
TITLE = {Speech recognition in adverse conditions by humans and machines},
JOURNAL = {JASA Express Letters},
PUBLISHER = {AIP Publishing},
YEAR = {2024},
VOLUME = {4},
NUMBER = {11},
}
@ARTICLE{petrov2024language,
AUTHOR = {Petrov, Aleksandar and La Malfa, Emanuele and Torr, Philip and Bibi, Adel},
TITLE = {Language model tokenizers introduce unfairness between languages},
JOURNAL = {Advances in Neural Information Processing Systems},
YEAR = {2024},
VOLUME = {36},
}
@INPROCEEDINGS{pmlr-v202-radford23a,
AUTHOR = {Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and Mcleavey, Christine and Sutskever, Ilya},
EDITOR = {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
TITLE = {Robust Speech Recognition via Large-Scale Weak Supervision},
BOOKTITLE = {Proceedings of the 40th International Conference on Machine Learning},
SERIES = {Proceedings of Machine Learning Research},
PUBLISHER = {PMLR},
YEAR = {2023},
VOLUME = {202},
PAGES = {28492--28518},
URL = {https://proceedings.mlr.press/v202/radford23a.html},
}
@ARTICLE{Rahbari2009,
AUTHOR = {Rahbari, Noriyeh and Sénéchal, Monique},
TITLE = {Lexical and nonlexical processes in the skilled reading and spelling of {P}ersian},
JOURNAL = {Reading and Writing},
YEAR = {2009},
VOLUME = {22},
NUMBER = {5},
PAGES = {511-530},
}
@BOOK{ShortSketchPersianGrammar,
AUTHOR = {Rastorgueva, Vera Sergeevna and Hill, Steven P.},
TITLE = {A Short sketch of the grammar of Persian},
SERIES = {Indiana University Research Center in Anthropology, Folklore and Linguistics Publication},
PUBLISHER = {Indiana University},
ADDRESS = {Bloomington},
YEAR = {1964},
NOTE = {"Also : Part II of the International Journal of Ammerican Linguistics"},
}
@INPROCEEDINGS{salehi-jacobs-2024-effect,
AUTHOR = {Salehi, Ali and Jacobs, Cassandra L.},
EDITOR = {Nicolai, Garrett and Chodroff, Eleanor and Mailhot, Frederic and {\c{C}}{\"o}ltekin, {\c{C}}a{\u{g}}r{\i}},
TITLE = {The Effect of Model Capacity and Script Diversity on Subword Tokenization for {S}orani {K}urdish},
BOOKTITLE = {Proceedings of the 21st SIGMORPHON workshop on Computational Research in Phonetics, Phonology, and Morphology},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Mexico City, Mexico},
YEAR = {2024},
PAGES = {51--56},
}
@INPROCEEDINGS{sanabria2023edinburgh,
AUTHOR = {Sanabria, Ramon and Bogoychev, Nikolay and Markl, Nina and Carmantini, Andrea and Klejch, Ondrej and Bell, Peter},
TITLE = {The {E}dinburgh international accents of {E}nglish corpus: Towards the democratization of {E}nglish {ASR}},
BOOKTITLE = {ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
YEAR = {2023},
PAGES = {1--5},
}
@INPROCEEDINGS{relation-extraction-models-by-data-augmentation-2021-moein-salimi,
AUTHOR = {Sartakhti, Moein Salimi and Etezadi, Romina and Shamsfard, Mehrnoush},
EDITOR = {Freihat, Abed Alhakim and Abbas, Mourad},
TITLE = {Improving {P}ersian Relation Extraction Models By Data Augmentation},
BOOKTITLE = {Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) co-located with ICNLSP 2021},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Trento, Italy},
YEAR = {2021},
PAGES = {32--37},
URL = {https://aclanthology.org/2021.nsurl-1.5},
}
@ARTICLE{DBLP:journals/corr/SennrichHB15,
AUTHOR = {Rico Sennrich and Barry Haddow and Alexandra Birch},
TITLE = {Neural Machine Translation of Rare Words with Subword Units},
JOURNAL = {CoRR},
YEAR = {2015},
VOLUME = {abs/1508.07909},
URL = {http://arxiv.org/abs/1508.07909},
}
@INPROCEEDINGS{sennrich-etal-2016-neural,
AUTHOR = {Sennrich, Rico and Haddow, Barry and Birch, Alexandra},
EDITOR = {Erk, Katrin and Smith, Noah A.},
TITLE = {Neural Machine Translation of Rare Words with Subword Units},
BOOKTITLE = {Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Berlin, Germany},
YEAR = {2016},
PAGES = {1715--1725},
}
@ARTICLE{zarifian2015,
AUTHOR = {Shakeri, N. and Soleymani, Z. and Zarifian, T. and Kamali, M.},
TITLE = {{Investigating Phonological Awareness in Persian-Speaking Children With Phonological Disorders}},
JOURNAL = {Middle East Journal of Rehabilitation and Health},
PUBLISHER = {},
YEAR = {2015},
VOLUME = {2},
NUMBER = {4},
PAGES = {e32200},
URL = {https://brieflands.com/articles/mejrh-21516},
}
@INPROCEEDINGS{song2020fast,
AUTHOR = {Song, Xinying and Salcianu, Alex and Song, Yang and Dopson, Dave and Zhou, Denny},
TITLE = {Fast wordpiece tokenization},
BOOKTITLE = {{Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing}},
YEAR = {2021},
PAGES = {2089--2103},
}
@ARTICLE{sproat2021taxonomy,
AUTHOR = {Sproat, Richard and Gutkin, Alexander},
TITLE = {The taxonomy of writing systems: How to measure how logographic a system is},
JOURNAL = {Computational Linguistics},
YEAR = {2021},
VOLUME = {47},
NUMBER = {3},
PAGES = {477--528},
}
@INPROCEEDINGS{nasrin2021,
AUTHOR = {Taghizadeh, Nasrin and Ebrahimi, Ali and Faili, Heshaam},
EDITOR = {Freihat, Abed Alhakim and Abbas, Mourad},
TITLE = {NSURL-2021 Shared Task 1: Semantic Relation Extraction in {P}ersian},
BOOKTITLE = {Proceedings of the Second International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2021) co-located with ICNLSP 2021},
PUBLISHER = {Association for Computational Linguistics},
ADDRESS = {Trento, Italy},
YEAR = {2021},
PAGES = {1--7},
URL = {https://aclanthology.org/2021.nsurl-1.1},
}
@UNPUBLISHED{taguchi2024language,
AUTHOR = {Taguchi, Chihiro and Chiang, David},
TITLE = {Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't},
YEAR = {2024},
NOTE = {preprint arXiv:2406.09202},
}
@INPROCEEDINGS{wisniewski2021screening,
AUTHOR = {Wisniewski, Guillaume and Zhu, Lichao and Ballier, Nicolas and Yvon, Fran{\c{c}}ois},
TITLE = {Screening Gender Transfer in Neural Machine Translation},
BOOKTITLE = {Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP},
YEAR = {2021},
PAGES = {311--321},
}
@BOOK{PersianSaeedYousef1097582,
AUTHOR = {Yousef, Saeed and Torabi, Hayedeh},
TITLE = {Persian},
SERIES = {Routledge comprehensive grammars},
PUBLISHER = {Routledge},
ADDRESS = {Abingdon, Oxon},
YEAR = {2018},
}
Behnoosh Namdarzadeh, Nicolas Ballier (2024), “Audio LLM subtokens as encapsulated knowledge: the case of Persian subtoken graphemic representations in Whisper,” in Proceedings of Grapholinguistics in the 21st Century, 2024 (Yannis Haralambous, Ed.), Grapholinguistics and Its Applications, Vol. 11, Brest: Fluxus Editions, 327–360.
@INPROCEEDINGS{gla11-namd,
AUTHOR = {Behnoosh Namdarzadeh and Nicolas Ballier},
EDITOR = {Haralambous, Yannis},
TITLE = {{Audio LLM subtokens as encapsulated knowledge: the case of Persian subtoken graphemic representations in Whisper}},
BOOKTITLE = {{Proceedings of Grapholinguistics in the 21st Century, 2024}},
SERIES = {{Grapholinguistics and Its Applications}},
VOLUME = {11},
PUBLISHER = {Fluxus Editions},
ADDRESS = {Brest},
YEAR = {2024},
PAGES = {327--360},
DOI = {https://doi.org/10.36824/2024-graf-namd},
}
|