Daniel Wu, Andrew Yang, Vinay Prabhu
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
We present Afro-MNIST, a set of synthetic MNIST-style datasets for four orthographies used in Afro-Asiatic and Niger-Congo languages: Ge'ez (Ethiopic), Vai, Osmanya, and N'Ko. These datasets serve as "drop-in" replacements for MNIST. We also describe and open-source a method for synthetic MNIST-style dataset generation from single examples of each digit. These datasets can be found at https://github.com/Daniel-Wu/AfroMNIST. We hope that MNIST-style datasets will be developed for other numeral systems, and that these datasets vitalize machine learning education in underrepresented nations in the research community. Figure 1: Afro-MNIST: Four synthetic MNIST-style datasets. 1 MOTIVATION Classifying MNIST Hindu-Arabic numerals (LeCun et al., 1998) has become the "Hello World!" challenge of the machine learning community. This task has excited a large number of prospective machine learning scientists and has led to practical advancements in optical character recognition. The Hindu-Arabic numeral system is the predominant numeral system used in the world today. However, there are a sizeable number of languages whose numeric glyphs are not inherited from the Hindu-Arabic numeral system. As it relates to human language, work in machine learning and artificial intelligence focuses almost exclusively on high-resource languages such as English and Mandarin Chinese. Unfortunately, these "mainstream" languages comprise only a tiny fraction of all extant languages. Indeed, of over 7,000 languages in the world, the vast majority are not represented in the machine learning research community. In particular, we note that there is a vast collection of alternative numeral systems for which an MNIST-style dataset is not available.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2009.13509")
get_code_for_paper("2009.13509")
have("2009.13509")
Connect an agent — have() is free.