Classifying cuneiform symbols using machine learning algorithms with unigram features on a balanced dataset

Problem: Recognizing written languages using symbols written in cuneiform is a tough endeavor due to the lack of information and the challenge of the process of tokenization. The Cuneiform Language Identification (CLI) dataset attempts to understand seven cuneiform languages and dialects, including Sumerian and six dialects of the Akkadian language: Old Babylonian, Middle Babylonian Peripheral, Standard Babylonian, Neo-Babylonian, Late Babylonian, and Neo-Assyrian. However, this dataset suffers from the problem of imbalanced categories. Aim: Therefore, this article aims to build a system capable of distinguishing between several cuneiform languages and solving the problem of unbalanced categories in the CLI dataset. Methods: Oversampling technique was used to balance the dataset, and the performance of machine learning algorithms such as Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Decision Tree (DT), Random Forest (RF), and deep learning such as deep neural networks (DNNs) using the unigram feature extraction method was investigated. Results: The proposed method using machine learning algorithms (SVM, KNN, DT, and RF) on a balanced dataset obtained an accuracy of 88.15, 88.14, 94.13, and 95.46%, respectively, while the DNN model got an accuracy of 93%. This proves improved performance compared to related works. Conclusion: This proves the improvement of classifiers when working on a balanced dataset. The use of unigram features also showed an improvement in the performance of the classifier as it reduced the size of the data and accelerated the processing process.

Location
Deutsche Nationalbibliothek Frankfurt am Main
Extent
Online-Ressource
Language
Englisch

Bibliographic citation
Classifying cuneiform symbols using machine learning algorithms with unigram features on a balanced dataset ; volume:32 ; number:1 ; year:2023 ; extent:11
Journal of intelligent systems ; 32, Heft 1 (2023) (gesamt 11)

Creator
Mahmood, Maha
Jasem, Farah Maath
Mukhlif, Abdulrahman Abbas
AL-Khateeb, Belal

DOI
10.1515/jisys-2023-0087
URN
urn:nbn:de:101:1-2023092514033130037118
Rights
Open Access; Der Zugriff auf das Objekt ist unbeschränkt möglich.
Last update
14.08.2025, 10:58 AM CEST

Data provider

This object is provided by:
Deutsche Nationalbibliothek. If you have any questions about the object, please contact the data provider.

Associated

  • Mahmood, Maha
  • Jasem, Farah Maath
  • Mukhlif, Abdulrahman Abbas
  • AL-Khateeb, Belal

Other Objects (12)