Skip to main navigation Skip to search Skip to main content

Quran Mining: Computational Stylometry of Quranic Texts

  • Mahmoud Shokrollahi Far

Research output: ThesisDoctoral Thesis

33 Downloads (Pure)

Abstract

The research reported in this thesis explores the application of computational methods, in particular methods that employ machine learning, in stylometric studies aiming to identify meaningful patterns and new insights in Quranic texts, i.e., texts in Classical Arabic from the Quran or from other writings inspired by the Quran in form and content. The primary objective is to statistically analyze variations in textual style, in particular across different authors and genres.

Drawing inspiration from the concept of the Human Stylome - a set of measurable linguistic traits proposed by Van Halteren et al. (2005), following the mapping of the Human Genome - this study investigates the reliability of Sequences of Morpho-Syntactic tags (SMS tags), i.e., sequences of codes that describe morphological and syntactic properties of grammatical units and relations between them, as a stylometric feature.

To facilitate this, the following tools were developed and implemented. First, we developed Regex Morphosyntax as a computational Quranic grammar, based on the use of regular expressions – i.e., sequences of characters that define a computationally effective search pattern for text. Second, we designed and implemented the MOBIN Parser-Tagger, which uses the Regex Morphosyntax as its knowledge base to generate SMS tags in Quranic corpora.

We investigated the computational reliability of SMS tags compared with other common feature types in stylometry in experiments concerning three tasks: (1) authorship attribution (AA), where the aim is to identify the author of a specific text; (2) authorship verification (AV), which is concerned with determining whether certain texts are written by the same author; and (3) source verification (SV), which aims at distinguishing texts by chronological, geographical, or other properties of their origin. In all experiments both large and small amounts of training data were employed; in the AA- and AV-experiments both long and (very) short Quranic texts were the object of classification.

In the AA experiments we found that SMS tags is a highly reliable stylome for robust classifications of short Quranic texts. The AV experiments confirm that SMS tags provide significant evidence for discussions regarding the authorship of the Quran. The results of the SV experiments demonstrated high reliability for verifying the geographical source (Mecca or Medina) of chapters and verses of the Quran.

From the results of these experiments we concluded that the sequences of morphosyntactic tags, generated by the MOBIN system, form a stylome with high computational reliability in representing the stylometric characteristics of short Quranic texts. We have consistently found evidence that grammatical information, especially in the form of SMS tags, is highly effective for stylometric classifications of Quranic texts of all sizes, great and small. These findings provide a foundation for future research to explore the implications of stylometric patters in Classical Arabic in general and in Quranic texts in particular.

In deze dissertatie wordt de toepassing onderzocht van computationele methoden, met name methoden die machine learning gebruiken, in stylometrische studies met als doel betekenisvolle patronen en nieuwe inzichten in Koranische teksten te identificeren, dat wil zeggen teksten in het Klassiek Arabisch uit de Koran of uit andere geschriften geïnspireerd door de Koran in vorm en inhoud. Het primaire doel is om variaties in tekstuele stijl statistisch te analyseren, met name tussen verschillende auteurs en genres.

Geïnspireerd door het concept van de Human Stylome wordt de betrouwbaarheid onderzocht van sequenties van Morpho-Syntactische tags (SMS-tags), oftewel sequenties van codes die morfologische en syntactische eigenschappen van grammaticale eenheden en de relaties daartussen beschrijven, als een stylometrisch kenmerk.

Om dit te vergemakkelijken werden de volgende tools ontwikkeld en geïmplementeerd. Ten eerste de Regex Morphosyntax als een computationele Koranische grammatica, gebaseerd op het gebruik van reguliere expressies – reeksen tekens die een computationeel effectief zoekpatroon voor tekst definiëren. Ten tweede werd de MOBIN Parser-Tagger ontworpen en geïmplementeerd, die de Regex Morphosyntax als kennisbank gebruikt om SMS-tags in Koranische corpora te genereren.

De computationele betrouwbaarheid van SMS-tags werd vergeleken met andere veel gebruikte kenmerken in de stylometrie, in experimenten met drie taken: (1) auteurschapstoekenning (AA), waarbij het doel is de auteur van een specifieke tekst te identificeren; (2) auteursverificatie (AV), die zich richt op het bepalen of zekere gegeven teksten door dezelfde auteur zijn geschreven; en (3) bronverificatie (SV), die gericht is op het onderscheiden van teksten op chronologische, geografische of andere eigenschappen van hun herkomst. In alle experimenten werden zowel grote als kleine hoeveelheden trainingsdata gebruikt; in de AA- en AV-experimenten waren zowel lange als (zeer) korte Koranische teksten het onderwerp van classificatie.

In de AA-experimenten ontdekten we dat SMS-tags een zeer betrouwbare stijl zijn voor robuuste classificaties van korte Koranteksten. De AV-experimenten bevestigen dat SMS-tags significant bewijs leveren voor discussies over het auteurschap van de Koran. De resultaten van de SV-experimenten toonden hoge betrouwbaarheid aan voor het verifiëren van de geografische bron (Mekka of Medina) van hoofdstukken en verzen uit de Koran.

De resultaten van deze experimenten leidden tot de conclusie dat de sequenties van morfosyntactische tags, gegenereerd door het MOBIN-systeem, een computationeel betrouwbaar styloom vormen voor het representeren van de stylometrische kenmerken van korte Koranische teksten. Voorts hebben we systematisch evidentie gevonden dat grammaticale informatie, in het bijzonder in de vorm van SMS-tags, zeer effectief is voor stylometrische classificaties van Koranische teksten van elke lengte. Deze bevindingen kunnen gezien worden als een basis voor toekomstig onderzoek naar de implicaties van stylometrische patronen in het Klassiek Arabisch in het algemeen en in Koranische teksten in het bijzonder.
Original languageEnglish
Award date16 Jun 2026
DOIs
Publication statusPublished - Jun 2026

Fingerprint

Dive into the research topics of 'Quran Mining: Computational Stylometry of Quranic Texts'. Together they form a unique fingerprint.

Cite this