Initial Explorations for Document Clustering Tasks in Latin Elegiac Poets
Initial Explorations for Document Clustering Tasks in Latin Elegiac Poets
Date
2025-6-19
Authors
Nusch, Carlos Javier
Del Río Riande, Gimena
Cagnina, Leticia Cecilia
Errecalde, Marcelo Luis
Antonelli, Ruben Leandro
Journal Title
Journal ISSN
Volume Title
Publisher
Springer, Cham
Abstract
This article describes various Automatic Text Analysis tasks
applying Natural Language Processing techniques on a corpus of Latin texts
from the 1st century BC and 1st century AD. The motivation behind this work
is to delve into and understand a historical literary trend revolving around the
themes of love, spanning from antiquity through to the medieval period. The
analyzed authors include Gaius Valerius Catullus, Albius Tibullus, and Sextus
Propertius, who represent the literary movement of the neoterics, as a group of
poets to be identified, and Publius Vergilius Maro and Marcus Annaeus
Lucanus, epic poets with remarkably distinct styles, as control samples. The
purpose of this preliminary and exploratory study is to investigate the potential
and best features for document clustering. The clustering tasks were carried out
using fixed ranges of character n-grams and word n-grams. For the clustering
tasks, the K-Means method and the Silhouette Index were used for determining
the optimal cluster sizes. Using optimal clusters as labels, decision trees were
trained for each range of n-grams, aiming to identify features with the highest
Information Gain and Information Gain Ratio. The trees were trained based on
the criterion of Entropy, and calculations of Feature Importance were
performed. Results show variations based on text preprocessing techniques:
simple filtering of stopwords in the corpus yields better Silhouette scores, with
one or two features showing potential classification value for the decision trees.
The application of TF-IDF weighting results in Silhouette indices closer to
zero, albeit with a more balanced distribution of Importance among different
features.
Description
Keywords
Latin Elegiac Poets,
document clustering,
K Means,
Silhouette Coefficient,
decision trees,
Feature Importance,
Information Gain Ratio
Citation
Nusch, C.J., del Rio Riande, G., Cagnina, L.C., Errecalde, M.L., Antonelli, L. (2025). Initial Explorations for Document Clustering Tasks in Latin Elegiac Poets. In: Agredo-Delgado, V., Ruiz, P.H., Meneses Escobar, C.A. (eds) Collaboration in Knowledge Discovery and Decision Making. DECISIONING 2024. Communications in Computer and Information Science, vol 2369.