Reference:
Chekhovich Y.V..
Artificial Intelligence in the Higher Education System: An Analysis of Institutional Response Scenarios
// Historical informatics.
2026. № 2.
P. 42-57.
DOI: 10.7256/2585-7797.2026.2.80266 EDN: WXIDIR URL: https://en.nbpublish.com/library_read_article.php?id=80266
Abstract:
The object of this study is the higher education (HE) system in the context of the proliferation of services based on generative artificial intelligence (GenAI) models. The subject of the study is the institutional responses of the HE system to the widespread use of GenAI in educational and research activities, including the impact of these technologies on academic writing, assessment, academic integrity, and university governance practices. The aim of the study is to systematize and compare scenarios of institutional response of higher education to the spread of GenAI, as well as to justify the transformational model as the most promising strategy for adapting the research and educational system. The relevance of the study is determined by the established high popularity of generative services among students, teachers, and researchers, as well as the observed increase in the proportion of scientific and academic works showing signs of GenAI use. The research methodology is based on empirical analysis of scientific articles published in academic journals, conference proceedings, and preprints that analyze the advantages, describe the methods and models of GenAI, examine the use of GenAI tools by students, teachers, and scientists, and present the results of surveys and other types of studies relevant to this work. The scientific novelty of the study lies in the fact that, for the first time within a single review, it systematizes the main scenarios of institutional response of higher education to the mass spread of GenAI; offers a comparative description of these scenarios through the lens of their applicability in educational practice; and shows that the limited reliability of AI-text detection makes prohibitive and strictly regulatory approaches insufficient, while reinforcing the importance of transforming assessment and pedagogical practices. The study analyzes the advantages that the use of GenAI offers to higher education, as well as the risks arising from the development of these technologies. It examines current methods for detecting machine-generated text and their limitations. It is shown that algorithmic detection methods have limited reliability and should be used as a supporting mechanism; the need to transform assignments, assessment forms, and tools for preparing scientific texts is justified. The study also analyzes various scenarios of institutional response to technological changes: ignoring, prohibition, strict regulation (partial prohibition), and system transformation. The paper concludes that experts favor the transformation of the research and educational system as the optimal strategy for higher education response.
Keywords:
artificial intelligence, generative AI, higher education, academic integrity, AI-text detection, large language models, assessment, plagiarism, digital transformation, academic system
Reference:
Tormozov V.S., Petrenko E.G..
Digital Anthropology of Emotions: A Sentiment Analysis of Diaries and Letters of Russian State Leaders in the 18th-19th Centuries
// Historical informatics.
2026. № 2.
P. 58-71.
DOI: 10.7256/2585-7797.2026.2.80576 EDN: QUMPTB URL: https://en.nbpublish.com/library_read_article.php?id=80576
Abstract:
This research is situated at the intersection of digital humanities, the history of emotions, and computational linguistics. The article presents the results of the sentiment analysis of the epistolary heritage and diary entries of three key figures of the Russian monarchy: Catherine II, Alexander I, and Nicholas I. The total corpus of analyzed texts amounted to over 2 million word usages. Using models of deep learning BERT (XLMRoBERTaLarge and Conversational RuBERT) adapted for historical texts, the authors reconstruct the emotional dynamics of communication throughout a turbulent century—from Enlightened absolutism to the crisis of the Nicholas system. The study confirms the hypothesis of a stable correlation between the genre of the document (official letter vs. private diary) and the degree of emotional expressiveness, as well as identifies specific lexical markers of anxiety during periods of political instability (on the eve of the Decembrist revolt and during the Crimean War). Methodologically, the research is based on three conceptual foundations. Firstly, it is the theory of emotional communities by B. Rosenwein, according to which emotions are constructed within social groups with shared values and norms of expression. Secondly, it adopts the semiotic approach of Yu. M. Lotman in studying the everyday behavior of the Russian nobility. Thirdly, it employs methodologies of computational text analysis. The hypothesis of this study is as follows: the emotional tone of the personal correspondence and diaries of Russian monarchs is not so much a spontaneous expression of an individual psychological state, but rather a ritualized social action, subject to the cultural codes of the era and genre canon. The aim of the study is to conduct a comprehensive historical-linguistic analysis of the emotional tone of the epistolary and diary heritage of Russian statesmen of the 18th-19th centuries using digital text processing methods, to identify stable emotional patterns and their connection with historical-biographical context. It has been established that the sentimentalist tradition of the late 18th century paradoxically combined with hypertrophied emotional restraint in official communication, creating an effect of "emotional dissonance," which was resolved in the literature and journalism of the 19th century. This work contributes to the methodology of analyzing historical texts, demonstrating the possibilities and limitations of NLP tools when working with archaic vocabulary and bilingual corpora (Russian-French linguistic dualism).
Keywords:
digital anthropology, emotions, analysis, tonality, diaries, letters, state, emperor, Russian Empire, 18th-19th centuries
Reference:
Pavlov A.V..
Development of a chat-bot prototype based on HybridRAG to simplify access to documents of the National Archives and Records Administration of the United States
// Historical informatics.
2025. № 4.
P. 177-202.
DOI: 10.7256/2585-7797.2025.4.76571 EDN: SCVTIE URL: https://en.nbpublish.com/library_read_article.php?id=76571
Abstract:
The subject of the research is the problem of effective information retrieval in archives, specifically in the catalog of the National Archives and Records Administration (NARA) of the United States. The aim of this work is to create a prototype chatbot capable of answering users' questions based on materials from the NARA catalog. Special attention is given to justifying the choice of a specific approach to generation, taking into account the specifics of the subject area, which is necessary to achieve the goal. The algorithm for creating a knowledge graph based on catalog materials and the architecture of the chatbot are discussed in detail. Various components used in the script for creating the knowledge graph and within the proposed chatbot architecture are analyzed. Examples of both correct and incorrect operation of the system using Mistral-7B-Instruct-v0.3 are examined. The methodology is based on the implementation of the HybridRAG approach, a variant of GraphRAG, which involves a combination of semantic search using vector representations and search via graph queries. The result of this work was the creation of a script that generates a knowledge graph based on catalog materials using Neo4j, a script for analyzing the generated graph, and a prototype chatbot that implements the HybridRAG approach using the generated graph. The area of application for the results is information retrieval systems for archives. The novelty lies in the adaptation of the HybridRAG approach to work with the complex hierarchical structure of the catalog. The prototype demonstrated the ability to provide correct answers based on catalog data. At the same time, key limitations of the prototype were identified: the large language model used within it must be resilient to "noise," refrain from answering questions when the necessary information is absent in the provided context, and be able to integrate information from multiple documents. The results of the work, including the created scripts, prototype code, and additional materials (tables, images, prompts, Cypher queries), are available in the GitHub repository at the following link: https://github.com/AlexanderPavlov36/NARA_ChatBot.
Keywords:
archival documents, archival search, large language models, knowledge graph, artificial intelligence, chatbot, GraphRAG, HybridRAG, NARA, Neo4j
Reference:
Kuznetsov A.V..
From Populus Romanus to Populus Christianus: The Concept of "People" in Thomas Aquinas through the Lens of Distributional Semantics
// Historical informatics.
2025. № 4.
P. 79-102.
DOI: 10.7256/2585-7797.2025.4.77131 EDN: MONANA URL: https://en.nbpublish.com/library_read_article.php?id=77131
Abstract:
The subject of the study is the structural transformation of the semantic field of lexemes denoting forms of human community (populus, plebs, gens, natio, vulgus, multitudo) in the works of Thomas Aquinas compared to the linguistic norm of Classical Latin. The transformation of the concept of populus in scholastic thought is documented in a number of special studies (I. Congar, P. Bolloni, A. V. Marey); however, the traditional historical-conceptual analysis based on the hermeneutics of individual textual passages faces the problem of representativeness. Conclusions drawn from a limited set of quotes carry the risk of subjective interpretation and do not allow for an assessment of the systemic nature of the identified changes. The aim of the research is to empirically verify and quantitatively measure these changes through the application of computational linguistics tools. The choice of Thomas Aquinas is determined by his central position in the scholastic tradition and the presence of detailed qualitative hypotheses subject to testing. The study is based on methods of distributive semantics. The analysis is conducted on two aligned word embedding models: Opera Latina (Classical Latin) and Opera Maiora (the works of Thomas Aquinas). The methodology includes calculating cosine similarity to measure semantic drift, visualizing the topology of semantic fields (PCA, heat maps), and comparative analysis of nearest semantic neighbors. The main findings of the study are as follows. The quantitative analysis has recorded a fundamental reconfiguration of the socio-political vocabulary in the language of Thomas Aquinas: negative or near-zero values of cosine similarity for key lexemes indicate a radical change in the contexts of their usage. It has been established that the concept of populus undergoes a triple transformation: sacralization, hierarchy, and de-subjectivation. The center of the new conceptual system becomes ecclesia, which conceptually encompasses the functions of the previous political categories. A greater stability of the vocabulary of vertical power compared to that of civil participation has been identified. The results obtained confirm and clarify the conclusions of previous studies conducted using traditional methods. The scientific novelty of the work lies in demonstrating the possibility of moving from qualitative interpretations of individual passages to statistically substantiated conclusions about the structure of semantic fields. The approach is applicable for verifying hypotheses of intellectual history based on the works of other authors and eras.
Keywords:
Thomas Aquinas, word embeddings, Latin, conceptual history, semantic shift, diachronic analysis, scholasticism, political vocabulary, distributional semantics, digital humanities
Reference:
Lyutikova L.A..
Hybrid model of speech act classification: combination of RuBERT embeddings and logical-algebraic implicants
// Historical informatics.
2025. № 4.
P. 103-114.
DOI: 10.7256/2585-7797.2025.4.77162 EDN: MPEBQM URL: https://en.nbpublish.com/library_read_article.php?id=77162
Abstract:
The subject of the study is the problem of interpretable automatic classification of speech acts in Russian-language utterances, which represents one of the key tasks of modern computational linguistics and applied natural language analysis. Traditional machine learning methods and transformer models demonstrate high quality; however, they remain insufficiently transparent, hindering their application in critical areas where explainability of decisions is required. This work explores the possibility of combining contextual embeddings from RuBERT with logical-algebraic mechanisms for feature interpretation aimed at identifying stable latent structures corresponding to various types of speech acts: assertions, questions, requests, and assumptions. The study encompasses the processes of corpus formation, extraction of vector representations, construction of a classification model, and analysis of how logical rules can enhance the reliability of interpretation of results. The methodology is based on a combination of RuBERT contextual embeddings, threshold binary feature binarization, and logical-algebraic extraction of minimal implicants, applied together with a linear classifier in a hybrid architecture. The scientific novelty of the research lies in the development of a hybrid approach that combines the strengths of neural network and symbolic methods: continuous embeddings from RuBERT are used to form informative representations of text, while logical-algebraic implicants provide transparent interpretation of the classifier’s decisions. The proposed model demonstrates that it is possible to construct compact logical rules over high-dimensional embeddings, describing stable regions of the feature space and enhancing trust in the system's operation. Experimental results show that the hybrid architecture surpasses the linear classifier by 3-4 percentage points, achieving high performance. The findings confirm that logical-symbolic structures can simultaneously improve the accuracy and interpretability of models, making the approach promising for Explainable AI and the development of dialog and analytical language systems.
Keywords:
classification of speech acts, RuBERT Contextual Embeddings, logical-algebraic implicants, data, analysis, hybrid neuro-symbolic models, computer linguistics, artificial intelligence, record, message
Reference:
Rinchinov O.S., Shelukheev A.A..
Buryat historical sources: digital infrastructure of machine translation
// Historical informatics.
2025. № 4.
P. 143-157.
DOI: 10.7256/2585-7797.2025.4.77283 EDN: QRCJHQ URL: https://en.nbpublish.com/library_read_article.php?id=77283
Abstract:
The study is dedicated to a vast yet still underexplored corpus of Buryat historical sources in the Old Written Mongolian language, preserved in academic and archival institutions in Russia. The introduction of these documents into scientific and public circulation, reflecting all aspects of Buryat society from the 18th to the early 20th centuries, requires the application of modern digital methods. The most promising approach to addressing this task is a comprehensive strategy employing digital humanities tools, which includes digitizing sources, full-text input of texts using romanized transliteration, creating a text Mongolian-Russian corpus with metatextual, structural, and morphological markup according to TEI standards, translating sources, and forming a balanced parallel Mongolian-Russian corpus. Special attention was paid to the genre, chronological, and territorial representativeness of the corpus data. The prepared digital materials serve as a foundation for applying machine learning methods to tackle tasks in optical text recognition for Mongolian script and machine translation. To assess the prospects of machine translation for Buryat historical sources, computational experiments were conducted with transformer models trained "from scratch" and large pre-trained multilingual models (mBART-50, mT5). The scientific novelty of the work is defined by the creation of the first methodologically grounded digital platform for studying Buryat historical sources, adapted to the characteristics of the local rendition of Old Written Mongolian. Important competencies have been gained for organizing the complete cycle of digital processing of historical sources from the digitization of archival documents to the formation of a balanced parallel corpus. For the first time, an online corpus of unique historical texts has been assembled and published, equipped with analytical tools, as well as specialized datasets for training AI models. It has been established that for low-resource language pairs, the most effective strategy is fine-tuning pre-trained multilingual models rather than modifying neural network architectures. The research lays the groundwork for creating comprehensive tools for digitizing Buryat written heritage, which will open new perspectives for historical, linguistic, and cultural studies.
Keywords:
Buryat historical sources, text corpus, deep learning, datasets, parallel corpus, machine translation, pretrained models, Computational experiment, Low-resource language, authority control
Reference:
Kotov A.S..
Fine-tuning a model based on the Transformer architecture for normalizing a corpus of medieval texts in German from the 14th-15th centuries from the Order of Prussia.
// Historical informatics.
2025. № 3.
P. 128-140.
DOI: 10.7256/2585-7797.2025.4.75275 EDN: XOHQXO URL: https://en.nbpublish.com/library_read_article.php?id=75275
Abstract:
The article is dedicated to the methods of automatic normalization of texts in Middle High German and Early New High German for the application of NLP in medieval history research. It provides an overview of existing approaches to the automatic normalization of historical texts in German. The problems of normalizing medieval German texts are identified: the peculiarities of using substitution dictionaries and replacement rules. The limitations of these approaches and the necessity of considering the goals of normalization are described. Neural language models are defined as the most promising for automatic normalization. The study compares the effectiveness of existing neural language models (NMT) with respect to texts in Middle High German and Early New High German. It demonstrates the low effectiveness of using NMT trained on texts from the New and Modern eras. Based on reviews presented in the literature, it asserts the need to prepare NMT according to specific goals and corpora. For the normalization of texts from the 14th-15th centuries created in monastic Prussia, a neural language model based on the Transformer architecture (BART) was further trained, and its effectiveness was presented in comparison with other models. The model was trained on a custom dataset of word pairs: original-normalized, consisting of 6,570 pairs. The conditions for retraining the model were: Epoch = 28; Batch = 50. For normalizing a corpus of texts in three historical forms of the German language, the DTAEC Type Normalizer model was chosen. The effectiveness of the retrained model's normalization was compared with existing models trained on German texts from the New and Modern eras based on the metrics of Accuracy, Accuracy OOV, CER, and Levenshtein distance. The retrained model shows significant effectiveness compared to other models. One normalized sentence using the model is proposed for review, and a comparison with a benchmark is conducted. Instances of "hallucinations" in the retrained model were identified. With an Accuracy OOV of 89.6, using this method is considered promising. However, the identified shortcomings in text normalization indicate the necessity of employing additional normalization methods, such as lemmatization.
Keywords:
normalization, AI, transformer, BART, Mittelhochdeutsch, Frühneuhochdeutsch, German Order, Prussia, Middle Ages, Digital Humanities
Reference:
Kuznetsov A.V..
Automatic information extraction from ego-documents: a comparative analysis of the effectiveness of large language models based on the example of K.A. Berezkin's diary.
// Historical informatics.
2025. № 3.
P. 99-127.
DOI: 10.7256/2585-7797.2025.3.75850 EDN: ZAYBBF URL: https://en.nbpublish.com/library_read_article.php?id=75850
Abstract:
The subject of the study is a comparative analysis of the performance, analytical strategies, and limitations of four large language models – Gemini-2.5-Pro, o3, Grok3, and Deepseek-v3 – in the task of extracting structured information from a historical ego-document. The analysis aims to determine the models' ability to work with complex narratives characterized by a high degree of subjectivity, an abundance of indirect evidence, multi-layered meanings, and emotional coloration. The key limitations of the models – over-interpretation, missing indirect evidence, and the trade-off between completeness and accuracy – are considered part of their analytical strategies. The material used was the diary of the Vologda gymnasium student K.A. Berezkin for the year 1849. The work addresses a complex task of developing and testing an approach that allows for the transformation of unstructured source text into a dataset suitable for solving a specific historiographical task – analyzing the perception of the European revolutions of 1848-1849 in the Russian province. The methodology is based on the automatic extraction of structured information using large language models. A comprehensive toolkit has been developed, including a domain-specific ontology, prompts, and a detailed JSON schema for data capture. The performance of the models was evaluated based on quantitative (completeness, accuracy, F1-score) and qualitative indicators (granularity, adherence to the ontology, understanding of historical context, typical errors). The scientific novelty lies in the first systematic testing and comparative analysis of the performance of leading language models in working with a historical ego-document in domestic historiography. It was established that the models implement various data extraction strategies: from exhaustive, but "noisy" coverage (Gemini-2.5-Pro) to highly accurate, but selective (Deepseek-v3), which directly determines the suitability of the resulting dataset for different research scenarios: from exploratory analysis to the creation of verified databases. The key conclusion of the study is that automated extraction is not merely a technical operation, but a form of digital hermeneutics. Accordingly, the final dataset is not objective data passively "discovered" in the source, but capta – a set of information selected for a specific task. The study shows that the application of artificial intelligence raises historian's requirements for critical expertise, shifting their role from information retrieval to verification and interpretation of machine results.
Keywords:
large language models, information extraction, artificial intelligence, digital humanities, digital hermeneutics, ego-documents, microhistory, revolutions of 1848-1849, Russian Empire, 19th century
Reference:
Latonov V.V., Latonova A.V..
Hierarchical clustering of the readings of the members of the Society of United Slavs using fuzzy set theory methods
// Historical informatics.
2025. № 3.
P. 141-150.
DOI: 10.7256/2585-7797.2025.4.75387 EDN: XSFVAJ URL: https://en.nbpublish.com/library_read_article.php?id=75387
Abstract:
The subject of the research presented in this article is the testimonies of members of the Society of United Slavs regarding the murder of the royal family. The article accumulates all testimonies from members of the Society that touch upon the issue of the intent to murder the royal family. The focus of the study in the article is the degree of radicalization among the members of the Society of United Slavs and the degree of similarity in their views on the proposed methods of the Society. The authors employ expert evaluation methods for an objective interpretation of each participant's testimony and to uncover their understandings of the goals of the Society of United Slavs. Subsequently, the authors apply methods from fuzzy set theory to construct an objective hierarchical clustering of the members of the Society, to demonstrate the internal connections that existed among the participants based on the similarity or dissimilarity of their views. The hierarchical clustering of the members of the Society is based on their testimonies. The authors establish an objective scale of radicalism for the testimonies of each member of the Society and introduce a measure of similarity of their testimonies, on the basis of which clustering is further constructed using the transitive closure of the introduced relation. The main conclusion of the presented work is that within the Society of United Slavs, two clusters were identified, wherein the Decembrists held diametrically opposed views regarding the permissible methods of achieving the Society's goals. The first cluster is centered on Decembrist I.I. Gorbachyovsky and includes Decembrists N.F. Lisovsky and I.V. Kiriev. Members of this cluster were convinced that the Society planned the murder of the royal family and were willing to adhere to this idea until the end. The second cluster included P.I. Borisov and A.I. Tyutchev, who were confident that the murder of the royal family was not intended. The scientific novelty of the work lies in the fact that it is the first time that fuzzy set theory has been applied to the method of hierarchical clustering of members of the Society of United Slavs.
Keywords:
Society of United Slavs, Decembrists, Secret society, Radicalism, Hierarchy, Fuzzy logic, Fuzzy set theory, Clustering, Transitive closure, Expert judgment method
Reference:
Debenova Z.A., TSipilova S.S., Tsyrenova N.D..
Monuments in Mongolian Writing: An Experience of Creating a Parallel Corpus
// Historical informatics.
2025. № 2.
P. 1-10.
DOI: 10.7256/2585-7797.2025.2.73930 EDN: MMDRBC URL: https://en.nbpublish.com/library_read_article.php?id=73930
Abstract:
This article highlights the results of the work on creating a parallel corpus of Buryat sources in Mongolian script. The project is being carried out with the support of the Russian Science Foundation, based on the archival materials from the Center for Eastern Manuscripts and Xylographs of the IMBT SB RAS. The subject of the research is the process of creating a database for the corpus, the specifics of compiling it, particularly the selection of materials. Currently, the developing corpus includes the following documents from the archival funds of the CVRK IMBT SB RAS: texts of historical content—"A Brief Outline of the History of Khori-Mongolian Buryats," "On the History of the Zugalai Region"; an official document "Protocol of the All-Buryat Assembly in Chita in 1917"; an ethnographic composition "Narrative of Samdan Noyon," a medical work "Notes of Tibetan Doctor Donduba Munkuyev"; a work of Buddhist didactic literature "Subhashita" translated by Galsan-Jimba Tuguldur. General scientific and source study methods were applied to the analysis of handwritten, printed, and xylographic texts in Mongolian script. The processes of material selection, their transliteration and translation, as well as substantive (thematic, lexical) and technical aspects (typos, pagination, numerals) were examined. The parallel Russian-language version is being created by the research group. The authors emphasize the significance of creating a parallel corpus as a resource for further research in the field of Buryat linguistics, translation studies, and cultural studies, as well as its role in promoting Old Mongolian script among the general public and preserving the intangible heritage of the Baikal region. The corpus represents a unique database for further research in various fields of science, etc. The texts considered will serve as a basis for the development of machine translation algorithms, and the work being conducted at this stage will help future developers create more effective algorithms. The creation of a specialized database that is open not only to researchers but also to representatives of the educational sector, professional translators, and anyone showing a scientific or cultural interest in written heritage appears promising.
Keywords:
Mongolian script, parallel corpus, written sources, Buryatia, Center of Oriental Manuscripts and Xylographs, Baikal region, intangible heritage, machine translation, digitization, text corpus
Reference:
Latonov V.V., Latonova A.V..
Determining the authorship of the "Notes of the Decembrist I.I. Gorbachevsky" by machine learning methods
// Historical informatics.
2025. № 1.
P. 122-133.
DOI: 10.7256/2585-7797.2025.1.72805 EDN: QALGAU URL: https://en.nbpublish.com/library_read_article.php?id=72805
Abstract:
In the presented work, the object of research is the "Notes of the Decembrist I.I. Gorbachevsky", which are one of the most valuable sources on the history of the Decembrist movement, created by its participants themselves. They highlight the formation and development of such a Decembrist organization as the Society of United Slavs, which later joined the Southern Society of Decembrists. Written in exile in Siberia, these notes represent not only a source of factual material, but also an original concept of the secret society's development, and a retrospective "inside look" at the mistakes made by the conspirators. However, Gorbachevsky's "Notes" are notable for another circumstance. Contrary to their well-established name in literature, we cannot unequivocally assert that their author was I.I. Gorbachevsky himself from among the Decembrists. The fact is that the first publication of the "Notes" – in the journal "Russian Archive" in 1882 – was presented under the heading "Notes of an Unknown Person from the Society of the United Slavs." The subject of the research in the presented work is the question of the authorship of the "Notes", which has no clear answer among historians today. In this paper, we propose a solution to the problem of determining the authorship of the "Notes of the Decembrist I.I. Gorbachevsky" using machine learning methods. I.I. Gorbachevsky himself, as well as the Decembrist P.I. Borisov, are considered as possible authors. The novelty of the research lies in the fact that machine learning methods were used to determine the authorship of the "Notes". The authors trained four types of models to predict the authorship of each of the sentences in the Notes. As a result, most of the proposals of the "Notes" were assessed as written by Gorbachev. The largest percentage of offers, 69.2%, was attributed to Gorbachev by the Count Vectorizer + SVC model. The accuracy of all models exceeded 80% on average, while those based on BERT coding averaged close to 90%. The main conclusion of the work, therefore, can be considered that the "Notes" were more likely to have been written by I.I. Gorbachevsky than by P.I. Borisov. The methods used in the framework of the presented study provide another argument in favor of this version. The code and dataset are available at the link: https://github.com/WLatonov/Gorbachevskiy_notes .
Keywords:
authorship definition, Attribution, Stylometry, Machine learning, Neural networks, Binary classification, BERT, The Decembrists, Gorbachevskiy's notes, Gorbachevskiy's letters
Reference:
Yumasheva J.Y..
The possibility of using artificial intelligence in historical research
// Historical informatics.
2025. № 1.
P. 95-121.
DOI: 10.7256/2585-7797.2025.1.73578 EDN: PQTZJT URL: https://en.nbpublish.com/library_read_article.php?id=73578
Abstract:
The article is devoted to the controversial problem of the use of artificial intelligence in historical research. The introduction briefly examines the history of the emergence of "artificial intelligence" (AI) as a field in computer science, the evolution of this definition and views on the application of AI; analyzes the place of artificial intelligence methods at different stages of specific historical research. In the main part of the article, based on the analysis of historiographical sources and his own experience of participating in foreign projects, the author analyzes the practice of implementing handwritten text recognition projects using various information technologies and AI methods, in particular, describes and justifies the requirements for creating electronic copies of recognizable sources, the need to take into account the texture of information carriers, writing materials, techniques and technologies for creating the text; varieties and methods of creating paleographic, codicological, diplomatic datasets, historical and lexicological dictionaries, the possibility of using large language models, etc. As a methodological basis, the author used a systematic approach, historical-comparative, historical-chronological and descriptive methods, as well as the analysis of historiographical sources. In conclusion, it is concluded that the use of artificial intelligence technologies is promising not only as an auxiliary tool, but also as research methods that help in establishing the authorship of historical sources, clarifying their dating, detecting forgeries, etc., as well as in creating new types of scientific reference search systems for archives and libraries. At the same time, the use of artificial intelligence technologies is highly expensive and capital intensive, which is a serious obstacle to the widespread introduction of these technologies into the practice of historical research.
Keywords:
artificial intelligence, historical sources, automated text recognition, paleography, codicology, diplomatics, historical lexicology, datasets, large linguistic models, information technologies
Reference:
Voronkova D.S..
Computerized content analysis of articles from the journal "Bulletin of Finance, Industry, and Trade" for the year 1917: testing the capabilities of the artificial intelligence module in the MAXQDA program.
// Historical informatics.
2025. № 1.
P. 134-161.
DOI: 10.7256/2585-7797.2025.1.73332 EDN: QEHIBU URL: https://en.nbpublish.com/library_read_article.php?id=73332
Abstract:
The subject of the research is the articles of the official printed organ of the Russian Ministry of Finance – the journal "Bulletin of Finance, Industry and Trade" – for the year 1917. Undoubtedly, this year was a turning point in domestic history. In this regard, it is important to use new approaches to uncover the informational potential of this largely unique source, which contains valuable information about the country's economy (including not only those areas highlighted in the journal's title but also, for example, about tax and customs policy, as well as preparations for a number of reforms, including agrarian reforms). Moreover, it is necessary to take into account that during this period the journal was published against the backdrop of the ongoing First World War, and the related issues were also reflected in its pages. Methodologically, the article is based on computerized content analysis. The main focus is on artificial intelligence tools within the specialized software MAXQDA. The novelty of the research lies in the fact that for the first time the capabilities of the AI Assist module and its latest component, MAXQDA Tailwind, which was in the beta version at the time of the article's publication, have been tested. The author received early access to all product features by invitation from the developers and provided feedback based on the work outcomes. The international virtual conference of MAXQDA users (MAXDAYS 2025), where the functionality of MAXQDA Tailwind will be presented, will take place on March 18-19 of this year. Thus, readers will be able to familiarize themselves with it before its official release. The article proves that artificial intelligence in no way replaces the historian but can assist them in deepening and making the analysis of historical sources more comprehensive.
Keywords:
Bulletin of Finance, Media, content analysis, MAXQDA, artificial intelligence, AI Assist, MAXQDA Tailwind, official press organ, First World War, February Revolution
Reference:
Borodkin L..
Historian in the world of neural networks: the second wave of artificial intelligence technology application.
// Historical informatics.
2025. № 1.
P. 83-94.
DOI: 10.7256/2585-7797.2025.1.74100 EDN: QXYMHF URL: https://en.nbpublish.com/library_read_article.php?id=74100
Abstract:
Over the last decade, artificial intelligence (AI) technologies have become one of the most sought-after areas of scientific and technological development. This process has also impacted historical science, where the first research in this area began in the 1980s (the so-called first wave) – both in our country and abroad. Then came the "AI winter," and at the beginning of the 2010s, the "second wave" of AI emerged. The subject of this article is the new opportunities for applying AI in history and the new problems arising in this process today, when the main focus of AI has shifted to artificial neural networks, machine learning (including deep learning), generative neural networks, large language models, etc. Based on the experience of historians applying AI, the article proposes the following seven directions for such research: recognition of handwritten and old printed texts, their transcription; attribution and dating of texts using AI; typological classification and clustering of data from statistical sources (particularly using fuzzy logic); source criticism tasks, data completion and enrichment, and reconstruction using AI; intelligent search for relevant information, utilizing generative neural networks for this purpose; using generative networks for text processing and analysis; and the use of AI in archives, museums, and other institutions that store cultural heritage. An analysis of the discussion of similar issues organized by the leading American historical journal AHR has been conducted. These are conceptual questions regarding the interaction between humans and machines ("historian in the world of artificial neural networks"), the possibilities for historians to use machine learning technologies (particularly deep learning), various AI tools in historical research, as well as the evolution of AI in the 21st century. Practical aspects were also touched upon, such as the experience of recognizing newspaper texts from past centuries using AI. In conclusion, the article addresses the problems related to the use of generative neural networks by historians.
Keywords:
Artificial Intelligence, artificial neural networks, machine learning, deep learning, generative neural networks, image recognition, text atribution, algorythms, data, historical source
Reference:
Mekhovskii V.A., Kizhner I.A..
The world through the eyes of an educated person in Minusinsk of the late XIX - early XX centuries: distribution of the frequency of geographical names in the books of the Minusinsk Public Library
// Historical informatics.
2025. № 1.
P. 174-189.
DOI: 10.7256/2585-7797.2025.1.72586 EDN: QCQWHG URL: https://en.nbpublish.com/library_read_article.php?id=72586
Abstract:
The subject of the study is the corpus of children's literature from the collection of the Minusinsk Public Library of the late XIX – early XX century, consisting of 121 works written between 1719 and 1905. These texts are a significant source for studying the formation of geographical perception among residents of a provincial Siberian city through fiction. Special attention is paid to the analysis of geographical names (toponyms) found in texts in order to identify their frequency and geographical distribution. This allows us to reconstruct the picture of the world presented in the books of that time and understand how it was perceived by the children's audience, forming their idea of countries, cities and cultural centers. The research is aimed at studying the role of children's literature as a cultural tool that reflects and forms geographical representations, as well as at identifying methodological challenges and limitations when working with historical buildings. The methodological basis includes bringing pre-reform texts to a machine-readable form using digitization tools and geoparsing to automatically identify geographical entities. The Spacy library was used for the analysis, followed by manual verification and correction of the data. The results of the study include the identification of 668 cities and 97 countries represented in the texts, as well as the construction of a cartographic visualization of the frequency distribution of mentions. The analysis revealed an uneven distribution of geographical names in various texts, where mentions of Russia, Poland and England prevail among countries, and Kiev, Moscow and St. Petersburg among cities. The scope of the results includes research in the field of digital humanities, library science and historical and cultural studies. The novelty of the work lies in the use of modern geoparsing methods for processing Russian-language texts of pre-reform spelling and in the analysis of the previously unexplored literature corpus of the Minusinsk Library. The conclusions emphasize the importance of text mapping for understanding the formation of geographical perception and the need for further development of NER tools for complex corpora. Despite the limitations, the research contributes to the development of NLP methods for historical texts.
Keywords:
Geoparsing, Mapping, Named-entity recognition, Historical Computer Science, Siberia, Minusinsk, World map, Children's literature, Minusinsk Public Library, Pre-reform orthography
Reference:
Mashchenko N.E., Gaidar E.V..
Artificial intelligence technologies in the formation of the archival environment: problems and prospects
// Historical informatics.
2025. № 1.
P. 162-173.
DOI: 10.7256/2585-7797.2025.1.73393 EDN: QEIGBR URL: https://en.nbpublish.com/library_read_article.php?id=73393
Abstract:
The authors studied the prospects of using artificial intelligence (AI) technologies to create and develop a digital archival environment, as well as their impact on the optimization and automation of archived data management processes. The main purpose of the work is to analyze modern digital solutions aimed at improving the processes of storing, searching and processing archival documents (including handwritten, damaged, multilingual). The paper explores key technologies used in digital archives, including intelligent scanning, natural language processing (NLP), computer vision, machine learning, and intelligent search methods. Special attention is paid to the problems of loss of archival materials, the need to restore them, ensure data security and accessibility, which is especially important in an unstable political situation and limited resources for new territories. The research is based on a systematic analysis of modern information technologies and their application in the archival business. The work uses methods of comparative analysis, classification and forecasting, which allows us to identify key areas of AI implementation in the archival field. The novelty of the work lies in an integrated approach to analyzing the use of AI in the archival field, identifying problematic aspects of archive digitalization, and proposing automation of the processes of storing, processing, and searching archival data. It is concluded that artificial intelligence technologies can significantly improve the efficiency of archives, providing accelerated document processing, intelligent classification, data protection and convenient access to information. In addition, the need to develop new algorithms based on machine learning is emphasized, which will improve the recognition of handwritten texts, the processing of corrupted documents and multilingual archival materials. The introduction of such technologies is becoming an important part of the digital transformation strategy of archival affairs and plays a key role in preserving historical heritage.
Keywords:
archives, digital archival environment, digital transformation, artificial intelligence, machine learning, computer vision, natural language processing, data security, intelligent scanning, predictive intelligence
Reference:
Orekhov B.V..
Text and knowledge in the aspect of large language models
// Historical informatics.
2023. № 4.
P. 104-113.
DOI: 10.7256/2585-7797.2023.4.44180 EDN: BJQBQB URL: https://en.nbpublish.com/library_read_article.php?id=44180
Abstract:
The focus of this text is on the influence of large linguistic models on the self-determination of the humanities. Large language models are able to generate plausible texts. It seems that they thus become on a par with other tools that, throughout the development of technology have freed people from routine. At the same time, for the humanities, the individualization of the generated texts is very great, and knowledge itself is closely related to its textual embodiment. If we agree that knowledge is a text, and embodied in another text, another knowledge appears before us, then humanities will have to answer the question of how a text generated by a person differs in value from the same text generated by a machine. The text of the work raises methodological and epistemological problems of the correlation of texts of natural and artificial origin if they are made in the genre of a scientific work. The difference between such artifacts is clearly visible only for some scientific disciplines, and raises questions about the rest. These issues should be resolved with the help of deep reflection, which was not so urgently needed in the last centuries of the development of the humanities, but which is now required from a humanitarian scientist. The humanitarian will have to explicitly oppose himself to large language models and prove the importance of his work compared to what a neural network can generate.
Keywords:
large language models, chatgpt, scientific publications, methodology of science, text generators, knowledge, the science, text, formal languages, Humanities