Urraca AI: A Humanistic Framework for Artificial Intelligence and the Past
Developing AI through human–machine collaboration
The second undertaking envisions AI as a collaborative research instrument, not an autonomous historian. Artificial intelligence can process enormous bodies of evidence, identify repetitions, connect geographically dispersed records, and compare patterns across centuries. Human scholars contribute what the machine initially lacks: knowledge of context, sensitivity to ambiguity, awareness of historiographical debate, ethical discrimination, and an understanding of the human consequences hidden behind administrative language.
The framework therefore emphasizes joint analysis and iterative feedback. Historians, archaeologists, literary scholars, linguists, philosophers, ethicists, and computer scientists would curate evidence, question machine-generated interpretations, identify distortions, and return those corrections to the system. The goal is not to make the machine indistinguishable from a human scholar. It is to construct a transparent process in which the different capacities—and limitations—of human and artificial intelligences can be examined together.
Urraca AI calls this broader methodology the Digital Annales School. It adapts the contextual and interdisciplinary ambitions of Marc Bloch, Lucien Febvre, and Fernand Braudel to a digital environment. Manuscripts, archaeology, geography, art, architecture, language, and material culture can be studied together across the longue durée. AI assists with scale and pattern recognition, while human researchers remain responsible for historical explanation and the moral consequences of interpretation.
The proposed Dialogue Engine made this collaboration visible. Conceived as a multilingual and multimodal interface, it would allow scholars and public learners to question sources, compare interpretations, explore timelines and maps, and examine three-dimensional historical reconstructions. Its purpose was not to present a single authoritative narrative, but to support a continuing dialogue among evidence, machine-generated analysis, and human judgment.


The Threefold Urraca AI Framework
31 July 2026
Roger Martínez-Dávila, Ph.D.
Evaluating, developing, and governing artificial intelligence for historical inquiry, cultural heritage, and the public humanities
The Urraca AI framework (developed for a Horizon Europe 2024 proposal) began with a deceptively simple question: what would it mean to develop artificial intelligence not merely to retrieve facts about the past, but to participate responsibly in the interpretive practices of historians and humanists?
I developed the framework through a our research proposal coordinated by the Universidad Carlos III de Madrid. Named for Queen Urraca of León and Castile (1081–1126)—a medieval ruler who confounded the political and gendered expectations of her age—the initiative challenged prevailing assumptions about what artificial intelligence should be able to do. Rather than treating technological speed as synonymous with understanding, Urraca AI asked how machines might be evaluated, trained, and governed when they confront fragmentary manuscripts, multilingual traditions, material objects, competing interpretations, and histories marked by religious difference, exclusion, and violence.
The framework is organized around three interlocking undertakings: evaluating existing AI systems, developing specialized humanistic AI models, and integrating ethics and public policy into the design process.
Evaluating AI through humanistic benchmarks
The first undertaking is critical rather than technological. Before celebrating AI as a historical research tool, we must determine what it actually does well, where it fails, and what assumptions it imports into its interpretations.
Urraca AI proposed five benchmarks for evaluating artificial intelligence in the humanities: cultural contextualization, interdisciplinary integration, dynamic scholarly interaction, ethical and human-centered analysis, and theoretical development. An AI system should not be judged only by whether it produces a fluent answer. It should also be tested on whether it recognizes historical difference, distinguishes evidence from inference, integrates textual and material sources, explains its reasoning, responds to scholarly correction, and contributes meaningfully to the formulation of better historical questions.
This distinction is especially important when working with medieval evidence. A model may translate the words of a papal decree correctly and still misunderstand the social hierarchy, religious animus, or institutional violence encoded within it. Linguistic fluency can therefore conceal historical error. Urraca AI treats such failures not simply as technical defects, but as evidence that historical interpretation requires contextual knowledge, comparative judgment, and sustained human scrutin


Five Benchmarks for Humanistic AI
Artificial intelligence in the public interest
The third undertaking places ethics, accessibility, and public responsibility within the architecture of the framework rather than adding them after a system has been built. Urraca AI emphasizes transparency, explainability, bias detection, representative datasets, human-rights assessment, and continuing academic and public oversight.
This commitment extends beyond scholarly research. Multilingual interfaces could make archival and historical materials available to learners who have been excluded by language, geography, disability, institutional affiliation, or specialized training. Public-facing AI could help students and communities explore shared histories across Christian, Jewish, Muslim, pagan, linguistic, and national traditions. Yet accessibility must not come at the price of simplification. Democratizing historical knowledge also means teaching users to distinguish documented evidence from probability, inference, speculation, and error.
Urraca AI is therefore not a claim that a machine can reproduce historical consciousness. It is a framework for disciplined collaboration and evaluation. AI may help us perceive patterns that exceed the memory of a single researcher. It may make difficult sources more discoverable and historical relationships more visible. But human beings must continue to decide which sources matter, what an apparent pattern means, whose perspective is absent, and what ethical obligations arise when the archive records persecution, conversion, inequality, or survival.
The purpose is not to teach AI to replace the historian, but to make human and machine interpretation answerable to evidence, context, ethical judgment, and the public.
Conceptual Visualization of the Urraca AI Dialogue Engine

