
In this article, we highlight five areas where data workflows typically stall – and the specific solutions we have built based on these principles.

Lots of data. In catalogs, digital libraries, spreadsheets, local databases, storage systems, and often still in paper cards and binders.
Each source was created at a different time, for a different purpose, and according to different rules. The real work begins the moment you need to find connections, check thousands of records, or process a large volume of material. Switching between systems, exporting, manual verification, retyping data, and finding ways to keep the entire process under control.
Over more than fifteen years of working with libraries and other memory institutions, we have encountered similar situations repeatedly. The principle is usually the same: the system handles the repeatable part of the process, while the expert handles exceptions, context, and decisions that require their experience.
The principles we have learned in this environment can be applied wherever there is a need to work more effectively with large datasets, documents, image materials, or specialized know-how.
In this article, we highlight five areas where data workflows typically stall – and the specific solutions we have built based on these principles.
Catalog. Digital library. Excel. Local database. Registration system. Image repository. Paper documentation. In one system, you find out that a document exists. In another, you verify its digitization. Elsewhere, you look for metadata, an identifier, or information about what has happened to the record. Answering a single question can thus involve piecing together information from several sources.
That is why we start by mapping where the necessary data originates, what its structure and quality are, and how it can be accessed securely. A common working layer can then be built over the relevant sources for searching, filtering, checking, and connecting the dots. Depending on the situation, it works with one-time exports, regularly updated data, or direct integration.
It was precisely from this need that Anakon was born.

Currently, Anakon works with data formats such as MARC 21, MODS, and Kramerius digital library indexes. It also allows for comparing data between the catalog and Kramerius. Specific solutions can be extended to other formats and sources according to project needs, such as UNIMARC. The key principle is to enable experts to analyze, verify, and link data regardless of the system in which it was originally created.
Imagine, for example, checking a collection before the next step of digitization. You need to find records with a specific combination of fields, identify where something is missing, or compare information from several sources. Anakon allows you to prepare such a query and run it repeatedly over your data.
This provides the expert with a tool to work with data the institution already possesses, allowing them to view it within a broader context.
With ten images, many things can be fixed manually. With tens or hundreds of thousands of pages, every click starts to have a cost. This is clearly visible when processing images in a digitization pipeline. Pages need to be identified, properly cropped, rotated, and checked. A large portion of the material goes through a very similar process.
This is exactly where there is room for automation. Cropilot uses an AI model to evaluate image data, recognize pages, and prepare them for cropping and rotation. A staff member can then check the result in an editing interface and focus their attention on cases that require intervention.
An expert's time is thus focused where it brings the most value – on quality control and handling exceptions.

Historical prints, newspapers, card catalogs, or personal documents can have different formats, specifics, and processing rules. That is why it is possible to work with custom training data and adapt the model to specific material.
In larger operations, Cropilot can fit into an existing digitization workflow via an integration interface. For smaller workplaces, it can function as a standalone tool. In this case, automation handles the technical step while simultaneously freeing up the capacity of people who can then focus on the more specialized parts of the digitization process.
Cataloging historical documents requires experience, knowledge of rules, and the ability to work with context. A cataloger assesses the title page, colophon, authorship, publication details, authorities, period language, and ambiguous information. Alongside this expert work, they spend time gathering background materials, transcribing data, and searching for potential matches in other databases.
We are currently testing ways to speed up this preparatory part using AI cataloging.

The system extracts bibliographic data from photos of the title page, colophon, cover, and other key parts of the document, processes it, and prepares a draft record in MARC 21 format. The process includes working with authority databases and existing catalog records. The system prepares potential options and supporting materials that the cataloger can use when making decisions. Every piece of data can be checked, edited, or supplemented.
The role of AI in this workflow is clearly defined. It handles the mechanical part of the preparation and creates a high-quality starting point. Professional assessment and responsibility remain with the cataloger.
We are currently developing and testing the tool in collaboration with catalogers from the Moravian Library. Their experience helps us fine-tune the interface and individual steps to match real-world cataloging practices. At the same time, we are identifying which types of documents and situations benefit most from assisted cataloging.
Institutions often use standard central systems that handle their core agenda well. Alongside these, additional institution-specific needs gradually emerge, such as:
In such a situation, it makes sense to create a separate layer focused specifically on the task at hand to save time, eliminate manual errors, or address capacity shortages.
First, we determine the integration options. Depending on the system, this could involve an API, a database, or data exports. We then build a workspace on top of these where the user can set parameters, run tasks, monitor progress, and work with the results, or have the process automated using artificial intelligence.
This principle is well illustrated by Divadelní fundus JAMU. The collection managers work with data in Airtable, which they are already familiar with from their daily operations. We built a web application on top of this data where students and staff can browse costumes, props, equipment, or musical instruments, filter them by available parameters, and create reservations.
One layer serves the people who manage the inventory, while the other serves those who need to quickly find and use items from it. All the while, the data remains part of an established working environment.

A similar principle is used by the Kramerius statistics module. It creates clear analytical views from digital library operational data, allowing traffic to be tracked by document type, license, year of publication, or author, for example. This creates a specialized tool for the institution's specific needs on top of the existing system.
The same logic is applied to other metadata tools, digitization workflows, and solutions for digital libraries. The main system continues to handle the tasks it was designed for. The specialized layer addresses a specific process that the user needs to complete faster, more clearly, or in greater volume.
Researchers, students, or visitors usually come with a topic, location, author, or a rough idea of what they are looking for. The internal structure of a collection works with precise terminology, cataloging rules, and metadata. Furthermore, historical documents introduce archaic language, different spelling, and varying OCR quality.
Effective accessibility therefore connects two different perspectives:
Balancing both needs within a single environment is a key part of the design process, especially for systems used by both professionals and the general public. We build presentation and search layers on top of the collections, tailored to who is using the content and what they need from it.
Individual collections may require completely different ways of discovery. For Moll's collection of rare maps the focus is on the image, geographical context, and the ability to study maps and vedute in high resolution. Users can browse the collection by territory, author, or period, treating it as a digital research environment.

Paměť novin addresses a different problem with historical newspapers. Semantic search allows users to ask questions in natural language, work with the meaning of the text, and locate the source documents the answer is based on. For historical materials, this is also useful where archaic terminology or OCR quality complicates traditional searching.

For the Kramerius digital library we have been designing and developing a web client for a long time – from designing new features and development (primarily frontend) to integrating artificial intelligence, which takes working with data to a whole new level.

Technology shortens the path between the person and the content. Users can find relevant documents faster, discover related material, and continue their research in a direction that matches their actual interests.
Cropilot, Anakon, AI cataloging, Divadelní fundus JAMU, tools for digital libraries and presentation projects for individual collections were born out of different situations.
The way we think about them, however, is the same. We start with a specific process and the people who work with it every day. We identify which steps are repetitive, where errors occur, what takes the most time, and which decisions require expert context. We work with real data, including historical layers, exceptions, and local specifics. We always focus on what is truly unique to the institution in question.
The result we aim for remains the same – less manual work, more control, and data, archives, and collections that are actually usable.
To start a conversation with us, just describe a specific process that currently takes too much time, ends up in spreadsheets, or regularly hits the limits of your existing tools. We will see if one of our products, a specialized add-on, data processing, or a quick pilot project using a real data sample makes sense for you.
Discuss your institution's specific challenge with us. Email our CEO, Jan Rychtář, at jan.rychtar@trinera.cz or use the form below.




