5. Data preparation

Once the company knowledge has been fully collected, the step begins that will largely determine how powerful, reliable, and useful the later AI will actually be:

the data preparation.

In this phase, a large number of individual files, texts, media, system extracts, and experience reports are turned for the first time into a coherent, structured, and professionally robust knowledge system.

This is where Vimmera Cortex comes into play. Our knowledge base with a very special data structure.

The collected information initially exists in very different qualities. Often there are multiple versions of the same content, contradictory statements, outdated versions, informal working methods, inconsistent terminology, or missing links between related topics. Much of it has grown historically, is not maintained consistently, or is designed for different target groups and situations.

A particularly important part of data preparation is the standardization and structuring of your company’s language. Technical terms, internal designations, product names, abbreviations, process names, and typical formulations are prepared and linked in such a way that the AI understands and correctly uses the “language” of your organization, your employees, and your customers. Only then can requests be interpreted correctly, content assigned unambiguously, and misunderstandings avoided. Without this linguistic clarification, many questions lead nowhere or are answered inaccurately, incorrectly, or incompletely.

If an AI were to work directly on the unstructured raw material (that is, in simple files and documents), it could find text passages and content, but it could not reliably decide which information is valid, current, and technically correct. It becomes even more difficult with complex questions in which several areas of knowledge, rules, products, or processes interact. Without preparation, the AI lacks the necessary clarity to recognize relationships with confidence, argue without contradictions, or provide complete, reliable answers.

Data preparation is therefore the step in which a solid knowledge base emerges from unstructured information. A knowledge base that is technically consistent, speaks your company’s language, and creates the prerequisite for AI not only to find, but to understand, evaluate, and provide meaningful support.

From document to knowledge

Another, often underestimated aspect is the quality of later answers when dealing with large volumes of data. The more information is placed into a (vector) database and the more data an AI has to consider at the same time, the greater the fuzziness of the matches and outputs becomes. This is a well-known problem in many systems:

As the amount of data grows, precision decreases, answers become more general, less accurate, or mix content that is technically unrelated.

The reason is simple: in classic systems, thousands of pages, PDFs, protocols, or manuals sit side by side. For the AI, these are equivalent text sources. Relevance, validity, responsibility, product reference, or technical context are not clearly represented in them.

The data preparation for Vimmera Cortex solves this problem fundamentally by treating information not as documents, but as knowledge. Content is broken down into its technical components: into facts, rules, terms, questions, answers, relationships, dependencies, validity periods, variants, and links. The individual document loses its role as a knowledge container; it is “dissolved.” What remains is the knowledge it contains in structured, abstract form.

The best way to imagine this process and storage is how we humans store memories. After all, we don’t store a PDF of an assembly manual in our brain when we want to remember how our television works.

The previously collected knowledge is clearly assigned to products, services, processes, functions, categories, or application scenarios. As a result, the total amount of data no longer matters. The AI does not access a large, fuzzy mass of text, but exactly the knowledge building blocks that are relevant to the respective request.

If desired, the original documents remain available. You can still look up where something is written, find documents in a targeted way, or search text passages.

What matters, however, is:

The AI from Vimmera AI does not work with documents – it works with knowledge from Vimmera Cortex, that is, with your knowledge.

Linking instead of filing

In data preparation, content is not only cleaned up and standardized, but above all linked with one another. Different sources of knowledge, such as documents, processes, products, rules, experience-based knowledge, and system data, are put into relation. This creates a knowledge network in which the AI knows not only individual facts, but also their meaning, validity, and context.

Among other things, it is defined:

  • which documents belong to which processes
  • which products, item numbers, variants, and rules belong together
  • which functions are intended for what and when they are not useful
  • which exceptions, alternatives, or dependencies exist
  • which information should be considered together in which situations

Only then can the AI later not only find information, but also classify, combine, and evaluate it correctly.

Technical logic instead of chance

In the data preparation for Vimmera Cortex, it is also defined which links should be applied automatically in which situations. This creates a technical logic according to which the AI works.

Examples:

  • For a price inquiry, the system automatically recognizes the product, the matching item number, associated variants, valid discount groups, and relevant conditions.
  • For a question about a function, it not only explains what it does, but also whether it is useful in this specific situation, which limitations apply, or which alternatives would be better suited.
  • For service requests, the device, error code, known causes, suitable spare parts, and proven solution steps are automatically linked with one another.
  • For process questions, responsibilities, forms, guidelines, and dependencies are considered simultaneously.

Such answers are only possible when knowledge has been structured, interconnected, and technically modeled.