8. Selection of the LLM or LLMs

Once knowledge has been structured, verified, and technically prepared so that it can be found and used precisely (i.e., once your company-specific Vimmera Cortex has been built), the next central step follows:

the selection of the Large Language Models (LLMs).

LLMs are the “thinking machines” behind the AI. They determine how language is understood, how texts are generated, how logically arguments are made, and how flexibly the system can respond.

Vimmera AI does not select these models in a blanket way, but together with you and based on your specific requirements. Because there is no single best model. There are large and small models, very creative and very precise models, fast and resource-efficient models, as well as highly specialized models for specific tasks. Depending on the use case, a single model may make sense, or an interplay of several specialized models.

What happens in this step?

In this step, it is determined which LLMs are used for which tasks. Together, it is decided which capabilities are needed: for example language quality, subject-matter expertise, computing power, speed, data protection, offline capability, or cost control.

Highly capable online models can be used, for example from providers such as OpenAI, Google, or Meta. Offline models can also be used, running on your own servers, in private cloud environments, or even locally on individual computers, such as models from OpenAI, Deepseek, Anthropic, or other providers. The choice depends on which security requirements, data protection regulations, performance goals, or budget constraints apply to your company.

Vimmera AI is not tied to individual manufacturers. All common and powerful systems can be integrated, combined, and orchestrated. This creates AI architectures that fit your organization exactly instead of forcing your organization to adapt to an AI.

Multiple models, one system

In many projects, not just a single LLM is used, but several specialized models. One model may be responsible for the actual subject-matter answers, another for preprocessing inputs, for example to anonymize sensitive data, filter unwanted content, or increase security. Additional models can be used for quality control, structuring outputs, or summarization and further processing.

These models are linked together and orchestrated so that they work as a shared system. For users, only a powerful, consistent AI assistant is visible; in the background, however, several specialized AI instances work together to maximize security, quality, and subject-matter expertise.

How much “knowledge” may the model contribute itself?

A particularly important point in this step is deciding what role the general world knowledge of the LLMs may play. Modern language models bring enormous prior knowledge from their training phase. This knowledge can be helpful, for example for general contexts, language, or logical inferences. In some scenarios, however, it is undesirable because only verified, validated company knowledge may be used.

Together with you, it is therefore determined whether a model may contribute its own knowledge or whether it is deliberately used as an “empty shell” that accesses almost exclusively your company data. This ensures that answers are not based on external, possibly incorrect or unapproved knowledge, but on exactly what your company specifies.

What you gain from this

Through the targeted selection and combination of LLMs, you do not get a standard AI system, but a tailor-made AI architecture. You receive exactly the mix of performance, security, cost control, and subject-matter expertise that fits your requirements.

You retain control over where your data is processed, which models are used, and how strongly external systems are integrated. At the same time, you benefit from state-of-the-art AI technology that can be flexibly expanded, replaced, or adapted as requirements change.