
Software & AISeptember 29, 2026
Beyond the chatbot: when a small AI model has to decide, not write
Many business systems do not need to generate text at every step. They need to make a decision. I3K is exploring local Decision Models with EuLLM: a demo at roughly 8 ms per decision shows how small AI models can drive agents, RAG and business workflows.

Beyond text generation
When people talk about Large Language Models, the first thought goes almost inevitably to ChatGPT and text generation. Questions, answers, documents, summaries, code.
That is only part of the problem.
Many business systems do not need to generate text at every step. They need to make a decision.
· Which tool should an AI agent use?
· Is the information retrieved by a RAG system sufficient?
· Should a request be handled automatically or passed to a human operator?
· Does a security event deserve deeper investigation?
· What is the next step in a workflow?
I3K is exploring this usage pattern through EuLLM, its open source platform for running and specialising models locally.
Roughly 8 milliseconds per decision
One of the most recent demos uses a small Jev-style Decision Model with 2 billion parameters to control the game Snake in real time. The game is the least important part of the experiment.
At every move the model receives the current state, evaluates the available actions and selects the one it considers best. The test runs entirely locally on an NVIDIA RTX 5070 Ti. The latency observed in the demo is in the order of 8 milliseconds per decision.
Alongside the model, a deterministic evaluator was implemented, able to independently identify the best choice for that particular game configuration. In the test shown in the demo, roughly 95% of the model's decisions match the choice indicated by the evaluator.
So this is not a generic accuracy percentage awarded by a second LLM. It is a comparison against deterministic logic specific to the problem.
The figures reported are observations from this demo on this hardware, not certified benchmarks.
Why Snake?
Because it makes visible a concept that is normally hidden inside software. When the model decides to go right, the result of the decision appears immediately on screen.
In a business application the same logic could decide something far less spectacular but decidedly more useful.
· An agent could decide to query a second source before answering.
· A document system could establish that the available evidence is not sufficient.
· An administrative process could choose between automatic processing and escalation to an operator.
· A cybersecurity system could classify an event and choose the next level of analysis.
The model does not necessarily have to produce an articulate answer. It has to choose correctly among a set of actions.
The problem with using huge models for tiny decisions
Many contemporary AI architectures use very large general-purpose models even for relatively simple steps. The model receives the context, reasons, and returns perhaps a single word:
continue or escalateTechnically it works. Architecturally it is not always efficient. Every call can mean network latency, token consumption, variable costs and information transfer to external infrastructure. If the workflow contains dozens of steps, the problem multiplies. A small, specialised model can be an interesting alternative. It does not have to know everything. It has to know its own task very well.
From generative AI to a decision layer
One possible future architecture for enterprise AI separates functions that today are entrusted to the same model. The more powerful models can keep handling problems that require genuine generative capability and complex reasoning. Smaller models can instead handle classification, routing and repetitive decisions. In between sits the application.
It is a structure that is particularly interesting for agents and RAG systems, where the model is called many times during the execution of a single request.
Consider an enterprise RAG. The first search returns four documents. The system can immediately generate an answer from them, or it can ask a small decision model: is the retrieved information sufficient?
If the answer is no, it can start a second retrieval. The same model could determine which archive to query, or whether to ask the user for clarification. The main generative model is involved only when it is genuinely needed.
Local AI also means controlling the infrastructure
The EuLLM demo uses no external API. Model, data and inference all stay on the local machine. This becomes particularly important in enterprise environments.
The requirement is not only about privacy. A local platform lets you know the hardware running the model, control versions, set update policies and keep infrastructure behaviour predictable.
Latency does not depend on the internet connection. An outage at an external provider does not necessarily block the process. And the cost of inference does not grow linearly with every single call to a commercial API.
EuLLM: Engine, Forge and Hub
EuLLM was designed with an architecture broader than a simple inference server.
Engine
The runtime that makes it possible to run models locally.Forge
Dedicated to verticalisation and teacher-student processes, aiming to turn larger generalist models into smaller, specialised ones.Hub
Conceived as a European layer for distribution, cataloguing and model information. Decision Models are therefore a particularly natural fit for the whole ecosystem. A large model can contribute to producing a specialised student. Forge manages the specialisation process. Engine runs the result on the customer's infrastructure. The code is on GitHub.Not everything has to be a chatbot
This is probably the most interesting conclusion from the demo. In recent years we have associated LLMs with the conversational interface.
But a language model can become an internal component of a piece of software without the user ever having to talk to it. It can observe. Interpret. Decide. And hand execution to the next component.
In many cases, the future of enterprise AI may be less visible than we imagine. Fewer chat windows. More specialised models inside processes.
The game of Snake only serves to show them at work.
For anyone looking to bring this kind of infrastructure in-house, the appliance range is described on the I3K Local AI page.
Interested?
Contact us to receive a personalized quote.
All articles
I3K Technologies Srl (single-member company) — i3k.eu