It will have escaped no one’s notice that we are in the midst of a technological revolution. Wherever AI combined with LLMs can support us, applications are emerging that attempt to solve problems or, at the very least, assist us with the help of AI. From now on, I will refer to these applications as AI apps.
AI apps are now being developed for all kinds of problems, such as automating repetitive tasks, answering customer questions, and supporting decision-making within companies. In everyday life too, we are seeing more and more applications that help with writing, translation, navigation, searching, and organizing. Wherever speed, convenience, and efficiency are important, AI apps are appearing as a possible solution.
Non-deterministic
All these AI apps share a fundamental problem: by definition, they are non-deterministic. If you ask an AI app the same question today as you did yesterday, the answer you receive will not necessarily be the same; the result may be slightly different at any moment.
Many teams within organizations are now building their own AI apps to solve specific problems, often without fully realizing it. For example, prompts are shared within and between teams to move from requirements to test cases. In practice, this is very simple: you open, for example, Microsoft Copilot, load your requirements document into it, and use a shared prompt to ask it to generate test cases. The quality of current models is so good that the resulting test cases can serve perfectly well as a basis for the definitive test cases that are subsequently refined.
A familiar thought among testing professionals
When test professionals assess a new AI app from the market, this familiar thought quickly arises. This belief also appears regularly on LinkedIn. Yet the visible interface tells only a small part of the story. An AI app is much more than what you see on the outside. It is a combination of system prompts, AI agents, a careful selection of the right LLM(s) for the task, settings such as the model’s temperature, and the way in which the entire chain is controlled and monitored. What this does make clear is that companies developing AI apps need a method for explicitly benchmarking the quality of AI apps against that quick “built-in-a-few-days” solution.
Evaluaite
To provide a solution precisely for these challenges, QA Company developed Evaluaite. Evaluaite is a method for testing AI apps based on clear, predefined quality criteria. QA Company considers Evaluaite to be the way in which AI apps should be tested systematically and consistently.
Zo werkt Evaluaite
1
Determine the metrics to use: The relevant quality measures for the AI app are defined, such as correctness, consistency, safety, or responsiveness. Evaluaite supports both traditional metrics and so-called “LLM-as-a-judge” metrics, in which one LLM is used to assess the output of another LLM.
2
Set thresholds: A threshold value is determined for each metric; this is the minimum level the AI app must achieve to be considered “good enough.”
3
Run tests: The AI app is tested using these metrics. Evaluaite can be used with well-known LLM evaluation tools such as DeepEval and Promptfoo, although QA Company currently has a clear preference for LangWatch, a platform for LLM evaluation, testing, and observability.
4
Assess the scores: The results are assessed against the established thresholds, making it immediately clear where the app meets the requirements and where it does not.
5
Improve through Evaluation-Driven Development: Based on the scores, the AI app is improved in a targeted way, with the aim of getting all scores structurally above the thresholds and keeping them there.
Evaluaite transforms a non-deterministic system into a manageable, quasi-deterministic system — through metrics, thresholds, testing and continuous improvement.
Predictable quality, not a predictable answer
What Evaluaite really helps teams do is turn a non-deterministic system into a manageable, quasi-deterministic system: the output varies, but the quality remains above the desired level. Evaluaite does not do this by requiring the same question to always produce exactly the same answer. Instead, it requires the AI app to consistently achieve a score above the established threshold for every metric. The content of the answers may vary, but the quality becomes predictable and repeatable.
By continuously measuring the quality of your AI app with Evaluaite, your AI app distinguishes itself from the “vibe-coded” AI apps mentioned earlier. Let us return to the example of an AI app that generates test cases from requirements. As a customer, I am willing to pay for demonstrably receiving, for example, 95% usable test cases, whereas an internally “vibe-coded” solution within my organization might achieve no more than 50% usable test cases.
That vibe-coded AI app may still be valuable for teams, but professionally developed AI apps tested with Evaluaite systematically achieve much higher scores. In my view, that difference in percentages is precisely why customers should choose a professionally developed AI app.
A concrete competitive advantage
If you use Evaluaite to test your AI apps, you gain not only clear insight into the quality of your application, but also a tangible competitive advantage. If you have an AI app that you want to bring to market commercially, Evaluaite offers a clear way to differentiate your product from the competition. You can do this by sharing your scores for the various metrics openly and transparently with the outside world.
Conclusion
AI apps are everywhere, but without structural evaluation they remain, in essence, non-deterministic black boxes whose quality is difficult to prove. By working with Evaluaite, you translate this non-deterministic behavior into predictable quality levels using clear metrics and thresholds. In doing so, you move from “vibe-coded” to a professional, evidence-based development process in which you not only build better AI apps, but also demonstrate why your solution is worth more than the rest of the field.

