
Most SMEs evaluate the wrong half of an AI tool.
The request arrives on Tuesday via Teams, two sentences long: "Can we use that meeting-notes tool for our sessions? It costs 18 francs a month." What happens next usually follows one of two paths. Either someone looks at the vendor's website, finds a trust centre with certification logos and says yes. Or someone sends out a questionnaire with forty items, gets forty answers back three weeks later and nobody reads them.
The outcome is the same either way. The question stays open, once quickly and once slowly.
Yet you can evaluate AI tools without turning it into a project. You do not need the vendor's security report. You need three statements about what your company grants the tool.
Why the evaluation of AI tools rarely gets finished
Expertise is seldom what is missing. Ownership is.
In most of the companies we see, nobody has settled who decides an AI request. IT feels responsible for the technical side, management for the money, data protection for the form, and the person who wants to use the tool waits. Because nobody owns the decision, there are two outcomes, and neither of them is a decision.
The first is the silent yes. Someone pays the 18 francs on their own credit card, files it as an expense, and the tool is in operation without anyone having approved it. The second is the silent no: the request fades away, and the person solves their task anyway, just with a private account. How that path forms, and why a ban tends to entrench it rather than close it, we described in our article on shadow AI in companies.
Neither is a discipline problem. It is the predictable result of an evaluation that has no end. That is why the first lever is not a better questionnaire, but an evaluation that fits into an hour and ends with a decision. For that you need to know what you are actually judging.
What the tool sees
The first statement is the one asked most often and answered cleanly least often. Not "is the vendor secure", but: which data goes into this tool, and what happens to it there.
One distinction helps. Some tools only see what someone actively puts in, a text or an audio file. Others reach into your existing data, the mailbox or the file share. That difference is an order of magnitude, and it determines how much scrutiny the matter deserves.
The legal position is clearer than many assume. On 8 May 2025 the Swiss Federal Data Protection and Information Commissioner stated that the Data Protection Act is formulated in a technology-neutral way and is "consequently also directly applicable to the use of AI-supported data processing". The Commissioner also names what users of such systems must make transparent: the purpose, the mode of operation and the data sources of the AI-based processing. And he names a legal right for users to learn whether their inputs are further processed to improve the self-learning programs or for other purposes.
Translated, that means: you have to be able to answer this question. Not because an auditor asks, but because there are people with a right to the answer. If nobody in the house can say whether the inputs to the transcription service feed into the training of the model, the tool has not been evaluated, however many logos sit on the vendor's page.
Which categories of data have no business being in such a tool is a separate question. We took it apart in which company data may go into ChatGPT. For evaluating a single tool the shorter version is enough: determine the data category, then decide.
What the tool may change
This is the statement missing from most evaluations, and it is the more important one.
A tool that reads can expose you. A tool that writes can keep you busy. "Summarise my meeting" and "create the contact in the CRM, send the summary and set the follow-up deadline" belong in two different risk classes. And the assistants currently growing into existing subscriptions are quietly sliding from reading into writing.
The OWASP project for generative AI lists this point as a risk of its own under the name Excessive Agency (LLM06:2025). The description is sober and usable: "The root cause of Excessive Agency is typically one or more of: excessive functionality; excessive permissions; excessive autonomy."
All three can be asked about without jargon, and OWASP supplies an example for each that you can put to a vendor.
- Too much functionality: the tool is meant to read documents, but the built-in integration also permits modification and deletion.
- Too many permissions: the integration works with an account that may do more than it needs to. A read-only service connects with rights for SELECT, UPDATE, INSERT and DELETE when SELECT would do.
- Too much autonomy: the tool carries out consequential actions without anyone confirming them. In the original text this is the failure to have high-impact actions "independently verify and approve".
The useful part: none of these three points is a property of the vendor. All three are properties of what you grant the tool. That is good news, because this is exactly where you have your hand on the matter, even with a vendor you will hardly ever negotiate with.
In our experience the evaluation often shrinks by itself at this point. Many tools a company hesitates over for weeks are not permitted to change anything at all. And some that get waved through in ten minutes may do more than the house realises. How injected text then exploits those rights, without the model having to be hacked, we described in prompt injection explained. The conclusion there is the same as here: the risk sits in the permissions.
Who answers for the result
The third statement is the most uncomfortable, because it calls for a name.
An AI tool produces results that get used further in the business. Meeting notes become the basis of a commitment, a pre-selection of applications becomes an invitation. If the result is wrong, the vendor does not carry that, and the tool certainly does not.
Legally the responsibility stays with you in any case. Article 9 paragraph 2 of the Data Protection Act sets out a duty that is often read past: the controller must in particular satisfy itself that the processor is able to guarantee data security. Satisfy itself, not agree with. A signature under a contract is the precondition for that, not the fulfilment of it. Which contract points count, and why the account tier often decides more than the price plan, is in the data processing agreement for AI tools.
One special case deserves attention, because it is easily overlooked: tools that speak or write to customers rather than only working in-house. In the same communication the Commissioner names a legal right for the people concerned to learn whether they are speaking or corresponding with a machine. Whoever puts a reply assistant on the contact form is deciding about data flow and permissions and is additionally taking on a disclosure duty. That is a point to settle before approval, not afterwards.
For evaluating a tool this means in practice: there is a person who answers for the result before it is used further, and that person is named. A person with a name, not a department. If nobody can be found to take that on for a tool, then that is already the result of the evaluation.
The evaluation that fits the tool
Now the opening question can be answered. How deep the check must go depends on the three statements.
If the tool only sees what someone pastes in, and may change nothing, it is a short matter: set the data category, company account instead of private account, name a person, done. If it reaches into your existing data, permissions and the contract come on top, and the evaluation takes a serious but manageable amount of time. If it may change or trigger something, that is the case that deserves work, specifically on the question of which actions need a confirmation.
Without that grading you check everything to the same depth, and in practice that means everything too superficially or nothing to the end. This grading is exactly what we build in mandates for AI governance and safe AI adoption, usually as an extension of what already exists in-house for suppliers and permissions, not as a second rulebook beside it. If you do not yet know which tools are in use in the house at all, it starts with an AI inventory.
Just as useful is the list of what you may leave out. The certificate on the vendor page tells you that someone there passed an audit, not which rights your account hands out. The model comparison tells you which tool writes better prose, not who answers for the result. Neither is wrong, it just answers none of the three questions. In our experience the greater part of the time in tool evaluations goes into material that changes nothing about the decision.
One side effect makes the effort worthwhile regardless of security. An approval with no expiry date rarely gets looked at again. We regularly find AI subscriptions that were approved once, are by now being paid for twice because the function comes included in an existing subscription, and that nobody misses. Tie the approval to the renewal and you check the risk again while tidying up the invoice.
The request from Tuesday
Take the request sitting open on your desk right now. Not the big platform question, the one with the 18 francs.
Three sentences are enough for the decision. This tool sees these and those data. It may change this and that, nothing else. This person answers for what comes out. If you can write all three sentences, you have evaluated the tool. If one is missing, you now know which, and that nobody outside the house will fill that gap for you.
Whether this becomes a rule for all future requests is the next question, and it is shorter to answer than most expect (do we need an AI policy). If you want to set the path up cleanly for the first few tools, talk to us. A first conversation is without obligation.
Which of the tools already running at your company could nobody tell you today what it is allowed to change?
Frequently asked questions
How do you evaluate an AI tool in a company?
Through three statements: which data the tool sees, what it may change or trigger, and which named person answers for its result. All three describe the access your company grants, not the properties of the vendor. They also determine how deep the check needs to go.
Does every AI tool need a security review?
Not to the same depth. A tool that only sees what someone pastes in, and may change nothing, is decided in fifteen minutes: data category, company account, responsible person. Tools reaching into the mailbox or file share need permissions and a contract, and tools with write access deserve the most attention.
What questions should you ask an AI vendor?
Whether inputs are further processed to improve the models, which functions the integration brings beyond the intended purpose, and which actions the tool performs without confirmation. OWASP lists excessive functionality, excessive permissions and excessive autonomy as the root causes of excessive agency.




