The test a wrapper cannot pass
Ask any AI vendor for a measured accuracy figure on your own data. Watch what happens.
The market is full of companies whose product is a prompt and a login. This is not a moral failing, and some of them are genuinely useful. It is a problem only because a buyer cannot tell them apart from a company that has actually built something, and both cost about the same in the first meeting.
There is one question that separates them cleanly. Ask for a measured accuracy or error figure on a sample of your own documents, your own images, your own recordings. Not a case study. Not a benchmark. Your data, a held-out set, a number you can re-run after the vendor leaves.
A demo is a claim about their data. An eval set is a claim about yours.
A company that has built a system will find this a normal request, because building the eval set is most of the work and they already have the harness. A company that has built a prompt will offer you a pilot instead, which is a way of saying that you should pay to find out.
We hold ourselves to this and it is written into how we quote. Every system with a model in it ships with an eval set and a figure. Sometimes the figure is worse than the customer hoped, and that conversation is uncomfortable, and it is still much cheaper than the one that happens after deployment.
The second question, if you want one: where does the data go, and what happens when you stop paying. A system that runs on your own keys and your own infrastructure has a boring answer. A wrapper usually does not.