
There are 132 firms listed under custom AI solutions in this directory, and they are not selling the same thing. One of them will configure a platform you could have bought yourself and bill you for the configuration. Another will build a retrieval layer over documents your own staff cannot find. A third will fine-tune a model on your labelled data. A fourth will take on a research problem where nobody involved knows at the outset whether the thing is possible. Those are four separate trades with four different skill sets, four different cost curves and four different ways of failing, and the phrase on the website is identical in all four cases.
That ambiguity is expensive in one specific direction. Buyers routinely pay for the fourth when the first would have done, because the fourth sounds more serious in a board meeting. Nobody has ever been fired for commissioning original work. The opposite mistake, buying a configuration when the problem genuinely needs new modelling, shows up faster and hurts less, because the vendor runs out of road inside a quarter and the money is still mostly in your account.
What custom AI solutions actually describes
Sort the work by what is new at the end of it, and the four trades separate cleanly.
| What is new at the end | What you are buying | When it is the right answer | What the first phase should hand you |
|---|---|---|---|
| Nothing. Configuration. | Someone else's platform, set up for your process and connected to your systems | A common task where the differentiator is your workflow, not your data | A working configuration in your own tenancy, and the documentation to change it without the vendor |
| A retrieval layer | Your own documents and records made answerable by a general model | The knowledge exists inside the company and staff cannot find it quickly | An evaluation set of real questions with graded answers, before any interface is built |
| A tuned model | A general model adapted to your labels, format or domain language | Output format or domain vocabulary matters more than raw capability, at volume | A measured comparison against the untuned model on the same test set |
| A novel system | Modelling or engineering that does not exist off the shelf at any price | The task is specific to your physics, your sensors or your regulated process | A written feasibility result, including the conditions under which the answer is no |
The right-hand column is the useful part of that table. A firm that cannot describe the first deliverable in those terms has not done the work before, whichever row it claims.
One question sorts most briefs in a single meeting
Ask what already exists that does seventy per cent of this, and then ask why the remaining thirty per cent cannot be a process change.
Both halves matter. The first half catches the expensive mistake, because a serious firm will name the two or three products that cover most of your problem and explain where they stop. A firm that claims nothing exists is either ignorant of its own market or counting on you being ignorant of it. The second half catches the subtler mistake, which is building a system to automate a step that should not happen at all. Plenty of document-extraction projects exist because a form was badly designed eleven years ago, and fixing the form is cheaper than teaching a model to read the mess.
The answer also tells you which of the four trades you are in. If seventy per cent is covered and the gap is integration, you want the AI integration conversation, not a modelling project. If nothing covers it because the data is yours alone, you are in row three or four and the budget conversation changes shape.
Four tests that decide build against configure
- Can the data leave, and does it exist yet? A model cannot be tuned on data that is scattered across four systems under different customer names. Entity resolution comes first, and it is unglamorous work that no proposal wants to lead with. Where the data cannot leave your estate at all, that constraint sets the architecture before anyone chooses a model.
- Is the task repeated enough to amortise? A build has a fixed cost and a running cost. A task that happens two hundred times a day pays for both. A task that happens weekly almost never does, and a human with a good checklist is the correct answer more often than vendors admit.
- Is accuracy limited by the model or by your data? This one is decidable in a fortnight and almost nobody spends the fortnight. If a general model plus good retrieval already gets most of the way there, the ceiling is your data, and a tuned model buys you very little. If the general model fails in a way that more context does not fix, then tuning has somewhere to go.
- Who reviews the output, and what happens when it is wrong? A system whose output a professional signs off can tolerate error rates that an automated system cannot. The review capacity you actually have is a hard constraint on the design, and the AI Risk Management Framework published by NIST is a reasonable common vocabulary for having that argument with a vendor who would rather not have it.
The costs that arrive after the invoice
A subscription hides its maintenance inside the price. A custom build hands you the maintenance.
Three items in particular get left out of first proposals. The first is evaluation: a build without a held-out test set and a scheduled re-run has no way of telling you it has got worse, and it will get worse, because your data changes and the underlying models change under you. The second is model deprecation. Every major provider retires model versions on a published schedule, so a system pinned to one version has a dated expiry and somebody has to own the migration. The third is the person. A custom system needs an owner inside your company who understands what it does, and if that person is not named in the plan the system stops being trusted about eight months in, without anyone deciding to stop trusting it.
None of this argues against building. It argues for pricing the second year at the same time as the first, which is a question worth putting to every firm on your shortlist.

Five things a proposal has to contain
- A test set you can read. Real examples from your business, with the right answer written down, agreed before any building starts. Without it, every later argument about quality is a matter of opinion.
- The comparison against doing nothing clever. What a keyword search, a rules engine or an existing product scores on that same test set. A proposal that will not state the baseline is hiding the size of the win.
- A named stopping condition. The result that would make the firm recommend abandoning the approach. Research work without one becomes an annuity.
- The handover artefacts, listed. Code, weights, prompts, the evaluation set and the scripts that run it, the labelled data and the runbook, each named as something you receive rather than something you can see.
- Second-year cost, in writing. Hosting, inference, re-evaluation, and the migration when a model version is retired.
Who owns the model, the weights and the prompts
This is where custom work differs most sharply from buying a product, and where contracts are most often vague. Four things can be owned separately: the code, the trained weights, the prompts and evaluation sets, and the labelled data. A firm can hand you working code while keeping the weights, or hand you weights you cannot retrain because the labelled data stayed with them. Both arrangements are common and neither is dishonest if it is stated. What is not acceptable is finding out at the end.
Ask the question in its blunt form. If we stopped working with you tomorrow, what do we have, and can we retrain it without you? The answer sorts vendors faster than any case study. Firms that do this work routinely have a ready answer, because every serious client has asked.
How to shortlist without wasting six weeks
Write the scope once, in a page, including the test set and the baseline. Then send the same page to three firms from different rows of that table and compare what comes back. Quotes for custom AI solutions that differ by a factor of five usually mean the firms have read the brief as different trades, which is the most useful thing a quote can tell you.
For sourcing, the custom AI solutions category is the obvious place to start, and it is worth also reading ML and data science for the tuned-model work and generative AI where the output is content rather than a prediction. Firms whose strength is connecting things rather than building them sit in AI integration, and strategy-first engagements sit in AI strategy consulting. The full list is in the directory, and by city if you want people who will come to the site, with San Francisco, New York and Atlanta carrying the deepest listings today.
Two companion pieces are worth the ten minutes before you send anything out. What to buy and how covers the same decision across every category, and the first-time hiring checklist covers the contract questions this piece only touches. If the build you are weighing is a chat interface over your own documents, custom GPT development is the narrower version of row two. If it is a data and modelling problem, AI and ML consulting covers what that engagement contains, and AI development services covers delivery. Larger organisations should read how to choose an enterprise AI company, and anyone weighing whether an ongoing partner is needed at all should read when one person is enough.
If writing the scope is the part you are stuck on, describe the problem here and the directory will return the firms whose listings match it. Agencies that build this kind of work and are not listed yet can add a profile.
Sources and further reading
Listing counts in this article are read from this directory's own category pages and change as agencies are added. For the wider picture on where adoption actually sits, the AI Index from Stanford HAI and the Business Trends and Outlook Survey from the US Census Bureau both publish on it regularly, and the Census survey has the advantage of asking firms directly rather than asking vendors. MIT Sloan Management Review publishes the most consistent coverage of why these projects stall after a pilot. NIST's AI Risk Management Framework is the reference for the review and oversight questions above.