How to Tell If an AI Worker Is Ready for a Specific Role (Before You Sign)

Most AI vendor evaluations follow a familiar arc: request a demo, watch it go well, call a reference, sign. That process was built for software that reliably does exactly what it was configured to do. An AI worker is a different kind of purchase. Its quality shows up in the cases nobody rehearsed — the odd input, the ambiguous request, the thing just outside its scope.

This guide is a practical method for answering one question: **is this AI worker actually ready to do *this* job in *my* business?** Not "is the vendor impressive," and not "is the demo smooth," but whether the worker is ready for a defined role with defined stakes.

Start with the role, not the vendor

The most common evaluation mistake is starting with a product tour. Before you talk to anyone, write down the role you're hiring for. Not the technology — the role. Treat it like a job description:

- **What the worker owns.** Which tasks, which queue, which inbox, which workflow step. - **What "good" means.** The quality bar a competent human would be held to. - **Volume and variability.** How many interactions, how varied, how much of it is genuinely routine. - **Blast radius.** What happens when it gets something wrong — a mildly annoying reply, or a compliance problem?

A worker that's ready for high-volume, low-stakes triage is not automatically ready to draft contract language or handle distressed customers. Stakes should set the bar for everything that follows.

The six-part readiness test

The research on agentic evaluation converges on a useful frame: a complete evaluation tests several dimensions, and weak evaluations over-index on the first one — the demo. Here's the frame, adapted to hiring an AI worker for a role.

1. Capability — can it do the job at all?

Ask for your own scenario, not the vendor's showreel. A polished demo proves the worker can handle the handful of inputs the sales team rehearsed. It tells you nothing about input six.

Better: give every vendor you're considering the *same* test set. Same knowledge base, same sample conversations or tickets, same scoring rubric. Comparisons across vendors are only meaningful when the inputs are held constant.

2. Failure behavior — what happens when it's wrong?

This is where readiness is actually decided, and where most evaluations go easy. Ask directly:

- What does the worker do when it's not confident? - How does it escalate, and to whom? - What does a bad output look like, and how would I know? - How do I see what it did and why?

Score vendors on how they handle failure, not on their best-case output. A worker that fails loudly and hands off cleanly is often worth more than one that fails quietly and confidently.

Remember that no single metric tells you whether an AI worker truly works — accuracy alone hides the shape and cost of its mistakes.

3. Fit with your data and tools

A worker is only as ready as the systems it can reach. Check whether it can actually access your knowledge sources, your CRM, your ticket system, your internal docs — and whether the answers it gives trace back to those sources. Data quality and access are readiness questions, not implementation details to sort out later.

4. Governance and oversight

Two things to pin down before you commit:

- **Who manages the inputs.** Clear guidelines for the data sources the worker draws on, and who owns keeping them accurate. - **How performance is measured.** Agree up front on what you'll track to show the worker is delivering value — not as a vanity dashboard, but as the thing that tells you whether to expand or pull back its scope.

If a vendor can't describe either of these concretely, that's a signal.

5. Onboarding cost — yours, not theirs

The real effort of standing up an AI worker is usually on your side: cleaning up knowledge, mapping the workflow, defining escalations, deciding who supervises. Get an honest estimate of that work in hours and in your team's attention. A worker with modest capability and light onboarding is sometimes the better hire.

6. Ability to grow the role

Roles expand. Ask how the worker's scope gets widened — what it takes to add a task, adjust its judgment, or raise its autonomy as trust builds. A worker you can't safely expand is a worker you'll outgrow.

A checklist you can run in one pass

Copy this into your notes and score each vendor against it.

- [ ] I wrote the role description before the first vendor call - [ ] Every vendor got the same test inputs and the same rubric - [ ] I tested at least one case the vendor didn't prepare for - [ ] I know what happens when the worker is uncertain - [ ] I know how errors surface and who catches them - [ ] I've seen a trace of how it reached an answer - [ ] I confirmed access to my actual data sources and tools - [ ] Someone is named as owner of its knowledge and data - [ ] We agreed on how performance will be measured - [ ] I have an hours estimate for my side of onboarding - [ ] I know how its scope expands as it earns trust - [ ] I've ranked every vendor against this list, not against the demo

If you can tick most of these for a given vendor — and be honest about the ones you can't — you're in a much better position than a reference call would have put you in.

A note on where to start

You don't need a big enterprise budget to benefit from AI workers, and you don't need to evaluate five vendors to learn what matters. Starting with one role, one clearly defined job, and a small test set will teach you more about readiness than a dozen polished demos. The operators who get this right tend to start small, prove value, and grow the role from there.

If you have a specific role in mind and want to see how we'd approach it, you can tell us about it on the application form at aiworker.today.

Reserve early access