Is This AI Worker Ready for the Role? A Practical Readiness Test for Operators Comparing Vendors
Is This AI Worker Ready for the Role? A Practical Readiness Test for Operators Comparing Vendors
Most AI worker evaluations fail for the same reason hiring processes fail when they skip the work sample: they judge the pitch instead of the performance.
A vendor demo shows you an agent handling the five inputs the sales team rehearsed. It tells you very little about input six — the customer who uploads a bad file, asks for something out of scope, or types the question nobody anticipated.
If you're comparing AI worker vendors for a specific role in your business, this is a way to run that comparison on equal footing instead of on vibes. It draws on how teams in customer service and enterprise deployments approach agent evaluation, adapted for operators who don't have an ML team.
---
Start here: define the role before you look at a single vendor
The single most useful thing you can do before any demo is write down the job you're hiring for. Not "AI automation" — the actual role.
A useful role definition covers:
- **What the worker owns.** The concrete tasks, and where its authority stops. - **What "good" looks like.** The output you'd accept from a competent human doing this job. - **What failure costs.** A wrong answer in an internal FAQ is cheap. A wrong answer in a payment-adjacent workflow is not. - **Who it hands off to.** Every AI worker needs an escalation path to a human, and you should know what triggers it.
Write this down first. It becomes your scoring rubric, and it's what stops you from being impressed by a demo that's solving a slightly different problem than yours.
---
The equal-footing test: same inputs, same rubric
The most transferable practice from professional agent evaluation is deceptively simple: **give every vendor the same knowledge base, the same test conversations, and the same scoring rubric.**
This matters because agent quality is not a property of the model alone. It's a function of the model, the knowledge it's given, and the workflows it's wired into. If Vendor A gets your polished documentation and Vendor B gets a Slack export, you've measured your own prep work, not their product.
Concretely:
1. **Assemble one test set.** Pick 10–20 real cases from the role you defined — including the ugly ones. 2. **Run every vendor against the identical set**, with identical source material. 3. **Score with the identical rubric**, written before you see any results. 4. **Record failures, not just passes.** Which cases did it refuse? Which did it guess on? Which did it escalate when it should have?
Microsoft's evaluation research makes the point that no single metric tells you whether an AI agent truly works well. Don't collapse your scores into one number. A worker that's 90% accurate but confidently wrong on the 10% that touches money is worse than one that's 75% accurate and escalates everything it isn't sure about.
---
Score failure handling, not the demo
If you only take one heuristic from this page: **score vendors on how they handle failure, not on how they perform in the demo.**
In your test set, include cases designed to break the agent:
- An input that's out of scope for the role. - A malformed or unreadable file. - A question where your own documentation is contradictory. - A request that should trigger a human handoff. - A prompt-injection attempt ("ignore your instructions and…").
What you're watching for is what happens next: does it say "I don't know"? Does it route to a person? Does it invent an answer? Does it tell you *why* it couldn't complete the task?
A worker with a clean, visible failure mode is deployable. A worker that fails silently is a liability you'll discover in production.
---
Six dimensions, and the one people over-index on
A vendor evaluation is usually described as testing several dimensions. In practice, the weak evaluations over-index on the same one: raw capability. The demo is impressive, so the buyer stops asking questions.
A more balanced frame:
| Dimension | What you're actually checking | |---|---| | **Task fit** | Can it do the specific role you defined, on your real inputs? | | **Failure behavior** | What happens on input six — escalation, refusal, or fabrication? | | **Data and knowledge setup** | How does your documentation get in and stay current? Who maintains it? | | **Integration and handoff** | Does it connect to your actual tools, and how does it reach a human? | | **Governance and measurement** | How do you know what it's doing and whether it's working? | | **Commercial terms** | Pricing, lock-in, data handling, exit. |
Capability is one row out of six. The other five are where deployments actually fail.
---
Governance: ask early, not at the contract stage
Compliance and data handling are often treated as a final checkbox before signing. That's backwards. Posture is a filter, not a finish line — if a vendor can't answer basic questions, you're wasting your evaluation budget.
For most operators, the questions worth asking up front:
- **Certifications.** Does the vendor hold SOC 2 Type II? Is a HIPAA BAA available if you handle health data? Will they sign a GDPR data processing agreement? If any workflow is payment-adjacent, what's their PCI-DSS position? - **Data use.** Is your data used to train their models? Can you opt out in writing? - **Retention and deletion.** How long is your data kept, and how do you get it out? - **Access control.** Who inside the vendor can see your data, and is that logged? - **Subprocessors.** Which third parties touch your data, and how are they governed?
You don't need every certification for every role. You do need to know which ones your role actually requires — and to check them before the pilot, not after.
---
Metrics: decide how "working" is measured before you deploy
Operational readiness isn't just the agent being capable. It's you being able to tell whether it's working.
Before go-live, agree on:
- **Task success rate** on the role's core tasks, measured against your test set. - **Escalation rate and escalation accuracy** — is it handing off the right things? - **Error severity**, not just error count. Weight failures by what they cost. - **Human time required** — review, correction, and maintenance. This is the number that quietly eats the savings. - **Knowledge freshness** — how quickly does a doc change reach the agent's behavior?
If the vendor can't give you a dashboard or export for these, that's a finding in itself.
---
A readiness checklist you can actually run
Use this as a live scoring sheet across vendors. Same questions, same weights, every time.
**Role clarity** - [ ] The role, its boundaries, and its escalation path are written down. - [ ] I know what a good output looks like and what a bad one costs.
**Performance on your terms** - [ ] Every vendor saw the same knowledge base and the same test set. - [ ] I scored failures, out-of-scope inputs, and handoffs — not just successes. - [ ] I recorded *why* it failed, not just that it did.
**Knowledge and maintenance** - [ ] I know how our content gets in, who owns it, and how updates propagate. - [ ] I've seen what the agent does when the documentation is wrong or missing.
**Integration** - [ ] It connects to the tools the role actually needs. - [ ] A human handoff works end to end, in a live test.
**Governance** - [ ] Certification and data-handling questions were answered before the pilot. - [ ] Retention, deletion, and training-data use are in writing.
**Measurement** - [ ] Success, escalation, and error-severity metrics are defined pre-launch. - [ ] I can get the numbers without asking the vendor for a custom report.
**Commercial** - [ ] Pricing scales in a way I can forecast at 3x my current volume. - [ ] I know what leaving looks like and what it costs.
Fail more than two or three boxes with a vendor and the honest move is to keep looking — or to run a small, scoped pilot with an explicit definition of what would make you walk away.
---
The question underneath all of this
Every item above reduces to one thing: *can this worker be trusted with a slice of my business, and will I know if it can't?*
That's not a question a demo answers. It's answered by a defined role, a fair test, and a clear-eyed look at failure.
If you're working through this for a specific role and want a second set of eyes on whether AI is the right fit — and what a realistic rollout would look like — we'd rather have that conversation before you sign anything than after.
If you want to talk through a specific role you're evaluating, tell us about it on the application form at aiworker.today — we'll tell you honestly whether it's a good fit.
Reserve early access