A/B Test a Tool Description
Try a second description for a tool and see which one gets it called, and called successfully, more often. Available on Pro.
Why the description matters
An AI agent never sees your tool's code. It reads the name, the description and the input schema, and decides from those whether this is the tool for the task in front of it. A vague description ("Search products") can make an agent pass the tool over or call it with the wrong input. A precise one ("Search the catalog by keyword. Returns name, price and link for each match.") tells it exactly when to use the tool and what it gets back.
An A/B test answers which wording works better, using calls from real agents on your site.
Starting a test
- Hover the tool on your site's page in the dashboard and click the split icon.
- The panel shows a text box filled with the current description. Rewrite it into the version you want to try.
- Click Start test.
The current description is A, your new one is B.
What happens during the test
Every time a page with the tool loads, the snippet picks A or B at random, each with an even chance, and registers the tool with that description. When an agent calls the tool, the call records which description the agent saw.
The draw holds for that page. Polls and client-side route changes keep it. A new page load draws again. Nothing is stored in the visitor's browser, so the test needs no cookie consent.
Reading the result
The panel shows, for A and B:
| Column | Meaning |
|---|---|
| Calls | How many times agents called the tool with that description |
| Succeeded | How many of those calls worked, with the success rate |
Below the table is the verdict:
- Collecting: fewer than 30 successful calls so far. The panel says how many more are needed.
- No clear difference so far: the split is close enough to even that it could be chance.
- A (or B) gets clearly more successful calls: the difference is too large to be chance (95% confidence).
How the winner is decided
Both descriptions are shown equally often, so if the wording made no difference, successful calls would split roughly 50/50. The test checks whether the actual split is too lopsided for chance.
It counts successful calls, not success rate. A description that gets the tool called twice as often at a slightly lower success rate still gets more done for your visitors.
Example: A got 12 successful calls and B got 31. That is 43 in total; an even split would be about 21 each. 31 is far enough above that for the panel to say B is clearly better.
Ending the test
- End test, keep A: B is discarded and the description stays as it was.
- End test, switch to B: B becomes the tool's description. The change is saved in the tool's history, like any edit, and can be reverted from there.
The button for the variant that is ahead is highlighted. You can end the test at any time, with or without a verdict.
Good to know
It takes real agent traffic. The verdict needs at least 30 successful calls. Few agents call WebMCP tools today, so on most sites a test runs for weeks. Test on the tools that get called most.
Changing A restarts the test. If you edit the current description while a test runs, calls made against the old wording no longer count, and the comparison starts over.
Change one thing at a time. If B differs from A in several ways, the test tells you which description is better but not why. Small, focused changes teach you more.
Calls outside the test are not counted. Only calls made after the test started, and only from pages that loaded a variant, are compared. Your regular analytics still show every call.