How the bench will test AI ad tools
No bench round has run, so every page is a research edition: official pages and terms, each fact dated. Once an approved round runs, each benched tool’s outputs from one fixed brief will sit beside the others’. Some tools’ terms keep them off the bench unless the vendor agrees in writing.
What a research edition rests on
A research edition rests on documents you can open yourself. Most are the vendor’s own: its live pricing page, help centre and terms. Some pages also count dated user reports. Every claim sits on one rung of the evidence ladder below, and its wording shows which one.
| Rung | What it rests on, and how a page words it |
|---|---|
| Stated | The vendor’s own words, in quotation marks and attributed |
| Read | A primary document such as a pricing page or terms clause, cited with the clause number and the date read; “on our reading” where we interpret |
| Computed | Our arithmetic on sourced figures, shown with the formula and its inputs |
| Reported | Forum threads and other third-party accounts, counted and linked with their dates, never averaged |
| Inferred | Our judgement from the rungs above, labelled as ours with the reason |
| Unknown | What only the bench can answer, in a line that starts “Not yet known:” |
Scroll sideways for all 2 columns
An example of each rung: editorial policy and corrections.
A price is verified when it was read on the date shown, on the vendor’s live pricing page, official help centre or developer documentation. A price we could not read there shows “Unverified”, and a vendor with no public price list shows “Not published”.
When this page was built, 59 of the 76 tools in the database had a verified price and 15 showed “Unverified”. Of the rest, 1 had no public price list and 1 had shut down.
A terms summary counts as verified only when the clause itself was read in the vendor’s terms. A plan card does not count.
How often prices and terms are re-checked
Prices are re-checked against the live official page at least every 30 days, and terms at least every 90 days. After 30 days a price stops counting as verified. Tables then replace the figure with “Unverified, last checked” and the date; sentences keep the figure beside the date it was last read.
Each re-check moves the checked date in a page’s meta line; “Updated” moves only when a fact changes, which the page’s list of changes records. The price tracker dates every tool’s prices.
The AI commercial use matrix does the same for terms: every row shows the date its clause was read.
What the bench status labels mean
Every tool page carries one of four bench status labels:
- “Not yet bench-tested”: the tool has not run the brief, so its pages are research editions.
- “Bench-tested” with a date: the tool ran the brief on that date, and its outputs are published.
- “Re-test due”: the tool ran the brief, but a new model or new pricing since then means its results may no longer hold.
- “Discontinued”: the tool has shut down, and its page says what happened and what replaces it.
On 24 September 2026 no tool in the database had run the brief, so every tool reads “Not yet bench-tested”, or “Discontinued” where it has shut down.
The fixed brief: what is decided and what is not
The bench brief will be one product, one hook line and a fixed set of ad formats, identical for every tool. Only the tool then changes between two sets of outputs. None of the three is chosen yet.
What the bench will measure and publish
Every output a tool makes on the brief will be published, failures included, next to the brief that produced it. The three measures below are proposals. Who judges an output usable, and against what checklist, is not decided.
| Measure (proposed) | How it would be worked out |
|---|---|
| Usable as it comes | Outputs that need no re-render, divided by all outputs made |
| Time to a first usable cut | From sign-up to the first usable output, so learning the tool counts |
| Cost per finished ad | What the round cost at list price, divided by the outputs usable as they came |
Scroll sideways for all 2 columns
Research editions quote a different figure today: the cost per output at list price. It divides a plan, top-up or API price by what that price buys, then multiplies by what one output uses.
Each tool’s method note shows its own sum. Where a plan counts credits, the figure holds only if every credit is used and nothing is re-rendered.
Tools whose terms keep them off the bench
The terms below bar benchmarking the service or publishing the results. Meta’s go further and bar any use of its output off its own platforms.
On our reading, each benchmark clause covers running a tool on a shared brief and publishing the comparison. None of these tools will run on the bench without its vendor’s written agreement, so their pages stay research editions.
| Tool, clause and date | The words that bar it |
|---|---|
| HeyGen Terms s.2, 23 Jul 2026 | “Access any portion of the services for benchmarking, comparative or competitive purposes” |
| Luma Terms 3.4(g), 14 May 2026 | “publish benchmarks or performance information about the Services” |
| AdCreative.ai Subscriber terms 2.2(f), 19 Dec 2025 | “use the Services for any competitive or benchmark purposes” |
| Pencil Core terms 5.1(i), 22 Apr 2026 | “engage in any competitive analysis or benchmarking of the Pencil Technology” |
| Claid.ai Terms s.3, item e, 10 Aug 2023 | “for benchmarking or competitive analysis of our Services” |
| Motion Terms, preamble, 12 May 2026 | “FOR ANY OTHER BENCHMARKING OR COMPETITIVE PURPOSES” |
| The Brief Acceptable Use Policy, “Engage in Competitive Analysis”, 1 Oct 2025 | “Use the Services for performance benchmarking or to build a competing product or service” |
| Canva Terms of Use 2(e)(iii), 19 Aug 2026 | “access the Service for purposes of performance benchmarking” |
| Jasper Usage Policies v1.2, A(iv), 24 Aug 2024 | “monitor the Services for any benchmarking or competitive purpose” |
| Meta Advantage+ creative Generative AI Terms s.1, 6 May 2024 | “Use or publication of Output outside of Meta’s platforms is unauthorized and a violation of these Terms” |
Scroll sideways for all 2 columns
Motion makes no ad creative to bench anyway.
Canva’s live terms page refused our requests on 24 September 2026. Its clause was read in the Wayback Machine capture of 23 September 2026, which carries the same effective date. We have not read whether Canva’s Enterprise agreement says the same.
Jasper’s clause does not apply to a use expressly permitted, for example in its Documentation or an Order Form.
Meta’s clause names no benchmark. Its ban still keeps every Advantage+ creative output off the bench and off every page here. Its words come from the Wayback capture of 22 November 2024; the live page still showed the 6 May 2024 date on 24 September 2026.
Vidu’s clause is unclear. Its Terms of Service, section 5, last updated 3 July 2026, name no benchmark but bar “Using the Services for competitive analysis, developing competing products or services, or any purpose that may be detrimental to our business interests”.
On our reading a published side-by-side comparison may count, so Vidu stays off the bench unless it agrees in writing.
Before each round, every tool’s terms will be re-read for benchmark clauses, and this section updated to match.
How tools enter the database, and what happens when they leave
A tool enters the database when it makes ad creatives, edits them or helps test them. That covers AI UGC and avatar video, generated video, product video, static ads, ad copy and creative research. No vendor can pay for a listing or a better place.
A tool that shuts down or leaves ads keeps its one page, with no child pages or comparisons. Which tools run in which round is not decided, and will be published here with the brief.
How we stay independent
Every bench account will be bought at the price checkout charges any buyer, with the list price recorded beside it. List price is a plan’s regular price on its official pricing page, before any coupon, partner deal or limited-time offer. Gifted plans and credits will be refused. No vendor will see an output or draft before it is public.
Verdicts are written without commission data, and each verdict’s wording is recorded before anyone checks which linked tools pay us. Who pays us lists every tool in the site’s database and whether it pays. A correction is dated and stays visible on the page it fixes; the about page says how to report one.
Method versions
- Draft 0.3, 24 September 2026: Canva, Jasper and Meta Advantage+ creative joined the tools kept off the bench, and Vidu’s clause was added as unclear.
- Draft 0.2, 24 September 2026: the unapproved product and hook line came off, and open parts are now labelled as proposals. The benchmark clauses on file that day were quoted.
- Draft 0.1, 23 September 2026: the first draft of the method.
Sources
- HeyGen: HeyGen Terms of Service, section 2, benchmarking bar, last updated 23 July 2026. Checked
- Luma AI: Luma Terms of Service, section 3.4(g), benchmarks and performance information, last updated 14 May 2026. Checked
- AdCreative.ai: AdCreative.ai terms of service for online subscribers, clause 2.2(f), competitive or benchmark purposes, updated 19 December 2025. Checked
- Pencil: Pencil core terms and conditions, clause 5.1(i), competitive analysis or benchmarking, updated 22 April 2026. Checked
- Claid.ai: Claid.ai Terms of Service, section 3, item e, benchmarking or competitive analysis, effective 10 August 2023. Checked
- Motion: Motion Terms of Service, preamble, benchmarking or competitive purposes, last updated 12 May 2026. Checked
- The Brief: The Brief Acceptable Use Policy, Engage in Competitive Analysis, performance benchmarking, last updated 1 October 2025. Checked
- Canva: Canva Terms of Use, section 2(e)(iii), performance benchmarking, effective 19 August 2026; the live page refused automated requests on 24 September 2026, so the clause was read in the Wayback Machine capture of 23 September 2026. Checked
- Jasper: Jasper Usage Policies, version 1.2, A Platform Guidelines, item (iv), benchmarking or competitive purpose, effective 24 August 2024. Checked
- Meta: Meta Ad Creative Generative AI Terms, section 1, Rights in Ads Content, output outside Meta’s platforms, effective 6 May 2024; English wording from the Wayback Machine capture of 22 November 2024, and the live page still showed the 6 May 2024 date on 24 September 2026. Checked
- Vidu: Vidu Terms of Service, section 5, Prohibited Conduct and Content, competitive analysis, last updated 3 July 2026; unclear whether it covers a published bench. Checked
What changed on this page
- Bench accounts will be bought at the price checkout charges any buyer, with the list price recorded beside it (was: at list price), since some tools apply a public offer at checkout.
- Corrected: a price read in the vendor’s developer documentation also counts as verified (was: the pricing page or help centre only).
- Method draft 0.3; see Method versions.
- Rewritten as Method draft 0.2; see Method versions.
- Method draft 0.1 written.