How we compare AI products
Method last changed:
In short
This page sets out exactly how the comparisons on this site are made. It was published before any ranking, and every ranking follows from it mechanically: nobody, neither the people nor the AI that build the site, picks winners or changes the order by opinion. The same criteria apply to every vendor.
What we compare
The free and paid plans that a private person can sign up for, for AI products in four groups: chatbots, agents, image generation and video generation. A plan appears in a group when the vendor's own page offers it for that use.
Which vendors: every organisation that has at least one model on the group's Arena leaderboard (the same one used for quality) and sells, under its own name, a plan that a private person in Italy can sign up for, according to the organisation's own pages. There is no ranking cut-off, and no vendor is added or left out by hand. Organisations without such a plan, such as research labs that only publish models, do not appear.
Quality: the Arena leaderboard
For quality we use one source: the public leaderboard of Arena (arena.ai, formerly LMArena), where people compare two anonymous answers and vote for the better one. Arena publishes its leaderboards as an open dataset under the CC BY 4.0 licence.
For each group we use the matching “overall” leaderboard: chatbots – text with style control (text_style_control); agents – agent; images – text_to_image; videos – text_to_video. We only use a leaderboard that Arena published in the last 30 days.
We chose Arena because it covers all four groups, publishes its results under a licence that allows their use on a commercial site, and is not one of the vendors compared. The Arena score measures which answers people preferred, not whether they are correct: it is a widely used public measure, not a final verdict.
Which model a plan gives you
The models a plan includes are those named on the vendor's own pages. The vendor's name for a model, in lower case with hyphens for spaces, must equal an Arena name; if it equals none, an Arena name without its configuration suffix. Configuration suffixes, at the end of the name, are: reasoning levels (-minimal, -low, -medium, -high, -xhigh, -max, also followed by a budget such as -32k), -thinking (also with a budget), -preview, and anything in brackets such as “(xHigh)”. Dates, -instruct and -chat are part of the name. When several Arena variants match the same name, we count the one with the worst rank, because the page does not say which configuration the plan uses. When no variant matches, or the vendor's pages contradict each other, the plan is shown with “no Arena score”.
A plan's quality is that of its best model on the leaderboard: the lowest Arena rank and, on equal rank, the highest score.
Order
Within each group, plans are ordered by: 1) Arena rank of the best model, best first; plans with no Arena score come after all others; 2) on a tie, Arena score, highest first; 3) when everything else is equal, the plan's code in alphabetical order. The last rule only separates plans that are equal in everything else and says nothing about quality. For now the price does not affect the order (see “Prices and usage limits”).
Labels
We use only two labels, worked out within each group: “Top-ranked” for every plan whose best model has the best Arena rank in the group; “Best free plan” for the free plan with the best Arena rank. We use no other labels, ratings or adjectives of quality, apart from the strengths and weaknesses worked out as described below.
Strengths and weaknesses
For chatbots, Arena publishes leaderboards for single tasks besides the overall one. We use seven of them: creative writing, coding, maths, following instructions, long conversations, hard questions and German. For each plan we look at the same model that sets its place in the order, under the same Arena name.
“Strong at” a task means that model has Arena rank 1 to 10 on the task's leaderboard; “weaker at” means a rank worse than 30. At rank 11 to 30, or if the model is not on the task's leaderboard, we say nothing. The thresholds are the same for every product and do not depend on the vendor; a plan with no Arena score gets neither strengths nor weaknesses.
Prices and usage limits
For now we do not show prices. Vendors show prices according to the visitor's country, and the computer that builds this site is outside Italy, so it cannot read or check the price for Italy. Instead, each plan links to the vendor's pricing page. A plan is marked as free only when the vendor's page offers it at no cost.
Usage limits (such as the number of messages or images) are quoted word for word from the vendor's page, with a link. They do not affect the order, because vendors state them in different units that cannot be compared mechanically.
Up-to-date or missing data
Every model list, limit and “free” status records the page it was read from, the date it was read and an expiry of at most 30 days. The site cannot be published with an expired figure: the figure is checked again or removed, never shown out of date.
When a figure is missing we say so (“no Arena score”). A plan without an Arena score gets no label and comes after the plans that have one; we never estimate a missing figure.
Affiliate links
Some links to vendors may be affiliate links: if you buy through them, this site may earn a commission, at no extra cost to you. They are marked as such where they appear. The code that works out the order cannot tell whether a plan has an affiliate link, and an automated test checks this.
Who builds this site
The code and texts of this site are written with the help of Claude, an AI model made by Anthropic. Anthropic sells Claude, one of the products compared here. That is why the method is fixed in advance and applied mechanically, why every figure about Anthropic's products is read from published pages like any other vendor's, and why every change to this page needs the site owner's approval.
Vendors with affiliate links
None at the moment.