METHODOLOGY · DATA AS OF 2026-09
How the Scores Work Every number on this site, derivable by hand
No ranking here is hand-ordered and none is sponsored. This page defines every input, every weight and every penalty, so any score can be recomputed from the published fields.
The match formula
Each tool carries two editorial ratings on a 0–100 scale: simplicity (how fast a team becomes productive with defaults) and power (how far configuration, automation and scale can be pushed). The priority slider sets a weight p between 0 and 1, and the base score is the blend:
score = simplicity × (1 − p) + power × p
Two sets of default inputs appear on the site, and they differ on purpose. The printed table columns on every category page use team 6, budget $12 per seat and priority 50 (p = 0.5, an even blend), so the static numbers favour neither camp. The live engine starts its priority slider at 25 (p = 0.25, leaning simplicity), because a reader who has not yet touched the slider is usually earlier in the search than a reader who has. Move the slider to 50 and the engine reproduces the printed column exactly; every difference between widget and table is those inputs, nothing else.
Two adjustments follow. If the tool's published per-seat price exceeds your budget setting, the score is multiplied by the ratio of budget to price, floored at 0.55: expensive tools are penalised, never hidden. If your team size is below the vendor's evident target scale, the score is multiplied by 0.8, because adopting a tool built for a different scale costs configuration time the ratings cannot see. Tools with a usable free plan at zero entry price receive a three-point bonus. The result is capped at 99: no tool is a perfect answer for constraints it has never seen.
The lock-in grade
Lock-in measures the cost of leaving, from four facts checkable in vendor documentation. Each fact maps to fixed penalty points:
- Self-hosting — yes 0 · partial 12 · no 25
- API — open 0 · limited 15 · none 25
- SSO — all plans 0 · paid tier 8 · enterprise only 18
- Data export — full 0 · partial 16 · limited 30
Totals under 20 grade LOW, under 45 MODERATE, otherwise HIGH. The facet letters shown on comparison pages use the same table. The weights are opinionated (export completeness weighs most because it is the fact you cannot negotiate after signing) and they are fixed across every category, so grades are comparable site-wide.
Where the data comes from
Pricing figures are the vendor's own published per-seat monthly price on annual billing, read from the public pricing page and normalised to USD; every ranking row links that page in its price column, so the source is one click from the figure it produced. Where a vendor prices flat per company (accounting tools) or by usage (email sends, error events, build minutes), the entry tier's published price is used and the category page says so in prose. Vendors change prices without notice; the dataset carries the reading date at the top of this page, and the comparison baseline matters more than the absolute cent. Free-plan availability, SSO placement, export scope and self-hosting status come from vendor documentation and plan matrices.
What is editorial, and what is not
The simplicity and power ratings are editorial judgments: two numbers per tool, set once per dataset revision, moved only with a reason. Everything else is either a vendor-published fact (prices, plans, SSO tiers) or a deterministic computation over those facts (match scores, lock-in grades, rankings). Vendors cannot pay for placement, for a rating change, or for exclusion of a competitor. If a figure on this site disagrees with a vendor's current pricing page, the vendor's page is right and the dataset is behind. The comparison logic still holds, because every competitor in the category is read at the same revision.
Why the adjustments are shaped the way they are
Each mechanism in the formula answers a specific failure mode of ranking sites. The budget penalty is a multiplier rather than a filter because hiding expensive tools would falsify the field: you deserve to see that the strongest tool is out of budget, priced at exactly how far out it is. The 0.55 floor exists so that no price gap can zero a tool out entirely: a score is a comparison, and a comparison needs both ends visible. The team-size discount is flat (0.8) instead of graduated because the dataset can state a tool's evident target scale but cannot honestly measure degrees of mismatch, so a single fixed discount claims only what the data supports. And the free-plan bonus is worth three points, not ten, because a free tier changes how cheaply you can be wrong, not how good the tool is. Where the formula is crude, it is crude in the open, which is the property the whole site is built on.
Limits worth stating
Scores compress trade-offs into one number, and compression loses information: a 78 hides which constraint bound. The category prose exists to restore that context, and the facet grids on comparison pages unpack the lock-in number into its parts. Treat any two scores within three points as a tie, and settle ties with a time-boxed trial: the formula ranks candidates; it does not run your pilot.
Frequently asked questions
04 QUESTIONSCan I recompute a score myself?
Yes, with a pocket calculator. Take the tool’s simplicity and power ratings from its category table, blend them by your priority weight, apply the budget ratio if the published price exceeds your budget, the 0.8 scale factor if your team is below the tool’s target size, and the free-plan bonus. The result matches the engine to the point.
Why are simplicity and power editorial rather than measured?
Because no honest measurement exists. Setup-minutes benchmarks measure the benchmark author, not the tool, and feature counts reward bloat. Two openly-editorial numbers with fixed definitions, moved only with a reason, are more auditable than a fake-precise metric, and everything downstream of them is deterministic.
How often does the dataset get revised?
Per revision, dated at the top of this page and in the dataset itself. Vendors change prices without notice, so between revisions a figure can lag the vendor’s pricing page; the ranking still holds as a comparison because every competitor is read at the same revision.
Why does data export weigh most in the lock-in grade?
Because it is the only facet you cannot negotiate after signing. SSO placement changes with plans, APIs open over time, self-hosting editions appear. But if your history exports incomplete today, the migration you might need in two years is already compromised. The penalty table prices that asymmetry.