Every tool is checked against recorded sources, tagged by category and job-to-be-done, and audited for evidence coverage. Where a claim cannot be traced, we reduce the score and show the gap rather than guessing.
Facts follow a tiered hierarchy: official product and pricing sources first, then publisher-owned videos, then independent evidence. Pricing and profile checks lose freshness points automatically as their dates age. A missing claim-level or independent citation can never receive full evidence credit.
The visible score is deterministic: feature evidence contributes 30%, pricing evidence 25%, decision-support depth 25%, and source breadth plus recency 20%. The score measures how much of our evidence rubric is supported. It is not a product-quality rating, a predicted outcome, or proof that one tool is better than another.
A dimension reaches 100% only when every required input is present. For pricing, that means complete plan records, a current plan-specific official source, and a separately recorded corroboration link. Official-source-only pricing can score highly, but it does not receive the corroboration points.
Rankings lead with qualitative fit labels (Best fit, Strong fit, Conditional fit, Weak fit, Insufficient evidence). The same tool can be a Best fit for one buyer and a Weak fit for another. No vendor can purchase a fit label.
AI recommendation scores are displayed only when a live, reproducible prompt-panel dataset is connected. Illustrative sample data is excluded from tool pages and does not affect the evidence audit.