A 200-row dataset scoring fictional local businesses across 10 trades on the seven content factors that matter for AI search — with the full methodology, the scoring formula, and a generator script you can audit and rerun.
CSV dataset + methodology report + generator script
Single-buyer license
CSV + report emailed within 24 hours of secure checkout
ai-search-visibility-benchmarks.csv — 200 rows: 10 trades (plumber, electrician, HVAC, roofer, landscaper, painter, dentist, lawyer, auto repair, garage door) × 20 sample businesses. Columns: business ID, trade, market, 7 binary factor columns, and the composite ai_visibility_score (0–100).generate_csv.py — the generator script. Fixed random seed, documented weights; re-running it reproduces the CSV byte-identically. Audit it, tweak the weights, generate your own variants.benchmark-report.md — full methodology: the 7-factor framework, the scoring formula, how the sample was built, per-trade findings, and an explicit section on the dataset's limits.| Factor | What it checks | Weight |
|---|---|---|
| Dedicated page per job | Each core service gets its own page | 0.20 |
| Prices in plain text | Prices, ranges, or pricing policy stated | 0.15 |
| Written-estimate policy | Estimate-in-writing stated on the site | 0.15 |
| Reviews name the job | Reviews describe the specific work done | 0.15 |
| Real photos / town pages | Real photos, town-specific pages | 0.15 |
| Quotable certifications | Licenses/certs with numbers or dates | 0.10 |
| Crisis wording | Emergency scenarios addressed | 0.10 |
Score = 100 × Σ(weight × factor). Linear, auditable, recomputable by hand from any row.
In this synthetic sample, service specificity and review texture separate high scorers from low scorers by roughly 20 points — the factors most aligned with how AI retrieval actually works move the score most. Per-trade averages range from 32.5 (landscapers) to 50.8 (dentists), reflecting the illustrative maturity assumptions baked into the generator — the machinery is demonstrated, not the real world. The full findings and their limits are in the report.
Single-buyer license: use it within your own business — score your sites and client sites, include findings in client reports, modify the weights for your own use. You may not resell, redistribute, or republish the dataset or report, or use them to build a competing data product. No warranty; the sample is illustrative and no claim is made that scores predict real AI-search performance.