Data methodology
Where every number on this site comes from, how big the sample behind it is, and what we have not measured yet. If a figure is a placeholder, this page says so and the page carrying it says so too.
Last updated July 2026
Read this first: the site is new
We launched with the collection infrastructure built and the datasets close to empty. That is deliberate — the alternative is publishing plausible numbers and backfilling evidence later, which is how every scraped directory in this industry already works. The counts below are live: they are read out of the content files when this page is built, so they go up as real data arrives and they cannot be inflated by editing this paragraph.
Current state of the data
| Dataset | Published | Held back |
|---|---|---|
| Fault code pages | 0 | 15 in technical review |
| Repair cost pages | 0 | 10 in review |
| Cost pages with at least one range from our own receipts | 0 | 10 provisional, sourced but unverified |
| Phone-verified shop listings | 0 | 38 seeded demonstration records, noindexed and labelled |
| City directory pages in the index | 0 | 8 built from seeded records only (noindex, not in the sitemap) · 4 below the listing minimum, rolled up to their state hub |
| State labor-rate pages | 0 | 10 states with no defensible sample yet |
Every one of those “held back” figures is a page we could have published today and chose not to.
1. Crowdsourced repair receipts
This is the dataset the rest of the site is built to justify. Anyone who has paid for a truck repair can send us the invoice detail — parts, labor hours, the rate they were charged, shop supplies, the total, and, critically, whether the repair actually fixed the problem. That last field is the one nobody else collects, and it is what turns a price list into a cause-frequency table.
What happens to a submission
- It lands in a moderation queue with
status = pending. Nothing a member of the public submits is ever rendered on this site automatically. The public site is a static export with no database connection at request time, so this is a structural guarantee, not a policy we promise to remember. - A human reads it. We check the arithmetic (parts + labor × rate + supplies against the stated total), sanity-check the labor hours against published flat-rate guides, and flag anything that looks like a shop submitting its own pricing.
- Approved rows are aggregated. Individual submissions are never published — not the shop name, not the invoice, not the submitter. Only the aggregate moves onto a page.
- The aggregate goes into the repair-cost page's content file, and the page starts rendering
verified: truewith the sample size printed next to the range.
What we exclude
- Estimates and quotes. Only invoices for work actually performed.
- Warranty and goodwill repairs, which have no meaningful customer price and would drag a range downward.
- Repairs bundled with unrelated work where the line items cannot be separated.
- Anything older than 36 months, because parts and labor pricing in this industry does not hold still that long. Age is recorded on every row so we can shorten that window later rather than guess at it now.
- Duplicate submissions of the same invoice, which are marked as duplicates rather than deleted so the deduplication itself is auditable.
How a range is calculated
We publish an interquartile range — the 25th to 75th percentile of approved submissions — not a minimum and maximum. One fleet that negotiated a parts discount and one owner-operator who got taken on the side of a highway are both real, and both are useless as the edges of a range someone is about to use to judge a quote. Alongside it we publish the median, the sample size, and the date of the most recent submission in the set.
The minimum sample rule
A cost page needs at least 5 approved receipts before it shows a range derived from them. Below that it falls back to the published sources it cites, renders verified: false, and displays a banner saying the figures are provisional. A percentage calculated from four invoices is noise wearing a decimal point.
Right now, no cost page has cleared that threshold. The rest cite their published sources and say plainly that the figures are provisional.
2. Regional labor rates
Shop labor rates vary more by geography than almost any other input to a repair bill, and there is no public dataset for heavy-truck rates. We are building one from two sources: the labor rate line on submitted invoices, and direct calls to shops in each state asking what they charge per hour for heavy-duty work.
A state publishes a labor-rate band only when it has a sample we would defend in an argument. Today that is 0 of 10 states we hold records for, which is why /repair-costs/labor-rates/ currently lists methodology rather than numbers. A state whose record is open still has a page, because the freight corridors, inspection regime and verified shop coverage there are worth reading — but it publishes no band, it is excluded from our sitemap, and it is marked noindex until the sample clears the floor.
The national placeholder
Cost pages need something to multiply labor hours by before our own rate data exists, so they use a provisional national band. Its published provenance reads, verbatim: provisional national placeholder — TODO(verify), not yet from our collected dataset. It appears in the visible byline under every range that uses it. We would rather show you an ugly disclosure than a clean number you cannot check.
Rates are recorded separately for dealer, independent, and mobile service, because averaging them produces a figure that describes nobody. Mobile and roadside carry a call-out charge and often a higher hourly rate; folding that into a single state average would make every dealer quote look like a rip-off and every mobile invoice look like a bargain.
3. Fault code and diagnostic content
Fault code pages are compiled from OEM service literature, the SAE J1939 standard's suspect parameter number and failure mode identifier definitions, technical service bulletins, and — where we have them — the diagnostic paths our reviewing technicians actually follow.
The publishing gate, enforced in code
A fault code page cannot publish unless it has either an annotation from a credentialed technician who has diagnosed that code, or at least three cited technical sources plus real parts and labor cost data. This is a schema validation that fails the build, not an editorial aspiration. Pages that do not clear it stay drafts: excluded from the sitemap, marked noindex, and labelled “in technical review” wherever they are listed.
All 15 fault code pages currently on the site are drafts, because our technical reviewer positions are contracted but unfilled. You can read them — they are linked and they work — but they are not in the sitemap and they carry a banner saying no technician has signed them off yet.
Cause frequencies
Where a code has several plausible causes, we intend to rank them by how often each one turned out to be the actual fault. That ranking requires outcome data, which is the did it fix it field on the receipt form. Until a code has at least 5 outcomes, the component that draws those bars renders nothing at all rather than a percentage. Internally the frequency is stored as -1, a sentinel meaning “not yet measured” that cannot be mistaken for zero.
Causes are still listed in that state, ordered by our reading of the service literature, and the page says the ordering is not yet evidence-based. That is the honest version of a ranking we cannot yet support.
4. Shop listings and verification
Every directory in this industry is scraped. Ours is phoned. A listing does not publish until someone calls the number, confirms the business exists, confirms what it works on, and confirms it is still open. The date of that call is printed on the listing, and the listing carries the method used.
- Verified — we called, someone answered, the details were confirmed. The call date is published.
- Unverified — submitted but not yet reached by phone. Shown with the badge withheld and a visible note. Never counted toward a city's listing minimum.
- Sample — 38 seeded demonstration records that exist so the directory logic can be reviewed before real coverage lands. They use reserved 555-01xx phone numbers and are labelled on every surface they appear on. They do satisfy the city minimum, deliberately, because the rollup and the service split have to be testable — and they are excluded from every public count, from
LocalBusinessstructured data, and from the index: a city page held up only by sample records isnoindexand stays out of the sitemap. They are not real businesses and are not a recommendation.
Why some cities have no page
A city directory page needs 3 verified listings to exist, and a city-plus-service page needs 2. Below the threshold the city rolls up to its state hub, where its shops still appear under “coverage in progress” — the listings are not hidden, they just do not get a page of their own. A page listing one shop is not a directory; it is a shop profile with a directory's URL, and mass-producing them is the fastest route to a manual action from Google.
8 cities currently clear the threshold and 4 do not. The same rule governs the links: the helper that decides whether a city page exists is the same helper every link on the site calls, so we cannot link to a page we declined to build.
Paid listings
Paid tiers add detail to a listing — photos, service descriptions, a website link. They do not affect ordering, they do not affect whether a shop appears, and they never will. Paid listings are labelled as such. If we ever change that, this paragraph changes with it and the change gets a date.
What is missing, by name
Anywhere we needed a fact we do not have, the content file carries a TODO(verify): marker rather than a guess. Those markers are collected into a single file in the repository, CONTENT-TODO.md, so the gap is tracked rather than forgotten. The categories currently outstanding:
- Flat-rate labor hours for most operations, which need a licensed guide, not a forum post.
- Part numbers and current parts pricing across engine families.
- Cause frequencies for every fault code — blocked on outcome data.
- Per-state labor rates — blocked on the calling program.
- Scan tool coverage by OEM, which needs testing rather than marketing copy.
- Service-desert mapping along the interstate corridors.
Corrections
If a number here is wrong, tell us and we will fix it. Corrections are handled as an editorial priority rather than a support ticket, and a substantive change is noted on the page rather than made quietly. We would much rather be corrected by a technician than be confidently wrong in front of one.