findings
observations, not conclusions. this project has run audits across a small number of pilot brands — far from a representative sample, and not a controlled experiment.
one before/after data point
one industrial-equipment brand we audited went from an average score of 5 to 39 across two audits, several months apart, after acting on several of the fixes this project surfaces. we don't have a holdout group or a way to isolate which fix mattered, or rule out other explanations — a competitor's content changing, or the underlying models updating between audits. we're reporting the number because it's the only before/after data point we have, not because it proves the fixes caused the change.
[placeholder — confirm specifics before publishing: aggregate patterns across the other pilot brands audited]
we don't yet have a case where the score got worse after intervention, or where multiple audits of the same brand disagreed sharply with each other — if either happens, it belongs here too.