Soultab personas are versioned prompt files, not humans. Each persona is a set of behavioral instructions — what to explore first, what triggers impatience, when to abandon — loaded at run time. Persona prompts are stored in the repository as versioned code (e.g. persona.busy_professional v2.0.0), promoted through eval gates that measure fidelity, differentiation, and drift.
Three personas browse on mobile profiles, two on desktop — matching the real-world distribution of SaaS landing page traffic (79% mobile, per Unbounce's benchmark of 464M visits). Every persona navigates the live site in a real browser session. They click, they scroll, they read. No screenshots-only shortcuts.
The Soultab Score is a 0–100 composite across five dimensions: comprehension (does the persona understand the product?), intent (is the persona motivated to act?), trust (does the persona trust the site enough to proceed?), discoverability (can the persona find what it's looking for?), and friction (where does the experience break down?).
Each run receives a verdict band: looking good, needs improvement, or critical. Bands are broad categories — we never render a score with decimal precision, because that would imply a level of accuracy we don't claim.
Every finding carries a chain of custody. The exact persona prompt version, reasoning model, code release, and browser context that produced it are recorded in the finding's provenance block. A run without provenance is invalid by system rule. This means any finding can be traced back to its inputs — and any disagreement between runs can be investigated, not shrugged at.
Findings are hypotheses to test, not facts. Each one includes a recommendation written as a testable statement (“Hypothesis: a visible pricing link in primary nav recovers price-first evaluators”), a severity verdict, and the persona's own words as the evidence. We label confidence and surface methodology limitations separately — when the harness itself couldn't test something, we say so instead of blaming the site.
Soultab is built as a repeated-measures instrument. The same five personas, the same way, every run — so when the score moves, it's the page that moved. The system runs a test-retest reliability gate against a fixed reference cohort of PLG and self-serve SaaS landing pages. Persona differentiation is measured under normalized conditions (same budget, same viewport, same entry point) to confirm that personas produce measurably different behavior, not just mechanically different action counts.
We publish this methodology because a truth-telling tool that hides its workings would be a very short-lived joke. If you have questions, email the founder.