Cumulative pool: ClickUp, Notion, PostHog, Attio, Linear, HubSpot, Webflow, Zendesk, Amplitude, Miro, Calendly, Airtable
The target: there isn't one worth reporting as damage. This session tested Airtable's AI cobuilder, Omni, with a real prompt — a usability-test tracker for FindTheTarget itself — and Omni handled it correctly at every stage that trips other tools in this pool: it stated a plan in plain language before touching anything, asked for explicit confirmation (“Build it”), stayed labeled through the entire generation, and produced a genuinely usable base with realistic sample data on the first try. No blank pages, no unlabeled spinners, no dead ends. The one honest note is smaller than a finding: neither the “Planning” nor the “Identifying the relevant data” stage gives a time estimate, so a first-time user has no way to know if a 20-second wait is normal or already too long. That's worth watching if it recurs elsewhere in the product. It is not, on its own, evidence of anything broken.
Full analysis in report run 9.
Rated in a previous run; not re-analyzed this cycle.
Full analysis in report run 2.
Full analysis in report run 2. Case-study finalist.
Full analysis in report run 2. Case-study finalist.
Full analysis in report run 3.
Full analysis in report run 4. Case-study finalist — targeted fix in report run 5.
Full analysis in report run 6.
Full analysis in report run 7.
Full analysis in report run 8.
Full analysis in report run 10. Targeted fix included.
~5 min 57 sec · ~30+ distinct product states · Google OAuth signup · AI-cobuilder session with a real, on-brand prompt
This is the strongest onboarding session in the pool since Linear, and for a related reason: it doesn't make the user guess what's happening at any point. The finding below is reported as a positive pattern worth naming, not a crack — consistent with the standing rule against manufacturing a flag where the evidence doesn't support one.
The resulting base (not pictured to save space, but verified directly) shipped with 15 realistic sample participants, linked Test Sessions, Tasks, and Test Results tables, and a working log-session form — a first-try result that matched the original prompt's structure exactly. Returning to the workspace list afterward correctly showed the built app as “Opened 2 minutes ago”; nothing was lost or orphaned by navigating away.
Copy holds up — every status label (“Planning,” “Identifying the relevant data,” “Build it”) accurately describes what's actually happening.
Hierarchy holds up — the plan-then-confirm sequence puts the one decision that matters (does this match what I asked for) in front of the user before any work is committed.
Placement holds up — nothing competes for attention during generation; the status is always where the user is already looking.
Load time is the one soft spot — both AI stages ran for roughly 20–30 seconds with no numeric progress or estimate. Labeled beats blank, but an estimate would beat a label alone. Noted, not treated as a finding on its own.
Motivation not applicable — nothing in this session gave the tester a reason to hesitate or reconsider.
Outside signal
This cycle's search was for disconfirming evidence rather than corroborating it — specifically, complaints about Omni or Airtable's AI builder producing broken, empty, or misleading results, which would undercut the positive read above. None were found in the course of this session's research. That's reported as an absence of contrary signal, not as independent proof the pattern holds at scale; a single clean session is evidence, not a guarantee.
Not rated “greatly help” the way Linear was, because this flow asks more of a new user — writing a real prompt is a higher-effort first action than Linear's near-zero-decision signup, and the two AI stages together cost close to a minute with no time estimate. But every one of those minutes was labeled, staged, and confirmed, and the result matched the request on the first attempt. That combination — genuine effort asked of the user, matched by genuine transparency about what's happening — is what “slightly help” is for.
Pool is now 12 tools (ClickUp, Notion, PostHog, Attio, Linear, HubSpot, Webflow, Zendesk, Amplitude, Miro, Calendly, Airtable).