Testing

How we test Queva before you get it.

We write pretend owners, such as a tile installer, a salon and a parent. They text Queva about ordinary tasks, and a separate AI grades whether each owner got what they wanted.

Results, October 4 and 5, 2026

Simulated owner chats

Rounds we kept a full record of. Round 2 is not included.

RoundDateSimulated chatsOwner got what they wantedTool builds that finishedBuilds the owner could useTypical wait for a reply
Round 1Oct 410071%10 of 103 of 1011 s
Round 3Oct 42383%2 of 20 of 2 (tool links were broken on the test server)17 s
Round 4Oct 42383%2 of 22 of 215 s
Round 5Oct 52378%2 of 22 of 221 s

Read this with the table

  • These are simulated owners we wrote, not real customers.
  • Another AI did the grading. It can be wrong.
  • Samples are small: 23 chats, so one chat moves the score about 4 points. A single run can move a few points by chance.
  • Round 1 used a different grader, so compare it with later rounds loosely.
  • Waits got longer in later rounds as replies got more careful.
  • This is not a promise about your results.

In round 5 we also tried businesses Queva had not seen in testing. 78% of those owners got what they wanted. Tech words per 100 words of reply went from 5.7 in round 1 to 1.2 in rounds 3 to 5.

What we changed because of it

Fixes that came out of these rounds

  • Plain words instead of tech terms.
  • One question per text, at the end.
  • A tappable link to each tool it builds, instead of code or file names.
  • Builds start from a page we already tested.
  • It learns the owner's trade early and does not stop at the first "no".

We will add newer rounds here with their dates. Back to Queva.