Testing
How we test Queva before you get it.
We write pretend owners, such as a tile installer, a salon and a parent. They text Queva about ordinary tasks, and a separate AI grades whether each owner got what they wanted.
Results, October 4 and 5, 2026
Simulated owner chats
Rounds we kept a full record of. Round 2 is not included.
| Round | Date | Simulated chats | Owner got what they wanted | Tool builds that finished | Builds the owner could use | Typical wait for a reply |
|---|---|---|---|---|---|---|
| Round 1 | Oct 4 | 100 | 71% | 10 of 10 | 3 of 10 | 11 s |
| Round 3 | Oct 4 | 23 | 83% | 2 of 2 | 0 of 2 (tool links were broken on the test server) | 17 s |
| Round 4 | Oct 4 | 23 | 83% | 2 of 2 | 2 of 2 | 15 s |
| Round 5 | Oct 5 | 23 | 78% | 2 of 2 | 2 of 2 | 21 s |
Read this with the table
- These are simulated owners we wrote, not real customers.
- Another AI did the grading. It can be wrong.
- Samples are small: 23 chats, so one chat moves the score about 4 points. A single run can move a few points by chance.
- Round 1 used a different grader, so compare it with later rounds loosely.
- Waits got longer in later rounds as replies got more careful.
- This is not a promise about your results.
In round 5 we also tried businesses Queva had not seen in testing. 78% of those owners got what they wanted. Tech words per 100 words of reply went from 5.7 in round 1 to 1.2 in rounds 3 to 5.
What we changed because of it
Fixes that came out of these rounds
- Plain words instead of tech terms.
- One question per text, at the end.
- A tappable link to each tool it builds, instead of code or file names.
- Builds start from a page we already tested.
- It learns the owner's trade early and does not stop at the first "no".
We will add newer rounds here with their dates. Back to Queva.