Written from your specs AI Test Case Generator Map the coverage, then write the cases you pick Do the test once Record a Test Beta Get readable steps and Playwright + pytest scripts It can only read MCP server Let Claude, ChatGPT or Cursor read your test data Compare Test Management OracleAI Pricing Blog Docs Login Start free
AI in Hawzu

Putting AI in everything is easy. Saying where we didn't is not.

Some answers need judgement. Some have to be the same twice. Every feature here that uses AI says so and shows what it read; the rest is plain arithmetic, explained further down. In QA, a number you can't reproduce is worse than no number.

Uses AI
  • Coverage map
  • Already-covered check
  • Test case writing
  • Documentation gaps
  • Readiness narrative
  • Atlas map proposal
  • Rewrite
  • Duplicate defects
  • Oracle
  • Your own agent not our model
No AI — fixed rules
  • Readiness score
  • Impact tiers
  • Chart picks
  • Observatory insights
  • Flaky detection

same input → same number, every time

Everything on the left shows what it read. Everything on the right gives the same answer every time.

The AI side

Each one settles a question someone actually asks

Not here are our AI features. Here is the question each one is in the room to answer — and, at the foot of every card, the thing it read in order to answer it.

What should we even be testing here?

It maps the coverage first

Before a single step is written it proposes up to forty one-line scenarios across ten focus lenses, scored for depth per requirement. Ten finished cases can't show you a hole; forty titles can.

reads your specs
More

Do we already test this?

A second pass reads your existing cases

Every proposed scenario is checked against what your repository already covers and ruled covered, partial or not covered — citing the cases it relied on. It errs toward proposing: you can discard a duplicate you can see, not a test you were never offered.

reads your repository

Can it write the ones I picked?

Full cases, only for what you ticked

Steps, expected results, preconditions, priority and the requirement each one covers. It's a preview — nothing is saved until you accept it, and accepted cases land as ordinary test cases with an AI tag.

reads your specs
More

What does our spec fail to say?

It reports the gaps instead of filling them

A behaviour the source names but never gives an outcome for isn't a test case — there's nothing to assert. It comes back as a documentation gap in its own list, kept out of the drafts so nobody ticks one by accident.

reads your specs

Are we ready to ship?

It writes the readiness narrative

When a release wraps it turns that release's own numbers into the plain-English summary and the top risks, judged against the criteria you wrote down. It puts the scorecard into words — it does not compute it.

reads release metrics your specs
More

What is this product even made of?

It drafts a map of your application

Atlas reads your Canon documents and suggests the product's areas, screens, routes and journeys — as a draft you review item by item. Nothing reaches the map until a person applies it.

reads your specs
More

Can someone tidy this wording?

Rewrite, tighten, expand or proofread

Pick a mode, see a few variations, preview the exact before-and-after, and apply step by step or all at once. It works on test cases, requirements, releases and defects alike, one item at a time.

reads your repository

Has someone already filed this?

It catches the duplicate as you type

Your draft is compared by meaning, not just keywords, against every defect in the project, with a match score, so you can close it as a duplicate or jump to a fix someone already found.

reads your defects
More

Can I just ask?

Oracle answers in whatever shape the answer is

Ask in plain English and it returns the matching records, a chart, an answer from your own documents, or the right form open and ready. It prints every condition it applied, and when nothing can answer you it says so rather than producing something plausible.

reads your project's shape your specs Hawzu's docs
More
not our model

Can I ask from the editor I already use?

Your agent reads your tests, over MCP

Connect Claude, ChatGPT, Cursor or anything else that speaks the Model Context Protocol, and it can read your test cases, requirements, defects and releases from inside the tool you already work in. It can only read, under exactly your own permissions — and it uses no Hawzu AI credits, because your own AI does the thinking.

reads your project's shape
More
Intake

Everything it reads, you gave it

The chip at the foot of every card above points at one of these. Nothing is based on general knowledge about software testing — every AI feature reads something out of your project, or our own published documentation, and says which.

What the AI can read

Your specifications

your specs

The PRDs, standards and exit criteria you added to Canon. Test case generation and the release summary read them, and show the document and section they used.

Canon

Your repository

your repository

Your existing test cases, so the second pass can tell you something is already covered — and the exact field you are standing in when you ask for a rewrite. Nothing wider than that.

Test cases

Your project's shape

your project's shape

What your fields are called and the values they hold — folder names, labels, release titles, people. It is what lets Oracle answer a question about your records without reading them.

Oracle

This release's metrics

release metrics

The pass rates, coverage and open blockers of the release in front of it — handed over as data, with an instruction never to invent a number.

Releases

Your defects

your defects

Your project's own native defects, compared by meaning rather than by keyword — so the duplicate shows up before you file the second one.

Defects

Hawzu's documentation

Hawzu's docs

The in-app assistant answers from our published docs and links the page it used. It reads our documentation, never your content.

Docs

Never general knowledge about software testing, and never another project's data.

In the product

What the boundary looks like on screen

Release 2026.8.1 Completed
Written by AI
Top risks
Arithmetic
70 At risk
Test Execution 88
Quality 74
Defect Health 40
Requirements 70
One release screen, and both halves of this page inside it. The paragraph was written by AI from that release's own numbers. The score beside it is fixed arithmetic over the same numbers.
Boundaries Negative Security +7
REQ-14 Checkout totals 7
REQ-15 Promo stacking 4
REQ-16 Refund window none

Nothing proposed for REQ-16 — the source gives no outcome to assert.

Stage one, before a single step is written. Depth scored per requirement, and the requirement it found nothing for named outright rather than quietly skipped.
New defect
Checkout total wrong when promo applied
Similar defects
DEF-318 Promo code doubles the discount 91%
DEF-402 Basket total ignores promo cap 78%
Shown while you're still writing the one that would have duplicated it — matched by meaning, not by keyword.
How do I bulk-import a Postman collection?

“I couldn't find that information in the Hawzu documentation.”

instead of a confident four-paragraph answer
The other half of the same promise, quoted exactly as the assistant says it. Every other answer it gives links the page it came from; this is what happens when there isn't one.
Where we stopped

Five things we didn't make AI

Every one of these would be easy to ship as an AI feature, and several tools do — the quoted line on each is the name it would carry. Here they are arithmetic and fixed rules, because the answer has to be the same twice and you have to be able to argue with it.

“AI-powered release readiness”

The release readiness score

A score out of 100, built from four parts of the release. Two people looking at the same release always get the same number. AI writes the summary underneath it; it never touches the score.

Test Execution — how much of the release has been runQuality — how much of what ran passedDefect Health — open blockers, old or overdue defects, reopensRequirements — how many are fully verified (left out if there are none)

The verdict is Ready, At risk, Not ready or Not started — and an open blocker makes it Not ready whatever the score says. Each release can set its own bars.

“AI impact analysis”

Atlas impact analysis

Pick the part of the map that changed and Atlas sorts the affected test cases into three fixed buckets, each with the reason it's there. No risk score — a number nobody can predict is worse than a bucket they can.

changed item → Must retest · Should retest · Consider
“AI chart recommendations”

Chart recommendations

Fixed rules match what your project contains to the insights worth adding. The same project always gets the same suggestions.

project shape → a fixed list of charts
“AI-generated insights”

Observatory insights

Fixed rules with set thresholds. Each insight says which threshold it crossed, so you can disagree with the threshold instead of arguing with a black box.

if metric crosses threshold → say which threshold
“AI flakiness detection”

Flaky test detection

A count of how often each test flips between pass and fail across its run history. No AI decides what's flaky.

count pass/fail flips per test over its history
The rules it works under

AI you can put in a sign-off

In QA a made-up number is worse than no number. So every AI feature works under the same rules.

Applies to every AI feature in Hawzu 04 clauses
  1. 01

    It cites what it used

    A generated case names the document and section it drew on. A docs answer links the page. An uncited claim is visible as an uncited claim.

  2. 02

    It returns fewer rather than padding

    Asked for thirty scenarios from a thin source, it returns the number the source actually supports and says why. The ceiling is a ceiling, not a quota.

  3. 03

    It says when it doesn't know

    No outcome in the spec means a documentation gap, not an invented assertion. No answer in the docs means it tells you so.

  4. 04

    Nothing is written without you

    Generated cases are a preview. Rewrites show the before-and-after. An Atlas suggestion is reviewed item by item. There is no background AI writing to your project.

Wednesday · 11:20 AM — release review

“Did the AI decide that, or did we?”

Someone points at the amber readiness band and asks the question that should be asked of every AI-shaped number in a review. Here the answer is short: the band came from fixed arithmetic over the release's own results, the paragraph under it was written by AI from those same numbers, and the exit criteria it judged against are a document in your Canon that anyone in the room can open. Three different kinds of answer, and you can tell which is which.

Asked and answered

The questions this page invites

A page that publishes a boundary should be willing to be asked about it. Every answer below is checkable in the product.

What in Hawzu actually uses AI?

Coverage mapping, a second pass that checks proposals against your existing cases, test case writing, documentation gaps, the release readiness narrative, Atlas map suggestions, rewrite, duplicate defect detection and Oracle. Duplicate detection compares defects by meaning rather than by keyword. Everything else that looks like AI isn't: the readiness score, Atlas impact tiers, chart recommendations, Observatory insights and flaky detection are arithmetic and rules.

Does AI decide whether a release is ready to ship?

No. The readiness score is fixed arithmetic over four parts of the release: how much has been run, how much of it passed, the state of its defects, and how many requirements are fully verified. Two people looking at the same release always get the same number. The verdict — Ready, At risk, Not ready or Not started — checks hard bars first, so an open blocker means Not ready whatever the score. AI writes the plain-English summary and the top risks underneath it; it never touches the score or the verdict.

What actually gets sent to the AI?

It depends on the feature, and every one of them shows what it read on screen. Generation and the readiness narrative send the relevant parts of your Canon documents, and the narrative also gets that release's own numbers. The coverage judge is shown what your repository already covers. Rewrite sends the field you are editing. Oracle's questions about records send only the shape of your project — what your fields are called and the values they hold — and never the contents of a record. Duplicate detection only compares your defects with each other by meaning.

Can it change my project on its own?

No. There is no background AI writing to your project. Generated cases are a preview you tick through, and nothing is saved until you accept it. A rewrite shows the exact before-and-after and applies only when you say so. An Atlas suggestion is reviewed item by item, and nothing reaches the map until a person applies it.

Can I tell later which test cases came from AI?

Yes, permanently. An accepted case carries an AI badge for the rest of its life, and opening it shows the scenario that produced it and the focus lens that proposed it. A case written by hand carries no marker at all, so a repository that has never used generation shows no extra chrome. It records where the case came from — it is not a review state anyone is waiting on.

Do I need Canon documents for the AI to work?

It depends on the feature. Generation uses your Canon by default and you can turn that off for a one-off, so it works either way. The release readiness narrative always reads it. Atlas map suggestions need Canon outright and refuse rather than guessing at your product's structure without it.

What does it do when it doesn't know?

It says so, in the shape the question arrived in. A behaviour your specification names without giving an outcome comes back as a documentation gap in its own list — kept out of the drafts, because there is nothing to assert yet. Asked for thirty scenarios from a thin source, it returns the number the source actually supports and says why. And the documentation assistant answers that it couldn't find the information rather than writing four confident paragraphs.

Are there AI credits or usage limits?

There is a monthly allowance — 50 AI credits on Free, 250 per member on Growth, pooled across the workspace — but there is nothing to buy and no balance to top up. You are never shown a token count or a running cost, nothing appears while you work unless you get close to the allowance, and if a workspace does reach it, AI picks up again at the start of the next month. You are never billed for going over, because you cannot go over.

Upstream

The AI is only as good as what it's given

Which is why the features that decide what it reads have pages of their own.

Give it your spec. Check its work.

Every AI feature on every plan. Free for five people, $20 per member after — no credit card.

Talk to Us

Tell us about your QA setup. We'll get back to you within 24 hours.

Book a demo

Pick a time that works — we'll confirm by email and send a calendar invite.

Select a date

Available times

Times shown in your timezone: —