AI Unit Testing, Explained: What Generated Tests Are Worth, How to Verify Them, and Where They Fail
AI unit testing explained: how tests are generated from the code that changed, why running them is the only proof, the mirror-test failure mode, coverage vs value, and the cost.
Written by Max Zeshut
Founder at Agentmelt
TL;DR: AI unit testing means a model writes the tests for a function — usually the functions that changed in a pull request and have no tests — and the only thing that makes those tests worth anything is running them. A generated test earns its place if it passes against the current code, adds coverage the suite did not have, and does not flake; everything else is discarded before a person sees it. The failure mode to design against is the mirror test: a test that restates the implementation and passes no matter what the code does. Generate from the function's contract, verify by running, and let the author review the survivors like any other pull request.
Buildable version: the unit test generation workflow — what arrives, what happens, who approves, a free template, and the price to have it run for you.
What "AI unit testing" covers, and what it does not
Three different things get called AI testing:
- Test generation — the model writes unit tests for existing or changed code. This article.
- Test maintenance — self-healing end-to-end tests that relocate an element when a selector changes. A different problem, covered in the automated testing guide.
- AI-assisted test design — a model proposing test cases from a requirement, for a person to write. Useful, and not automation.
Generation is where the risk is lowest and the value clearest, because the output is code that can be executed. A generated test is not an opinion; it either passes or it does not, covers new lines or it does not.
How generation works when it works
The pipeline that produces tests worth keeping, as run in the workflow:
- Find the untested change. The diff of the pull request mapped against the CI coverage report: functions and branches that changed and have no covering test. Generating tests for everything is how you get a thousand tests nobody reads.
- Assemble the context. The changed file, its imports, the existing test file for the module — for the framework, the fixtures, the mocks and the naming — and the test configuration. Tests written without the module's existing tests in view use different fixtures and look foreign.
- Generate from the contract. The prompt is built from the function's signature, its documentation and its callers — what it promises — rather than from its body. This is the single most important choice: a model shown the body writes tests that assert what the body does, including the bug.
- Cover three paths per function: the happy path, the edge cases visible in the signature and the callers (empty input, boundaries, nulls), and the error path.
- Run in CI on a branch with the existing suite, and collect the coverage delta.
- Keep only what earns it: tests that fail are discarded, or retried once with the failure output; tests that pass but add no coverage are dropped; tests that pass on one run and fail on another are flagged as flaky and dropped.
- Open a companion pull request with the survivors, a summary of what they cover and the coverage delta, assigned to the original author.
The author reviews and merges, or edits. Their edits are collected weekly and fed back into the prompts — the reviewer's corrections are the training data.
The mirror test, and the other failure modes
The mirror test asserts the implementation: expect(calculateTotal([10, 20])).toBe(30) is fine; expect(calculateTotal(items)).toBe(items.reduce((a, b) => a + b)) is the function rewritten as an assertion, and it passes when the function is wrong. Models produce these readily when shown the body. Defences: generate from the contract, reject tests whose assertion duplicates the implementation's logic (a reviewer spots it in seconds; a heuristic catches the obvious cases), and prefer concrete expected values.
Over-mocking: mocking the collaborator so thoroughly that the test exercises nothing. Detected by the coverage delta — a test that adds no covered lines is dropped.
Testing the framework: asserting that a getter returns what the setter set. Adds coverage, adds nothing. Reviewers tune this out; the acceptance rate shows whether it is happening.
Flakiness: tests with time, randomness or ordering assumptions. Two runs before acceptance catch most; a monthly flake report catches the rest.
Coverage as the goal: a suite at 85% coverage with 40% mirror tests is worse than one at 65% with none. Track coverage on changed code and the acceptance rate of generated tests, not the headline number.
What to expect
From teams running generation on every pull request, with the pipeline above:
| Measure | Typical range | What moves it |
|---|---|---|
| Generated tests that pass on first run | 60–80% | Context quality, framework maturity |
| Survivors after coverage and flake filters | 40–60% of generated | Strictness of the filters |
| Accepted by the author without edits | 60–80% of survivors | Prompt tuning from prior edits |
| Coverage on changed code, before → after | 40–60% → 75–90% | Which functions are targeted |
| Time from pull request to companion PR | 3–8 minutes | CI speed |
The honest headline: roughly half of what the model writes survives, and most of what survives is accepted. That is a good trade for a step that costs cents and minutes, and a bad one if nobody reviews the survivors.
Languages and frameworks
Anything with a coverage report CI can produce: TypeScript and JavaScript (Jest, Vitest), Python (pytest), Go (go test), Java and Kotlin (JUnit), C# (xUnit, NUnit), Ruby (RSpec). The generated tests follow the module's existing tests — fixtures, mocks, naming — so they read as the team's own. Monorepos with a custom build graph, integration tests with seeded data or external services, and repositories that cannot send code to a model provider are the custom builds.
Cost and privacy
Model usage per pull request is cents to a few dollars, driven by the size of the changed files and their tests. As an installed workflow: $249 one-time in your own CI, or a monthly subscription for up to 500 pull requests a month. Privacy: the changed file, its direct imports and the module's tests go to the model provider for the generation call under a zero-retention agreement, and nothing else; teams that cannot send code run the same workflow against a model in their own cloud. The CI, the branch and the pull request never leave your version-control system either way.
Questions, answered
Does generative AI write good unit tests?
It writes plausible ones; running them is what makes them good. A generated test is worth keeping only if it passes against the current code, adds coverage the suite did not have, and does not flake across two runs. Tests that restate the implementation are the main failure mode, so the generator works from the function's contract — signature, documentation, callers — rather than its body, and the reviewers' edits tune it weekly.
Can AI unit testing replace writing tests by hand?
For the mechanical majority — the happy path, the visible edge cases, the error path of a function that changed — yes, and that is most of the tests a team never gets round to writing. For tests that encode a business rule the code does not make obvious, or a regression that needed a specific scenario, a person still writes them, and the generated ones give that person time.
How do you use AI to write unit tests safely in a real codebase?
Target only changed, untested functions; give the model the module's existing tests as the style guide; generate from the contract; run everything in CI before anyone reads it; discard failures, zero-coverage tests and flakes; and deliver the rest as a pull request the author reviews. Never merge generated tests automatically — a test suite is a promise, and promises are made by people.
Which unit test frameworks work with AI test generation?
Jest, Vitest, pytest, Go's test package, JUnit, xUnit, NUnit and RSpec, among others — the requirement is a coverage report CI can produce and an existing test file the generator can imitate. End-to-end frameworks such as Playwright and Cypress are a different pipeline.