QA engineers are getting AI from two directions at once. It’s arriving as a tool that writes test cases and fixes flaky locators, and as a new kind of software to test: features built on language models, which don’t return the same answer twice and can be talked into misbehaving. Testing LLM features also requires checking variation between responses, factual support, and attempts to bypass instructions.
The checks below separate using AI in a test workflow from evaluating a feature built on AI.
Generative AI for everyday testing work
| Use case | What the AI does | What you still check |
|---|---|---|
| Test plan drafts | Builds a test plan outline from a requirements document or user story | Coverage of the risks that matter most to this release |
| Test case generation | Writes positive, negative, and boundary cases from acceptance criteria | Edge cases the requirements don’t mention, which AI rarely invents on its own |
| Test data creation | Generates realistic but fake data sets, including awkward inputs | That no real customer data was used to produce it |
| Automation scripts | Converts a manual test case into Playwright, Selenium, or Cypress code | Assertions, since AI often checks that a page loads rather than that it’s correct |
| Self-healing locators | Suggests a new selector when the UI changes and a test breaks | That the test still checks the right element, not just any element that passes |
| Bug report drafting | Turns rough notes, logs, and screenshots into a clear reproducible report | Steps to reproduce actually reproduce the bug |
Agents for larger testing jobs
| Use case | What the agent does | Where you stay involved |
|---|---|---|
| Regression selection | Reads a change set and picks which regression tests are most likely affected | Approving a reduced suite before a release |
| Exploratory test runs | Clicks through an app toward a goal and reports anything unexpected | Deciding which findings are real defects |
| Performance test analysis | Runs load tests and explains where response times degrade | Setting the thresholds that count as failure |
| Security probing | Tries common attack inputs against forms and APIs in a test environment | Keeping it pointed only at environments you’re allowed to test |
| Bug triage | Groups duplicate reports, suggests severity, and routes to the right team | Final severity on anything customer-facing |
| Quality reporting | Builds a release-readiness summary from test results and open defects | The go/no-go recommendation |
The new job: testing features built on LLMs
When your product includes a chatbot, a summarizer, or an agent, exact-match assertions alone are insufficient because the same input can produce several acceptable outputs. Keep deterministic checks for schemas and permissions, and add evaluations for answer quality.
| What to test | The question you’re answering | How teams test it |
|---|---|---|
| Faithfulness | Does the answer stick to the source material, or does it invent facts? | Score answers against reference documents with tools like RAGAS or DeepEval |
| Relevance | Does it answer the question that was asked? | Graded evaluation sets with expected answers |
| Prompt injection | Can a user talk it into ignoring its instructions or leaking data? | A library of attack prompts run before every release |
| Consistency | Does it give materially different answers to the same question? | Repeated runs and comparing results |
| Refusals | Does it refuse things it should, and only those? | Test sets covering both harmful and harmless requests |
| Cost and latency | Is each answer fast and cheap enough at real traffic? | Load tests that record tokens and response time |
A repeatable evaluation set lets the team compare model or prompt changes against the same cases. Record failures and keep adding examples from production incidents.
Where to be careful
- Generated tests can pass while testing nothing meaningful. Review assertions with the same care you’d give a junior colleague’s first pull request.
- Self-healing can hide real bugs by quietly re-pointing a test at a different element. Log every heal and review them weekly.
- Don’t paste production data into public AI tools to generate test data.
- Build a test set with reference answers and known failure cases.
- Run the same cases against the proposed model or prompt.
- Check factual support, safety, response time, and deterministic rules.
- Review failures and rerun the set before approving a release.
Suggested workflow based on the use cases in this article.
Where to start
Pick test case generation for one feature and measure honestly how much you had to fix. Then, if your product has any LLM feature, build a small evaluation set of twenty to thirty questions with expected answers. Run the same set after each prompt or model change, and investigate any regression before release.
AgileFever’s Gen AI and Agentic AI for QA and Test Engineers course covers six areas in 12 hours: test plan and case generation, automation and self-healing scripts, bug report generation, regression, performance and security testing agents, AI and LLM evaluation testing, and bug triage and quality reporting agents. It currently costs $425 in the US (regular $650) or ₹21,000 in India (regular ₹35,000).
FAQ
Will AI replace manual testers?
It takes over a lot of repetitive test writing. It’s weak at noticing that something feels wrong to a real user, and at deciding what’s worth testing in the first place. Those parts of the job become more valuable.
Do I need to know Python?
For using AI to write tests, no. For building evaluation sets for LLM features, basic Python helps, since most evaluation tools are Python libraries.
What’s the difference between testing AI and using AI to test?
Using AI to test means AI writes or runs your tests. Testing AI means checking a product feature that itself uses a language model. The second needs new methods, covered in the table above.
What to read next
For use cases across other delivery roles, see 20 Use Cases for Generative AI and Agentic AI Across Every Agile Role. If you want to build AI systems rather than test them, GenAI and Agentic AI BootCamp: The Complete Guide covers the longer program.