Home / Content Hub / Blog

When AI Changes the Answer, QA Has to Change the Question 

AI systems can return different valid answers to the same input. Verica Solin on why testing AI systems means asking a new question about software quality.

Verica Solin Verica Solin Published 17 September 2026 · 4 min read
Share
When AI Changes the Answer, QA Has to Change the Question 

What does it mean for software to be correct?

After more than 15 years in QA, I thought I had the answer. Then I started testing AI systems, and an assumption I’d carried my whole career stopped holding.

Why traditional testing worked

Traditional testing rests on a simple contract: define the expected behavior, run the input, compare the result. If the result matches, the test passes. If it doesn’t, we investigate. That model gives us something concrete to verify: a known expected result, and a clear line between pass and fail.

It has served QA well for decades. It served me well too.

Working with AI made me question something I had taken for granted for years: what exactly are we testing when there is no single answer to compare against?

What changes when AI answers differently?

AI-powered systems can give a different response to the same input, and a different response does not automatically mean a wrong one. Two very different answers can both be perfectly acceptable. That was the first thing that stood out to me when I started looking at AI from a QA perspective.

In traditional testing, a difference from the expected result is a strong signal that something needs investigation. With AI, the situation is more nuanced. Detecting a difference stops being the problem. The harder part is understanding whether the difference matters.

For years, I was comfortable asking: “Did the system do what we expected?”

With AI, I increasingly find myself asking: “How do we know when the answer is good enough?”

That sounds like a small shift. What it means for testing is anything but.

Is an accurate answer enough?

Accuracy matters, but an accurate answer is not automatically a good one. When there isn’t one predefined correct answer, quality has to mean more than matching a value. That became clear to me the moment I tried to judge real AI outputs consistently.

What if the answer is accurate but irrelevant to what the user asked? What if it contains useful information but ignores important context? What if it is technically correct but doesn’t help the user achieve what they wanted?

These questions became particularly interesting while I was working on an AI-focused testing initiative. We quickly moved past asking whether an answer was technically correct. The harder problem was how to judge the quality of an answer consistently, when reasonable responses could look very different.

That experience changed the way I think about AI testing.

Rethinking what quality means

The real challenge goes beyond learning to test a new type of technology: AI forced me to reconsider what we mean by quality. As QA professionals, we have always had to understand what good software should do. With AI, that responsibility grows, because the output itself may not be fully predictable.

We need to think about what makes an outcome acceptable, useful, relevant, and trustworthy, even when there isn’t a single predefined answer.

This is where I believe the role of QA becomes even more interesting. AI does not make QA less important. It makes some of our old assumptions worth questioning.

If AI changes the way software behaves, QA has to rethink how it defines and evaluates quality. The shift is from asking only whether the output matches what we expected, to asking whether the output meets the quality we need.

The question I’m bringing to SEETEST 2026

If there isn’t always one correct answer, how do we decide whether an AI system has produced a good one? That is one of the questions I’ll be exploring at SEETEST 2026 in Ljubljana. I’ll share a practical perspective on how QA can approach the evaluation of AI-generated outputs, and what this shift means for the way we think about software quality.

My session, “Evaluation Prompts: The Future of QA,” runs on October 7 from 16:05 to 16:40, in person and online. Register for SEETEST 2026 and join me there.

Maybe AI isn’t just changing the answers.

Maybe it is time for QA to change the question.

Verica Solin

Verica Solin

Verica Solin is a Senior QA Consultant at IWConnect with over 15 years of experience in software quality engineering across e-commerce, insurance, IoT and telecom. She specializes in test automation, AI-assisted testing and quality strategy, and has worked in both QA leadership and Product Management roles. She speaks at SEETEST 2026 on October 7.

Curious how this applies to your numbers? Let's find out.

Share where things are getting stuck today and we will walk you through what a fix could look like.

Talk to our team