I remember shipping a new AI-driven feature where I had written painstaking acceptance criteria. Every item on the Definition of Done checklist was checked off, yet users still hated the result. We had, for example, required “the system returns the correct answer 90% of the time.” Our QA passed all test cases, but in real life the AI answers were half-baked, inconsistent, or just plain wrong. What went wrong? We had treated an AI output like deterministic software, assuming a single “correct” response. But AI features are non-deterministic: the same input can yield different outputs each run. A binary “yes/no” acceptance criterion simply doesn’t capture that.
Today’s product owners face this exact challenge: traditional Scrum-style acceptance criteria and checklist Definitions of Done break down under AI. Instead of pass/fail rules, we now need graded evaluation rubrics, often called evals, that score AI output quality. This article will explain why old checklists fail, what these eval rubrics look like, and how to write them as a product owner. We’ll draw on industry examples (like OpenAI’s own deprecation of its Evals platform in 2026) and foundational work (Hamel Husain’s 2024 guide on AI evals) to chart a practical path forward.
Even if you’ve trained in agile best practices (like the Refonte Learning Product Owner Program), AI feature quality brings new pitfalls. The program covers backlog management and data-driven metrics, but it doesn’t yet teach how to define “good enough” for AI outputs. This article fills that gap. You’ll learn why “the system returns the correct answer” no longer suffices, and how to replace it with eval-based criteria that set clear, measurable quality targets.
The Checklist That Passed and the Feature That Failed
I’ve seen it too often: a product-owner-crafted feature that ticks all the boxes but fails the user. For example, I once led development of an AI-powered search feature. The user story’s acceptance criteria were straightforward: “Given a query, the AI returns a relevant answer.” We wrote unit tests for key queries, marked them done, and demo’d to stakeholders. The Definition of Done on our Scrum board was green. Yet beta users reported that search results were off-topic or incomplete. We realized the AI’s output was only mostly correct, not the guaranteed correct answer our criteria assumed.
This failure happened because our acceptance criteria were binary and deterministic: they assumed the feature would either meet a condition or not. In reality, the AI’s answers were probabilistic and often on a spectrum of quality. We needed a more nuanced way to judge success. A checklist item like “returns the correct answer” doesn’t define how often or to what degree the answer must be correct. In other words, we had a definition of done, but it was inadequate for AI.
The takeaway: an AI feature can satisfy every checklist item and still disappoint. Traditional Scrum views features as “done” only if all criteria are met, but for AI those criteria must change. They need to capture acceptable levels of correctness, relevance, and safety, not just existence of a response. In the next sections, we’ll dig into exactly why AI breaks our old rules and what we write instead.
Why AI Features Break Traditional Acceptance Criteria
Most acceptance criteria assume a deterministic system: if input X is given, output Y must be produced. An AC might say “Given ticket ID 123, the search returns the ticket with that ID.” For AI features, this logic fails because the output is probabilistic, not fixed. The same user prompt can yield different answers each time, and correctness itself can be fuzzy.
Probabilistic vs. Deterministic. Traditional features (databases, APIs, business logic) return consistent results. AI models (LLMs, generative systems) return a distribution of possible outputs. A criterion like “returns the correct answer” implies a single correct reference, which doesn’t exist in open-ended AI tasks. Instead, we must ask “How often and by what measure is the answer acceptable?”
Quality over Binary. With AI we care about how good an answer is, not just right/wrong. A chatbot may give a partially correct response or mention extraneous details. Our AC must shift from “Answer is correct” to something like “Answer is relevant and accurate at least X% of the time, and contains no harmful content.” This is a graded standard, not a yes/no pass/fail.
Ambiguity in Output. AI language and image outputs are subjective. Two developers or stakeholders might disagree if a given response meets the bar. An AC like “the tone is friendly” has no single truth. We need objective metrics or rubrics to make these judgments.
In short, non-deterministic AI product features demand probabilistic, metric-based criteria. Instead of “this checkbox is either checked or not,” we ask “does this feature achieve our quality goals consistently enough?” We’ll see that this leads to writing evals, essentially scoring rubrics, rather than simple acceptance criteria.
Probabilistic Output vs. Deterministic Requirements
Consider a concrete example: an AI summarizer. A traditional AC might have been: “When given an article, the system returns the key points.” That sounds reasonable, but it is too vague: Which key points, and how do we judge correctness? Each run of the summarizer might highlight different details. The old-style criteria would say “done” if there is some summary, but we really need to quantify quality. In other words, we must accept that the output is probabilistic and set thresholds. The revised requirement could be: “The summary must cover at least 80% of user-identified key points and no factual inaccuracies.” This explicitly acknowledges variation and requires measurement.
When acceptance criteria become metric-driven, we can also track them over time. For example, we might say: “Over 100 sample documents, at least 85% of them produce a summary that human reviewers score at least 4/5 on relevance.” This is clearly not a simple checklist item, but a quality target. In the parlance of modern AI product management, we are moving toward an LLM evaluation framework for product teams. Instead of binary rules, we use statistical or scoring criteria.
Common Agile frameworks like Scrum don’t currently prescribe how to write such criteria. The Scrum Definition of Done traditionally covers coding, testing, documentation, etc. For an AI feature, we have to augment that DoD with performance targets. For example, “Answers must score at least 0.8 on our relevance rubric for 95% of test cases” could be part of Done. This is a very different mindset: we are accepting that “100% always right” is unrealistic, and we define what “good enough” means.
In practice, this means acceptance criteria for AI features look more like quality-of-service metrics. That’s why the term AI feature quality assurance is becoming popular. We no longer say “it must return answer Y” but rather “it should achieve at least X accuracy or usability score.” This idea is sometimes called eval-driven development: we build evaluation into the process from day one, not just at final test. Next, we’ll define what these “evals” actually are.
What an Eval Actually Is
An eval (short for evaluation) is essentially a structured test suite for AI, combined with a scoring rubric. It replaces the old pass/fail checklist. In a product context, think of an eval as a set of usage scenarios plus criteria for judging the output. For each test case (e.g. a user prompt or input), we don’t just check “did it match expected output?” Instead, we apply one or more graders or metrics to the output. Those could be automated (like comparing answers to a reference key) or manual (human rating). The result is a score or label that reflects output quality.
Hamel Husain’s influential 2024 essay “Your AI Product Needs Evals” crystallized this concept. He describes evals as the core of a robust AI development process: “AI evals show where an AI product fails so you can improve how it works in production.” In practice, an eval might include: test inputs, model outputs, and expected behavior. For example, for a QA feature, test inputs could be a set of questions; the model outputs are answers; and the eval could measure whether each answer is factually correct, relevant, and delivered promptly.
As a product owner (PO), writing an eval means creating acceptance criteria in the form of a graded rubric. Instead of “return correct answer,” you’d define “correct enough.” For instance, an eval for a summarization feature might say: “Each generated summary should have at least 90% overlap in key points with the reference summary, as judged by a human or LLM grader.” The criterion is explicitly measurable and has a numeric target. This addresses the nondeterminism: some variation is allowed, but scores below the threshold would trigger investigation.
Industry commentary confirms that top PMs are embracing this language. A June 30, 2026, Lenny’s Newsletter guest post explicitly told PMs to “set up evals to automate and improve the quality of your work.” In other words, the phrase “evals for product managers” is already entering the standard vocabulary of savvy product teams. That newsletter lists evals alongside prototyping tools and AI analytics as part of a modern PM toolkit. We treat it as evidence that leading PMs are thinking this way. Product owners should take the cue.
In short, an eval is an AI-specific deliverable: a test plan plus scoring system. It is the new artifact we produce in place of (or alongside) a traditional acceptance criterion. Writing an eval is a different skill: you must decide what metrics matter, how to measure them, and what target constitutes “done.” The rest of this article will guide you through exactly how to do that.
The Three-Level Framework: Assertions, Graded Review, Live Testing
Husain’s framework divides evals into three levels. This helps teams choose the right approach at different stages:
Level 1: Unit-Test Style Assertions: Write fast, automated checks for obvious cases. For example, if your AI feature calls an external API or does simple data processing, you might assert that certain edge cases always produce a known response. These are like software unit tests: they must pass every time (or the build fails). POs typically define the scenarios (e.g. “when given input A, at least one correct keyword appears”) and engineers implement the quick checks.
Level 2: Model/LLM or Human Evals: This is where most AI evaluation happens. We collect a set of realistic inputs (like a variety of user questions) and get outputs. Then each output is graded. The grader might be a human team member, a crowd worker, or even an LLM acting as a judge. For instance, you might have a rubric that rates answer correctness on a 1-to-5 scale, or flags any unsafe content. These evaluations provide a numeric or categorical score for each case. They’re slower and more expensive than Level 1, but they capture the nuance of AI output.
Level 3: Live A/B Testing or Monitoring: Ultimately, real-world user feedback is the gold standard. This involves rolling out the feature (perhaps only to a test segment) and measuring key business metrics. For example, you might compare user engagement or error rates between the old and new AI versions. Or you might have a “rate our answer” prompt for users. Level 3 tells you if your AI actually improves the product experience. It’s costly and usually done after reaching some confidence with Levels 1 and 2.
These levels have increasing cost and impact. In practice, POs often focus first on Level 1 and 2 during feature development, and reserve Level 3 for a launch or continuous monitoring phase.
Where a Product Owner Fits in Each Level
As a PO, your role shifts at each level: you’re the architect of what to evaluate.
Level 1 (Assertions): Here you define the critical scenarios or edge cases. For instance, if a chatbot should answer math questions, a PO might specify a few math problems with known answers and say, “the sum must equal X.” The PO decides what to check; engineers write the code and automated tests. Your task is to ensure these assertions cover the core feature behavior.
Level 2 (Graded Evals): You decide what to score and how. You create or curate test cases (e.g. representative user queries), then define the rubric. Maybe it’s a checklist: “Answer is correct / not correct,” “Tone is friendly / not friendly,” etc. You may ask colleagues to grade outputs or use a model to grade. The PO must review the scores, analyze failure patterns, and iterate on the prompts or requirements. You own the definition of “pass” at this level, even if engineers run the evals.
Level 3 (Live Testing): You set the evaluation metrics (e.g. user satisfaction, retention). After a rollout, you own the interpretation: “At 85% accuracy, should we ship or refine more?” You translate A/B test results into product decisions. Here, engineers focus on instrumentation and data collection, while you focus on business interpretation.
In summary, the PO is responsible for defining the metrics and goals at each level. Engineers may handle the technical mechanics (writing the test scripts, running CI jobs, or instrumenting analytics), but the PO guides what gets measured. This collaboration ensures the team tests the right things.
Level | Who Defines It | Who Runs It | What’s Checked |
1. Unit-Test Assertions | Product Owner writes scenarios and expected outcomes. | Engineers implement fast, automated tests. | Key edge cases (e.g. “normalized answer contains expected value”). |
2. Graded Evals | Product Owner sets criteria and scoring rubric. | Engineers (or PMO) gather outputs; humans or LLMs grade them. | Quality of responses (relevance, accuracy, safety, style, etc.) in realistic scenarios. |
3. Live Testing | Product Owner decides success metrics (e.g. user ratings, error rate). | Engineers deploy feature for A/B test; analytics team collects data. | Business/UX impact (user satisfaction, conversions, support tickets). |
Table: Three levels of AI feature evaluation. The PO defines “what good looks like” at every level.
Why OpenAI Just Killed Its Own Evals Platform
If you needed a sign that evals are important, consider this news: on June 3, 2026, OpenAI announced in its developer changelog that it would deprecate the Evals platform. OpenAI’s deprecations documentation confirms that existing evals will become read-only on October 31, 2026, and that the Evals dashboard and API are scheduled to shut down on November 30, 2026. In the same announcement, OpenAI directed users to migrate to Promptfoo, an open-source alternative discussed next.
This means that even OpenAI, which originally built a tool for running evals, is moving away from a proprietary solution. The platform it launched to help engineers test models is being retired, giving users roughly six months’ notice. The implication is that evals have become so integral that companies will support them through external tools. For product owners, this underscores a new reality: building and maintaining evals is not an optional engineering add-on; it is a core part of AI product development. We cannot rely on built-in vendor tools forever; we must own our eval process.
To be clear, this announcement is part of a broader shift. OpenAI also deprecated its Agent Builder and “reusable prompts” on the same date, all pointing toward a future where open tools (or homegrown ones) handle these tasks. The takeaway for PMs is: invest in evals now, because industry momentum (and platform support) is moving in that direction.
What Replaces It: Promptfoo and the Open-Source Shift
With OpenAI sunsetting Evals, many teams are considering Promptfoo, an open-source CLI and library for AI evaluation. Promptfoo is not built by OpenAI; it is community-driven. OpenAI’s official migration guide recommends Promptfoo for continuing and extending evaluation workflows after the hosted Evals product is retired. Promptfoo lets teams define evals in YAML configuration files and run them locally or in CI, creating a portable, code-centric workflow.
Promptfoo positions itself for developers, security teams, and product teams that care about AI quality. Its website says that 156 of the Fortune 500 use Promptfoo in their AI development lifecycle and advertises more than 300,000 developers worldwide. Those figures are Promptfoo’s own marketing claims, not independently verified adoption data. The company also highlights “zero vendor lock-in,” which appeals to buyers wary of being tied to one cloud vendor.
The move to Promptfoo reflects a larger trend: the industry is embracing open, flexible tools for evaluation. As a PO, you don’t need to care about the tech internals of Promptfoo, but you should know it exists. The key point is that the concept of evals has outgrown any single platform. Product teams will either use something like Promptfoo or build their own evaluation pipelines. In either case, our job as owners is the same: create and maintain those eval definitions.
It is wise to read vendor claims like Promptfoo’s skeptically. The “156 of the Fortune 500” figure is a vendor marketing claim, not an independent study. Still, the underlying message is real: companies are investing in AI evaluation. For product owners, the lesson is that moving away from rigid checklists to evals is becoming standard practice. Whether a team uses Promptfoo or another tool, product teams increasingly need evaluation suites as part of feature development.
Writing Your First Eval as a Product Owner
Now let’s get practical. How do you, as a PO, write an eval? It starts with your user story. Suppose the story is: “As a customer, I want the chatbot to answer product questions accurately.” Instead of an AC like “the answer is correct,” we define an eval around that goal.
Choose test cases. Pick realistic inputs to represent user queries. For example, select 20 common product questions customers ask. These will be your eval prompts.
Define quality criteria. Decide what makes an answer “good.” This becomes your rubric. Maybe criteria include correctness, relevance, completeness, and no offensive language.
Set target thresholds. For each criterion, set a pass threshold. For example, “At least 90% of answers should be fully correct,” or “No answer should be rated below 3/5 on quality.” These turn vagueness into numbers.
Select the grader. Determine how you’ll grade. Will you have team members review outputs? Or use an LLM judge (e.g. GPT with a judgment prompt)? This influences how you structure the eval.
Write the eval test. Create a spreadsheet or config file that lists each test input, expected behavior, and how to score it. For a simple test, you might note the “correct answer” and mark it manually. For more complex, you might mark key facts to check.
For example, one eval item might be:
Test Input: “What is your return policy on electronics?”
Criteria: “Answer mentions free returns within 30 days and 24/7 support.”
Rubric: The answer is correct (1) if both points are covered; partially correct (0.5) if one; incorrect (0) if neither.
During sprint planning, you’d add “Write chatbot eval” as a task, just like writing unit tests. This is a very different task than drafting a user story or backlog with Jira AI tools (see AI Backlog Tools for Product Owners). The backlog tools help create the story, but writing an eval requires thinking through quality scenarios. It’s a creative, human-driven deliverable.
To write your first eval, think of it as a test suite with scoring. You are no longer saying “done when X happens once,” but “done when our quality metrics pass.” The resulting artifact might be a spreadsheet, a Markdown document, or a Promptfoo configuration file, but its purpose is clear: it defines “good enough” for the feature.
From Acceptance Criteria to Grading Rubrics
A useful way to see the shift is in a side-by-side comparison. On one side is the old acceptance criterion; on the other is the new eval-based rubric.
Traditional AC: “The AI returns the correct answer.”
Eval Rubric: “Run 50 test questions. An answer is correct if it matches the reference on all key facts. At least 45/50 answers must be correct (90% accuracy).”Traditional AC: “The image generator produces a realistic human face.”
Eval Rubric: “For 100 generated images, a third-party tool or human must rate image quality. At least 90% should score at least 4/5 on realism, with no uncensored sensitive content.”Traditional AC: “No known bugs in critical path.”
Eval Rubric: “Automated tests cover critical queries and assertions. All unit-test style assertions (Level 1 evals) must pass on every build. Additionally, run the Level 2 human eval suite weekly and track any issues.”
We can also tabulate differences:
Aspect | Traditional AC (Deterministic) | Eval-Based Rubric (Probabilistic) |
Output expectation | One fixed “correct” response. | Range of acceptable responses, scored by criteria. |
Condition form | Boolean (pass/fail). | Numeric or categorical (score, percentage). |
Testing process | Typically manual or simple automation (does X happen?). | Mix of automation and human review, possibly CI integration. |
Pass threshold | 100% (every time). | Tuned target (e.g. ≥80% on metrics) plus error tolerance. |
Scope | Single example or scenario. | Aggregated over many cases, measuring distribution. |
This comparison shows why we can’t just copy old practices. Under the old model, either “returning any answer” passes or fails. The new rubric asks: “Does the set of answers meet our quality target?” In Agile terms, the Definition of Done must expand. A story is done when its results meet these statistical criteria. In other words, we’ve turned acceptance criteria into something like “acceptance fractions” or “scorecards” instead of a yes/no list.
Importantly, writing an eval rubric requires collaboration. As a PO, you bring customer perspective (what quality means). Engineers and QA bring technical insight (how to measure it). Together, you specify things like how many test cases, how to score partial credit, and what to do if targets aren’t met. This is inherently more complex than ticking a box, but it’s how you gain confidence in an AI feature’s performance.
How This Changes Sprint Planning and Definition of Done
Introducing evals into your workflow will shake up sprint rituals. Suddenly, sprint planning must include tasks like “Define evaluation metrics” and “Create test prompts and expected outcomes.” These were not traditional backlog items, but they should be. Team velocity now needs to account for the time spent on evaluation design and analysis, not just coding.
The Definition of Done (DoD) in a sprint also changes. In classic Scrum, the DoD might include “All acceptance criteria passed,” “Code reviewed,” and “QA approved.” With AI features, the DoD should also include meeting specific evaluation targets. For example:
At least 95% of unit-test style assertions must pass (Level 1).
Average human eval score at least 4/5 on 20 example cases (Level 2).
No critical failures on core user scenarios.
These become new exit conditions. If the AI results do not meet them, the story is not done, even if “all tests pass.” This shifts some work into the “done” column: instead of handing off a feature that merely runs, we hand off a feature that is demonstrably high quality by our own standards.
In sprint reviews and demos, you will present evaluation summaries alongside feature screenshots. For instance, “This week’s chatbot returned the correct policy information in 47 out of 50 cases, exceeding our 90% target.” The Scrum board might show an “Evaluation” column after “Development” and “Testing.” Stakeholders start asking for these metrics, turning the Definition of Done from a checklist into a compact performance report.
This evolution also affects acceptance sign-off. Rather than saying “QA passes, so product is shipped,” you’ll say “We ran X eval tests, and quality meets the rubric.” In effect, evaluation results become part of the acceptance criteria itself. This ensures that finishing a sprint is about delivering value, validated by data, rather than merely delivering code.
Working With Engineers on Shared Eval Ownership
Eval-driven development is a team sport. As a product owner, you lead on defining the eval, but engineers bring it to life. You own the what (what to measure) and why, and engineers own the how (how to implement it). Clear collaboration is key.
For example, the PO might list “User intent recognition accuracy” as a metric. Engineers will then write the code to run test queries and collect results (maybe using Promptfoo or a custom script). If a human grading task is needed, you’ll coordinate who grades it or integrate it into QA time. In agile terms, tasks under your user story might look like:
PO writes 20 eval prompts and a scoring guide.
Engineer automates running the model against these prompts and logs outputs.
QA tester or analyst scores the outputs based on the rubric.
Team reviews scores and either closes the story or iterates.
Where the PO’s Job Ends and the Engineer’s Begins
Product Owner Responsibilities:
Define what matters: Which behaviors or qualities to evaluate (e.g. “no profanity”, “relevance to query”).
Create or approve test cases: Provide realistic inputs and expected criteria (could be a list of key facts or a grading checklist).
Set targets: Decide acceptable scores or percentages.
Review evaluation outcomes: Analyze scores, spot failure modes, and decide if more work is needed or if the feature is good enough.
Engineer Responsibilities:
Implement infrastructure: Set up evaluation harness, CI jobs, or use tools like Promptfoo to run tests.
Automate data collection: Ensure the model outputs are captured and logged for review.
Assist with rubric definition: Advise on what can be measured automatically vs. manually.
Iterate on system: Use eval feedback to improve prompts, model parameters, or code, then re-run evals to check progress.
For instance, in an evaluation collaboration a PO might say, “We need the model to correctly parse shipping addresses from text 90% of the time.” The engineer might propose splitting that into automated regex checks (Level 1) and human review of ambiguous cases (Level 2). Then the PO would use the results of those tests to accept the story or refine the requirement.
The key is communication. Product owners should not simply hand off a document and step back, and engineers should not unilaterally decide “done” without consulting the PO’s goals. An effective workflow is iterative: run the eval early and often, share the results, then adapt requirements together. This shared ownership of eval outcomes helps ensure the final product meets the intended quality bar.
Common Mistakes Product Owners Make With AI Features
When teams are new to this, several pitfalls often occur:
Treating AI like traditional software. Writing plain “the correct answer” ACs, then being surprised when outputs vary.
Writing only one test case. Assuming one example is enough. In reality, AI can fail many different ways, so you need a suite of varied test cases.
Ignoring distributional performance. Only looking at the average score or a single percentile. It’s better to check the worst-case or outliers too (e.g. “What if the AI goes off track?”).
Neglecting to involve users or SMEs in grading. Assuming developer intuition is enough to define “good.” Human or expert judgment is crucial at Level 2.
Relying entirely on automatic LLM judges. LLM-as-judge is powerful, but if you just plug a model to evaluate outputs without oversight, you might miss subtle errors. Always validate a new LLM judge with a human-in-the-loop at first.
A particularly common mistake is treating an eval score as a pass/fail gate. For example, a team might say, “We set accuracy at 90% or higher; if it is 89%, we stop the feature.” This misunderstands the purpose of evals. The score is a guide, not an absolute decision-maker. An 89% result may be acceptable if the remaining mistakes are inconsequential, while 95% may still be unacceptable if the 5% of errors are catastrophic.
In practice, the eval results should spark discussion, not unthinking gating. Use them to understand where the feature fails and then iterate. Do not ignore a 90% pass rate; it is telling you something. The true “pass” condition is: Does the evaluated performance meet the product’s needs? Sometimes that means adjusting expectations, not blindly throwing out the build.
In summary, avoid the checklist mentality. Embrace evals as living criteria. Use the scores as metrics to improve and decide how to iterate, rather than as a binary pass/fail stamp.
How This Fits Into Stakeholder Communication
Defining feature quality with evals also changes how you talk to stakeholders. Instead of saying “the feature is done,” you’ll be saying things like: “Our chatbot is 92% accurate on our test set, and average user feedback score is 4.3/5.” This shift from qualitative to quantitative updates can build trust because stakeholders see data, not just anecdotes.
It also means setting the right expectations upfront. When scoping an AI feature, involve stakeholders in agreeing on targets: “If we hit 80% accuracy, is that success?” This way, the Definition of Done includes their buy-in. If a stakeholder was expecting 100% always, you’ll need to coach them: explain that 100% is unrealistic for generative AI, but 80% with continual improvement is our approach. Aligning on these numbers early avoids surprises later.
You might draw a parallel to something Refonte Learning has already emphasized: its article on AI-Moderated Customer Interviews and Continuous Discovery shows how AI can gather user insights, while evals address the delivery side by validating the solution. In stakeholder conversations, you can say, “Before release, we are using an AI evaluation framework to verify that the feature meets user needs.” Stakeholders appreciate a data-driven method for defining quality.
Finally, share eval results visually. Include screenshots of evaluation dashboards, charts of score distributions, or examples of model output successes and failures. This transparency is powerful. Instead of a bug report or “we found edge cases,” you can say “We ran 100 test cases; 90 passed the rubric. Here are the 10 that did not, along with our plan to fix them.” That kind of communication demystifies AI features and keeps everyone on the same page.
Product Owner Skills and Salaries in 2026
Becoming adept at this new style of QA is a matter of evolving your skill set. It draws on competencies like data-driven decision-making and backlog prioritization, which are among the skills the Refonte Learning Product Owner Program teaches. Thinking in metrics, crafting test scenarios, and analyzing results are all part of being a data-savvy PO. The program also covers stakeholder management and agile frameworks, which will help you convince others to invest in this extra work.
In terms of market demand, product owners are in demand, but pay should be discussed realistically. Refonte’s program page markets a “$220,000+” starting salary and “175,000+” jobs annually for Product Owner positions. Those are Refonte Learning’s own promotional figures, and the salary claim is notably higher than typical public reporting. As verified on August 26, 2026, Glassdoor lists an average U.S. Product Owner salary of $142,356 per year, while ZipRecruiter lists an average of $112,891 as of August 25, 2026. ZipRecruiter reports that most salaries fall between $93,500 and $129,500, with top earners around $150,000, which places the program’s $220,000+ figure well outside the typical range.
Regardless of the exact figures, product owners who can handle AI features may command a premium. The public salary sources do not explain how Refonte calculated its higher figure, so readers should not treat it as a guaranteed starting salary. In practice, salaries vary by location, industry, and experience. Learning how to handle AI features, including writing evals, remains a differentiator because it builds on agile and analytics skills such as backlog prioritization, sprint planning, and data-driven metrics.
If you’re comparing certifications (CSPO vs PSPO vs SAFe POPM), keep in mind this new trend. Refonte’s recent blog CSPO vs. PSPO vs. SAFe POPM explains certification options, but regardless of your track, incorporating AI evaluation into your repertoire will set you apart. Companies seeking Product Owners for AI products will look for experience with AI QA.
Building This Skill Set: The Refonte Learning Product Owner Program
The Refonte Learning Product Owner Program is a three-month program requiring 8-10 hours per week. Its curriculum includes three core modules: Introduction to Product Ownership; Backlog Management and Prioritization; and Stakeholder Communication and Collaboration. The coursework covers competencies such as backlog prioritization, sprint planning, and data-driven decision making, all of which underpin a disciplined approach to feature quality.
In this program, you’ll also develop skills in collaboration and Agile metrics. These are crucial when you start defining and measuring AI feature success. For example, the Backlog Management module teaches how to maintain a healthy backlog, where you’ll now include eval-writing tasks. The Data-Driven Decision Making competency trains you to use analytics, which ties directly into evaluating AI outputs.
The program is led by Professor Kevin Harris, a seasoned Agile Product Owner with over 10 years of experience at Fortune 500 companies. He mentors students on real-world projects and best practices. You’ll also get hands-on practice with the tools named in the program: Jira, Confluence, roadmapping tools, and Agile frameworks like Scrum and Kanban. (Think of these as your toolkit for managing the process; the principles you learn apply whether you’re using Atlassian or other platforms.)
Upon completion, graduates of the program can pursue roles like Product Owner, Scrum Product Manager, or Agile Business Analyst. Earning this certificate signals employers that you’ve mastered the core PO skill set. As AI features become common, those skills, augmented by an understanding of eval-driven QA, will make you even more competitive.
The Refonte Learning Product Owner Program does not currently list writing evals or AI-specific quality assurance in its curriculum. Its Backlog Management and Prioritization module and Data-Driven Decision Making competency are the legitimate foundations for this skill, not evidence that the syllabus already teaches it. The program builds skills in backlog prioritization, user story mapping, sprint planning, stakeholder communication, and metric-based decisions. From there, you can extend those foundations into AI-specific practices such as eval-driven development and defining quality for probabilistic outputs.
If you’re ready to master the fundamentals of product ownership (and then take on the challenge of AI features), check out the Refonte Learning Product Owner Program. In just three months, you’ll learn the vision, tools, and metrics that modern Product Owners use (even as they prepare to adopt evals for AI). Let this program be your springboard: we’ll teach you how to manage the backlog and data analysis, and you bring those skills to the next level by applying them to AI evaluation.
