Acceptance and Evaluation Criteria
Acceptance and evaluation criteria are the written measures used to judge a solution. An acceptance criterion states a condition the solution must meet to be accepted by the stakeholders, in a verifiable form that settles as a pass or a fail. An evaluation criterion is a measurement scale on which several candidate solutions are scored, weighted and then aggregated into a value ranking. Both rest on the same material, the value attributes, those characteristics of a solution on which the value it delivers to stakeholders depends. Acceptance criteria are written when a single solution is in play and the question is whether it passes, evaluation criteria when several are still in the running and the question is which one is worth most. The technique can apply at all levels of a project, from the most general to the most detailed.
Purpose
Acceptance and evaluation criteria are the measures that make a judgement on a solution verifiable by a third party. They answer two questions. When a single solution is on the table, the question is binary: does it meet the conditions the stakeholders set for accepting it? When several candidates are still in the running, the question is comparative: which one delivers the most value and on what measures is that established? Both uses fit into one technique because they are built on the same material, the value attributes. Only the form of the assessment changes with the number of candidates.
The technique produces two deliverables, one per use. On the acceptance side, a list of verifiable statements attached to the requirements or user stories they qualify, each written to settle as a pass or a fail, then the record of the user acceptance test results. On the evaluation side, a grid crossing the weighted value attributes with the candidate solutions, where every column settles into a weighted total, then the ranking that comes out of it. Both are written before the thing they judge. A criterion written after the fact records what the solution already does: it can no longer make anything fail.
The technique earns its place by moving the discussion earlier in time. With no written criterion, acceptance is negotiated at delivery, when the budget is spent and the balance of power runs against whoever finds the problem. With a written criterion, it is negotiated at definition time, when the correction costs one sentence. The same shift holds for evaluation: the weights given to the attributes are a statement of what the organisation values, and the grid puts them down while they can still be argued.
Usage
When to use it
- Contractual obligation: sign-off and payment hang on a written condition both parties can check.
- User acceptance test to prepare: every requirement has to become a statement that settles as a pass or a fail.
- Several solutions to separate: software packages, design variants or competing bids compare on common measures.
- Agile backlog to prepare: an agile method may require every requirement to be expressed as testable criteria.
- Non-functional requirement to qualify: a threshold on response time, availability or volume makes it verifiable.
- Priorities to arbitrate across diverse needs: common measures rank demands that nothing else made comparable.
When not to use it
- Need still under exploration: no stable statement to make verifiable, frame the need first in a workshop.
- Choice dominated by uncertainty: the value of the candidates hangs on future events, go through decision analysis, which handles risk.
Description
The material common to both uses is the value attribute. BABOK defines it as a characteristic of a solution that determines or substantially influences its value for stakeholders. It presents it as a decomposition of the value proposition into its constituent parts, agreed on by the stakeholders. The guide gives six families of examples: the ability to provide specific information, the ability to perform or support specific operations, performance and responsiveness characteristics, the applicability of the solution in specific situations and contexts, the availability of specific features and capabilities, as well as usability, security, scalability and reliability. It puts the business analyst in charge of making sure the definition of every attribute is agreed by all stakeholders.
The number of solutions in play decides the form the assessment takes. A single candidate calls for a threshold and a binary verdict; several call for a scale and a ranking. The two tracks start from the same attributes and diverge at the next step.
Defining the value attributes
The attributes derive from the value proposition and from the business objectives. A list copied out of a product's feature catalogue reduces the assessment to what that product can do, which favours its vendor. A grid holds to a handful of attributes, four to eight in most cases: beyond that, each weight becomes too small to weigh and the grid adds up noise.
Every attribute gets a written definition saying how it is measured. The word "usability" settles nothing while the definition names neither the task, nor the user profile, nor the threshold. Where the solution is software, ISO/IEC 25010 gives a product quality model in which usability, security, scalability and reliability are broken down into measurable characteristics and sub-characteristics, which saves reinventing their meaning on every project. The correspondence with BABOK's wording is not term for term: the 2023 edition treats usability under the name interaction capability and places scalability among the sub-characteristics of flexibility. For a solution that is not software, these definitions are built case by case. The same attributes feed the non-functional requirements, where they already carry their thresholds. The glossary freezes the vocabulary so that two assessors read the same word the same way. For attributes that no instrument measures directly, BABOK points to expert judgement or to scoring techniques.
Writing a verifiable acceptance criterion
An acceptance criterion is a statement that can be established as true or false. BABOK states that acceptance criteria are expressed in a testable form, that this may require breaking the requirement down to an atomic form so that a test case can be written and that the verification often runs through user acceptance testing. ISO/IEC/IEEE 29148 places verifiability among the properties of a well-formed requirement, and an unverifiable criterion leaves the requirement as vague as it was.
A numeric threshold qualifies a non-functional requirement: the monthly payroll run completes in under five minutes. A sign-off clause qualifies a contractual obligation: a cap on implementation cost, a go-live deadline, a service level owed. A scenario written in "given, when, then" qualifies a user story, a form agile teams use because it transposes straight into a test case.
The commonest pitfall is the criterion that restates the requirement without adding anything to it. "The system handles payroll" becomes verifiable only once the input data, the action and the expected result are set down. A criterion with no threshold is settled at delivery, in favour of whoever speaks loudest: "the system is fast" holds nothing against a vendor who judges its product fast. A criterion written by the vendor measures what the vendor already knows how to do. A criterion frozen into a contract costs the most: BABOK holds acceptance criteria to be necessary where requirements express contractual obligations, and it notes that such criteria may on that account be difficult to change for legal or political reasons. A criterion that has become inadequate is handled by an amendment, under the change control procedure agreed in the contract: the observed gap is documented with its effect on price and on schedule, then both parties sign. A unilateral rewrite has no effect on what is owed.
Building an evaluation criterion
An evaluation criterion is a scale. BABOK describes it as a parameter measurable against a continuous or discrete scale, whose definition opens the measurement to various methods, among them benchmarking or expert judgement. The guide adds that defining evaluation criteria may involve designing the tools and instructions for the assessment itself, as well as the way its results are recorded and processed.
Three decisions make the grid, all of them taken before the results are seen. The choice of attributes says what the judgement is about. The weighting says what counts most: weights summing to 1.00 spread the importance out and record the stakeholders' value judgement. Where the weights are argued without converging, they are derived from pairwise comparisons with the analytic hierarchy process rather than negotiated as a block. The scoring scale, 1 to 5 for instance, states its direction: on cost the high score goes to the cheapest, and a grid silent on this point is read backwards.
The weighted score of a cell is the weight multiplied by the score; a candidate's total is the sum of its column. That total ranks the candidates without deciding in their place. A gap of a few hundredths between two totals, on scores of 1 to 5, carries no certainty: it is tested by moving the most debatable weight one notch to see whether the ranking holds.
The commonest bias is the weighting set once the scores are known, so as to recover the candidate that was already preferred. The remedy is the order of operations: the weights first, in session and in the open, the scores afterwards. The second bias is the knock-out threshold drowned in the average. A condition that is indispensable, data hosted in Switzerland or regulatory compliance, is set at preselection and removes the candidate that fails it; weighted among the others, it lets itself be offset by good scores elsewhere and stops being indispensable.
The evaluation grid shares its mechanics with the decision matrix, where weighted scoring is treated in its own right. Vendor assessment applies those mechanics to the choice of a supplier. What is proper to acceptance and evaluation criteria is where the dimensions of the grid come from: value attributes agreed by the stakeholders and derived from the value proposition.
Running the technique
- Break the value proposition down into a short set of attributes and write a measurable definition for each.
- Have those definitions agreed by all the stakeholders concerned, before drawing a single criterion from them.
- Count the solutions in play, which decides the track: threshold and verdict for one, scale and ranking for several.
- Acceptance track: translate each attribute into verifiable statements, broken down until a test case can be written.
- Evaluation track: weight the attributes in the open, freeze the weights, then score each candidate on evidence.
- Run the assessment: user acceptance testing on one side, scoring and aggregation on the other.
- Record the results and attach them to the requirements they qualify, so the trace survives delivery.
AI considerations
The first use that pays is putting a requirement into testable form. From a requirement written in prose, a language model breaks it into atomic statements, proposes the input data, the action and the expected result for each and flags those that stay unverifiable for want of a threshold. It also spots the criteria that contain two, whose statement carries an "and" that makes the verdict ambiguous when one half passes and the other fails. On a backlog, it produces a first pass of "given, when, then" scenarios that the team corrects faster than it would write them.
On the evaluation track, the gain runs through the calculation tool. A model writes the formula or the script that recomputes the weighted totals, sweeps a range of weights and flags the rank that flips at the first hundredth moved. Executing that script is what guarantees the result; the arithmetic a model does in its head stays unreliable and gets recomputed. A model also rereads the grid to find two attributes that measure the same thing or a scale whose direction contradicts its label.
What AI does not do follows from the nature of the technique. It does not set the weights: they carry the stakeholders' value judgement, and automating that removes the one thing the grid exists to make explicit. It does not pronounce acceptance: the verdict commits whoever issues it, and a pass produced by a model commits nobody. It does not know the implicit conditions of the business, the ones an AVS compensation fund, an auditor or a regulator will apply without having written them down and that an experienced user states in three minutes of interview. The test data of a payroll acceptance test or of customer files hold personal data: they do not go off into an uncontrolled third-party service, and the tool chosen has to process that data in line with the Federal Act on Data Protection.
Examples
A logistics SME in the canton of Vaud, thirty-four employees, is replacing its payroll software. The stakeholders have agreed four value attributes: cost, functionality, usability and performance. Three bids are left in the running after preselection.
Evaluation criteria
Payroll software, three bids ranked
| Value attribute | Weight | Bid A | Bid B | Bid C |
|---|---|---|---|---|
| CostTotal cost over 5 years · A CHF 42'000 · B CHF 68'000 · C CHF 95'000 | 0.30 | 5= 1.50 | 4= 1.20 | 2= 0.60 |
| FunctionalityCoverage of the company's payroll cases | 0.30 | 2= 0.60 | 4= 1.20 | 5= 1.50 |
| UsabilityEntry of a new hire by an HR administrator | 0.20 | 4= 0.80 | 4= 0.80 | 3= 0.60 |
| PerformanceDuration of the monthly payroll run | 0.20 | 4= 0.80 | 3= 0.60 | 4= 0.80 |
| Weighted total | 3.702nd | 3.801st | 3.503rd | |
The margins are thin. A is only twenty hundredths ahead of C, and its price advantage there is exactly cancelled by its functionality shortfall: usability is what puts it ahead of C. Moving five hundredths from cost to functionality leaves B at 3.80, lifts C to 3.65 and drops A to 3.55. The lead holds, the rest of the ranking inverts. A margin that one notch of weighting overturns measures the weighting before it measures the bids, and the question put to the decision maker becomes the weight they give to price. With bid B selected, the same four attributes are rewritten as acceptance conditions for the user acceptance test.
Acceptance criteria
User acceptance test of the selected bid
| Value attribute | Acceptance criterion | Result |
|---|---|---|
| Functionality | A gross monthly salary of CHF 6'500 produces the AVS/AI/APG, AC, LPP and non-occupational accident insurance (AANP) deductions borne by the employee at the configured rates, and the net figure on the payslip equals the gross less those deductions. | Pass |
| Functionality | The salary declarations of the thirty-four employees are transmitted through the unified salary reporting procedure (ELM) and the acknowledgement of receipt from the AVS compensation fund is recorded. | Pass |
| Performance | The monthly payroll run for the thirty-four employees completes in under five minutes. | Pass |
| Usability | An HR administrator trained for half a day enters a complete new hire without assistance and without opening the manual. | Fail |
| Cost | The implementation cost invoiced stays at or below the cap of CHF 18'000 set in the contract. | Pass |
The same attribute changes nature between the two tables. In the grid, a low score is offset: B wins despite a performance score of 3. In the acceptance test, the threshold replaces the scale and no average softens a shortfall. That crossing is the moment when the team has to say what is indispensable, a rougher conversation than the weighting.
Visualisations
The fork between the two tracks is a position in space: two parallel paths that start from the same material and end in verdicts of a different nature. The diagram carries four steps per track, the value attributes at the head of each, the binary verdict on one side and the value ranking on the other.
The weighted grid and the list of verdicts are made of rows and columns: they are carried as HTML in the body, where they stay legible at zoom and readable by a screen reader. The two tables carry the same four attributes and are read against each other: the same word becomes a score in the grid and a threshold in the acceptance test. In the acceptance table, the result column takes only two values: an in-between cell, "partially compliant", would reopen the negotiation the criterion existed to close.
Cost
| Phase | Level | Rationale |
|---|---|---|
| Preparation | High | Breaking the value proposition down into attributes, writing a measurable definition for each and obtaining the agreement of all the stakeholders. The weighting needs a session of its own, and BABOK notes that this agreement among stakeholders with diverse needs can be hard to reach. |
| Execution | Medium | Writing the verifiable statements or scoring the candidates goes fast once the attributes are agreed. The cost is the cost of the evidence: an acceptance test campaign, demonstrations, performance measurements, expert judgement on the attributes that resist direct measurement. |
| Documentation | Medium | The criteria live attached to the requirements they qualify and follow their changes. Under contract they are enforceable, and their history is kept together with the acceptance test results. |
Tools
A spreadsheet is enough for the evaluation grid and keeps it verifiable and repeatable, which counts when an award has to be justified. The weighted contributions and the totals are formulas, and the sensitivity check is done by changing one weight cell to watch the ranking recompute. Conditional formatting marks out the leading candidate. Under a public tender, procurement platforms such as simap.ch carry the same grid with the timestamping and the audit trail the procedure demands.
On the acceptance side, the requirements management tool is the right place: a criterion recorded in the minutes of a meeting gets lost, whereas attached to the requirement in the tool it follows its changes, turns up at the acceptance test and links to the test cases that verify it. Jira with Xray or Zephyr, Azure DevOps, Polarion and Jama hold that traceability, from the criterion to the test case and then to the execution result. Test management tools carry the acceptance campaign itself and produce the record of passes and fails that the client or the steering committee signs off.
Where the criteria are written in "given, when, then", behaviour-driven development tools, Cucumber, SpecFlow or Behave, execute them as they stand: the criterion becomes the test, and the question of whether the test matches the criterion disappears. The anti-tool is the scoring grid supplied by the candidate, which weights the attributes its own bid is strong on and presents a commercial preference as an assessment.
Sources
- IIBA, A Guide to the Business Analysis Body of Knowledge (BABOK Guide) v3, §10.1 Acceptance and Evaluation Criteria: the purpose of the two uses, the definition of value attributes and their six families of examples, the testable form of acceptance criteria, the continuous or discrete scale of evaluation criteria, together with the recognised strengths and limitations, among them contractual rigidity and the difficulty of reaching agreement among stakeholders with diverse needs.
- ISO/IEC/IEEE, ISO/IEC/IEEE 29148:2018, requirements engineering: verifiability among the properties of a well-formed requirement, which grounds the testability demanded of an acceptance criterion.
- ISO/IEC, ISO/IEC 25010:2023, product quality model: the operational definitions of the qualities BABOK cites as examples, for software solutions, with interaction capability in place of usability and scalability filed under flexibility.

