Your Training Partner
Techniques Toolbox
Kraljic quadrant crossing supply risk and impact on results, showing when a full vendor assessment is warranted.

Vendor Assessment

Vendor assessment measures a supplier's capacity to meet its delivery and continuity commitments, before the organisation selects or contracts with it. As soon as part of a solution is bought, built or operated by a third party, that dependency introduces a risk of capacity, continuity and data handling that does not exist when the work is done in-house. The technique structures that judgement: it defines the criteria a supplier is judged on, assigns each a weight, scores every candidate on evidence and aggregates the whole into a weighted ranking that supports the sourcing decision. Its calculation engine is a weighted decision matrix, specialised to the problem of choosing a supplier, with the criteria families BABOK describes as its backbone and the pieces of a tender as evidence.

Purpose

Vendor assessment answers a precise question, which of the candidates best serves the need once its requirements are stated and weighted, and it answers it before the commitment, while the choice is still open and reversible. It comes into play when a solution is designed, built, deployed, maintained or outsourced, and where the make-vs-buy leans towards buying. Handing part of the solution to a third party moves a risk outside the organisation without removing it: the supplier must be able to deliver, to sustain a service level over time, to stay solvent, to keep qualified staff and to protect the data entrusted to it.

The central deliverable is a weighted scoring grid that crosses the evaluation criteria with the candidate vendors, and where each column resolves into a weighted total. Around it, four pieces make the decision defensible: the written definition of the criteria and their weights, which says what is judged and what matters most; the shortlisting record, which keeps track of the candidates dropped and why; the log of evidence and due-diligence findings by criterion; and the recommendation, together with its sensitivity analysis. Under a formal tender, this set is the audit trail that justifies the award.

The value of the technique lies in making explicit a judgement that is made anyway. The criteria retained, the relative importance given to them and the score placed against each vendor are three decisions, and the grid poses them while they are still open to debate. It raises the probability of a lasting relationship with a reliable, fitting supplier. It does not remove the risk of failure as the partnership evolves, and its share of subjectivity can bias the result if the criteria are scored by gut feel.

Usage

When to use it

  • Component handed to a third party: packaged software, SaaS or outsourced operation, the supplier's capability conditions the solution.
  • Heavy or hard-to-reverse decision: high spend, long contract, deep integration or a component critical to the solution.
  • Several candidates to rank: compare offers on one shared grid or make a sole supplier clear a capability threshold.
  • Confidential data handed to the supplier: it will host personal data, continuity and compliance become weighted criteria.
  • Formal tender: RFI, RFP or public procurement, the award must rest on an objective, traceable basis.

When not to use it

  • Low-stakes commodity purchase: the assessment costs more than the decision, buy on price and availability and reserve the effort for strategic or bottleneck purchases.
  • Imposed or sole supplier: regulation, framework agreement or the only viable supplier, there is nothing to compare, structure the risk instead with decision analysis.
  • Choice between internal options: the make-vs-buy is settled on "make", score the variants with the decision matrix, there is no supplier to assess.

Description

The depth of the assessment is set by the stakes of the purchase. An interchangeable consumable does not deserve the same diligence as a packaged product that will host personal data for five years. Two dimensions locate the purchase: supply risk, that is the difficulty of finding and replacing the supplier, and impact on results, that is the spend and the criticality of the component. A high-impact, high-risk purchase justifies the full assessment; a low-impact, low-risk purchase is handled on price. Positioning the purchase on these two axes is the first decision, the one that says whether the technique is worth its cost.

Impact on resultsSupply risklowhighlowhighLeverageStrategicNon-criticalBottleneckBuy on priceFull assessment warranted
Position the purchase before assessing it: only strategic and bottleneck buys justify the full weighted assessment; a non-critical buy is decided on price.

Frame the decision and shortlist

First one states what is being bought, the make-vs-buy premise and the list of candidates. A long list is narrowed to a short one by a binary shortlist on the non-negotiable criteria: certification held, hosting jurisdiction, financial-strength floor. Shortlisting removes the out-of-scope candidates before scoring, which is the expensive part. Skipping it means scoring in detail vendors that already failed an eliminating requirement.

Define and weight the criteria

The criteria derive from five families BABOK describes and from the non-functional requirements specific to the solution, in particular the expected service levels. Knowledge and expertise cover the capability the company lacks in-house. Licensing and pricing carry the cost, to be analysed on usage scenarios rather than the list price. Market position measures longevity and the fit between the organisation's profile and the supplier's customer base. Contractual terms cover intellectual property, liability for the data entrusted, the update schedule and the risk of lock-in on exit. Experience, reputation and stability cover references, compliance with external quality or security standards and guarantees against financial difficulty at the supplier. Five to eight weighted criteria are retained: beyond that, the weighting dilutes and each weight stops carrying weight. Each criterion gets a measurable descriptor, without which two evaluators will score the same offer differently and the grid will aggregate noise.

Weighting distributes importance, for instance weights that sum to 1.00. This is the strategic conversation of the exercise, the one that encodes what the organisation truly values. It is held before the scores are seen, in session and in the open, then frozen. Setting the weights afterwards to recover the supplier one already preferred is the most common bias of the technique, and the only remedy is the order of operations: weights first, scores second.

Gather the evidence and score

Each criterion attaches to a source of evidence. RFI or RFP responses, customer references, due-diligence findings, financial statements, certification evidence, a demonstration or a proof of concept, a site visit feed the scoring. Each vendor is scored on each criterion on a fixed scale, for instance 1 to 5, ideally by more than one evaluator to blunt individual bias, with the gaps then reconciled. On the licensing family, one scores the total cost of ownership, licence plus implementation plus integration plus operation plus exit: a licensing model that is cheap on paper often tips once the real usage scenario is applied. For a Swiss buyer entrusting personal data, the hosting jurisdiction and compliance with the revised Federal Act on Data Protection (revFADP) are scored and weighted criteria, often heavily.

Aggregate, read and decide

The weighted score of a cell is the weight times the score; a vendor's weighted total is the sum of its column. That total ranks the candidates, but one reads the grid, one does not obey it. A vendor can come top on the total and fail a requirement that had been set as indispensable: shortlisting should have removed it, and if it did not, the total does not buy back the breach. A gap of a few hundredths between two totals is not decisive when it rests on subjective scores; one tests it with a sensitivity analysis, shifting the most debatable weights by one notch to see whether the ranking holds. If it flips, the decision is close and will be settled on a criterion outside the grid. Finally one documents the recommendation and its reasoning, because the assessment does not immunise against a later failure: it is redone as the relationship evolves, a reassessment that ISO 9001 in fact writes into the control of external providers.

The calculation engine is a weighted decision matrix

The scoring grid of a vendor assessment is a decision matrix whose options are suppliers and whose weighted criteria are the evaluation dimensions. The decision matrix holds the general method, criteria, weights, scores and weighted sums; vendor assessment is that method specialised to the sourcing problem, with the five families as the criteria backbone and the evidence of a tender as inputs. The whole sits within decision analysis, the approach that frames a decision, defines the alternatives, evaluates them and arbitrates under uncertainty. When the supplier is imposed and there is no matrix to fill, that approach still structures the risk and the negotiating posture around the fixed choice.

Working through the technique

  1. Frame the purchase and position its stakes, to decide whether the full assessment is warranted or the purchase is handled on price.
  2. Shortlist on the eliminating criteria, certification, jurisdiction, financial floor, so that only live candidates are scored.
  3. Define five to eight weighted criteria, derived from the five families and the service levels, each with a measurable descriptor.
  4. Weight before seeing the scores, in session and in the open, then freeze the weights.
  5. Gather the evidence, RFP responses, references, due diligence, financial statements, certifications, demonstration.
  6. Score each vendor on each criterion, with several evaluators, and reconcile the gaps.
  7. Aggregate, run the sensitivity and document the recommendation with its audit trail.

AI considerations

The first useful use is preparing the criteria and processing the offers. A language model given a specification and the non-functional requirements proposes a list of candidate criteria and spots overlaps, the two criteria that at bottom measure the same thing and, left in place, would count one concern twice. Faced with bulky RFP responses, it extracts per criterion what each vendor states and lines up the claims in a first pass of the grid, which the evaluators then only have to challenge. It also speeds up documentary due diligence: summarising a financial report, aggregating public references, flagging the certifications to verify.

The arithmetic and the sensitivity are the most rewarding ground. Recomputing the weighted totals, sweeping a range of weights, flagging where the ranking is fragile, holding the coherence of the scores when one score changes, all of this is calculation the machine does without error where a room miscounts a twenty-line grid.

What AI does not do goes to what makes the technique exist. It does not set the weights: they are the stakeholders' value judgement, and automating it removes the one thing the grid exists to make explicit. It does not verify that a vendor's claim is true: a certification stated in an RFP response is checked at its source. It does not judge the trust or the fit of a long-term relationship, which are read in references and direct contact. Finally, a vendor scoring grid contains confidential commercial data, sometimes protected by contract, and sometimes the buyer's own personal data: it is not dropped into an uncontrolled public service, but into a tool whose data handling complies with the revFADP.

Examples

A cantonal administration in French-speaking Switzerland is selecting a SaaS electronic document-management solution that will host personal data, among three shortlisted offers. Six criteria grouped by family, weights summing to 1.00, scores from 1 to 5.

Weighted scoring grid

SaaS document-management solution

Normalised weights, sum = 1.00. Scores from 1 to 5 (5 = best; cost is inverted, 5 = cheapest). Below each score is its weighted contribution (weight × score). The weighted total, at the foot of each column, adds these contributions.
CriterionWeightVendor AVendor BVendor C
Functionality and fitKnowledge and expertise0.254= 1.005= 1.253= 0.75
revFADP compliance and Swiss hostingContractual terms0.205= 1.003= 0.604= 0.80
Total cost of ownership over 5 yearsLicensing and pricing · A CHF 480'000 · B CHF 360'000 · C CHF 300'0000.203= 0.604= 0.805= 1.00
Financial strength and referencesExperience, reputation, stability0.154= 0.605= 0.753= 0.45
Reversibility and data portabilityContractual terms0.124= 0.482= 0.243= 0.36
Market position and longevityMarket position0.083= 0.245= 0.404= 0.32
Weighted total3.922nd4.041st3.683rd
Vendor B leads on the weighted total (4.04), carried by functionality and reputation. A follows closely (3.92), stronger on compliance and reversibility. C is the cheapest but weakest on data protection.

The concept the grid makes visible is that a winning total does not excuse reading the column. B dominates the weighted total, yet it earns the lowest score on reversibility (2): had data portability been set as an eliminating threshold at shortlisting, B would have dropped out despite its first rank. The gap between B and A is only 0.12, a margin that scores from 1 to 5 do not carry with certainty. The sensitivity analysis confirms it: since the solution hosts personal data, it is legitimate to load compliance. Shifting eight hundredths of weight from functionality to revFADP compliance, A rises to 4.00 and B falls to 3.88, and the ranking inverts. The grid has therefore not named a clear winner, it has shown that the choice turns on the importance given to data protection, and it hands that question back to the person who decides.

Visualisations

The scoring grid is made of rows and columns: the deliverable is the table itself. Carrying it in HTML rather than as an image has practical consequences: the totals recompute when a score changes, the grid stays legible on zoom and for a screen reader, and the leading vendor stands out to the eye in the foot of the column. Two conventions make it legible. The weights appear in the clear next to each criterion, since they carry the value judgement and a grid that hides them cannot be checked. The weighted contribution shows below each score, so that the reader sees where the total comes from without recomputing it.

The prior decision, whether the full assessment is worth its cost, is drawn better than it is tabulated, because it is a position in a two-dimensional space rather than a value in a cell. A quadrant crossing supply risk with impact on results locates the purchase and indicates when the heavy assessment is warranted: strategic and bottleneck purchases deserve it, non-critical purchases are handled on price.

Cost

PhaseLevelRationale
PreparationHighThe real work is here: building the criteria from the five families and the service levels, shortlisting the long list and above all eliciting the weights from the decision-makers. Drafting the RFI or the RFP adds to it.
ExecutionMedium to highGathering the evidence, RFP responses, references, due diligence, demonstrations, then scoring with several people and reconciling the gaps. This is the step BABOK notes as one that may consume time and resources.
DocumentationMediumThe grid, the definition of the criteria and weights, the shortlisting record and the recommendation form the audit trail, required under a formal tender and to be kept up to date at reassessments.

Tools

The spreadsheet is the honest and sufficient choice for most cases. The weighted totals are formulas, the sensitivity is done by changing one weight cell to watch the ranking recompute before your eyes, and conditional formatting colours the leading vendor. It keeps the grid checkable and replayable, which is the essential thing when the award has to be justified.

E-sourcing and tender-management platforms, from e-procurement suites to public-procurement portals such as simap.ch for a Swiss public buyer, take over when the procedure is formal, the candidates are numerous or the audit trail must be enforceable. They circulate the RFP, collect the responses in a comparable form and timestamp each exchange. Supplier-intelligence services, financial data, certification registers, security ratings, feed the stability and security criteria with evidence independent of the candidate's word.

For the criteria-and-weights workshop, a whiteboard, physical or shared, remains the best support: the criteria are written up, the points distributed in full view and everything frozen before moving to the scores. The anti-tool is the scoring sheet supplied by the vendor itself, which weights the criteria on which its offer is strong and presents a commercial preference as an assessment.

Sources

  • IIBA, A Guide to the Business Analysis Body of Knowledge (BABOK Guide) v3, §10.49 Vendor Assessment: the purpose of the technique, the five families of evaluation criteria, the formality levels of the tender from informal to RFP and the recognised strengths and limitations, including its resource-consuming nature and the risk of bias through subjectivity.
  • ISO, ISO 9001:2015, clause 8.4 Control of externally provided processes, products and services: the selection, evaluation and re-evaluation of external providers against defined criteria.
  • Peter Kraljic, Purchasing Must Become Supply Management, Harvard Business Review, 1983: the purchasing matrix crossing supply risk and impact on results, which locates when a full assessment is warranted.
  • PMI, PMBOK Guide, source selection criteria and bid weighting system: which anchor the scoring engine in established procurement practice.
Value Stream Mapping (VSM)
All techniques
Visioning