
Somewhere in the middle of a selection, someone finally builds the spreadsheet. That spreadsheet is the vendor evaluation scorecard: criteria down the left, vendors across the top, weights invented in a meeting where one person had already made up their mind. The numbers come out the way everyone expected, and nobody says the obvious thing: the scorecard didn't decide anything, it ratified something.
Key takeaways
Ratifying a decision already made is the normal case, not the failure case. Forrester's 2024 Buyers' Journey Survey, as reported by Digital Commerce 360, found 92% of buyers start with at least one vendor in mind, and 41% already have a single preferred vendor before formal evaluation begins. Forrester's conclusion was that "B2B buying today is a process of confirmation, not selection" [A16]. 6sense's 2025 report asked buyers whether they could rank their shortlist before engaging any seller: "a resounding 94% of buyers answered yes" [A17].
A weighted vendor evaluation scorecard is a fixed list of evaluation criteria, each carrying a percentage weight, scored on one written scale and summed to a single total out of 100 [A2]. It is valid rather than decorative when three things hold: the criteria and weights are written down before the bids arrive [A6], each bid is scored against your stated requirement rather than against the other bids [A8], and anything that could end the engagement is a pass/fail gate instead of a weighted line [A4].
So a scorecard's job isn't precision. It's pre-commitment: fixing what you care about, and how much, before you can see which vendor it favours. Which is why how to evaluate software development vendors is a sequencing problem before it's a scoring one. What follows is that vendor evaluation template, rebuilt from the public method, with the derivation of every weight shown, traced to a named public source where one exists, and labelled as our own judgement with the reason given where none does, and a last section where we run it on ourselves and land a 4 on seven of the eight lines.
Search for a software vendor evaluation scorecard and two serious weighted templates come back. Both were published by companies selling into the market they score.
Pangea.ai publishes eight dimensions — Technical Vetting Quality 20%, Talent Pool Depth 15%, Speed to Match 15%, Replacement Guarantees 15%, Compliance and Legal Infrastructure 10%, Contract Flexibility 10%, Communication and Account Management 10%, Pricing Transparency 5% — scored 1 to 5 and multiplied by weight, published 7 May 2026. It does state a derivation, the sheet is "built from analysis of vendor selection processes across hundreds of technology companies", but that analysis isn't published, so there is nothing to check the percentages against. And Pangea.ai is a staff augmentation marketplace competing in the market its scorecard evaluates [A34].
ZTABS publishes six categories — Technical Capability 25%, Portfolio and Experience 20%, Communication and Process 20%, Pricing and Value 15%, Cultural Fit 10%, References and Reputation 10% — published 30 March 2026. Its only note on provenance: "you can adjust the weights to match your priorities". ZTABS is itself a software development vendor [A35].
Neither is dishonest, and both give you usable structure. But look at Pangea's shape: Speed to Match and Replacement Guarantees at 15% each, Pricing Transparency at 5% [A34]. That weights what a marketplace is structurally good at and discounts what marketplaces are worst at. The problem isn't the numbers. It's that neither derivation is checkable: one points at unpublished internal analysis, the other at your own priorities. We went looking for a published, non-vendor weighted scorecard for software development vendors to compare them against and didn't find one, which is a statement about our search, not a proof that none exists.
Second disclosure: this one is published by a vendor too, on you-source.com. The difference on offer is procedure, not neutrality. Every weight is either traced to a named public source or stated as our own judgement with the reason given, one of the eight is the latter, and it's flagged where it appears. Several criteria are ones You-Source falls short of the top score on, and the last section runs the sheet against You-Source in public. Who those competitors are is a separate exercise.
Public procurement has spent decades scoring bids defensibly, because there a losing bidder can sue you over the method. Five documents carry most of the weight.
The World Bank's Evaluating Bids and Proposals with Rated Criteria, third edition, February 2025 is the closest thing to a manual. Weights come from a pairwise matrix: compare every criterion against every other, count the wins, rank, agree percentages. They "should add up to 100% in total" [A2]. Criteria "should be kept to the essential minimum" because too many "serves to dilute the important characteristics of Bid/Proposals" [A1]. A dominant criterion may take "perhaps 50% of the total technical weighting", with a worked 40/30/20/10 split [A3]. And a minimum quality threshold can sit on a single criterion that "carries relatively low weighting, but is important or critical to the Procurement outcomes" [A4]. That's where the gates come from.
The UKUPC Guide to Evaluation, November 2020 supplies the 0–5 scale with written descriptors [A7], the rule to compare a bid "against requirements, not against other tenders" [A8], and moderation to "a consensus result" [A9].
GOV.UK's Digital Outcomes and Specialists buyer guidance adds three rules worth copying: "You can't change the criteria or their weightings after your requirements are published"; one scoring scheme across all questions; and "document the reasons for awarding scores to each supplier and the reasons for choosing the criteria and weightings" [A6]. Its companion page gave a worked weighting: "30% on price and 60% on technical expertise, 10% on cultural fit" [A5]. Published 1 March 2021, withdrawn 31 December 2025 as outdated: the clearest non-vendor worked example for digital services, and no longer policy.
Cabinet Office Procurement Policy Note 04/15, 25 March 2015 required buyers to satisfy themselves that a supplier's "principal relevant contracts in the last three years" had been "satisfactorily performed", via written performance certificates stating reasons for any unsatisfactory rating: delays, incomplete scope, missed service levels, other contractual failures [A26, A27]. It is a 2015 policy note and not something you can hold a private vendor to, borrow the method, not the legal status.
Kiiver and Kodym, Journal of Public Procurement 15(3), 2015 — two European Parliament procurement officers on how scoring formulas break: "any formula that makes the score of one tender dependent on the content of another tender includes an element of arbitrariness in the evaluation and, at worst, exposes the process to deliberate manipulation by colluding tenderers" [A10].
One caveat before you spend a week on this. There is no measured evidence that structured vendor evaluation produces better outcomes than unstructured judgement. No such study exists. The nearest support is from another field: a meta-analysis of 136 studies found mechanical prediction "about 10% more accurate than clinical predictions", holding "regardless of the judgment task, type of judges, judges' amounts of experience, or the types of data being combined" [A14]. That's clinical prediction, not procurement. An analogy, and it should be read as one. Dawes' 1979 work on improper linear models points the same way, as context rather than proof [A15].
Weighted totals have a failure mode: a vendor can be catastrophic on one line and still win on points. The World Bank's fix is a minimum quality threshold on a single criterion [A4]. Three belong there. Each one is binary, cheap to check, and capable of ending the engagement badly however good the code was.
One thing to be clear about before you use them: all three gates are questions you ask a vendor during the selection, and none of them can be scored from a website. A contract clause, a client's permission to be phoned, and a written exit term are things a vendor tells you in writing when you ask. A marketing page can indicate a position; it cannot clear a gate, not the vendors you're evaluating, and not the one publishing this.
Gate 1 — a written IP assignment of all deliverables. The UK Intellectual Property Office is blunt: for commissioned work "the first legal owner of copyright is the person or organisation that created the work and not you the commissioner, unless you otherwise agree it in writing" [A29]. No written assignment, and the vendor owns the code you paid for.
Gate 2 — two named, contactable clients from the last three years, with permission to call them. PPN 04/15's standard, scaled down [A26, A27]. A logo wall isn't a reference. Neither is a marketplace listing: Crown Commercial Service states of the Digital Marketplace that "Crown Commercial Service (CCS) don't evaluate suppliers or their individual services" [A28].
Gate 3 — a stated exit: notice period, handover scope, no charge to return code and data. Under the EU Data Act's switching rules the parties "must contractually agree a notice period that does not exceed two months", with a transitional period of "a maximum of 30 calendar days" [A31], and from 12 January 2027 in-scope providers "won't be able to charge their customers for the operations that are necessary to facilitate switching or for data egress" [A30]. Read the scope: those rules bind data processing services (cloud, SaaS, PaaS, IaaS), not development agencies [A30, A31]. Use it as a benchmark, never as an obligation you can claim a dev vendor is already under.
A vendor failing any gate isn't scored low. It isn't scored.
Eight already pushes the World Bank's "essential minimum" advice [A1], and the weights sum to 100 as required [A2]. Each weight is a declared preference with a stated reason. Change them if your reason is better. But write it down, before the bids arrive [A6].
| # | Criterion | Weight | Why this weight |
|---|---|---|---|
| 1 | Verifiable delivery evidence | 20% | The only criterion with government policy behind it: named contracts in the last three years, written performance certificates, defined failure categories [A26, A27]. Also the only line a third party can falsify. World Bank guidance illustrates a vastly-more-important criterion taking "perhaps 50% of the total technical weighting" [A3] — an example, not a ceiling, and half of a technical score that excludes cost, so it isn't directly comparable to 20% of a sheet that includes commercial terms. It's the shape of the advice we're following, not the number. |
| 2 | People you actually get | 15% | US federal contracting treats key personnel as a controlled contract term: 30 days' notice, a named replacement of equal or better credentials, and no substitution "without the written consent of the Contracting Officer" [A32]. Scoring this closes the bait-and-switch gap a "team quality" question leaves open. |
| 3 | Technical practice | 15% | The 2025 DORA report, covering nearly 5,000 technology professionals, found 90% use AI at work, 30% report little or no trust in AI-generated code, and a negative relationship between AI adoption and delivery stability alongside a positive one with throughput [A33]. So score it from artefacts, not answers. |
| 4 | Commercial terms | 15% | UK guidance at paragraph 7.2.1 states "relative price scoring should be treated with caution and not be used unless there is a specific business reason which has been approved by the commercial lead and the project SRO" [A12], and scoring one bid against another invites manipulation [A10, A11]. The 2021 UK example put price at 30% [A5]; 15% is a stated choice, because a vendor who won't publish a number can't be scored against a budget. The objection, from the same article we cite here: Kiiver and Kodym also warn that in an additive price-plus-quality model "the weight of the price criterion must never drop below 50%", because below that the price criterion "can get cancelled out entirely" and the sheet "imposes an implicit minimum quality threshold, restricting competition at the low end of the market" [A10, A11]. We depart from it knowingly. Their model is a commodity bid where the requirement is fixed and price is the live variable; this sheet buys an ongoing supplier relationship where the variable is who does the work, and it accepts the consequence they name — the cheapest vendor cannot win this sheet on price alone. If your requirement really is a fixed, commoditised scope, follow them and not us. |
| 5 | Exit terms | 10% | A concrete external benchmark exists [A30, A31], which is the reason this line is on the sheet at all. WorldCC and Deloitte report contract value erosion of 8.6% across 1,236 organizations, best performers a little over 3%, worst more than 20% [A25] — that measures the whole contracting lifecycle rather than exit terms specifically, so read it as a reason to take post-signature terms seriously, not as a derivation of this number. |
| 6 | Security and data-access model | 10% | The one weight with no public source behind it: 10% is our own judgement, and here's the reason. It's gate-and-verify territory — an attestation either exists and is checkable, or it doesn't — so the top of the line rewards a binary you can confirm rather than a description you can't. Certifications, sub-processors and Article 28 terms belong in the security questionnaire you send every vendor. If security is the thing that would end your engagement, raise it, or make it a gate. |
| 7 | Communication cadence and reporting | 10% | The one line both vendor scorecards agree on: Pangea.ai weights it 10%, ZTABS 20% [A34, A35]. Agreement between competitors is weak evidence, but it is evidence, and 10% sits at the bottom of their range. |
| 8 | Working-model fit | 5% | The 2021 UK guidance made cultural fit a required dimension at 10% [A5]. Halved because it's the least falsifiable line on the sheet and the most exposed to the halo effect moderation exists to catch [A9]. |
To derive your own, use the World Bank pairwise method [A2]. And if you're still deciding between an agency, a subscription and a permanent hire, settle the staffing model first. The weights follow from it.
Use the UKUPC anchors, written on the sheet before anyone scores: 0 cannot be scored, 1 unsatisfactory, 2 satisfactory, 3 good, 4 very good, 5 excellent [A7]. Add a sentence per criterion saying what a 3 and a 5 look like for your requirement.
Score independently, then moderate to a consensus, with challenges encouraged [A9]. Write down the reason for each final score, because UK guidance requires documenting both the scores and the choice of criteria and weightings [A6].
Then the rule that matters most: score each bid against your requirement, never against the other bids. UKUPC states it twice, including in moderation, where "each individual tender must be considered independently of the others and the scoring process must not be influenced by comparable scores for other tenders" [A8].
Which is why this template scores price absolutely, against your stated budget, rather than giving the cheapest bid five points and scaling everyone else down. UK guidance discourages that without senior approval [A12], and Kiiver and Kodym show what non-linear point curves do: "if all tenders get nearly the same price score in the end, the advantage of low-range prices is almost cancelled out" [A11]. Set a budget, anchored on what the work actually costs. Describe a 3 and a 5 against it. Score each vendor alone.
A method you can't criticise is a method you shouldn't trust. Seven ways this one fails.
You over-tune the weights. The flat-maximum effect holds that "the predictive ability of linear models is insensitive to large variations in the size of regression weights" and to the number of predictors; tested against ten credit unions' models from 1984 to 1988, a generic weighted-average model performed "very close to that of the empirically derived models" [A13]. Arguing 15% versus 18% is wasted time. Choosing the right eight criteria isn't.
You score bids against each other. Any formula making one bid's score depend on another's content "includes an element of arbitrariness in the evaluation and, at worst, exposes the process to deliberate manipulation by colluding tenderers" [A10].
Your scores compress. If every vendor lands between 3.2 and 3.6, the sheet has stopped discriminating [A11]. Widen the descriptors, or accept you have three acceptable vendors.
You add criteria. Every extra line dilutes the important characteristics [A1].
You change criteria after publishing them. UK guidance forbids it outright [A6], and with 41% of buyers walking in with a favourite [A16], the temptation isn't hypothetical.
One person scores everything. No independent scoring, no moderation, no consensus [A9], and one impressive demo drags every unrelated line up with it.
A fatal low-weight line passes on total. That's what gates are for [A4]. If something would end the engagement, it doesn't belong in a weighted average at any weight.
And the caveat again, beside the criticisms, not buried at the end: the evidence that mechanical combination beats unstructured judgement comes from clinical prediction [A14]. Nobody has measured whether it picks better software vendors.
Lift this.
Gates — pass or fail, before any scoring [A4]. Ask each one of the vendor and get the answer in writing; none of them is scoreable from a website.
| Gate | Test | Fail condition |
|---|---|---|
| IP assignment | Written assignment of all deliverables to you, in the contract, no carve-outs [A29] | Absent, verbal, or "on final payment" with final undefined |
| References | Two named, contactable clients from the last 3 years, with permission to ask about performance [A26, A27] | Logos only, NDA'd out, or nothing inside three years |
| Exit | Stated notice, handover scope, no charge to return code and data. Benchmark: ≤2 months' notice, ≤30 days' transition, no egress fee from 12 Jan 2027 [A30, A31] | No stated notice, or a charge to get your repository back |
Criteria and weights [A2]
| # | Criterion | Weight | What a 5 looks like |
|---|---|---|---|
| 1 | Verifiable delivery evidence | 20% | 5: published case studies with production numbers a reader can check without asking — plus three-plus comparable contracts in 3 years, references taken, written detail on what went wrong. 4: named, published references, plus detailed delivery evidence available on request |
| 2 | People you actually get | 15% | Named individuals with CVs, substitution only by your written consent, documented cover for the lead [A32] |
| 3 | Technical practice | 15% | 5: a published engineering standard, or a technical case study with production numbers — plus the artefacts on request: a real PR, a test suite, a CI config, a written policy on reviewing AI-generated code [A33]. 4: credible published practice statements and a demonstrable review offering, without the standard or the case study |
| 4 | Commercial terms | 15% | A complete price you can act on, a defined change-order mechanism, no unpriced assumptions |
| 5 | Exit terms | 10% | The gate 3 benchmark met or beaten, named transition assistance |
| 6 | Security and data-access model | 10% | 5: a current third-party attestation (ISO 27001, SOC 2 Type II or equivalent), plus work inside your environment under your controls and sub-processors listed. 4: a published operating model that controls access and data location, without third-party attestation |
| 7 | Communication cadence | 10% | Stated frequency, named channel, named escalation path, demonstrated response time |
| 8 | Working-model fit | 5% | The overlap you need, decision latency in hours, evidence they'll push back |
Scale [A7]: 0 cannot be scored · 1 unsatisfactory · 2 satisfactory · 3 good · 4 very good · 5 excellent.
Totalling: weighted points = weight × (score ÷ 5); sum the eight for a score out of 100. Score independently, moderate to consensus, document the reason [A6, A9].
A trial, pilot or paid proof-of-concept scores under commercial terms for every vendor — a competitor offering a free pilot earns the same credit here that our Proof of Quality does.
Same sheet, run on its publisher. Everything below is what You-Source publishes, so you can check it.
Gates. No pass marks here, because the gates are answered in a selection process and this section only has a website to work from. What You-Source publishes against each, and what it doesn't:
Three published positions, one cleared gate element, and the rest to be asked for. That is where every vendor on your shortlist starts, including this one.
| Criterion | Weight | Score | Basis |
|---|---|---|---|
| Verifiable delivery evidence | 20% | 4/5 | Against the 4 anchor — named, published references, plus detailed delivery evidence available on request. Four named client testimonials, each with a named company: Kevin Conlon (Kinnect), Derek Sturdy (Robur), Michael Hanna (Solargain), Eric Peterson (Stack Moxie) [A40], with case studies available on request. The caveat stands: "60+ Projects Delivered", "5 Continents served" and a "98.4% Renewal rate" are all self-reported [A41]. Not a 5: nothing published here is a case study with production numbers a reader can check without asking |
| People you actually get | 15% | 4/5 | Named dedicated engineer, "Same-day match" [A38], "5-day replacement guarantee", swapped at no cost [A37]. Bench depth is thin at 1–2 engineers per subscription [A37] |
| Technical practice | 15% | 4/5 | Against the 4 anchor — credible published practice statements plus a demonstrable review offering. The statements: "Production-grade work — reviewed, tested, deployed. Built to ship, not to demo." [A42] and a task-by-task approval gate [A37]. The offering: a Code Audit that maps "what's fragile, what won't scale, and what's blocking production" and hands over "a prioritized fix plan", sold as "3 days • Senior review • Fixed price" [A43], specified elsewhere as "Three days, two senior engineers, one production-worthy plan" covering "Architecture, security, data review" with "Risks ranked by blast radius" [A48]. Not a 5: no published engineering standard and no technical case study with production numbers |
| Commercial terms | 15% | 5/5 | A complete published monthly price ("Single stream $3,495/mo", "Dual stream $6,795/mo") [A37], "Flat monthly. No timesheet theatre." [A45], plus "If we're a fit, we'll set up a Proof of Quality — one real task, you judge the engineer before subscribing" [A44]. Not published: a change-order mechanism, or what overflow beyond the subscription costs |
| Exit terms | 10% | 4/5 | "Cancel any time", "No lock-in", "no notice period required" [A37] beats the notice benchmark outright [A30, A31]. Not a 5: the anchor also names transition assistance, and no handover scope or code-and-data return terms are published |
| Security | 10% | 4/5 | Against the 4 anchor — a published operating model that controls access and data location: "NDA before access", engineers "operate inside your environment under your access controls", and "Code and data don't leave your environment without your explicit approval" [A46]. Not a 5: there is no third-party attestation. Compliance is described as frameworks applied when required, which is a readiness statement and not a certificate [A38], and it earns nothing on this line |
| Communication cadence | 10% | 4/5 | "3-day task cycle, daily async updates" and a task-by-task approval gate [A37], plus a stated "4-hour response" to email [A47]. Not a 5: the anchor names an escalation path, and none is published |
| Working-model fit | 5% | 4/5 | "Dev On Demand runs async by design — submit a task, 3-day cycle, daily updates — so timezone alignment isn't required" [A47]. Not a fit for on-site or heavy-overlap buyers |
The arithmetic, so you can check it: 20 × 4/5 = 16 · 15 × 4/5 = 12 · 15 × 4/5 = 12 · 15 × 5/5 = 15 · 10 × 4/5 = 8 · 10 × 4/5 = 8 · 10 × 4/5 = 8 · 5 × 4/5 = 4. Totalling 16 + 12 + 12 + 15 + 8 + 8 + 8 + 4 = 83/100.
One 5 and seven 4s, and nothing below a 4 — which is a statement about where the anchors sit, not a claim of excellence. The 5 is commercial terms, where a complete monthly price is published and you can act on it without booking a call. Every other line clears its 4 anchor and stops there.
Where the seven 4s stop short of 5. Delivery evidence has named, published references and detail available on request, and no published case study with production numbers you can check without asking. Technical practice clears the 4 anchor and not the 5 one: no published engineering standard, no technical case study with production numbers. People, and working-model fit, rest on published positions rather than a record you can audit. Exit terms beat the notice benchmark and then run out of published detail: no transition assistance, nothing in writing about returning code and data. Security publishes an operating model and holds no third-party attestation. Communication publishes cadence and a response target but no escalation path.
What would move the scores, each one from a 4 to a 5: published case studies with production numbers a reader can check without asking would take delivery evidence to 5; a published engineering standard or a technical case study with production numbers would take technical practice; a third-party attestation — ISO 27001, SOC 2 Type II or equivalent — would take security; published transition-assistance and data-return terms would move exit; a published escalation path would move communication. None of them exists today, and none of them is promised here. If any of them arrives, it arrives as a document, not an adjective.
Which is why the 5 anchors stay where they are. A competitor publishing checkable case studies and holding a current third-party attestation scores above You-Source on this sheet, on the publisher's own arithmetic.
Take the two tables above, write your own 3-and-5 sentence under each criterion, and score three vendors before you've decided which one you want. If one of them sells a subscription, a single real ticket is the cheapest way to fill in criterion 3. That ordering is the hard part, not the arithmetic.
Three things. The criteria and weights are written down and published before the bids arrive, and UK guidance forbids changing them afterwards [A6]. Each bid is scored against your stated requirement, not against the other bids [A8]. And anything that could end the engagement is a pass/fail gate rather than a weighted line [A4]. A sheet built after the shortlist exists is a record of a decision already made, not an evaluation.
Score it absolutely, against your own stated budget, rather than giving the cheapest bid five points and scaling everyone else down. UK guidance at paragraph 7.2.1 states "relative price scoring should be treated with caution and not be used unless there is a specific business reason which has been approved by the commercial lead and the project SRO" [A12], and non-linear point curves flatten the field: "if all tenders get nearly the same price score in the end, the advantage of low-range prices is almost cancelled out" [A11].
Not if the failure is one that would end the engagement. A weighted total lets a vendor be catastrophic on one line and still win on points, which is why the World Bank permits a minimum quality threshold on a single criterion even when that criterion "carries relatively low weighting, but is important or critical to the Procurement outcomes" [A4]. A vendor failing a gate isn't scored low — it isn't scored.
Nobody has measured it. No study shows that structured vendor evaluation produces better outcomes than unstructured judgement. The nearest support comes from a different field entirely: a meta-analysis of 136 studies found mechanical prediction "about 10% more accurate than clinical predictions", holding "regardless of the judgment task, type of judges, judges' amounts of experience, or the types of data being combined" [A14]. That is clinical prediction, and it should be read as an analogy rather than proof.
Only as far as its derivation goes. Both published weighted templates found in this research come from companies selling into the market they score, and neither derivation can be checked: one rests on unpublished internal analysis, the other tells you to adjust the weights yourself [A34, A35]. This one is published by a vendor too, so the test is the same: every weight is either traced to a named public source or stated as our own judgement with the reason given, and the publisher runs the sheet on itself in public. You-Source scores 83/100 on its own scorecard — one 5, seven 4s, nothing lower — and claims no gate pass on published evidence. A vendor with a current SOC 2 Type II or ISO 27001 certificate scores 5 on security and beats it there, and one publishing case studies with production numbers a reader can check beats it on delivery evidence.