Hire Software Developers 7
Back to blogs

The Vendor Evaluation Scorecard, With Every Weight Shown

Eight blocks of different sizes stacked into a single column, representing weighted criteria summing to one score

The Vendor Evaluation Scorecard, With Every Weight Shown

Somewhere in the middle of a selection, someone finally builds the spreadsheet. That spreadsheet is the vendor evaluation scorecard: criteria down the left, vendors across the top, weights invented in a meeting where one person had already made up their mind. The numbers come out the way everyone expected, and nobody says the obvious thing: the scorecard didn't decide anything, it ratified something.

Key takeaways

  • A weighted vendor evaluation scorecard is only worth running if the criteria and weights are fixed and documented before the bids arrive. UK guidance is explicit: "you can't change the criteria or their weightings after your requirements are published" [A6].
  • Buyers arrive pre-committed. 41% already have a single preferred vendor before formal evaluation begins, and 92% start with at least one vendor in mind, per Forrester's 2024 Buyers' Journey Survey [A16].
  • Three checks are pass/fail before any scoring: a written IP assignment of all deliverables [A29], two named contactable clients from the last three years [A26, A27], and a stated exit with no charge to return code and data [A30, A31]. All three are questions you put to a vendor during the selection and get answered in writing — none of them can be scored off a website. A vendor failing any gate isn't scored low — it isn't scored [A4].
  • Score each bid against your requirement, never against the other bids. Any formula making one bid's score depend on another's content "includes an element of arbitrariness in the evaluation" [A10].
  • There is no measured evidence that structured vendor evaluation produces better outcomes than unstructured judgement. The nearest support is a meta-analysis of 136 studies finding mechanical prediction "about 10% more accurate than clinical predictions" [A14], and it comes from a different field. An analogy, not proof.
  • Disclosure: You-Source publishes this scorecard and scores itself with it in public. It gets 83/100 — one 5 and seven 4s, nothing below a 4 — and it does not claim a pass on the three gates from published evidence. A vendor holding a current SOC 2 Type II or ISO 27001 certificate scores 5 on security and beats it outright, and one publishing case studies with production numbers a reader can check beats it on delivery evidence too.

Ratifying a decision already made is the normal case, not the failure case. Forrester's 2024 Buyers' Journey Survey, as reported by Digital Commerce 360, found 92% of buyers start with at least one vendor in mind, and 41% already have a single preferred vendor before formal evaluation begins. Forrester's conclusion was that "B2B buying today is a process of confirmation, not selection" [A16]. 6sense's 2025 report asked buyers whether they could rank their shortlist before engaging any seller: "a resounding 94% of buyers answered yes" [A17].

A weighted vendor evaluation scorecard is a fixed list of evaluation criteria, each carrying a percentage weight, scored on one written scale and summed to a single total out of 100 [A2]. It is valid rather than decorative when three things hold: the criteria and weights are written down before the bids arrive [A6], each bid is scored against your stated requirement rather than against the other bids [A8], and anything that could end the engagement is a pass/fail gate instead of a weighted line [A4].

So a scorecard's job isn't precision. It's pre-commitment: fixing what you care about, and how much, before you can see which vendor it favours. Which is why how to evaluate software development vendors is a sequencing problem before it's a scoring one. What follows is that vendor evaluation template, rebuilt from the public method, with the derivation of every weight shown, traced to a named public source where one exists, and labelled as our own judgement with the reason given where none does, and a last section where we run it on ourselves and land a 4 on seven of the eight lines.

The vendor scorecards you'll find, and why you can't use them

Search for a software vendor evaluation scorecard and two serious weighted templates come back. Both were published by companies selling into the market they score.

Pangea.ai publishes eight dimensions — Technical Vetting Quality 20%, Talent Pool Depth 15%, Speed to Match 15%, Replacement Guarantees 15%, Compliance and Legal Infrastructure 10%, Contract Flexibility 10%, Communication and Account Management 10%, Pricing Transparency 5% — scored 1 to 5 and multiplied by weight, published 7 May 2026. It does state a derivation, the sheet is "built from analysis of vendor selection processes across hundreds of technology companies", but that analysis isn't published, so there is nothing to check the percentages against. And Pangea.ai is a staff augmentation marketplace competing in the market its scorecard evaluates [A34].

ZTABS publishes six categories — Technical Capability 25%, Portfolio and Experience 20%, Communication and Process 20%, Pricing and Value 15%, Cultural Fit 10%, References and Reputation 10% — published 30 March 2026. Its only note on provenance: "you can adjust the weights to match your priorities". ZTABS is itself a software development vendor [A35].

Neither is dishonest, and both give you usable structure. But look at Pangea's shape: Speed to Match and Replacement Guarantees at 15% each, Pricing Transparency at 5% [A34]. That weights what a marketplace is structurally good at and discounts what marketplaces are worst at. The problem isn't the numbers. It's that neither derivation is checkable: one points at unpublished internal analysis, the other at your own priorities. We went looking for a published, non-vendor weighted scorecard for software development vendors to compare them against and didn't find one, which is a statement about our search, not a proof that none exists.

Second disclosure: this one is published by a vendor too, on you-source.com. The difference on offer is procedure, not neutrality. Every weight is either traced to a named public source or stated as our own judgement with the reason given, one of the eight is the latter, and it's flagged where it appears. Several criteria are ones You-Source falls short of the top score on, and the last section runs the sheet against You-Source in public. Who those competitors are is a separate exercise.

Where the method actually comes from

Public procurement has spent decades scoring bids defensibly, because there a losing bidder can sue you over the method. Five documents carry most of the weight.

The World Bank's Evaluating Bids and Proposals with Rated Criteria, third edition, February 2025 is the closest thing to a manual. Weights come from a pairwise matrix: compare every criterion against every other, count the wins, rank, agree percentages. They "should add up to 100% in total" [A2]. Criteria "should be kept to the essential minimum" because too many "serves to dilute the important characteristics of Bid/Proposals" [A1]. A dominant criterion may take "perhaps 50% of the total technical weighting", with a worked 40/30/20/10 split [A3]. And a minimum quality threshold can sit on a single criterion that "carries relatively low weighting, but is important or critical to the Procurement outcomes" [A4]. That's where the gates come from.

The UKUPC Guide to Evaluation, November 2020 supplies the 0–5 scale with written descriptors [A7], the rule to compare a bid "against requirements, not against other tenders" [A8], and moderation to "a consensus result" [A9].

GOV.UK's Digital Outcomes and Specialists buyer guidance adds three rules worth copying: "You can't change the criteria or their weightings after your requirements are published"; one scoring scheme across all questions; and "document the reasons for awarding scores to each supplier and the reasons for choosing the criteria and weightings" [A6]. Its companion page gave a worked weighting: "30% on price and 60% on technical expertise, 10% on cultural fit" [A5]. Published 1 March 2021, withdrawn 31 December 2025 as outdated: the clearest non-vendor worked example for digital services, and no longer policy.

Cabinet Office Procurement Policy Note 04/15, 25 March 2015 required buyers to satisfy themselves that a supplier's "principal relevant contracts in the last three years" had been "satisfactorily performed", via written performance certificates stating reasons for any unsatisfactory rating: delays, incomplete scope, missed service levels, other contractual failures [A26, A27]. It is a 2015 policy note and not something you can hold a private vendor to, borrow the method, not the legal status.

Kiiver and Kodym, Journal of Public Procurement 15(3), 2015 — two European Parliament procurement officers on how scoring formulas break: "any formula that makes the score of one tender dependent on the content of another tender includes an element of arbitrariness in the evaluation and, at worst, exposes the process to deliberate manipulation by colluding tenderers" [A10].

One caveat before you spend a week on this. There is no measured evidence that structured vendor evaluation produces better outcomes than unstructured judgement. No such study exists. The nearest support is from another field: a meta-analysis of 136 studies found mechanical prediction "about 10% more accurate than clinical predictions", holding "regardless of the judgment task, type of judges, judges' amounts of experience, or the types of data being combined" [A14]. That's clinical prediction, not procurement. An analogy, and it should be read as one. Dawes' 1979 work on improper linear models points the same way, as context rather than proof [A15].

Three gates before you score anything

Weighted totals have a failure mode: a vendor can be catastrophic on one line and still win on points. The World Bank's fix is a minimum quality threshold on a single criterion [A4]. Three belong there. Each one is binary, cheap to check, and capable of ending the engagement badly however good the code was.

One thing to be clear about before you use them: all three gates are questions you ask a vendor during the selection, and none of them can be scored from a website. A contract clause, a client's permission to be phoned, and a written exit term are things a vendor tells you in writing when you ask. A marketing page can indicate a position; it cannot clear a gate, not the vendors you're evaluating, and not the one publishing this.

Gate 1 — a written IP assignment of all deliverables. The UK Intellectual Property Office is blunt: for commissioned work "the first legal owner of copyright is the person or organisation that created the work and not you the commissioner, unless you otherwise agree it in writing" [A29]. No written assignment, and the vendor owns the code you paid for.

Gate 2 — two named, contactable clients from the last three years, with permission to call them. PPN 04/15's standard, scaled down [A26, A27]. A logo wall isn't a reference. Neither is a marketplace listing: Crown Commercial Service states of the Digital Marketplace that "Crown Commercial Service (CCS) don't evaluate suppliers or their individual services" [A28].

Gate 3 — a stated exit: notice period, handover scope, no charge to return code and data. Under the EU Data Act's switching rules the parties "must contractually agree a notice period that does not exceed two months", with a transitional period of "a maximum of 30 calendar days" [A31], and from 12 January 2027 in-scope providers "won't be able to charge their customers for the operations that are necessary to facilitate switching or for data egress" [A30]. Read the scope: those rules bind data processing services (cloud, SaaS, PaaS, IaaS), not development agencies [A30, A31]. Use it as a benchmark, never as an obligation you can claim a dev vendor is already under.

A vendor failing any gate isn't scored low. It isn't scored.

Eight software vendor selection criteria, and why each weight

Eight already pushes the World Bank's "essential minimum" advice [A1], and the weights sum to 100 as required [A2]. Each weight is a declared preference with a stated reason. Change them if your reason is better. But write it down, before the bids arrive [A6].

#CriterionWeightWhy this weight
1Verifiable delivery evidence20%The only criterion with government policy behind it: named contracts in the last three years, written performance certificates, defined failure categories [A26, A27]. Also the only line a third party can falsify. World Bank guidance illustrates a vastly-more-important criterion taking "perhaps 50% of the total technical weighting" [A3] — an example, not a ceiling, and half of a technical score that excludes cost, so it isn't directly comparable to 20% of a sheet that includes commercial terms. It's the shape of the advice we're following, not the number.
2People you actually get15%US federal contracting treats key personnel as a controlled contract term: 30 days' notice, a named replacement of equal or better credentials, and no substitution "without the written consent of the Contracting Officer" [A32]. Scoring this closes the bait-and-switch gap a "team quality" question leaves open.
3Technical practice15%The 2025 DORA report, covering nearly 5,000 technology professionals, found 90% use AI at work, 30% report little or no trust in AI-generated code, and a negative relationship between AI adoption and delivery stability alongside a positive one with throughput [A33]. So score it from artefacts, not answers.
4Commercial terms15%UK guidance at paragraph 7.2.1 states "relative price scoring should be treated with caution and not be used unless there is a specific business reason which has been approved by the commercial lead and the project SRO" [A12], and scoring one bid against another invites manipulation [A10, A11]. The 2021 UK example put price at 30% [A5]; 15% is a stated choice, because a vendor who won't publish a number can't be scored against a budget. The objection, from the same article we cite here: Kiiver and Kodym also warn that in an additive price-plus-quality model "the weight of the price criterion must never drop below 50%", because below that the price criterion "can get cancelled out entirely" and the sheet "imposes an implicit minimum quality threshold, restricting competition at the low end of the market" [A10, A11]. We depart from it knowingly. Their model is a commodity bid where the requirement is fixed and price is the live variable; this sheet buys an ongoing supplier relationship where the variable is who does the work, and it accepts the consequence they name — the cheapest vendor cannot win this sheet on price alone. If your requirement really is a fixed, commoditised scope, follow them and not us.
5Exit terms10%A concrete external benchmark exists [A30, A31], which is the reason this line is on the sheet at all. WorldCC and Deloitte report contract value erosion of 8.6% across 1,236 organizations, best performers a little over 3%, worst more than 20% [A25] — that measures the whole contracting lifecycle rather than exit terms specifically, so read it as a reason to take post-signature terms seriously, not as a derivation of this number.
6Security and data-access model10%The one weight with no public source behind it: 10% is our own judgement, and here's the reason. It's gate-and-verify territory — an attestation either exists and is checkable, or it doesn't — so the top of the line rewards a binary you can confirm rather than a description you can't. Certifications, sub-processors and Article 28 terms belong in the security questionnaire you send every vendor. If security is the thing that would end your engagement, raise it, or make it a gate.
7Communication cadence and reporting10%The one line both vendor scorecards agree on: Pangea.ai weights it 10%, ZTABS 20% [A34, A35]. Agreement between competitors is weak evidence, but it is evidence, and 10% sits at the bottom of their range.
8Working-model fit5%The 2021 UK guidance made cultural fit a required dimension at 10% [A5]. Halved because it's the least falsifiable line on the sheet and the most exposed to the halo effect moderation exists to catch [A9].

To derive your own, use the World Bank pairwise method [A2]. And if you're still deciding between an agency, a subscription and a permanent hire, settle the staffing model first. The weights follow from it.

What scoring scale should you use?

Use the UKUPC anchors, written on the sheet before anyone scores: 0 cannot be scored, 1 unsatisfactory, 2 satisfactory, 3 good, 4 very good, 5 excellent [A7]. Add a sentence per criterion saying what a 3 and a 5 look like for your requirement.

Score independently, then moderate to a consensus, with challenges encouraged [A9]. Write down the reason for each final score, because UK guidance requires documenting both the scores and the choice of criteria and weightings [A6].

Then the rule that matters most: score each bid against your requirement, never against the other bids. UKUPC states it twice, including in moderation, where "each individual tender must be considered independently of the others and the scoring process must not be influenced by comparable scores for other tenders" [A8].

Which is why this template scores price absolutely, against your stated budget, rather than giving the cheapest bid five points and scaling everyone else down. UK guidance discourages that without senior approval [A12], and Kiiver and Kodym show what non-linear point curves do: "if all tenders get nearly the same price score in the end, the advantage of low-range prices is almost cancelled out" [A11]. Set a budget, anchored on what the work actually costs. Describe a 3 and a 5 against it. Score each vendor alone.

Where a weighted vendor scoring matrix breaks

A method you can't criticise is a method you shouldn't trust. Seven ways this one fails.

You over-tune the weights. The flat-maximum effect holds that "the predictive ability of linear models is insensitive to large variations in the size of regression weights" and to the number of predictors; tested against ten credit unions' models from 1984 to 1988, a generic weighted-average model performed "very close to that of the empirically derived models" [A13]. Arguing 15% versus 18% is wasted time. Choosing the right eight criteria isn't.

You score bids against each other. Any formula making one bid's score depend on another's content "includes an element of arbitrariness in the evaluation and, at worst, exposes the process to deliberate manipulation by colluding tenderers" [A10].

Your scores compress. If every vendor lands between 3.2 and 3.6, the sheet has stopped discriminating [A11]. Widen the descriptors, or accept you have three acceptable vendors.

You add criteria. Every extra line dilutes the important characteristics [A1].

You change criteria after publishing them. UK guidance forbids it outright [A6], and with 41% of buyers walking in with a favourite [A16], the temptation isn't hypothetical.

One person scores everything. No independent scoring, no moderation, no consensus [A9], and one impressive demo drags every unrelated line up with it.

A fatal low-weight line passes on total. That's what gates are for [A4]. If something would end the engagement, it doesn't belong in a weighted average at any weight.

And the caveat again, beside the criticisms, not buried at the end: the evidence that mechanical combination beats unstructured judgement comes from clinical prediction [A14]. Nobody has measured whether it picks better software vendors.

The vendor evaluation scorecard template

Lift this.

Gates — pass or fail, before any scoring [A4]. Ask each one of the vendor and get the answer in writing; none of them is scoreable from a website.

GateTestFail condition
IP assignmentWritten assignment of all deliverables to you, in the contract, no carve-outs [A29]Absent, verbal, or "on final payment" with final undefined
ReferencesTwo named, contactable clients from the last 3 years, with permission to ask about performance [A26, A27]Logos only, NDA'd out, or nothing inside three years
ExitStated notice, handover scope, no charge to return code and data. Benchmark: ≤2 months' notice, ≤30 days' transition, no egress fee from 12 Jan 2027 [A30, A31]No stated notice, or a charge to get your repository back

Criteria and weights [A2]

#CriterionWeightWhat a 5 looks like
1Verifiable delivery evidence20%5: published case studies with production numbers a reader can check without asking — plus three-plus comparable contracts in 3 years, references taken, written detail on what went wrong. 4: named, published references, plus detailed delivery evidence available on request
2People you actually get15%Named individuals with CVs, substitution only by your written consent, documented cover for the lead [A32]
3Technical practice15%5: a published engineering standard, or a technical case study with production numbers — plus the artefacts on request: a real PR, a test suite, a CI config, a written policy on reviewing AI-generated code [A33]. 4: credible published practice statements and a demonstrable review offering, without the standard or the case study
4Commercial terms15%A complete price you can act on, a defined change-order mechanism, no unpriced assumptions
5Exit terms10%The gate 3 benchmark met or beaten, named transition assistance
6Security and data-access model10%5: a current third-party attestation (ISO 27001, SOC 2 Type II or equivalent), plus work inside your environment under your controls and sub-processors listed. 4: a published operating model that controls access and data location, without third-party attestation
7Communication cadence10%Stated frequency, named channel, named escalation path, demonstrated response time
8Working-model fit5%The overlap you need, decision latency in hours, evidence they'll push back

Scale [A7]: 0 cannot be scored · 1 unsatisfactory · 2 satisfactory · 3 good · 4 very good · 5 excellent.

Totalling: weighted points = weight × (score ÷ 5); sum the eight for a score out of 100. Score independently, moderate to consensus, document the reason [A6, A9].

A trial, pilot or paid proof-of-concept scores under commercial terms for every vendor — a competitor offering a free pilot earns the same credit here that our Proof of Quality does.

Scoring ourselves with our own vendor evaluation scorecard

Same sheet, run on its publisher. Everything below is what You-Source publishes, so you can check it.

Gates. No pass marks here, because the gates are answered in a selection process and this section only has a website to work from. What You-Source publishes against each, and what it doesn't:

  • IP assignment. The page states "100% you own the code" [A37] and answers ownership as "You do. Always. All code written for your project belongs to you, with no hidden clauses or carve-outs" [A39]. That is a published position. The gate is a written assignment clause in your contract, which is a document, not a web page — ask for it.
  • References. Four named client testimonials, each with a person and a company: Kevin Conlon (Kinnect), Derek Sturdy (Robur), Michael Hanna (Solargain), Eric Peterson (Stack Moxie) [A40]. A published testimonial is not a contactable reference. Two clients you can phone, from the last three years, with permission to ask about performance, are supplied on request — as they should be by everyone on your list.
  • Exit. "Cancel any time", "No lock-in", no notice period required [A37]. That covers one of the gate's three elements. No transition-assistance scope and no terms on returning code and data are published; those are answered on request.

Three published positions, one cleared gate element, and the rest to be asked for. That is where every vendor on your shortlist starts, including this one.

CriterionWeightScoreBasis
Verifiable delivery evidence20%4/5Against the 4 anchor — named, published references, plus detailed delivery evidence available on request. Four named client testimonials, each with a named company: Kevin Conlon (Kinnect), Derek Sturdy (Robur), Michael Hanna (Solargain), Eric Peterson (Stack Moxie) [A40], with case studies available on request. The caveat stands: "60+ Projects Delivered", "5 Continents served" and a "98.4% Renewal rate" are all self-reported [A41]. Not a 5: nothing published here is a case study with production numbers a reader can check without asking
People you actually get15%4/5Named dedicated engineer, "Same-day match" [A38], "5-day replacement guarantee", swapped at no cost [A37]. Bench depth is thin at 1–2 engineers per subscription [A37]
Technical practice15%4/5Against the 4 anchor — credible published practice statements plus a demonstrable review offering. The statements: "Production-grade work — reviewed, tested, deployed. Built to ship, not to demo." [A42] and a task-by-task approval gate [A37]. The offering: a Code Audit that maps "what's fragile, what won't scale, and what's blocking production" and hands over "a prioritized fix plan", sold as "3 days • Senior review • Fixed price" [A43], specified elsewhere as "Three days, two senior engineers, one production-worthy plan" covering "Architecture, security, data review" with "Risks ranked by blast radius" [A48]. Not a 5: no published engineering standard and no technical case study with production numbers
Commercial terms15%5/5A complete published monthly price ("Single stream $3,495/mo", "Dual stream $6,795/mo") [A37], "Flat monthly. No timesheet theatre." [A45], plus "If we're a fit, we'll set up a Proof of Quality — one real task, you judge the engineer before subscribing" [A44]. Not published: a change-order mechanism, or what overflow beyond the subscription costs
Exit terms10%4/5"Cancel any time", "No lock-in", "no notice period required" [A37] beats the notice benchmark outright [A30, A31]. Not a 5: the anchor also names transition assistance, and no handover scope or code-and-data return terms are published
Security10%4/5Against the 4 anchor — a published operating model that controls access and data location: "NDA before access", engineers "operate inside your environment under your access controls", and "Code and data don't leave your environment without your explicit approval" [A46]. Not a 5: there is no third-party attestation. Compliance is described as frameworks applied when required, which is a readiness statement and not a certificate [A38], and it earns nothing on this line
Communication cadence10%4/5"3-day task cycle, daily async updates" and a task-by-task approval gate [A37], plus a stated "4-hour response" to email [A47]. Not a 5: the anchor names an escalation path, and none is published
Working-model fit5%4/5"Dev On Demand runs async by design — submit a task, 3-day cycle, daily updates — so timezone alignment isn't required" [A47]. Not a fit for on-site or heavy-overlap buyers

The arithmetic, so you can check it: 20 × 4/5 = 16 · 15 × 4/5 = 12 · 15 × 4/5 = 12 · 15 × 5/5 = 15 · 10 × 4/5 = 8 · 10 × 4/5 = 8 · 10 × 4/5 = 8 · 5 × 4/5 = 4. Totalling 16 + 12 + 12 + 15 + 8 + 8 + 8 + 4 = 83/100.

One 5 and seven 4s, and nothing below a 4 — which is a statement about where the anchors sit, not a claim of excellence. The 5 is commercial terms, where a complete monthly price is published and you can act on it without booking a call. Every other line clears its 4 anchor and stops there.

Where the seven 4s stop short of 5. Delivery evidence has named, published references and detail available on request, and no published case study with production numbers you can check without asking. Technical practice clears the 4 anchor and not the 5 one: no published engineering standard, no technical case study with production numbers. People, and working-model fit, rest on published positions rather than a record you can audit. Exit terms beat the notice benchmark and then run out of published detail: no transition assistance, nothing in writing about returning code and data. Security publishes an operating model and holds no third-party attestation. Communication publishes cadence and a response target but no escalation path.

What would move the scores, each one from a 4 to a 5: published case studies with production numbers a reader can check without asking would take delivery evidence to 5; a published engineering standard or a technical case study with production numbers would take technical practice; a third-party attestation — ISO 27001, SOC 2 Type II or equivalent — would take security; published transition-assistance and data-return terms would move exit; a published escalation path would move communication. None of them exists today, and none of them is promised here. If any of them arrives, it arrives as a document, not an adjective.

Which is why the 5 anchors stay where they are. A competitor publishing checkable case studies and holding a current third-party attestation scores above You-Source on this sheet, on the publisher's own arithmetic.

Take the two tables above, write your own 3-and-5 sentence under each criterion, and score three vendors before you've decided which one you want. If one of them sells a subscription, a single real ticket is the cheapest way to fill in criterion 3. That ordering is the hard part, not the arithmetic.

Frequently asked questions

What makes a vendor evaluation scorecard valid rather than decorative?

Three things. The criteria and weights are written down and published before the bids arrive, and UK guidance forbids changing them afterwards [A6]. Each bid is scored against your stated requirement, not against the other bids [A8]. And anything that could end the engagement is a pass/fail gate rather than a weighted line [A4]. A sheet built after the shortlist exists is a record of a decision already made, not an evaluation.

How should you score price on a vendor scorecard?

Score it absolutely, against your own stated budget, rather than giving the cheapest bid five points and scaling everyone else down. UK guidance at paragraph 7.2.1 states "relative price scoring should be treated with caution and not be used unless there is a specific business reason which has been approved by the commercial lead and the project SRO" [A12], and non-linear point curves flatten the field: "if all tenders get nearly the same price score in the end, the advantage of low-range prices is almost cancelled out" [A11].

Should a vendor that fails badly on one criterion still be scored?

Not if the failure is one that would end the engagement. A weighted total lets a vendor be catastrophic on one line and still win on points, which is why the World Bank permits a minimum quality threshold on a single criterion even when that criterion "carries relatively low weighting, but is important or critical to the Procurement outcomes" [A4]. A vendor failing a gate isn't scored low — it isn't scored.

Does a structured scorecard actually pick better software vendors than judgement?

Nobody has measured it. No study shows that structured vendor evaluation produces better outcomes than unstructured judgement. The nearest support comes from a different field entirely: a meta-analysis of 136 studies found mechanical prediction "about 10% more accurate than clinical predictions", holding "regardless of the judgment task, type of judges, judges' amounts of experience, or the types of data being combined" [A14]. That is clinical prediction, and it should be read as an analogy rather than proof.

Can you trust a vendor evaluation scorecard published by a vendor?

Only as far as its derivation goes. Both published weighted templates found in this research come from companies selling into the market they score, and neither derivation can be checked: one rests on unpublished internal analysis, the other tells you to adjust the weights yourself [A34, A35]. This one is published by a vendor too, so the test is the same: every weight is either traced to a named public source or stated as our own judgement with the reason given, and the publisher runs the sheet on itself in public. You-Source scores 83/100 on its own scorecard — one 5, seven 4s, nothing lower — and claims no gate pass on published evidence. A vendor with a current SOC 2 Type II or ISO 27001 certificate scores 5 on security and beats it there, and one publishing case studies with production numbers a reader can check beats it on delivery evidence.

Sources

  • [A1] World Bank procurement guidance states "Generally, the overall number of Rated Criteria should be kept to the essential minimum. Having too many Rated Criteria often serves to dilute the important characteristics of Bid/Proposals and makes identification of the optimal Bidder/Proposer more difficult." — https://thedocs.worldbank.org/en/doc/9dcb7971706bf29b2732779c39922b77-0290012025/original/Evaluating-Bids-and-Proposals-with-Rated-Criteria-Feb-4-2025.pdf
  • [A2] The same World Bank guidance sets out a pairwise prioritisation matrix for deriving weights — each criterion is compared head-to-head against every other, the letters are counted to produce a priority order, and the team then agrees percentages — and states "The weightings of all Rated Criteria should add up to 100% in total." — https://thedocs.worldbank.org/en/doc/9dcb7971706bf29b2732779c39922b77-0290012025/original/Evaluating-Bids-and-Proposals-with-Rated-Criteria-Feb-4-2025.pdf
  • [A3] The World Bank guidance advises that if one criterion is vastly more important than the others it "might receive perhaps 50% of the total technical weighting", and gives a worked four-criterion split of 40%, 30%, 20% and 10%. — https://thedocs.worldbank.org/en/doc/9dcb7971706bf29b2732779c39922b77-0290012025/original/Evaluating-Bids-and-Proposals-with-Rated-Criteria-Feb-4-2025.pdf
  • [A4] The World Bank guidance describes minimum quality thresholds that can be applied to the total score, to a subset of criteria, or to a single criterion — the last being useful "where the specific Rated Criteria/Subcriteria carries relatively low weighting, but it is important or critical to the Procurement outcomes" — with bids below the threshold rejected before cost is even considered. — https://thedocs.worldbank.org/en/doc/9dcb7971706bf29b2732779c39922b77-0290012025/original/Evaluating-Bids-and-Proposals-with-Rated-Criteria-Feb-4-2025.pdf
  • [A5] UK Government guidance for buying digital specialists required buyers to evaluate technical competence, cultural fit and price, and gave a worked example weighting of "30% on price and 60% on technical expertise, 10% on cultural fit"; the page was published 1 March 2021 and withdrawn on 31 December 2025 as outdated. — https://www.gov.uk/guidance/how-to-find-a-digital-specialist-on-digital-marketplace
  • [A6] UK Digital Outcomes and Specialists buyer guidance states "You can't change the criteria or their weightings after your requirements are published", requires that buyers "mark all questions using the same scoring scheme and apply the relevant weighting for that particular criteria to get a score", and requires buyers to "Document the reasons for awarding scores to each supplier and the reasons for choosing the criteria and weightings." — https://www.gov.uk/guidance/digital-outcomes-and-specialists-buyers-guide
  • [A7] The UKUPC Guide to Evaluation publishes a 0–5 scoring scale with written descriptors — 0 cannot be scored, 1 unsatisfactory, 2 satisfactory, 3 good, 4 very good, 5 excellent — and notes buyers "may decide to use alternative values for example 0-3 or 0-6" provided the scale "must be clear to understand for all evaluators." — https://www.lupc.ac.uk/sites/default/files/UKUPC%20-%20Guide%20to%20Evaluation%20-%20Nov%202020%20(002).pdf
  • [A8] The UKUPC Guide to Evaluation instructs evaluators that "when evaluating you ensure you compare the bid against requirements, not against other tenders", and that in moderation "Each individual tender must be considered independently of the others and the scoring process must not be influenced by comparable scores for other tenders." — https://www.lupc.ac.uk/sites/default/files/UKUPC%20-%20Guide%20to%20Evaluation%20-%20Nov%202020%20(002).pdf
  • [A9] The UKUPC guide describes moderation as a consensus meeting in which "Moderators are responsible for directing a review of all the evaluation answers to provide a consensus result by selecting the most representative score", with challenges to individual scores encouraged. — https://www.lupc.ac.uk/sites/default/files/UKUPC%20-%20Guide%20to%20Evaluation%20-%20Nov%202020%20(002).pdf
  • [A10] Kiiver and Kodym, procurement officers at the European Parliament writing in the Journal of Public Procurement, warn that "any formula that makes the score of one tender dependent on the content of another tender includes an element of arbitrariness in the evaluation and, at worst, exposes the process to deliberate manipulation by colluding tenderers." — https://www.ippa.org/images/JOPP/vol15/issue-3/Article_1_Kiiver_Kodym.pdf
  • [A11] The same Journal of Public Procurement article describes score compression, noting that where point distribution curves are not linear "some tenders will be disadvantaged and will have to offer more value for money than others in order to compensate", and that "if all tenders get nearly the same price score in the end, the advantage of low-range prices is almost cancelled out." — https://www.ippa.org/images/JOPP/vol15/issue-3/Article_1_Kiiver_Kodym.pdf
  • [A12] UK government guidance quoted at paragraph 7.2.1 states "Relative price scoring should be treated with caution and not be used unless there is a specific business reason which has been approved by the commercial lead and the project SRO", with value-for-money ratio scoring proposed by the Government Commercial Function as the alternative. — https://www.shma.co.uk/our-thoughts/evaluating-price-in-tenders-new-guidance-on-new-practices/
  • [A13] The flat-maximum effect holds that "the predictive ability of linear models is insensitive to large variations in the size of regression weights" and the number of predictors used, so that seemingly different scoring systems generate comparable results; testing ten credit unions' scoring models from 1984 to 1988, the authors found a generic weighted-average model performed "very close to that of the empirically derived models". — https://academic.oup.com/imaman/article-abstract/4/1/97/656018
  • [A14] A meta-analysis of 136 studies found that "mechanical-prediction techniques were about 10% more accurate than clinical predictions", that mechanical prediction substantially outperformed clinical judgement in 33%–47% of studies while clinical judgement was substantially more accurate in only 6%–16%, and that this superiority held "regardless of the judgment task, type of judges, judges' amounts of experience, or the types of data being combined." — https://experts.umn.edu/en/publications/clinical-versus-mechanical-prediction-a-meta-analysis/
  • [A15] Robyn Dawes' paper on improper linear models, which found that equal ("unit") weighting is robust and that improper linear models outperform clinical intuition, was published in American Psychologist, 1979, volume 34, issue 7, pages 571–582. — https://web.stanford.edu/~knutson/nfc/dawes79.pdf
  • [A16] Forrester's 2024 Buyers' Journey Survey found that "92% of buyers start with at least one vendor in mind" and "41% already have a single preferred vendor selected before formal evaluation begins", with Forrester concluding "B2B buying today is a process of confirmation, not selection." — https://www.digitalcommerce360.com/2025/07/07/forrester-b2b-buyers-choose-vendors-before-the-buying-process-begins/
  • [A17] 6sense's 2025 B2B Buyer Experience Report found that when asked whether their team was able to put the shortlist in order of preference prior to engaging with sellers or SDRs, "A resounding 94% of buyers answered yes". — https://6sense.com/science-of-b2b/buyer-experience-report-2025/
  • [A25] The WorldCC and Deloitte report "The ROI of contracting excellence" states that 2014 research by World Commerce & Contracting, then named IACCM, "indicated average value erosion of 9.2% of contract value", and that on updated data "erosion now stands at 8.6%, with the best performers operating at a little over 3% and the worst more than 20%"; the update drew on 1,236 organizations with data collected between April 2021 and December 2022. — https://www.deloitte.com/content/dam/assets-zone3/us/en/docs/services/tax/2024/us-tax-roi-of-contracting-excellence.pdf
  • [A26] UK Cabinet Office Procurement Policy Note 04/15 requires in-scope organisations to satisfy themselves "that suppliers' principal relevant contracts in the last three years are being or have been satisfactorily performed in accordance with their terms", to obtain a list of past contracts and written Certificates of performance, and where appropriate to "verify information provided by any supplier in relation to past performance." — https://assets.publishing.service.gov.uk/media/5a807c70ed915d74e33fab47/PPN04-15_Supplier_Past_Performance_.pdf
  • [A27] PPN 04/15 specifies that where a performance Certificate records unsatisfactory performance it should give reasons, which "may include" delays in providing goods or services, failure to supply everything in the contracted scope, failures to meet service levels or quality standards, and any other failure to comply with contractual obligations; the policy applied to ICT, facilities management and business process outsourcing contracts with a total anticipated value of £20 million or greater. — https://assets.publishing.service.gov.uk/media/5a807c70ed915d74e33fab47/PPN04-15_Supplier_Past_Performance_.pdf
  • [A28] Crown Commercial Service states plainly of the Digital Marketplace that "Crown Commercial Service (CCS) don't evaluate suppliers or their individual services", and that buyers "must write clear requirements and evaluate suppliers' services against them." — https://www.gov.uk/guidance/how-digital-marketplace-suppliers-have-been-evaluated
  • [A29] The UK Intellectual Property Office states that for commissioned work "the first legal owner of copyright is the person or organisation that created the work and not you the commissioner, unless you otherwise agree it in writing", and that a person working under a contract for services "will usually retain copyright in any works he produces, unless there is a contractual agreement to the contrary." — https://www.gov.uk/guidance/ownership-of-copyright-works
  • [A30] The EU Data Act has applied since 12 September 2025, and from 12 January 2027 providers of in-scope data processing services "won't be able to charge their customers for the operations that are necessary to facilitate switching or for data egress". — https://digital-strategy.ec.europa.eu/en/factpages/data-act-explained
  • [A31] Under the EU Data Act's switching rules, "the provider and the customer must contractually agree a notice period that does not exceed two months", and the mandatory transitional period is "a maximum of 30 calendar days" extendable to up to seven months only where the provider demonstrates technical unfeasibility. — https://www.pinsentmasons.com/out-law/guides/switching-porting-rules-eu-data-act
  • [A32] US federal contract clause HHSAR 352.237-75 Key Personnel requires the contractor to give at least 30 days' prior notice before diverting or replacing key personnel, to identify the proposed replacement and explain how their skills, experience and credentials meet or exceed the requirements, and states "The Contractor shall not divert, replace, or announce any such change to key personnel without the written consent of the Contracting Officer." — https://www.acquisition.gov/hhsar/352.237-75-key-personnel
  • [A33] The 2025 DORA report, based on nearly 5,000 technology professionals surveyed worldwide plus over 100 hours of qualitative data, found 90% of respondents report using AI at work, more than 80% believe it has increased their productivity, and 30% report little or no trust in AI-generated code; it also found a positive relationship between AI adoption and software delivery throughput but a negative relationship between AI adoption and software delivery stability. — https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report
  • [A34] Pangea.ai publishes an 8-dimension weighted vendor evaluation scorecard — Technical Vetting Quality 20%, Talent Pool Depth 15%, Speed to Match 15%, Replacement Guarantees 15%, Compliance and Legal Infrastructure 10%, Contract Flexibility 10%, Communication and Account Management 10%, Pricing Transparency 5% — scored 1 to 5 and multiplied by weight; the page, and Pangea.ai is itself a staff augmentation marketplace competing in the market its scorecard evaluates. Pangea states a derivation - "built from analysis of vendor selection processes across hundreds of technology companies" - but does not publish the analysis, so the weights cannot be checked. — https://pangea.ai/resources/staff-augmentation-vendor-selection
  • [A35] ZTABS publishes a six-category weighted scorecard — Technical Capability 25%, Portfolio and Experience 20%, Communication and Process 20%, Pricing and Value 15%, Cultural Fit 10%, References and Reputation 10% — on a 1 to 5 scale, states only that "you can adjust the weights to match your priorities" and gives no derivation for the weights; ZTABS is itself a software development vendor. — https://ztabs.co/blog/software-vendor-evaluation-scorecard
  • [A37] You-Source publishes Dev On Demand at "$3,495/mo" for a single stream with "1 dedicated AI-augmented engineer" and "$6,795/mo" for a dual stream with "2 dedicated AI-augmented engineers in parallel", both including "5-day first ship, 100% you own the code", on a "3-day task cycle, daily async updates" with "Task-by-task approval gate", a "5-day replacement guarantee", "Cancel any time", "No lock-in" and "No notice period required". — https://you-source.com/dev-on-demand
  • [A38] The same You-Source page states engineers are "Same-day match" with a "Kickoff call within 24 hours of signing", claims "60+ Projects Delivered" and a "98.4% Renewal rate", and describes compliance as "SOC 2 and GDPR frameworks when required" rather than as a held certification. — https://you-source.com/dev-on-demand
  • [A39] You-Source states code ownership on the Dev On Demand page: "You do. Always. All code written for your project belongs to you, with no hidden clauses or carve-outs." — https://www.you-source.com/dev-on-demand
  • [A40] The Dev On Demand page carries four named client testimonials, each with a person and a company: Kevin Conlon (Kinnect), Derek Sturdy (Robur), Michael Hanna (Solargain) and Eric Peterson (Stack Moxie). — https://www.you-source.com/dev-on-demand
  • [A41] You-Source publishes the self-reported metrics "60+ Projects Delivered", "5 Continents served" and "98.4% Renewal rate". — https://www.you-source.com/
  • [A42] The You-Source homepage states "Production-grade work — reviewed, tested, deployed. Built to ship, not to demo." — https://www.you-source.com/
  • [A43] The You-Source Code Audit is described as "Built with AI tools? Let's make it production-ready. We map what's fragile, what won't scale, and what's blocking production — then hand you a prioritized fix plan. 3 days • Senior review • Fixed price". — https://www.you-source.com/
  • [A44] You-Source states: "If we're a fit, we'll set up a Proof of Quality — one real task, you judge the engineer before subscribing." — https://www.you-source.com/dev-on-demand
  • [A45] You-Source states "Flat monthly. No timesheet theatre." and "Flat monthly subscription. Avoid unpredictable invoices". — https://www.you-source.com/dev-on-demand
  • [A46] You-Source states: "NDA before access. Engineers operate inside your environment under your access controls — including SOC 2 and GDPR frameworks when required. Code and data don't leave your environment without your explicit approval." — https://www.you-source.com/dev-on-demand
  • [A47] You-Source states a "4-hour response" to email, and "Dev On Demand runs async by design — submit a task, 3-day cycle, daily updates — so timezone alignment isn't required." — https://www.you-source.com/dev-on-demand
  • [A48] The You-Source /development page describes the Code Audit as "Three days, two senior engineers, one production-worthy plan", covering "Architecture, security, data review", with "Risks ranked by blast radius" and a "Fixed-price quote to ship it prod-worthy". — https://www.you-source.com/development
back to top

Related Articles

Book 30 min with Albert
Smiling man with short dark hair and glasses wearing a black suit, white shirt, and black tie against blue background.
Tell Albert what you're shipping.
He'll read this before joining the call. Phone number comes next, on the calendar step.
↳ info@you-source.com
↳ 4-hour response
Please wait while we retrieve meeting schedules.
Oops! There's a problem with your request. We're working on fixing it. Please try again later.