ESC
Type to search guides, tutorials, and reference documentation.

Vendor Management

Selecting, contracting for and living with software vendors: evaluation that survives the demo, what an SLA actually guarantees, pricing mechanics, lock-in and exit cost, renewal leverage, and third-party risk.

Vendor management is the work of choosing third-party software and services, contracting for them on terms you can live with, and operating alongside them for years — including on the day you decide to leave. In most engineering organisations the majority of systems running in production were written by somebody else, which makes this a core engineering competency rather than a procurement formality to be delegated and forgotten.


What Vendor Management Covers

The lifecycle has four phases, each with its own failure modes: selection, contracting, operation, and exit. Nearly all organisational attention goes to selection, because that is where the excitement and the decision-making authority are. Nearly all of the pain comes from the other three.

The structural reason is that leverage is highest before signature and declines steadily afterwards. Every term you did not ask for during selection becomes something you must request as a favour later, and every exit provision you did not negotiate becomes a cost you discover at the worst possible time.


Evaluation That Survives the Demo

A vendor demonstration is a rehearsed performance, given by the vendor's strongest engineer, on the vendor's own curated data, along the vendor's happy path. It is genuinely informative about what the product can do. It tells you very little about what the product will do in your environment, at your data volumes, with your edge cases, operated by your team.

Requirements Before Shortlist

Write down what you need before looking at products, and separate must-have from nice-to-have while you still have no emotional investment in an answer. Doing it in the other order means reverse-engineering the requirements from the product somebody already liked. The diagnostic sign is a requirements list that happens to match one vendor's feature page almost exactly.

Decide in advance what "good enough" means and what would make you stop. Evaluations without a stopping rule do not converge; they continue until the participants are exhausted, and the winner is whichever option still had an advocate with energy left.

The Proof of Concept

Run the evaluation against your own data, your own volume shape and your own integration points, and aim it deliberately at the ugliest case you have rather than a representative one. Representative cases are what the demo already covered.

Four things are worth testing explicitly because they rarely appear in a demo. The failure path: what the product does when it is unavailable, rate-limited or degraded, and what your systems do in response. The operational surface: what logs, metrics and audit trails you actually get, and whether you can debug a problem without opening a support ticket. The integration seam: how data gets in and out, whether the export is as capable as the import, and what the API does under real concurrency. The administration model: how permissions, tenancy and provisioning work, which is usually where the ugly surprises live.

Have the people who will operate the system in production run the evaluation. An architect who will hand it over has no incentive to discover the operational problems.


Contract Terms That Matter Later

What an SLA Actually Guarantees

A service level agreement is usually a refund schedule rather than an availability guarantee. The remedy for a breach is typically a credit against future fees, which is rarely proportionate to the loss the outage caused you, and which you generally receive only if you file a claim — within a defined window, on your own initiative, because the vendor does not file it for you.

Read what is being measured rather than the headline number. Exclusions commonly cover scheduled maintenance, degradation attributed to your configuration, problems in a dependency the vendor does not own, and anything classified as force majeure. Also check the measurement window, because the same nominal target behaves very differently measured monthly versus annually.

The correct way to treat an SLA is as a signal of the vendor's own confidence, not as insurance. If the service being unavailable would materially harm you, the mitigation is architectural — graceful degradation, caching, a fallback path, or a second supplier — because no credit schedule restores a lost day of trading.

Pricing Mechanics

The shape of a price matters more than its initial magnitude. What meter are you billed on, and is that meter something you control? A per-seat price behaves very differently from a price per request, per gigabyte ingested, or per unit of compute, and the difference shows up when your usage pattern changes rather than at signature.

Ask two symmetrical questions: what happens to the bill if usage doubles, and what happens if it halves. Most contracts answer these asymmetrically — growth is frictionless, contraction is not, because committed minimums and term lengths prevent you from reducing mid-term. That asymmetry is the point of the structure, and it is negotiable before signature and essentially not negotiable afterwards.

Get renewal uplift caps in writing. Without a cap, the renewal price is set by how difficult you would be to replace, which is a number that rises every year you stay.

Data and Exit

Establish before signing whether you can get your data out, in what format, how completely, and on what notice. Completeness is the term that hides the problems: an export that returns current records but not history, not attachments, not audit logs, and not the derived artefacts the product generated is not a usable exit.

Actually perform an export during the evaluation, while you still have leverage and attention, rather than discovering its limitations at the point you have decided to leave. Settle deletion at the same time: what happens to your data after termination, on what timeline, and whether the vendor will certify that it has been destroyed.


Lock-In and Exit Cost

Lock-in is not a binary property, and treating it as one produces bad decisions in both directions. It is a cost, and it comes from several distinct sources, roughly in increasing order of difficulty to unwind.

  1. Commercial. The remaining contract term and any committed spend. Bounded, knowable, and the easiest to reason about.
  2. Data gravity. The volume of data held by the vendor and the time, cost and risk of moving it somewhere else.
  3. Integration surface. How many of your systems call the vendor's interfaces, and how deeply. A single well-defined seam is cheap to replace; a hundred call sites spread across four services is not.
  4. Process and skill. The workflows your organisation has built around the product and the expertise your staff have accumulated in it. This is real cost that rarely appears in any spreadsheet.
  5. Semantic. The vendor's data model has quietly become your data model — its entities, its identifiers, its assumptions about how your business works. This is the hardest to undo and is usually invisible until somebody attempts a migration.

Accepting lock-in is often the right decision. The alternative is frequently a lowest-common-denominator abstraction that discards the specific capability you bought the product for, plus the ongoing cost of maintaining the abstraction and the near-certainty that it will leak at the worst moment. The mistake is not accepting lock-in; it is accepting it without knowing you did. Estimate the exit cost at selection time and re-estimate it at each renewal, so that the number is available when it is needed rather than being constructed under pressure.


Living With the Vendor

Give every significant vendor one named internal owner. Relationships with no owner fail quietly: the contract auto-renews unexamined, nobody tracks the incidents, and the relationship accumulates no institutional memory — so each renewal negotiation starts from zero while the vendor's account team starts from a full history.

Keep a running record: outages and how they were handled, support responsiveness against what was promised, commitments made about the roadmap and whether they landed. At renewal this record is the only evidence you have. Without it the conversation is your anecdotes against their account plan, and anecdotes lose.

Agree an escalation path before you need it, including a named human rather than a queue, and confirm it still exists when the account team changes. Then watch the calendar: renewal leverage is a function of remaining time and credible alternatives. Begin the renewal conversation early enough that switching could genuinely be executed if it came to that. A negotiation opened a month before expiry is not a negotiation. Diarise auto-renewal notice windows specifically, because missing one is the most common way an organisation silently forfeits all of its leverage for another full term.


Third-Party Risk

A vendor's security posture, availability and compliance status become part of yours, and so do those of their own subprocessors. The questions worth answering for anything holding meaningful data are narrow and concrete: what data do they hold and in which jurisdictions, who are their subprocessors, how and how quickly do they notify you of a breach, and what independent assurance can they produce rather than assert. Where regulated data is involved, bring your compliance function in at evaluation rather than at signature — see Compliance and Security.

Concentration risk deserves separate attention because it is easy to miss. Several critical vendors may resolve to the same underlying infrastructure provider or the same upstream service, which means a single failure can take out supposedly independent systems simultaneously. Map dependencies down to the provider level rather than stopping at the vendor level, or your redundancy is nominal.


Common Failure Modes

Failure modeWhat it looks like in practiceWhat to do instead
Requirements written after the demoThe requirement list matches one vendor's feature page almost exactlyWrite and freeze requirements before the first vendor conversation
Evaluating on the happy pathThe product performs well until it meets real data volume or a real edge caseRun the evaluation against your ugliest case, with the team who will operate it
Treating an SLA as insuranceAn outage costs far more than the service credit that compensates itMitigate architecturally; read the SLA as a confidence signal
Optimising the headline priceThe bill grows unexpectedly because the meter is not one you controlModel the price shape at double and half your current usage
Never testing the exportThe exit reveals that history, attachments or derived data cannot be retrievedPerform a full export during evaluation, while leverage is highest
No internal ownerThe contract auto-renews, incidents go unrecorded, the relationship has no memoryName one owner and keep a running performance record
Late renewal conversationNegotiating with no time to execute an alternative, so there is nothing to negotiate withDiarise notice windows and open the conversation while switching is still feasible
Unmapped concentrationIndependent-looking vendors fail together because they share an upstream providerMap dependencies to the provider level, not the vendor level

Key Takeaways

  • Leverage peaks before signature and declines steadily afterwards — ask for everything you will later need while you still can.
  • A demo shows capability; only an evaluation on your own worst data shows fit.
  • An SLA is a refund schedule, not an availability guarantee. Mitigate outages architecturally.
  • The shape of the price matters more than its size: know your meter and model it in both directions.
  • Test the export before you need it, and settle deletion terms at the same time.
  • Lock-in is a cost to quantify, not a state to avoid. Semantic lock-in is the expensive kind.
  • Give every vendor an owner, a record, and a renewal conversation that starts early enough to matter.
  • Map third-party dependencies to the provider level, or your redundancy is only nominal.