Own the system, then choose who runs it
A vendor you cannot replace is a dependency, not a supplier. Ownership is the product; who operates the system afterwards is a choice.
Thomas
Founder · Forward-Deployed Engineer
Here is a story you have probably lived. A vendor ships an AI system. It demos beautifully. Six months later it is still running, but only because the vendor is still running it. Every retrain goes through them. Every config change goes through them. When the data drifts, you wait for their ticket queue. The system works, for as long as you keep paying the people who built it.
The problem in that story is not that the vendor is still there. It is that you could not make them leave. What you signed for was a dependency, and in machine learning that is the default outcome.
The model is the small part
The intuition that gets clients into trouble is that the model is the product. It is not. In a real-world ML system, only a small fraction of the code is the machine-learning code itself; the surrounding infrastructure (data collection, feature extraction, configuration, serving, monitoring) is vast and complex.1 Google’s own MLOps guidance adopts the same point and states the operational thesis bluntly: the real challenge is not building the model but building an integrated system and continuously operating it in production.2
This matters because the maintenance burden sits in everything around the model. The same research frames ML systems through the lens of technical debt and finds it is common to incur massive ongoing maintenance costs in real-world deployments, driven in part by the outside world’s background rate of change.1 The data keeps moving after launch, so the system that was correct on day one becomes wrong without anyone changing a line of code.
The model can be handed over in an afternoon. Everything that keeps it correct has to be transferred on purpose.
The last mile is where ownership is won or lost
The gap between a model that works in a notebook and a system that works in production is the last mile of ML, and it is where most engagements quietly fail to transfer ownership. Google describes the lowest maturity level of ML operations as one where the model-building and operations teams are disconnected, which is the source of handoff failures, training-serving skew, and silent model decay.2
That disconnection is exactly what a badly built vendor relationship institutionalises. When the people who built the system and the people who must run it are different organisations, the seam between them hardens. Every operational decision has to cross a contract boundary. The client’s team never learns the system, and a year in they could not take it back even if they wanted to.
The cost of not owning it
The cost is rarely on the invoice. It shows up later, in three forms.
- 01
Lock-in and switching cost. A system wired into proprietary services and formats becomes expensive and difficult to move. NIST warns that non-interoperable applications create the danger of isolated islands of cloud solutions and vendor lock-in, and treats workload and data portability as the standards-based remedy.3 Ownership has a plain test: can the workload and the data move somewhere else? As long as the answer is no, whoever hosts them sets the terms.
- 02
Bus factor. Operational knowledge concentrates dangerously fast. The truck factor is the minimal number of developers who must leave before a project is incapacitated, and an empirical study of 133 popular systems found 65% have a truck factor of two or lower.4 When one party is the only one who understands the pipeline, your bus factor is effectively zero on your own side of the table.
- 03
Compounding dependency. Every change you cannot make yourself is a paid request in someone else’s queue. Retrains, config changes, drift fixes: the older the system gets, the more of these there are, and each one widens the gap between what the vendor knows about your system and what your team does.
None of this is exotic. Each risk follows predictably from treating the model as the deliverable and the system as the vendor’s private property.
Ownership is the product. Operation is a choice.
So we separate two things the industry usually sells as one. Ownership is not negotiable: the code, the schema, the documentation, and the governance posture are yours from the first commit, configured in your accounts and under your control. Who operates the system after it ships is a different question, and it belongs to the client rather than to us.
Some clients have the team and want the keys, so we hand over and step back. Some would rather we kept the pager, so we keep running it. Some want to take the system to their own market with us and share what it earns. The arrangement follows the situation. What does not vary is that they could change their mind.
A vendor you can replace is a supplier. A vendor you cannot replace is a tax.
Concretely, the same four things transfer deliberately throughout the engagement, whether or not we stay to run it. They are what make the choice real rather than a line in a contract.
- 01
Documentation that matches reality. Architecture decisions, data contracts, and runbooks written as the system is built, so they describe what is actually deployed rather than what someone remembers. A wiki assembled the week before a transition fails that test.
- 02
A governance posture you own. Onshore residency, zero-retention, and row-level isolation are configured in your accounts, under your control, and none of it is bolted onto ours. The compliance story survives our departure intact because it never depended on our presence.
- 03
Operational knowledge, transferred by doing. Your engineers run the retrain, trip the alert, and fix the drift while we are still in the room. Nobody learns a runbook by watching someone else follow it. This happens even when we are the ones holding the pager, because a bus factor of zero on your side of the table is unacceptable in either arrangement.4
- 04
Portability by construction. Standard formats, movable workloads, no hidden umbilical to our infrastructure. The portability that NIST prescribes against lock-in is built in from the start,3 so there is nothing to retrofit under duress later.
This is the Immerse, Spot, Execute loop run to completion: we immerse until we know the system you actually run, we spot the piece that belongs in production, and the execute step is not finished until your team could carry it alone. Whether they have to is a separate decision.
The test is whether you could leave
Most AI never reaches production. The part that does is mostly the surrounding system; the model is the small fraction of the code.1 A system counts as delivered when someone other than its builder could keep it alive. Note the tense. The test is not whether they do, it is whether they could.
That is the standard we build to, and it is why we are comfortable staying on when a client asks us to. A long tail of billable support is a failure when the client is stuck with it and an ordinary commercial arrangement when they are not. The difference is not the invoice. It is whether the documentation, the knowledge, and the portability are real enough that leaving would take a fortnight rather than a rebuild.
References
- 01D. Sculley, Gary Holt, Daniel Golovin, et al., “Hidden Technical Debt in Machine Learning Systems,” Advances in Neural Information Processing Systems 28 (NeurIPS/NIPS 2015). Link ↗ ↩
- 02Google Cloud Architecture Center, “MLOps: Continuous delivery and automation pipelines in machine learning.” Link ↗ ↩
- 03NIST Special Publication 500-291, Version 2 — NIST Cloud Computing Standards Roadmap, National Institute of Standards and Technology. Link ↗ ↩
- 04G. Avelino, L. Passos, A. Hora, M. T. Valente, “A Novel Approach for Estimating Truck Factors,” 24th IEEE International Conference on Program Comprehension (ICPC 2016). Link ↗ ↩