A deployment platform can make releases look faster without making the business safer or the engineering organization more effective. A CTO choosing between building one and buying one should therefore reject demonstrations built around pipeline speed alone. My position: buy the deployment control plane unless an in-house pilot produces a measurable advantage in incident-adjusted delivery and total cost. The measurement system must be independent of either option.
Pipeline speed is the wrong purchasing metric
A pipeline timer starts too late and stops too early: it misses changes waiting for a release and defects discovered after production. The case for faster releases is persuasive (compare DevOps Deployment Best Practices for Faster Releases), but a shorter pipeline is not a win if engineers spend the saved time recovering from releases. Measure a service’s merge-to-production time from the Git merge timestamp to the first production deployment containing that commit. Keep the underlying timestamps as well as the summary, because a median can improve while the longest waits become intolerable.
Pair that measure with deployment frequency, change failure rate, and recovery time, using written definitions before the trial begins. Count a production deployment when the new version first receives real traffic, rather than when GitHub Actions or GitLab CI reports a successful job; that distinction matters when approval gates or progressive delivery delay exposure. Define a failed change as a deployment that causes a rollback, hotfix, or breach of a service-level objective within a fixed observation window. An adjustable 24-hour window is a reasonable pilot definition, but the CTO should lengthen it for products whose defects appear only after slower customer activity.
Use a service-level objective alongside these delivery measures. For example, define availability with successful requests divided by eligible requests, and inspect the error budget consumed after each deployment in Prometheus or Datadog. That makes the denominator explicit: a rollback after a handful of canary requests is different from a defect that reaches the whole customer base. Record exposure as request count or affected users where possible, because “one failed deployment” otherwise hides a large difference in harm.
I would not rank tools by deployments per developer, because splitting one change into many deploys improves that number without delivering additional customer value. Nor would I let a vendor’s dashboard be the sole source of success data: a product cannot be the independent judge of events it chooses to record. GitHub’s Deployments API, Argo CD application history, incident records in PagerDuty, and request metrics should remain available for reconciliation whether the final choice is built or bought.
A measurement contract must exist before either platform is piloted
Define one deployment event with a service identifier, immutable artifact digest, production environment, commit range, start time, traffic-exposure time, and outcome. Record timestamps in UTC and preserve the original event ID, because retries can otherwise masquerade as additional releases. OCI image digests provide a stronger artifact identifier than a mutable tag such as latest; the digest lets an investigator connect a customer-facing version to the build that produced it. CloudEvents 1.0 can provide an event envelope, but agreeing on the fields matters more than adopting the format.
Instrument the application separately from the pipeline. OpenTelemetry traces can attach a deployment version to affected requests, while Prometheus counters can show whether errors rose after exposure. Grafana can display both, but the underlying events should stay queryable outside its dashboard. Join deployments to incidents by service, exposure interval, and evidence of impact; do not infer causation merely because an alert followed a release. PagerDuty acknowledgments and Jira incident tickets can establish investigation and resolution times, provided the team agrees which timestamp ends recovery.
The following self-contained Python 3 check uses synthetic records to demonstrate the merge-to-production calculation. Its printed 2-hour median is an example, not a claim about any team; replace the embedded CSV with exported event rows before using it for a decision.
python3 - <<'PY'
import csv, io, statistics
from datetime import datetime
rows = list(csv.DictReader(io.StringIO("""service,merged_at,prod_at
api,2026-09-01T09:00:00+00:00,2026-09-01T10:00:00+00:00
api,2026-09-02T09:00:00+00:00,2026-09-02T12:00:00+00:00
""")))
hours = [(datetime.fromisoformat(r["prod_at"]) -
datetime.fromisoformat(r["merged_at"])).total_seconds() / 3600
for r in rows]
print("median merge-to-prod hours:", statistics.median(hours))
PY
For the real scorecard, retain every row and report the median together with the 90th percentile, deployment count, and the number of records missing a production timestamp. A missing timestamp is not an infinitely slow release or a successful one; it is a data-quality failure that can bias either platform’s result. Audit a sample of artifact digests against production before accepting the dashboard, because a clean chart built from incomplete deployment events offers false precision.
A controlled trial beats a before-and-after success story
A platform introduced alongside staffing changes, a quieter release calendar, or a new test suite cannot claim all subsequent improvement. Select comparable services, capture a baseline, then move some to the candidate platform while leaving others on the current path. A tunable 12-week evaluation window gives the CTO time to observe routine releases rather than a launch demonstration; extend it if deployments are too infrequent to support a useful comparison. Pre-register the primary outcome and incident definition so the team cannot select whichever graph looks best afterward.
Compare changes over time within each group, then compare those changes between groups. If both groups improve, the platform may deserve little credit because a shared change could explain the gain. Segment results by service and change size, since moving only simple services first makes the candidate appear safer than it is. Report how many releases contributed to each rate and show uncertainty rather than declaring victory from a small difference in percentages.
Release strategies also need an outcome test. Traffic shaping, automated checks, and rollback are useful only when they reduce customer exposure or engineer recovery effort enough to justify their overhead. That is a stricter test than counting strategy adoption (contrast DevOps Deployment Strategies for Faster Releases). For each failed release, record the requests exposed before mitigation, the time from first bad signal to action, and the engineer-hours spent investigating. A rollback button that is fast in a demo but rarely usable in an actual incident has little operational value.
Set a decision threshold in advance. One illustrative policy would require a 20% reduction in median merge-to-production time with no deterioration in error-budget consumption or incident-adjusted engineer-hours; that percentage is a management target to tune, not a published benchmark. Add a stopping rule for severe customer harm, because an experiment should not continue merely to finish its calendar window. Interview the engineers who did the releases, but use their accounts to explain measured results rather than substitute for them.
Buying wins unless building clears a higher economic bar
Compare a Harness CD purchase with an in-house GitHub Actions and Argo CD control plane against the same event contract. Harness wins when its controls work with little custom integration and its recurring fee is lower than the engineering capacity it frees; it costs subscription fees, migration work, and some dependence on vendor behavior. The in-house option wins when unusual deployment requirements make product workarounds expensive and the organization can maintain its own control plane; it costs engineering time for authorization, auditability, upgrades, support, and on-call ownership. Neither option wins because its interface looks simpler in a demonstration.
Make the costs comparable with a documented model. Suppose a hypothetical purchase quote is $180,000 annually and integration consumes the equivalent of half an engineer at an assumed fully loaded cost of $200,000 per engineer-year: the modeled first-year operating cost is $280,000, before migration. If building needs two engineers at that same assumed rate, its modeled staffing cost is $400,000 before hosting and support. Those figures are planning inputs, not vendor-published prices or measured savings; replace them with a written quote and tracked internal time before approving spend.
Include the work that disappears from feature delivery. Measure engineer-hours spent maintaining deployment code, responding to platform faults, and explaining audit evidence, because those hours are the principal opportunity cost of an in-house system. Include the purchase’s migration and exit costs too: exporting deployment history and preserving an independent event stream make a later switch less disruptive. Price both options over the same horizon and apply the trial’s observed reliability and delivery outcomes, rather than assigning a speculative dollar value to every minute of lead time.
I would not build merely to avoid a license line item, because an internal platform still incurs maintenance and support costs even when they are spread across payroll budgets. I would build if the trial showed a persistent, material advantage under the pre-registered measures and the staffing plan named who would own it after its initial developers moved on. That is a contestable standard, but it prevents an attractive prototype from being mistaken for a durable operating model.
The first decision is which production event counts
Ask one service owner and one incident responder to agree on the exact event that marks customer exposure, then extract a recent deployment and verify its artifact digest against production. If they cannot reconcile that event with Git history and request metrics, commission the measurement fix before a vendor trial or an internal build. A purchasing decision made without that shared record will be decided by demonstrations, not outcomes.


