The 78-Disease Library

Tier 2 of the three-tier diagnostic model — beneath the 9 CEO-Facing Symptoms and above the 451 root causes. Seventy-eight named technology diseases, organised into nine disease families. Each entry carries its definition, the typical signs the doctor looks for, the symptoms it causes, and the treatment direction.

The Three-Tier Diagnostic Model

Symptoms Are What CEOs See. Diseases Are What's Underneath.

The 9 CEO-Facing Symptoms are visible at the surface (slow growth, rising cost, weak AI ROI). The 78 Diseases below are the named conditions that produce them — and beneath each disease sit its specific root causes (451 in total). The relationship is many-to-many: one symptom usually has several diseases beneath it, and one disease can cause more than one symptom. Treatment addresses the root causes, not the symptoms.

9
CEO-Facing Symptoms (Tier 1)
78
Named Diseases (Tier 2)
451
Root Causes (Tier 3)
9
Disease Families
Nine Disease Families

The Map of Where Disease Lives

Click into any family for its full disease list below. Each family groups diseases by where in the technology operating system they appear — growth, quality, architecture, delivery, cost, operating model, leadership, observability, AI & data.

Family 01

Growth & Product Diseases

Funnel leakage, slow experimentation, weak personalisation — diseases that limit growth at the customer-acquisition and product-execution layer.

8 diseases in this family →

Family 02

Quality & Reliability Diseases

Manual QA, defect leakage, brittle APIs — diseases that erode product quality, release stability and customer trust.

7 diseases in this family →

Family 03

Architecture Diseases

Tight coupling, missing domain boundaries, fragile integrations — diseases inside the technology architecture that limit scale and resilience.

10 diseases in this family →

Family 04

Delivery & DevOps Diseases

Weak CI/CD, manual deployment, no delivery metrics — diseases in the engineering delivery pipeline that slow throughput.

6 diseases in this family →

Family 05

Cost & FinOps Diseases

Cloud waste, vendor sprawl, no cost ownership — diseases that drive technology cost faster than business value.

9 diseases in this family →

Family 06

Operating Model Diseases

No shared operating model, weak product discovery, unclear ownership — diseases at the operating model layer that block execution.

8 diseases in this family →

Family 07

Leadership & Capability Diseases

No diagnostic framework, reactive decisions, key-person dependency — diseases at the leadership and technology-capability layer.

10 diseases in this family →

Family 08

Observability Diseases

Weak instrumentation, tool sprawl, no signal ownership — diseases in the observability layer that hide problems until customers find them.

10 diseases in this family →

Family 09

AI & Data Diseases

No AI strategy, poor data quality, hype-driven use cases — diseases in the AI and data layer that block measurable AI ROI.

10 diseases in this family →

The Full Catalog

78 Diseases, 9 Families

Jump to any family or scroll through the full catalog. Click any disease name to expand its definition, signs, symptom mapping, and treatment direction.

Family 01

Growth & Product Diseases

8 diseases in this family · click any to expand

Funnel Leakage

Funnel Leakage is the condition where customers drop out of the acquisition, activation, conversion, retention, or expansion funnel and the company cannot reliably see where, why, or how much. The visible result is weak conversion, weak retention, and growth spend that is allocated on guesswork because the company is making decisions in the dark.

  • Funnel stages exist in the product but are not consistently instrumented end-to-end.
  • Marketing, product, and growth teams each report different funnel numbers; nobody agrees on a single truth.
  • The company cannot show a clean conversion rate by segment, channel, cohort, or campaign.
  • Drop-off points in onboarding, checkout, payment, configuration, or search cannot be ranked by business impact.
  • A/B tests cannot be reliably evaluated because baseline funnel metrics are unstable.

Primarily causes Symptom 1 (Slow Growth). Often contributes to Symptom 9 (AI and Automation Not Producing Business Value) because data-driven growth use cases (personalisation, recommendations, predictive churn) cannot be built on a leaking, unmeasured funnel.

  1. No end-to-end funnel instrumentation — funnel stages exist in the product but are not consistently instrumented from acquisition to expansion.
  2. No single source of funnel truth — marketing, product, and growth each report different numbers; no owned definition exists.
  3. Inconsistent event schemas and identifiers — events and customer IDs differ across stages, so the funnel cannot be reconstructed.
  4. No segment/channel/cohort decomposition — conversion cannot be sliced by segment, channel, cohort, or campaign.
  5. Unranked drop-off points — onboarding, checkout, payment, and search drop-offs cannot be ordered by business impact.
  6. Unstable baseline metrics — A/B tests cannot be trusted because the underlying funnel baseline drifts.

Establish a single, owned definition of the funnel from acquisition through activation, conversion, retention, and expansion. Instrument every stage with consistent event schemas and shared identifiers. Define one source of truth for funnel metrics across marketing, product, and growth. Make stage-level drop-off rates visible at the executive level. Only then invest in optimisation, personalisation, or AI on top of the funnel — never before.

Slow Experimentation

Slow Experimentation is the condition where the company cannot run, measure, and ship growth experiments fast enough to learn what works. Hypotheses queue up, results take weeks, and decisions fall back on opinion because data-led iteration is too slow to keep up with the business.

  • A/B testing requires engineering effort for every variant; product and growth cannot self-serve.
  • Experiment results take days or weeks because the data pipeline is slow or untrusted.
  • Most "experiments" launch without a hypothesis, success metric, or stop condition.
  • Failed experiments are not documented; the same idea gets re-tested by different teams.

Primarily causes Symptom 1 (Slow Growth). Contributes to Symptom 4 (Delayed Product Releases) because every change becomes a release rather than a test.

  1. No self-serve experimentation infrastructure — every variant needs engineering; product and growth cannot run tests themselves.
  2. Slow or untrusted data pipeline — results take days or weeks, so the loop is too slow to inform decisions.
  3. No hypothesis discipline — experiments launch without a hypothesis, success metric, or stop condition.
  4. No experiment knowledge base — failed experiments go undocumented and the same idea is re-tested by different teams.
  5. No feature-flag / targeting capability — there is no flagging or audience-targeting layer to safely vary experiences.

Build self-serve experimentation infrastructure (feature flags, audience targeting, automated metric pipelines). Establish a hypothesis template with a metric, sample size, and stop condition. Increase experimentation throughput before engineering throughput — the company must learn before it builds.

Performance Drag

Performance Drag is the condition where slow page loads, slow API responses, and unreliable journeys quietly cost conversions at every step of the funnel. Each small slowdown looks minor, but the cumulative effect across millions of sessions is meaningful revenue leakage the company cannot easily attribute to technology.

  • p95 and p99 latencies of customer-facing journeys exceed segment benchmarks.
  • Performance regressions occur after releases but are not caught before deployment.
  • Frontend performance (Core Web Vitals) is not measured or owned.
  • Conversion drops on slower devices, networks, or geographies but no one investigates.

Primarily causes Symptom 1 (Slow Growth) and Symptom 2 (Poor Customer Experience and Quality Issues). Contributes to Symptom 5 (Rising Cost) when teams compensate with more infrastructure.

  1. No journey-level performance measurement — performance is watched at the endpoint, not across the customer journey.
  2. No release performance gating — regressions ship because performance is not checked before deployment.
  3. Unowned, unmeasured frontend performance — Core Web Vitals are neither measured nor owned by anyone.
  4. No performance budgets — there are no per-journey budgets to block changes that exceed them.
  5. Performance treated as background task — performance is an engineering afterthought, not a product feature with an owner.
  6. No segment-aware investigation — conversion drops on slow devices, networks, or geographies go uninvestigated.

Measure performance at the journey level, not just the endpoint level. Set performance budgets per journey and block releases that exceed them. Treat performance as a product feature with an owner, not an engineering background task. Fix worst-impact paths first using real user data.

Onboarding Friction

Onboarding Friction is the condition where customers sign up but never reach the point of value. They get stuck in setup, configuration, identity verification, integration, or first-use friction. Acquisition spend is wasted because the activation moment never happens.

  • Time-to-first-value is unmeasured or worse than the company assumes.
  • Onboarding has more than 3–4 mandatory steps before value is delivered.
  • Drop-off concentrates at specific steps (identity, payment, integration) and those steps are unchanged for months.
  • Support tickets cluster around the same onboarding blockers.

Primarily causes Symptom 1 (Slow Growth). Strong contributor to Symptom 2 (Poor Customer Experience and Quality Issues).

  1. Time-to-first-value unmeasured — TTFV is unknown or worse than assumed, so the activation gap is invisible.
  2. No defined activation event — there is no per-segment definition of the moment value is delivered.
  3. Too many mandatory pre-value steps — onboarding forces more than 3–4 mandatory steps before value.
  4. Blockers in the critical path — identity, payment, and integration sit on the critical path instead of being deferred.
  5. Stale, unfixed drop-off steps — high-drop steps remain unchanged for months despite known friction.
  6. Onboarding not owned as a product — no team owns onboarding, so clustered support tickets never get resolved.

Define the activation event for each customer segment. Measure time-to-first-value end to end. Cut every mandatory step that does not directly unlock value. Move identity, payment, and configuration out of the critical path where possible. Treat onboarding as a dedicated product.

Missing Personalization

Missing Personalization is the condition where every customer receives the same experience regardless of segment, behaviour, intent, or context. High-intent segments are under-converted because the platform cannot adapt to the customer in front of it.

  • The product cannot vary content, offers, or recommendations by segment without an engineering project.
  • Customer segments exist in marketing but are not exposed to the product runtime.
  • No identity stitching across channels — the same customer appears as different people on web, mobile, and email.
  • No clear owner for personalisation as a product capability.

Primarily causes Symptom 1 (Slow Growth). Contributes to Symptom 9 (AI and Automation Not Producing Business Value) since personalisation is one of the most valuable AI use cases.

  1. No runtime personalisation capability — content, offers, and recommendations cannot vary by segment without an engineering project.
  2. Segments trapped in marketing — segments exist in marketing tools but are not exposed to the product runtime.
  3. No customer identity layer — no identity stitching, so the same customer looks like different people across channels.
  4. No owner for personalisation — personalisation has no owner as a product capability.
  5. No segment-level measurement — lift is measured in aggregate, hiding per-segment performance.

Establish customer identity as a first-class platform capability. Expose segment, behaviour, and intent signals to the product runtime, not just to dashboards. Start with rule-based personalisation on highest-impact surfaces, then move to model-based. Measure lift per segment, not aggregate.

Pricing Mismatch

Pricing Mismatch is the condition where the pricing and packaging model does not match the value the customer perceives. Conversion is depressed because the offer is hard to evaluate or too expensive at the decision point; expansion is depressed because there is no clean upgrade path.

  • Win-loss analysis cites price as a top blocker but pricing has not been revisited in 12+ months.
  • Pricing tiers differ by features rather than by differentiated outcomes.
  • High discount rates suggest list pricing is misaligned with willingness to pay.
  • Expansion revenue is weak because upgrade paths are unclear or require sales touch.

Primarily causes Symptom 1 (Slow Growth). Indirect contributor to Symptom 5 (Rising Cost) when low-value segments are subsidised.

  1. Stale pricing model — price is cited as a top blocker in win-loss but has not been revisited in 12+ months.
  2. Feature-based, not outcome-based tiers — tiers differ by feature counts rather than differentiated outcomes.
  3. Misaligned willingness-to-pay — high discount rates show list pricing is out of step with willingness to pay.
  4. No clean upgrade path — expansion is weak because upgrade paths are unclear or require sales touch.
  5. Wrong segmentation basis — customers are segmented by company size rather than by outcome and willingness to pay.
  6. Pricing outside the experimentation system — pricing is never tested, so it is set by opinion not evidence.

Re-segment customers by outcome and willingness to pay, not by company size alone. Redesign tiers around value milestones, not feature counts. Build pricing into the experimentation system. Make upgrade and downgrade paths self-serve where possible.

Disconnected Roadmap

Disconnected Roadmap is the condition where the product roadmap is not measurably tied to revenue, retention, conversion, expansion, or cost outcomes. Engineering builds what is asked for, not what moves the business. The roadmap becomes a feature factory output rather than a business instrument.

  • Roadmap items are described by feature, not by business outcome.
  • No post-launch measurement of business impact for shipped features.
  • The same items recur quarter after quarter because impact is unclear.
  • Sales, marketing, and support each lobby for their own list; the roadmap is the union of asks.

Primarily causes Symptom 1 (Slow Growth) and Symptom 6 (Business, Product and Technology Misalignment). Contributes to Symptom 4 (Delayed Product Releases).

  1. No outcome linkage on roadmap items — items are described by feature, not by a measurable business outcome.
  2. No post-launch impact measurement — shipped features are never measured for business impact.
  3. No priority cap / trade-off discipline — concurrent priorities are uncapped, so trade-offs are never forced.
  4. Roadmap as union of asks — sales, marketing, and support each lobby their list and the roadmap is their sum.
  5. Recurring items from unclear impact — the same items recur quarter after quarter because impact is never proven.
  6. Roadmap not trusted as a forecast — the CEO/CFO cannot use the roadmap to forecast outcomes.

Tie every roadmap item to a measurable business outcome with a target and a measurement window. Cap concurrent priorities to force trade-offs. Run post-launch reviews that measure business impact, not feature completion. Make the roadmap a tool the CEO and CFO trust to forecast outcomes.

Growth Silos

Growth Silos is the condition where Growth, Product, Marketing, and Engineering each optimise their own metrics, and no single function owns the end-to-end growth system. Each team can show local progress, but company-level conversion does not move because the system is not managed as one.

  • Growth, product, marketing, and engineering report to different leaders with no shared growth KPI.
  • Each function has its own dashboard with its own definitions of activation, conversion, retention, and churn.
  • Hand-offs between marketing → product → support → success have visible drop-offs that no one owns.
  • Quarterly reviews show local improvements but flat or declining company-level conversion.

Primarily causes Symptom 1 (Slow Growth). Strong contributor to Symptom 6 (Business, Product and Technology Misalignment).

  1. No shared growth KPI — functions report to different leaders with no common growth KPI.
  2. Conflicting local metric definitions — each function defines activation, conversion, retention, and churn differently.
  3. No end-to-end conversion owner — no single accountable owner for end-to-end conversion.
  4. Unowned cross-function hand-offs — drop-offs at marketing→product→support→success hand-offs have no owner.
  5. Misaligned incentives — a team can win on its local metric while the overall system loses.
  6. No journey-based growth operating rhythm — growth is not managed by journey (Acquisition, Activation, Retention, Expansion).

Establish a shared growth KPI hierarchy from company outcome down to function inputs. Appoint one accountable owner for end-to-end conversion. Move to journey-based growth meetings (Acquisition, Activation, Retention, Expansion). Align incentives so a team cannot win on its local metric while the system loses.

Family 02

Quality & Reliability Diseases

7 diseases in this family · click any to expand

Manual QA

Manual QA is the condition where quality depends primarily on human checks at the end of the development cycle. Releases slow because manual testing is the bottleneck, defects leak because manual coverage is incomplete, and release confidence is fragile because the company is always one missed case away from a customer-facing failure.

  • Test automation coverage is low or concentrated on trivial paths; high-risk paths are manual.
  • QA is a separate function rather than an engineering discipline.
  • Releases are gated by manual sign-off that takes days, not minutes.
  • Regression testing for every release is incomplete because the test matrix is too large for manual coverage.

Primarily causes Symptom 2 (Poor Customer Experience and Quality Issues). Strong contributor to Symptom 4 (Delayed Product Releases).

  1. Low / mis-targeted test automation — coverage is low or concentrated on trivial paths while high-risk paths stay manual.
  2. QA siloed from engineering — QA is a separate function, not an engineering discipline owned by builders.
  3. Manual sign-off gating — releases are gated by manual sign-off that takes days, not minutes.
  4. Test matrix too large for humans — regression coverage is incomplete because the matrix cannot be covered manually.
  5. Quality not shifted left — quality is checked at the end, not built into design, code review, and tests.
  6. No test pyramid — there is no unit-heavy test structure, so fast automated signals are missing.

Shift quality left — into design, code review, and automated tests. Establish a test pyramid with most tests at the unit level. Treat test automation as production code with its own quality standards. Phase out manual gating in favour of automated quality signals.

Defect Leakage

Defect Leakage is the condition where teams fix the visible defects but do not prevent the underlying classes of defects from returning. The same kind of bug appears in different parts of the system, and the company is stuck in a cycle of repair rather than prevention.

  • Postmortems list root cause as "human error" rather than systemic gap.
  • The same defect category (null pointer, race condition, validation gap, state corruption) appears repeatedly.
  • No standing engineering practice to harden the system against entire defect classes.
  • Defect taxonomy and trend analysis are absent; teams discuss individual defects, not patterns.

Primarily causes Symptom 2 (Poor Customer Experience and Quality Issues). Contributes to Symptom 4 (Delayed Product Releases) through firefighting load.

  1. Blame-the-human postmortems — root cause is logged as "human error" instead of a systemic gap.
  2. No defect taxonomy / trend analysis — defects are discussed individually, never as patterns.
  3. No class-level hardening practice — there is no standing practice to harden against entire defect classes.
  4. Missing structural safeguards — types, contracts, validations, state machines, and invariants are absent where they would prevent classes.
  5. Instance fixes over system fixes — the team fixes the instance, not the system that produced it.

Build a defect taxonomy and track repeat classes. Treat every recurring class as a systemic gap (types, contracts, validations, state machines, invariants) and fix the system, not the instance. Hold blameless postmortems that ask "what would prevent this class of defect entirely?"

Unstable Releases

Unstable Releases is the condition where releases ship without consistent quality, performance, security, and journey checks. Bad changes reach customers because the release process trusts goodwill rather than evidence. Every release feels like a gamble.

  • Release process is documented loosely or differs by team.
  • Pre-release checks are inconsistent across services; some teams run full validation, others skip steps.
  • Rollback path is unclear or untested.
  • Post-release validation in production is missing; teams discover failures from customer reports.

Primarily causes Symptom 2 (Poor Customer Experience and Quality Issues) and Symptom 4 (Delayed Product Releases).

  1. No standardised release pipeline — the release process is loosely documented and differs by team.
  2. Inconsistent pre-release checks — some teams run full validation, others skip steps.
  3. Untested rollback path — the rollback path is unclear or never exercised.
  4. No post-release validation — failures are discovered from customer reports, not automated checks.
  5. Release safety left to team discipline — safety depends on individual goodwill rather than a platform property.
  6. No mandatory quality/perf/security/journey gates — required checks are not enforced in the pipeline.

Standardise the release pipeline across services with mandatory quality, performance, security, and journey checks. Make rollback a first-class capability tested regularly. Run automated post-release validation against key journeys before declaring success. Make release safety a property of the platform, not of individual team discipline.

No Journey Ownership

No Journey Ownership is the condition where customer journeys cross multiple teams but no single team owns the end-to-end experience. Each team optimises its slice; the journey as a whole degrades because no one is accountable for it.

  • Critical journeys (checkout, onboarding, search, payment) cross 3+ teams without a journey owner.
  • When the journey breaks, leadership cannot identify a single accountable person.
  • Each team's roadmap optimises its own slice; the journey is the union of those slices.
  • Cross-team coordination requires escalation rather than a standing forum.

Primarily causes Symptom 2 (Poor Customer Experience and Quality Issues) and Symptom 6 (Business, Product and Technology Misalignment).

  1. Journeys span teams with no owner — critical journeys cross 3+ teams without a journey owner.
  2. No accountable person on breakage — when the journey breaks, leadership cannot name one accountable owner.
  3. Slice-optimised roadmaps — each team optimises its slice; the journey is only the union of slices.
  4. Escalation instead of standing forum — cross-team coordination relies on escalation, not a standing mechanism.
  5. Service-owned, not journey-owned metrics — metrics belong to services, so no one is measured on the whole journey.
  6. No end-to-end journey SLOs — there are no journey-level SLOs or improvement roadmap to hold an owner to.

Map critical customer journeys and appoint a single accountable journey owner with cross-team authority. Move from service-owned to journey-owned metrics for those journeys. Make the journey owner accountable for end-to-end SLOs, business outcomes, and improvement roadmap.

Missing Telemetry

Missing Telemetry is the condition where the company has no real-user monitoring data for the frontend, mobile, or customer-facing surfaces. Customer friction is invisible until customers complain. The team operates on synthetic data and assumptions rather than on what real users actually experience.

  • No real-user monitoring (RUM) on web or mobile.
  • Frontend errors are not captured, or are captured into a queue nobody reviews.
  • Session replay, error tracking, and performance metrics are missing or fragmented.
  • Customer-reported issues cannot be reproduced because there is no session context.

Primarily causes Symptom 2 (Poor Customer Experience and Quality Issues). Strong contributor to Symptom 8 (Missing Observability).

  1. No real-user monitoring (RUM) — there is no RUM on web or mobile to see actual experience.
  2. Uncaptured / ignored frontend errors — frontend errors are not captured or land in a queue nobody reviews.
  3. Missing session context — no session replay or context, so reported issues cannot be reproduced.
  4. Fragmented frontend signals — error tracking and performance metrics are scattered or absent.
  5. Frontend telemetry not release-blocking — instrumentation is optional, so it is skipped under deadline.
  6. Frontend-to-backend traces unlinked — frontend signals are not tied to backend traces for end-to-end reproduction.

Instrument the frontend with real-user monitoring, structured error capture, and session context. Make frontend telemetry a release-blocking requirement. Tie frontend signals to backend traces so customer-reported issues can be reproduced end to end. Frontend is not someone else's job.

Brittle APIs

Brittle APIs is the condition where API contracts change unpredictably and break customer-facing experiences or integrations. Internal teams cannot rely on a service's interface, and external integrators get burned every release.

  • Breaking API changes ship without versioning, deprecation notices, or consumer migration paths.
  • Consumer-driven contract tests are missing; producers learn they broke consumers in production.
  • API documentation lags the implementation or is autogenerated and out of date.
  • The same fields have different meanings across services or endpoints.

Primarily causes Symptom 2 (Poor Customer Experience and Quality Issues) and Symptom 4 (Delayed Product Releases).

  1. No versioning / deprecation policy — breaking changes ship without versions, deprecation notices, or migration paths.
  2. No consumer-driven contract tests — producers discover they broke consumers only in production.
  3. Stale / autogenerated docs — API documentation lags the implementation and cannot be trusted.
  4. Inconsistent field semantics — the same fields mean different things across services and endpoints.
  5. APIs treated as internal details — APIs are not treated as products with consumers and standards.
  6. No CI contract enforcement — nothing blocks breaking changes before they ship.

Treat APIs as products with versioning, deprecation policy, and consumer contracts. Add consumer-driven contract testing in CI to block breaking changes before they ship. Establish API design standards and review for cross-service consistency. APIs are interfaces between teams, not internal implementation details.

Untreated Tech Debt

Untreated Tech Debt is the condition where technical debt accumulates without active treatment. Every new feature touches debt and slows down. Defect probability rises with every change. The team spends increasing capacity working around debt instead of fixing it.

  • No explicit budget or capacity for debt reduction; debt work happens only when forced.
  • The team can name areas they are "afraid to touch."
  • Estimates for new work include hidden time for working around debt.
  • Code health metrics (complexity, duplication, coverage, age) trend negative quarter over quarter.

Primarily causes Symptom 2 (Poor Customer Experience and Quality Issues) and Symptom 4 (Delayed Product Releases). Contributes to Symptom 5 (Rising Cost) through hidden rework.

  1. No debt budget / capacity — there is no explicit capacity for debt reduction; it happens only when forced.
  2. Fear-to-touch areas — the team can name code they are afraid to modify.
  3. Hidden work-around time in estimates — new-work estimates secretly include time to navigate debt.
  4. Negative code-health trend — complexity, duplication, coverage, and age trend worse quarter over quarter.
  5. Debt invisible in the backlog — debt is not tracked as work with business impact, so it is never prioritised.
  6. No refactor practice — neither continuous refactoring nor planned strategic refactors are in place.

Allocate a fixed percentage of every cycle to debt reduction (typically 15–25%). Make debt visible — track it as a backlog with business impact. Treat debt reduction as feature work, not chores. Refactor as you go for small debt; plan deliberate strategic refactors for structural debt.

Family 03

Architecture Diseases

10 diseases in this family · click any to expand

Tightly Coupled Code

Tightly Coupled Code is the condition where small changes touch many areas of the system. The codebase cannot be modified safely or quickly because modules share state, logic, and assumptions in ways that propagate change across boundaries.

  • A change to one module requires changes in several others to compile or pass tests.
  • Modules have circular dependencies or implicit shared state.
  • New developers take months to understand the change blast radius of any modification.
  • The team is "afraid to touch" certain areas because the impact is unpredictable.

Primarily causes Symptom 3 (Architecture Not Fit) and Symptom 4 (Delayed Product Releases).

  1. Change ripple across modules — one change forces edits in several modules to compile or pass tests.
  2. Circular dependencies / implicit shared state — modules depend on each other in cycles or share hidden state.
  3. Unpredictable blast radius — new developers take months to learn the impact of any change.
  4. Unpredictable change blast-radius — certain areas are avoided because their change impact is unpredictable.
  5. No clear interfaces / seams — there are no extracted interfaces to decouple implementations behind.
  6. No architecture rules in CI — nothing prevents reintroduction of cycles and coupling.

Identify the strongest coupling hotspots through change-correlation analysis. Apply the Strangler Fig pattern to decouple incrementally — extract clear interfaces, then replace implementations behind them. Establish architecture rules in CI that prevent reintroduction of cycles. Decoupling is a continuous discipline, not a one-time refactor.

Missing Domain Boundaries

Missing Domain Boundaries is the condition where services and modules share data and logic without clear ownership. Business concepts (Customer, Order, Payment, Product) have multiple definitions across the system. Scale and change become risky because nobody owns the truth of what a domain means.

  • The same concept (Customer, Order, Product) has different schemas, IDs, or fields in different services.
  • Cross-team data fetches are common because data is in the "wrong" service.
  • No clear owner for core domain concepts; multiple teams modify the same model.
  • Domain models do not map to business language; engineering and business use different terms.

Primarily causes Symptom 3 (Architecture Not Fit) and Symptom 6 (Business, Product and Technology Misalignment).

  1. Duplicated concept definitions — Customer, Order, Product have different schemas, IDs, or fields across services.
  2. Chatty cross-service data fetching — data lives in the wrong service, forcing constant cross-team fetches.
  3. No owner per domain concept — multiple teams modify the same model with no single owner.
  4. Model–business language mismatch — domain models do not map to business language; teams use different nouns.
  5. No bounded contexts — there is no DDD partitioning into bounded contexts with assigned ownership.
  6. No event-driven propagation — there is no event mechanism to propagate domain changes cleanly.

Apply Domain-Driven Design to identify bounded contexts and assign ownership. Establish one source of truth per domain concept with a clear owning service. Move from chatty cross-service data fetching to event-driven propagation. Align domain models to business language so engineering and business operate on the same nouns.

Mixed Workloads

Mixed Workloads is the condition where critical and non-critical workloads share the same resources. Non-critical workloads (batch jobs, analytics queries, internal admin) can degrade or take down critical customer-facing paths because the platform does not isolate them.

  • Customer-facing traffic and batch/analytics traffic share the same database, compute, or queue.
  • Incidents in non-critical paths cascade into customer-facing degradation.
  • No workload classification or priority scheduling at the infrastructure level.
  • Capacity decisions are made for the worst-case workload, inflating cost for the rest.

Primarily causes Symptom 3 (Architecture Not Fit) and Symptom 5 (Rising Cost). Contributes to Symptom 2 (Poor CX) during peaks.

  1. Shared resources across criticalities — customer traffic and batch/analytics share the same DB, compute, or queue.
  2. Cascading non-critical failures — incidents in non-critical paths spill into customer-facing degradation.
  3. No workload classification — there is no classification or priority scheduling at the infrastructure level.
  4. No critical-path isolation — critical workloads lack dedicated resources or quotas.
  5. Worst-case capacity for all — capacity is sized for the worst workload, inflating cost for the rest.

Classify workloads by criticality, latency sensitivity, and consistency needs. Isolate critical workloads with dedicated resources or quotas. Use read replicas, dedicated queues, and tiered compute to separate workload classes. Architecture must protect critical paths from non-critical noise.

Missing Resilience

Missing Resilience is the condition where the system has no built-in patterns for handling failure. One failure cascades into many. There are no timeouts, retries, circuit breakers, bulkheads, or graceful degradation, so the system fails hard rather than failing safely.

  • A dependency failure (third party, database, queue) takes down the entire service rather than degrading it.
  • No circuit breakers; calls to failing dependencies pile up until thread exhaustion.
  • Retries are unbounded or naive, amplifying load on failing dependencies.
  • No idempotency guarantees; retries cause duplicate side effects.
  • Customer-facing features lack a fallback for backend failure.

Primarily causes Symptom 3 (Architecture Not Fit) and Symptom 2 (Poor CX).

  1. No timeouts on calls — calls lack timeouts, so failures hang and propagate.
  2. No circuit breakers — calls to failing dependencies pile up until thread exhaustion.
  3. Naive / unbounded retries — retries are unbounded and amplify load on failing dependencies.
  4. No idempotency — retries cause duplicate side effects because operations are not idempotent.
  5. No graceful degradation / fallbacks — customer features have no fallback when a backend fails.
  6. No bulkheads / failure isolation — there is nothing to contain failure to one domain.
  7. No resilience testing — failure modes are never exercised through chaos engineering.

Establish resilience patterns as platform defaults: timeouts on every call, bounded retries with backoff, circuit breakers on shared dependencies, bulkheads to isolate failure domains, idempotency on side-effecting operations, graceful degradation on customer-facing features. Test resilience with chaos engineering — failure is a property to be designed for, not avoided.

Database Bottleneck

Database Bottleneck is the condition where the database design does not match the business growth pattern. Queries, locks, or contention become the throttle on the entire system. Scaling the application does not help because the database is the limit.

  • The database is the hot spot in every performance investigation.
  • Lock contention or long-running transactions block other queries.
  • Indexes are missing on frequent queries, or the wrong indexes exist.
  • Schema design forces application-side joins that the database could do better, or vice versa.
  • Read and write workloads share the same database without separation.

Primarily causes Symptom 3 (Architecture Not Fit) and Symptom 5 (Rising Cost).

  1. Persistent DB hot spot — the database is the hot spot in every performance investigation.
  2. Lock contention / long transactions — long-running transactions and locks block other queries.
  3. Missing or wrong indexes — frequent queries lack indexes, or the wrong indexes exist.
  4. Schema mismatched to workload — schema forces inefficient application-side joins or vice versa.
  5. Read/write workloads not separated — transactional reads and writes share one database without separation.
  6. OLTP and OLAP collision — analytical workloads run on the transactional store.

Profile and fix the worst queries first using real production load data. Add and maintain indexes deliberately, not reactively. Separate read and write workloads. Move analytical workloads off the transactional database. Where structural, redesign the schema or move to a database technology that matches the workload (OLTP vs OLAP, document vs relational, search vs key-value).

No Caching

No Caching is the condition where repeated reads and computations hit the database or compute layer instead of a cache. The system pays full cost for work it has already done, inflating both cost and latency.

  • Hot read paths hit the database every request.
  • No caching layer (in-memory, distributed, CDN) at obvious natural boundaries.
  • Cache invalidation strategy is ad hoc or absent; data freshness vs cost is not consciously designed.
  • CDN is used only for static assets, not for cacheable API responses.

Primarily causes Symptom 3 (Architecture Not Fit) and Symptom 5 (Rising Cost).

  1. Hot reads hit the DB every request — frequently read paths go to the database on every request.
  2. No caching layer — there is no in-memory, distributed, or CDN cache at natural boundaries.
  3. Ad hoc / absent invalidation strategy — freshness vs cost is not consciously designed.
  4. CDN under-used — the CDN serves only static assets, not cacheable API responses.
  5. No caching at the right layer — caching is not placed deliberately by access pattern (CDN/edge/application).

Identify hot read paths and high-cost computations. Cache at the right layer (CDN for public, edge for personalised-but-cacheable, application for hot internal reads). Design cache invalidation explicitly — choose between TTL, event-driven, or read-through. Caching is a design decision, not a performance band-aid.

Brittle Integrations

Brittle Integrations is the condition where internal and third-party service contracts are unstable. Integration changes break production. Every new partner becomes a custom project because there is no integration discipline.

  • Integrations are point-to-point and tightly coupled to specific partners.
  • No retry, idempotency, or backpressure for third-party calls.
  • Partner outages take the system down rather than degrading gracefully.
  • Each new partner takes weeks of engineering effort to onboard.
  • No anti-corruption layer to insulate the domain from external schemas.

Primarily causes Symptom 3 (Architecture Not Fit) and Symptom 4 (Delayed Product Releases).

  1. Point-to-point coupling — integrations are point-to-point and tightly coupled to specific partners.
  2. No resilience on third-party calls — no retry, idempotency, or backpressure for external calls.
  3. Partner outages cascade — a partner outage takes the system down instead of degrading.
  4. No anti-corruption layer — external schemas leak directly into the domain.
  5. No standardised integration patterns — each new partner takes weeks because there is no template.
  6. Third parties trusted as reliable — external dependencies are not treated as unreliable by default.

Build an integration layer that insulates the domain from external schemas (anti-corruption layer). Treat third-party calls as unreliable by default — timeouts, retries, circuit breakers, queues. Standardise integration patterns (webhooks, polling, event-driven) so each new partner follows the template. Make integration a platform capability, not a custom project.

Weak Identity Layer

Weak Identity Layer is the condition where authentication, authorisation, and access design are inconsistent across the system. Security depends on tools and ad-hoc checks rather than on coherent design. The blast radius of any compromise is unbounded because access is not segmented.

  • Authentication is inconsistent across services; some use modern protocols, some use legacy patterns.
  • Authorisation logic is scattered across services rather than centralised.
  • Service-to-service calls trust the network rather than the identity.
  • No clear role and permission model; access is granted ad hoc.
  • Internal admin tools have weaker auth than customer-facing surfaces.

Primarily causes Symptom 3 (Architecture Not Fit).

  1. Inconsistent authentication — some services use modern protocols, others legacy patterns.
  2. Scattered authorisation logic — authorisation is duplicated across services, not centralised in a policy layer.
  3. Network-trusting service calls — service-to-service calls trust the network instead of verifying identity.
  4. No role / permission model — access is granted ad hoc with no coherent model.
  5. Weak internal-tool auth — admin tools have weaker auth than customer-facing surfaces.
  6. No access reviews — access is never re-verified, so entitlements drift.

Standardise authentication on modern protocols (OAuth, OIDC, SAML). Centralise authorisation in a policy layer that all services consult. Adopt Zero Trust principles — verify every request, regardless of network position. Establish a clear role and permission model with regular access reviews. Identity is the architecture, not an afterthought.

Scattered Secrets

Scattered Secrets is the condition where API keys, tokens, credentials, and secrets are spread across code, configuration files, environment variables, and developer machines. Rotation is painful, leakage is likely, and the company cannot prove who has access to what.

  • Secrets appear in source control (current or git history).
  • Secrets live in environment variables managed manually per environment.
  • Rotation requires coordinated changes across multiple systems.
  • No central secrets manager with audit logging and access control.
  • Developer machines hold long-lived production credentials.

Primarily causes Symptom 3 (Architecture Not Fit).

  1. Secrets in source control — secrets appear in current code or git history.
  2. Manual per-environment secrets — secrets live in environment variables managed by hand.
  3. No central secrets manager — there is no vault with audit logging and access control.
  4. Painful rotation — rotation needs coordinated changes across multiple systems.
  5. Long-lived static credentials — credentials are long-lived rather than short, scoped, identity-issued.
  6. Production creds on dev machines — developer machines hold long-lived production credentials.

Adopt a central secrets manager (Vault, AWS Secrets Manager, equivalent) with audit logging. Move all secrets out of source and environment files. Use short-lived, scoped credentials issued by identity rather than long-lived static keys. Make rotation automated and frequent. Remove production credentials from developer machines.

Wrong Architecture

Wrong Architecture is the condition where architecture decisions were made based on trends, fashion, or under-investment rather than on workload, scale, and resilience needs. The platform may look modern (microservices, Kubernetes, serverless) but does not actually fit what the business needs. Or it may be under-engineered for the scale it now serves.

  • Microservices were adopted before domain boundaries were understood, producing a "distributed monolith."
  • A modern stack was chosen because it was new, not because it matched the workload.
  • Architecture was chosen by an exiting leader and never re-evaluated by current leadership.
  • The team cannot explain *why* a major architectural choice was made — only that it exists.
  • Workload, scale, latency, consistency, and cost requirements were not explicitly considered in the original decision.

Primarily causes Symptom 3 (Architecture Not Fit). Contributes to Symptom 5 (Rising Cost) and Symptom 4 (Delayed Releases).

  1. Premature microservices — microservices were adopted before domain boundaries, producing a distributed monolith.
  2. Fashion-driven stack choice — a modern stack was chosen because it was new, not because it fit the workload.
  3. Never re-evaluated decisions — architecture chosen by an exiting leader is never reassessed by current leadership.
  4. Unexplained architectural choices — the team cannot say *why* a major choice was made, only that it exists.
  5. Requirements never considered — workload, scale, latency, consistency, and cost were not explicit in the decision.
  6. No Architecture Decision Records — trade-offs were never captured, so decisions cannot be revisited.
  7. Under- or over-engineering — the platform is too complex or too thin for the scale it now serves.

Re-evaluate architecture decisions against current workload, scale, resilience, and cost requirements. Capture decisions in Architecture Decision Records with the trade-offs that were considered. Fit-for-purpose beats fashionable — adopt complexity only where it earns its keep against measurable requirements. Be willing to walk back wrong decisions; the cost of correction grows over time.

Family 04

Delivery & DevOps Diseases

6 diseases in this family · click any to expand

Weak CI/CD

Weak CI/CD is the condition where build, test, and deploy pipelines are slow, flaky, or partially manual. Each release takes coordination and luck. The pipeline that should give teams confidence to ship instead becomes a source of fear.

  • Pipeline runs take long enough that developers context-switch while waiting.
  • Pipeline flakiness is normalised; "re-run the failed job" is standard practice.
  • Some steps require manual triggering or approval where automation is feasible.
  • Deployment frequency is weekly or monthly rather than daily or hourly.
  • Different services have different pipelines with different quality.

Primarily causes Symptom 4 (Delayed Product Releases) and Symptom 2 (Poor CX).

  1. Slow pipeline runs — runs are long enough that developers context-switch while waiting.
  2. Normalised flakiness — "re-run the failed job" is standard practice.
  3. Manual pipeline steps — some steps require manual triggering or approval where automation is feasible.
  4. Low deployment frequency — deployments are weekly or monthly, not daily or hourly.
  5. Inconsistent pipelines across services — different services have different pipelines of different quality.
  6. Pipeline not treated as a product — the pipeline has no owner or quality standard of its own.

Treat the pipeline as a product with its own quality standards. Reduce pipeline time aggressively — target under 15 minutes for a full build and test. Quarantine and fix flaky tests rather than tolerating them. Standardise pipelines across services. Make deployment a non-event that happens many times per day.

Manual Deployment

Manual Deployment is the condition where releases depend on manual steps and human coordination. Deployment frequency is low. Releases are stressful events rather than routine operations. The blast radius of any human mistake is large.

  • Deployment requires a runbook with multiple manual steps.
  • Releases happen during off-hours because they are risky.
  • Rollback is manual and slow.
  • A specific person or team is required to deploy; others cannot.
  • Deployment artifacts are built differently per environment.

Primarily causes Symptom 4 (Delayed Product Releases) and Symptom 2 (Poor CX).

  1. Runbook-driven manual deploys — deployment requires a multi-step manual runbook.
  2. Off-hours risky releases — releases happen off-hours because they are risky.
  3. Slow manual rollback — rollback is manual and slow.
  4. Single-person deploy dependency — only a specific person or team can deploy.
  5. Per-environment build differences — artifacts are built differently per environment, so behaviour diverges.
  6. No progressive delivery — there is no canary, blue-green, or flag-based mechanism to bound risk.

Automate the full deployment path end to end. Make builds reproducible and environment-agnostic. Make rollback a single button that anyone can press safely. Adopt progressive delivery (canary, blue-green, feature flags) so deployment risk is bounded. Deployment is a platform capability, not a heroic operation.

Flaky Tests

Flaky Tests is the condition where test automation is unreliable. Tests pass and fail intermittently for reasons unrelated to the code under test. Teams stop trusting the test suite and fall back to manual checks. The investment in automation has been silently negated.

  • Tests fail and then pass on re-run without code changes.
  • "Re-run the failed job" is a normalised practice in PR reviews.
  • Specific tests have known flakiness flags but remain in the suite.
  • New tests are not added because the team has lost faith in the system.
  • CI green status no longer means the code is actually working.

Primarily causes Symptom 4 (Delayed Product Releases) and Symptom 2 (Poor CX).

  1. Non-deterministic tests — tests fail then pass on re-run without code changes.
  2. Normalised "re-run" culture — re-running failed jobs is accepted practice in reviews.
  3. Known-flaky tests left in suite — flaky tests are flagged but never removed or fixed.
  4. Lost faith stops new tests — the team stops adding tests because the suite is not trusted.
  5. Green no longer means working — CI green status no longer guarantees correctness.
  6. Unaddressed flakiness sources — time, ordering, shared state, and external dependencies are not isolated.

Quarantine flaky tests immediately and treat fixing them as priority engineering work. Investigate root cause — most flakiness comes from time, ordering, shared state, or external dependencies. Establish a flakiness budget — when it is exceeded, ship freezes until reduced. A flaky suite is worse than no suite because it teaches the team to ignore failures.

Overloaded Roadmap

Overloaded Roadmap is the condition where there are too many priorities and none are sharp. Teams switch context constantly without finishing. Stakeholders each get a token win on the roadmap, but the company makes no decisive progress on anything.

  • Every team is "fully loaded" with priorities, yet completion rates are low.
  • The roadmap is a list of asks, not a ranked set of decisions.
  • Context switching across initiatives consumes a significant share of capacity.
  • Each quarterly review carries forward most of the previous quarter's items.
  • Leadership cannot name the top 3 priorities without disagreement.

Primarily causes Symptom 4 (Delayed Product Releases) and Symptom 6 (Business, Product and Technology Misalignment).

  1. Uncapped concurrent priorities — teams are "fully loaded" yet completion rates are low.
  2. Roadmap as a list of asks — the roadmap is asks, not a ranked set of yes/no decisions.
  3. High context-switching cost — switching across initiatives consumes significant capacity.
  4. Quarter-over-quarter carryover — most items roll forward each quarter.
  5. No agreed top priorities — leadership cannot name the top 3 without disagreement.
  6. No single intake process — new asks do not compete against existing priorities through one gate.

Cap concurrent priorities per team to force trade-offs. Move from "everything is important" to a ranked list with explicit yes/no decisions. Establish a single intake process where new asks compete against existing priorities. Make the trade-off visible: adding X means delaying Y.

No Delivery Metrics

No Delivery Metrics is the condition where cycle time, lead time, deployment frequency, change failure rate, and business impact are not measured. Improvement is invisible. The company cannot tell if delivery is getting better or worse, or why.

  • DORA metrics (deployment frequency, lead time, change failure rate, time to restore) are not tracked.
  • Delivery discussions rely on opinion and anecdote rather than data.
  • Post-release business impact is not measured.
  • Teams cannot answer "is our delivery improving?" with evidence.

Primarily causes Symptom 4 (Delayed Product Releases) and Symptom 6 (Business, Product and Technology Misalignment).

  1. DORA metrics untracked — deployment frequency, lead time, change-failure rate, and MTTR are not measured.
  2. Opinion-led delivery discussion — delivery is debated on anecdote, not data.
  3. No post-release impact measurement — business impact after release is not measured.
  4. No baseline measurement — there is no baseline to tell whether delivery is improving.
  5. Delivery treated as art, not system — delivery is not managed as a measurable, improvable system.

Establish baseline measurements of cycle time, lead time, deployment frequency, change failure rate, and time to restore. Add business-outcome measurement post-release. Make these metrics visible at the team and executive level. Treat delivery as a system that can be measured and improved, not as an art.

Wrong Release Governance

Wrong Release Governance is the condition where the release process is either too loose (chaos, broken releases) or too bureaucratic (paralysis, weeks-long approval). The governance is rarely right-sized to the actual risk of the change.

  • Every release requires the same approval regardless of risk (paralysis), or no release requires any approval (chaos).
  • Change approval boards meet weekly and become a delivery bottleneck.
  • High-risk changes (database migrations, identity changes) get the same process as low-risk changes (copy updates).
  • Approval is reduced to a rubber stamp because nobody actually reviews the substance.

Primarily causes Symptom 4 (Delayed Product Releases) and Symptom 2 (Poor CX).

  1. Governance not sized to risk — every release gets the same approval, or none does.
  2. Approval-board bottleneck — change boards meet weekly and become a delivery throttle.
  3. No risk differentiation — high-risk changes get the same process as trivial copy edits.
  4. Rubber-stamp approvals — approval is a formality; nobody reviews substance.
  5. No premortem / rollback review for high-risk — risky changes ship without explicit premortem or rollback review.
  6. No progressive delivery to bound risk — there is no canary/flag mechanism to make most changes low-risk by design.

Right-size governance to risk. Low-risk changes ship continuously without approval. High-risk changes (data migrations, identity, payments, contracts) get explicit pre-mortem and rollback plan review. Use progressive delivery (canary, feature flags) to bound risk so most changes can be low-risk by design.

Family 05

Cost & FinOps Diseases

9 diseases in this family · click any to expand

Cloud Waste

Cloud Waste is the condition where infrastructure is overprovisioned, idle, or wrongly sized. The company pays for capacity it does not use. Spend grows because nobody owns the question of whether each resource is actually earning its cost.

  • Resource utilisation (CPU, memory, storage IOPS) sits well below provisioned capacity at average load.
  • Idle or zombie resources (unattached volumes, stopped instances, unused load balancers) accumulate.
  • Right-sizing has not been done in 6+ months.
  • Reserved or savings-plan coverage is low; most spend is on-demand.
  • Storage tiers do not match access patterns (hot data on cold tier or vice versa).

Primarily causes Symptom 5 (Rising Cost).

  1. Low utilisation vs provisioned — CPU, memory, and storage IOPS sit well below provisioned capacity.
  2. Accumulated zombie resources — unattached volumes, stopped instances, and unused load balancers pile up.
  3. No regular right-sizing — right-sizing has not been done in 6+ months.
  4. Low reserved / savings coverage — most spend is on-demand rather than reserved or savings-plan.
  5. Mismatched storage tiers — hot data sits on cold tiers or vice versa.
  6. No utilisation targets / alerts — nothing flags resources that fall below a utilisation floor.

Run a right-sizing review against real utilisation data. Automate cleanup of zombie resources. Move steady-state workloads to reserved or savings plans. Match storage tiers to access patterns. Set utilisation targets and alert when resources fall below them. Treat cloud capacity as a continuously managed inventory, not a set-it-and-forget-it.

No Cost Ownership

No Cost Ownership is the condition where no one in engineering or product is accountable for the cost their decisions create. Cost is "Finance's problem." Engineers build without seeing the cost impact of their choices. Product ships features without unit economics.

  • Engineering teams cannot see their own cloud spend.
  • Cost is reviewed only at the company level, not the team or feature level.
  • Product roadmap decisions do not consider infrastructure cost.
  • No FinOps function or it exists only in Finance.
  • "Cost reduction" is a periodic project rather than a continuous discipline.

Primarily causes Symptom 5 (Rising Cost).

  1. Cost invisible to teams — engineering teams cannot see their own cloud spend.
  2. Company-level-only cost review — cost is reviewed only at company level, not team or feature level.
  3. No unit economics in product — roadmap decisions ignore infrastructure cost.
  4. No FinOps inside engineering — FinOps is absent or lives only in Finance.
  5. Cost as periodic project — cost reduction is an occasional project, not a continuous discipline.
  6. Cost not an engineering quality attribute — cost sits outside the quality attributes engineers own.

Make cost visible at the team level — every engineering team sees its own spend weekly. Establish a FinOps practice that sits inside engineering, not outside it. Include unit economics in feature and roadmap decisions. Tie cost to business value at the team level. Cost is an engineering quality attribute alongside performance, reliability, and security.

Weak Tagging

Weak Tagging is the condition where cloud spend cannot be allocated to product, feature, customer, tenant, or workflow. Leadership cannot see where cost goes. Cost optimisation is impossible because the data does not exist to make trade-offs.

  • A meaningful percentage of cloud spend is untagged or tagged inconsistently.
  • Cost cannot be reported by product, feature, customer cohort, or workflow.
  • Tag standards exist on paper but are not enforced in infrastructure-as-code.
  • "Showback" or chargeback to teams is impossible because allocation is unreliable.

Primarily causes Symptom 5 (Rising Cost).

  1. Untagged / inconsistent spend — a meaningful share of spend is untagged or tagged inconsistently.
  2. No allocation dimensions — cost cannot be reported by product, feature, cohort, or workflow.
  3. Tag standard not enforced — tag standards exist on paper but are not enforced in IaC.
  4. No showback / chargeback — allocation is too unreliable to attribute cost to teams.
  5. No tag backfill — historical resources are never tagged retroactively.

Define a tagging standard (product, team, environment, customer-tier, workflow). Enforce tags in infrastructure-as-code and block untagged resources at provisioning. Backfill historical resources. Build cost dashboards that slice by tag. Until cost can be allocated, it cannot be governed.

No Autoscaling

No Autoscaling is the condition where capacity does not flex with demand. The system runs at peak even at low load, or fails to scale up under unexpected load. Cost is paid 24/7 for capacity needed only at peaks.

  • Compute, container, or database capacity is fixed regardless of load pattern.
  • Capacity is provisioned for the worst-case peak and never scaled down.
  • Scale events require manual intervention.
  • Off-hours and weekend cost is the same as peak hours.

Primarily causes Symptom 5 (Rising Cost). Contributes to Symptom 3 (Architecture Not Fit) during unexpected peaks.

  1. Fixed capacity regardless of load — compute, container, or DB capacity is static.
  2. Peak-provisioned, never scaled down — capacity is sized for worst-case peak and never reduced.
  3. Manual scale events — scaling requires manual intervention.
  4. Flat off-hours cost — weekend and off-hours cost equals peak cost.
  5. No demand-tied scaling policies — there are no policies tied to real demand signals.
  6. No spot / preemptible use — fault-tolerant workloads do not use cheaper spot capacity.

Implement autoscaling at the right granularity (instance, container, pod, database read replica) with sensible scaling policies tied to real demand signals. Test scaling behaviour deliberately. Move from peak-provisioning to elastic-provisioning. Use spot or preemptible capacity for fault-tolerant workloads.

Vendor Sprawl

Vendor Sprawl is the condition where multiple vendors and tools solve overlapping problems. Contracts renew without value review. The company pays for capabilities it duplicates internally, and nobody can answer whether each vendor still earns its cost.

  • Multiple tools cover the same use case (observability, CI/CD, security scanning, project management).
  • Contracts auto-renew without an explicit value review.
  • Usage data per vendor is unknown or unreviewed.
  • New tool adoption happens without retiring an older one.
  • No vendor management function with authority to consolidate.

Primarily causes Symptom 5 (Rising Cost).

  1. Overlapping tools per category — multiple tools cover the same use case.
  2. Auto-renewal without review — contracts renew without an explicit value check.
  3. Unknown per-vendor usage — usage data per vendor is unknown or unreviewed.
  4. Adopt-without-retire pattern — new tools are added without retiring older ones.
  5. No vendor-management authority — no function has authority to consolidate vendors.
  6. No "one tool per category" policy — there is no default policy forcing justification for duplicates.

Audit vendor portfolio by use case and identify overlaps. Establish a renewal review process that requires evidence of usage and value before each renewal. Set a "one tool per category unless justified" policy. Create a vendor management function with authority to consolidate. Renewal season is the moment of leverage — use it.

Expensive Queries

Expensive Queries is the condition where inefficient database queries, missing indexes, and unnecessary data movement inflate compute and storage cost. The application is paying full price for work it should not be doing.

  • A small number of queries consume a disproportionate share of database load.
  • Full table scans appear in production query plans.
  • N+1 query patterns persist despite being well-known anti-patterns.
  • Data is moved across regions or services unnecessarily, incurring transfer cost.
  • ORM-generated queries are not reviewed for efficiency.

Primarily causes Symptom 5 (Rising Cost). Contributes to Symptom 3 (Architecture Not Fit).

  1. Disproportionate query load — a few queries consume an outsized share of database load.
  2. Full table scans in production — production query plans show full scans.
  3. Persistent N+1 patterns — known N+1 anti-patterns remain in place.
  4. Unnecessary data movement — data crosses regions or services needlessly, incurring transfer cost.
  5. Unreviewed ORM queries — ORM-generated queries are never reviewed for efficiency.
  6. Missing / unmaintained indexes — required indexes are absent or stale.

Profile and fix the worst queries first using production query analytics. Add missing indexes and review existing ones for utility. Eliminate N+1 patterns systematically. Minimise cross-region and cross-service data movement. Review ORM-generated queries as you would review handwritten ones.

Premium Service Overuse

Premium Service Overuse is the condition where premium managed services are used by default rather than by deliberate business justification. The team picks the most convenient or fashionable service tier without weighing cost against value. The company pays for managed convenience it does not need.

  • Production workloads run on premium tiers where standard tiers would suffice.
  • Managed services are chosen reflexively over self-managed alternatives without cost analysis.
  • "Serverless" or "fully managed" is preferred even where workload patterns make it more expensive.
  • No tier-selection criteria documented; tier choice is left to individual engineers.

Primarily causes Symptom 5 (Rising Cost).

  1. Premium tiers by default — production runs on premium tiers where standard would suffice.
  2. Reflexive managed-service choice — managed is chosen over self-managed without cost analysis.
  3. Serverless regardless of fit — fully managed is preferred even where workload patterns make it costlier.
  4. No tier-selection criteria — there is no documented basis for tier choice; it is left to individuals.
  5. No periodic right-tiering — existing workloads are never re-tiered against actual usage.

Document tier-selection criteria per service category. Require explicit cost-benefit justification for premium tiers. Periodically right-tier existing workloads based on actual usage and growth. Recognise that managed convenience has a price — pay it deliberately, not reflexively.

Hidden Rework Cost

Hidden Rework Cost is the condition where defects, incidents, and refactors consume significant engineering capacity but are never measured as cost. The company pays for the same work twice (build, then fix, then fix again) without seeing it as a cost line.

  • No tracking of engineering capacity spent on defects, incidents, and unplanned work.
  • Sprint reviews show low completion rates with no analysis of where time actually went.
  • The same engineers are repeatedly pulled into incident response.
  • Refactor work is invisible in the roadmap but real in the capacity.

Primarily causes Symptom 5 (Rising Cost) and Symptom 4 (Delayed Product Releases).

  1. Untracked unplanned work — capacity spent on defects, incidents, and unplanned work is not tracked.
  2. Unexplained low completion — sprint reviews show low completion with no analysis of where time went.
  3. Repeated incident pull-ins — the same engineers are repeatedly pulled into incident response.
  4. Invisible refactor capacity — refactor work is absent from the roadmap but real in capacity.
  5. No capacity-by-category accounting — capacity is not split across feature, debt, defect, incident, support, refactor.
  6. No prevention business case — rework cost is invisible, so investment in prevention cannot be justified.

Track engineering capacity by category — feature, debt, defect, incident, support, refactor, meeting. Make rework cost visible at the team and executive level. Use the data to justify investment in prevention (test automation, quality, observability, debt reduction). What is not measured cannot be reduced.

No Value Linkage

No Value Linkage is the condition where technology spend is not tracked against business outcome. The company cannot tell which spend earned its keep. Budget conversations become political because they are not grounded in evidence.

  • Technology budget is allocated by team or function, not by business outcome served.
  • No view of cost per customer, per product line, per feature, or per workflow.
  • "ROI of engineering" is impossible to answer with data.
  • Cost cuts are made without knowing which spend was earning its value.

Primarily causes Symptom 5 (Rising Cost).

  1. Function-based budget allocation — budget is allocated by team or function, not by outcome served.
  2. No per-outcome cost view — there is no cost per customer, product line, feature, or workflow.
  3. Unanswerable engineering ROI — "ROI of engineering" cannot be answered with data.
  4. Blind cost cuts — cuts are made without knowing which spend was earning value.
  5. No cost-to-value model — there is no model linking spend to revenue, cost-per-transaction, or margin.
  6. Value linkage not in the cadence — the spend-to-outcome connection is not part of the operating rhythm.

Build a cost-to-value model that links technology spend to business outcomes (revenue per customer, cost per transaction, margin per product line). Review spend through this lens at every budget cycle. Make the connection between spend and outcome part of the operating cadence. Without value linkage, all cost discussions become opinion.

Family 06

Operating Model Diseases

8 diseases in this family · click any to expand

No Shared Operating Model

No Shared Operating Model is the condition where there is no common operating system connecting business strategy, product roadmap, architecture decisions, and engineering execution. Each function has its own model. The company is several organisations pretending to be one.

  • Each function (business, product, engineering, design, ops) operates on different planning horizons, cycles, and metrics.
  • Strategy reviews and execution reviews happen separately, with no bridge.
  • Cross-functional decisions take weeks because no single forum owns them.
  • Functional leaders cannot describe how their function fits into the company's operating system.

Primarily causes Symptom 6 (Misalignment) and Symptom 4 (Delayed Releases).

  1. Divergent planning horizons — each function plans on different horizons, cycles, and metrics.
  2. Strategy and execution disconnected — strategy reviews and execution reviews happen with no bridge.
  3. No forum for cross-functional decisions — cross-functional decisions take weeks because no forum owns them.
  4. No shared operating system — leaders cannot describe how their function fits the company's operating model.
  5. No single cadence — there is no one rhythm on which the whole company plans and reviews.

Define one operating model that connects company strategy → product strategy → architecture strategy → engineering execution → measurement. Run it on one cadence. Hold one set of reviews where trade-offs across functions get resolved. The operating model is what turns several functions into one company.

Missing KPI Hierarchy

Missing KPI Hierarchy is the condition where business KPIs, product KPIs, and system signals are not linked. Teams optimise things that do not move the business. Leadership has metrics but no diagnostic chain from outcome to driver to system.

  • Team-level metrics cannot be aggregated to product or business KPIs.
  • Different teams optimise metrics that conflict with each other.
  • When a business KPI moves, leadership cannot decompose the cause to a system signal.
  • "Vanity metrics" persist because their connection to outcomes is not interrogated.

Primarily causes Symptom 6 (Misalignment). Contributes to Symptom 1 (Slow Growth) and Symptom 8 (Missing Observability).

  1. Non-aggregating team metrics — team metrics cannot roll up into product or business KPIs.
  2. Conflicting metrics across teams — teams optimise metrics that work against each other.
  3. No decomposition chain — when a business KPI moves, leadership cannot trace it to a system signal.
  4. Persisting vanity metrics — metrics with no outcome chain are never interrogated or retired.
  5. No outcome→metric→signal map — there is no defined hierarchy from outcome to driver to system.

Define the KPI hierarchy from business outcome → product metric → system signal. Map every team-level metric back to a business KPI it influences. Eliminate metrics with no chain to outcome. Use the hierarchy in every operating review so leadership can decompose change end to end.

Weak Product Discovery

Weak Product Discovery is the condition where features are decided by opinion or stakeholder pressure rather than by validated business and customer signals. The product team becomes an order-taker. Investment goes to whichever voice is loudest, not whichever opportunity is highest-value.

  • Features are committed to before discovery (customer research, data, prototyping, sizing).
  • Product managers spend more time managing stakeholders than understanding customers.
  • Discovery is treated as optional or compressed to "we have a Figma."
  • Hypothesis-driven prioritisation is absent; the loudest voice wins.

Primarily causes Symptom 6 (Misalignment) and Symptom 1 (Slow Growth).

  1. Commitment before discovery — features are committed before research, data, prototyping, or sizing.
  2. Stakeholder-management over customers — PMs spend more time managing stakeholders than understanding customers.
  3. Discovery treated as optional — discovery is skipped or compressed to "we have a Figma."
  4. Loudest-voice prioritisation — the loudest stakeholder wins, not the strongest hypothesis.
  5. No discovery cadence / phases — there is no staged discovery (opportunity→insight→prototype→validation→commit).
  6. PM under-resourced for discovery — product has no protected time for discovery, only delivery.

Establish a discovery cadence with explicit phases — opportunity, customer insight, prototype, validation, commitment. Require evidence at each phase. Resource product management with time for discovery, not just delivery. Replace stakeholder-loudest with hypothesis-strongest.

No Platform Mindset

No Platform Mindset is the condition where reusable platform capabilities are not built. Every team rebuilds the same plumbing differently. The company pays for the same work many times and ends up with inconsistent, fragmented capability.

  • Authentication, identity, payments, notifications, search, observability are implemented differently across teams.
  • No platform engineering function with explicit charter to build internal products.
  • Teams reject reusable platforms because "ours is special" without testing the assumption.
  • Onboarding new engineers takes long because every service is structurally different.

Primarily causes Symptom 6 (Misalignment) and Symptom 4 (Delayed Releases). Contributes to Symptom 5 (Rising Cost).

  1. Duplicated plumbing across teams — auth, identity, payments, notifications, search, observability are built differently per team.
  2. No platform engineering function — there is no platform team chartered to build internal products.
  3. "Ours is special" rejection — teams reject reusable platforms without testing the assumption.
  4. Structurally different services — every service differs structurally, so onboarding engineers is slow.
  5. Platforms not treated as products — internal platforms are not built to be easier than rebuilding.
  6. No adoption / time-saved measurement — platform adoption and savings across teams are not measured.

Establish a platform engineering function with internal products for the shared capabilities (auth, identity, payments, notifications, observability, deployment, data). Treat internal teams as customers — platforms must be easier than rebuilding. Measure platform adoption and the time saved across consuming teams.

Status-Only Meetings

Status-Only Meetings is the condition where leadership reviews report status instead of resolving trade-offs and ranking priorities. The forums that should produce decisions produce updates. The decisions get pushed elsewhere, slowing the system.

  • Standing reviews end without a decision logged.
  • The same trade-off is discussed in multiple meetings without resolution.
  • Pre-reads describe what happened, not what decision is needed.
  • Action items are status updates ("update the board") rather than decisions ("approve X over Y").

Primarily causes Symptom 6 (Misalignment) and Symptom 4 (Delayed Releases).

  1. Reviews end without decisions — standing reviews close with no decision logged.
  2. Trade-offs re-litigated everywhere — the same trade-off recurs across meetings unresolved.
  3. Progress-only pre-reads — pre-reads describe what happened, not what decision is needed.
  4. Status-as-action-item — action items are updates ("update the board"), not decisions.
  5. Forums not designed around decisions — meetings are not structured to begin with the decision required.
  6. No decision log — decisions are not recorded with owner and rationale.

Redesign forums around decisions, not status. Every meeting starts with the decision required. Pre-reads describe options and trade-offs, not just progress. Decisions get logged with owner and rationale. If a meeting has no decision pending, it does not need to happen.

Hidden Architecture Constraints

Hidden Architecture Constraints is the condition where architecture limits are invisible to business leaders. Commitments are made the system cannot keep. Promises are made to customers, partners, or investors that the platform cannot deliver in the timeframe expected.

  • Sales or marketing commits to capabilities the platform cannot deliver.
  • "Can we scale to X by Y?" is answered with opinion, not architecture evidence.
  • Enterprise deals are won with conditions the platform must scramble to meet.
  • Roadmap promises are made without architecture review.

Primarily causes Symptom 6 (Misalignment) and Symptom 4 (Delayed Releases).

  1. Commitments beyond platform capability — sales/marketing commit to capabilities the platform cannot deliver.
  2. Opinion-based scale answers — "can we scale to X by Y?" is answered by opinion, not architecture evidence.
  3. Deals won with unmet conditions — enterprise deals are won with conditions the platform must scramble to meet.
  4. No architecture review before commitment — roadmap promises are made without architecture review.
  5. Constraints not translated for leaders — architecture limits are not expressed in language leaders can use.
  6. No capability map — there is no leadership-facing map of what the platform supports today.

Surface architecture constraints to business leaders in language they can use. Establish a "can we?" architecture review before commitments are made. Maintain a public-to-leadership map of what the architecture can support today and what it would take to extend. Architecture must be a partner in commercial decisions, not an after-the-fact problem.

Function-Based Teams

Function-Based Teams is the condition where teams are organised by function (frontend, backend, QA, mobile) rather than by business capability. Every outcome requires coordination across multiple functional teams. Speed dies in the hand-offs.

  • Feature delivery requires sequential handoffs (backend → frontend → QA → release).
  • Each functional team has its own backlog with its own priorities that may conflict.
  • Cross-functional work is "everyone's responsibility, no one's accountability."
  • Status meetings are dominated by coordination instead of progress.

Primarily causes Symptom 6 (Misalignment) and Symptom 4 (Delayed Releases).

  1. Sequential functional hand-offs — delivery flows backend→frontend→QA→release through hand-offs.
  2. Conflicting functional backlogs — each functional team has its own, sometimes conflicting, priorities.
  3. Diffused cross-functional accountability — cross-functional work is everyone's responsibility, no one's accountability.
  4. Coordination-dominated meetings — status meetings are dominated by coordination, not progress.
  5. No capability-aligned teams — teams are not organised around end-to-end business capabilities.
  6. Functional purity over end-to-end ownership — the model optimises functional specialism over ownership of outcomes.

Reorganise teams around business capabilities (Onboarding, Checkout, Payments, Search, Recommendations) — each team owns its slice end to end. Functional expertise (design, QA, security, data) embeds in capability teams or operates as a guild. Optimise for end-to-end ownership over functional purity.

Unclear Ownership

Unclear Ownership is the condition where customer journeys, services, and business outcomes cross teams without a single owner. When the journey breaks, no one is accountable. When it needs improvement, no one drives it.

  • Critical services or journeys cannot be assigned to a single owning team.
  • "Tragedy of the commons" — shared services degrade because no one owns them.
  • Escalations bounce between teams.
  • The same gap is identified in multiple incident reviews but never closed.

Primarily causes Symptom 6 (Misalignment) and Symptom 2 (Poor CX).

  1. Unassignable critical services — critical services or journeys cannot be assigned to one owning team.
  2. Tragedy of the commons — shared services degrade because no one owns them.
  3. Bouncing escalations — escalations bounce between teams with no resolver.
  4. Repeated unclosed gaps — the same gap recurs in incident reviews but is never closed.
  5. No ownership visibility — there is no service catalog with owner, on-call, and SLOs.
  6. No end-to-end accountability — owners are held to their slice, not to end-to-end health.

Map every critical service, journey, and outcome to a single owning team. Make ownership visible (a service catalog with owner, on-call, SLOs). Hold owners accountable for end-to-end health, not just their slice. Unowned things rot — make ownership the default, not the exception.

Family 07

Leadership & Capability Diseases

10 diseases in this family · click any to expand

No Diagnostic Framework

No Diagnostic Framework is the condition where there is no structured way to connect business pain to technology root cause. Decisions are reactive. Major investments are approved without clear diagnosis. The company spends on wrong treatments because there is no discipline to find the right ones.

  • Major decisions (cloud migration, microservices, AI programme) are approved without root-cause analysis.
  • Different leaders explain the same problem differently and there is no method to converge.
  • Treatment is chosen by trend, vendor, or executive opinion rather than evidence.
  • The company has no shared language for diagnosing technology root causes.

Primarily causes Symptom 7 (Weak Leadership and Capability). Indirect cause of every other symptom.

  1. Decisions without root-cause analysis — major moves (migration, microservices, AI) are approved with no diagnosis.
  2. No convergence method — leaders explain the same problem differently with no way to converge.
  3. Trend / vendor / opinion-led treatment — treatment is chosen by trend, vendor, or executive opinion, not evidence.
  4. No shared diagnostic language — there is no common language for technology root causes.
  5. No symptom→disease→evidence mapping — there is no framework linking symptoms to candidate diseases to evidence.
  6. Diagnosis not in the operating cadence — diagnostic discipline is not built into how leadership reasons.

Adopt a diagnostic framework that maps symptoms to candidate diseases to evidence. Make every major treatment decision pass through the diagnostic. Build the discipline into the operating cadence so it becomes the default way leadership reasons about technology. *Symptoms → Diagnosis → Prescription → Treatment* is the framework.

Reactive Decisions

Reactive Decisions is the condition where major decisions are made under pressure, urgency, or trend, not from evidence. Each crisis produces a response. Each board meeting produces a commitment. The cumulative effect is a portfolio of half-finished initiatives with no coherent direction.

  • Strategic direction shifts each quarter based on the latest pressure.
  • Major commitments are made without a clear trigger criterion ("we should do AI because everyone is doing AI").
  • Past decisions are not revisited even when conditions change.
  • The strategic backlog grows faster than the strategic execution.

Primarily causes Symptom 7 (Weak Leadership and Capability).

  1. Pressure-driven strategy shifts — direction changes each quarter with the latest pressure.
  2. No commitment trigger criteria — commitments lack explicit "what would have to be true" criteria.
  3. Decisions never revisited — past decisions are not re-examined when conditions change.
  4. Backlog outgrows execution — the strategic backlog grows faster than strategic execution.
  5. Trend-following commitments — "we should do X because everyone is" substitutes for analysis.
  6. No "no" discipline — saying no is not treated as a normal strategic act.

Slow down strategic decisions. Require explicit criteria (what would have to be true to commit) before major investments. Maintain a strategic review cadence where decisions are revisited against actual evidence. Saying no is a strategic act — make it normal.

No Decision Records

No Decision Records is the condition where architecture and strategic decisions are not recorded. Past decisions cannot be revisited or learned from. The company repeats mistakes because the reasoning is lost. New leaders cannot understand the system because the why is not written down.

  • Architecture Decision Records (ADRs) are missing or stale.
  • Strategic decisions live in slack threads and meeting notes that nobody can find.
  • When a leader leaves, their decisions become un-debuggable.
  • The same architectural debates recur every 18–24 months.

Primarily causes Symptom 7 (Weak Leadership and Capability).

  1. Missing / stale ADRs — Architecture Decision Records are absent or out of date.
  2. Decisions buried in chat — strategic decisions live in Slack threads and notes nobody can find.
  3. Un-debuggable departed-leader decisions — when a leader leaves, their decisions become opaque.
  4. Recurring architectural debates — the same debates recur every 18–24 months for lack of a record.
  5. No record template — there is no lightweight ADR/decision template capturing context, options, and consequences.
  6. Records never reviewed on change — old decisions are not reviewed when conditions shift.

Adopt lightweight ADRs for architecture and similar templates for strategic decisions. Record context, options considered, decision, and consequences. Make ADRs the default for any decision worth defending later. Review old ADRs when conditions change. Decisions are organisational assets — write them down.

Missing Operating Cadence

Missing Operating Cadence is the condition where there is no regular operating rhythm between CEO, Product, Technology, Finance, and Operations to surface and resolve trade-offs. Cross-functional issues are handled ad hoc. The company runs on heroics and escalations instead of cadence.

  • No standing forum for resolving cross-functional trade-offs.
  • Quarterly planning is the only cross-functional moment; intra-quarter trade-offs become escalations.
  • Cross-functional projects depend on personal relationships to move.
  • Leaders complain about being "in too many meetings" but the meetings are status, not decisions.

Primarily causes Symptom 7 (Weak Leadership and Capability) and Symptom 6 (Misalignment).

  1. No cross-functional trade-off forum — there is no standing forum to resolve cross-functional trade-offs.
  2. Quarterly-only cross-functional moment — intra-quarter trade-offs become escalations.
  3. Relationship-dependent execution — cross-functional projects move only on personal relationships.
  4. Meetings are status, not decisions — leaders are "in too many meetings" but the meetings produce no decisions.
  5. No short, decision-focused review — there is no regular evidence-led operating review against the operating model.

Establish a weekly or bi-weekly operating review where CEO, Product, Tech, Finance, and Ops resolve cross-functional trade-offs against the operating model. Keep it short, decision-focused, evidence-led. The operating cadence is what turns strategy into execution.

CTO Stuck in Execution

CTO Stuck in Execution is the condition where the technology leader is buried in delivery and cannot provide strategic diagnosis. They are valuable in the trenches but invisible in the boardroom. The company loses the strategic voice it needs because its highest technology leader is firefighting.

  • The CTO's calendar is dominated by execution, incidents, and code reviews.
  • The CTO does not have time to engage with strategy, board, customers, or M&A.
  • Reports to the CTO say "I miss having strategic conversations."
  • The CTO cannot delegate execution because the bench beneath is thin.

Primarily causes Symptom 7 (Weak Leadership and Capability).

  1. Execution-dominated CTO calendar — the CTO's time is consumed by execution, incidents, and code reviews.
  2. No time for strategy/board/customers/M&A — the CTO cannot engage strategically.
  3. Thin bench beneath — the CTO cannot delegate execution because the bench is thin.
  4. No senior execution leaders — there are no VPs of Engineering/Architecture/Platform to own execution.
  5. Strategic conversations lost — reports miss strategic engagement that no longer happens.
  6. CTO operating below required level — the role runs at the level the company allows, not the level it needs.

Free the CTO from execution by building the leadership bench beneath. Promote or hire VPs of Engineering, Architecture, and Platform who can own execution. Reset the CTO's calendar around strategy, board, customers, and M&A. The CTO must operate at the level the company needs, not the level the company allows.

Key-Person Dependency

Key-Person Dependency is the condition where critical knowledge or capability sits with one or two people. Their absence is an operational risk. The company is one resignation away from a serious problem.

  • Specific systems have only one person who understands them deeply.
  • Vacation calendars are scrutinised because certain people cannot be away.
  • Documentation exists in heads, not in writing.
  • Bus-factor of one for critical capabilities.

Primarily causes Symptom 7 (Weak Leadership and Capability).

  1. Bus-factor of one — critical systems have only one person who understands them.
  2. Knowledge only in heads — documentation lives in people's heads, not in writing.
  3. No pairing on critical systems — no other engineers are paired onto critical systems.
  4. Constrained by key-person availability — vacation calendars are scrutinised because certain people cannot be away.
  5. No failure drills — there are no drills simulating the key person's absence to surface gaps.

Identify the key-person dependencies explicitly. Document the critical knowledge in writing. Pair other engineers with the key people on critical systems. Run drills where the key person is "unavailable" to surface gaps. Bench is not a luxury — it is operational risk management.

Weak Engineering Ladder

Weak Engineering Ladder is the condition where there is no clear career framework, so growth, accountability, and bench-building all weaken. Engineers cannot see how to grow. Managers cannot calibrate fairly. Hiring loses standards because there is nothing to hire against.

  • No documented engineering levels with criteria for each.
  • Promotion decisions vary by manager rather than by standard.
  • Engineers leave because "there is no path."
  • Hiring loops produce inconsistent calibration.

Primarily causes Symptom 7 (Weak Leadership and Capability).

  1. No documented levels — there are no engineering levels with criteria per level.
  2. Manager-dependent promotions — promotion decisions vary by manager, not by standard.
  3. No growth path drives attrition — engineers leave because "there is no path."
  4. Inconsistent hiring calibration — hiring loops produce inconsistent calibration with nothing to hire against.
  5. No manager calibration training — managers are not trained to calibrate fairly.
  6. Ladder not visible to engineers — growth is not made public and aspirational.

Adopt or adapt an engineering ladder with clear criteria per level. Train managers on calibration. Run regular promotion cycles with explicit standards. Make the ladder public to engineers so growth is visible and aspirational. A clear ladder is one of the highest-leverage retention investments.

Weak Hiring Brand

Weak Hiring Brand is the condition where the company struggles to attract senior engineering, product, and data talent. Roles stay open. Offers get declined. The bar drops because the pipeline is thin.

  • Time-to-fill for senior roles exceeds market norms.
  • Offer-acceptance rate is low.
  • Inbound applicant quality is poor; sourcing is the only channel that works.
  • The engineering brand is invisible (no blog, no conferences, no open source, no public technical voice).

Primarily causes Symptom 7 (Weak Leadership and Capability).

  1. Slow time-to-fill — senior roles take longer than market norms to fill.
  2. Low offer-acceptance — offers are declined at a high rate.
  3. Thin inbound pipeline — applicant quality is poor; only sourcing works.
  4. Invisible engineering brand — no blog, conferences, open source, or public technical voice.
  5. Below-market compensation — pay targets the talent the company can afford, not the talent it wants.
  6. Slow / disrespectful hiring process — the process is not fast, respectful, or predictive.

Invest in technical brand — engineering blog, conference presence, open source, public technical leadership. Make the hiring process respectful, fast, and predictive. Compensate at market for the talent you want, not the talent you can afford. Brand compounds; start now.

No Learning System

No Learning System is the condition where there is no structured capability-building for engineers, leads, and leaders. People grow accidentally, if at all. Skills atrophy. The capability of the organisation falls behind the demands placed on it.

  • No learning budget per engineer, or budget is unused.
  • No internal teaching, mentoring, or rotation programmes.
  • Leadership development is absent or limited to a few executives.
  • Skills critical for the next stage of the business are not being built deliberately.

Primarily causes Symptom 7 (Weak Leadership and Capability).

  1. No / unused learning budget — there is no learning budget per engineer, or it goes unused.
  2. No internal teaching or mentoring — there are no teaching, mentoring, or rotation programmes.
  3. Absent leadership development — leadership development is missing or limited to a few executives.
  4. Next-stage skills not built — skills the next stage of the business needs are not built deliberately.
  5. No rotation paths — there are no rotation paths to build cross-team capability.
  6. Capability treated as cost — capability is tolerated as a cost, not invested in as an asset.

Establish a learning system — budget, time, structured programmes for technical and leadership growth. Run internal teaching forums. Create rotation paths that build cross-team capability. Treat capability as an asset the company invests in, not a cost it tolerates.

Thin Bench

Thin Bench is the condition where critical roles have no clear successor. One departure can stall a major initiative. The company has senior leaders but no second line ready to step up.

  • Succession map is absent or theoretical.
  • Critical roles have no internal candidate within 12–18 months of readiness.
  • Internal promotions to senior roles are rare; senior hires are external.
  • Reorganisations expose unprepared leaders because there was no one else.

Primarily causes Symptom 7 (Weak Leadership and Capability).

  1. Absent succession map — the succession map is missing or theoretical.
  2. No ready internal candidates — critical roles have no internal candidate within 12–18 months of readiness.
  3. Rare internal promotions — senior roles are filled externally; internal promotion is rare.
  4. Reorgs expose unprepared leaders — reorganisations reveal there was no second line.
  5. No stretch investment — high-potential second-line leaders get no stretch opportunities.
  6. Bench-building not a measured responsibility — bench depth is not a measurable leadership responsibility.

Build a succession map for every critical role with current bench depth. Invest in stretch opportunities for high-potential second-line leaders. Make bench-building a measurable leadership responsibility. Bench is built years in advance, not in the moment of need.

Family 08

Observability Diseases

10 diseases in this family · click any to expand

No Journey Observability

No Journey Observability is the condition where there is no end-to-end visibility of customer journeys across frontend, backend, database, cache, queue, CDN, and third-party systems. The team can see each component but cannot see the customer's path through them.

  • Tracing is missing or stops at service boundaries.
  • Customer-reported issues cannot be reproduced because the journey cannot be reconstructed.
  • Frontend, backend, and infrastructure signals are not stitched together.
  • "Which release caused the issue?" cannot be answered in minutes.

Primarily causes Symptom 8 (Missing Observability) and Symptom 2 (Poor CX).

  1. Tracing stops at service boundaries — distributed tracing is missing or does not cross services.
  2. Journeys cannot be reconstructed — customer-reported issues cannot be reproduced as a path.
  3. Unstitched layer signals — frontend, backend, and infrastructure signals are not joined.
  4. "Which release?" unanswerable — the team cannot identify the offending release in minutes.
  5. No consistent trace IDs — there are no trace IDs flowing browser-to-backend.
  6. Journey observability not release-blocking — new services ship without end-to-end visibility.

Adopt distributed tracing across frontend, backend, and infrastructure. Use consistent trace IDs from browser to backend. Build journey dashboards that show end-to-end flow for critical paths. Make journey visibility a release-blocking requirement for new services.

Weak Instrumentation

Weak Instrumentation is the condition where application code, infrastructure, and third-party integrations are inconsistently instrumented. Signals are missing, low quality, or untrusted. The observability system reflects what was easy to instrument, not what matters.

  • Logging is inconsistent across services in format, structure, and level.
  • Critical business events (purchase, signup, payment, churn) are not consistently emitted as observability signals.
  • Third-party calls and external dependencies are observability black holes.
  • New services ship without an instrumentation standard.

Primarily causes Symptom 8 (Missing Observability).

  1. Inconsistent logging — logging differs across services in format, structure, and level.
  2. Business events not emitted — purchase, signup, payment, and churn are not emitted as signals.
  3. Third-party black holes — external calls and dependencies are not instrumented.
  4. No instrumentation standard — new services ship without an enforced instrumentation standard.
  5. No CI enforcement — instrumentation is not enforced in CI as part of "done."
  6. Instrument-what-is-easy bias — instrumentation reflects what was easy, not what matters.

Establish an instrumentation standard (log format, span attributes, metric naming, required events) and enforce it in CI. Instrument business events as first-class signals, not just technical ones. Cover third-party calls. Treat instrumentation as part of "done."

Disconnected Signals

Disconnected Signals is the condition where logs, metrics, traces, real-user monitoring, and product analytics exist but are not correlated. Root-cause analysis takes too long because each signal tells part of the story and the team has to assemble the rest manually.

  • Logs, metrics, and traces live in different tools without shared identifiers.
  • Real-user monitoring and backend tracing are not linked.
  • Product analytics (Mixpanel, Amplitude) and system observability are different worlds.
  • Incident timelines have to be manually assembled from multiple consoles.

Primarily causes Symptom 8 (Missing Observability).

  1. No shared identifiers — logs, metrics, and traces lack shared trace/session/user IDs.
  2. RUM not linked to backend — real-user monitoring and backend tracing are not joined.
  3. Product analytics siloed — product analytics and system observability are separate worlds.
  4. Manual incident-timeline assembly — timelines must be assembled by hand from multiple consoles.
  5. Tools without a joinable data model — tools cannot be joined because the data model is not unified.

Adopt a unified observability strategy with shared identifiers (trace ID, session ID, user ID) across logs, metrics, traces, RUM, and product analytics. Where tools cannot be unified, ensure the data model can be joined. Make correlation a design property of the observability platform.

Tool Sprawl

Tool Sprawl is the condition where multiple observability tools overlap. There is no unified strategy or shared data model. Teams use different tools, lose context across them, and pay multiple vendors for overlapping capability.

  • Multiple tools cover the same observability category (APM, log management, RUM, error tracking, infrastructure monitoring).
  • Different teams use different tools; switching context between them is friction.
  • Vendor cost grows with no consolidation review.
  • Observability data cannot be joined across tools.

Primarily causes Symptom 8 (Missing Observability) and Symptom 5 (Rising Cost).

  1. Overlapping observability tools — multiple tools cover the same category (APM, logs, RUM, error, infra).
  2. Different tools per team — teams use different tools, losing context across them.
  3. Unreviewed growing vendor cost — observability vendor cost grows with no consolidation review.
  4. Data cannot be joined across tools — observability data is not joinable across tools.
  5. No tooling strategy / data-model standard — there is no strategy or data-model standard to consolidate against.

Define an observability tooling strategy and consolidate where possible. Pick one tool per category, or a unified platform that spans them. Establish data-model standards so any future tool addition fits the existing model. Tool consolidation is also a cost lever.

Poor Alert Design

Poor Alert Design is the condition where alerts are noisy, ignored, or unconnected to business impact. Severity classification is weak. Engineers stop trusting alerts because most are false alarms. Real issues hide in the noise.

  • Alert volume is high; most alerts are routinely ignored or auto-resolved.
  • Alert severity is uniform rather than tiered.
  • Alerts fire on technical symptoms, not on business impact.
  • Incident response is delayed because the alert that mattered was lost in noise.

Primarily causes Symptom 8 (Missing Observability) and Symptom 2 (Poor CX).

  1. High-volume ignored alerts — alert volume is high and most are routinely ignored or auto-resolved.
  2. Uniform severity — alert severity is flat rather than tiered (page vs ticket vs FYI).
  3. Technical-threshold alerts — alerts fire on technical symptoms, not on business impact.
  4. Real alert lost in noise — incident response is delayed because the alert that mattered was buried.
  5. No alert-quality tracking — precision, recall, and ack time are not measured.
  6. Ignored alerts not tuned — repeatedly ignored alerts are never deleted or tuned.

Redesign alerts around customer impact and business outcomes, not raw technical thresholds. Tier severity sharply (page vs ticket vs FYI). Track alert quality (precision, recall, ack time). Delete or tune any alert that is repeatedly ignored. Alerts should earn the right to wake people up.

Missing Reliability Targets

Missing Reliability Targets is the condition where there are no service-level objectives or error budgets tied to customer journeys. Engineering optimises whatever feels broken, not what the business needs to be reliable.

  • No SLOs for customer-facing journeys.
  • No error budgets to govern the trade-off between reliability and velocity.
  • Reliability investment decisions are made by opinion, not by SLO miss.
  • Engineering and business have no shared language for "reliable enough."

Primarily causes Symptom 8 (Missing Observability) and Symptom 2 (Poor CX).

  1. No journey SLOs — there are no SLOs for customer-facing journeys.
  2. No error budgets — there are no error budgets to govern reliability-vs-velocity trade-offs.
  3. Opinion-based reliability investment — reliability decisions are made by opinion, not by SLO miss.
  4. No shared "reliable enough" language — engineering and business have no common definition of reliable.
  5. SLOs absent from the cadence — SLO performance is not part of the operating rhythm.

Define SLOs for customer-facing journeys with business stakeholders. Track error budgets and enforce velocity vs reliability trade-offs against them. Make SLO performance part of the operating cadence. SLOs are how reliability becomes a managed variable, not an aspirational adjective.

No Executive Dashboard

No Executive Dashboard is the condition where leadership has no health dashboard connecting technology signals to business outcomes. The CEO and board see business metrics; they cannot see the technology system driving them.

  • Executive reviews include business KPIs but no technology health.
  • When a business KPI moves, leadership cannot decompose to a system signal.
  • Technology health is reported as narrative ("things are good") rather than as data.
  • The CEO has no single page that says "is the system healthy enough for the strategy?"

Primarily causes Symptom 8 (Missing Observability) and Symptom 7 (Weak Leadership).

  1. No tech health in exec reviews — executive reviews include business KPIs but no technology health.
  2. No KPI-to-signal decomposition — when a business KPI moves, leadership cannot trace it to a system signal.
  3. Narrative-only health reporting — technology health is reported as "things are good," not as data.
  4. No single health page — there is no one page answering "is the system healthy enough for the strategy?"
  5. No off-cadence refresh — there is no real cadence keeping the dashboard current.
  6. No multi-dimension health coverage — reliability, performance, security, cost, velocity, and defects are not covered together.

Build an executive technology health dashboard covering reliability, performance, security posture, cost efficiency, delivery velocity, defect leakage, and key business KPIs the technology supports. Refresh it on a real cadence. Make it part of every operating review. Leadership cannot manage what it cannot see.

No Postmortem Discipline

No Postmortem Discipline is the condition where incidents are fixed but not systemically learned from. The same failure classes return. Each incident is treated as a one-off rather than as evidence of a pattern.

  • Postmortems are not written for most incidents, or they are written but not actioned.
  • Action items from postmortems are not tracked to completion.
  • Root cause is attributed to "human error" rather than systemic gap.
  • The same incident category recurs.

Primarily causes Symptom 8 (Missing Observability) and Symptom 2 (Poor CX).

  1. Postmortems not written / not actioned — most incidents get no postmortem, or postmortems are not actioned.
  2. Untracked action items — postmortem actions are not tracked to completion.
  3. Blame-the-human root cause — root cause is attributed to human error, not systemic gap.
  4. Recurring incident categories — the same incident category recurs without prevention.
  5. No incident categorisation — incidents are not categorised to surface patterns.
  6. No systemic prevention investment — repeat categories get no investment in testing, observability, or design.

Hold blameless postmortems for every meaningful incident. Track action items to completion. Categorise incidents to surface recurring patterns. Invest in systemic prevention (testing, observability, design) for repeat categories. Postmortem discipline is how an organisation learns.

No Signal Ownership

No Signal Ownership is the condition where signals exist but no team owns their quality, coverage, or business linkage. Dashboards rot. Alerts go stale. Coverage drifts. Observability becomes a fossil rather than a living system.

  • Dashboards are unmaintained; nobody is sure if the metrics still mean what they used to.
  • Alert thresholds are stale; alerts fire on conditions that no longer matter.
  • New services ship without observability and nobody notices.
  • "Who owns this dashboard?" is unanswerable.

Primarily causes Symptom 8 (Missing Observability).

  1. Unmaintained dashboards — dashboards rot and metrics no longer mean what they used to.
  2. Stale alert thresholds — alerts fire on conditions that no longer matter.
  3. New services ship without observability — coverage drifts because nobody notices the gaps.
  4. Unanswerable "who owns this?" — dashboard and signal ownership cannot be named.
  5. No ownership review cadence — there is no quarterly review of signal ownership.
  6. No retire-unused discipline — unused dashboards and alerts are never retired.

Assign every dashboard, alert, and signal to a named owner. Review ownership quarterly. Retire dashboards and alerts that nobody is using. Treat the observability system as a product with maintenance discipline.

No Self-Healing

No Self-Healing is the condition where known failure patterns are not automated. Recovery still depends on human action. Engineers wake up at night for failures the system could have handled itself.

  • Known failure patterns (full disks, queue backlogs, deadlocked databases, expired certs) cause incidents repeatedly.
  • Runbooks describe manual recovery for problems that could be automated.
  • The same human action is required for the same failure week after week.
  • "We need to add automation for that" is a recurring postmortem item that never gets done.

Primarily causes Symptom 8 (Missing Observability) and Symptom 2 (Poor CX).

  1. Repeated known-pattern incidents — full disks, queue backlogs, deadlocks, and expired certs cause incidents repeatedly.
  2. Manual runbook recovery — runbooks describe manual recovery for automatable problems.
  3. Same human action each time — the same manual fix is required week after week.
  4. Automation deferred indefinitely — "we need automation for that" is a recurring postmortem item never done.
  5. No automated failure responses — auto-scaling, auto-restart, queue drain, cert rotation, and fail-over are not in place.

Automate the most common, well-understood failure responses (auto-scaling, auto-restart, queue drain, certificate rotation, fail-over). Treat every recurring manual recovery as a candidate for automation. Self-healing is the natural endpoint of postmortem discipline.

Family 09

AI & Data Diseases

10 diseases in this family · click any to expand

No AI Strategy

No AI Strategy is the condition where there is no AI strategy connected to specific business outcomes. AI activity exists — pilots, demos, vendor pitches — but no measurable link to revenue, cost, productivity, decisioning, or customer experience.

  • AI initiatives are launched without business cases.
  • Leadership cannot name the top 3 AI use cases by business value.
  • AI roadmap is a list of technologies, not a list of outcomes.
  • AI investment is justified by fear of falling behind, not by expected value.

Primarily causes Symptom 9 (AI and Automation Not Producing Business Value).

  1. AI initiatives without business cases — initiatives launch with no business case.
  2. No ranked, value-based use cases — leadership cannot name the top 3 AI use cases by value.
  3. Technology-list roadmap — the AI roadmap is a list of technologies, not outcomes.
  4. Fear-of-falling-behind justification — investment is justified by fear, not expected value.
  5. No value×feasibility×time ranking — use cases are not ranked by value, feasibility, and time-to-impact.
  6. No discipline to say no — fashionable but low-value use cases are not declined.

Define an AI strategy starting from business outcomes — where can AI move revenue, cost, productivity, decisioning, or experience by a measurable amount? Rank use cases by value, feasibility, and time-to-impact. Concentrate investment on the few that earn it. Say no to the rest, including the fashionable ones.

Poor Data Quality

Poor Data Quality is the condition where data is incomplete, inconsistent, or fragmented. AI built on it cannot be trusted. Reports built on it conflict. Decisions built on it are wrong in ways the company does not realise.

  • Common entities (customer, order, product) have different definitions across systems.
  • Missing values, duplicates, and stale data are widespread.
  • No data quality metrics or owners.
  • Teams routinely "clean" data manually before using it.

Primarily causes Symptom 9 (AI and Automation Not Producing Business Value) and Symptom 8 (Missing Observability).

  1. Conflicting entity definitions — customer, order, and product differ across systems.
  2. Widespread missing / duplicate / stale data — gaps, duplicates, and stale records are pervasive.
  3. No data quality metrics or owners — quality is neither measured nor owned.
  4. Manual cleaning before use — teams routinely clean data by hand before using it.
  5. No golden entities / canonical schemas — there are no canonical schemas for core entities.
  6. Quality checked downstream, not upstream — quality is not enforced where data is created.

Establish data quality as a first-class engineering discipline with metrics, owners, and SLAs. Define golden entities with canonical schemas. Build quality checks into pipelines, not after them. Quality is upstream — fix it where the data is created, not where it is consumed.

No Single Source of Truth

No Single Source of Truth is the condition where business metrics conflict across systems. Teams cannot agree on basic numbers. Every leadership review starts with reconciling reports instead of making decisions.

  • Revenue, customer count, churn, and active-user numbers differ across teams and tools.
  • Multiple BI tools each produce different versions of the "same" metric.
  • Reconciliation is a recurring meeting.
  • Trust in dashboards is low; teams shadow their own.

Primarily causes Symptom 9 (AI and Automation Not Producing Business Value) and Symptom 6 (Misalignment).

  1. Conflicting core numbers — revenue, customer count, churn, and active users differ across teams and tools.
  2. Multiple BI tools, multiple truths — each BI tool produces a different version of the same metric.
  3. Recurring reconciliation meeting — reconciliation is a standing meeting, not an exception.
  4. Low trust drives shadow metrics — teams distrust dashboards and shadow their own.
  5. No metric owner / lineage — there is no owning team or clean upstream lineage per metric.
  6. No semantic layer — there is no shared semantic layer that all dashboards consume.

Establish a single source of truth per business metric with one owning team and a clean upstream lineage. Build a semantic layer that all dashboards and queries consume. Retire conflicting alternatives. Without a single truth, every other data investment compounds the confusion.

AI Without Data

AI Without Data is the condition where AI pilots run on data that is not ready. Results are unreliable and not safe to act on. The output may look impressive in demo but cannot survive production accuracy, governance, or compliance requirements.

  • AI projects start before data quality is validated for that use case.
  • Model accuracy in demo does not hold in production.
  • Pilots succeed on hand-curated data; productionisation fails on real data.
  • AI failures are blamed on the model when the cause is the data.

Primarily causes Symptom 9 (AI and Automation Not Producing Business Value).

  1. Build before data validation — AI projects start before data quality is validated for the use case.
  2. Demo accuracy fails in production — model accuracy in demo does not hold on real data.
  3. Hand-curated pilot data — pilots succeed on curated data; productionisation fails on real data.
  4. Model blamed for data failures — failures are blamed on the model when the cause is the data.
  5. No data-readiness gates — there are no completeness, freshness, and label-quality gates before model investment.

Validate data readiness for each specific AI use case before building the model. Treat data preparation as the first 70% of the AI project. Make data quality, completeness, freshness, and label quality explicit gates before model investment.

Hype-Driven Use Cases

Hype-Driven Use Cases is the condition where use cases are chosen by trend or excitement rather than by business value or feasibility. The company chases what is in the news instead of what would move its numbers.

  • Use cases are introduced as "we should do X because the market is doing X."
  • ROI is justified by analogy ("Company Y did this") rather than by the company's own economics.
  • Use case selection skips feasibility analysis (data, integration, compliance, change management).
  • Successful pilots cannot be scaled because the productionisation cost was not considered.

Primarily causes Symptom 9 (AI and Automation Not Producing Business Value).

  1. Market-imitation selection — use cases are introduced as "we should do X because the market is."
  2. ROI by analogy — ROI is justified by "Company Y did this," not by the company's own economics.
  3. Skipped feasibility analysis — selection skips data, integration, compliance, and change-management feasibility.
  4. Productionisation cost ignored — pilots cannot scale because productionisation cost was never considered.
  5. No use-case scoring framework — there is no value×feasibility×time-to-impact scoring to reject hype.

Adopt a use-case selection framework that scores business value × feasibility × time-to-impact. Reject use cases that fail feasibility regardless of how fashionable. Be willing to skip a trend if it does not fit the business — patience beats imitation.

Non-Standard Processes

Non-Standard Processes is the condition where underlying business processes are too inconsistent to automate cleanly. Automation amplifies whatever process it touches; an inconsistent process automated becomes inconsistent at scale.

  • The "same" process is performed differently by different people or teams.
  • No documented standard operating procedure for the process being targeted for automation.
  • Automation pilots stall because the underlying process changes mid-build.
  • Automation succeeds in one team but fails to spread because each team's process is different.

Primarily causes Symptom 9 (AI and Automation Not Producing Business Value).

  1. Same process done differently — the "same" process is performed differently by different people or teams.
  2. No documented SOP — there is no standard operating procedure for the targeted process.
  3. Process changes mid-build — automation stalls because the process shifts during the build.
  4. Automation fails to spread — it works in one team but not others because each team's process differs.
  5. Automating before standardising — the process is automated before it is standardised, trained, and measured.

Standardise processes before automating them. Document the standard, train against it, measure adherence. Automate the standardised version. Automating a chaotic process produces chaos at machine speed.

No Build-vs-Buy Model

No Build-vs-Buy Model is the condition where build, buy, partner, and integrate decisions are made ad hoc without an evaluation framework. The company builds what it should buy and buys what it should build. Investment goes to the wrong layer of the stack.

  • Major build decisions skip evaluation of buy alternatives.
  • Vendor purchases skip evaluation of build feasibility and TCO.
  • Teams default to "build" (engineering instinct) or "buy" (procurement instinct) without analysis.
  • The same build-vs-buy debate recurs without resolution because there is no method.

Primarily causes Symptom 9 (AI and Automation Not Producing Business Value) and Symptom 5 (Rising Cost).

  1. Build decisions skip buy evaluation — major build decisions never evaluate buy alternatives.
  2. Buy decisions skip build / TCO — vendor purchases skip build feasibility and total-cost-of-ownership.
  3. Instinct-driven defaults — teams default to "build" or "buy" by instinct, without analysis.
  4. Recurring unresolved debate — the same build-vs-buy debate recurs for lack of a method.
  5. No build-vs-buy scoring framework — there is no scoring of differentiation, TCO, time-to-value, integration, and switching cost.
  6. Decisions not recorded — build-vs-buy decisions are not documented to be revisited.

Adopt a build-vs-buy framework that scores strategic differentiation, total cost of ownership, time-to-value, integration cost, and switching cost. Use the framework on every major investment. Document the decision in a record that can be revisited. The right answer is rarely obvious without the framework.

No Production Pipeline

No Production Pipeline is the condition where models are built but not deployed, monitored, retrained, or governed at production grade. The company has a research function but no MLOps. Models live in notebooks rather than in the product.

  • Models are built in notebooks and never reach production.
  • Production models are not monitored for drift, accuracy, or freshness.
  • Retraining is manual and infrequent.
  • No A/B testing infrastructure for model variants.
  • Model deployments do not follow the same rigor as application deployments.

Primarily causes Symptom 9 (AI and Automation Not Producing Business Value).

  1. Models stuck in notebooks — models are built in notebooks and never reach production.
  2. No drift / accuracy monitoring — production models are not monitored for drift, accuracy, or freshness.
  3. Manual, infrequent retraining — retraining is manual and rare.
  4. No model A/B infrastructure — there is no infrastructure to test model variants.
  5. Lower rigor than app deploys — model deployments do not follow the same rigor as application deployments.
  6. No MLOps pipeline — versioning, deployment, monitoring, retraining, and rollback are absent.

Build the MLOps pipeline before scaling model production — versioning, deployment, monitoring, drift detection, retraining, rollback. Treat models as production systems with the same operational discipline as application code. Pilots that cannot graduate to production are not pilots; they are demos.

No AI ROI Framework

No AI ROI Framework is the condition where there is no consistent way to measure whether an AI investment produced business value. Pilots are declared successful based on demo quality. Production AI runs without measurement. Investment continues because nobody can tell whether it should stop.

  • AI investments are approved without defined success metrics.
  • Pilot reviews celebrate model accuracy without measuring business outcome.
  • Production AI is not measured for revenue, cost, productivity, or decisioning impact.
  • "AI is working" is asserted rather than evidenced.

Primarily causes Symptom 9 (AI and Automation Not Producing Business Value).

  1. No success metrics at approval — AI investments are approved without defined success metrics.
  2. Accuracy celebrated over outcome — pilot reviews celebrate model accuracy, not business outcome.
  3. Unmeasured production AI — production AI is not measured for revenue, cost, productivity, or decisioning impact.
  4. Assertion over evidence — "AI is working" is asserted, not evidenced.
  5. No baseline / measurement window — initiatives lack a target outcome, baseline, and measurement window.
  6. No willingness to kill — value-less initiatives are not stopped because nobody can tell they should be.

Establish an AI ROI framework — every initiative defines its target business outcome, baseline, and measurement window before investment. Measure outcomes consistently in production. Be willing to kill AI initiatives that do not produce value. Measurement is what separates AI as an investment from AI as a gesture.

Weak AI Governance

Weak AI Governance is the condition where security, privacy, compliance, auditability, and human-in-the-loop controls are missing or immature. AI failures damage customer trust, create regulatory exposure, and may breach contracts.

  • No policy on what data may be used to train or prompt AI systems.
  • No human-in-the-loop for high-stakes AI decisions (lending, hiring, eligibility, content moderation).
  • Prompt-injection, data-leakage, and model-poisoning risks are not assessed.
  • Customers and regulators cannot be answered when they ask "how does your AI decide?"

Primarily causes Symptom 9 (AI and Automation Not Producing Business Value).

  1. No data-use policy — there is no policy on what data may be used to train or prompt AI.
  2. No human-in-the-loop for high-stakes — high-stakes decisions (lending, hiring, eligibility, moderation) lack human review.
  3. Unassessed AI security risks — prompt-injection, data-leakage, and model-poisoning risks are not assessed.
  4. Unexplainable decisions — customers and regulators cannot be answered on "how does your AI decide?"
  5. No risk-matched governance — governance rigor is not matched to use-case risk.
  6. Governance as afterthought — governance is bolted on rather than built into the AI lifecycle.

Establish AI governance covering data use, model risk, security (including prompt injection), privacy, compliance, auditability, and human-in-the-loop. Match governance rigor to use case risk. Make governance part of the AI lifecycle, not an afterthought.

From Catalog to Diagnosis

Reading About Diseases Doesn't Treat Them.

The 78 entries above are the diagnostic vocabulary. A Growth Blocker MRI Dx™ engagement identifies which of these diseases are active in your company — and prescribes the right sequence of treatments.