Industrial digitalization rarely fails with a dramatic crash. More often, a pilot works in a demonstration, earns polite approval, and then stalls when it meets shift work, mixed equipment, awkward records, cybersecurity controls, or a manager whose attention has moved elsewhere. The software may still function. The project has failed if it cannot scale beyond the pilot or cannot sustain operational value after the launch team leaves.
That definition is intentionally practical. It separates a promising prototype from a capability the plant can own. It also avoids treating every cancelled experiment as waste: a disciplined pilot can stop early and still succeed by disproving an assumption before the organization spends more. The costly failures are harder to spot. They keep consuming licenses, integration effort, operator patience, and management time while nobody can show that a recurring plant decision has improved.
The five mistakes below are not laws of engineering. They are recurring mechanisms to test. Industry, site maturity, risk, regulation, and the problem itself can change the answer. Any decision involving process changes, cybersecurity, safety, quality release, capital approval, or financial commitment remains with the plant’s authorized roles.
What failure means after the pilot demo
A working demo proves that something worked once under the conditions of the demo. It does not prove that the same result will survive another line, product family, shift pattern, source-system change, or quarter of ownership.
For this guide, industrial digitalization has failed when the pilot does not scale to its intended operating boundary or when the organization cannot sustain the operational value it was meant to create. A pilot can therefore meet its technical acceptance test and still fail as an operating change.
The World Economic Forum used the phrase “pilot purgatory” for this gap.
Its 2019 Lighthouse work reported that more than 70% of businesses investing in technologies such as big-data analytics, artificial intelligence, or 3D printing had not moved beyond the pilot phase (WEF Beacons report). That figure came from the Forum’s earlier work and is not a universal failure rate for every sector, geography, project type, or year.
It is useful here as evidence that scaling is a separate problem from proving a technology in one setting.
Judge the pilot against an operating contract. Name the problem, the population in scope, the current baseline, the decision or task that should change, and the person accountable for the result. Then name the evidence window. A reduction in investigation time may need several comparable events; a scheduling aid may need multiple product mixes; a maintenance workflow may need enough repetitions to show whether people keep using it after the novelty fades. One successful workshop is not a sustained outcome.
Scaling also has more than one dimension. Technical scaling asks whether interfaces, performance, security, and support work outside the test environment. Operational scaling asks whether the output fits real work and exceptions.
Organizational scaling asks whether ownership, training, incentives, and approval routes still function when the project team is absent. Economic scaling asks whether the repeatable benefit exceeds the full recurring burden, not merely the initial license. A pilot is ready to expand only when the dimensions relevant to its risk are evidenced together.
A 42-project manufacturing study likewise describes different combinations of social and technical conditions rather than one standard technology journey (Clausen et al., 2025). It is not a universal scoring model.
Warning signs appear early: the team reports model accuracy but no changed decision; every new line needs bespoke mapping; operators keep a parallel spreadsheet; data cleaning depends on one analyst; the benefit disappears when demand or product mix changes; or nobody owns the feed after handover. Prevention starts with scale criteria written before configuration. Define what would count as stop, repair, limited continuation, or expansion. Keep the original process available during a controlled comparison where risk requires it.
There is an important limit. A first pilot should not be expected to prove the economics of an entire network. Some infrastructure benefits emerge only after reuse. The honest response is to separate demonstrated local value from a portfolio hypothesis, state the assumptions behind reuse, and fund the next boundary as another test. Calling the hypothesis “transformation” does not make it evidence.
Mistake 1: choosing technology before the problem
Technology-first projects begin with a capability looking for a plant problem: a sensor package, computer-vision model, data platform, digital twin, or generative assistant has already attracted attention, so the team is asked to find a use. The mechanism is subtle.
Because the product defines what can be seen, the project frames success around deployment milestones or tool output. The plant’s delayed decision, repeated loss, or manual burden becomes supporting scenery. A good demo then validates the purchase without showing that the original work improved.
Michael Brundage describes a manufacturer that bought sensors before deciding how the new data would be used, while useful maintenance knowledge was already present in difficult-to-analyze work orders (NIST, “How Do We Get Smart?”). The example does not mean sensors were inherently wrong. It shows why a data inventory should precede an assumption that the plant lacks data, and why the first question should be operational.
The warning signs are familiar. The scope is a noun such as “AI” or “platform” rather than a decision. The success metric is devices connected, dashboards released, or users trained. No owner can state the current cost or delay of the problem. The team cannot say which action may change after the new output appears. Requirements expand whenever the tool reveals another feature. Meanwhile, a spreadsheet, shift conversation, or existing maintenance record continues to carry the work that matters.
Prevention starts with a one-page problem contract. Write the recurring event, who encounters it, how they handle it now, what evidence they lack, what decision remains human, and how the outcome will be compared.
The smallest credible use case may be unglamorous: reconstruct one stoppage window, match work-order language to asset aliases, or show why one production report needs manual reconciliation. The guide to industrial data context explains why records need relationships, ownership, and scope before an assistant can assemble them responsibly.
Next, test three alternatives against the same problem: improve the existing method, add a narrow digital aid, or introduce the larger platform capability. Estimate implementation and recurring effort for each with finance, IT/OT, and plant owners. This is a screening comparison, not investment authorization. A rough estimate can eliminate an obviously mismatched concept; it cannot approve capital or establish a safe design.
An eight-manufacturer study placed data-driven use cases alongside strategy, organization, collaboration, and cross-functional work (Budde et al., 2022). Even a sound technical use case needs an organizational route.
The limit is that technology scouting still has value. Teams sometimes discover a capability before they understand its best application. Keep that work explicitly exploratory, time-boxed, and separate from a value claim. The mistake is not curiosity. It is allowing curiosity to inherit a budget, production dependency, or scale promise before a real problem and owner exist.
Mistake 2: designing without the plant
Excluding the plant does not always mean excluding every operator from every meeting. A project can hold interviews and still treat the shop floor as a source of requirements rather than a co-designer of changed work.
The failure mechanism appears when formal process maps replace local reality. The design misses how people handle abnormal modes, interrupted tasks, unofficial asset names, glove use, shared terminals, radio calls, temporary repairs, or the moment when a supervisor must choose between recording detail and restoring flow.
The World Economic Forum’s 2024 front-line report was based on more than 85 interviews with operators, mechanics, electricians, manufacturing engineers, and supervisors in eight factories across the United States, Europe, and Asia (WEF front-line report). The participating sites belonged to large international corporations, so the findings should not be treated as a representative survey of all manufacturers. Their value lies in showing concrete worker concerns across the preparation, introduction, and continued use of technology.
Watch for a pilot that works only when the project engineer stands beside the user. Other warnings include late requests for operator feedback, training scheduled after workflow decisions are fixed, low use on nights or weekends, shared accounts, paper notes kept “temporarily,” and complaints dismissed as resistance. A more serious sign is silence. People may stop reporting defects if earlier feedback produced no explanation, no visible change, or blame for slowing the rollout.
Prevention means involving the roles that perform, supervise, maintain, and depend on the task before the interface is fixed. Ask them to walk through a normal case and an ugly one. Observe the current workflow where local rules permit it. Let users challenge the asset names, timing, handoffs, exceptions, and explanation of why the change is being made. The WEF report found that workers’ perspectives are often overlooked even though they are essential to effective technology introduction (WEF front-line report).
Make plant participation consequential. Record decisions changed by the review, unresolved objections, and the owner of each response. Run shadow mode across representative shifts. Measure completion, rework, unanswered cases, time away from the physical task, and whether the new record is trusted downstream. Involve the IT and OT owners as well; the practical IT/OT integration guide explains why a useful connection also needs a safe route, field authority, timestamps, and maintainers.
The eight-company study identified a combined top-down and bottom-up approach in its case synthesis (Budde et al., 2022). Its qualitative design does not prescribe one governance structure.
Participation has limits. It does not transfer safety, quality, cybersecurity, labor, or engineering authority to an informal workshop, and one vocal user does not represent every shift. Some design choices remain constrained by regulation or controlled procedures. Explain those boundaries plainly. People can accept a constraint more readily than a consultation whose outcome was decided in advance.
Mistake 3: treating the vendor story as the business case
A vendor narrative describes what a product can do and, at its best, shows how another customer used it. A business case must explain what this plant expects to change, from which baseline, through which mechanism, at what total burden, and under whose accountability. When the first substitutes for the second, the project borrows someone else’s problem, economics, data readiness, and implementation conditions. The benefit may be real elsewhere and still be irrelevant here.
This usually starts with a headline percentage or polished case study. That number enters a slide, is multiplied by local production or labor cost, and becomes expected value. Missing from the calculation are the comparison boundary, product mix, adoption path, integration work, validation, support, retraining, cybersecurity, downtime for change, and the cost of keeping the source data usable. Upside compounds while uncertainty disappears.
Warning signs include benefits stated without a local denominator, savings counted across all hours when the problem occurs only occasionally, labor time presented as cash release without a staffing decision, and a return calculation that excludes internal plant effort.
Another is the absence of a counterfactual: what would happen if the plant improved the existing process, repaired a known reliability issue, or did nothing for one more quarter? If the only alternative is the vendor proposal, the case is a sales comparison rather than a plant decision.
Research on eight manufacturing companies in Western Europe identified strategy and organization, collaboration, cross-functionality, and data-driven use cases as four aggregated managerial practices for digital transformation (Budde et al., 2022). The study used qualitative case analysis and described productivity effects qualitatively, so it does not supply a universal financial model. It does support a narrower conclusion: technology selection sits inside managerial and organizational practice rather than replacing it.
Build the business case from a traceable baseline. State the affected process and decision, frequency, current time or loss, measurement source, known uncertainty, and owner.
Describe the causal chain in plain language: the tool changes this task; that change affects this operational measure; finance recognizes value only under these conditions. Include one-time and recurring costs, internal effort, maintenance of mappings and models, training, controls, and the cost of failure or reversal.
The article on plant investment decisions provides a fuller evidence packet for comparing a constraint, alternatives, and success measures.
Set stop conditions before approval. If adoption remains below the agreed operating threshold, data exceptions exceed what the team can maintain, or the measured outcome does not move after a credible window, the sponsor should review whether to repair, narrow, or stop. Those thresholds must be chosen locally; this article does not authorize an expenditure or prescribe a financial hurdle rate.
The WEF Lighthouse report examined selected sites that had moved use cases beyond pilots (WEF Beacons report). These advanced sites are not a benchmark for an average plant.
Vendor evidence remains useful. It can reveal implementation patterns, integration demands, and questions worth testing. Treat claimed benefits as hypotheses until the source, method, population, and local transfer conditions are understood. A transparent vendor can help build that test. It still cannot own the customer’s business case.
Mistake 4: discovering too late that the data is unusable
“We have the data” can mean that values exist somewhere. A pilot needs more: the records must represent the right objects and time windows, remain accessible under approved controls, preserve units and lineage, and survive the joins required by the decision. Discovering those gaps after a model or dashboard has been promised creates pressure to hide exclusions, accept manual patches, or reduce the question until the output looks complete.
Trouble usually begins with a source list instead of a data test. A historian, MES, ERP, laboratory system, or CMMS is marked available. Nobody yet checks whether asset aliases match, clocks align, manual states are recorded, product genealogy survives rework, or sufficient comparable events exist. The first end-to-end join happens after architecture and licenses are committed. By then, every missing record feels like an implementation problem rather than evidence that the use case may not be ready.
A 2024 IFAC case study at a Swiss engine-component manufacturer examined deviations in manually collected assembly-line data from technical and behavioral perspectives (Thurnheer et al., 2024). It developed a model that compared manually collected records with planning-system data to identify discrepancies. This is one case in manual assembly, not proof that all manual records are inaccurate or that planning data is automatically authoritative.
Warning signs include timestamps with unknown timezone, fields whose owner cannot define them, changing units, generic reason codes, unexplained gaps, overwritten corrections, and matches based only on nearest time. If the pilot excludes most failure events because they do not join cleanly, the remaining dataset may tell a reassuring story about the easy cases. A sophisticated model cannot restore events that were never recorded or determine which conflicting source had authority without a rule.
Prevention is a thin-slice data proof before solution design. Select a small but representative set of normal, abnormal, and ambiguous events. Trace each from source to proposed output. Reconcile counts, identifiers, timestamps, units, operating state, access rights, and quality flags. Preserve raw values beside transformations. Record unmatched and excluded cases instead of deleting them. Assign an owner to every mapping and a response when a source changes.
Then test fitness for the decision. Data can be accurate for accounting and unsuitable for second-by-second process analysis. Historian values can be precise and lack material genealogy. A maintenance closure can confirm administrative completion without proving the equipment returned to a required condition. The industrial data context guide and the IT/OT integration guide show how to keep source authority and system boundaries visible.
In NIST’s manufacturer example, difficult maintenance text became useful when the team analyzed variants in technicians’ language (NIST, “How Do We Get Smart?”). “Unusable” is a decision-specific finding.
The limit matters most for AI. An assistant can find gaps, compare records, or draft a traceable exception list; it cannot manufacture trustworthy evidence that the plant never captured. The guide to plant AI capabilities and limits keeps those outputs inside human review and approval. If data repair would require a controlled process, security, or quality change, authorized plant roles must design and approve it.
Mistake 5: confusing budget approval with sponsorship
Budget approval answers a narrow question: may the organization spend within an approved envelope and purpose? Sponsorship is continuing work. A sponsor protects the operational problem when priorities compete, assigns decision owners, resolves cross-functional conflict, insists on evidence, and accepts the responsibility to stop or scale. Treating the signature as sponsorship leaves the project funded but politically homeless.
The gap appears after kickoff. The sponsor delegates attendance, the steering meeting becomes a status recital, and unresolved issues circulate between plant, IT, procurement, finance, cybersecurity, and the vendor. Nobody has enough authority to settle ownership or narrow scope. The project team responds by optimizing what it can control: configuration, milestones, and presentation. Operational value becomes a future dependency.
A study of 42 digital manufacturing projects conducted from 2018 to 2023 in one large wind-turbine manufacturer found diverse sociotechnical configurations, with social variables contributing to project success (Clausen et al., 2025). Its single-company design limits generalization, but the project-level comparison is a useful warning against treating technical sophistication as the only explanation for outcomes.
Look for decisions with no named due date or owner, repeated requests to “align offline,” plant time promised but never released, benefits owned collectively, and risks escalated only when a milestone slips. Another sign is a sponsor who can describe the platform but not the operational problem, baseline, or stop condition. Money may still be available while the people needed to change the work are measured on competing priorities.
Prevention starts with an explicit sponsor contract. Name which conflicts the sponsor will resolve, which evidence they will review, how often, and which decisions they cannot delegate. Pair that role with an operational owner who lives with the result and a technical owner who maintains the capability. Give finance, IT/OT, cybersecurity, quality, safety, and workforce representatives defined gates where their authority is relevant. Do not create a committee in which everyone can object and nobody can decide.
Use a short decision record at each gate: the claim being tested, evidence observed, unresolved limits, owner, next boundary, and decision to stop, repair, continue, or scale. Budget status belongs in the record, but it is not the decision itself. The process prevents a sunk-cost story from replacing current evidence. It also gives the sponsor a legitimate way to end a weak pilot without calling every experiment a failure.
The eight-manufacturer study describes top management coordinating knowledge exchange, standards, and conflicts between autonomy and standardization (Budde et al., 2022). Those are examples of active sponsorship, not a universal structure.
The limit is organizational scale. A local sponsor may not control enterprise architecture, labor agreements, regulated quality systems, or capital allocation. Those constraints should be mapped before the pilot promises a rollout. Escalation is part of sponsorship; bypassing authority is not. No sponsor can waive technical, security, financial, safety, or quality approval simply by supporting the project.
Before the next steering meeting, ask for one page that joins the five tests: the operating problem, plant participation, local business case, proven data path, and active sponsor decisions. If one is missing, reduce the next commitment to the smallest step that can produce that evidence. Do not scale ambiguity.
Frequently asked questions
Why do industrial digitalization pilots fail to scale?
They often fail to scale because a local demonstration has not proved a repeatable operating result, an owned data path, adoption by plant roles, or a credible case for ongoing support. A demo can validate a feature while leaving mappings, shift exceptions, training, security, maintenance, and economics unresolved. WEF documented “pilot purgatory,” while a 42-project study examined social and technical conditions around success (WEF Beacons report; Clausen et al., 2025). Define scale criteria before configuration.
How should a plant choose its first digitalization use case?
Choose one recurring operational problem with a named owner, a measurable baseline, an explicit decision boundary, and records that can be tested before selecting the technology. NIST’s sensor example and eight-manufacturer research support starting from the use and its management system (NIST, “How Do We Get Smart?”; Budde et al., 2022). Prefer a question that can run in shadow mode. “Explain these repeated short stops” is a use case; “implement AI” is not.
Who should be involved in an industrial digitalization project?
Include the people who perform and supervise the work, the process and maintenance specialists, data and IT/OT owners, cybersecurity, and the roles accountable for quality, safety, finance, or change approval where relevant.
The WEF worker interviews explain the shop-floor case; the 42-project study supplies a project-level sociotechnical lens (WEF front-line report; Clausen et al., 2025). The exact group depends on scope and risk. Participation should affect design decisions, exception handling, and acceptance criteria; it should not blur formal authority or allow one shift’s experience to stand for the whole site.
What should a business case say beyond vendor benefits?
It should state the baseline, value mechanism, total operating burden, assumptions, alternative explanations, stop conditions, accountable owner, and evidence required to verify value after deployment.
The managerial-practices study warns against isolating technology from organization, while the WEF Lighthouse work offers evidence from selected at-scale sites rather than a ready-made forecast (Budde et al., 2022; WEF Beacons report). Separate time saved from cash released and a local observation from a repeatable forecast.
Vendor cases can supply questions and implementation evidence, but local finance and operational owners must decide whether the assumptions transfer. This guide does not set an investment threshold.
What makes industrial data usable for a pilot?
The data must fit the decision: identifiers, timestamps, units, operating states, lineage, access, quality flags, and enough comparable history must survive the joins used by the pilot.
The IFAC assembly case compared manual and planning records, while NIST showed that difficult work-order text could still contain useful maintenance patterns (Thurnheer et al., 2024; NIST, “How Do We Get Smart?”). Test normal and awkward cases end to end, reconcile exclusions, and keep transformations reviewable.
A dataset may be valid for its source system and still be unsuitable for a different operational question. Missing evidence should remain visible rather than being filled by inference.
Is budget approval the same as executive sponsorship?
No. Budget approval releases money; sponsorship supplies continuing authority to resolve priorities, assign owners, remove organizational blockers, review evidence, and stop or scale the work.
The project study connects success to sociotechnical configurations, and the managerial-practices research describes active coordination across organizational levels (Clausen et al., 2025; Budde et al., 2022). The sponsor should know the operational problem and the criteria for the next decision.
Sponsorship does not override controlled approvals, cybersecurity rules, safety responsibilities, quality authority, labor obligations, or finance governance.
Sources
- World Economic Forum: Views from the Manufacturing Front Line
- World Economic Forum: Fourth Industrial Revolution Beacons of Technology and Innovation in Manufacturing
- NIST: When a Manufacturer Asks “How Do We Get Smart?”
- Budde et al.: Managerial Practices for the Digital Transformation of Manufacturers
- Clausen et al.: Why project success in manufacturing digitalization remains elusive
- Thurnheer et al.: Manual Data Collection in Assembly Lines