Line stoppage analysis in minutes

Turn a line stoppage into a timely, traceable review by joining SCADA alarms, historian context, approved SOPs, and changeover records.

Operator reviewing a stopped production line with alarm, historian, and procedure context

The stoppage investigation begins after the useful part has evaporated. The line stops, an operator records what can be remembered while recovery is under way, and the next shift inherits a note plus a reason code. By the time someone opens the historian the following day, alarms have been acknowledged, the format has changed again, and the people who saw the event are busy elsewhere. The record is not empty. It is simply cold.

That delay leaves the reviewer with data, but without the operating mode, live display, current procedure, or people able to explain the decision under pressure when it mattered most.

Here, real-time means a review assembled minutes after a stoppage, not continuous streaming, while operational context remains available. It joins the SCADA alarm sequence, historian trends, the approved SOP, and the format-change record. It does not infer a root cause from proximity. APQC defines its unplanned machine or equipment downtime measure around disruption of scheduled run time, which is a useful boundary for deciding what the packet is explaining (APQC downtime measure). The packet then gives the next qualified person a defensible place to look.

Capture the context before the shift ends

The familiar end-of-turn reconstruction fails for a mundane reason: important details change state. The operator may recall that a temperature alarm appeared before the stop but not whether it cleared before the first restart attempt. A supervisor may remember a difficult format change but not the exact line mode. A later review can retrieve timestamps, yet it cannot reliably recreate what was visible at the console, which controlled procedure applied, or which workaround was considered and rejected.

HSE describes shift handover as accurate, reliable communication of task-relevant information across shifts or teams, with preparation, exchange, and incoming cross-checking as separate elements (HSE shift handover). That principle applies before the handover itself. The analyst should preserve the event while the outgoing crew can still correct it: last known normal state, first abnormal signal, stop declaration, operator actions, restart attempts, and stable running. A reason code belongs in that record, but it is a classification, not a verdict.

The first useful packet can be short. Capture the asset and product context, the event owner, the applicable time zone, the scheduled run window, the selected downtime reason, and direct links to the SCADA and historian views. APQC excludes setup and changeover from its machine-downtime measure, while the observed stop may still need to be investigated in relation to a changeover (APQC measure scope). Keeping those categories separate prevents a format change from being silently counted as a failure or a failure from being excused as setup.

This is the direct answer for a production manager: do not wait to solve the event. Record enough context to keep the question alive. WizeeMind can assemble the packet and show its sources, but the operator, engineer, or supervisor decides what the record means. That boundary matters when production pressure makes a neat explanation attractive.

Put every signal on one honest clock

SCADA, historians, MES records, electronic logbooks, and changeover sheets rarely arrive with identical clocks or identifiers. Treating their timestamps as interchangeable creates a false sequence. A historian can report a tag in server time, an operator can enter a note in local shift time, and a changeover can be closed after the physical work was completed. When order matters, the packet must show the source clock and any known offset instead of silently normalizing the evidence.

ISA-95 provides an abstract model for information exchange between manufacturing control functions and business functions; it exists partly to give integration projects a shared vocabulary and boundary between systems (ISA-95 purpose and model). NIST similarly describes operational performance measurement as work that needs well-defined methods and standards to collect and analyze manufacturing data (NIST operations-driven measurement). Neither source promises that a plant’s systems are already synchronized. They explain why the packet has to make its mapping explicit.

Start with one anchor: a line-state transition, a common batch marker, or another event that appears in more than one system. Record the native timestamp, timezone, source identifier, asset name, and confidence for each event. Then annotate, rather than overwrite, any offset used for comparison. If no reliable anchor exists, say so and retain a wider window. A wider but honest window is more useful than a precise-looking chronology built from untested assumptions.

This step is also where name mismatches surface. Line 3 in SCADA may be L3, its historian tags may use HX02, and the changeover system may call the equipment a packaging cell. Build an alias map for the packet, not a permanent data model on the fly. The industrial data context guide explains why a tag without asset, time, and operating context remains weak evidence. The immediate decision is simpler: can a reviewer tell which records refer to the same physical process?

Read alarms as a sequence, not a list

An alarm list is often the first evidence available after a stop and the least reliable thing to interpret in isolation. Repeats, clears, acknowledgments, consequence alarms, and standing alarms can all appear together. The loudest group may be downstream of the initiating condition. The right question is not which tag appeared most often. It is which signal first required timely assessment, in which operating mode, and what happened next.

HSE says alarms should direct attention to plant conditions that require timely assessment or action (HSE alarm management). ANSI/ISA-18.2 is the management-of-alarm-systems standard named for process industries (ISA-18.2). Those references justify a careful review of the sequence; they do not let an analyst alter priorities, limits, or suppression settings. Any such change remains in the plant’s authorized alarm and management-of-change process.

For each relevant tag, keep the native alarm timestamp, state transition, priority, acknowledgment time where available, configured response reference, and relationship to line state. Pair it with the historian trend for the same interval. The alarm tells you that a condition crossed its configured boundary. The trend can show whether the variable moved abruptly, drifted, or remained stable while another signal changed. A signal pair can strengthen a hypothesis, but it still does not establish mechanism.

The review should also retain alarms that cleared before the stop. A cleared alarm is a timestamped observation. Removing it from the packet because it is no longer active breaks the sequence.

WizeeMind should therefore label its output plainly: observation, correlation, hypothesis, or confirmed cause. The alarm-management evidence guide uses the same discipline when it reconstructs first abnormal signals. In practice, the packet may say: “The temperature alarm preceded the stop in this event window. The trend rose during the same period. No verified mechanism yet links the rise to a specific changeover action.” That sentence is useful because it identifies a test without pretending the test has passed.

Add the approved SOP and changeover record

The SOP belongs in the evidence packet because an alarm is never interpreted outside an operating context. The applicable revision can define permitted states, expected checks, escalation routes, and restart boundaries for that line and product. It can also settle a basic question: was the equipment in normal production, planned changeover, cleaning, or recovery? That context changes which comparisons are fair.

HSE’s maintenance-procedure guidance treats maintenance procedures and communication between maintenance and production as part of the safety case for work, including fault recognition and marginal performance criteria (HSE maintenance procedures). The relevant local SOP remains controlling; ISA-95 helps distinguish systems. Public guidance cannot replace it. Nor can a tool turn an SOP into permission to make a new parameter adjustment, bypass a hold, or return equipment to service.

Attach the approved document identifier, revision, effective status, scope, and the steps relevant to the operating state. Then attach the changeover record: outgoing SKU, incoming SKU, physical completion time if known, declared completion time, person or role, and any sign-off or deviation. The procedure version control guide explains why a newer-looking file may not be the applicable instruction. This is not paperwork for its own sake. It prevents the team from comparing a temperature trend to the wrong recipe, mode, or procedure revision.

Where a changeover record gives only a completion status, ask whether it represents the physical task, the electronic sign-off, or both. That answer narrows the comparison without adding facts absent from the record.

The decision rule is modest. If the SOP or changeover record is missing, record that as a limitation and route it to the document or operations owner. Do not invent the missing sequence from a generic best practice. If the record shows a prescribed post-changeover check, verify whether that check is documented; do not assume it occurred. That distinction protects the analysis from both hindsight and blame.

Use Case 1 to create a testable question

Case 1 is an anonymized WizeeMind example of the packet working as a timeline, not as a root-cause report. At 06:12, Line 3 changed to SKU-204. At 06:48, ALM_TEMP_L3-HX02 appeared. At 07:15, the line stopped unexpectedly for 23 minutes. Later that day, the line changed to SKU-118 at 14:30, and the same alarm appeared at 15:02. The documented review groups that recurrence as three appearances in seven days and observes a 32-to-36-minute window after format changes.

The arithmetic is worth stating plainly. The first alarm is 36 minutes after the 06:12 change; the second is 32 minutes after the 14:30 change. This is an observed temporal correlation. It does not show that either SKU caused the alarm, that the alarm caused the 23-minute stop, or that a parameter adjustment is safe. It may point to thermal settling, a changeover step, a recipe difference, instrumentation behavior, or an unrelated pattern. The record alone cannot choose among them.

NIST’s operations-driven project describes using system characterization and data analysis to identify performance issues and establish the frame of reference for evaluating a system (NIST system characterization). Douglas Thomas’s NIST report examines the data needed to estimate costs and losses from manufacturing maintenance approaches, including the feasibility of collecting that data (NIST maintenance evidence). Together, they support the discipline here: preserve the data and limits before claiming an explanation.

For Case 1, the next task is specific. A qualified controls and operations review can compare the two changeovers against the approved procedure, recipe and setpoint history, equipment state, historian trend, alarm configuration, and earlier instances in the seven-day window. The output should state whether the pattern persists after those checks. It should never say that a recommendation is a confirmed cause merely because its timestamps look persuasive.

Turn the correlation into work during the shift

The point of a minutes-scale analysis is not a faster postmortem. It is a smaller, assigned check while the people and records are still available. Once the Case 1 window is visible, a supervisor can assign a review before the next comparable changeover rather than asking tomorrow’s meeting to reconstruct it. The work is bounded: confirm the applicable SOP, compare the planned and actual changeover record, inspect the historian window, and capture the alarm context.

HSE’s handover guidance calls for preparation, communication, and cross-checking as responsibility changes hands (HSE handover elements). ISA-95’s information-exchange model is useful for the same reason: the answer spans operations, controls, maintenance, and planning systems (ISA-95 information exchange). Assigning an owner does not erase those boundaries. It makes the next question and required evidence explicit at each boundary.

A good assignment has a window, a question, an owner, and a closeout condition. For example: before the next Line 3 format change, the controls owner compares ALM_TEMP_L3-HX02 against recipe and trend history for the preceding 45 minutes; the operations owner verifies the applicable SOP revision and completion records; the supervisor records whether the 32-to-36-minute pattern reappears. Any proposed parameter, alarm, or procedure change goes through the site’s approval route. No automatic adjustment follows from the packet.

This approach stays useful even when the pattern disappears. A non-repeat is evidence against a simple recurring explanation, not a reason to rewrite history. Close the event with what was observed, what was checked, what remains uncertain, and who accepted the conclusion. For a broader method of preserving stop evidence, see downtime analysis with evidence. The goal is not to make every stop look solved. It is to replace a cold reconstruction with a timely, reviewable question.

Frequently asked questions

What data is needed for a line stoppage analysis?

Use the line state and stop interval, SCADA alarms, historian trends, the applicable SOP revision, changeover records, and operator notes, with each record linked to its source and timestamp. ISA-95 is useful for describing the system boundaries and shared identifiers that make those records comparable (ISA-95 model). Include the downtime classification, but retain the evidence that supports or limits it. APQC’s definition keeps the scheduled-run-time measure separate from setup and changeover categories (APQC definition).

How should SCADA and historian timestamps be aligned?

Record each system clock and timezone, identify a common event marker, disclose any offset, and preserve the source timestamp rather than silently rewriting the sequence. Use a line-state transition, batch marker, or another shared event as the anchor only after checking that it genuinely refers to the same equipment and operating period. NIST notes that performance analysis needs a defined frame of reference and methods for collecting operational data (NIST measurement context). If the offset cannot be established, keep a wider correlation window and say that the order is uncertain.

What does the SOP add to a stoppage review?

The approved SOP defines the permitted operating and recovery context; it does not prove why the stoppage happened or authorize a new adjustment. Link the exact approved revision, its scope, and relevant operating-state steps. HSE places procedures, fault recognition, communication, and competence within maintenance controls (HSE procedure context). A draft, obsolete, or mismatched procedure should be recorded as a limitation and escalated to its owner, not treated as an instruction.

Does a repeating time window prove that a changeover caused a stop?

No. A repeated interval is a correlation that creates a focused verification task; cause requires evidence of a mechanism and review by accountable plant roles. In Case 1, the 32-to-36-minute timing after two changes and three observations in seven days narrows the next comparison, but it does not establish that a format change, an alarm, or a setting caused the stop. HSE’s alarm guidance focuses on timely assessment and action, which is the appropriate immediate response to such a pattern (HSE alarm purpose).