Intermittent quality deviations are awkward because the process can look normal between failures. Averages stay inside familiar bands, most lots pass, and the few affected units may be separated by hours or campaigns. The investigation then drifts toward the last alarm, the most visible operator action, or a correlation calculated across unlike production. None of those is a defensible pattern until affected and acceptable material have been compared under conditions that could reasonably be the same.
The work starts by aligning quality results to the material that experienced the process. Operating context must remain attached while the analyst builds traceable features and checks whether an association survives segmentation and time checks. An intermittent quality-deviation analysis is an auditable comparison of affected and acceptable production units under comparable conditions. Its purpose is not to find the highest coefficient.
For U.S. drug manufacturing, 21 CFR 211.192 requires unexplained discrepancies and specification failures to be thoroughly investigated, with written conclusions and follow-up (21 CFR 211.192). Other industries must apply their own quality and legal requirements.
This article goes beyond the broader five plant historian analyses. That guide helps choose a useful historian question. Here the question is already fixed: why does a quality deviation appear only sometimes, and what evidence would justify the next controlled test? The answer requires more than historian tags. It requires historian, MES, laboratory or quality, raw-material, environmental, and change-log records to meet at the correct lot, campaign, equipment, state, and time for the material.
Define the comparison before collecting variables
Start with the quality response and production unit. The response might be a laboratory value, an inspection category, a defect count, or a pass/fail disposition. The production unit might be a batch, coil, roll, pallet, serialized item, or a bounded continuous-production interval. Do not let timestamp convenience decide the unit. If the quality record belongs to a lot, the analytical table must retain that lot identity.
Write one comparison statement before opening a trend client: “For this product and specification, compare affected and acceptable lots made on the same equipment during equivalent operating states, accounting for material residence time.” That statement names the population. Add campaign when cleaning history, equipment condition, or accumulating material can carry over across lots. Add product grade, recipe, packaging format, or customer specification when any of them changes the expected response.
The first analytical error is usually pooled context. A coefficient calculated across products may capture recipe differences. A line-level average may hide one equipment path. A campaign effect may be mistaken for a raw-material effect because the material lot changed at the same boundary.
ICH Q9(R1) says risk questions should state pertinent assumptions and background information. It also connects the formality of quality risk work to uncertainty, importance, and complexity (ICH Q9(R1)). Those principles apply directly to the comparison definition.
For U.S. drug manufacturing, the comparison boundary also determines what evidence enters the written investigation required by 21 CFR 211.192. FDA guidance likewise recommends an ongoing program to collect and analyze product and process data related to quality. The chosen population and time window should therefore remain reproducible as new production arrives (FDA process validation).
Define affected, acceptable, and indeterminate states separately. “Acceptable” should mean the same test method and applicable specification, not merely absence from a deviation list. An untested lot is not a good lot. A retested result should not silently replace the original result. If an investigation later changes disposition, preserve both the original observation and the controlled decision.
Set the time boundary from process knowledge. Include enough history to cover relevant campaigns, material changes, maintenance, seasonal conditions, and known acceptable operation. Avoid a fixed rule such as thirty days. A short high-throughput line and a seasonal batch process have different evidence windows. State what the window excludes and why.
This boundary matters because quality decisions may affect product safety, release, and regulatory compliance. The article does not provide release, legal, or process-change advice. In regulated settings, accountable quality and engineering roles must apply the approved procedure and jurisdiction-specific requirements. ICH Q9(R1) also says quality risk management does not remove the obligation to comply with regulatory requirements (ICH Q9(R1)).
Build one lot-aligned evidence table
A defensible analytical record has one row per quality-bearing production unit and enough provenance to reconstruct every value. Keep identifiers for lot, campaign, product, equipment path, operating state, quality method, specimen time, and result. Add the source system, query or export version, extraction time, units, and any conversion applied. A wide table may be convenient for analysis, but lineage should remain available in a companion dictionary or long-form feature table.
Combine six evidence families. Historian data supplies process values and equipment states. MES supplies order, route, recipe, lot genealogy, phase boundaries, and production events. Laboratory or quality systems supply specimen identity, method, result, specification, retest, and disposition. Material records supply supplier and raw-material lots. Environmental systems supply conditions relevant to the process. Change logs supply setpoint edits, recipes, calibration, maintenance, software, procedures, and temporary instructions.
Wayne Matthews describes modern historians as receiving data from control and monitoring, laboratory, enterprise, and asset systems (ISA on historian inputs). That breadth does not make a historian record the master for every field. Record which system owns each identifier and how conflicts are resolved. A laboratory specimen time and a result-entry time, for example, answer different questions.
Align the specimen to the material, not to the moment the result became visible. ISA measurement guidance says laboratory results should retain the time the specimen was taken because the analysis may occur later, and it connects correct specimen timing with comparison to online measurements (ISA measurement timing). Then account for residence time between the measured process location and the specimen or quality observation.
Residence time may be fixed, estimated from flow and inventory, defined by batch phases, or represented as a plausible interval. Document the method and uncertainty. For queues, recirculation, blended vessels, and parallel paths, a single offset may be false precision. Use an interval, path-specific logic, or exclude the feature until engineering can justify the mapping.
FDA process validation guidance recommends an ongoing program to collect and analyze product and process data related to product quality. It names incoming materials, in-process material, finished products, relevant process trends, and statistically reviewed data (FDA process validation). The guidance applies to drug manufacturing, but the data categories offer a disciplined cross-check for any plant evidence table.
Preserve bad data as status, not as a quiet deletion. Useful flags include missing, stale, substituted, manually entered, out of service, compressed, outside calibration, and timestamp corrected. Keep the raw extract read-only. Store the join logic, timezone, daylight-saving treatment, duplicate policy, unit conversion, and exclusion list. Someone reviewing the investigation should be able to reproduce why Lot A received one feature value and Lot B received another.
Finish with a row-count reconciliation. Count source quality units, unmatched units, duplicate matches, excluded units, and final analytical rows. Split those counts by affected and acceptable status. A model built on a clean-looking subset can be misleading if most affected units failed to join.
Engineer features that another reviewer can audit
Raw historian points rarely match the mechanism under investigation. Convert them into features tied to engineering hypotheses: phase mean, range, standard deviation, slope, dwell time, time above an approved threshold, number of state transitions, peak, time since cleaning, or difference between upstream and downstream measurements. Do not create hundreds of anonymous aggregates first. Each feature needs a name, physical meaning, source tags, time window, aggregation rule, units, missing-data rule, and version.
Use process phases and material windows. A batch temperature mean across charging, reaction, hold, and discharge may erase the only relevant variation. Calculate phase-specific values where the process definition supports them. For continuous production, shift the window by residence time and test nearby plausible lags. Label an exploratory lag sweep as exploratory; choosing the lag with the strongest correlation after inspecting many lags inflates the chance of a chance finding.
Range and variability often matter more than the mean. Two lots can have the same average temperature while one experienced oscillation or a short excursion. Preserve the minimum and maximum, but also consider robust spread, rate of change, time outside a defined operating region, and controller output behavior. Do not confuse a specification limit with a statistical threshold or a validated process range.
Time dependence changes the effective evidence. NIST explains that an autocorrelation plot compares observations at different lags and that many standard statistical conclusions depend on the randomness assumption (NIST autocorrelation plot). Autocorrelation means adjacent historian values do not provide the same independent information as equally many separated production units. Check autocorrelation in raw features and, more importantly, in model residuals.
If the same sensor scan contributes hundreds of near-duplicate points to one lot, do not let that lot dominate. Aggregate at the production-unit level or use a time-series method that represents dependence. Report the number of lots or units as the analytical unit count, alongside the number of raw points. The two counts answer different questions.
Change-derived features need the same discipline. “After maintenance” is too vague. Store the change record ID, affected asset, effective timestamp, approved scope, and the first production unit exposed. For calibration or sensor replacement, preserve the old and new configuration. For raw material, keep genealogy across blends and substitutions. For environment, use the condition at the process step where it could matter, not a daily site average chosen because it is easy to export.
Feature engineering should end in a reviewable catalog. A quality engineer checks the response and disposition logic. A process engineer checks mechanism, phase, lag, and units. The data owner checks lineage and timestamps. A statistician or trained analyst checks dependence, multiplicity, and model assumptions. FDA process validation guidance recommends that personnel with adequate statistical process control training develop the data collection plan and methods used to evaluate process stability and capability (FDA process validation).
Test whether the pattern survives context
Begin with plots, not a ranked correlation table. For each plausible feature, plot the quality response against the feature and encode affected status, product, equipment, campaign, and state. Show raw points. A fitted line without the points can hide clusters, gaps, nonlinear shapes, and one influential lot. NIST says scatter plots can expose linear or nonlinear relationships, changing variation, and outliers (NIST scatter plots).
A scatter plot establishes association, never causation. NIST states that association does not imply causality and that a scatter plot cannot prove cause and effect (NIST scatter plots). Correlation never authorizes a process change. It can justify a better question, a data check, or an approved test.
Next, condition the relationship. Replot within product, equipment, operating state, campaign, and meaningful ranges of a third variable. Calculate range-conditioned correlations only when each range has enough comparable units and was not chosen to manufacture a result. NIST defines a conditioning plot as a view of two variables within groups of a third variable and notes that it can reveal whether the relationship depends on that third variable (NIST conditioning plots).
Segmentation makes confounding easier to spot. A pooled positive correlation may weaken, disappear, or reverse within each product. A material lot may appear associated with failure because it ran only on one equipment path. Ambient humidity may track a seasonal recipe schedule rather than a physical moisture mechanism. The segmented result is not automatically causal either. It shows which explanations remain compatible with the records.
Check time lags against process knowledge, then test sensitivity. If the feature-result association exists only at one implausible lag, treat it skeptically. If several nearby, physically reasonable lags give similar results, report the range rather than the most favorable coefficient. Keep holdout campaigns or later production for confirmation when data volume permits. Never call the same rows discovery and confirmation.
Practical multivariate work can start with multiple regression for a continuous response or logistic regression for a binary deviation. Include variables because the process model supports them, not because automated selection retained them. Inspect residuals, nonlinear terms, interactions, correlation among predictors, influential units, missing-data patterns, and performance by product and equipment. Report uncertainty and error modes, not just fit.
Principal component methods can summarize correlated process features, and partial least squares can relate many correlated predictors to a response. Their scores and loadings still need engineering interpretation. NIST lists principal components, clustering, and classification among multivariate techniques for identifying correlation structure, while noting that its handbook covers multivariate analysis only lightly (NIST multivariate overview). A latent component is a screening aid, not a root cause.
Use resampling or validation splits that respect time and campaign boundaries. Randomly splitting adjacent rows can leak nearly identical conditions into training and test sets. Compare a multivariate model with a simple segmented baseline. If the complex model adds little or fails on a later campaign, prefer the simpler explanation. The analysis must remain repeatable and useful to the investigation.
Move from exploratory work to a controlled decision
Case 4 is an illustrative scenario only. A plant sees intermittent quality deviations across several campaigns. A trend review suggests that a process feature and a material attribute move with the result, but only on one equipment path and within one operating range. No plant measurements, coefficients, outcomes, customer evidence, or causal conclusion are claimed in this scenario.
The team first verifies lot genealogy, specimen timing, residence-time alignment, sensor status, and the effective dates of material and maintenance changes. It then reproduces the scatter plots by product, equipment, state, and campaign. The pooled association weakens after segmentation, while one range-conditioned pattern remains. That remainder becomes a hypothesis about a mechanism. It does not become an instruction to change the process.
The tool should follow the maturity of the method. Use the historian trend client to inspect raw timing, states, and anomalies. Use versioned SQL to extract and join repeatable tables. Excel is suitable for bounded filters, pivots, scatter plots, and reconciliations. Python or Jupyter becomes useful for audited feature pipelines, lag sensitivity, diagnostics, and multivariate reruns. Publish in BI only after definitions, ownership, refresh behavior, and exclusions are stable.
Do not turn the first notebook into an automatic decision system. Freeze the input version, feature catalog, analysis code, package versions, and output. Save plots with the filters and unit counts. Link every candidate association to contradictory evidence and missing checks. Record who reviewed the data mapping, statistics, process mechanism, quality impact, and proposed test.
An approved test should distinguish the hypothesis from credible alternatives while staying inside safety, quality, validated-state, and change-control boundaries. NIST separates correlation from causality and points to designed experiments when the aim is to establish causal relationships (NIST experimental design). The design, operating window, measurements, stopping rules, product disposition, and approvals depend on the plant’s procedures and risk.
FDA process validation guidance says unexpected observations and manufacturing nonconformances should be evaluated and that reports should describe corrective actions or changes in sufficient detail, with appropriate review and approvals (FDA process validation). ICH Q9(R1) calls for documented, transparent, reproducible risk methods and continuing review when new knowledge or failure-investigation results arise (ICH Q9(R1)).
The operational chain ends here: association → hypothesis → approved test → controlled change (NIST experimental design). Preserve the conditions and data when a test refutes the hypothesis. If a test supports it, controlled review is still required before implementation. Correlation, statistical significance, model accuracy, or repeated visual similarity never supplies change authority.
Frequently asked questions
How many good and bad lots do I need?
There is no universal minimum. Use all comparable affected and acceptable units, disclose the resulting unit count, and treat sparse groups as leads rather than stable estimates.
Treat the production unit as the evidence count. Hundreds of time-adjacent sensor values do not turn a small number of lots into a large independent evidence base. NIST notes that many standard statistical conclusions depend on randomness, so check autocorrelation and state both the lot count and raw-point count (NIST autocorrelation plot).
Can a strong correlation identify the root cause?
No. A strong correlation is an association that may support a hypothesis. Confounding, time dependence, measurement error, or control action can produce the same pattern.
NIST makes the limit explicit: association does not imply causality, and a scatter plot cannot prove cause and effect (NIST scatter plots). When causal evidence is needed, NIST points to designed experiments (NIST experimental design); plant-specific approval, safety, quality, and change-control limits still govern the test.
For quality-risk work, ICH Q9(R1) separately calls for documented, transparent, reproducible methods. A coefficient cannot replace that review (ICH Q9(R1)).
Should I remove startup and changeover data?
Do not delete them by default. Label operating state, analyze it separately where appropriate, and document any exclusion with its reason, owner, and effect on the result.
Startup and changeover may define meaningful operating states rather than noise. ICH Q9(R1) says risk questions should state pertinent assumptions and background information, so keep the state label and make any exclusion reviewable (ICH Q9(R1)).
FDA’s continued process-verification approach includes relevant process trends, another reason to preserve operating state before comparing units (FDA process validation).
When should I move the analysis from Excel to Python?
Move when joins, lag calculations, feature versions, repeated reruns, or model diagnostics become difficult to reproduce and review reliably in the workbook.
What permits a process setting to be changed?
Only the plant’s approved test and change-control process can authorize a change. Correlation, a dashboard, or a model score cannot provide that authority.
FDA process-validation guidance says unexpected observations and manufacturing nonconformances should be evaluated, and reports should describe corrective actions or changes with appropriate review and approvals (FDA process validation). That documented route, not the analytical output, carries the decision.
If the aim is causal evidence, NIST points to designed experiments; the plant’s approved procedure still defines the design and its limits (NIST experimental design).