Line Stop Escalation and Recovery

A line stop is a decision as much as an event. Continuing to build while a problem is unresolved produces suspect product that has to be sorted later; stopping the line loses output and may lose the shift. The framework that makes the decision consistent is a set of triggers, an escalation path and a defined set of criteria for restarting, all of them agreed before the event rather than during it.

Line stopped with an andon signal showing

Triggers for Stopping

A trigger is a condition that requires the line to stop or to be formally escalated. The obvious ones are a safety issue, a machine fault that produces scrap and a defect that cannot be contained by inspection. Less obvious but equally important are a repeated defect at the same station within a short period, a material problem such as a wrong part in a feeder and a process parameter outside its window with no clear cause.

The triggers should be specific enough to be applied without judgement. Three defects of the same type within an hour, a safety interlock that has been bypassed, or a measurement outside the control limit are all conditions that an operator can recognise. Where the trigger requires interpretation, it will be interpreted differently on each shift. Shift handover should carry the state of any trigger that was close to being met.

The Escalation Path

The escalation path should name the roles rather than the people, with a response time for each level. The operator escalates to the supervisor, the supervisor to the process engineer, the process engineer to the engineering manager, and each level has a defined period in which to respond. Where the response does not arrive, the next level is engaged automatically rather than by a further request.

The path should also state who has the authority to stop the line and who has the authority to restart it. In many factories the authority to stop is clear and the authority to restart is not, which produces a line that stays down longer than necessary while decisions are sought. Naming both, and communicating them, shortens the recovery without weakening the control.

Supervisor and engineer discussing at a halted line

Containment

Containment begins the moment the problem is recognised. The units built since the last known good point are suspect, and they should be identified and held rather than allowed to continue down the line. The last known good point is usually the last inspection or test that the product passed, which is why the traceability record matters: without it, the suspect range has to be estimated generously.

Containment also covers material. A feeder loaded with the wrong part, a reel of suspect components or a batch of boards with a known issue should be quarantined at the same time as the units, because the material will otherwise be used again on the next build. The containment record should state what was held, how much and where. Batch records are the basis for defining the range.

Diagnosis Under Pressure

The diagnosis should follow the standard route: establish what changed, when it changed and what else changed at the same time. The last change is usually the cause, and it may be a material lot, a setup adjustment, a maintenance activity or an environmental change. Recording changes as they happen, even small ones, is what makes this question answerable in minutes.

Where the cause is not immediately clear, a bounded experiment is better than a series of adjustments. Changing one parameter, observing the result and recording it keeps the investigation traceable; changing several at once may restore the line and leave the cause unknown, which guarantees that the problem returns. Yield data should be captured during the investigation so that the recovery can be assessed afterwards.

Restart Criteria

Restarting should require evidence rather than confidence. The typical criteria are that the cause has been identified, the corrective action has been applied, a first article or a verification build has passed, and the containment has been completed. Where the cause cannot be identified but the problem has stopped, the restart should be conditional, with additional inspection for a defined period.

The conditional restart should specify what is inspected, by whom and for how long, and what evidence would allow the condition to be lifted. A conditional restart that is never closed becomes the permanent process, which is usually worse than the problem it was meant to manage, because it consumes capacity without being visible in any plan.

Communication

Communication during a stop should be short, frequent and factual: what has stopped, what is known, what is being done and when the next update will be. Silence during a stop is filled with speculation, and speculation consumes the attention of people who could be helping. A single channel and a defined update interval is usually enough.

After the stop, the communication changes to a summary: the cause, the corrective action, the affected units, the disposition and the change to the standard. That summary is what prevents the same event from being rediscovered, and it should be circulated to the shifts that were not present rather than only discussed by the people who were there. Inspection standards may need updating as a result.

Recovery and Catch-Up

Recovery should be planned rather than improvised. The queued work, the affected units and the maintenance that was deferred all compete for the same time, and the pressure to catch up creates the risk that the recovery produces a second problem. A short plan with priorities, agreed at the time, prevents the recovery from becoming a source of defects.

Catch-up should not begin until the containment is complete and the restart criteria are met, because a line that catches up with a bad process produces more of the same problem. Where the schedule pressure is severe, the honest options are additional capacity, a revised commitment or overtime, and those are management decisions rather than production ones.

Learning from the Stop

Every stop above a defined threshold should produce a short review: what happened, why it was not detected earlier, what would prevent it and what has been changed. The review should be brief and should end with actions rather than with observations. Where the same stop recurs, the original review was incomplete.

The reviews are also a measure of the escalation system. Frequent stops that are resolved quickly suggest that the triggers are working; few stops but long durations suggest that the escalation is late or that the restart decision is unclear. Both are visible in the downtime data, provided the stops are recorded with their duration and their cause, which is the first requirement of improving either.

FAQ

Who decides to stop the line? Anybody who observes a trigger, with the escalation path defining the response. Hesitation to stop is more expensive than a short stop.

How is the suspect range defined? From the last known good point, using the traceability record. Without it, the range must be estimated generously.

What is needed before restarting? An identified cause, a corrective action, a verification build and completed containment. Where the cause is unknown, a conditional restart with additional inspection.

How do we prevent a recurrence? By reviewing every significant stop, changing the standard and communicating the change to all shifts.

Leave A Comment