Weibull Analysis for PCB Reliability Data: Reading the Slope

A reliability test produces a set of failure times, and reducing them to an average throws away the information that the test was run to obtain. Weibull analysis describes the failures as a distribution with a shape and a scale, and the shape, plotted as the slope on the probability chart, is the parameter that identifies whether the failures come from wear out, from a random defect or from an early life weakness of the assembly.

The analysis is applied to board level tests, from thermal cycling and thermal shock to vibration and humidity bias, and it is the standard way to compare a design change or a process change. This article describes what each parameter means in that context and the practical conditions that make a result defensible.

What a Distribution Adds

A single failure time from a test describes one unit. A set of failure times describes the population, and the failure distribution is what the product will be sampled from in the field. Weibull analysis fits a distribution to those times and reports a scale parameter, the characteristic life, and a shape parameter, the slope.

The characteristic life is the time at which about 63 percent of the population has failed. It is a location on the time axis and it answers how long the product lasts. The slope answers a different question, which is why it fails, and it is the more useful of the two when the purpose of the test is to improve a process.

The Meaning of the Slope

A slope below about 1 describes a decreasing failure rate, which is characteristic of early life failures in a population with a distribution of defects. A slope near 1 describes a random failure rate, where failures are caused by external events rather than by the product. A slope well above 1 describes wear out, where the failure rate rises with time and the mechanism is progressive.

In a thermal cycling test, a slope in the range of 2 to 4 is typical for a plated barrel that fails by fatigue, and a much lower slope usually indicates that the population contained units with different initial quality, such as a batch with a variable plated thickness. Reading the slope as a statement about uniformity is the practical skill, and it is what turns the test from a pass or fail into a process signal.

Where the Data Comes From

The data is a set of failure times and a set of censored times. A unit that survived to the end of the test contributes a lower bound rather than a failure, and a unit that failed from a cause unrelated to the mechanism contributes an unrelated event. Both have to be handled correctly, and the failure mode of each unit should be confirmed by a section or by a resistance measurement rather than inferred from the fact that it stopped working.

Mixed failure modes are the commonest reason for a slope that makes no sense. A test in which some units fail by barrel fatigue and others by a cracked solder joint fits a single line to two populations and reports a slope that describes neither. Separating the modes by inspection is a prerequisite of the analysis.

Sample Size and Confidence

The slope is estimated from the data, and its uncertainty depends on the number of failures rather than on the number of units tested. A test with twenty units and two failures estimates the slope poorly, while one with twelve units and ten failures estimates it well. That asymmetry is why a reliability test is often run to a defined number of failures rather than to a fixed time.

Plot of thermal cycling failures on a Weibull probability chart

The confidence interval on each parameter should be reported with the estimate, because the comparison of two materials is a comparison of intervals rather than of point values. Two characteristic lives that differ by ten percent are indistinguishable if the intervals overlap widely, and reporting only the point values invites a decision that the data does not support.

Comparing Materials and Processes

A comparison is made on the same test, with the same conditions and the same failure definition, and with sufficient units of each variant to produce a usable number of failures. Where one variant survives the whole test without failing, the comparison is a statement about a limit rather than about a life.

The strongest form of the comparison is a change in the slope rather than in the location. A process change that raises the characteristic life but leaves the slope unchanged has shifted the distribution, while a change that reduces the slope has removed defective units from the population. The second is often more valuable, because it changes the shape of the failure pattern in the field.

Thermal Cycling and Thermal Shock Data

Thermal cycling tests are run to a defined profile with a defined dwell, and the choice of profile is part of the test definition rather than a detail. A cycle from minus 40 to plus 125 °C with a ten minute dwell loads a plated barrel differently from a shock between the same two temperatures with a short dwell, and the data from the two tests cannot be compared directly.

The failure criterion matters as much as the profile. A resistance increase of ten percent measured in situ detects a barrel crack earlier than a functional failure does, and the analysis is more informative because the failure time is closer to the physical event.

Laboratory chamber with test boards under thermal cycling

The test methods behind those criteria are described in the notes on thermal cycling and in the notes on thermal shock testing of plated holes.

Humidity and Corrosion Data

Humidity tests produce failures that are usually driven by contamination or by a coating defect rather than by fatigue, so their distributions behave differently. The failure times often scatter widely, the slope is low, and the analysis shows that the population contains a range of surface conditions. That is a statement about cleanliness control, and it is useful even when the absolute life is not.

The bias voltage and the test temperature should be reported with the data, because a different condition produces a different distribution. When a test result is used to qualify a coating, the comparison should be made at the same humidity and bias, and the analytical method for the contamination side is described in the notes on biased humidity testing.

Limits of the Analysis

Weibull analysis assumes that the failures come from a single mechanism and that the units are statistically alike, and neither is guaranteed. A population that contains two production lots, a design change made part way through the build or a repair is not alike, and mixing those units produces a slope that describes the mixture rather than the product.

The analysis also cannot predict a mechanism that has not appeared in the test. A test that produces no failures establishes a lower bound on the life at a stated confidence, which is a legitimate and useful result, and it does not establish that the failure mode is absent. The limitations are best handled by reporting the evidence rather than by fitting a distribution to it and presenting the fit as a life prediction.

Process Control and Records

Where the analysis is used for process control, the records are the test conditions, the unit history including lot and process route, the failure mode confirmed on each unit, and the fitted parameters over time. Trending the parameters rather than the individual failures shows a change in the population before it appears in the field.

That trend is the practical output of the whole exercise. A characteristic life that is stable while the slope falls is a warning that variability has entered the process, and the measurement behind it is described in the notes on delamination testing as well as in the mechanical tests that share this reporting discipline.

FAQ

How many failures are needed for a meaningful slope? Several. A slope estimated from two or three failures has an interval so wide that it distinguishes little, and the test is usually run to a defined failure count.

What does a low slope mean in a board level test? Usually that the population contained units with different initial quality, which points to a process control issue rather than to a fatigue mechanism.

Can a test with no failures be reported? Yes, as a lower bound on the life at a stated confidence. It cannot be reported as a characteristic life.

Leave A Comment