The Significance of Significance
Perhaps the most misunderstood single word from the world of Statistics is significant. It is a leading example of using simple words for complicated ideas. The problem is the English language. Outside Statistics, significant means important. A “significant event” is one that changed history; a “significant other” is an important person in one’s life; a “significant number” can refer to a large amount. Inside Statistics, significant means real; that’s something entirely different.
The misconception appears widely in the public, especially when the word is used in the phrase “statistically significant”. Media use it, researchers celebrate it, and companies make vital decisions based on it. Yet many people—including some who work with data every day—mistakenly assume that “statistically significant” means “important”. It doesn't.
There is a critical difference between “statistical significance” and “practical significance”. Understanding that difference is essential for good decision-making in a data-centric world. The distinction becomes even more important today, when organizations routinely analyze millions, or even billions, of observations. Let’s explore the difference.
Statistical significance is a foundational concept in statistical inference. Hypothesis testing and confidence intervals are based on random samples, and account for sampling variability, namely, that results will change when new random samples are taken. Thus, the first way to think of statistical significance is to ask whether an observed result would persist if the study were repeated. Can the result be explained by random chance, or would the result be expected to happen again? If chance seems an unlikely explanation, then look for another explanation. If the result isn’t just a fluke of sampling, it must represent something real. Note that it does not tell us whether the result matters. We’ll return to that point soon.
Imagine tossing a fair coin ten times. Getting seven heads isn't too surprising, but getting ten heads in a row seems unlikely, and would make you begin to wonder whether the coin was indeed fair. Note the use of the word “unlikely” rather than “impossible”. Sometimes extraordinary-looking results do occur entirely by chance. We haven’t yet quantified “unlikely”; we will, below.
Practical significance addresses whether an observed result is large enough to matter. Is it sufficiently important to change practice or behaviour or make a different decision. A headline may proclaim that a new diet significantly reduces cholesterol. Even if the result were statistically significant, would anyone notice or care if cholesterol fell by only one point? It is doubtful whether a physician would change treatment or a patient would benefit. In medical and health care settings, practical significance is often referred to as clinical significance.
Imagine a new toothpaste advertised as reducing cavities by 0.2%. The clinical trial includes 100,000 participants and achieves a statistically significant result. That seems like good news, but would you switch brands? Now imagine an automobile manufacturer announcing a 0.2% reduction in brake failures. Suddenly that tiny improvement becomes enormously important. The same number can be trivial in one setting and lifesaving in another. Practical significance always depends on context. Statistics tells us whether an effect probably exists. Judgment tells us whether it matters.
Let’s return to statistical significance. Most statistical studies begin with a skeptical assumption. Suppose a pharmaceutical company tests whether a new medication is effective in reducing blood pressure. The statistician starts by assuming it is not effective, an assumption known as the null (i.e., “nothing happened”) hypothesis. The data are then examined to see how compatible they are with that assumption. If the observed improvement would occur less than about 5% of the time purely by chance, the result is called statistically significant. Where did the 5% threshold come from? Tradition. In effect it has quantified “unlikelihood” as something that happens less than 5% of the time. And the probability of the observed important occurring purely by chance is called the p-value. That will be the subject of a future article.
Suppose the data do indeed show that the blood pressure reduction with the new medication is statistically significant. It simply means that random chance is an unlikely explanation but notice what has not been said. The researchers have not “proved” that the medication works. Nor have they shown that it is worthwhile.
A tiny improvement measured with great precision may be statistically significant while remaining practically unimportant. Conversely, a large and potentially important improvement observed in a small pilot study may fail to reach statistical significance because too little data were collected. Looking only at statistical significance hides this important distinction.
Scientific papers often place an asterisk beside statistically significant results. Researchers sometimes joke that they spend their careers chasing stars. It is an understandable temptation since statistically significant results are more likely to be published. They attract headlines and generate excitement. But that singular focus encourages researchers to ignore the questions that matter more.
Now let’s put statistical and practical significance together. Modern statistical reporting increasingly emphasizes effect size rather than simply saying whether a result is statistically significant. Effect size answers questions like how much lower was blood pressure? How much faster was the new process? How much larger were sales? How much better did students perform?
Putting effect size and uncertainty due to sampling variability together is where confidence intervals come in. In the blood pressure medication study, the conclusion might be reported as, “We estimate the average reduction in blood pressure to be between six and ten points.” Now we know that the reduction is statistically significant, and we know how large it is expected to be. It’s up to the physicians to decide whether a reduction between six and ten points is clinically important and relevant. Confidence intervals shift the emphasis from passing an arbitrary threshold to understanding the result. They will also be a subject for a future article.
Bigger Isn't always better. Decades ago, it was expensive and time-consuming to collect data. Researchers celebrated studies involving a few hundred observations. Today companies routinely analyze millions of customer transactions every day. Hospitals maintain databases containing decades of medical records. Internet companies monitor billions of clicks. Ironically, with so much information statistical significance becomes much less useful. When I started my consulting practice, I proposed a tongue-in-cheek advertising slogan, “Significance guaranteed, or double your data back.”
The larger the sample, the easier it becomes to detect tiny differences. It’s analogous to increasing the resolution on a telescope or microscope. What is not visible to the named eye is visible under magnification. Suppose one version of a website increases purchases by just 0.03%. With 100 visitors, you would never notice. With 100 million visitors, the difference becomes overwhelmingly statistically significant. But should the company redesign its website? That depends on annual sales. At $20 million, the change may not be worth the programming costs. At $20 billion, that same tiny improvement could be worth millions. Big data can detect microscopic effects with extraordinary precision. It cannot decide whether those effects deserve our attention.
Statistical significance still matters with huge datasets, but not in the same way. Millions of observations cannot rescue poor study design. If data are systematically biased, adding more observations simply produces a more precise estimate of the wrong answer. As the statistician John Tukey famously observed: "An approximate answer to the right question is worth a great deal more than an exact answer to the wrong question."
One more misconception deserves mention. Statistical significance does not establish cause and effect. Every summer, both ice cream sales and drowning deaths increase. The correlation is highly statistically significant. However, ice cream does not cause drowning; warm weather causes both rates to increase. Statistics can identify patterns but cannot, by itself, explain why those patterns exist. Good science still requires thoughtful study design, careful reasoning, and healthy skepticism.
Whenever you read that a study found a statistically significant result, pause before accepting the conclusion. Ask a few additional questions. How large was the effect? Would anyone actually notice the difference? Could bias explain the result? Was the study designed well? Would the finding change a doctor’s treatment, a business decision, or a public policy?
Statistical significance is a great achievement in the science of understanding data. It gives us a disciplined way to distinguish genuine patterns from random noise, but it was never meant to be the final word. In the age of artificial intelligence, machine learning, and massive databases, our computers have become extraordinarily good at finding tiny differences. Humans must decide whether those differences are worth acting upon.
Perhaps the simplest way to remember the distinction is this: Statistical significance tells us whether we should believe the effect exists. Practical significance tells us whether we should care. Good statistics requires both.