{
  "abstract": "In many areas of quantitative research, particularly in clinical research, it is common practice for continuous data (eg, personal information such as height, chemical levels in the body, or questionnaire total scores) to be summarised by the mean and standard deviation (SD). When using skewed continuous data, while still having merit, these summary measures should not be presented as the sole measures available. We argue that the reporting of median and interquartile interval for skewed data should be the predominant choice of summary measure, while the mean or SD can be presented as an additional measure. We also highlight how calculating a 95% reference range using the mean and SD can assess normality when the raw data are unavailable, for example, when reviewing a published article. We illustrate these concepts through examples and provide some practical guidance for authors and reviewers. The concepts discussed in this article highlight not only frequent author oversights when summarising data values, but how researchers, peer reviewers, and students can use the presented information to assess normality as a critical appraisal tool.",
  "authors": [
    {
      "affiliations": [
        "College of Health and Life Sciences, Aston University, Birmingham, UK"
      ],
      "name": "Dan J Green"
    },
    {
      "affiliations": [
        "Division of Medicine and Public Health, University of Sheffield, Sheffield, England, UK"
      ],
      "name": "Michael J Campbell"
    },
    {
      "affiliations": [
        "Population Policy Practice Research and Teaching Programme, University College London, London, England, UK"
      ],
      "name": "Eirini Koutoumanou"
    }
  ],
  "full_text": "The mean and standard deviation are commonly used to summarise continuous data, often without regard for the distribution of the data. For data in a non-normal distribution, the median and interquartile interval are often more appropriate. This article outlines circumstances where the mean and standard deviation alone are insufficient summary variables, and provides a simple appraisal method for reviewers without the raw data.Key messages Means and standard deviations are often used alone for summarising continuous dataReference ranges derived from the mean and standard deviation can illustrate implausible values when compared to the variable’s contextThis approach can indicate non-normality and skewed direction in the dataUniform distributions are a special case where all values may fall within the 95% reference range, contradicting Altman and Bland’s 2005 statistical paper in The BMJIntroduction In many areas of quantitative research, particularly in clinical research, it is common practice for continuous data (eg, personal information such as height, chemical levels in the body, or questionnaire total scores) to be summarised by the mean and standard deviation (SD). When using skewed continuous data, while still having merit, these summary measures should not be presented as the sole measures available. We argue that the reporting of median and interquartile interval for skewed data should be the predominant choice of summary measure, while the mean or SD can be presented as an additional measure. We also highlight how calculating a 95% reference range using the mean and SD can assess normality when the raw data are unavailable, for example, when reviewing a published article. We illustrate these concepts through examples and provide some practical guidance for authors and reviewers. The concepts discussed in this article highlight not only frequent author oversights when summarising data values, but how researchers, peer reviewers, and students can use the presented information to assess normality as a critical appraisal tool.Summary measures When summarising quantitative data, researchers have various options, especially when summarising continuous measures. The most common option for reflecting the central tendency will be a mean (the sum of all values divided by the number of values) or a median (the middle value when all values are arranged in sequential order). The sample mean is highly influenced by extreme values (because every data point is used in its calculation), especially in smaller samples, and so should not be used if there is concern that observations are not normally distributed. In those instances, the median is a better estimate of the population median, because it is not affected by extreme values. The spread of the data is equally important and of interest: when a mean is used, the SD (the square root of the sum of the squared distances of every value from the mean divided by the number of values minus one) is the likely reported measure of spread. Like the mean, the SD also uses every data point in its calculation, so the same concerns around extreme values also remain. In general, a smaller SD would indicate that most observed data values are clustered around the mean, and a larger SD would indicate that the observed data values are more spread out around the mean. The optimal measure of spread to pair with the median is the interquartile interval, reflected by two values, where 25% of the data fall below the lower value and 25% above the upper value, irrespective of the shape of the distribution. Figure 1 shows visual examples of different distributions: normal, uniform, left skewed, and right skewed.Figure 1Randomly generated examples of four different data distributions, based on a randomly generated sample size of n=106 in every scenarioDominance of standard deviations in clinical research The SD appears to be the sole measure of spread used in most published articles, irrespective of the underlying distribution of data. Authors may feel supported by Altman and Bland’s influential 2005 paper in The BMJ,1 where they state, “Contrary to popular misconception, the standard deviation is a valid measure of variability regardless of the distribution.” However, Altman and Bland later quantify a key part of their message: “We may choose a different summary statistic, however, when data have a skewed distribution.”1 Campbell, among others, has suggested that for non-normally distributed variables, the SD alone is not a useful summary measure, and indeed its use may hide useful information.2 The SD can have some use as a summary statistic, as can the mean, even for non-normal distributions. For example, researchers might wish to use the mean and SD in their reviews and meta-analyses, so including these measures will support and streamline their work. Usually, for non-normal distributions, it is helpful to give both the mean and the median, because their proximity is also a measure of normality or skewness.Recent examples of insufficient published summaries Although Altman and Bland would choose “a different summary statistic … when data have a skewed distribution,” 1 authors frequently do not do the same. For example, one study investigated the questionnaire scoring tool CAPS-5, a 30 item tool used by clinicians to diagnose current or lifetime post traumatic stress disorder.3 This tool has a potential range from 0 (no issue) to 80 (extreme issue), and was completed by 83 individuals. The CAPS-5 tool scored a mean of 13.0 (SD 11.1) in one group, and a mean of 13.1 (11.7) in the other. Both groups reflected a right skewed distribution (box 1, figure 2).Box 1Worked example with interpretation of the CAPS-5 scoring tool3CAPS-5, a scoring tool to diagnose current and lifetime post traumatic stress disorder, has potential values of 0 to 80. CAPS-5 was completed by 83 individuals; no negative values were possible, with a reported mean of 13.0 (standard deviation (SD) 11.1). To calculate the 95% reference range limits, two SDs* were subtracted and added on either side of the mean:Lower limit: 13.0 − (2×11.1) = 13.0 − 22.2 = −9.2Upper limit: 13.0 + (2×11.1) = 13.0 + 22.2 = 35.2If we assume a normal distribution, we expect 95% of data points to lie between −9.2 and 35.3. Thus, 5% of data points will be outside this range, with 2.5% below −9.2, and 2.5% above 35.3. Therefore, with a population of 83 in this example, 5% of data points equates to about two individuals expected to have values below −9.2 and two individuals expected to have values above 35.3.Observations above 35.3 do not raise concerns, but negative values are not possible. The scenario here is that most values were relatively small, possibly clustered in the 0-20 range, and that a small number of individuals had a larger score, which inflated the SD and resulted in a right skewed distribution (can also be termed as positive or upward distribution).Figure 2 (top panel) reflects simulated data of the CAPS-5 variable described above with the mean (13.0) and 95% reference range noted (−9.2 and 35.3). Calculating the reference range reveals that, should the variable be normally distributed, there would be values less than −9.2, which are not plausible values. Figure 2 (bottom panel) reflects the same simulated data of the CAPS-5 variable, but with the median and interquartile interval values instead, giving a more accurate reflection of the central tendency and spread of the data.*Technically, the calculation should use 1.96 × SD to go with 95% reference range, but two SDs is a good approximation of this value and helpful for a quick calculation.Figure 2Simulated data using scores from the CAPS-5 tool to diagnose current and lifetime post traumatic stress disorder, based on conditions described in box 1. Top panel: calculation of reference range limits with mean of 13.0 and 95% reference range of −9.2 and 35.3. Bottom panel: calculation of interquartile interval (IQI) with median of 10.9 and IQI limits of 3.2 and 20.5In another example, the PRECISE-DAPT risk calculator is a validated clinical tool to predict the risk of bleeding in patients undergoing coronary stenting followed by dual antiplatelet therapy (DAPT), and scores range from 0 to 100.4 One group had a mean of 10.5 (SD 7.9) while a second group had a mean of 10.3 (7.8) (box 2; figure 3). This scenario again reflects a right skewed distribution, where the mean and SD have been pulled in the direction of extreme value.Box 2Worked example with interpretation of the PRECISE-DAPT risk calculator4PRECISE-DAPT, a validated tool to predict bleeding risk during coronary stenting followed by dual antiplatelet therapy, has a potential range of 0 to 100 points; scores for 878 individuals had a mean of 10.5 and standard deviation of 7.9. Reference range limits were calculated using the same method as used in box 1:Lower limit = 10.5 − (2 × 7.9) = 10.5 − 15.8= −5.3Upper limit = 10.5 + (2 × 7.9) = 10.5 + 15.8 = 26.395% reference range calculated as −5.3 to 35.3Similiar to the example in box 1, this scenario is another example of a right skewed distribution because negative values are not plausible, rather than a normal distribution (figure 3)Figure 3Simulated data of the PRECISE-DAPT risk calculator to predict bleeding risk during coronary stenting followed by dual antiplatelet therapy, based on conditions described in box 2. Top panel: simulated PRECISE-DAPT variable distribution. Bottom panel: simulated PRECISE-DAPT data with mean (solid line) and 95% reference range limits (dashed lines)A more extreme example involves cost data from Medicaid, a US public health insurance programme providing free or low cost health insurance cover to eligible individuals. A population of 5796 Medicaid beneficiaries had a mean of $2296 (€1937; £1676; SD $8987) in one group and a mean of $1947 (SD $7995) in the other group (box 3).5 This showed another example of a right skewed distribution, where the SD is larger than the mean—which is an automatic potential warning sign (box 3; figure 4).Box 3Worked example of US Medicaid costs5For 5796 Medicaid beneficiaries, the sum of total charges has a potential range of $0 to ∞ (infinity), with a mean of $2296 and standard deviation (SD) of $8987. Reference range limits were calculated using the same method as used in box 1:Lower limit = 2296 – (2 × 8987) = 2296 – 17 974 = –15 678Upper limit = 2296 + (2 × 8987) = 2296 + 17 974 = 20 27095% reference range: –15 678 to 20 270Because negative values are not feasible (unless treatments were generating money rather than costing money), this is another example of a right skewed distribution (similar to the example in box 1), with a few treatments costing a high amounts. We can already see that the SD is larger than the mean, and when negative values are not feasible, performing the calculation is redundant; the distribution cannot be normally distributed (figure 4).Figure 4Simulated data of the Medicaid variable distribution, based on conditions described in box 3. Top panel: simulated data of Medicaid costs distribution. Bottom panel: simulated data of Medicaid costs with mean (solid line) and 95% reference limits (dashed lines)Figure 5Uniform distribution of simulated data showing various summary measures (mean 36.8, standard deviation (SD) 11.9, n=100). Top left panel: mean 36.8 (solid line). Top right panel: mean 36.8 (solid line) with lower and upper 95% reference range limits 13.5 and 60.1, respectively (dashed lines). Bottom left panel: mean 36.8 (solid line), highlighting approximately 2 times SD region (1.96×11.9) on one side. Bottom right: mean 36.8 (solid line) with lower and upper 95% reference range limits 13.5 and 60.1, respectively (dashed lines), highlighting full 95% reference range (mean ±1.96×SD (or ±1.96×11.9)), which encapsulates 100% of data valuesAll three worked examples illustrate scenarios where the mean and SD alone are insufficient to describe the distribution, and the median and interquartile interval should have been presented to offer a more informative reflection of the data.Use of reference ranges to assess normal distribution with no access to raw data Although assessing normal distributions is more straightforward because authors have access to the raw data (by plotting histograms, QQ plots, or other relevant graphics), we propose an approach for reviewers when assessing a published article with only the mean and SD available. For a normal distribution, the 95% reference range of observations must lie 1.96 SD values either side of the mean. Therefore, we can calculate the 95% range (mean ±1.96×SD) and examine the limits within the context of the variable. Authors can use this approach as a critical appraisal tool when reviewing the appropriateness of published data summarised by the mean and SD. We have illustrated these calculations and interpreted the findings in context in boxes 1-3. For instance, the CAPS-5 variable has possible values from 0 to 80. Calculating the mean±1.96×SD would indicate negative values that are implausible, as described in box 1.Do reference ranges apply to all data distributions: a special case An important factor to consider is the uniform distribution, which expands on Bland and Altman's suggestion that “95% of observations of any distribution usually fall within the two standard deviation limits.” 1 A uniform distribution has a unique characteristic of equal probability assigned to each value, and for larger samples, a 95% reference range calculation can result in 100% of all data observations falling within the two limits. For a uniform distribution over the range of A to B, the expected SD is (B–A)/√12. Thus, if the range is 0 to 1, and the expected mean and SD are 0.5 and 0.29, respectively, then the expected 95% range would be −0.08 to 1.08, which includes all the data. We have described this scenario using a randomly generated dataset in figure 5, which illustrates that 100% of the observed values fall within 1.96 SDs either side of the mean. Because a uniform distribution is different to a normal distribution, the calculation of the proposed reference range further highlights that the mean and SD alone are insufficient to fully summarise central tendency and spread.Why standard deviations are commonly used Non-normally distributed variables are often not the main outcome variable; therefore, authors might not be overly concerned about showing them in useful ways. Also, even if they are the main outcome, non-normality does not necessarily invalidate parametric models, which have normality assumptions. Some parametric models assume normality in the residuals (the difference between observed and model-predicted values) after accounting for covariates and not in the distribution of the data, while parametric tests are quite robust to normality assumptions, particularly for large datasets and where the data are randomised. 2We theorise two reasons why authors may display the SD inappropriately: firstly, because everyone else does; and secondly, because it is produced routinely by statistical computational packages. We suggest that by authors giving more thought about which measures to use, papers can become more informative.Recommendations on reporting continuous data When reporting continuous data:Report the median and interquartile interval for skewed dataInclude the mean and SD, because they are useful for normal data and may still be useful with non-normal data (eg, to inform data extraction for meta-analysis or other purposes)Avoid interpreting the mean with two SDs as a range, unless normality is plausibleConsider graphical summaries, such as dot plots, for clarity.Table 1 gives an example of how data might be optimally presented.Table 1Recommended presentation of continuous dataVariableMedian (interquartile interval), unless otherwise statedMean (standard deviation)Median (range) age (years)40 (18-65)46 (34)Continuous outcome18 (10-45)20 (22)Continuous outcome 25.5 (4-7)5.4 (2)Conclusions Although the mean and SD remain useful for many purposes, including for meta-analysis and calculation of standard errors, they are not sufficient to describe skewed distributions. If data are not normally distributed, reviewers and authors should consider additional measures (eg, median, interquartile interval, and graphical summaries) to ensure accurate representation and interpretation of data. The approach outlined here is intended as a practical tool for reviewers when raw data are unavailable, but not as a replacement for full statistical analysis.",
  "title": "When means and standard deviations are an incomplete summary of a continuous variable: problems, solutions, and utilising the reference ranges to check normality",
  "uid": "30a9332c-4f6a-54e3-9391-6bb8fc25e2cc"
}
