{
  "abstract": "Objective To examine the efficacy of conservative (non-surgical) treatments, usual care, and no treatment for chronic radicular and non-specific back pain.Design Time course network meta-analysis.Data sources Six electronic databases (Medline, SPORTDiscus, CINAHL, PsycINFO, Embase, and CENTRAL), searched from inception to 24 July 2020, and 302 previous systematic reviews.Eligibility criteria for selecting studies Full peer reviewed publications in English or German of randomised controlled trials, randomised clinical trials, randomised controlled cluster trials, or randomised crossover trials in adults (aged ≥18 years) receiving common conservative treatments for non-specific and radicular chronic low back pain. Treatments examined were acupuncture, education or advice, electrotherapy (including heat and ice electrotherapeutic modalities applied non-invasively), exercise training, manual treatments or manipulation, massage, the McKenzie method, pharmacotherapy, psychological treatments, traction, physical therapy (otherwise not falling into specific treatment combinations), placebo, multidisciplinary pain management, usual care (eg, management by a doctor), and no treatment (true control).Results Back pain intensity, leg pain intensity, disability, and mental health outcomes were reported immediately (<1 day), and at short term (≥1 day and ≤3 months), intermediate term (>3 and <12 months), and long term (≥12 months) time points. 581 reports of 551 studies (71 126 patients) were included. 510 trials included people with non-specific chronic low back pain and 41 trials included those with radicular chronic low back pain. For back pain (0-100 scale), acupuncture (mean difference −20.91, 95% credible interval −24.00 to −11.95), electrotherapy (−18.98, −21.84 to−10.95), exercise (−15.59, −17.51 to −10.05), manual treatment (−19.48, −22.17 to −11.74), massage (−25.61, −30.42 to −10.91), and multidisciplinary pain management (−18.96, −22.26 to −9.58) exceeded the minimal clinically important difference (set at 0.5 standard deviation) in the short term. For disability (0-100 scale), acupuncture (mean difference −10.52, 95% credible intervals −11.84 to −6.59), massage (−9.95, −11.45 to −5.50), and multidisciplinary pain management (−12.56, −13.91 to −8.55) were clinically effective in the short term. In the immediate and intermediate term only, the McKenzie method and massage, respectively, exceeded the minimal clinically important difference. In the long term, although two of the 14 treatments for back pain and nine of 14 treatments for disability had statistically significant benefits compared with no treatment, the effects were not clinically significant. The certainty of the evidence based on the GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) framework was low (1.4%) to very low (98.6%) across interventions and time points. Findings for massage and the McKenzie method were not stable in the sensitivity analyses. Treatment effects for radicular chronic low back pain did not seem to differ from those for non-specific chronic low back pain.Conclusions Some treatments were effective for pain and function in non-specific chronic low back pain, but improvements did not persist long term. Most of the evidence was for non-specific chronic low back pain; the evidence base for radicular chronic low back pain was limited. Although sensitivity analyses did not provide evidence for a different response in radicular chronic low back pain, an evidence gap remains for this subpopulation. Future work should explore strategies to establish the long term efficacy of modifications to lifestyle and behaviour.Systematic review registration PROSPERO CRD42020182039",
  "authors": [
    {
      "affiliations": [
        "Department of Nursing, Midwifery, and Therapeutic Sciences, Hochschule Bochum, Bochum, Germany",
        "Institute for Physical Activity and Nutrition, Deakin University, Melbourne, Victoria, Australia"
      ],
      "name": "Daniel L Belavý"
    },
    {
      "affiliations": [
        "Physio Meets Science, Heidelberg, Germany"
      ],
      "name": "Tobias Saueressig"
    },
    {
      "affiliations": [
        "Department of Nursing, Midwifery, and Therapeutic Sciences, Hochschule Bochum, Bochum, Germany"
      ],
      "name": "Nitin Kumar Arora"
    },
    {
      "affiliations": [
        "School of Allied Health, Human Services, and Sport, La Trobe University, Melbourne, Victoria, Australia"
      ],
      "name": "Arun Prasad Balasundaram"
    },
    {
      "affiliations": [
        "Orthopaedic Surgery, Xuanwu Hospital Capital Medical University, Beijing, China",
        "Orthopaedic Surgery, University of New South Wales-Saint George Campus, Sydney, New South Wales, Australia"
      ],
      "name": "Xiaolong Chen"
    },
    {
      "affiliations": [
        "Orthopaedic Surgery, University of New South Wales-Saint George Campus, Sydney, New South Wales, Australia",
        "School of Medicine, University of Adelaide, Adelaide, South Australia, Australia",
        "Department of Orthopaedic Surgery, Royal Adelaide Hospital, Adelaide, South Australia, Australia"
      ],
      "name": "Ashish D Diwan"
    },
    {
      "affiliations": [
        "School of Allied Health, Human Services, and Sport, La Trobe University, Melbourne, Victoria, Australia",
        "Advance Healthcare, Bundoora, Victoria, Australia"
      ],
      "name": "Jon J Ford"
    },
    {
      "affiliations": [
        "School of Allied Health, Human Services, and Sport, La Trobe University, Melbourne, Victoria, Australia"
      ],
      "name": "Andrew J Hahne"
    },
    {
      "affiliations": [
        "Department of Nursing, Midwifery, and Therapeutic Sciences, Hochschule Bochum, Bochum, Germany"
      ],
      "name": "Svenja Kaczorowski"
    },
    {
      "affiliations": [
        "Institute for Physical Activity and Nutrition, Deakin University, Melbourne, Victoria, Australia"
      ],
      "name": "Clint T Miller"
    },
    {
      "affiliations": [
        "Institute for Physical Activity and Nutrition, Deakin University, Melbourne, Victoria, Australia"
      ],
      "name": "Niamh L Mundell"
    },
    {
      "affiliations": [
        "Bristol Medical School, University of Bristol, Bristol, UK"
      ],
      "name": "Hugo Pedder"
    },
    {
      "affiliations": [
        "Department of Nursing, Midwifery, and Therapeutic Sciences, Hochschule Bochum, Bochum, Germany",
        "Graduate School for Applied Research in North Rhine-Westphalia, Bochum, Germany",
        "Clinical Research Unit, Parker Institute, Frederiksberg, Denmark"
      ],
      "name": "Tim Schleimer"
    },
    {
      "affiliations": [
        "Institute for Physical Activity and Nutrition, Deakin University, Melbourne, Victoria, Australia",
        "Centre for Youth Mental Health, University of Melbourne, Parkville, Victoria, Australia",
        "Orygen National Centre of Excellence in Youth Mental Health, Parkville, Victoria, Australia"
      ],
      "name": "Scott D Tagliaferri"
    },
    {
      "affiliations": [
        "Department of Nursing, Midwifery, and Therapeutic Sciences, Hochschule Bochum, Bochum, Germany",
        "Graduate School for Applied Research in North Rhine-Westphalia, Bochum, Germany"
      ],
      "name": "Florian Teichert"
    },
    {
      "affiliations": [
        "Xi’an University of Architecture and Technology, Xi’an, Shaanxi, China"
      ],
      "name": "Xiaohui Zhao"
    },
    {
      "affiliations": [
        "Emergency Medicine Programme, Eastern Health, Box Hill, Victoria, Australia",
        "Monash University Eastern Health Clinical School, Box Hill, Victoria, Australia"
      ],
      "name": "Patrick J Owen"
    }
  ],
  "full_text": "WHAT IS ALREADY KNOWN ON THIS TOPIC Several conservative treatments are recommended in clinical practice guidelines for chronic non-specific and radicular low back painPrevious meta-analyses have typically focused on individual treatment approaches and at discrete time points, limiting understanding of the relative effectiveness of treatments and their performance over timeWHAT THIS STUDY ADDS For non-specific low back pain, a range of conservative treatments provided pain relief beyond a minimal clinically important difference in the short termClinically important effects in the long term were not apparent, with effects typically lasting little beyond the typical treatment durationConservative treatments for chronic radicular low back pain represent an evidence gapHOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE, OR POLICY Future policy and research efforts should focus on long term maintenance of patient improvements for people with chronic low back painAlthough booster treatments with or without lifestyle and behavioural modification strategies for self-management may be better for chronic back pain, more primary randomised controlled trials are required to evaluate this approachThese findings can inform evidence based decisions by healthcare providers and guideline groups on prioritising conservative treatments for chronic low back pain, given their lack of long term benefits and low to very low certainty evidence for short term benefitsIntroduction Low back pain is the greatest cause of disability and lost productivity worldwide, 1 and chronic low back pain (>12 weeks since onset) generates the most economic burden.2 In the triage of low back pain, when serious causes are excluded (eg, malignancy or fracture), radicular syndromes (ie, radicular pain, radiculopathy, and spinal stenosis) and non-specific low back pain remain.3 To reduce the global burden of disease of chronic low back pain, identifying effective treatment is important.First line management of both non-specific and radicular chronic low back pain is conservative.4 5 Conservative management was defined as non-surgical interventions, including both pharmacological and non-pharmacological treatments,6 7 while excluding invasive interventional procedures, such as epidural injections, radiofrequency denervation, and spinal cord stimulation. High quality international guidelines support treatment recommendations that are similar for both chronic radicular syndromes4 and chronic non-specific low back pain.5 Consistent evidence indicates that many conservative interventions can provide short term symptom relief. Whether these benefits are sustained long term, however, is unclear, or whether the time course or durability of effects differ across interventions or between non-specific and radicular forms of chronic low back pain. Although hundreds of meta-analyses of randomised controlled trials have been performed for conservative treatments of chronic low back pain, whether any treatment is superior or effective beyond the short term is not clear. Understanding the comparative effectiveness of interventions is necessary for informing guidelines and the allocation of healthcare resources.A time course, model based network meta-analysis is an extension of a network meta-analysis.8 9 A network meta-analysis allows the ranking and comparison of interventions as more or less effective,10 and incorporates data on multiple treatments simultaneously from randomised controlled trials that do not have similar comparator groups.11 A limitation is that a network meta-analysis cannot readily integrate information from the same randomised controlled trial over time or from different time points across various randomised controlled trials in estimating comparative efficacy.8 9 A time course network meta-analysis overcomes this limitation by modelling treatment trajectories across follow-up durations, allowing comparison of onset, duration, and persistence of effects.8 9 This approach helps in predicting the onset, duration, and overall efficacy of treatments, leading to better informed healthcare decisions.8 9 The main aim of this work was to examine the comparative efficacy of common conservative treatments for non-specific and radicular chronic low back pain by time course network meta-analysis.Methods This systematic review was conducted and reported in accordance with Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA; online supplemental data 1)12 and the PRISMA extension for network meta-analyses (PRISMA-network meta-analysis; online supplemental data 2).13 This systematic review was prospectively registered on PROSPERO (CRD42020182039). The study protocol was published14 and we present an abridged version here. Online supplemental data 3 shows the changes to the protocol.SP110.1136/bmjmed-2025-001908.supp1Supplementary dataEligibility criteria We included full peer reviewed publications in English or German of randomised controlled trials, randomised clinical trials, randomised controlled cluster trials, or randomised crossover trials. Online supplemental data 4 has details on the eligibility criteria. Studies in adults (aged ≥18 years) with non-specific or radicular chronic low back pain were included. Recurrent pain (ie, <12 weeks’ duration of symptoms and a pain free period of at least six months15) was excluded. Specific spinal pathologies (ie, vertebral fracture, malignancy, spinal infection, axial spondyloarthritis, and cauda equina syndrome3) were excluded. Spondylolisthesis, spondylosis, disc herniation, disc degeneration, scoliosis, deformity (eg, hemivertebrae) were classified as non-specific back pain,3 16 and radicular syndromes (eg, radicular pain (leg pain or sciatica), radiculopathy, and spinal stenosis) were included.3For interventions and comparators, we examined acupuncture, education or advice, electrotherapy (including heat and ice electrotherapeutic modalities applied non-invasively), exercise training, manual treatments or manipulation, massage, the McKenzie method, pharmacotherapy, psychological treatments, traction, physical therapy (otherwise not falling into specific treatment combination), placebo, multidisciplinary pain management, usual care (eg, managed by a general practitioner), and no treatment (true control). Treatment combinations were defined according to their primary and secondary treatment intervention components (online supplemental data 5). Outcomes of interest were back or leg pain intensity, disability, and mental health.Information sources and search strategy Six databases were searched on 24 July 2020, with no restriction on publication dates ( online supplemental data 6). We identified 285 previous systematic (including Cochrane) reviews of any type of treatment for chronic low back disorders published in the past 10 years. The reference lists of these reviews were included in screening. We also screened the reference lists of 17 relevant Cochrane reviews published between January 1990 and July 2019. To assess whether new trials would meaningfully change previous conclusions, we conducted trial sequential analyses.17 These analyses (online supplemental data 7) showed that the required information size has already been surpassed in >80% of comparisons for both pain and disability, indicating robust evidence that is unlikely to be altered by future studies.Selection process Covidence ( https://www.covidence.org) was used for screening of titles and abstracts, full text screening, and removing duplicates. Pilot testing with all screeners and adjudicators was performed before the beginning of each phase, with 100 title and abstract records and 20 full texts. Discrepancies from the pilot tests were discussed among the team. For each record and full text, two independent assessors (from a pool of wider screeners; online supplemental data 7) screened the studies. Disagreements that could not be resolved among the assessors were decided by an adjudicator.Data collection process and data items To pilot test the extraction process, two senior team members (DLB and SDT) extracted the same 10 publications and achieved consensus with their extraction. Subsequently, the remaining extractors extracted the same 10 publications and their extraction was compared with the first team, feedback given, and extraction updated to clarify misunderstandings. For each record, two independent assessors (see Contributor statement) extracted the data. Disagreements were resolved by discussion between the extractors, and the adjudicators consulted if necessary. As an additional quality control step, an additional third independent extractor (see Contributor statement) cross checked the extraction of each study and updated the extracted data if errors were detected. To check for multiple reports from the same study, a semi-automated approach was used: user written code (in R; www.r-project.org) was used to flag reports where we found ≥30% overlap in author surnames, the number of patients enrolled and duration of intervention were, in the extracted data, within 10% of each other, and the type of back pain and primary intervention in the reports were the same. These reports were checked by extractors to determine if they were from the same study. During independent cross checking by a third extractor, the extractor manually checked for potential multiple reports from the same study (eg, by citation within a report, similar sample sizes, and interventions). To cross check for retracted studies, the included reports were imported into Zotero (www.zotero.org) to cross check with the integrated Retraction Watch database. Online supplemental data 8 describes the extracted data, and an exemplary extraction sheet is available online (online supplemental data 9). Online supplemental data 10 has information on data handling, including handling of cluster randomised and cross over trials according to guidance from Cochrane.Study risk of bias assessment During extraction, two independent assessors used the Cochrane Collaboration risk of bias version 1 (RoB-1) tool 18 to examine potential selection bias (random sequence generation and allocation concealment), performance bias (masking of patients and staff), detection bias (masking of outcome assessment), attrition bias (incomplete outcome data), reporting bias (selective outcome reporting), and other biases. Cluster randomised trials were assessed as recommended by the Cochrane Collaboration.19 This choice was based on a prespecified methodological plan14 and aligned with guidance from the Cochrane Collaboration guidance at the time of the start of the study.20 In 2025, RoB-1 is still preferred by the Cochrane Back and Neck Group over the revised version when workload and time considerations are important.21 Using RoB2 would not influence our finding because risk of bias assessments informed the certainty of evidence within the GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) framework through a prespecified rule (downgrading for >50% of participants from studies with selection or performance bias). Reassessment with RoB2 would not have altered these judgments, because these domains directly map to the randomisation process of RoB2 and deviations-from-intervention domains.22 Moreover, RoB2 is highly time and resource intensive, requiring an estimated 358 min for each study compared with 10-60 min for RoB-1.23 24 Adjudication was performed by discussion with an adjudicator, if necessary. Also, independent cross checking by a third person was performed.Effect measures Because all of the outcomes of interest were continuous or ordinal but could be measured on different scales, we used standardised mean differences with internal reference standard deviations (SDs) as the effect estimates. 25 A minimum of 50 participants were required for each class of treatment to be included in the network meta-analysis. The number of participants was limited to mitigate the impact of small study effects on the results of any particular class. We used change from baseline mean differences and the corresponding change from baseline SD to calculate the standardised mean difference. When the correlation of before treatment and after treatment SD was not available, a correlational value of ρ=0.59 was used, which was derived from an empirical analysis of before and after treatment continuous outcome trials.26 When studies were reverse scaled, meaning that higher values indicated better outcomes instead of lower values, the recommended approach, as stated in the Cochrane Handbook, was to multiply the mean in each group by −1.27 We back transformed all standardised mean differences to natural units by multiplying by the average of the corresponding SD for the natural scale.28We presented the effect sizes and corresponding 95% credible intervals with forest plots for the discrete time point network meta-analysis, and the results of the time course network meta-analysis model with prediction plots and forest plots. For standard network meta-analysis models, we ranked the relative effects of each treatment or class and calculated the surface under the cumulative ranking curve. For time course models, we ranked the relative effects of each treatment or class for each time course parameter. We also ranked the full area under the time course function for each treatment or class at 0-10 weeks, 0-26 weeks, and 0-52 weeks. Cumulative rankograms were plotted to show the range of rankings of different treatments or classes for each ranked parameter. These intervals (0-10, 0-26, and 0-52 weeks) were not discrete data extraction time points but represented clinically meaningful segments along the time scale used for ranking and summarising modelled treatment effects.In the time course network meta-analysis, the evolution of treatment effects was represented by a smooth spline based function, with knots placed within these intervals to allow the slope of the curve to change where necessary and better capture non-linear trends over time. This approach differed from the discrete time point analyses (immediate, and short, intermediate, and long term), which were based on grouped trial data. Hence although the extraction time points defined when the data were collected, the spline knots were technical features of the statistical model that enabled a flexible estimation of treatment effects across the whole follow-up period.Geometry of the network Network connectivity was explored by network plots. Network plots helped to visualise how the evidence in the network was connected and allowed identification of which studies compared which treatments. 11 This visualisation helped in understanding which treatment effects could be estimated. The time course relation was examined by a time plot, which was a plot of the raw study responses over time. The time plots aimed to elucidate the underlying time course of the treatment effects and identify which statistical time course model was appropriate.8Planned methods of analysis (synthesis methods) The time course network meta-analysis was conducted with the R package MBNMAtime. 29 With this package, we could incorporate multiple time points for each study in a bayesian network meta-analysis to inform estimates of effect size over time. The model represents these changes, with flexible functional forms (eg, spline or parametric based time functions), with knots or parameters that allow the effect trajectory to vary smoothly across the follow-up period. To account for variation in follow-up times across studies, time was modelled as a continuous variable within the random effects framework, allowing each trial to contribute data at its exact reported time point. This approach minimises bias arising from different reporting schedules and avoids the need to collapse outcomes into broad time categories. For interpretability, estimated effects were also summarised at predefined clinical intervals (immediate, and short, intermediate, and long term). Arm level means and change from baseline mean differences were used in the model. Multi-arm trials for a time course network meta-analysis were adjusted as described in the NICE DSU (National Institute for Health and Care Excellence Decision Support Unit) technical support documents.30A bayesian network meta-analysis was performed at discrete time points (immediate (<1 day) effect of treatment, short term (≥1 day and ≤3 months), intermediate term (>3 and <12 months), and long term (≥12 months)) with the R package multinma.31 Standardised mean differences were estimated for each treatment contrast in each study. The standard error in the base or reference arm was computed to account for within trial correlations in multi-arm trials (≥3 conditions) for the bayesian network meta-analysis at discrete time points. We used the square root of the covariance between the computed effects, with the assumption that the effect sizes had a correlation of 0.5.32Convergence for all models was evaluated by visual inspection of trace plots, posterior densities, and the potential scale reduction factor (R-hat value).11 Model fit was assessed with the deviance information criterion, posterior mean residual deviance, and plotting the residual deviance contributions veruss the data points.11 We used 95% credible intervals and the between study SD (τ) to measure heterogeneity.11We selected vague previous distributions for all analyses.11 For the discrete time point network meta-analysis, we used a normal prior with location zero and scale 100 for the treatment effect and a half normal distribution for the heterogeneity prior with location zero and scale 5. For the time course models, we used a normal distribution as prior with a mean of zero and SD of 100 for all time course study specific baseline effect (µ) and pooled time course relative effect (d) parameters. For the heterogeneity priors, we used a half normal distribution with a mean of zero and an SD of 4.47, corresponding to a 95% prior interval for τ of about 0-7 on the standardised mean difference scale (based on established guidance for a bayesian network meta-analysis11 30).Assessment of inconsistency Assessment of similarity and homogeneity assumptions Assessment of similarity is crucial for ensuring that the transitivity assumption holds in the network meta-analysis. 33 Strong effect modifiers in chronic low back pain trials are not proven, to our knowledge,34 and thus we hypothesised possible effect modifiers based on methodological and clinical experience: baseline pain intensity, baseline disability, baseline leg pain intensity, baseline mental health, and pain duration. All baseline values, except pain duration, were scaled to a common 0-100 scale.35 We tabulated these values for each comparison and also plotted the values with box plots. Between study SDs were estimated from random effects models, and the effects of subgrouping or meta-regression were examined. Pairwise meta-analysis of the data was synthesised by standardised mean differences with accompanying 95% confidence intervals, with a frequentist random effects model with a restricted maximum likelihood estimator for the between study variance (τ²). These analyses were carried out with the R package meta.36 Visual inspection of the forest plots, statistical estimates of heterogeneity (I², τ), and 95% prediction intervals were used to assess the validity of the homogeneity assumptions.Outlier and influential study analysis was performed with dmetar for pairwise meta-analyses that had ≥10 studies to further detect potential heterogeneity.37 Meta-regression with potential effect modifiers (baseline pain intensity, baseline disability, baseline leg pain, baseline mental health, pain duration, co-intervention, and type of low back pain)38–40 was used to further check for potential heterogeneity among the pairwise comparisons if the number of studies was ≥10.41 For effect modification in pairwise comparisons (identified by meta-regression), network meta-regression with these potential effect modifiers was also fitted for network meta-analyses conducted at each time point with the package multinma.3131 To further evaluate the transitivity assumption, we examined the comparability of studies across treatment nodes by visually inspecting the distributions of key effect modifiers (baseline pain intensity, baseline disability, baseline leg pain, baseline mental health, pain duration, and co-interventions) across comparisons. In line with current methodological guidance,42 43 these visual assessments, together with the meta-regression analyses described above, were used to judge whether differences in study populations or design characteristics might threaten the assumption of transitivityConsistency assumptions For the bayesian approach, consistency assumptions were first checked by an unrelated mean effects model, which did not assume consistency. The unrelated mean effects model only synthesised direct relative effects between each arm in a study and the study reference treatment. 44 If the consistency assumption held, then the results from the unrelated mean effects and network meta-analysis models were similar. Our assessment focused on evaluating patterns in individual residual deviances, between study heterogeneity (τ), and dev-dev plots. When the consistency assumption is reasonable, these diagnostics show similar patterns across the consistency and unrelated mean effects models and no systematic departures. Although summary quantities, such as total residual deviance or deviance information criterion, were examined, these were not used as formal criteria for determining inconsistency, because small numerical differences are expected and do not reliably indicate inconsistency, particularly in larger networks. If comparison between unrelated mean effects and network meta-analysis models was suggestive of inconsistency, node splitting was performed. In node splitting, network contrasts were split into direct and indirect evidence contributions, which were then compared to examine their similarity.45Risk of bias across studies (reporting bias assessment) Small study effects and publication bias were assessed for every pairwise comparison that had ≥10 studies by visual inspection of a funnel plot and testing for small study bias by Egger's test. 46 To examine the influence of specific studies or comparisons on treatment rankings, we conducted a threshold analysis with the R package network meta-analysisthresh for the discrete time points network meta-analysis at both the study and contrast level.47 48 By using a threshold analysis, we can estimate the precise extent to which biases or other factors, such as sample variance, could affect the evidence and how that would affect recommendations. If the evidence is not expected to alter by more than this projected amount, the recommendation is deemed strong. On the other hand, the recommendation is considered sensitive to reasonable fluctuations in the evidence if it is likely that the evidence will change beyond this point.Additional analyses Subgroup analyses We performed subgroup analyses to explore whether inconsistency or heterogeneity and group differences in the outcomes were influenced by the type of low back disorder (non-specific chronic low back pain v radicular syndrome), and by exclusion of the multidisciplinary pain management, pharmacological, and physical therapy nodes from the analyses. The treatment nodes may have been a source of significant heterogeneity or inconsistency for the overall network meta-analysis because of the variability of this treatment definition compared with other interventions. A subgroup analysis focusing on key participant or study characteristics produced smaller, more homogeneous networks and was a good strategy to analyse inconsistency or heterogeneity with fewer assumptions and pitfalls than a network meta-analysis meta-regression.49Sensitivity analyses We conducted the following sensitivity analyses.Excluding studies with imputed missing SD values and imputed median values.Study sample size. Impact of studies of <20 participants in all study arms.Dropout numbers and handling of dropouts within studies. Effect of the proportion of dropouts (if reported) and the type of analysis performed in individual studies (eg, analysing all participants with imputation of missing data v analysing complete cases only).Comparison of class effect models with a model with fully independent treatment effects that assumed no within class similarity, to assess the statistical validity of class assumptions.Secondary treatment components (online supplemental data 5). We considered the impact of treatment combinations where secondary classes of treatment were present in all arms by fitting models that incorporated combinations as different nodes in the network. This approach was used to assess the assumption of additivity of combined treatments. We also investigated the impact of the ordering of primary or secondary treatment components by fitting a model in which the order was ignored (eg, physical therapy with massage was assumed to be equivalent to massage with physical therapy).Secondary treatment components (online supplemental data 5). We assessed the impact on effect estimates of when secondary treatments were included by conducting a sensitivity analysis excluding those interventions with a secondary treatment component.Because some osteopathic interventions may have included visceral techniques not declared in the original methods of the study, the impact of removing these interventions from the manual treatment node was examined.In response to reviewer feedback, and to further evaluate the robustness of the findings to potential inconsistency, we conducted an additional sensitivity analysis based on the comparison of individual deviance contributions under the base case consistency model and the unrelated mean effects model. For each data point, we calculated the ratio of its mean deviance contribution under the unrelated mean effects model to its contribution under the consistency model. A ratio >1.5 was considered indicative of potential inconsistency. Studies containing these data points were removed, and the network meta-analysis was re-estimated.Certainty assessment We used user written R code to semi-automatically perform the GRADE approach for a network meta-analysis 50 51 to assess the quality of the evidence (criteria listed in online supplemental data 11). The rating was finalised and checked by two independent assessors and any discrepancies were resolved by discussion. A range of equivalence of standardised mean differences, from −0.5 to 0.5,52 was used to evaluate imprecision and inconsistency.53 All GRADE evaluations were performed on back transformed scales; the corresponding outcomes and the range of equivalence were multiplied by the pooled SD of the corresponding scale.54 55 Publication bias was assessed by statistical and non-statistical methods.53 Indirectness was judged with Schünemann's approach.56 The risk of bias was downgraded by one level if >50% of participants were from studies with selection bias and performance bias. This criterion was selected because inadequate randomisation and lack of masking may exaggerate the intervention effect estimates.57–59 For the categories of imprecision and inconsistency, we downgraded by one level if some concerns and by two levels if major concerns existed. Indirectness was downgraded by one level if deemed serious and by two levels if deemed very serious. We downgraded by one level if publication bias was suspected.Drawing conclusions Our method for deriving conclusions from network meta-analyses followed recommendations from the GRADE working group for a minimally contextualised framework 60 and the process had six steps.Choosing a reference intervention. To guarantee reliable comparisons, the reference intervention should be the treatment with the highest degree of connectivity within the network. In our network, we anticipated no treatment control to be the most connected.Creating a decision threshold. To assess clinical significance, we used a minimal clinically important difference with a standardised mean difference of 0.5 SD, where the SD was pooled from the baseline (before the intervention) data of all included studies with a frequently occurring scale for each outcome (back pain and leg pain, visual analogue scale 0-100; disability, Oswestry disability index 0-100; mental health, Short Form 36 (SF-36) mental health component score). A 0.5 SD threshold has been established52 as a general distribution based benchmark. The use of 0.5 SD as a commonly chosen threshold for meaningful change has been established in subsequent methodological reviews61 62 and closely approximates established anchor based minimal clinically important differences.55Classification of interventions at first. Based on comparisons with the reference intervention, interventions were categorised into two categories (effective and ineffective) depending on whether the minimal clinically important difference was included in the credible interval of the treatment effect.Reclassifying interventions with pairwise comparisons. To find any interventions that were noticeably more effective, additional comparisons were made within the effective category. This reclassification was based on comparing the 95% credible intervals of each treatment with the minimal clinically important difference to see if any intervention exceeded the minimal clinically important difference, which would have moved the intervention to a higher rated group.Grouping interventions based on the certainty of evidence. Reviewers evaluated the certainty of evidence for each intervention in relation to the reference. Groups of interventions with low or very low certainty of evidence were separated from those with high or moderate certainty of evidence.Verifying consistency through pairwise comparisons and rankings. To confirm consistency in classification, we lastly looked at pairwise comparisons that had not been taken into consideration previously. To ensure that the interventions with the highest rankings were also among the most successful, we also looked over the rankings.Patient and public involvement Patients and members of the public were not involved in this study. The findings of this study will be disseminated to the public through mainstream and social media. Results will also be disseminated by all of the co-authors through their institutions.Results Study selection After removing 20 828 duplicate records with Covidence, 19 253 records remained for screening ( figure 1). During the screening phase, 15 221 records were excluded, leaving 4032 reports sought for retrieval. Of these, 50 reports could not be retrieved. In the next phase, 3982 reports were assessed for eligibility; 3401 reports were excluded for various reasons, such as population criteria, intervention concerns, publication type, study design, intervention or comparator problems, duplication, language, and outcomes (online supplemental data 12). Ultimately, 551 studies were included in the review, with 581 reports of these included studies (online supplemental data 13). The trials included people with non-specific chronic low back pain (n=510) pain and radicular chronic low back pain (n=41; table 1). Of the included reports, 194 were flagged as requiring input from the authors: we found contact email addresses for 155 (80%) authors, 49 (25%) responded, and 26 (13%) provided the requested data in part or in full. Tables 1 and 2 present the characteristics of the included trials. Median duration across all interventions was six weeks (interquartile range 6.9; table 1).Figure 1Study selection process. *Studies that were otherwise eligible, but had the same primary intervention in all study arms (eg, massage in both arms of a two arm study)Table 1Summary of included randomised controlled trials, and population and intervention characteristicsCharacteristicsOverallBack pain intensityLeg pain intensityDisabilityMental healthTotal No of unique trials included55144935357135Population Range of mean age (years); No of trials reporting(20.4-77.2); 535(20.4-77.2); 441(34.7-72.8); 34(20.4-77.2); 352(22.6-73.9); 132 Range of No of men (%); No of trials reporting(0-100); 518(0-100); 425(25-100); 34(0-100); 337(0-100); 130 Range of average pain duration (weeks); No of trials reporting(12-1058); 310(13-1058); 254(13-1028); 21(13-1028); 200(13-1058); 87 Type of back pain (No of trials; No. of participants)  Non-specific*510; 68 049425; 55 90112; 1612332; 45 834127; 23 804  Radicular syndromes†41; 307724; 150923; 170726; 17278; 587 Subtype of back pain (No of trials; No of participants)  Non-specific disc degeneration11; 6529; 5241; 1067; 3783; 81  Non-specific disc herniation, protrusion, or extrusion11; (774)11; (774)2; (117)7; 4251; 44  Non-specific failed back surgery4; (404)4; (404)1; (41)2; 1412; 141  Non-specific196; 19 767175; 17 6724; 652143; 14 05549; 6925  Non-specific not-stated (eg, back pain or mechanical back pain)286; 46 369224; 36 4444; 696172; 30 77271; 16 593  Non-specific spondylosis2; 832; 83NA1; 631; 20  Radicular pain, sciatica, or leg pain18; 13408; 55410; 7379; 6054; 300  Radiculopathy11; 8527; 4477; 5427; 5322; 97  Spinal stenosis12; 8859; 5086; 42810; 5902; 190Intervention or comparator No of trials with these treatment nodes  acu (Acupuncture)40324258  edu (Education)675735521  elc (Electrophysical agents)13010787625  exe (Exercise)2061661015054  man (Manual treatments and manipulation)786166017  mas (Massage)38301238  mck (McKenzie method)18176133  mul (Multidisciplinary pain management)312482413  pha (Pharmacotherapy)1048695026  pio (Physical therapy; otherwise not falling into specific treatment combination)7061105619  pla (Placebo or sham)19314429331  psy (Psychological treatments, including cognitive-behavioural therapies)353032618  tra (Traction)12939  tru (No treatment or true control)94705425  usu (Usual care, eg, management by doctor)42343115 Range of intervention comparator duration (weeks); No of trials reporting(0-52); 548(0-52); 447(1-26); 33(0-52); 354(1-26); 135 Median (IQR) intervention comparator duration (weeks); No of trials reporting(6±6.9); 548(6±6); 447(6±6.5); 33(6±6); 354(8±8); 135Key population personal (age, sex, pain duration, and type and subtype of back pain) and intervention characteristics are presented, including treatment types and duration. Overall refers to reports on all included trials; outcome specific columns report trials that provided data for a meta-analysis of that outcome.*Includes spondylosis, spondylolisthesis, disc herniation, disc degeneration, and failed back surgery.†Includes radicular pain, sciatica, leg pain, radiculopathy, and spinal stenosis.IQR, interquartile range; NA, not available.Table 2Summary of methodological and funding characteristics of included randomised controlled trialsCharacteristicsOverallBack pain intensityLeg pain intensityDisabilityMental healthDesign No of intervention arms included  Two arms3873172524790  Three arms1279898531  Four arms33242011  Five arms4331  Range of sample size per trial arm(3-1363)(3-1363)(4-141)(3-1363)(3-1363) Study design of included trials  Cluster randomised trial1  Randomised controlled or clinical trial52443333352132  Randomised crossover trial2616263Funding Industry or treatment provider pays554922912 No funding or non-profit (eg, public funding agency or own university)3192621921282 Not stated1741371411640 Patient pays (patient themselves, Medicare, work cover, Department of Veterans Affairs, or insurance company)3111Study design features (number of arms, sample size, and trial type) and sources of funding are reported across included trials. Overall refers to reports on all included trials; outcome specific columns report trials that provided data for a meta-analysis of that outcome.Some studies contributed data to more than one outcome domain: 232 trials contributed to two outcome domains, 112 trials contributed to three outcome domains, and nine trials contributed to all four outcome domains. Online supplemental data 13 provides detailed study level information on outcome reporting.Presentation of network structure, summary of network geometry, and study characteristics Figure 2 shows network plots for the included outcomes (online supplemental data 14 shows networks for each discrete time point network meta-analysis). Online supplemental data 15 provides a summary of the network geometry. Online supplemental data 16 gives the characteristics of each study in a summary table.Figure 2Network plots for (A) back pain intensity (k=442 for time course model), (B) leg pain intensity (k=34 for time course model), (C) disability (k=354 for time course model), and (D) mental health (k=133 for time course model). Line thickness is proportional to the number of studies for a given comparison and the size of the nodes is proportional to the number of participants at the earliest time point reported in each study. acu=acupuncture; edu=education; elc=electrophysical agents; exe=exercise; man=manual treatments and manipulation; mas=massage; mck=McKenzie method; mul=multidisciplinary pain management; pha=pharmacotherapy; pio=physical therapy; pla=placebo or sham; psy=psychological treatments; tra=traction; tru=no treatment or true control; usu=usual careRisk of bias within studies Online supplemental data 17 presents the risk of bias assessment for each study, with corresponding summary plots for each discrete time point and outcome as well as for all time points combined. For all outcomes, most studies showed a low risk of bias in random sequence generation (70.9% rated as low) and incomplete outcome data (80.9% rated as low). We found some concerns with allocation concealment (57.8% rated as unclear), however, and a high risk of bias in masking participants and staff (70.2% rated high), as well as masking the outcome assessment (61.6% rated as high), indicating potential influences on subjective outcomes. Most studies adhered to their prespecified outcomes without selective reporting (94.9% rated as low) and showed low risk from other bias sources (99.1% rated as low).Synthesis of results The minimal clinically important differences of 0.5 SD were calculated as 8.53 points, 8.98 points, 5.23 points, and 5.2 for back pain intensity, leg pain intensity, disability, and mental health, respectively, on a 0-100 scale. Table 3 provides an overview of the subgroup and sensitivity analyses and the impact on findings. Online supplemental data 18–20 provide details of the assessment for consistency, homogeneity, and transitivity, which are summarised for each outcome in the following sections. To model how treatment effects changed over time, we used spline functions, which allow smooth, flexible curves to describe non-linear improvement patterns rather than assuming a single linear trend. For back pain intensity and disability, random effects linear spline models with two random time course parameters and three knots at 0.2, 10, and 20 weeks provided the best fit, with deviance information criterion values of 178.131 and −184.146, respectively. For mental health, a similar model with one random time course parameter and the same knot placement showed the best model fit (deviance information criterion −133.918). All models converged successfully, and the results were robust to alternative functional forms and knot placements (online supplemental data 19).Table 3Overview of the impact on findings for back pain intensity, leg pain intensity, disability, and mental health from subgroup and sensitivity analysesBack pain intensityLeg pain intensityDisabilityMental healthImpact of baseline pain intensity on findings (by network meta-regression)No impact on findingsNot estimableNo evidence from pairwise meta-regression that disability was an effect modifierNo impact on findingsImpact of baseline disability on findings (by network meta-regression)No evidence from pairwise meta-regression that back pain was an effect modifierNot estimableNo evidence from pairwise meta-regression that disability was an effect modifierNo evidence from pairwise meta-regression that mental health was an effect modifierImpact of pain duration on findings (by network meta-regression)No impact on findingsNot estimableNo evidence from pairwise meta-regression that disability was an effect modifierNo evidence from pairwise meta-regression that mental health was an effect modifierNon-specific low back pain subgroup onlyMcKenzie no longer clinically relevant at immediate term. Multidisciplinary pain management no longer clinically relevant at short term and massage no longer clinically relevant at intermediate termNot estimableNo impact on findingsNo impact on findingsRadicular low back pain subgroup onlyTime course network meta-analysis not possible. McKenzie remains clinically relevant at short termNot estimableTime course network meta-analysis not possible. All treatments no longer clinically relevantNot estimableDo treatment effects of radicular low back pain differ from non-specific low back pain (by network meta-regression)?No impact on findingsNot estimableNo impact on findingsNot estimableWhether having co-interventions (ie, a secondary treatment component in contrast with one primary treatment) influenced treatment effects (by network meta-regression)No impact on findingsNot estimableNo impact on findingsNot availableExclusion of studies with imputed SDs or imputed meansMcKenzie no longer clinically relevant at immediate term. Massage and multidisciplinary pain management no longer clinically relevant at short term and massage also no longer relevant at intermediate termNot estimableMassage no longer clinically relevant at short termMultidisciplinary pain management clinically relevant at short termDoes excluding small studies (ie, inclusion of only studies with ≥20 participants in total) change the findings?Massage, exercise, and multidisciplinary pain management no longer clinically relevant at short term and massage also not relevant at intermediate termNot estimableMassage no longer clinically relevant at short termNot estimableAnalysis only of studies that had no dropouts or correctly imputed missing dataNot estimableNot estimableNot estimableNot estimableDoes having a model with treatments within class (class effects model) impact the findings?Not estimableNot estimableNot estimableNot estimableExclusion of physical therapy (otherwise not falling into specific treatment combination) nodeNot estimableNot estimableMassage no longer clinically relevant at short termNo impact on findingsExclusion of multidisciplinary pain management nodeNo impact on findingsNot estimablePharmacotherapy clinically ineffective at intermediate term (became worse than true control)No impact on findingsExclusion of osteopathic and chiropractic interventionsMassage no longer clinically relevant at short termNot estimableNo impact on findingsNo impact on findingsNot estimable is because of model non-convergence or lack of data, or both. No impact on findings means the model fit was similar or worse after the change in the analysis and conclusions based on mean estimates remained unchanged.SD, standard deviation.Back pain intensity For back pain intensity, 442 unique studies were available for analysis. Figure 3 shows the results for the predicted time course of back pain intensity. In the immediate term, only the McKenzie method was more effective than the reference treatment (mean difference −16.32, 95% credible interval −18.66 to −9.65; GRADE evaluation of the quality of the evidence, very low). The analysis identified acupuncture (mean difference −20.91, 95% credible interval −24.00 to −11.95; GRADE very low), electrotherapy (mean difference −18.98, −21.84 to −10.95; GRADE low), exercise therapy (mean difference −15.59, −17.51 to −10.05; GRADE very low), manual treatment (mean difference −19.48, −22.17 to −11.74; GRADE very low), massage (mean difference −25.61, −30.42 to −10.91; GRADE very low), and multidisciplinary pain management (mean difference −18.96, −22.26 to −9.58; GRADE very low) as exceeding the minimal clinically important difference at the short term follow-up but the evidence was very uncertain. All other treatments did not show a clinically important treatment effect at the short term follow-up. For intermediate follow-up times, only massage showed a clinically significant effect (mean difference −22.89, 95% credible interval −26.59 to −11.83; GRADE very low) compared with true control but with very low certainty. At long term follow-up times, although no treatment provided clinically important benefits, the effects for manual treatment and massage were significant (figure 3). Online supplemental data 19 shows rankings for the network for different time courses and forest plots for discrete time points.Figure 3Plot of time course predictions relative to true control from network meta-analysis: back pain intensity. Blue shaded areas indicate the number of data points supporting the predictions. The minimal clinically important difference of 0.5 SD is represented by red dotted horizontal lines and set at 8.5 points for back pain intensity (range 0-100). acu=acupuncture; edu=education; elc=electrophysical agents; exe=exercise; man=manual treatments and manipulation; mas=massage; mck=McKenzie method; mul=multidisciplinary pain management; pha=pharmacotherapy; pio=physical therapy; pla=placebo or sham; psy=psychological treatments; tra=traction; tru=no treatment or true control; usu=usual care; MBNMA=model based network meta-analysisIn exploring potential reasons for heterogeneity, we examined possible effect modifiers, subgroups, and performed sensitivity analyses.Effect modifiers. Although pairwise meta-regression suggested that baseline pain intensity and pain duration might be effect modifiers, subsequent network meta-regression showed that these variables did not affect the findings (online supplemental data 18).Non-specific and radicular chronic low back pain separately. We evaluated non-specific chronic low back pain and radicular chronic low back pain separately in subgroup analyses. Because the bulk of the evidence was for non-specific chronic low back pain, the overall findings (short term effectiveness only, but not long term) were not affected in this subgroup. In terms of individual treatments, the McKenzie method was no longer clinically relevant in the immediate term, multidisciplinary pain management was no longer clinically relevant in the short term and massage was no longer clinically relevant in the intermediate term (online supplemental data 19).Radicular versus non-specific chronic low back pain. By network meta-regression for short term follow-up, we could not find evidence that treatment responses for radicular chronic low back pain differed from non-specific chronic low back pain, and the 95% credible intervals were very wide (online supplemental data 18).Exclusion of nodes with potentially higher levels of treatment heterogeneity (multidisciplinary pain management, pharmacological, and physical therapy nodes). Exclusion of nodes with potentially higher treatment heterogeneity (multidisciplinary pain management, pharmacological, and physical therapy nodes) produced results that were broadly consistent with the main analysis, with minor, but not clinically meaningful, changes in the effectiveness of some interventions (online supplemental data 19)Only one treatment component versus more than one treatment component. We evaluated whether the use of more than one treatment component (ie, co-interventions or secondary treatments) affected the findings in the short term. This effect was assessed by network meta-regression, because studies with secondary treatment components did not converge in the time course model (online supplemental data 18 and 19). For the back pain outcome, network meta-regression showed that having a secondary treatment component in a study arm did not influence the findings for the short term.Publication bias and small study effects. We assessed publication bias and small study effects by examining funnel plots and Egger's tests (online supplemental data 18). We found evidence indicating small study effects only for education versus exercise (Egger's P=0.001) and electrotherapy versus placebo (Egger's P=0.002; online supplemental data 18). When small studies (n <20) were excluded from the analysis, massage, exercise, and multidisciplinary pain management were no longer clinically relevant in the short term and massage was no longer clinically relevant in the intermediate term (table 3; online supplemental data 19).Imputed data. Excluding studies with imputed or missing SDs and imputed means resulted in the McKenzie method being no longer clinically relevant in the immediate term. Massage and multidisciplinary pain management were no longer clinically relevant in the short-term and massage was also no longer relevant at the intermediate term (online supplemental data 19).Stability of treatment rankings. Treatment rankings for the time points immediate term, short term, and long term were sensitive to a change in the data for the threshold analyses at the study and contrast levels (online supplemental data 14), showing that the individual treatment rankings should be treated with caution.To assess the consistency assumption, we compared individual residual deviances, between study heterogeneity estimates, and dev-dev plots between the base case consistency model and the unrelated mean effects model. For these diagnostics, we did not identify any coherent pattern indicating inconsistency: residual deviance contributions were similar across models, τ estimates were stable, and dev-dev plots showed no systematic departures. Small numerical fluctuations in total residual deviance were expected given the model uncertainty and the size of the network, and were not interpreted as evidence of inconsistency. Sensitivity analyses removing data points with notably better fit under the unrelated mean effects model did not substantially change the results (online supplemental data 19). Overall, these findings indicate that the assumption of consistency was adequately met for this outcome.Leg pain intensity A time course network meta-analysis for leg pain intensity could not be performed because of lack of data and hence we present the analysis of the discrete time point network meta-analysis ( online supplemental data 18 and figures 4 and 5). Thirty five studies were included in the analysis and data were available to analyse at the short term, intermediate term, and long term time points only. In the short term, acupuncture treatment exceeded the minimal clinically important difference (mean difference −37.15, 95% credible interval −63.52 to −10.7; GRADE evaluation of the quality of the evidence, very low) but with very low certainty. Based on intermediate term and long term results, no treatment provided clinically important reductions in leg pain intensity compared with no treatment (true control).Figure 4Forest plot for back pain intensity. Back pain intensity data are from exemplary predicted time points from the time course network meta-analysis. Mean differences are relative to no treatment control. Minimal clinically important difference=8.53 points. Reference treatment is no treatment or true control. acu=acupuncture; edu=education; elc=electrophysical agents; exe=exercise; man=manual treatments and manipulation; mas=massage; mck=McKenzie method; mul=multidisciplinary pain management; pha=pharmacotherapy; pio=physical therapy; pla=placebo or sham; psy=psychological treatments; tra=traction; tru=no treatment or true control; usu=usual care; NE=not estimable; NA=not availableFigure 5Forest plot for leg pain intensity. Leg pain intensity data are from the discrete time point network meta-analysis. Mean differences are relative to no treatment control. Minimal clinically important difference=8.98 points. acu=acupuncture; edu=education; elc=electrophysical agents; exe=exercise; man=manual treatments and manipulation; mas=massage; mck=McKenzie method; mul=multidisciplinary pain management; pha=pharmacotherapy; pio=physical therapy; pla=placebo or sham; psy=psychological treatments; tra=traction; tru=no treatment or true control; usu=usual carePlots of the SD between studies for each comparison indicated variability between the studies and showed significant heterogeneity for each discrete time point analysis (online supplemental data 18). We could not explore the potential reasons for heterogeneity or perform sensitivity analyses (online supplemental data 18–20) because of lack of data (table 3).In assessing the consistency assumption for the discrete time point network meta-analyses, plots for comparing the residual deviance of the base case model with the unrelated mean effects model for the discrete time points indicated no inconsistency. Comparison of the residual deviances and deviance information criterion values also showed no indication of inconsistency. Between study SDs were lower for the unrelated mean effects models for the intermediate term and long term time points (online supplemental data 14).Disability In the analysis for disability, 354 unique studies were included ( figure 6, figure 7). Acupuncture (mean difference −10.52, 95% credible interval −11.84 to −6.59; GRADE evaluation of the quality of the evidence, very low), massage (mean difference −9.95, −11.45 to −5.50; GRADE very low), and multidisciplinary pain management (mean difference −12.56, −13.91 to −8.55; GRADE very low) provided clinically important effects at the short term follow-up compared with true control but with very low certainty. All other treatments did not provide a clinically important benefit for reducing disability. No treatment provided clinically important benefits in the immediate, intermediate, or long term. At the long term follow-up, acupuncture, education, electrotherapy, exercise, manual treatment, massage, multidisciplinary pain management, and psychological treatments showed significant benefits that did not exceed the minimal clinically important difference (figure 6). Pharmacotherapy had significantly poorer outcomes than no treatment in the long term, which also remained below the minimal clinically important difference (figure 6). Online supplemental data 19 has details on model selection, rankings for the network, and forest plots.Figure 6Plot of time course predictions relative to true control from network meta-analysis: disability. Blue shaded areas indicate the number of data points supporting the predictions. The minimal clinically important difference of 0.5 SD is represented by red dotted horizontal lines and set at 5.24 points for disability (range 0-100). acu=acupuncture; edu=education; elc=electrophysical agents; exe=exercise; man=manual treatments and manipulation; mas=massage; mck=McKenzie method; mul=multidisciplinary pain management; pha=pharmacotherapy; pio=physical therapy; pla=placebo or sham; psy=psychological treatments; tra=traction; tru=no treatment or true control; usu=usual care; MBNMA=model based network meta-analysisFigure 7Forest plot for disability. Data are from exemplary predicted time points from a time course network meta-analysis. Mean differences are relative to no treatment control. Minimal clinically important difference=5.24 points. Reference treatment is no treatment or true control. acu=acupuncture; edu=education; elc=electrophysical agents; exe=exercise; man=manual treatments and manipulation; mas=massage; mck=McKenzie method; mul=multidisciplinary pain management; pha=pharmacotherapy; pio=physical therapy; pla=placebo or sham; psy=psychological treatments; tra=traction; tru=no treatment or true control; usu=usual care; NE=not estimable; NA=not availableTo explore potential reasons for heterogeneity, we examined possible effect modifiers, subgroups, and performed sensitivity analyses (online supplemental data 18–20).Effect modifiers. Baseline back pain intensity, baseline disability, and pain duration were not found to be effect modifiers (ie, higher or lower values of these variables in the included studies did not affect the effect sizes; table 3 and online supplemental data 18).Non-specific and radicular chronic low back pain separately. Focusing on the non-specific low back pain subgroup only did not change the findings (online supplemental data 19) because the bulk of the data came from this subgroup. Similar to the back pain outcome, a time course network meta-analysis was not possible for radicular low back pain.Radicular versus non-specific chronic low back pain. Attempting network meta-regression of treatment effects of radicular low back pain versus non-specific low back pain at the short term follow-up did not affect the findings, but uncertainty was very high (95% credible intervals were wide; online supplemental data 18).Exclusion of nodes with potentially higher levels of treatment heterogeneity (multidisciplinary pain management, pharmacological, and physical therapy nodes). When these were excluded, the results remained largely consistent with the main analysis, except for a reduced effectiveness of massage in the short term and pharmacotherapy performing worse than true control (online supplemental data 19).Only one treatment component versus more than one treatment component. In assessing the use of more than one treatment component (ie, co-interventions or secondary treatments) we found that, similar to the back pain intensity outcome, the presence of co-interventions did not modify the effects (online supplemental data 18).Publication bias and small study effects. Evidence of publication bias or small study effects, or both, was present for manual treatment versus placebo (Egger's P=0.01), education versus exercise (Egger's P=0.0005), and exercise versus true control (Egger's P=0.04) (online supplemental data 18). When small studies (n <20) were excluded from the analysis, massage was no longer clinically relevant at the short term follow-up (table 3 and online supplemental data 19).Imputed data. Excluding studies with imputed or missing SDs and imputed means resulted in massage no longer being clinically effective in the short term (online supplemental data 19).Stability of treatment rankings. Rankings of individual treatments against each other were not robust to threshold analyses (online supplemental data 14) and, similar to the back pain intensity outcome, should be treated with caution.Consistency was evaluated with residual deviance patterns, heterogeneity estimates, and dev-dev plots. Although we saw minor numerical differences in residual deviances between the consistency and unrelated mean effects models, particularly for some time course parameters, these differences did not form a coherent pattern or suggest meaningful inconsistency. Estimates of τ were highly similar across models. As recommended in current methodological guidance, these small model fit differences were not used as evidence for or against inconsistency. Sensitivity analyses based on deviance ratio screening criteria identified a small number of potentially influential data points, but exclusion of these studies did not substantially change the conclusions (online supplemental data 19). Overall, the diagnostics indicated good consistency between direct and indirect evidence.Mental health In the analysis for mental health outcomes, 133 unique studies were included ( figures 8 and 9). The key finding of the analysis was that none of the treatments had a credible interval that exceeded the minimal clinically important difference at any follow-up time point. Multidisciplinary pain management (mean difference −8.52, 95% credible interval −9.95 to −4.19; GRADE evaluation of the quality of the evidence, very low) was the only treatment that was close to exceeding the minimal clinically important difference at the short term follow-up. Online supplemental data 19 has details on model selection and additional forest plots for discrete time points, as well as rankings across the 52 week period.Figure 8Plot of time course predictions relative to true control from network meta-analysis: mental health. Blue shaded areas indicate the number of data points supporting the predictions. The minimal clinically important difference of 0.5 SD is represented by red dotted horizontal lines set at 5.2 points for mental health (range 0-100). acu=acupuncture; edu=education; elc=electrophysical agents; exe=exercise; man=manual treatments and manipulation; mas=massage; mck=McKenzie method; mul=multidisciplinary pain management; pha=pharmacotherapy; pio=physical therapy; pla=placebo or sham; psy=psychological treatments; tra=traction; tru=no treatment or true control; usu=usual care; MBNMA=model based network meta-analysisFigure 9Forest plot for mental health. Data are from exemplary predicted time points from a time course network meta-analysis. Mean differences are relative to no treatment control. Minimal clinically important difference=4.81 points. Reference treatment is no treatment or true control. acu=acupuncture; edu=education; elc=electrophysical agents; exe=exercise; man=manual treatments and manipulation; mas=massage; mck=McKenzie method; mul=multidisciplinary pain management; pha=pharmacotherapy; pio=physical therapy; pla=placebo or sham; psy=psychological treatments; tra=traction; tru=no treatment or true control; usu=usual care; NA=not availableTo explore potential reasons for heterogeneity, we performed analyses for possible effect modifiers, subgroups, and sensitivity analyses:Effect modifiers. Although baseline back pain intensity, but not baseline disability or pain duration, was identified as a potential effect modifier in pairwise meta-regression, network meta-regression did not find evidence for an effect on the findings (table 3 and online supplemental data 18).Non-specific and radicular chronic low back pain separately. Analysing only studies with non-specific low back pain did not affect the findings. Analysis of the radicular low back pain subgroup only and network meta-regression was not possible because of lack of data.Exclusion of nodes with potentially higher levels of treatment heterogeneity (multidisciplinary pain management, pharmacological, and physical therapy nodes). Excluding these nodes did not result in any clinically meaningful differences compared with the main analysis (online supplemental data 19).Only one treatment component versus more than one treatment component. Having more than one treatment component (ie, co-interventions or secondary treatments) did not influence the findings. Model comparison between the base case model and the network meta-regression model showed that the presence of co-interventions was not an effect modifier (online supplemental data 18).Publication bias and small study effects. We found no evidence of publication bias or small study effects for any pairwise comparison (online supplemental data 18).Imputed data. Excluding studies with imputed or missing SDs and imputed means did not affect the findings (online supplemental data 19).Stability of treatment rankings. Similar to the other outcomes, threshold analysis showed that the rankings of individual treatments against each other were not robust (online supplemental data 14).For the mental health outcome, we examined residual deviance contributions, heterogeneity estimates, and dev-dev plots to evaluate consistency. Some individual deviance contributions differed between the consistency and unrelated mean effects models, but these fluctuations were small, showed no systematic direction, and are expected given estimation uncertainty. Total residual deviance and τ estimates were largely comparable across models, and the dev-dev plots did not show consistent departures suggestive of true inconsistency. Sensitivity analyses excluding data points with better unrelated mean effects fit (>1.5 deviance ratio) did not substantially change the results (online supplemental data 19). Taken together, these findings indicate that the consistency assumption was reasonably met for this outcome.Certainty of evidence Online supplemental data 21 shows the overall certainty of the evidence available for each comparison in the network meta-analyses with discrete time points, assessed with the GRADE method. Almost all comparisons were rated with very low certainty (98.6%; 1.4% were rated with low certainty), mainly from inconsistency (94.2% of all comparisons), risk of bias (91.7%), and imprecision (87.2%). Incoherence (27.9%), indirectness (0%), and publication bias (1.1%) did not contribute frequently.Discussion Principal findings In this systematic review with network meta-analysis, we evaluated the efficacy of multiple conservative treatments for chronic low back pain over time. The bulk of the evidence was for non-specific low back pain. The key finding was that many treatments were similarly effective in the short term (≥1 day and ≤3 months) for some outcomes, with peak treatment efficacy at about 10 weeks after randomisation, but no individual approach was more effective than others. In the long term (≥12 months), treatments may not be more effective than having had no treatment at all. Multicomponent treatment (ie, with a secondary treatment component, such as education or acupuncture in addition to another primary treatment mode, such as exercise or manual treatment) did not yield better results. The findings of a clinically significant effect of massage and multidisciplinary pain management were not robust to subgroup and sensitivity analyses. Although sensitivity analyses did not provide evidence of a different response in radicular chronic low back pain, the evidence base for radicular chronic low back pain in isolation was limited and inconclusive.Comparison with other studies Few previous reviews have investigated the long term effects of conservative treatments. Reviews for individual treatment approaches, however, including psychological, 63 exercise,64 and multidisciplinary pain management65 treatments, indicated a similar attrition in relative treatment effects over time. Our study added the estimated trajectory of changes over time, with peak treatment efficacy at about 10 weeks after randomisation, corresponding to the duration of interventions which were typically 6-8 weeks in the included studies. The short lasting effect sizes of conservative treatments were in line with results from previous systematic reviews.66 Another recently published work examined the long term effects of interventions for back pain67 and found that some interventions could produce long term improvements. This conclusion was based on statistical, not clinical, significance. Further, without the time course approach of our work, the study was less able to account for patterns that indicated the change in treatment efficacy over time.Based on our findings, a single course of treatment of this duration may provide clinically important effects, but a prolonged effect on chronic symptoms is unlikely. Chronic back pain typically persists and is not self-limiting.68 Methods of boosting long term treatment efficacy are needed, and could include stronger treatment fidelity in trials, facilitating sustained patient compliance, a longer initial treatment delivery phase, scheduling booster treatment sessions across a longer time period, implementing sustainable self-management strategies, or individualising treatment intensity and titrating treatment received based on patient needs. A recent randomised controlled trial69 on cognitive functional therapy incorporating booster sessions and high treatment fidelity showed clinically important effects on pain and disability at the one year follow-up, but these findings could not be attributed to the booster sessions. Existing evidence from the few available randomised controlled trials incorporating booster treatments,70 71 however, suggests only limited additional efficacy of booster treatment approaches. Furthermore, a recent meta-analysis of health coaching72 and a new large scale randomised controlled trial of a lifestyle intervention73 both reported no clinically meaningful benefit compared with usual care. We argue that although lifestyle changes, and potentially booster treatments, may be important for managing chronic low back pain, they remain an important evidence gap.Education, advice, shared decision making, and long term behavioural and lifestyle changes should be part of the standard care of chronic pain.74 Our findings provide further arguments that individual courses of treatment should be deprioritised in favour of long term management (or self-management) and behavioural change. Other chronic diseases (eg, hypertension and diabetes) are managed with a view to long term management and behavioural change.75 76 Future randomised controlled trials should consider the long term trajectory of symptoms and design adequate interventions to target long-term symptom management.Our findings showed that highly disparate treatment approaches gave similar outcomes for patients. A treatment such as acupuncture has a markedly different purported mechanism of action77 than, for example, exercise.78 79 Given our findings of similar outcomes, whether these disparate postulated mechanisms of action can simultaneously be true is questionable. Research on contextual factors (eg, patient-practitioner interaction), non-specific effects (eg, natural history), and specific effects (ie, actual physiological mechanism of action) show that specific mechanistic effects have only a small role.80 81 Assuming that these effects explain the similar efficacy found in our study is reasonable. Implementing double blinding and controls to study the efficacy or mechanisms of treatment has historically been difficult for some conservative treatments.82 Nonetheless, to assess whether specific conservative treatments for chronic low back pain have a role in patient care, robustly implementing these in future randomised controlled trials is necessary. Future research should consider well designed, blinded, efficacy, or mechanistic trials of conservative treatments for chronic low back pain, with assessment of long term follow-up after treatment, alongside an economic analysis. Also, although randomised controlled trials are the most robust design for establishing causal effects, they may not fully capture the contextual and multifactorial nature of conservative pain interventions. Pragmatic trials and well designed observational or comparative effectiveness studies could therefore provide valuable complementary evidence, particularly in evaluating real world implementation and accounting for the strong placebo and contextual components inherent in pain management.83In our review, we did not assess safety because adverse events are typically poorly reported in randomised controlled trials and broader efforts looking at multiple study types are recommended.27 Nonetheless, the conservative interventions we examined typically have a good safety profile. A recent overview of Cochrane reviews84 evaluated non-pharmacological and non-surgical interventions for back pain. This review noted that whether these treatments had a higher incidence of adverse events compared with no treatment or usual care was not clear. Reviews of specific interventions noted that minor adverse events have been observed, such as with massage (eg, increased pain85), traction (eg, aggravation of symptoms86), acupuncture (eg, bleeding and pain at the needling site87), and exercise (ie, increase in non-serious but not serious adverse events in general88). A review of spinal manipulative therapy89 noted that assessing the safety of this intervention on the basis of the available data was difficult.Guidelines commonly recommend active treatment modalities, such as exercise, with passive treatment modalities (eg, manual treatment, acupuncture, electrotherapy, and massage) being deprioritised.5 Based on our findings, however, if outcomes such as pain and disability are the main rationale for choosing treatment modalities, a preference for active over passive treatments does not seem reasonable. As our findings showed, acupuncture, electrophysical agents, and analgesia have been tested more often against placebo controls than against treatments such as exercise and multidisciplinary pain management. Thus some treatments may benefit in guideline recommendations from not having been tested against adequately designed control groups that aim to test the efficacy or mechanism of the intervention. For sustained improvements to outcomes, management approaches enabling a patient to self-manage their back pain may be preferable to avoid patient dependency in chronic conditions.90 In the absence of causal evidence that passive treatments increase dependency or healthcare utilisation, however, and supported by the findings of our study, passive treatments seem to be a feasible treatment option to consider. Also, previous systematic reviews have suggested that active as well as passive interventions can be considered cost effective91 and safe88 89 in treating chronic low back pain. Because efficacy seemed to be similar across the treatment options and outcomes that we studied, clinicians and guideline author groups may need to reconsider what approach, or argument, is used to determine treatment preference. Given our findings, patient preference can drive many treatment decisions, assuming that healthcare systems have the resources to support multiple treatment options that are effective in the short-term only.The evidence was most limited by inconsistency (heterogeneity), risk of bias of the included studies, and imprecision. Risk of bias in future randomised controlled trials can be improved by implementing masking and adequately designed controls robustly. Inconsistency and imprecision could be improved by implementing and including in systematic reviews only randomised controlled trials with large sample sizes, enhanced and consistent reporting in randomised controlled trials, such as detailed description of population and interventions, and use of core outcome measures.92 Further reductions in heterogeneity could be achieved by improving diagnostic accuracy to differentiate the heterogeneous group of patients with non-specific low back pain and better target treatments, or the use of more granular methodology, such as individual patient data meta-analysis, dose-response, or component network meta-analysis.Strengths and limitations of this study A key strength of this study was the use of time course models that allowed us to study the relation between study efficacy and time. Also, our review was comprehensive, accepting all key conservative treatments, with a large number of studies and participants analysed. Potential limitations included pooling of similar treatments into one treatment class (eg, antidepressants and paracetamol in the pharmacotherapy class) which may mask subgroups within each class. Evidence for this finding is weak, however, and current evidence suggests similar efficacy across treatments within the same class, such as exercise treatments. 93 Despite the large number of treatments considered in our study, we acknowledge that we did not examine every possible intervention for low back pain, including behavioural and lifestyle modification based approaches. Not all patients are willing to participate in trials.Selection bias can limit the generalisability of findings from trials and thus from meta-analyses on the basis of these trials. Although potentially important based on our findings, we did not include interventions on or extract data for long term repeated (booster) or sustained self-management strategies, because these strategies are inconsistently recommended in international guidelines and seldom supported in healthcare systems. Existing evidence70 71 from the few available randomised controlled trials suggests only limited additional efficacy of booster treatment approaches. We followed robust guidance for choosing our minimal clinically important difference, but future work should consider that a minimal clinically important difference, or a smallest worthwhile effect, may differ across short, intermediate, and long term follow-up times and depend on the perspective of an individual versus societal benefit.We planned a priori to combine evidence on non-specific low back pain and radicular syndromes. Sensitivity analyses excluding radicular syndromes changed the relative clinical effectiveness of some treatments (eg, the McKenzie method, multidisciplinary pain management, and massage) at specific time points, suggesting that the presence of radicular syndromes may influence the comparative effectiveness of these interventions. Network meta-regression analyses, however, found no evidence that radicular syndromes consistently changed relative treatment effects across all comparisons, because model fit did not improve compared with the base case model. Given that radicular syndromes are often associated with higher baseline pain intensity because of nerve root involvement (eg, sciatica), these syndromes may rather act as a prognostic factor,94 influencing absolute outcomes, such as pain severity or disability. Further research is needed to clarify whether radicular syndromes show a different treatment response.The literature search for this review was completed on 24 July 2020. Although more recent randomised controlled trials have since been published, the comprehensiveness of our dataset, comprising 551 trials and more than 71 000 participants, makes it the largest synthesis of conservative treatments for chronic low back pain so far. To evaluate whether additional evidence would likely change the conclusions, we conducted trial sequential analyses (online supplemental data 7), which showed that the required information size had already been reached for most comparisons and outcomes, indicating that further small or moderate sized trials are unlikely to meaningfully change effect estimates. Moreover, the more recent studies largely investigated interventions already well represented in our dataset. For these reasons, in addition to workload feasibility, a full update was not performed. Nonetheless, we acknowledge that the search period represents a limitation and that future updates should incorporate new evidence as it accumulates to ensure continued relevance of the findings.Conclusions Current treatments for non-specific chronic low back pain were effective only in the short term, which corresponded to the end of the intervention in most included randomised controlled trials. Improvements did not persist long term. Although our analyses showed that treatment effects did not differ for radicular chronic low back pain, the evidence base was limited for this subpopulation. Both passive and active treatments had similar short term efficacy. Facilitating long term efficacy needs to be considered in future work, such as improving treatment fidelity and patient compliance, or adding booster treatment sessions. Reframing the management of chronic low back pain in line with other chronic conditions (eg, hypertension and diabetes) should also be considered, with self-management approaches, shared decision making, education, and lifestyle modifications.SP210.1136/bmjmed-2025-001908.supp2Supplementary data",
  "title": "Conservative treatments for chronic non-specific low back pain: time course network meta-analysis",
  "uid": "f627f57a-20f7-57e1-b302-dc1c6dd507ee"
}
