{
  "abstract": "Introduction Type 1 diabetes mellitus (T1DM) is a chronic autoimmune disease targeting insulin-producing cells in the pancreas. The rising global incidence, particularly in early childhood, suggests environmental triggers, such as infections, may contribute to its pathogenesis. Prior studies have reported spatiotemporal clustering of T1DM, and we aimed to further investigate spatial and spatiotemporal clustering in Finnish children using high-quality data with complete residential histories.Research design and methods We included patients under 18 diagnosed with T1DM between 1990 and 2019, identified from the Finnish Social Insurance Institution, based on insulin reimbursement. Each case was assigned three age-matched and sex-matched controls. Clustering was analyzed using the Cuzick-Edwards test, Knox test, and Jacquez’s Q statistic. Multiple testing adjustments were applied using the Benjamini-Hochberg correction.Results The study included 16 307 cases and 48 914 controls (median age at diagnosis: 8.9 years; 56% male). The Cuzick-Edwards test identified modest spatial clustering among males 1 year prior to diagnosis, while the Knox test revealed significant spatiotemporal clustering across all cases. Analyses incorporating full residential histories confirmed these findings, with more pronounced spatiotemporal clustering in children over 6 years old.Conclusions These results demonstrate evidence of spatiotemporal clustering of T1DM in Finnish children, supporting the hypothesis of environmental triggers in T1DM etiology. These findings highlight the need for further research to identify the specific environmental factors and mechanisms behind the clustering.",
  "authors": [
    {
      "affiliations": [
        "Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland"
      ],
      "name": "Julia Ventelä"
    },
    {
      "affiliations": [
        "Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland"
      ],
      "name": "Heikki Hyoty"
    },
    {
      "affiliations": [
        "Tampere Center for Child, Adolescent, Maternal Health Research, Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland",
        "Tays Cancer Center, Tampere University Hospital, Tampere, Finland"
      ],
      "name": "Olli Lohi"
    },
    {
      "affiliations": [
        "Tampere Center for Child, Adolescent, Maternal Health Research, Faculty of Medicine and Health Technology, Tampere University, Tampere, Finland",
        "Tays Cancer Center, Tampere University Hospital, Tampere, Finland"
      ],
      "name": "Atte Nikkilä"
    }
  ],
  "full_text": "WHAT IS ALREADY KNOWN ON THIS TOPIC Incidence of childhood type 1 diabetes is rising globally, suggesting a role for environmental triggers.Most previous studies of spatiotemporal clustering have used single residential location, and no prior research has examined the first year of life using residential histories.WHAT THIS STUDY ADDS We examined whether childhood type 1 diabetes cases in Finland show spatial or spatiotemporal clustering, which would indicate environmental influences.Using complete residential histories from more than 16 000 children diagnosed between 1990 and 2019 and matched controls, we found strong spatiotemporal clustering, particularly among children older than 6 years.HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY These findings reinforce the role of environmental factors in type 1 diabetes development and underscore the need to pinpoint specific triggers and critical exposure windows in future research.Introduction Type 1 diabetes mellitus (T1DM) is an autoimmune disorder marked by the targeted destruction of insulin-producing β-cells within the pancreatic islets, resulting in an absolute deficiency of insulin and necessitating lifelong exogenous insulin therapy. 1 In 2021, the prevalence of T1DM was 8.4 million people globally, with nearly 1.5 million cases diagnosed in children and adolescents.2 Finland has one of the highest incidence rates worldwide, with a rate of 56.8 per 100 000 person-years in individuals under 20 years from 2000 to 2022.3The pathogenesis of T1DM is complex, involving genetic predispositions, most notably within the HLA region, and environmental triggers. A key role is played by autoantibodies directed against islet cell antigens. Several autoantibodies have been identified as markers of T1DM pathogenesis and prediction. These include anti-islet antibodies to insulin (IAA), glutamic acid decarboxylase (GADA), tyrosine phosphatase-like protein IA-2 (IA-2A) and zinc transporter 8 (ZnT8A).4 5 The timing of autoantibody appearance varies: IAA typically emerges between 9 and 24 months of age, GADA around 3 years, while IA-2A and ZnT8A often appear later.6–8The rise in global T1DM prevalence has been widely studied with growing interest in potential environmental triggers. Nutritional factors, especially vitamin D deficiency, are suspected of contributing to beta-cell autoimmunity and T1DM development.9–11 Viral infections, particularly enteroviruses, have also been suggested as potential etiological agents.12–14 Additionally, T1DM incidence shows distinct seasonal patterns, with a higher occurrence during autumn and winter, coinciding with the seasonal prevalence of viral infections.15–17 Conversely, the hygiene hypothesis proposes that early childhood exposure to common pathogens may offer protection against T1DM development.18 19Several studies have examined the spatiotemporal clustering of childhood T1DM, particularly in high-prevalence regions.20–29 These studies consistently identified clustering both temporally and spatially. Studies have identified clustering, with certain clusters showing associations with sex, age, and socioeconomic factors.24 27 28 Although incidence is generally higher in males,30 31 some studies report more pronounced clustering in females.24 Additionally, a study on adults found significant T1DM clusters with a slightly higher proportion of women.32 No study has evaluated the pathogenetically important first year separately with complete residential histories.Previous studies have primarily assessed T1DM clustering at specific time points or spatial locations. In this study, we investigated spatial and spatiotemporal clustering of childhood T1DM within the Finnish population using comprehensive residential histories. Furthermore, we separately analyzed residential locations from the first year of life to explore possible early-life spatiotemporal clustering of T1DM.Research design and methods T1DM cases were identified through the Social Insurance Institution of Finland using special reimbursements for insulin medication. The inclusion criteria covered individuals aged 0 to <18 years who were diagnosed between 1990 and 2019. The cases included individuals who received reimbursement for insulin therapy (reimbursement codes 103, 171, 177, 371, and 382). To improve data accuracy, T1DM diagnoses were manually verified using International Classification of Diseases, 10th revision (ICD-10) code group E10* and International Classification of Diseases, 9th revision codes 250, 2500, 2501, 2500B, 2502B, and 2507B (with “B” indicating insulin-dependent diabetes) from the special reimbursement decisions. Individuals with no contradicting ICD-10 entries were included.For each case, three age-matched and sex-matched individual controls were randomly selected from the Finnish Population Information System. Each control was assigned a reference date at which they were the same ages as their individual case when diagnosed. The study period ended on the diagnosis date for cases and the corresponding reference date for controls.Complete residential histories for cases and controls were obtained from the Finnish Digital and Population Data Services Agency (DVV). The dataset included start and end dates for each residence period, as well as municipality, postal code, and street address information. Each address was geocoded by DVV and linked to latitude and longitude coordinates based on the ETRS-TM35FIN planar coordinate system, the official national coordinate reference system in Finland. These coordinates represent the centroid of each residential building location and were used for all spatial and spatiotemporal analyses.Data were limited to addresses from birth to the date of diabetes diagnosis, excluding addresses associated with omitted cases and their matched controls. This resulted in a dataset of 161 326 rows of residential history data. Among cases, 6.1% (n=998) lacked coordinates for certain residential periods, with a similar proportion observed in controls (6.9%, n=3373). Consequently, 5213 rows of data were excluded from the analysis. The final dataset comprised 156 113 rows, with 38 625 entries for 16 272 cases and 117 488 entries for 48 763 controls. Residential histories were complete, with no missing coordinates or dates, for 15 104 cases (92.6%) and 43 946 controls (89.8%) (figure 1).Figure 1Flow chart of the study population and their residential data. †Study cohort following the exclusion of cases with non-diabetes-related medications and duplicate records, along with their matched controls.Statistical analysis All statistical analyses were conducted using R (V.4.0.5, R Core Team, 2021, Vienna) and Spacestat (BioMedware, Ann Arbor, MI, V.4.0.21) in a secure remote environment managed by the Finnish Social and Health Data Permit Authority, Findata. 33We analyzed clustering with five specific approaches: complete residential histories, first-year residential history, the addresses at the time of diagnosis, the addresses 1 year preceding the diagnosis, and the addresses at birth. Clustering was assessed with a significance level set at p<0.05.To assess purely spatial clustering, we applied the Cuzick-Edwards test.34 This test analyzes the distribution of cases relative to controls among the k=5 nearest neighbors. Clustering is identified when the observed number of cases exceeds expectations. We performed the analysis using the ZceTk() function from the nnspat R package (V.0.1.2). Due to computational limitations in the secure analysis environment, a 1:1 case–control sampling strategy was used for the primary analysis, while the complete dataset at a 1:3 ratio was used for subgroup analyses. For the overall population, a random sample comprising 50% of cases and an equal number of controls was selected, and this process was repeated 30 times. The final result represents the median of these iterations.The Knox test examines spatial and temporal clustering of cases within predefined distance (0.25, 0.5, 1, 5, 10 km) and time (2, 6, 12, 18, 24 months) thresholds.35 This test was implemented using the knox() function from the surveillance R package (V.1.20.0), with 1000 iterations. Statistical significance was assessed by deriving p values from Monte Carlo permutation tests.Jacquez’s Q statistic was applied to assess spatiotemporal clustering by examining the k=5 and k=15 nearest neighbors of cases relative to controls, using their complete residential histories. The analysis was performed using the commercial software SpaceStat (V.4.0.21, BioMedware).36 37 This method incorporates four indicators of clustering, with global Q evaluating the overall clustering pattern across the study period. A higher global Q value indicates more persistent clustering of cases over time. The local Q-statistic evaluates clustering at the individual residential level, while the Qi statistic examines clustering over an individual’s lifetime. The Qt statistic assesses clustering at a specific point in time, and Qit analyzes clustering at a particular time for a specific individual. To address computational constraints of the available secure environment, significance testing was conducted with 99 iterations, which restricted the precision of p value with lowest possible estimate to 0.01.Due to computational and software limitations when analyzing complete residential histories, the full dataset could not be processed simultaneously. Therefore, the study population was divided into two approximately equal subsets based on latitude for data handling purposes. To minimize the risk of missing clusters near the division line, a 10% spatial overlap was introduced between the subsets. This overlap was intended to ensure continuity of cluster detection across boundaries. As an additional sensitivity analysis, we also performed segmentation based on residential time periods using the same overlap approach. Furthermore, we analyzed each individual’s residential location history during their first year of life using two latitude-based and two period-based subsets, each incorporating a 10% overlap.To control the false discovery rate arising for multiple comparisons, the Benjamini-Hochberg (BH) procedure was used for the Cuzick-Edwards test and Knox test.38 To address the multiple comparisons issue with Jacquez Q statistic, we applied the multiple-testing parameter to specify the number of randomizations to use to 99.39For potential heterogeneity in clustering patterns, the dataset was stratified by sex and age groups (0–5.99, 6–11.99, and 12–17.99 years) with statistical analyses repeated independently for each subgroup. Additionally, we conducted the Knox test on subgroups of individuals with a single residence prior to diagnosis and on those diagnosed during specific time periods (1990–1999, 2000–2009, and 2010–2019), providing further insight into spatial clustering across demographic and temporal strata.Results Dataset The final dataset included 16 307 diabetes cases and 48 914 controls. The median age at diagnosis was 8.9 years (IQR: 5.1–12.5), with a male predominance of 56.1% (n=9142) ( table 1). Analysis of seasonal trends in diagnosis revealed a significant increase in cases diagnosed during autumn and winter months compared with summer (figure 2). On average, participants resided in 2.40 locations during the study period. Cases had slightly fewer residencies (mean 2.37) compared with controls (mean 2.41).Figure 2Monthly distribution of new type 1 diabetes diagnoses, shown as a percentage of total annual cases.Table 1Characteristics of the study populationCasesControlsTotal16 30748 914Sex Female7165 (43.9%)21 495 (43.9%) Male9142 (56.1%)27 419 (56.1%)Age (years) 0–5.995062 (31.0%)15 185 (31.0%) 6–11.996640 (40.7%)19 917 (40.7%) 12–17.994605 (28.2%)13 812 (28.2%)Residential history information Complete15 104 (92.6%)43 946 (89.8%) Residence at birth15 845 (97.2%)46 878 (95.8%) Residence at 1 year prior to diagnosis16 086 (98.6%)48 363 (98.9%) Residence at diagnosis16 164 (99.1%)48 695 (99.6%)Clustering We evaluated potential clustering of T1DM cases using three methods. Spatial clustering analysis with the Cuzick-Edwards test, using the five nearest neighbors, identified significant but modest clustering only among male cases at 1 year before diagnosis (observed/expected ratio [Obs/Exp]=1.02, 95% CI 1.01 to 1.03, BH corrected p=0.004). This corresponds to a 2% increase over the expected value, indicating that the spatial clustering detected by this method was limited in magnitude. Clustering related to birth among male cases showed borderline significance (Obs/Exp=1.01, 95% CI 1.00 to 1.03, p=0.08). No significant clustering was identified across age groups using the Cuzick-Edwards test ( online supplemental table S1).SP110.1136/bmjdrc-2025-005547.supp1Supplementary dataThe Knox test, which evaluates clustering based on the proximity of cases in both space and time, revealed significant clustering among cases, even at short distances and time intervals. Clustering was detected across nearly all tested time and distance thresholds when analyzed relative to diagnosis, the previous year, and birth date (online supplemental table S2). However, when the spatial thresholds were expanded to 20, 50, and 100 km, significant clustering by birth date was no longer observed.SP210.1136/bmjdrc-2025-005547.supp2Supplementary dataFemale cases demonstrated more pronounced clustering across various timeframes and distances compared with male cases, regardless of the date considered (online supplemental table S3). Age-based analysis indicated exclusive birth-related clustering for the 12–17.99 age group, while younger subgroups exhibited clustering at all three time points (online supplemental table S3). Additionally, clustering was observed during certain diagnosis periods (1990–1999 and 2010–2019), as well as among cases where individuals had a single residential location, with clustering related to the diagnosis date and the year before.SP310.1136/bmjdrc-2025-005547.supp3Supplementary dataJacquez’s Q statistics, which incorporate full residential histories, revealed significant spatiotemporal clustering across all subgroups when analyzing the 15 nearest neighbors (online supplemental table S4). For females, the global Q statistic, which measures clustering over time, was 225 263 case-years (p=0.01), while for males, it was 320 107 case-years (p=0.01). Both sexes also exhibited significant clustering at the individual residential level (local Q), with the strongest clustering observed in 1993. Notably, reducing the number of nearest neighbors to five produced similar results (data not shown).SP410.1136/bmjdrc-2025-005547.supp4Supplementary dataAge-stratified analysis indicated more pronounced clustering among older children. No overall clustering was found for the youngest group (<6 years) (global Q=67 584 case-years, p=0.63), though significant local clustering (Qit) was detected (n=632, p=0.02). Children aged 6–11.99 and 12–17.99 years displayed significant clustering (global Q=226 160 and 251 191 case-years, respectively, p=0.01 for both). Consistent local clustering was seen in the oldest group, with peak clustering in 1989 for this group and in 2008 for the 6–11.99 year group.Geographically, the latitude-based division into two overlapping subgroups runs through central Finland, approximately across the regions of Pirkanmaa and South Savo (latitude ranges: 6 638 446–6 819 645 and 6 763 700–7 775 031, using ETRS-TM35FIN). Significant clustering was observed in both groups, with global Q values of 292 091 and 315 626, respectively (p=0.01). Both subgroups also displayed significant local Q values. Temporal analysis with the same overlap strategy for the periods 1972–2000 and 1998–2019 also revealed clustering (global Q=289 350 and 251 377, p=0.01 for both), with peak clustering observed in 1993–1995 for the earlier period and in 2006–2008 for the latter.First-year residential location analysis divided by latitude revealed significant clustering only in the northern regions (latitude 6 765 202–7 775 031), with a global Q value of 38 247 (p=0.03), and significant Qi and Qit values. No significant clustering was observed when dividing the dataset by time periods (online supplemental table S4).To illustrate spatial clustering, we mapped locations of significant Qit values across Finland for the period 1972–2019 (figure 3). The strongest clustering is located in the most densely populated area in the capital region. A summary of cluster analysis results across methods and residential time points is shown in figure 4.Figure 3Heat map illustrating the spatial distribution of significant local spatiotemporal clusters (Qit values) identified using Jacquez’s Q statistic for all T1DM cases combined between 1972 and 2019. The density surface was smoothed using a 30 km search radius. Darker shades indicate higher concentrations of significant clustering locations. Major Finnish cities are shown for geographic reference. T1DM, type 1 diabetes mellitusFigure 4Summary of spatial and spatiotemporal clustering results across methods, residential timing, and subgroups. Residence timing categories: A=At the time of diagnosis; B=1 year prior to diagnosis; C=At birth. Color codes: Red=Significant clustering, p<0.05, Light gray=No clustering, p≥0.05, Dark gray=Not analyzed. †The Knox test was applied for individuals with a single residence prior to diagnosis.Discussion We conducted an extensive investigation into the spatiotemporal clustering tendency of T1DM among children and adolescents in Finland, a country with a notably high incidence of T1DM. Using a large dataset with detailed residential histories from birth to diagnosis, we applied multiple clustering analysis methods to explore the geographic and temporal distribution of T1DM cases. Our findings revealed significant spatiotemporal clustering, suggesting that environmental factors and early-life exposures may contribute to disease risk. One plausible explanation for these clustering patterns is the involvement of infectious agents that spread within communities over limited spatial and temporal scales. Enteroviruses, in particular, have long been proposed as potential environmental triggers for β-cell autoimmunity, and localized outbreaks could account for the clustering observed in our data. 12–14While a seasonal pattern in T1DM incidence, with peaks in autumn and winter, was also observed, this aligns with previous research and likely reflects known seasonal variation linked to infection dynamics.15–17 The key novel contribution of our study, however, is the demonstration of spatiotemporal clustering patterns that extend beyond these expected seasonal effects.The Cuzick-Edwards test identified limited overall spatial clustering. Although the magnitude of the effect was small (Obs/Exp=1.02) and does not by itself indicate clinical significance, there is no established threshold that defines a clinically meaningful degree of clustering. In contrast, the stronger and more consistent clustering signals identified by the Knox test and Jacquez’s Q statistic, both of which incorporate spatiotemporal components, indicate that timing and location together may play a greater role in disease aggregation than geography alone.For all T1DM cases, the Knox test revealed significant clustering at all three examined time points within a 5 km radius and a 2-year time window, but also detectable at shorter spatial and temporal scales, for example, within 500 m and 6 months. The observation that clustering was most pronounced over short distances and time intervals supports the hypothesis that transmissible infections may contribute to localized disease onset in genetically susceptible individuals. Additionally, the clustering observed at birth locations among children diagnosed after 12 years of age is consistent with the hypothesis that environmental exposures during early life may initiate or modulate the autoimmune processes that culminate in T1DM onset years later.These findings are consistent with previous studies examining T1DM clustering. For example, a Swedish study identified significant clustering at birth with optimal cut-off values of 5 km and 7 months, while research in Chile reported the strongest clustering at diagnosis with thresholds of 750 m and 60 days.22 29 In England, a study showed significant spatiotemporal clustering at diagnosis with thresholds of 25, 35, and 50 km and time intervals of 90, 270, and 350 days (all p<0.05).21 Similarly, a Norwegian study using spatial scan statistics identified two clusters of elevated T1DM incidence among children in southern regions during the 1960s and late 1980s, with increases of 2-fold to 2.6-fold.26Given the critical role of the early years in T1DM etiology, studying exposures during this period could help clarify the environmental mechanisms underlying T1DM development. We conducted a separate analysis of residential locations during the first year of life. Interestingly, clustering was observed only in the northern regions (latitude 6 765 202–7 775 031) where the population is more sparsely distributed compared with the capital region. This finding may reflect latitude-related differences in ultraviolet exposure and vitamin D synthesis, regional variation in infection dynamics and viral circulation, or the greater environmental homogeneity of sparsely populated areas, which can amplify localized exposure patterns. A recent study in Utah, a US state with a population heavily concentrated in the northern region, identified more than 40 spatial clusters for T1DM and found a positive correlation between increased latitude and T1DM risk, while population density and median household income were negatively associated with T1DM incidence.28Subgroup analyses revealed that both sexes experienced notable clustering, with a peak clustering period in 1993. Based on full residential history analysis, older children, aged over 6 years, demonstrated the most significant clustering. McNally et al used the K-function method to examine T1DM clustering in England and found significant space-time clustering in the 10–14 and 15–19 year-old age groups in Yorkshire, while in north-east England, clustering was more prominent among case pairs involving at least one female or at least one individual from a densely populated area.23 24Regarding temporal trends, there was no consistent evidence of changes in clustering over time. While the Knox test indicated significant clusters for T1DM diagnoses in 1990–1999 and 2010–2019, the intervening years showed no clustering, even when broader spatial (10 km) or temporal (2 year) thresholds were applied. When considering full residential histories, clustering was most prominent in 1993–1995 and 2006–2008.The primary strength of this study is the inclusion of detailed residential histories which enabled the identification of specific periods and locations associated with T1DM risk, providing new insights into potential environmental contributions to disease onset. The important events for pathogenesis are not necessarily related to the address at birth or at diagnosis, and thus it is highly important to use complete residential histories when looking for clustering. Second, the dataset has comprehensive national coverage, leveraging data from the Social Insurance Institution of Finland and the Finnish Digital and Population Data Services Agency, which both have vast national coverage. The large sample size, comprising over 16 000 T1DM cases and nearly 49 000 controls, enhanced the statistical power of the study and allowed extensive analysis across various subgroups. Also, only a small percentage of residential history was excluded due to missing coordinates or documentation. The use of multiple methods to analyze clustering provided a thorough examination of potential spatiotemporal clustering. To address the issue of multiple comparisons, we applied the BH correction method to reduce the likelihood of false positives.The results of the study are slightly constrained by computational limitations of the Finnish secure environment, which required the adoption of a 1:1 case–control sampling strategy for the Cuzick-Edwards spatial analysis and a limited number of iterations for significance testing. These limitations reduce the depth of p value estimates, especially in the analysis using Jacquez’s Q statistics, where fewer iterations were performed and the lowest p value possible was 0.01. Additionally, dividing the study population by latitude and time period might limit the ability to detect more subtle clustering patterns, especially in the border zone.Conclusions Our findings of spatiotemporal clustering and seasonal variation in T1DM diagnoses support the hypothesis that environmental factors play a role in the development of the disease. Notably, clustering was observed during the first year of life, particularly in the northern regions of Finland. These results suggest that early-life exposures may play a critical role in T1DM onset. Future research should investigate potential contributing factors such as genetic predispositions, socioeconomic conditions, or environmental exposures to understand the mechanisms driving the observed spatiotemporal clustering.",
  "title": "Clustering patterns in Finnish type 1 diabetes patients: a nationwide register-based study",
  "uid": "cb454423-c89f-5e9c-9ff3-9172591baaf5"
}
