![]() | Content Disclaimer Copyright @2020. All Rights Reserved. |
Links : Home Index (Subjects) Contact StatsToDo
|
Layout of this page
This page presents all the discussions and programs related to cluster randomisation, in sample size determination during the planning stages, and the analysis of the data collected.
Explanations and References
The large amount of information are arranged in a number of nested collapsible panels, the contents of which can be shown or hidden by clicking on the panel header. Layout: this panel Explanations and References
Introduction & References
Javascript Programs Calculations
Cluster randomisation experiments are used in situations where individual research subjects cannot easily be randomly allocated
to receive different treatments. In this situation, research subjects are firstly grouped into clusters, and
experimental treatments are randomly allocated to clusters so that all members of a cluster receive the same experimental treatment.
Sample Size
Some of the reasons for using the cluster randomisation experiments are
The main difference between individual randomisation and cluster randomisation is that members of a cluster may be more similar to each other than to those from different clusters, so the effects of experimental treatment and cluster membership are confounded. There is therefore a need to introduce a correction for this possible confounding, the parameter Intraclass Correlation Coefficient ρ. ρ, conceptually, is the average of correlations between all possible pairs within a cluster. If subjects are randomly allocated to different clusters and the environments of all clusters are identical, then there should be no correlation between cases in any clusters, and ρ=0. If all members of a cluster produce the same results, then ρ=1. In both estimating sample size requirement and in the analysis of the data, therefore, the Intraclass Correlation Coefficient ρ is estimated. This is used to adjust the results of standard statistical procedures that are based on individual randomisation, so that the final results are appropriate for cluster randomisation. Cluster randomisation is a large subject, and StatsToDo provides but the two most basic and commonly used models, that of two group comparison for normally distributed measurements and binomially distributed proportions, as carried out by the algorithms in the calculations panels. ReferencesMachin D, Campbell M, Fayers, P, Pinol A (1997) Sample Size Tables for Clinical Studies. Second Ed. Blackwell Science IBSN 0-86542-870-0 p. 27-28Donner A and Klar N (2000) Design and Analysis of Cluster Randomisation Trials in Health Research. Arnold London ISBN 0 340 69153 0
This panel provides a discussion, and supports the calculations of sample size estimation for cluster randomization. These can be to calculate the number of clusters required, when the number of cases within each cluster is defined, or the number of individuals in each cluster when the number of cluster is defined
Analysis
Sample size can be estimated in two ways
The data collected from a cluster randomisation model is usually summarised. For example, when the outcome is a
normally distributed measurement, the sample size, mean, and Standard Deviation from each cluster is used for analysis.
When the outcome is a binomially distributed parameter (no/yes, true/false), the numbers of cases with positive and
negative outcomes from each cluster are usually used for analysis.
Examples
The mean value of a cluster, or the proportion of positive responses in a cluster, can be used as a measurement, and these can be used for statistical analysis using standard statistical algorithms, using the cluster as the basic sampling unit. A concern of such an approach is that it assumes all clusters to have the same sample size. This however is usually the case in cluster randomisation experiments as all clusters should have the same sample size at start, and only data loss during collection results in minor differences in sample size from different clusters. Donner's book suggests 3 methods of analysis at the cluster level that can be used.
When the outcome in a cluster randomisation experiment is a binomially distributed proportion, three additional statistical tests can be applied.
Example for Two Means
Technical Considerations
The conduct of a cluster randomisation exercise comparing two means is best demonstrated with the default example data used in both the sample size and analysis program panels. Please note that the data used is computer generated to demonstrate the procedures, and not real observation.
Example for Two Proportions
We wish to conduct a controlled trial on the effect of introducing additional fertiliser to the feeding ground of calves on their weight gain. As we cannot randomize the calves because fertilisers can only be applied to paddocks, we will use the cluster randomisation model, using each paddock as the cluster for randomisation. Step 1 : Find ssizindividual We use the following parameters to determine the sample size based on individual randomisation
We looked up the sample size requirement table in the sample size for unpaired means table page, and found that we will require 65 calves per group if we were to randomize on individual calves. In other words, ssizindividual, s = 65 Step 2 : Find Intraclass Correlation Coefficient ρ
We decided to obtain the likely Intraclass Correlation Coefficient ρ for our experiment by a pilot study, where we placed some calves into a number of paddocks, and measure their growth (weight gain over 3 months), and found the results as in the table to the left. We use the first program in the sample size panel of this page to obtain a workable Intraclass Correlation Coefficient, which is ρ = 0.0889. This is, as expected a very low level of correlation, as the calves are allocated at random. Step 3 : Sample Size Adjustment
Using the third program in the sample size calculation panel, we can view the sample size required in each cluster for a range of cluster numbers, as shown in the table to the left. It can be seen that, as the number of paddocks (clusters) to be used decreases, the number of calves per paddock (sample size per cluster) increases exponentially. The sample size reaches infinity when the number of cluster is below 6. Although the sample size per cluster continues to decrease as the number of clusters increases, the changes in the total group sample size becomes increasingly minor. From such an analysis, and depending on the costs (financial, time, and effort) of different aspects of the experiment, the most efficient combination with the same power can be selected. In this example, the best combinations would seem to be from 9 paddocks with 19 calves per paddock (total 171 calves per group) to 18 paddocks with 5 claves per paddock (total 90 calves per group), depending on whether managing calves or paddocks to be more costly. Step 4 : Analysis of Results
Following calculations in the previous section, we decided to use 9 paddocks (clusters) per group, placing 20 calves in each paddock. Those in Grp 2 were controls, and additional fertilisers were added to the Grp 1 paddocks. The calves were weighed at the beginning of the experiment and 3 months later. The weight gain in Kg was the outcome. The table to the right shows means and SDs in weight gain from each paddock (cluster). The results are analysed using the two sample t test, with and without adjustments for intraclass correlation (ρ)
The overall mean and SD of the two groups are as in the table to the left. The difference between the two means is 3.7Kg, the Standard Error of the difference is 0.62, and probability of Type I error (α) p<0.0001. These results have not been adjusted by the intraclass correlation coefficient (ρ), so are references only, for comparisons with the final results. Intraclass Correlation Coefficient, calculated from the data, is ρ = 0.0084. The adjusted Standard Error of the difference is now 0.67, and the 95% confidence interval of the difference is 2.27 to 5.11 Kgs. This is the final results, showing a significant difference between the two groups, and we can conclude that adding fertilisers to the feeding paddocks increases the growth of calves. Please note that, in this example, the correction is so small as to be trivial, because of low ρ value. This is expected as allocation to different paddockes are randomized, and there is not much difference between the paddocks, so there would not be much interaclass correlation. Options to analyse the data using non-parametric statistical methods: The first and the third columns of the input data, group designations and cluster means in the table above and to the right, can be used for non-parametric statistical tests. These tests can be used as a check and comparison with the adjusted two sample t test, because of uncertainties of normal distribution, or very small sample size within the clusters. StatsToDo provides these test is the unpaired difference programs page
The conduct of a cluster randomization exercise comparing two proportions is best demonstrated with the default example data used in both the sample size and analysis program panels. Please note again that the data is computer generated to demonstrate the procedures, and does not represent any real observations.
We wish to conduct an experiment on the effect of introduce a student encouragement protocol into schools to reduce absenteeism, defined as having missed at least 1 scheduled class in a term. Given that such a protocol has to be introduced to a whole school, we decide to use the cluster randomization model. Step 1 : Find ssizindividual We use the following parameters to obtain the required sample size as if randomization is based on individuals.
We looked up the sample size requirement table in the sample size for comparing two proportions page, and found that we will require 121 students per group. In other words, ssizindividual s = 121 Step 2 : Find Intraclass Correlation Coefficient ρ
We decided to carry out a pilot simulation, by examining the absenteeism in a number of schools, and found the results as in the table to the left (Pos=number with absenteeism present, Neg = number with no absenteeism). Please note : In a real pilot study, many more clusters would be used to obtain a stable and robust ρ.
We used the second program in the sample size panel to calculate the Intraclass Correlation Coefficient ρ. The program firstly convert the number of positives and negatives into 1s and 0s, then calculate the means and SDs for each cluster, as shown in the table to the right. From this table, we estimate the Intraclass Correlation Coefficient ρ to be 0.197 Step 3 : Sample Size Adjustment
Although we can use the third or fourth programs of the sample size calculations to calculate one at a time the required number of individual within each cluster when the number of cluster in each group is pre-determined, or the number of individuals in each cluster when the number of clusters per group is pre-determined, we can also use either program to test a range of combination. Using the third program to determine the number of individuals in each cluster, with a range of numbers of clusters, we produced the results as shown in the table to the left. It can be seen that, as the number of schools (clusters) to be included decreases, the number of students per school (sample size per cluster) increases exponentially. The sample size reaches infinity when the number of cluster is below 24. Although the sample size per cluster continues to decrease as the number of clusters increases, the changes in the total group sample size becomes increasingly minor. As the main costs of such a program are related to selecting and introducing the encouragement protocol into schools, the decision can be based on finding the minimum number of schools (clusters) that contain sufficient number of individuals (students). As most schools have more than 84 students, a reasonable decision can be to have 25 schools (clusters) per group. Step 4 : Analysis of Data
We recruited 50 schools that had at least 84 students each that are available for the study, and randomly divided the schools into 2 groups. Those in Grp 2 were controls and those in Grp 1 were introduced to the encouragement program. The number of absentees in one term were collated from each school and presented in the table to the right. For comparing two proportions, 4 options are available, 3 of which adjusted for intraclass correlation coefficient (ρ), and can be used to draw conclusions from the data
Options to analyse the data using non-parametric statistical methods: The group designation and proportion of positives in the table above and to the right, can be used for non-parametric statistical tests. These tests can be used as a check and comparison with the adjusted two sample t test, because of uncertainties of normal distribution, or very small sample size within the clusters. StatsToDo provides these test is the unpaired difference programs page
The adjustment of sample size for individual randomization ssizintividual by the Intraclass Correlation coefficient
ρ uses two formulae, one to estimate the number of clusters per group needed if the sample size per cluster is pre-determined, and the other the sample size in each cluster needed if the number of cluster is pre-determined.
The formulae for these calculations are obtained from the text book by Pinol et.al. (see references), and the results are checked against the tables provided in that text book. However the following three points should be noted.
All other calculations are based on the text book by Donner and Klar (see references) and the following points should be noted.
Sample size
Hints for data entry (Sample Size)
Data Analysis
The sample size panel provides 4 programs. The first 2 calculates Intraclass probability, and the last two calculates cample size
Javascript Program (Sample Size)
Estimate Intraclass Correlation Coefficient ρρ represents correlations between indivuduals within each cluster. This is used to correct the sample size required in cluster randomization. ρ is calculated using data collected in a pilot study
Estimate Sample Size
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
Data |
Calculate Intraclass Correlation Coefficient ρ from clusters of n, mean, and SD - Data a table with 3 columns - Each row is calculation for a single study - Col 1 is sample size of cluster - Col 2 is the mean of cluster - Col 3 is the Standard Deviation of cluster Calculate Intraclass Correlation Coefficient ρ from clusters of from clusters of binomial data - Data a table with 3 columns - Each row is calculation for a single study - Col 1 is group designation and should be 1 or 2 - Col 2 is the number of positives in the cluster - Col 3 is the number of negatives in the cluster Calculate sample size from Intraclass Correlation Coefficient ρ and number of cluster - Data a table with 3 columns - Each row the data for a study - Col 1 is sample size per group according to individual randomization - Col 2 is Intraclass Correlation Coefficient ρ - Col 3 is number of clusters proposed Calculate sample size from Intraclass Correlation Coefficient ρ and numbers in each cluster - Data a table with 3 columns - Each row the data for a study - Col 1 is sample size per group according to individual randomization - Col 2 is Intraclass Correlation Coefficient ρ - Col 3 is sample size within each cluster proposed |
The difference is mean(group 1) - mean (group 2), with the two groups in alphabetical order
α<0.05, or the 95% confidence interval not traversing the null value of 0, can be used to identify that the difference is statistically significant.
The program also produces a table with 2 columns, the group designation and mean for each cluster. This table can be used for non-parametric comparisons using programs in the unpaired difference programs page
The program next performs 3 tests to compare the proportions in the two groups. The 3 tests are similar, but assumes different distribution of data.
The program also produces a table with 2 columns, the group designation and proportion of positive cases for each cluster. This table can be used for non-parametric comparisons using programs in the unpaired difference programs page
|
Data |
Data for Cluster randomization Comparing 2 Means
Data is in 4 columns separated by spaces or tabs. - Each row the data summary from a cluster - Col 1 is group designation and should be 1 or 2 - Col 2 is the sample size of cluster - Col 3 is the mean value cluster - Col 4 is the Standard Deviation value of cluster Data for Cluster randomization Comparing 2 Proportions
|
