Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Finding stable clusterings of single-cell RNA-seq data [version 1; peer review: awaiting peer review]

Дата публикации: 30-07-2026 11:04:02

Background There is evidently no consensus on how to find stable clusterings of cells in scRNA-seq UMI count matrices. Methods Run a count matrix through a pipeline to obtain n cell clusters. Suppose that counts for more cells from the same experiment become available. Would including them change the result? Form the matrix containing both sets of counts, obtain n clusters, restrict this clustering to the initial cells and compare it with the initial clustering. If they are not consistent, conclude that the initial clustering is unstable. This is unrealistic, but reverse the perspective: given a clustering, process samples of half of the cells. If their clusters are consistent with those of all cells restricted to the samples, conclude that the clustering is stable. Divisive hierarchical spectral clustering is used. The mapping of the dendrogram to nested clusterings may be novel. Counts are transformed to points in Euclidean space. Positive affinities are defined for points that are k-nearest neighbors. The affinity equals the inverse of the distance between points. Ng, Jordan, and Weiss’ algorithm divides the points into two clusters. The normalized cut measures the clusters’ separation. Recursion generates a dendrogram. Set the length of the branch between a node and its daughters to the normalized cut. Nodes’ distances from the root define the mapping to nested clusterings. Analyze for all cells and multiple pairs of complementary samples. For a given number of clusters, compare each sample’s clustering and clusters with those of the full data set, providing stability measures. Results For three large data sets, this found clusterings compatible with published results, though with fewer clusters. Clusterings of two were judged to be stable. Conclusions It is feasible to identify stable clusterings of as many as 100,000 cells. Future research should explore using differential expression for validation.

Основное содержимое страницы с новостью.

3.1 Impact of data reduction

Table 4 summarizes the impact of initial data reduction. This does not include the effect of excluding outliers as discussed in section 2.8.

  • - Columns 1 and 2: the number of cells and genes retained after preliminary filtering. The Pearson residuals matrix has a column for each cell. Compared to Table 1, the number of genes is smaller because genes with nonzero counts on fewer than 50 cells were excluded. The number of cells is smaller for the Zhengmixeq data because some barcodes have duplicate values (see section 3.5). Cells with high count contributions from mitochondrial genes were excluded from the PBMC and monocyte data. Blood cells were excluded from the lung data.

  • - Column 3: the number of analysis genes – genes that are highly variable in every sample and the full data set. This is the number of rows of the Pearson residuals matrix.

  • - Column 4: the rank of the Pearson residuals matrix estimated by the optht program. This is the number of rows in the SVD representation matrix.

  • - Column 5: the number of cells retained for clustering after excluding kNN outliers.

Table 4. Impact of data reduction.Preliminary filtering AnalysisEstimatedCellsData set Cells GenesGenes Rank RetainedZhengmix4eq3,9805,837646353,902Zhengmix8eq3,9746,075565373,90668k PBMC68,28612,5156524867,792CD14 monocytes2,5583,726719112,53725k retinal24,76913,5521,0815124,10165k lung60,99317,4701,65930560,114100k breast cancer100,06421,3541,70443498,681
3.2 Genes that are highly variable in the full set of data may not be highly variable in every sample

The plots in Figure 1 show the relation between the mean SSQ of Pearson residuals calculated with all cells (Mg) and the number of cells with nonzero counts for genes retained after preliminary filtering. Black points represent analysis genes retained for downstream analysis: Sg and Sg (s) are among the 2,000 largest when calculated with all cells and with each sample s, respectively.

Some genes that are highly variable in the full set of data (large vertical coordinates) are not highly variable in every sample (red points). It is possible that some of these genes would be retained after excluding outlier cells as described in section 2.8.

3.3 Analytic estimates of the rank of the Pearson residuals

Ranks of Pearson residuals matrices were estimated with the optht program. Four other programs were considered. Although all five gave reasonable results on toy problems (matrices of known low rank with added noise) each of the others gave problematic results:

  • - Wide variation between results depending on the user-selected algorithm (one program)

  • - Rank estimates differing by an order for magnitude for very similar input matrices (one program)

  • - Long run times (two programs)

  • - Failure to find a solution (one program)

Results are tabulated in column 4 of Table 4 and plotted in Figure 2 using the format of the screeplot program in the Bioconductor PCAtools package [38]. The estimated ranks of the Pearson residuals for the lung and breast cancer data (305 and 434, respectively) are larger than values we have seen in the literature. The maximum value plotted on the horizontal axis is the number of analysis genes – the number of rows in the Pearson residuals matrix (column 3 of Table 4).

3.4 Variation of Euclidean distance between k-nearest neighbors

Our interest in the relation between kNN and Euclidean distance was motivated by Meila’s recommendation to exclude outliers before performing spectral clustering and by the work of Cooley et al. [39]

We illustrate with the PBMC and breast cancer data. The variation of Euclidean distance between k-nearest neighbors in the SVD representation of the PBMC data is summarized in Table 5. The first column contains statistics for the distance from a cell to its nearest neighbor, which ranges from 1.3 to 294, with a mean of 5.0 and standard deviation of 7.4. Subsequent columns list statistics for 2nd -nearest neighbor distance, 4th-nearest, continuing to 64th-nearest, and finally the maximum distance between cells. The diameter of the set equals 823 – the bottom right entry.

Table 5. 68k PBMC data: distributions of distance between kNN.1248163264Maximummean 5.0 5.4 5.8 6.1 6.6 7.1 7.7 707.9 stddev 7.4 8.0 8.7 9.7 10.7 11.8 13.4 6.3 min1.31.61.82.02.12.32.4507.110%2.32.42.62.72.93.03.2707.750%3.84.04.24.54.75.05.3708.275%5.55.96.36.67.07.47.9708.490%8.28.89.410.010.711.512.4708.795%10.611.412.213.214.315.717.3708.999%18.220.022.024.427.230.635.2711.3max294.3384.4389.0433.6461.9507.1551.0822.6

Clearly, kNN neighborhoods may not resemble neighborhoods defined with the Euclidean metric. For half of the cells, the distance to the nearest neighbor is less than 4 units – less than 0.5% of the diameter of the set of cells. However, there is a cell whose nearest neighbor is 75 times further away – 294 units distant, 36% of the diameter of the set of cells.

Cells that are exceptionally distant from their kNN are identified as outliers to be excluded. Outliers are defined by distances at least three standard deviations larger than the mean. For nearest neighbors, this threshold equals 27.1. Fewer than 1% of cells are outliers based on this criterion. Applying this to 2nd, 4th,…, and 64th-nearest neighbors excludes a total of 494 cells, retaining 67,792. The distributions of distances for the retained cells are summarized in Table 6. Excluding kNN outliers reduces the range of nearest neighbor distances by an order of magnitude and the diameter of the set by 80%.

Table 6. 68k PBMC data: distributions of distance between kNN after excluding outliers.1248163264Maximummean 4.6 4.9 5.2 5.5 5.9 6.3 6.8 129.3 stddev 2.7 3.0 3.2 3.5 3.8 4.2 4.8 3.5 min1.31.61.82.02.12.32.4103.910%2.32.42.62.72.93.03.2126.550%3.74.04.24.44.75.05.3130.075%5.55.86.26.66.97.37.8130.690%8.08.69.19.710.411.212.1131.295%10.210.911.712.613.614.816.3131.999%15.116.217.519.121.123.426.7136.0max28.231.634.237.640.644.659.7163.9

For the breast cancer data, the variation is greater. Before excluding kNN outliers, nearest neighbor distance varies by a factor of 580 ( Table 7). Even after excluding 1.4% of the cells, the maximum is 50 times larger than the minimum ( Table 8).

Table 7. 100k breast cancer data: distributions of distance between kNN.1248163264Maximummean 48.6 54.0 59.4 65.1 71.1 77.2 83.1 6,394.7 stddev 95.6 112.6 129.8 148.7 168.1 184.7 197.7 47.6 min8.59.39.610.210.711.512.56,013.610%18.119.620.821.822.924.025.16,390.150%33.435.938.040.042.044.246.56,391.075%50.654.858.662.165.769.773.86,392.290%79.386.994.2101.5108.9116.7124.96,394.295%107.9120.1133.8148.2163.7180.5196.96,397.299%254.8316.4400.1490.4598.5705.0803.26,454.4max4,918.75,131.65,195.45,547.05,810.06,046.16,290.68,740.9

Table 8. 100k breast cancer data: distributions of distance between kNN after excluding outliers.1248163264Maximummean 42.2 45.8 49.2 52.8 56.6 60.7 65.1 1,189.1 stddev 31.0 34.6 38.9 44.0 50.1 57.1 65.2 18.6 min8.59.39.610.210.711.512.51,050.610%18.019.520.721.822.823.925.11,181.650%33.135.637.639.641.643.746.01,185.375%49.653.557.160.664.167.971.91,189.890%75.782.389.095.6102.3109.4116.81,198.295%98.4108.4118.4129.2141.3153.6166.21,207.799%168.5190.1215.0246.6285.3328.3385.91,266.6max427.9429.6502.6535.2604.7640.8892.81,566.0
3.5 Clustering results for individual data sets

In section 2.1 we proposed considering a clustering stable if the 90th percentile of normalized MED is less than or equal to 0.10. A cluster is judged stable if the 90th percentile of normalized CMER is less than or equal to 0.50. A stable clustering is admissible for downstream analysis if its unstable clusters have fewer than 500 cells.

Begin with the three small data sets. For the Zhengmixeq data we are interested in (1) the relation between the ground truth labels and our method’s clusterings and (2) the stability of specific clusterings and clusters. For the Zhengmix4eq data, agreement with ground truth labels is excellent; for the Zhengmix8eq data less so, though typical of what we have found in publications. For the monocytes, our results suggest that there are no stable clusterings.

For each of the four large data sets, as outlined in section 2.8, three sets of analyses were performed, progressively excluding outlier cells and genes. Six clusterings are reviewed:

  • - PBMC: two clusterings; an admissible 12-clustering and an unstable 9-clustering

  • - retinal: an admissible 11-clustering

  • - lung: two admissible clusterings; one with 19 clusters, the other with 16

  • - breast cancer: an inadmissible clustering compatible with published results

Zhengmix4eq

The data set contains counts for 15,568 genes and 3,994 cells. They represent four cell types. Ground truth labels are provided with the data.

Seven barcodes appear twice. The corresponding 14 columns were dropped, retaining counts for 3,980 cells. Two EnsemblIDs have the same gene symbol (SRSF10). Data for the EnsemblID with nonzero counts on fewer cells were excluded.

Filtering to exclude genes with nonzero counts on fewer than 50 cells retained 5,837 genes. Because data were not input through Seurat, cells were not screened for high mitochondrial DNA levels. 646 genes were found to be highly variable in the full data set and in all 40 samples.

The rank of the Pearson residuals matrix was estimated as 35 by optht. After mapping to a 35-dimensional SVD representation and excluding kNN outliers, 3,902 cells were retained.

Figure 3 displays the distributions of normalized MED for the clusterings of sizes 2–10. There is one line per clustering. The 40 vertically jittered dots in each line show the samples’ MED. For each clustering, the blue vertical segment marks the median of the distribution. The 75th percentile marker is green. The 90th percentile marker is red. The clusterings of sizes 2–5 are stable. Their 90th percentiles are less than or equal to 0.10. For the clusterings of sizes 2–4, the plotted percentiles equal 0.00. Only the blue markers are visible because we favor the lower percentile markers.

Because there are 4 ground truth labels, we review the 4-clustering. The green highlighted line displays its MED values.

Figure 4 displays the distributions of normalized CMER for each cluster. All four are stable. The largest value of CMER equals 0.017. The plot shows the median, 75th, and 90th percentiles of CMER for each cluster. They are indistinguishable for clusters 0, 2, and 3. In addition, the colored vertical lines that extend from the bottom to the top of the plot (all very close to 0 on the horizontal axis) indicate the clustering’s percentiles – the ones marked in Figure 3.

Table 9 compares the clusters with the ground truth labels.

Table 9. Zhengmix4eq data: compare ground truth with the 4-clustering.Ground truthHierarchical cluster0123 Totalnaive.cytotoxic9771600993regulatory.t598600991b.cells009970997cd14.monocytes092910921Total 982 1,011 999 910 3,902

Zhengmix8eq

The data set contains counts for 15,716 genes and 3,994 cells. They represent eight cell types or subtypes. Ground truth labels are provided with the data.

Ten barcodes appear twice. The corresponding 20 columns were dropped, retaining counts for 3,974 cells. Two EnsemblIDs have the same gene symbol (SRSF10). Data for the EnsemblID with nonzero counts on fewer cells were excluded.

Filtering to exclude genes with nonzero counts on fewer than 50 cells retained 6,075 genes. Because data were not input through Seurat, cells were not screened for high mitochondrial DNA levels. 565 genes were found to be highly variable in the full data set and in all 40 samples.

The rank of the Pearson residuals matrix was estimated as 37 by optht. After mapping to a 37-dimensional SVD representation and excluding kNN outliers, 3,906 cells were retained.

Figure 5 displays the distributions of normalized MED for the clusterings of sizes 2–10. Differences with Figure 3 are immediate: values for the clusterings of sizes 2 and 3 are large. The clusterings of sizes 4–8 are stable. The clusterings of sizes 7 and 8 are reviewed (green lines).

Figure 6 displays the distributions of normalized CMER for the 7 clusters. All are stable.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure6.gif

Figure 6. Zhengmix8eq data: distributions of normalized CMER for the 7-clustering.

Table 10 compares the clusters with the ground truth labels. Results are very accurate for cd56.nk, b.cells, and cd14.monocytes; less accurate for memory.t and naive.cytotoxic cells. Three subtypes – cd4.t.helper, naive.t, and regulatory.t – are commingled in clusters 0 and 1. The adjusted Rand index equals 0.74.

Table 10. Zhengmix8eq data: compare ground truth with the 7-clustering.Ground truthHierarchical cluster0123456Totalcd4.t.helper25813820000398naive.t485702000494regulatory.t18630810000495memory.t13343526000495naive.cytotoxic115389010397cd56.nk005258910597b.cells000004930493cd14.monocytes050011530537Total 931 492 448 419 590 496 530 3,906

The 8-clustering is formed from the 7-clustering by splitting its cluster 4 (590 cells) into the two clusters 0 (339 cells) and 1 (251 cells). Figure 7 displays the distributions of normalized CMER for the 8 clusters. All are stable, but the two new clusters (top two lines) are much less stable than the cluster from which they were formed.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure7.gif

Figure 7. Zhengmix8eq data: distributions of normalized CMER for the 8-clustering.

Table 11 compares the clusters with the ground truth labels. The adjusted Rand index equals 0.68.

Table 11. Zhengmix8eq data: compare ground truth with the 8-clustering.Ground truthHierarchical cluster01234567Totalcd56.nk339250005210597cd4.t.helper002581382000398naive.t0048570200494regulatory.t001863081000495memory.t001334352600495naive.cytotoxic0011538910397b.cells0000004930493cd14.monocytes0105001530537Total 339 251 931 492 448 419 496 530 3,906

CD14 Monocytes

The data set contains counts for 32,738 genes and 2,612 cells. Filtering to exclude genes with nonzero counts on fewer than 50 cells and to exclude cells with more than 5% of counts due to mitochondrial DNA retained 3,726 genes and 2,558 cells. 719 genes were found to be highly variable in the full data set and in all 40 samples.

The rank of the Pearson residuals matrix was estimated as 11 by optht. After mapping to an 11-dimensional SVD representation and excluding kNN outliers, 2,537 cells were retained.

Figure 8 displays the distributions of normalized MED for the clusterings of sizes 2–10. No other data set reviewed in this article has so many large values for small clusterings. The median is greater than 0.50 for all clusterings. This is consistent with all cells being of the same type, so that any clustering would be spurious.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure8.gif

Figure 8. CD14 Monocytes: distributions of normalized MED for clusterings of sizes 2–10.

68k PBMC

The data set contains counts for 32,738 genes and 68,579 cells. Filtering to exclude genes with nonzero counts on fewer than 50 cells and to exclude cells with more than 5% of counts due to mitochondrial DNA retained 12,515 genes and 68,286 cells. 652 genes were found to be highly variable in the full data set and in all 40 samples.

The rank of the Pearson residuals matrix was estimated as 48 by optht. After mapping to a 48-dimensional SVD representation and excluding kNN outliers, 67,792 cells were retained.

Our objective was to evaluate clusterings of sizes up to 25 – larger than the 10 reported in the paper [18]. Clusterings of all sizes in the range 2–25 were found in each of the three sets of analyses. For the first iteration, the clusterings of sizes 2, 3, 6, 11, and 12 are stable. The largest is admissible, as shown below. The second iteration, using fewer cells, also yields a stable 12-clustering, but it is not admissible. It has an unstable cluster of 1,716 cells. The third iteration yields two stable clusterings, but they are small, with 2 and 3 clusters.

Figure 9 displays the distributions of normalized MED for the clusterings of sizes 2–25 found with the first iteration. The green line indicates the 12-clustering.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure9.gif

Figure 9. 68k PBMC data (iteration 1): distributions of normalized MED for clusterings of sizes 2–25.

Figure 10 displays the distributions of normalized CMER for the 12-clustering. The two smallest clusters, 8 and 11, are unstable. Their data lines are highlighted red because CMER = 1 with all samples. The remaining clusters are stable. The clustering is admissible for downstream analysis because the unstable clusters have fewer than 500 cells. The blue, green, and red lines extending from the bottom to the top of the plot indicate the median, 75th, and 90th percentiles of MED for the clustering.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure10.gif

Figure 10. 68k PBMC data (iteration 1): distributions of normalized CMER for the 12-clustering.

Cell assignments to the ten k-means clusters summarized in Figure 3 of the paper [18] are evidently not publicly available. To attempt to compare clusterings found by our method with published results, the input data (68,579 cells) were clustered (k-means) using code from a 10x Genomics Github repository (see the section “Data and software availability” below). The sizes of the 10 clusters we obtained do not agree with the percentages in Figure 3b (nor is it obvious how these clusters correspond to the ones discussed in the paper). In particular, we obtained one small cluster – of 176 cells. The remaining 9 clusters range in size from three thousand to eighteen thousand cells. When the clusters obtained with our process were matched to these, none of the 176 cells in the smallest cluster were included – filtering discarded all of them.

Table 12 compares the 9 k-means clusters with the 12 hierarchical clusters. The adjusted Rand index equals 0.55.

  • - The majority of cells in k-means cluster 10 are split three ways. Most are in hierarchical cluster 7, which is stable. Approximately 300 cells each are in clusters 8 and 11, which are unstable.

  • - K-means cluster 9 is split four ways. Cells are divided among hierarchical clusters 0,1,9, and 10. Differential expression may help determine if these clusters are meaningful or spurious.

  • - K-means clusters 6 and 1 correspond to hierarchical clusters 2 and 3, respectively.

  • - 60% of the cells of k-means cluster 5 account for 80% of hierarchical cluster 6.

  • - 95% of cells in k-means cluster 4 are grouped with 90% of the cells in k-means cluster 7 in hierarchical cluster 5.

  • - 90% of cells in k-means cluster 2 are combined with nearly all of k-means cluster 3 and cells from other clusters into hierarchical cluster 4.

Table 12. 68k PBMC data (iteration 1): compare 9k-means clusters with the 12-clustering.k-means clusterHierarchical cluster01234567891011Total6005,211274208271204006,330100323,871001201003,9072002310,6381,59869310880112,4303001042,8480604325322,959400011,01017,1932190100118,236700034846,724246211007,254500894542,57785,0981709018,2531000031003,4353317192914,08796541,09102000001,98060904,336Total 654 1,091 5,344 4,615 17,560 25,523 6,081 3,526 336 2,135 631 296 67,792

We review a second analysis of the 68k PBMC data because it yields a clustering more compatible with the k-means clusters. It is the result of the third iteration described in section 2.8. After deleting outliers, 63,281 cells were retained. Figure 11 displays the distributions of normalized MED. Only the 2 and 3-clusterings are stable.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure11.gif

Figure 11. 68k PBMC data (iteration 3): distributions of normalized MED for clusterings of sizes 2–25.

Preliminary exploratory analysis using the median of MED instead of the 90th percentile to define stable clusterings led to consideration of the 9-clustering (green highlighted line) because it is the largest with median MED (blue marker) less than or equal to 0.10. It fails to satisfy the criterion we now propose to define a stable clustering. The 90th percentile of MED equals 0.21.

The confusion matrices in Tables 2 and 3 are for two of the samples summarized in Figure 12, which displays the distributions of normalized CMER for the 9 clusters.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure12.gif

Figure 12. 68k PBMC data (iteration 3): distributions of normalized CMER for the 9-clustering.

Cluster 3 (top line) is unstable. CMER = 1 with 39 samples. Cluster 1 (6th line) is also unstable. The 75th percentile of normalized CMER equals 1. Cluster 6 is unstable by a narrow margin. The remaining clusters are stable.

Table 13 compares the 9 k-means clusters with the 9 hierarchical clusters. The adjusted Rand index equals 0.66.

  • - Most of the cells in the unstable cluster 3 belong to k-means cluster 1. Recall that almost all of the cells in the two unstable clusters of the 12-clustering belong to k-means cluster 10.

  • - Each k-means cluster except 3 is unambiguously associated with a hierarchical cluster.

Table 13. 68k PBMC data (iteration 3): compare 9 k-means clusters with the 9-clustering. k-means clusterHierarchical cluster012345678Total416,3712981001,0423101917,74471,0585,499100481251357,0821003,0862862801003,401600173205,1672591055,95821,3885130010,47765862312,0933005072,725491532,80451743211042,1345,2225147,920900000002,77202,77210000000003,5073,507Total 18,818 5,855 3,701 307 5,306 16,861 5,956 2,901 3,576 63,281

25k retinal

The counts downloaded from the Gene Expression Omnibus contain data for 49,300 cells. We followed Lause et al. [19] by restricting to replicates p1 and r4-r6, netting counts for 22,292 genes and 24,769 cells. Thirty-nine cell clusters were reported ( Figure 5D of) [20]. Restricted to the cells we analyzed, cluster sizes range from 14 to 15,709 cells.

Filtering to exclude genes with nonzero counts on fewer than 50 cells retained 13,552 genes. Because data were not input through Seurat, cells were not screened for high mitochondrial DNA levels. 1,081 genes were found to be highly variable in the full data set and in all 40 samples.

The rank of the Pearson residuals matrix was estimated as 51 by optht. After mapping to a 51-dimensional SVD representation and excluding kNN outliers, 24,101 cells were retained. These cells belong to 37 of the 39 reported clusters.

Our objective was to evaluate clusterings of sizes 2–70 – the largest being greater than the number of reported cell clusters (39). The stopping conditions limited the largest clustering for at least one sample to a smaller size: 61 for iteration 1, 62 for iteration 2, 58 for iteration 3.

The largest stable clusterings found in the first and second iterations have 5 clusters. The largest stable clustering found in the third iteration has 11. The third iteration retained 22,416 cells, which belong to 36 of the 39 published cell clusters.

Figure 13 displays the distributions of normalized MED for the clusterings of sizes 2–58 found with the third iteration.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure13.gif

Figure 13. 25k retinal data (iteration 3): distributions of normalized MED for clusterings of sizes 2–58.

Figure 14 displays the distributions of normalized CMER for the 11 clusters. Cluster 9 is unstable. Because CMER = 1 with all samples, the line is marked red. Cluster 6 is also unstable. The remaining clusters are stable. The clustering is admissible for downstream analysis.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure14.gif

Figure 14. 25k retinal data (iteration 3): distributions of normalized CMER for the 11-clustering.

Table 14 illustrates the compatibility of the 11 hierarchical clusters with the reported cell clusters. The adjusted Rand index equals 0.49.

  • - Most members of cell cluster 24 (rods) belong to hierarchical clusters 0 and 1. We anticipate using differential expression analysis to evaluate this split.

  • - Similarly, cell cluster 26 is split between hierarchical clusters 2 and 3.

  • - Cell cluster 25 (cones) agrees closely with hierarchical cluster 8.

  • - Cell cluster 3 agrees almost exactly with the unstable hierarchical cluster 9.

  • - 80% of the cells in cluster 27 belong to the unstable hierarchical cluster 6.

  • - Cell cluster 34 agrees closely with the very stable hierarchical cluster 10.

Table 14. 25k retinal data (iteration 3): compare 36 of 39 published clusters with the 11-clustering.Cell clusterHierarchical clusterTotal01234567891036100000000012245,6469,2211413659384616015,13138001000000001262161467601010001,2951000011000002161001152013100159170010220001000222181000410000004219100074000000752011172017800200021021781011620300013722412001165121011423900001100000222200077000008140100110000001250100013000001460100199000001017100001600210016481100080000010902000189000001911011001540100058110100111801000121120111031360010015213000002400000241400000100000123291001400210015527172814401259200003432841010000162300002702993623201124930131630110521103150013363127343232550002793241210001790001873312232762404104014892521293010051,053001,11231000020001350138344111000000435442Total 5,756 9,434 702 699 951 1,108 294 1,765 1,071 136 500 22,416

65k lung

The paper [22] describes separately clustering the data for each patient. Our analysis included batch correction. Each patient’s data were treated as an independent batch.

The data set contains counts for 26,485 genes and 65,662 cells. The accompanying metadata file provides 57 cell type annotations.

Excluding data for blood cells retained 60,993 cells. Filtering to exclude genes with nonzero counts on fewer than 50 cells retained 17,470 genes. Because data were not input through Seurat, cells were not screened for high mitochondrial DNA levels. 1,659 genes were found to be highly variable in the full set of data and all 40 samples.

The rank of the Pearson residuals matrix was estimated as 305 by optht. After mapping to a 305-dimensional SVD representation and excluding kNN outliers, 60,114 cells were retained.

Our objective was to evaluate clusterings of sizes 2–70 – the largest being greater than the number of reported cell types (57). The stopping conditions limited the largest clustering for at least one sample to a smaller size: 69 for iteration 1, 48 for iteration 2, 41 for iteration 3. The largest stable clustering found in the first iteration has 19 clusters. The largest stable clusterings found in the second and third iterations have 13 clusters.

Figure 15 displays the distributions of normalized MED for the clusterings of sizes 2–69 found with the first iteration. The green lines highlight data for the clusterings of sizes 16 and 19, which are reviewed.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure15.gif

Figure 15. 65k lung data (iteration 1): distributions of normalized MED for clusterings of sizes 2–69.

Figure 16 displays the distributions of normalized CMER for the 19 clusters. Clusters 1 and 4 (first and third lines) are unstable. The remaining clusters are stable. The clustering is admissible for downstream analysis.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure16.gif

Figure 16. 65k lung data (iteration 1): distributions of normalized CMER for the 19-clustering.

Because blood cells were excluded from our analysis, data for one cell type were eliminated. Table 15 illustrates the compatibility of the 19 hierarchical clusters with the 56 retained cell types. The adjusted Rand index equals 0.65. The majority of macrophages are divided between clusters 2 and 3.

Table 15. 65k lung data (iteration 1): compare 56 of 57 reported cell types with the 19-clustering. Cell typeHierarchical cluster 01234567891011121314151617 18TotalCD4+ Memory/Effector T2,38560000020000210000002,414CD4+ Naive T287100000000000000000288CD8+ Memory/Effector T6930000000000077000000770Mesothelial701041200104000000020Natural Killer T2520000000000083000000335B223900000000000000010242Plasma0930000000000001000094Plasmacytoid Dendritic011401000000000010000116TREM2+ Dendritic0184460000001100160000149Ionocyte2001400200001002010022Macrophage805,1329,4680003113610660111014,701Proliferating Macrophage0038187000000000000010226Fibromyocyte00006200003001000000093Myofibroblast0000207110020031000000224Adventitial Fibroblast000053540000090000000368Alveolar Fibroblast000171,24000100720300001,261Lipofibroblast00001315000000000000028Basal3007003681900130050100407Differentiating Basal0006001844400010070100243Goblet000400114300000000001122Proliferating Basal0001003600000001020040Proximal Basal000400149100000010000155Club1001006816001102003610865Mucous0001001633200200000000351Neuroendocrine00000000300000000003Pericyte100201001,585150001000001,605Airway Smooth Muscle000010006679000010000687Vascular Smooth Muscle2000140045410000000000462Capillary Aerocyte00001000004,2509800001104,351Capillary Intermediate 100000000004751520000010628Artery000001001071,47010000001,480Bronchial Vessel 1010001001014310002000437Bronchial Vessel 2000002000082170010000228Capillary3103000001547,32120000107,386Capillary Intermediate 2000000000174560000000464Vein000000001071,14100001001,150CD8+ Naive T625300000000007760000201,406Natural Killer15000000000004,2320000004,247Proliferating NK/T220000000000078000010101Alveolar Epithelial Type 1000000130093094300210962Classical Monocyte200100000000001,01000001,013EREG+ Dendritic0001100000001001290000141IGSF21+ Dendritic700600000000002630010277Intermediate Monocyte000000000000001870000187Myeloid Dendritic Type 198010000000010940000113Myeloid Dendritic Type 2602700000000202130000230Nonclassical Monocyte000000000000006031000604OLR1+ Classical Monocyte110200000001001960030204Lymphatic000001000011000452000455Alveolar Epithelial Type 210000031750025236103,564203,791Signaling Alveolar Epithelial Type 20000000310020011063400669Basophil/Mast 1200000000000001001,33601,339Basophil/Mast 2000100000000000005470548Ciliated000130002700011000101,2701,313Platelet/Megakaryocyte002200110000010000411Proximal Ciliated0000000000000000008888Total 4,336 468 5,259 9,790 301 1,631 882 1,457 1,646 1,138 4,831 11,335 5,280 984 2,803 455 4,245 1,910 1,363 60,114

The 16-clustering of the same data is reviewed because the results – for both the clustering and its clusters – are the most stable found in the four large data sets. Referring to Figure 15, the maximum value of normalized MED equals 0.01. We cannot explain this why it is exceptionally small.

Figure 17 displays the distributions of normalized CMER for the 16 clusters. All are stable.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure17.gif

Figure 17. 65k lung data (iteration 1): distributions of normalized CMER for the 16-clustering.

Table 16 compares the published cell types with the 16 hierarchical clusters. The adjusted Rand index equals 0.81. The majority of macrophages belong to cluster 10.

Table 16. 65k lung data (iteration 1): compare 56 of 57 reported cell types with the 16-clustering. Cell typeHierarchical cluster01234 5 67891011121314 15TotalBasal3681900133005701000407Differentiating Basal1844400010007601000243Goblet114300000000400001122Proliferating Basal3600000000110200040Proximal Basal149100000001400000155Club6816001110201036010865Mucous1633200200000100000351Neuroendocrine00300000000000003Pericyte001,585150010102001001,605Airway Smooth Muscle006679000001000100687Vascular Smooth Muscle0045410002000000500462Capillary Aerocyte00004,2509800000011104,351Capillary Intermediate 100004751520000000010628Artery001071,47001000001001,480Bronchial Vessel 1001014311000020100437Bronchial Vessel 2000082170001000200228Capillary0001547,32142003000107,386Capillary Intermediate 2000174560000000000464Vein001071,14100000010001,150B000000241000000010242CD4+ Memory/Effector T0200002,39121000000002,414CD4+ Naive T000000288000000000288CD8+ Memory/Effector T0000006937700000000770Mesothelial200104700010050020Natural Killer T0000002528300000000335Plasma0000009300100000094Plasmacytoid Dendritic000000114001100000116CD8+ Naive T000000628776000000201,406Natural Killer000000154,232000000004,247Proliferating NK/T000000227800000010101Alveolar Epithelial Type 1130093009430002010962Classical Monocyte0000002001,0101000001,013EREG+ Dendritic0000010001291100000141IGSF21+ Dendritic000000700263600010277Intermediate Monocyte000000000187000000187Myeloid Dendritic Type 1000000171094100000113Myeloid Dendritic Type 2000000620213900000230Nonclassical Monocyte000000000603010000604OLR1+ Classical Monocyte000001200196200030204Ionocyte2000012002140100022Macrophage0311368106614,60001011014,701Platelet/Megakaryocyte110000001040000411Proliferating Macrophage000000000022500010226TREM2+ Dendritic0000111001613000000149Lymphatic000011000004520100455Alveolar Epithelial Type 23175002512361003,5640203,791Signaling Alveolar Epithelial Type 20310020001100634000669Adventitial Fibroblast000009000000035900368Alveolar Fibroblast00100702031001,247001,261Fibromyocyte00030010000000620093Lipofibroblast0000000000000280028Myofibroblast002003010000021800224Basophil/Mast 1000000200100001,33601,339Basophil/Mast 2000000000010005470548Ciliated027000101001301001,2701,313Proximal Ciliated0000000000000008888Total 882 1,457 1,646 1,138 4,831 11,335 4,804 5,280 984 2,803 15,049 455 4,245 1,932 1,910 1,363 60,114

100k breast cancer

Discussion in [24] and programs posted at [40] suggest that batch correction was performed for each patient’s data for some, if not all, analyses. In our analysis, each patient’s data were treated as an independent batch.

The data set contains counts for 29,733 genes and 100,064 cells. Nine major cell types were reported, as well as 29 minor cell types and 49 cell type subsets.

Filtering to exclude genes with nonzero counts on fewer than 50 cells retained 21,354 genes. The data include cells for which up to 20% of counts are due to mitochondrial DNA. (The percentage exceeds 5% for half of the cells.) 1,704 genes were found to be highly variable in the set of counts for all cells and in all 40 samples.

The rank of the Pearson residuals matrix was estimated as 434 by optht. After mapping to a 434-dimensional SVD representation and excluding kNN outliers, 98,681 cells were retained.

Our objective was to evaluate clusterings of sizes 2–70 – the largest being greater than the number of reported cell type subsets (49). The stopping conditions limited the largest clustering for at least one sample to a smaller size: 58 for iteration 1, 39 for iteration 2, 51 for iteration 3. No stable clusterings were found. The smallest values found for the 90th percentile of MED are 0.15 (first iteration), 0.16 (second), and 0.107 (third).

Figure 18 displays the distributions of normalized MED for the clusterings of sizes 2–51 found with the third iteration. We review the 9-clustering, which has the smallest 90th percentile (0.107). This near miss from being judged stable illustrates a disadvantage of hard thresholds. We describe the clustering as marginally unstable.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure18.gif

Figure 18. 100k breast cancer data (iteration 3): distributions of normalized MED for clusterings of sizes 2–51.

Figure 19 displays the distributions of normalized CMER for the 9 clusters. Cluster 6 (top line) is unstable. CMER equals 1 with 39 of the 40 samples. Cluster 0 (fifth line) is also unstable. The 90th percentile of CMER equals 0.55, larger than the threshold of 0.50 that defines stable clusters. This is a second reminder of the disadvantage of hard thresholding. The remaining clusters are stable. The clustering is not admissible for downstream analysis.

444620e6-cd4a-47c2-bbe7-829bb4ff3d44_figure19.gif

Figure 19. 100k breast cancer data (iteration 3): distributions of normalized CMER for the 9-clustering.

.

The third iteration retained 92,182 cells, which include all 49 published cell types. Filtering had a disproportionate impact on plasmablasts, one of the 9 major cell types. The data downloaded from the GEO included counts for 100,064 cells ( Table 1) of which 3,524 are plasmablasts. Excluding kNN outliers before clustering retained 98,681 cells ( Table 4) of which 2,290 are plasmablasts. They account for 1,234 of the 1,383 cells excluded as outliers in the first iteration. Only 801 plasmablasts were retained in the third iteration.

Table 17 compares the 9 hierarchical clusters with the major cell types. The adjusted Rand index equals 0.86.

  • - The hierarchical clusters do not separate normal from cancer epithelial cells, though approximately a quarter of the normal cells are assigned to cluster 2, accounting for more than 95% of its membership.

  • - The unstable cluster 6 represents a segment of myeloid cells.

  • - 90% of the plasmablasts that survived the filtering iterations are split between clusters 3 and 5.

Table 17. 100k breast cancer data (iteration 3): compare major cell types with the 9-clustering.Major cell typeHierarchical cluster012345678TotalCAFs5,3641572734108135,622PVL394,9154462600245,036Cancer Epithelial165611921,868182301,40214423,700Normal Epithelial139092,91012022103,858Plasmablasts69343143300270801Endothelial12220597,10920347,211B-cells 3148254982,425325452,875T-cells 4722051285033,8311334,159Myeloid281811144013266538,3878,920Total 5,622 5,206 948 25,960 7,351 2,851 271 35,353 8,620 92,182

Table 18 illustrates the compatibility of the 9 hierarchical clusters with the 49 cell type subsets identified in the paper. The adjusted Rand index equals 0.20.

  • - Cluster 2, mostly normal epithelial cells, corresponds to myoepithelial cells.

  • - Cluster 6 (unstable) corresponds to myeloid c4 DCs pDC IRF7 cells.

Table 18. 100k breast cancer data (iteration 3): compare 49 reported cell type subsets with the 9-clustering.Cell type subsetHierarchical cluster012345678TotalCAFs MSC iCAF-like s11,89322119000101,936CAFs MSC iCAF-like s259410213431046745CAFs Transitioning s347420600001483CAFs myCAF like s45002701200035547CAFs myCAF like s51,903402100011,911Cycling PVL537002000044PVL Differentiated s313,235326200233,272PVL Immature s1279971161400011,056PVL Immature s266460480000664Myoepithelial128992810001932Cancer Basal SC1083083,609270258494,071Cancer Cycling281174,826660249485,181Cancer Her2 SC3823,46424047293,559Cancer LumA SC19517,405740143157,599Cancer LumB SC7712,56412070533,290Luminal Progenitors00101,807010881,834Mature Luminal0101,0750101411,092Plasmablasts69343143300270801Endothelial ACKR1820224,39510144,433Endothelial CXCL1211071,53710001,547Endothelial Lymphatic LYVE1100241380020165Endothelial RGS5219061,03900001,066B cells Memory3145231781,930325452,334B cells Naive0032320495000541Myeloid c4 DCs pDC IRF72300501266310308T cells c0 CD4+ CCR7000201204,85804,872T cells c1 CD4+ IL7R12011331807,44937,589T cells c10 NKT cells FCGR3A00031101,03801,043T cells c11 MKI672205721801,34631,430T cells c2 CD4+ T-regs FOXP3020311004,08114,098T cells c3 CD4+ Tfh CXCL1300110302,16102,166T cells c4 CD8+ ZFP3600172504,96714,983T cells c5 CD8+ GZMK00000002730273T cells c6 IFIT110072309702985T cells c7 CD8+ IFNG00050803,06713,081T cells c8 CD8+ LAG300041201,87321,882T cells c9 NK cells AREG01030501,74801,757Cycling Myeloid0001838015384428Myeloid c0 DC LAMP3012029420359109Myeloid c1 LAM1 FABP501091000131,8591,892Myeloid c10 Macrophage 1 EGR10111110041,9761,994Myeloid c11 cDC2 CD1C10040103310319Myeloid c12 Monocyte 1 IL1B0001430071,2541,278Myeloid c2 LAM2 APOE24019180011,0101,054Myeloid c3 cDC1 CLEC9A00001002150153Myeloid c5 Macrophage 3 SIGLEC1100000002425Myeloid c7 Monocyte 3 FCGR3A000000005353Myeloid c8 Monocyte 2 S100A900050001859865Myeloid c9 Macrophage 2 CXCL1010000101439442Total 5,622 5,206 948 25,960 7,351 2,851 271 35,353 8,620 92,182

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1scCoExpress: A Sparsity-Aware R Package for Gene Co-Expression Analysis in Single-Cell Datasets [version 1; peer review: awaiting peer review]012.1627-07-2026
2On the limits of inferring biophysical parameters of RBP-RNA interactions from in vitro RNA Bind’n Seq data [version 3; peer review: 3 approved, 1 not approved]010.7417-07-2026
3Inferring and simulating a gene regulatory network for the sympathoadrenal differentiation from single-cell transcriptomics in human. [version 2; peer review: 1 approved]08.9923-07-2026
4Some New Generalized Types of J-spaces and Metacompact spaces [version 1; peer review: 1 approved with reservations]05.5323-07-2026
5DAVE: how to use explainable AI to interpret missense variants for genome diagnostics based on functional protein modeling [version 1; peer review: awaiting peer review]07.3710-08-2026
6Advertisement Characterizations by Ambiguous Graphs [version 1; peer review: awaiting peer review]05.0203-08-2026
7Exosome-Mediated Communication in The Tumor Microenvironment: Mechanism and Therapeutic Challenges [version 1; peer review: awaiting peer review]014.406-08-2026
8The Concentration-Fragility Nexus: Early-Warning Systems and Portfolio Implications in Concentrated Markets [version 2; peer review: 1 approved with reservations, 2 not approved]016.0207-08-2026
9How Microsoft is governing thousands of Kubernetes clusters without manual intervention013.3507-05-2026
10R 4- STANDBY- STRENGTH-STRESS MODEL [version 2; peer review: 1 approved with reservations, 3 not approved]010.2210-08-2026

Классификация: . Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 8.06. Источник: f1000research.com.