Balance of library re-pool 20260810_repool_wash_spikein_lysis

Analysis of the barcode sequencing of the miscellaneous plate 20260810_repool_wash_spikein_lysis from 2026-08-10. The plots are interactive: mouse over points for details.

Experimental description

The rebalanced pool made from the corrective volumes calculated for 20260731_repool and adjusted further based on the initial plates 5 and 6 data. It was used to infect MDCK-SIAT1 cells over a serial dilution. The barcodes were sequenced to check the strain representation that resulted. Additionally, in this experiment, the cells were washed 1x with PBS prior to lysis and the spikein was added to the lysis buffer.

Fates of the sequencing reads

The "fates" of the reads parsed for each well. If most reads are not "valid barcode", that could indicate a problem with the sequencing or with the barcodes being parsed.

Barcode counts per well

Wells need enough counts per barcode for the composition of the pool to be estimated accurately, and enough counts from the neutralization standard for the amount of pool in the well to be measured. Wells are flagged when they have an average of fewer than 500 counts per barcode or when fewer than 0.001 of their counts are from the neutralization standard.

The linear range of the dilution series

Diluting the pool should divide its viral counts by the dilution factor while leaving the neutralization standard alone, so the neutralization standard's share of the counts, in odds form, rises along a straight line against the dilution factor on log-log axes for as long as that holds. It is over the dilutions where that is true that the composition of the pool can be read.

Runs of 3 consecutive dilutions are read from concentrated to dilute, and the first whose fitted slope is within 0.3 of ideal marks where the pool starts responding. Wells failing either QC threshold above are drawn below but left out of the fits, as is any well with no viral counts, whose odds are undefined.

The slope fit to each run of dilutions:

dilutions slope standard error off ideal by within tolerance chosen
4-16 -0.2023 0.08874 0.7977 False False
8-32 -0.2363 0.1256 0.7637 False False
16-64 -0.5336 0.1318 0.4664 False False
32-128 -0.9222 0.1224 0.07775 True True
64-256 -1.059 0.03233 0.0587 True False
128-512 -1.248 0.09079 0.2483 True False
256-1024 -1.047 0.1348 0.04721 True False
512-2048 -0.9678 0.1251 0.03222 True False
1024-4096 -1.254 0.06999 0.2539 True False
2048-8192 -0.824 0.1282 0.176 True False

The pool starts responding to dilution at 32-fold: over 32 to 128 the slope is -0.922 +/- 0.122, off the ideal -1.0 by 0.078. The composition would be calculated from the middle dilution of that run, 64-fold (wells A5, B5, marked by the dashed rule).

The composition below is read from those wells.

Representation of each strain in the re-pool

The left panel shows each strain's share of the viral counts (the neutralization standard is excluded) in wells A5, B5. A strain's counts are summed over its barcodes, so the shares sum to one across strains and the dashed line at 0.68% marks where every strain would be equally represented. The black points are the individual wells, so replicate disagreement stays visible rather than being averaged away. The right panel shows how a strain's counts are split among its barcodes, which should be roughly even unless a barcode is poorly represented in the rescued virus.

A strain is called over-represented when its share exceeds that equal share by more than the over_representation_factor of 1.5 set in config.yml (above 1.01%), and under-represented when its share falls below that share divided by the under_representation_factor of 1.5 (below 0.45%). The two thresholds are reciprocal, so being this many fold too abundant and this many fold too scarce count as equally far from equal representation.

11 of the 148 strains are over-represented and 40 are under-represented. The most abundant, A/Croatia/10136RV/2023_H3N2, is at 4.35% of the pool, which is 6.4x its equal share; the least abundant, A/Hawaii/ISC-1140/2025_H1N1, is at 0.19%, or 0.28x.

The 10 most over-represented strains:

shortname strain subpool mean_fraction_strain x_expected
flu-seqneut-2025_H3N2_77 A/Croatia/10136RV/2023_H3N2 old_h3_vax 0.04355 6.445
flu-seqneut-H3-2023to2024_H3N2_79 A/Thailand/8/2022_H3N2 old_h3_vax 0.04023 5.953
flu-seqneut-2026_H3N2_27 A/StPetersburg/RII-25-2506S/2026_H3N2 flu-seqneut-2026_h3 0.03308 4.896
flu-seqneut-25to26_H3N2_25 A/Bangkok/P2391/2025_H3N2 old_h3_vax 0.01574 2.329
flu-seqneut-2025_H3N2_59 A/Lisboa/216/2023_H3N2 old_h3_vax 0.01304 1.93
flu-seqneut-2026_H3N2_48 A/Netherlands/888/2026_H3N2 flu-seqneut-2026_h3 0.01241 1.836
flu-seqneut-2026_H3N2_32 A/Trieste/257/2025_H3N2 flu-seqneut-2026_h3 0.01221 1.808
flu-seqneut-2026_H3N2_54 A/France/GES-IPP01077/2026_H3N2 flu-seqneut-2026_h3 0.0117 1.732
flu-seqneut-2026_H3N2_22 A/Singapore/NTF0577/2025_H3N2 flu-seqneut-2026_h3 0.0108 1.599
flu-seqneut-2025_H3N2_9 A/Colombia/1851/2024_H3N2 old_h3_vax 0.01031 1.526

The 10 most under-represented strains:

shortname strain subpool mean_fraction_strain x_expected
flu-seqneut-2025_H1N1_37 A/Hawaii/ISC-1140/2025_H1N1 old_h1_vax 0.001858 0.2751
flu-seqneut-2026_H1N1_37 A/Galicia/GA-CHUAC-612/2025_H1N1 flu-seqneut-2026_h1 0.002081 0.308
flu-seqneut-2025_H1N1_25 A/Colorado/218/2024_H1N1 old_h1_vax 0.002183 0.323
flu-seqneut-2026_H1N1_6 A/France/NAQ-HCL026046169102/2025_H1N1 flu-seqneut-2026_h1 0.00261 0.3863
flu-seqneut-25to26_H1N1_5 A/Norway/9556/2025_H1N1 old_h1_vax 0.002661 0.3938
flu-seqneut-2026_H1N1_34 A/Andalucia/PMC-01217/2025_H1N1 flu-seqneut-2026_h1 0.002786 0.4124
flu-seqneut-2025_H1N1_27 A/Ohio/259/2024_H1N1 old_h1_vax 0.003027 0.4479
flu-seqneut-2026_H1N1_2 A/Nebraska/34/2026_H1N1 flu-seqneut-2026_h1 0.003129 0.4631
flu-seqneut-2026_H1N1_38 A/Denmark/4774/2025_H1N1 flu-seqneut-2026_h1 0.003139 0.4645
flu-seqneut-25to26_H3N2_35 A/South_Australia/2527213276/2025_H3N2 old_h3_vax 0.003261 0.4826

Representation by subpool

As rates rather than counts: the subpools differ several-fold in size, so the number of flagged strains alone would make the large subpools look worse than they are.

subpool n_strains n_over_represented n_under_represented mean_fraction_strain percent_over_represented percent_under_represented
old_h3_vax 16 5 2 0.0119 31.25 12.5
flu-seqneut-2026_h3 66 6 5 0.007621 9.091 7.576
flu-seqneut-2026_h1 53 0 23 0.004823 0 43.4
old_h1_vax 13 0 10 0.003923 0 76.92

What this gives

No corrective re-pool is being calculated, corrective_repool being false for this pool in config.yml, so this report measures how the pool came out and stops there. Every strain's share of the pool, and how many fold that is off an equal share, is in 20260810_repool_wash_spikein_lysis_strain_representation.csv.

Across all 148 strains the most and least abundant differ 23.4-fold. Set corrective_repool to true, along with the pipetting keys it requires, to also calculate the volumes for a further re-pool.