Balance of library re-pool 20260731_repool

Analysis of the barcode sequencing of the miscellaneous plate 20260731_repool from 2026-07-31. The plots are interactive: mouse over points for details.

Experimental description

The balanced re-pool made by adding each strain at the volume calculated from the initial equal-volume pool, then used to infect MDCK-SIAT1 cells over a serial dilution. The barcodes were sequenced to check whether the re-pool actually came out balanced, to find which strains are still over-represented, and to calculate the volumes for a second corrective re-pool.

Fates of the sequencing reads

The "fates" of the reads parsed for each well. If most reads are not "valid barcode", that could indicate a problem with the sequencing or with the barcodes being parsed.

Barcode counts per well

Wells need enough counts per barcode for the composition of the pool to be estimated accurately, and enough counts from the neutralization standard for the amount of pool in the well to be measured. Wells are flagged when they have an average of fewer than 500 counts per barcode or when fewer than 0.001 of their counts are from the neutralization standard.

The linear range of the dilution series

Diluting the pool should divide its viral counts by the dilution factor while leaving the neutralization standard alone, so the neutralization standard's share of the counts, in odds form, rises along a straight line against the dilution factor on log-log axes for as long as that holds. It is over the dilutions where that is true that the composition of the pool can be read.

Runs of 3 consecutive dilutions are read from concentrated to dilute, and the first whose fitted slope is within 0.3 of ideal marks where the pool starts responding. Wells failing either QC threshold above are drawn below but left out of the fits, as is any well with no viral counts, whose odds are undefined.

The slope fit to each run of dilutions:

dilutions slope standard error off ideal by within tolerance chosen
4-16 -0.5661 0.2967 0.4339 False False
8-32 -0.476 0.2854 0.524 False False
16-64 -0.2605 0.2219 0.7395 False False
32-128 -0.9033 0.2084 0.09668 True True
64-256 -2.01 0.3549 1.01 False False
128-512 -1.129 0.5066 0.1293 True False
256-1024 -0.7629 0.4315 0.2371 True False
512-2048 -1.274 0.2418 0.2739 True False
1024-4096 -0.9655 0.3176 0.03453 True False
2048-8192 -0.7966 0.3394 0.2034 True False

The pool starts responding to dilution at 32-fold: over 32 to 128 the slope is -0.903 +/- 0.208, off the ideal -1.0 by 0.097. The composition would be calculated from the middle dilution of that run, 64-fold (wells A5, B5, marked by the dashed rule).

The composition below is read from those wells.

Representation of each strain in the re-pool

The left panel shows each strain's share of the viral counts (the neutralization standard is excluded) in wells A5, B5. A strain's counts are summed over its barcodes, so the shares sum to one across strains and the dashed line at 0.68% marks where every strain would be equally represented. The black points are the individual wells, so replicate disagreement stays visible rather than being averaged away. The right panel shows how a strain's counts are split among its barcodes, which should be roughly even unless a barcode is poorly represented in the rescued virus.

A strain is called over-represented when its share exceeds that equal share by more than the over_representation_factor of 1.5 set in config.yml (above 1.01%), and under-represented when its share falls below that share divided by the under_representation_factor of 1.5 (below 0.45%). The two thresholds are reciprocal, so being this many fold too abundant and this many fold too scarce count as equally far from equal representation.

10 of the 148 strains are over-represented and 36 are under-represented. The most abundant, A/Bangkok/P2391/2025_H3N2, is at 7.92% of the pool, which is 11.7x its equal share; the least abundant, A/France/NAQ-HCL026046169102/2025_H1N1, is at 0.26%, or 0.38x.

The 10 most over-represented strains:

shortname strain subpool mean_fraction_strain x_expected
flu-seqneut-25to26_H3N2_25 A/Bangkok/P2391/2025_H3N2 old_h3_vax 0.07915 11.71
flu-seqneut-25to26_H3N2_19 A/Galicia/GA-CHUAC-451/2025_H3N2 old_h3_vax 0.02491 3.687
flu-seqneut-2026_H3N2_27 A/StPetersburg/RII-25-2506S/2026_H3N2 flu-seqneut-2026_h3 0.02205 3.264
flu-seqneut-2025_H3N2_59 A/Lisboa/216/2023_H3N2 old_h3_vax 0.02205 3.264
flu_seqneut_pdmH1N1_2023-2024_VS4 A/Wisconsin/588/2019_H1N1 old_h1_vax 0.01169 1.729
flu-seqneut-2026_H3N2_48 A/Netherlands/888/2026_H3N2 flu-seqneut-2026_h3 0.01112 1.645
flu_seqneut_pdmH1N1_2023-2024_VS5 A/Hawaii/70/2019_H1N1 old_h1_vax 0.01075 1.591
flu-seqneut-2026_H3N2_35 A/Michigan/UM-10068134090/2025_H3N2 flu-seqneut-2026_h3 0.01064 1.574
flu-seqneut-2025_H3N2_77 A/Croatia/10136RV/2023_H3N2 old_h3_vax 0.01038 1.536
flu-seqneut-DRIVE-2018to2026_H3N2_29 A/Vietnam/F7324/2024_H3N2 old_h3_vax 0.0102 1.51

The 10 most under-represented strains:

shortname strain subpool mean_fraction_strain x_expected
flu-seqneut-2026_H1N1_6 A/France/NAQ-HCL026046169102/2025_H1N1 flu-seqneut-2026_h1 0.002561 0.379
flu-seqneut-2026_H3N2_42 A/Michigan/UM-10068747355/2026_H3N2 flu-seqneut-2026_h3 0.002816 0.4167
flu-seqneut-2026_H3N2_67 A/Navarra/260009/2025_H3N2 flu-seqneut-2026_h3 0.003158 0.4674
flu-seqneut-2026_H1N1_34 A/Andalucia/PMC-01217/2025_H1N1 flu-seqneut-2026_h1 0.003161 0.4679
flu-seqneut-2026_H3N2_1 A/Darwin/1454/2025_H3N2 flu-seqneut-2026_h3 0.00321 0.4751
flu-seqneut-2026_H1N1_37 A/Galicia/GA-CHUAC-612/2025_H1N1 flu-seqneut-2026_h1 0.003244 0.4801
flu-seqneut-2026_H3N2_41 A/Victoria/2948/2025_H3N2 flu-seqneut-2026_h3 0.003246 0.4804
flu-seqneut-2026_H3N2_7 A/StPetersburg/RII-25-2600S/2026_H3N2 flu-seqneut-2026_h3 0.003274 0.4845
flu-seqneut-2026_H3N2_38 A/SouthCarolina/USAFSAM-17046/2026_H3N2 flu-seqneut-2026_h3 0.003348 0.4955
flu-seqneut-2025_H1N1_37 A/Hawaii/ISC-1140/2025_H1N1 old_h1_vax 0.00336 0.4973

Representation by subpool

As rates rather than counts: the subpools differ several-fold in size, so the number of flagged strains alone would make the large subpools look worse than they are.

subpool n_strains n_over_represented n_under_represented mean_fraction_strain percent_over_represented percent_under_represented
old_h3_vax 16 5 0 0.01355 31.25 0
old_h1_vax 13 2 2 0.007125 15.38 15.38
flu-seqneut-2026_h3 66 3 20 0.005875 4.545 30.3
flu-seqneut-2026_h1 53 0 14 0.005714 0 26.42

Against the volumes that made this pool

Each strain's representation here against the volume of it that went into the pool, taken from the analyze_pool report for 20260715_equal_vol_pool. A strain that needed little volume was growing well, so if the imbalance tracks stock titer the over-represented strains gather at the low-volume end.

Corrective re-pooling

The strains are held in subpools, which are combined to make the final pool, so the imbalance has two separate causes and two separate remedies:

  1. Within a subpool, strains can be out of balance with each other. Fixing that means making the subpool again from the individual strain stocks, which is pipetting work proportional to the number of strains in it. A subpool is remade only where remake is set for it in config.yml.
  2. Between subpools, a subpool can contribute too much or too little to the final pool. Fixing that is just a matter of how much of each subpool is added when they are combined, so it is done for every subpool whether or not it is being remade.

The target is an equal share for every strain: each subpool has to supply virus in proportion to the number of strains it holds, so that all 148 strains aim for the same 0.68% of the final pool.

The arithmetic rests on the same premise as before, stated rather than inferred because nothing in the data reveals it: a strain's share of the reads is proportional to the volume of it that went in times the titer of its stock, so the stock titer is proportional to fraction_strain / previous_volume. Within a remade subpool the volume that equalizes its strains is therefore proportional to previous_volume / fraction_strain.

How much virus a subpool has to supply and how much of it to pipette are different numbers, related by its titer, the virus it holds per uL. A concentrated subpool supplies its share of the virus from less of its volume than a dilute one does. Each subpool's titer is therefore worked out below and the volumes divided by it, which is what makes the two remedies above independent of each other.

Note this predicts the composition only. It says nothing about the titer of the corrective pool, so unlike the initial pool no dilution is recommended here: the corrective pool has to be titrated.

What this gives

Subpools being remade from strain stocks: old_h3_vax. Every subpool is then combined in the proportion below.

Across the 148 strains that are kept (see below for the 0 that are dropped), the most and least abundant strain currently differ 30.9-fold. After the corrective re-pool they are predicted to differ 8.38-fold.

The volumes to pipette are in 20260731_repool_subpool_repooling_math/: one CSV per subpool being remade, most volume first, where volume_to_add_uL is what to pipette and dilution_factor how far the stock is diluted first, so a strain with a factor of 1 is added neat; plus combine_subpools.csv, holding the fractions the finished subpools are combined in. The measurement those came from is in 20260731_repool_repooling_math.csv, which also carries predicted_fraction_strain, what each strain is expected to come out at. The strains dropped from the re-pooling are in 20260731_repool_dropped_strains.csv.

What this assumes. The predictions hold only if each remade subpool is mixed from the individual strain stocks, those stocks have not changed titer since the previous pool was made, and the volumes recorded for that pool are what was actually pipetted. None of those can be checked from the sequencing data, so they are premises rather than findings. If any does not hold the numbers will be wrong in a way that looks perfectly reasonable, and the corrective pool should be sequenced again to find out. Nothing here predicts the titer of the corrective pool either, so it has to be titrated rather than diluted by calculation.

Combining the subpools

fraction_of_pool is the number to work from: multiply it by whatever volume of pool is being made. The fractions sum to one, so a 100 uL test pool is 58.52 uL of flu-seqneut-2026_h1, 17.57 uL of flu-seqneut-2026_h3, 17.21 uL of old_h1_vax, 6.70 uL of old_h3_vax. volume_uL is the same thing for the full 10307 uL.

fraction_of_pool is a share of the pool's volume, and it is not the same as target_virus_share, the share of the pool's virus that subpool has to supply. The two differ by relative_titer, how much virus the subpool holds per uL against the average of them, so that a subpool twice as concentrated as the average supplies its share of the virus from half as much liquid:

fraction_of_pool = (target_virus_share / relative_titer), renormalized to sum to one

For a subpool that is not being remade the titer is a property of the liquid already in the tube, worked out from the share of the reads it accounts for over the volume it was mixed at. For one that is being remade it comes from its own build below, and so already accounts for the extra carrier liquid of any strain diluted to make it pipettable.

subpool n_strains remake current_fraction_of_pool target_virus_share relative_titer fraction_of_pool volume_uL
flu-seqneut-2026_h1 53 False 0.3028 0.3581 0.4641 0.5852 6032
flu-seqneut-2026_h3 66 False 0.3878 0.4459 1.925 0.1757 1811
old_h1_vax 13 False 0.09263 0.08784 0.3871 0.1721 1774
old_h3_vax 16 True 0.2168 0.1081 1.224 0.06699 690.5

Remaking old_h3_vax

Made from 16 strains, giving 787 uL of subpool. The final pool takes 690 uL of that, so it is built 1.14x over what is needed to leave some to spare for what is lost in pipetting it.

3 of the 16 strains need less than 10 uL of stock, so they are diluted first and that much more of the dilution added. The virus delivered is the same either way; the subpool just ends up holding more liquid, which its titer above already accounts for.

shortname strain neat_volume_uL dilution volume_to_add_uL
flu-seqneut-25to26_H3N2_19 A/Galicia/GA-CHUAC-451/2025_H3N2 2.071 1:5 10.35
flu-seqneut-25to26_H3N2_25 A/Bangkok/P2391/2025_H3N2 0.6509 1:20 13.02
flu-seqneut-2025_H3N2_59 A/Lisboa/216/2023_H3N2 6.342 1:2 12.68

The 5 strains needing the most volume:

shortname strain mean_fraction_strain volume_to_add_uL
flu-seqneut-H3-2023to2024_H3N2_79 A/Thailand/8/2022_H3N2 0.008754 190.7
flu-seqneut-25to26_H3N2_46 A/South_Australia/2523812314/2025_H3N2 0.004583 105.9
flu-seqneut-2025_H3N2_77 A/Croatia/10136RV/2023_H3N2 0.01038 98.01
flu-seqneut-H3-2023to2024_H3N2_32 A/Massachusetts/18/2022_H3N2 0.008098 65.4
flu-seqneut-25to26_H3N2_21 A/Sydney/1359/2024_H3N2 0.00493 41.97

No strains are dropped from the corrective re-pooling.