20260731_repoolAnalysis of the barcode sequencing of the miscellaneous plate
20260731_repool from 2026-07-31.
The plots are interactive: mouse over points for details.
The balanced re-pool made by adding each strain at the volume calculated from the initial equal-volume pool, then used to infect MDCK-SIAT1 cells over a serial dilution. The barcodes were sequenced to check whether the re-pool actually came out balanced, to find which strains are still over-represented, and to calculate the volumes for a second corrective re-pool.
The "fates" of the reads parsed for each well. If most reads are not "valid barcode", that could indicate a problem with the sequencing or with the barcodes being parsed.
Wells need enough counts per barcode for the composition of the pool to be estimated accurately, and enough counts from the neutralization standard for the amount of pool in the well to be measured. Wells are flagged when they have an average of fewer than 500 counts per barcode or when fewer than 0.001 of their counts are from the neutralization standard.
Diluting the pool should divide its viral counts by the dilution factor while leaving the neutralization standard alone, so the neutralization standard's share of the counts, in odds form, rises along a straight line against the dilution factor on log-log axes for as long as that holds. It is over the dilutions where that is true that the composition of the pool can be read.
Runs of 3 consecutive dilutions are read from concentrated to dilute, and the first whose fitted slope is within 0.3 of ideal marks where the pool starts responding. Wells failing either QC threshold above are drawn below but left out of the fits, as is any well with no viral counts, whose odds are undefined.
The slope fit to each run of dilutions:
| dilutions | slope | standard error | off ideal by | within tolerance | chosen |
|---|---|---|---|---|---|
| 4-16 | -0.5661 | 0.2967 | 0.4339 | False | False |
| 8-32 | -0.476 | 0.2854 | 0.524 | False | False |
| 16-64 | -0.2605 | 0.2219 | 0.7395 | False | False |
| 32-128 | -0.9033 | 0.2084 | 0.09668 | True | True |
| 64-256 | -2.01 | 0.3549 | 1.01 | False | False |
| 128-512 | -1.129 | 0.5066 | 0.1293 | True | False |
| 256-1024 | -0.7629 | 0.4315 | 0.2371 | True | False |
| 512-2048 | -1.274 | 0.2418 | 0.2739 | True | False |
| 1024-4096 | -0.9655 | 0.3176 | 0.03453 | True | False |
| 2048-8192 | -0.7966 | 0.3394 | 0.2034 | True | False |
The pool starts responding to dilution at 32-fold: over
32 to 128 the slope is
-0.903 +/- 0.208, off the ideal -1.0 by
0.097. The composition would be calculated from the middle
dilution of that run, 64-fold (wells
A5, B5, marked by the dashed rule).
The composition below is read from those wells.
The left panel shows each strain's share of the viral counts (the neutralization
standard is excluded) in wells
A5, B5. A strain's counts are summed
over its barcodes, so the shares sum to one across strains and the dashed line at
0.68% marks where every strain would be equally represented. The black
points are the individual wells, so replicate disagreement stays visible rather than
being averaged away. The right panel shows how a strain's counts are split among its
barcodes, which should be roughly even unless a barcode is poorly represented in the
rescued virus.
A strain is called over-represented when its share exceeds that equal share by more
than the over_representation_factor of 1.5 set in config.yml (above
1.01%), and under-represented when its share falls
below that share divided by the under_representation_factor of 1.5
(below 0.45%). The two thresholds are reciprocal, so
being this many fold too abundant and this many fold too scarce count as equally far
from equal representation.
10 of the 148 strains are over-represented and 36 are
under-represented. The most abundant, A/Bangkok/P2391/2025_H3N2, is at
7.92% of the pool, which is
11.7x its equal share; the least
abundant, A/France/NAQ-HCL026046169102/2025_H1N1, is at 0.26%, or
0.38x.
The 10 most over-represented strains:
| shortname | strain | subpool | mean_fraction_strain | x_expected |
|---|---|---|---|---|
| flu-seqneut-25to26_H3N2_25 | A/Bangkok/P2391/2025_H3N2 | old_h3_vax | 0.07915 | 11.71 |
| flu-seqneut-25to26_H3N2_19 | A/Galicia/GA-CHUAC-451/2025_H3N2 | old_h3_vax | 0.02491 | 3.687 |
| flu-seqneut-2026_H3N2_27 | A/StPetersburg/RII-25-2506S/2026_H3N2 | flu-seqneut-2026_h3 | 0.02205 | 3.264 |
| flu-seqneut-2025_H3N2_59 | A/Lisboa/216/2023_H3N2 | old_h3_vax | 0.02205 | 3.264 |
| flu_seqneut_pdmH1N1_2023-2024_VS4 | A/Wisconsin/588/2019_H1N1 | old_h1_vax | 0.01169 | 1.729 |
| flu-seqneut-2026_H3N2_48 | A/Netherlands/888/2026_H3N2 | flu-seqneut-2026_h3 | 0.01112 | 1.645 |
| flu_seqneut_pdmH1N1_2023-2024_VS5 | A/Hawaii/70/2019_H1N1 | old_h1_vax | 0.01075 | 1.591 |
| flu-seqneut-2026_H3N2_35 | A/Michigan/UM-10068134090/2025_H3N2 | flu-seqneut-2026_h3 | 0.01064 | 1.574 |
| flu-seqneut-2025_H3N2_77 | A/Croatia/10136RV/2023_H3N2 | old_h3_vax | 0.01038 | 1.536 |
| flu-seqneut-DRIVE-2018to2026_H3N2_29 | A/Vietnam/F7324/2024_H3N2 | old_h3_vax | 0.0102 | 1.51 |
The 10 most under-represented strains:
| shortname | strain | subpool | mean_fraction_strain | x_expected |
|---|---|---|---|---|
| flu-seqneut-2026_H1N1_6 | A/France/NAQ-HCL026046169102/2025_H1N1 | flu-seqneut-2026_h1 | 0.002561 | 0.379 |
| flu-seqneut-2026_H3N2_42 | A/Michigan/UM-10068747355/2026_H3N2 | flu-seqneut-2026_h3 | 0.002816 | 0.4167 |
| flu-seqneut-2026_H3N2_67 | A/Navarra/260009/2025_H3N2 | flu-seqneut-2026_h3 | 0.003158 | 0.4674 |
| flu-seqneut-2026_H1N1_34 | A/Andalucia/PMC-01217/2025_H1N1 | flu-seqneut-2026_h1 | 0.003161 | 0.4679 |
| flu-seqneut-2026_H3N2_1 | A/Darwin/1454/2025_H3N2 | flu-seqneut-2026_h3 | 0.00321 | 0.4751 |
| flu-seqneut-2026_H1N1_37 | A/Galicia/GA-CHUAC-612/2025_H1N1 | flu-seqneut-2026_h1 | 0.003244 | 0.4801 |
| flu-seqneut-2026_H3N2_41 | A/Victoria/2948/2025_H3N2 | flu-seqneut-2026_h3 | 0.003246 | 0.4804 |
| flu-seqneut-2026_H3N2_7 | A/StPetersburg/RII-25-2600S/2026_H3N2 | flu-seqneut-2026_h3 | 0.003274 | 0.4845 |
| flu-seqneut-2026_H3N2_38 | A/SouthCarolina/USAFSAM-17046/2026_H3N2 | flu-seqneut-2026_h3 | 0.003348 | 0.4955 |
| flu-seqneut-2025_H1N1_37 | A/Hawaii/ISC-1140/2025_H1N1 | old_h1_vax | 0.00336 | 0.4973 |
As rates rather than counts: the subpools differ several-fold in size, so the number of flagged strains alone would make the large subpools look worse than they are.
| subpool | n_strains | n_over_represented | n_under_represented | mean_fraction_strain | percent_over_represented | percent_under_represented |
|---|---|---|---|---|---|---|
| old_h3_vax | 16 | 5 | 0 | 0.01355 | 31.25 | 0 |
| old_h1_vax | 13 | 2 | 2 | 0.007125 | 15.38 | 15.38 |
| flu-seqneut-2026_h3 | 66 | 3 | 20 | 0.005875 | 4.545 | 30.3 |
| flu-seqneut-2026_h1 | 53 | 0 | 14 | 0.005714 | 0 | 26.42 |
Each strain's representation here against the volume of it that went into the
pool, taken from the analyze_pool report for 20260715_equal_vol_pool. A strain that
needed little volume was growing well, so if the imbalance tracks stock titer the
over-represented strains gather at the low-volume end.
The strains are held in subpools, which are combined to make the final pool, so the imbalance has two separate causes and two separate remedies:
remake is set for it in config.yml.The target is an equal share for every strain: each subpool has to supply virus in proportion to the number of strains it holds, so that all 148 strains aim for the same 0.68% of the final pool.
The arithmetic rests on the same premise as before, stated rather than inferred
because nothing in the data reveals it: a strain's share of the reads is proportional
to the volume of it that went in times the titer of its stock, so the stock titer is
proportional to fraction_strain / previous_volume. Within a remade subpool the volume
that equalizes its strains is therefore proportional to
previous_volume / fraction_strain.
How much virus a subpool has to supply and how much of it to pipette are different numbers, related by its titer, the virus it holds per uL. A concentrated subpool supplies its share of the virus from less of its volume than a dilute one does. Each subpool's titer is therefore worked out below and the volumes divided by it, which is what makes the two remedies above independent of each other.
Note this predicts the composition only. It says nothing about the titer of the corrective pool, so unlike the initial pool no dilution is recommended here: the corrective pool has to be titrated.
Subpools being remade from strain stocks: old_h3_vax. Every subpool is then
combined in the proportion below.
Across the 148 strains that are kept (see below for the 0 that are dropped), the most and least abundant strain currently differ 30.9-fold. After the corrective re-pool they are predicted to differ 8.38-fold.
The volumes to pipette are in
20260731_repool_subpool_repooling_math/: one CSV per subpool being
remade, most volume first, where volume_to_add_uL is what to pipette and
dilution_factor how far the stock is diluted first, so a strain with a factor of 1
is added neat; plus combine_subpools.csv, holding the fractions the finished
subpools are combined in. The measurement those came from is in
20260731_repool_repooling_math.csv, which also carries
predicted_fraction_strain, what each strain is expected to come out at. The strains
dropped from the re-pooling are in
20260731_repool_dropped_strains.csv.
What this assumes. The predictions hold only if each remade subpool is mixed from the individual strain stocks, those stocks have not changed titer since the previous pool was made, and the volumes recorded for that pool are what was actually pipetted. None of those can be checked from the sequencing data, so they are premises rather than findings. If any does not hold the numbers will be wrong in a way that looks perfectly reasonable, and the corrective pool should be sequenced again to find out. Nothing here predicts the titer of the corrective pool either, so it has to be titrated rather than diluted by calculation.
fraction_of_pool is the number to work from: multiply it by whatever volume of
pool is being made. The fractions sum to one, so a 100 uL test pool is
58.52 uL of flu-seqneut-2026_h1, 17.57 uL of flu-seqneut-2026_h3, 17.21 uL of old_h1_vax, 6.70 uL of old_h3_vax. volume_uL is the same thing for the full
10307 uL.
fraction_of_pool is a share of the pool's volume, and it is not the same as
target_virus_share, the share of the pool's virus that subpool has to supply.
The two differ by relative_titer, how much virus the subpool holds per uL against
the average of them, so that a subpool twice as concentrated as the average supplies
its share of the virus from half as much liquid:
fraction_of_pool = (target_virus_share / relative_titer), renormalized to sum to one
For a subpool that is not being remade the titer is a property of the liquid already in the tube, worked out from the share of the reads it accounts for over the volume it was mixed at. For one that is being remade it comes from its own build below, and so already accounts for the extra carrier liquid of any strain diluted to make it pipettable.
| subpool | n_strains | remake | current_fraction_of_pool | target_virus_share | relative_titer | fraction_of_pool | volume_uL |
|---|---|---|---|---|---|---|---|
| flu-seqneut-2026_h1 | 53 | False | 0.3028 | 0.3581 | 0.4641 | 0.5852 | 6032 |
| flu-seqneut-2026_h3 | 66 | False | 0.3878 | 0.4459 | 1.925 | 0.1757 | 1811 |
| old_h1_vax | 13 | False | 0.09263 | 0.08784 | 0.3871 | 0.1721 | 1774 |
| old_h3_vax | 16 | True | 0.2168 | 0.1081 | 1.224 | 0.06699 | 690.5 |
old_h3_vaxMade from 16 strains, giving 787 uL of subpool. The final pool takes 690 uL of that, so it is built 1.14x over what is needed to leave some to spare for what is lost in pipetting it.
3 of the 16 strains need less than 10 uL of stock, so they are diluted first and that much more of the dilution added. The virus delivered is the same either way; the subpool just ends up holding more liquid, which its titer above already accounts for.
| shortname | strain | neat_volume_uL | dilution | volume_to_add_uL |
|---|---|---|---|---|
| flu-seqneut-25to26_H3N2_19 | A/Galicia/GA-CHUAC-451/2025_H3N2 | 2.071 | 1:5 | 10.35 |
| flu-seqneut-25to26_H3N2_25 | A/Bangkok/P2391/2025_H3N2 | 0.6509 | 1:20 | 13.02 |
| flu-seqneut-2025_H3N2_59 | A/Lisboa/216/2023_H3N2 | 6.342 | 1:2 | 12.68 |
The 5 strains needing the most volume:
| shortname | strain | mean_fraction_strain | volume_to_add_uL |
|---|---|---|---|
| flu-seqneut-H3-2023to2024_H3N2_79 | A/Thailand/8/2022_H3N2 | 0.008754 | 190.7 |
| flu-seqneut-25to26_H3N2_46 | A/South_Australia/2523812314/2025_H3N2 | 0.004583 | 105.9 |
| flu-seqneut-2025_H3N2_77 | A/Croatia/10136RV/2023_H3N2 | 0.01038 | 98.01 |
| flu-seqneut-H3-2023to2024_H3N2_32 | A/Massachusetts/18/2022_H3N2 | 0.008098 | 65.4 |
| flu-seqneut-25to26_H3N2_21 | A/Sydney/1359/2024_H3N2 | 0.00493 | 41.97 |
No strains are dropped from the corrective re-pooling.