20260715_equal_vol_poolAnalysis of the barcode sequencing of the miscellaneous plate
20260715_equal_vol_pool from 2026-07-15.
The plots are interactive: mouse over points for details.
An initial pool was made by adding equal volumes of all strains, and this pool was then used to infect MDCK-SIAT1 cells over a serial dilution of the pool. The barcodes were sequenced to determine the representation of each strain in the pool, how much of each strain to add to make a balanced pool, and the pool dilution (that is, the MOI) to use in the neutralization assays.
linear_range_wells ['A8', 'B8'] are not the wells ['A7', 'B7'] at the 256-fold dilution that would have been calculated from the linear rangeThese do not stop the analysis, but should be looked into.
The "fates" of the reads parsed for each well. If most reads are not "valid barcode", that could indicate a problem with the sequencing or with the barcodes being parsed.
Wells need enough counts per barcode for the composition of the pool to be estimated accurately, and enough counts from the neutralization standard for the amount of pool in the well to be measured. Wells are flagged when they have an average of fewer than 500 counts per barcode or when fewer than 0.001 of their counts are from the neutralization standard.
Diluting the pool should divide its viral counts by the dilution factor while leaving the neutralization standard alone, so the neutralization standard's share of the counts, in odds form, rises along a straight line of slope 1 against the dilution factor on log-log axes for as long as that holds. It is over the dilutions where that is true that the composition of the pool can be read, since a well too concentrated for its counts to fall with dilution measures the assay rather than the pool.
Runs of 3 consecutive dilutions are read from concentrated to dilute, and the first whose fitted slope is within 0.3 of ideal marks where the pool starts responding. Wells failing either QC threshold above are drawn below but left out of the fits, as is any well with no viral counts, whose odds are undefined.
The slope fit to each run of dilutions:
| dilutions | slope | standard error | off ideal by | within tolerance | chosen |
|---|---|---|---|---|---|
| 4-16 | -0.2857 | 0.1114 | 0.7143 | False | False |
| 8-32 | 0.01451 | 0.1299 | 1.015 | False | False |
| 16-64 | 0.05048 | 0.1357 | 1.05 | False | False |
| 32-128 | -0.1994 | 0.1351 | 0.8006 | False | False |
| 64-256 | -0.5735 | 0.1205 | 0.4265 | False | False |
| 128-512 | -0.8479 | 0.08604 | 0.1521 | True | True |
| 256-1024 | -1.059 | 0.09786 | 0.05916 | True | False |
| 512-2048 | -0.8101 | 0.2779 | 0.1899 | True | False |
| 1024-4096 | -1.047 | 0.3113 | 0.04708 | True | False |
| 2048-8192 | -1.215 | 0.2865 | 0.2153 | True | False |
The pool starts responding to dilution at 128-fold: over 128 to 512 the slope is -0.848 +/- 0.086, off the ideal -1.0 by 0.152. The solid line is the ideal slope placed over that run, drawn faint across the rest of the series.
The composition is calculated from the middle dilution of that run,
256-fold (wells
A7, B7, marked by the dashed rule), rather
than from its most concentrated edge. That leaves a dilution of margin on either
side of the wells used: at the edge of the run the pool has only just started
responding, so a slightly misplaced onset would put the wells outside the linear
range altogether.
That calculation is overridden by linear_range_wells in config.yml, which
sets the wells to A8, B8. The composition
below is read from those wells.
Reads that parsed as a barcode but match neither the viral library nor the
neutralization standard, over the wells the composition is read from. One within
1 nucleotide of a known barcode is sequencing error off
that barcode, while one further away is a genuinely different barcode, and so material
that does not belong in the pool rather than noise. The distance to the closest known
barcode is the one computed by seqneut-pipeline, not recomputed here; that barcode
itself is named only where it is close enough to be what the reads came from.
1.55% of the
3,597,410 parsed reads in wells
A8, B8 had an invalid barcode,
spread over 6,856 distinct barcodes. Of those reads,
64.7% are within
1 nucleotide of a known barcode and the remaining
35.3% are novel.
The most abundant of them:
| barcode | count | fraction_of_reads | barcode_is | hamming_distance | closest_valid_barcode | closest_is |
|---|---|---|---|---|---|---|
| GGTCCATCTCAGATCG | 12170 | 0.003383 | novel | 6 | ||
| CTTAGGTATTACATGC | 6167 | 0.001714 | sequencing error | 1 | CTTAGGTATTATATGC | A/Lisboa/216/2023_H3N2 |
| ATGAACCCGAACCACC | 3819 | 0.001062 | sequencing error | 1 | ATGAACCCGGACCACC | A/France/GES-IPP01077/2026_H3N2 |
| GTGGTATCAAGCCGGG | 3759 | 0.001045 | novel | 6 | ||
| CAGATAATATAGAGAC | 1574 | 0.0004375 | novel | 7 |
The left panel shows each strain's fraction of the viral counts (the neutralization
standard is excluded) in wells
A8, B8, divided by the number of
barcodes for that strain: the barcodes of a strain are pooled before it is rescued, so
they cannot be balanced against each other. The right panel shows how a strain's
counts are split among its barcodes, which should be roughly even unless a barcode is
poorly represented in the rescued virus.
Each strain is added at a volume proportional to the reciprocal of its current
representation, so that all of the strains that are kept end up equally represented.
That volume relative to a typical strain is ratio_to_add, normalized so its
geometric mean over the kept strains is one: a strain with a ratio_to_add of 2 needs
twice the volume of a typical strain, and one with 0.5 needs half. The volumes are
then scaled to sum to the total_pool_volume of
10307.409 uL set in config.yml.
Re-pooling the 148 strains that are kept (see below for the 6 that are dropped) gives 10307 uL of pool, adding between 6.5 and 982 uL of each strain.
The volume of each strain to add is in
20260715_equal_vol_pool_repooling_math.csv, and the strains dropped from the
re-pooling are in 20260715_equal_vol_pool_dropped_strains.csv.
Wells A8, B8 were chosen as being
in the linear range, so a 512-fold dilution of the current pool
gives the amount of infection that the neutralization assays should use. The re-pooled
library will not have the same titer, as balancing the strains means adding a lot of
volume of the poorly growing ones.
The current pool was made by mixing equal volumes of all 154 strains, so each
strain's share g of the counts is proportional to the titer of its stock, and the
current pool's titer is the mean of those 154 stock titers. The re-pool
instead mixes volume V of each strain, so its titer is the V-weighted mean of the
same stock titers, and the ratio of the two is:
re-pool titer / current titer = 154 x sum(V x g) / sum(V) = 0.444
So the re-pool is expected to be 2.25-fold weaker, and therefore needs proportionally less diluting to give the same infection:
512 x 0.444 = 227
Use the re-pooled library at about a 227-fold dilution.
This assumes that the counts are proportional to infectivity, which is why wells in the linear range are used, and that the strain stocks have not changed titer since the current pool was made. It is a starting point rather than a substitute for titrating the re-pooled library.
The strains are combined in subpools that are then combined into the final pool, so that a problem with one subpool does not require remaking all of them.
| subpool | n_strains | volume_uL | fraction_of_pool |
|---|---|---|---|
| flu-seqneut-2026_h1 | 53 | 5656 | 0.5488 |
| flu-seqneut-2026_h3 | 66 | 1746 | 0.1694 |
| old_h1_vax | 13 | 2074 | 0.2013 |
| old_h3_vax | 16 | 830.7 | 0.08059 |
The strains needing the most volume are the poorest growing ones that are kept, and they dominate the volume of the pool. Those needing the least are the ones already best represented in the current pool.
The 5 strains needing the most volume:
| shortname | strain | subpool | mean_fraction_strain | ratio_to_add | volume_to_add_uL |
|---|---|---|---|---|---|
| flu-seqneut-2026_H1N1_1 | A/Missouri/11/2025_egg_H1N1 | flu-seqneut-2026_h1 | 9.915e-05 | 22.71 | 982.4 |
| flu_seqneut_pdmH1N1_2023-2024_VS6 | A/Wisconsin/67/2022_H1N1 | old_h1_vax | 0.0002499 | 9.009 | 389.8 |
| flu-seqneut-2026_H1N1_15 | A/SouthAfrica/PATH-CERI-C075119/2026_H1N1 | flu-seqneut-2026_h1 | 0.0003423 | 6.577 | 284.6 |
| flu-seqneut-2026_H1N1_16 | A/SouthAfrica/PATH-CERI-C073358/2025_H1N1 | flu-seqneut-2026_h1 | 0.0003618 | 6.223 | 269.2 |
| flu-seqneut-2026_H1N1_9 | A/Sydney/50/2026_H1N1 | flu-seqneut-2026_h1 | 0.0003771 | 5.97 | 258.3 |
The 5 strains needing the least volume:
| shortname | strain | subpool | mean_fraction_strain | ratio_to_add | volume_to_add_uL |
|---|---|---|---|---|---|
| flu-seqneut-2026_H3N2_27 | A/StPetersburg/RII-25-2506S/2026_H3N2 | flu-seqneut-2026_h3 | 0.01508 | 0.1493 | 6.461 |
| flu-seqneut-2026_H3N2_35 | A/Michigan/UM-10068134090/2025_H3N2 | flu-seqneut-2026_h3 | 0.01346 | 0.1672 | 7.235 |
| flu-seqneut-25to26_H3N2_25 | A/Bangkok/P2391/2025_H3N2 | old_h3_vax | 0.01312 | 0.1716 | 7.423 |
| flu-seqneut-25to26_H3N2_19 | A/Galicia/GA-CHUAC-451/2025_H3N2 | old_h3_vax | 0.01311 | 0.1718 | 7.432 |
| flu-seqneut-2026_H3N2_58 | A/Peru/ANC-INS-062/2026_H3N2 | flu-seqneut-2026_h3 | 0.01256 | 0.1792 | 7.754 |
These strains are excluded from the re-pooling for the reasons given.
| shortname | strain | subpool | mean_fraction_strain | reason |
|---|---|---|---|---|
| flu-seqneut-2026_H1N1_43 | A/Netherlands/446/2026_H1N1 | flu-seqneut-2026_h1 | 2.862e-05 | low titer |
| flu-seqneut-2026_H3N2_44 | A/California/LACPHL-INF02113/2025_H3N2 | flu-seqneut-2026_h3 | 0 | failed rescue |
| flu-seqneut-25to26_H1N1_2 | A/Galicia/GA-CHUAC-449/2025_H1N1 | old_h1_vax | 0 | no rescue |
| flu_seqneut_pdmH1N1_2023-2024_H1_22 | A/Netherlands/1739/2023_H1N1 | old_h1_vax | 5.318e-05 | low titer |
| flu-seqneut-25to26_H3N2_28 | A/Bangkok/P2323/2025_H3N2 | old_h3_vax | 0 | no rescue |
| flu-seqneut-25to26_H3N2_58 | A/England/1845724/2025_H3N2 | old_h3_vax | 0 | no rescue |