Composition of library pool 20260715_equal_vol_pool

Analysis of the barcode sequencing of the miscellaneous plate 20260715_equal_vol_pool from 2026-07-15. The plots are interactive: mouse over points for details.

Experimental description

An initial pool was made by adding equal volumes of all strains, and this pool was then used to infect MDCK-SIAT1 cells over a serial dilution of the pool. The barcodes were sequenced to determine the representation of each strain in the pool, how much of each strain to add to make a balanced pool, and the pool dilution (that is, the MOI) to use in the neutralization assays.

Warnings

These do not stop the analysis, but should be looked into.

Fates of the sequencing reads

The "fates" of the reads parsed for each well. If most reads are not "valid barcode", that could indicate a problem with the sequencing or with the barcodes being parsed.

Barcode counts per well

Wells need enough counts per barcode for the composition of the pool to be estimated accurately, and enough counts from the neutralization standard for the amount of pool in the well to be measured. Wells are flagged when they have an average of fewer than 500 counts per barcode or when fewer than 0.001 of their counts are from the neutralization standard.

The linear range of the dilution series

Diluting the pool should divide its viral counts by the dilution factor while leaving the neutralization standard alone, so the neutralization standard's share of the counts, in odds form, rises along a straight line of slope 1 against the dilution factor on log-log axes for as long as that holds. It is over the dilutions where that is true that the composition of the pool can be read, since a well too concentrated for its counts to fall with dilution measures the assay rather than the pool.

Runs of 3 consecutive dilutions are read from concentrated to dilute, and the first whose fitted slope is within 0.3 of ideal marks where the pool starts responding. Wells failing either QC threshold above are drawn below but left out of the fits, as is any well with no viral counts, whose odds are undefined.

The slope fit to each run of dilutions:

dilutions slope standard error off ideal by within tolerance chosen
4-16 -0.2857 0.1114 0.7143 False False
8-32 0.01451 0.1299 1.015 False False
16-64 0.05048 0.1357 1.05 False False
32-128 -0.1994 0.1351 0.8006 False False
64-256 -0.5735 0.1205 0.4265 False False
128-512 -0.8479 0.08604 0.1521 True True
256-1024 -1.059 0.09786 0.05916 True False
512-2048 -0.8101 0.2779 0.1899 True False
1024-4096 -1.047 0.3113 0.04708 True False
2048-8192 -1.215 0.2865 0.2153 True False

The pool starts responding to dilution at 128-fold: over 128 to 512 the slope is -0.848 +/- 0.086, off the ideal -1.0 by 0.152. The solid line is the ideal slope placed over that run, drawn faint across the rest of the series.

The composition is calculated from the middle dilution of that run, 256-fold (wells A7, B7, marked by the dashed rule), rather than from its most concentrated edge. That leaves a dilution of margin on either side of the wells used: at the edge of the run the pool has only just started responding, so a slightly misplaced onset would put the wells outside the linear range altogether.

That calculation is overridden by linear_range_wells in config.yml, which sets the wells to A8, B8. The composition below is read from those wells.

Invalid barcodes

Reads that parsed as a barcode but match neither the viral library nor the neutralization standard, over the wells the composition is read from. One within 1 nucleotide of a known barcode is sequencing error off that barcode, while one further away is a genuinely different barcode, and so material that does not belong in the pool rather than noise. The distance to the closest known barcode is the one computed by seqneut-pipeline, not recomputed here; that barcode itself is named only where it is close enough to be what the reads came from.

1.55% of the 3,597,410 parsed reads in wells A8, B8 had an invalid barcode, spread over 6,856 distinct barcodes. Of those reads, 64.7% are within 1 nucleotide of a known barcode and the remaining 35.3% are novel.

The most abundant of them:

barcode count fraction_of_reads barcode_is hamming_distance closest_valid_barcode closest_is
GGTCCATCTCAGATCG 12170 0.003383 novel 6
CTTAGGTATTACATGC 6167 0.001714 sequencing error 1 CTTAGGTATTATATGC A/Lisboa/216/2023_H3N2
ATGAACCCGAACCACC 3819 0.001062 sequencing error 1 ATGAACCCGGACCACC A/France/GES-IPP01077/2026_H3N2
GTGGTATCAAGCCGGG 3759 0.001045 novel 6
CAGATAATATAGAGAC 1574 0.0004375 novel 7

Representation of each strain in the pool

The left panel shows each strain's fraction of the viral counts (the neutralization standard is excluded) in wells A8, B8, divided by the number of barcodes for that strain: the barcodes of a strain are pooled before it is rescued, so they cannot be balanced against each other. The right panel shows how a strain's counts are split among its barcodes, which should be roughly even unless a barcode is poorly represented in the rescued virus.

Re-pooling calculations

Each strain is added at a volume proportional to the reciprocal of its current representation, so that all of the strains that are kept end up equally represented. That volume relative to a typical strain is ratio_to_add, normalized so its geometric mean over the kept strains is one: a strain with a ratio_to_add of 2 needs twice the volume of a typical strain, and one with 0.5 needs half. The volumes are then scaled to sum to the total_pool_volume of 10307.409 uL set in config.yml.

Re-pooling the 148 strains that are kept (see below for the 6 that are dropped) gives 10307 uL of pool, adding between 6.5 and 982 uL of each strain.

The volume of each strain to add is in 20260715_equal_vol_pool_repooling_math.csv, and the strains dropped from the re-pooling are in 20260715_equal_vol_pool_dropped_strains.csv.

Dilution to use for the re-pooled library

Wells A8, B8 were chosen as being in the linear range, so a 512-fold dilution of the current pool gives the amount of infection that the neutralization assays should use. The re-pooled library will not have the same titer, as balancing the strains means adding a lot of volume of the poorly growing ones.

The current pool was made by mixing equal volumes of all 154 strains, so each strain's share g of the counts is proportional to the titer of its stock, and the current pool's titer is the mean of those 154 stock titers. The re-pool instead mixes volume V of each strain, so its titer is the V-weighted mean of the same stock titers, and the ratio of the two is:

re-pool titer / current titer = 154 x sum(V x g) / sum(V) = 0.444

So the re-pool is expected to be 2.25-fold weaker, and therefore needs proportionally less diluting to give the same infection:

512 x 0.444 = 227

Use the re-pooled library at about a 227-fold dilution.

This assumes that the counts are proportional to infectivity, which is why wells in the linear range are used, and that the strain stocks have not changed titer since the current pool was made. It is a starting point rather than a substitute for titrating the re-pooled library.

Volume of each subpool

The strains are combined in subpools that are then combined into the final pool, so that a problem with one subpool does not require remaking all of them.

subpool n_strains volume_uL fraction_of_pool
flu-seqneut-2026_h1 53 5656 0.5488
flu-seqneut-2026_h3 66 1746 0.1694
old_h1_vax 13 2074 0.2013
old_h3_vax 16 830.7 0.08059

Strains needing the most and least volume

The strains needing the most volume are the poorest growing ones that are kept, and they dominate the volume of the pool. Those needing the least are the ones already best represented in the current pool.

The 5 strains needing the most volume:

shortname strain subpool mean_fraction_strain ratio_to_add volume_to_add_uL
flu-seqneut-2026_H1N1_1 A/Missouri/11/2025_egg_H1N1 flu-seqneut-2026_h1 9.915e-05 22.71 982.4
flu_seqneut_pdmH1N1_2023-2024_VS6 A/Wisconsin/67/2022_H1N1 old_h1_vax 0.0002499 9.009 389.8
flu-seqneut-2026_H1N1_15 A/SouthAfrica/PATH-CERI-C075119/2026_H1N1 flu-seqneut-2026_h1 0.0003423 6.577 284.6
flu-seqneut-2026_H1N1_16 A/SouthAfrica/PATH-CERI-C073358/2025_H1N1 flu-seqneut-2026_h1 0.0003618 6.223 269.2
flu-seqneut-2026_H1N1_9 A/Sydney/50/2026_H1N1 flu-seqneut-2026_h1 0.0003771 5.97 258.3

The 5 strains needing the least volume:

shortname strain subpool mean_fraction_strain ratio_to_add volume_to_add_uL
flu-seqneut-2026_H3N2_27 A/StPetersburg/RII-25-2506S/2026_H3N2 flu-seqneut-2026_h3 0.01508 0.1493 6.461
flu-seqneut-2026_H3N2_35 A/Michigan/UM-10068134090/2025_H3N2 flu-seqneut-2026_h3 0.01346 0.1672 7.235
flu-seqneut-25to26_H3N2_25 A/Bangkok/P2391/2025_H3N2 old_h3_vax 0.01312 0.1716 7.423
flu-seqneut-25to26_H3N2_19 A/Galicia/GA-CHUAC-451/2025_H3N2 old_h3_vax 0.01311 0.1718 7.432
flu-seqneut-2026_H3N2_58 A/Peru/ANC-INS-062/2026_H3N2 flu-seqneut-2026_h3 0.01256 0.1792 7.754

Dropped strains

These strains are excluded from the re-pooling for the reasons given.

shortname strain subpool mean_fraction_strain reason
flu-seqneut-2026_H1N1_43 A/Netherlands/446/2026_H1N1 flu-seqneut-2026_h1 2.862e-05 low titer
flu-seqneut-2026_H3N2_44 A/California/LACPHL-INF02113/2025_H3N2 flu-seqneut-2026_h3 0 failed rescue
flu-seqneut-25to26_H1N1_2 A/Galicia/GA-CHUAC-449/2025_H1N1 old_h1_vax 0 no rescue
flu_seqneut_pdmH1N1_2023-2024_H1_22 A/Netherlands/1739/2023_H1N1 old_h1_vax 5.318e-05 low titer
flu-seqneut-25to26_H3N2_28 A/Bangkok/P2323/2025_H3N2 old_h3_vax 0 no rescue
flu-seqneut-25to26_H3N2_58 A/England/1845724/2025_H3N2 old_h3_vax 0 no rescue