- Home
- Angry by Choice
- Catalogue of Organisms
- Chinleana
- Doc Madhattan
- Games with Words
- Genomics, Medicine, and Pseudoscience
- History of Geology
- Moss Plants and More
- Pleiotropy
- Plektix
- RRResearch
- Skeptic Wonder
- The Culture of Chemistry
- The Curious Wavefunction
- The Phytophactor
- The View from a Microbiologist
- Variety of Life
Field of Science
-
-
Change of address1 year ago in Variety of Life
-
Change of address1 year ago in Catalogue of Organisms
-
-
Earth Day: Pogo and our responsibility1 year ago in Doc Madhattan
-
What I Read 20241 year ago in Angry by Choice
-
I've moved to Substack. Come join me there.1 year ago in Genomics, Medicine, and Pseudoscience
-
-
-
-
Histological Evidence of Trauma in Dicynodont Tusks7 years ago in Chinleana
-
Posted: July 21, 2018 at 03:03PM8 years ago in Field Notes
-
Why doesn't all the GTA get taken up?8 years ago in RRResearch
-
-
Harnessing innate immunity to cure HIV10 years ago in Rule of 6ix
-
-
-
-
-
-
post doc job opportunity on ribosome biochemistry!11 years ago in Protein Evolution and Other Musings
-
Blogging Microbes- Communicating Microbiology to Netizens11 years ago in Memoirs of a Defective Brain
-
Re-Blog: June Was 6th Warmest Globally12 years ago in The View from a Microbiologist
-
-
-
The Lure of the Obscure? Guest Post by Frank Stahl14 years ago in Sex, Genes & Evolution
-
-
Lab Rat Moving House14 years ago in Life of a Lab Rat
-
Goodbye FoS, thanks for all the laughs15 years ago in Disease Prone
-
-
Slideshow of NASA's Stardust-NExT Mission Comet Tempel 1 Flyby15 years ago in The Large Picture Blog
-
in The Biology Files
Not your typical science blog, but an 'open science' research blog. Watch me fumbling my way towards understanding how and why bacteria take up DNA, and getting distracted by other cool questions.
Learning from the experts, part I
To get evidence for USS-dependent binding, we should collect data for a histogram of rupture forces (i.e. how strong a pull is needed to break the bead free of the cell), for DNAs with and without USS. If binding is distinct from uptake, this distribution should be bimodal, probably with a much higher peak for the low 'binding' rupture force. She said this would be quite a lot of work (weeks, at least).
Can we stick H. influenzae cells onto a glass coverslip, (as the B. subtilis people did) rather than restraining them with a micropipette? This would make the manipulations much easier. She recommended trying BSA-coated coverslips as well as silane-coated ones. The coating will also help reduce non-specific sticking of DNA to the glass.
To test whether uptake is due to 'diffusive motor' or a 'power stroke' (I'll need to read up on these concepts) we should create a force-velocity curve. If we find that, below a critical stalling point, the velocity of DNA uptake is independent of the opposing force pulling the bead back into the center of the laser trap, we would conclude that uptake was due to machinery exerting a power stroke, whereas if we observed a gradual slowing of uptake with increasing opposing force, it would be due to a diffusive motor. I think Tfp use a power stroke, and a process driven by flow of ions (maybe the PMF (= proton motive force?) would be diffusive. The B. subtilis rate is largely independent of the opposing force on the bead, although it appears to somehow be generated by the PMF.
If uptake is indeed mediated by a short pseudopilus that acts as a ratchet on the DNA, we might be able to see the ratcheting in a velocity vs time (or force vs time?) graph. This would depend on the resolution of the velocity (or force?) measurements. If the pseudopilus is only, say 20 nm long (a guess of the thickness of the periplasm), we would need to detect fluctuations on this scale.
In principle optical traps can exert forces greater than 100 pN on trapped beads. However DNA deforms at 65 pN, allowing it to stretch. Not only will such stretching complicate length/velocity measurements but the deformation may perturb interactions with the uptake proteins, causing them to lose their grip on the DNA. So we should keep our measurements below about 60 pN.
She recommended making a flow chart to lay out the tests we will do and how the results will affect our hypotheses and guide our next steps.
Six sentences in Science!
IN HIS NEWS FOCUS STORY “ON THE ORIGIN of sexual reproduction” (5 June, p. 1254), C. Zimmer highlights the importance of the phylogenetic perspective championed by John Logsdon, but by considering only eukaryotes he overlooks an important bacterial clue to the evolution-of-sex puzzle.
Until recently, bacteria were thought to be sexual; they have well-characterized processes that cause recombination of chromosomal alleles, and these parasexual processes were assumed to have evolved for recombination in the same way as meiotic sex in eukaryotes. However, a more critical analysis of the genes responsible for the parasexual processes suggests that they did not evolve for sex after all. Instead, the chromosomal recombination they cause appears to arise as unselected effects of related processes, the evolutionary functions of which are well established (1).
The fact that bacteria lack genes evolved for recombination indicates that meiotic sex must have evolved in eukaryotes to solve a problem that bacteria don’t have. Bacteria apparently get whatever recombination they need by accident—why do eukaryotes need so much more?
ROSEMARY J. REDFIELD
Department of Zoology, University of British Columbia, Vancouver, BC V6T 3Z4, Canada. E-mail: redfield@interchange.ubc.ca
We need a hypothesis.
I can come up with a hypothesis for the first parts (the ones covered by the current post-doc's NIH fellowship proposal). "Every stage of transformational recombination is limited by sequence biases." This hypothesis might well be false - we only know of biases at two steps: DNA uptake and mismatch repair. But it provides a clear explicit framework for finding out where and what the biases are.
But I would like to extend the research to include the biases due to the differences in different isolates. Well, that's a way of saying it that sort-of fits the hypothesis but doesn't really capture what I want to propose. Instead I should first describe what I would like to propose, and then try to find a way to describe it that can be part of a unified hypothesis.
I would like to propose to find out the molecular reasons why different isolates of H. influenzae differ so much in their ability to transform. This should be fairly straightforward, given the molecular and genetic/genomic tools we have. We can also try to estimate how long these strains have had their specific transformation phenotypes - did a very recent ancestor get a mutation (or recombination!) that inactivated a competence/transformation gene, or is it descended from a long lineage of non-competent ancestors.
But I would also like to find out how different are the recombination histories of strains that can and can't transform - looking not just at the individual sequences they have acquired, replaced, or lost, but at the pattern underlying these events. Do strains that don't transform at all in the lab have a history of greatly reduced recombination, relative to strains that transform very well in the lab?
I don't know how do-able this is. If we get genome sequences of enough strains, can we build up a detailed picture of their recombination histories? The previous post-doc used a program called Mauve to identify recombination tracts...
I also don't know to what extent this analysis will be facilitated by the information we'll get from the first parts described above. We'll first build up a complete picture of the relationship between (i) a pair of donor and recipient genomes and (ii) the recombinant genomes that result from DNA uptake. (If resources permit we can do more than one pair of genomes.) Then we try to use this knowledge to infer the recombinational histories of other genomes. Is there a hypothesis in there? Maybe the hypothesis is that this will be possible. We could hypothesize that "Understanding the limits to transformational recombination in lab cultures will let us identify the recombination histories of natural isolates."
But as science this is what I consider a 'pseudo-hypothesis'. It's not a hypothesis about the nature of reality, but a hypothesis about the nature of our abilities. How about "The limits to transformational recombination in lab cultures explain the recombination histories of natural isolates."
Hmm, if that's our hypothesis then maybe this proposal does belong in the Evolution of Infections Disease program after all, and not in the Prokaryotic Cell Growth, Differentiation and Adaptation program. I think I'll handle this question by putting together a decent summary (once we have a single hypothesis) and asking both Program Directors where they think it belongs.
Competence variation: causes and consequences
A potential postdoc is interested in the role of gene exchange in natural populations, and I think this might be a very nice way to build on other work we've done and will be doing. It would also fit well into the NIH proposal we're preparing.
Work we've done: We know that DNA uptake and transformability are very variable in natural populations, but we don't know why (neither the molecular nor the evolutionary/ecological 'why'). The best study was by a recent postdoc, who analyzed 30+ natural isolates of H. influenzae for both DNA uptake and transformation. She found that very few took up as much DNA or transformed as well as the lab strain KW20. Some of these strains had sequenced genomes, but in most cases the genotype did not explain the phenotype.
Work we will be doing: A new postdoc is investigating the molecular constraints on transformation, to fully characterize the sequence biases that affect each stage of DNA uptake and recombination. This should give us a good picture of what transformational recombination looks like, at the genome-wide level.
The nice extension has two parts:
Part 1: Before we can hope to understand the evolutionary cause of the variation in competence and transformability, we need to find out its molecular causes - what are the genetic changes underlying the differences in DNA uptake and transformation? This is probably best done by a combination of genetics and genome sequencing. Genome sequencing has gotten so cheap that we can afford to do a lot of strains, not probably to the assembly stage (too much person-work), but enough to get the sequences of all the homologs of genes known from KW20 and other strains with assembled genomes. We could certainly do all the strains the postdoc characterized, but we might be able to do a much more comprehensive survey. (Maybe someone else is already doing this... I've just emailed the likely suspect so we can coordinate our efforts.)
With lots of genome sequences of strains of known phenotype, we can look for the individual causes of the different competence phenotypes and also look for patterns in these causes. Identifying the causes for individual strains will entail benchwork, using natural transformation to confirm that specific alleles are responsible for specific phenotypes. I would think that some of this could be a project for an undergraduate or M.Sc. student, but the new post-doc is bursting with ideas of how to do this efficiently for many strains - I’ll let him post these on his blog. Then we can ask whether the causes appear to be a random subset of the possible mutations, or whether there are consistent causes in the different strains. We can also ask how recent the changes are (have multiple mutations accumulated in the competence genes?), and maybe even ask whether the most recent ancestors were indeed competent (but this may not be possible - see our recent PLoS One paper).
Part 2: Once we have genome sequences and the results of the other post-doc’s genome-wide recombination analysis, we can not only examine the genomes for evidence of recombination, but also ask whether the recombination fits the pattern expected for transformation, and whether the differences in recombination correlate in any way with the differences in transformability. For example, do nontransformable strains have a history of reduced recombination? If yes, can we tell how recent this is? And is only the transformational component of recombination reduced? Is there evidence that transformation by other causes (conjugation, transduction) is increased in strains that don't transform? Can we tell whether the changes causing reduced competence and transformability are themselves due to recombination?
This is a great project because it can go in lots of directions and because all of the possible answers are interesting.
Do competent cells bind DNA?
Years ago I had a grad student working on a related question, and part of our motivation was the issue of whether sequence specificity is intrinsic to the process of DNA uptake, as is predicted by our hypothesis that it's a way to help get stiff inflexible DNAs across the outer membrane. If instead the specificity is caused by a protein external to the uptake machinery, this would be consistent with hypotheses positing a benefit of screening DNA for relatedness before letting it into the cell.
The practical obstacle to answering this question is that we have no good assay for DNA binding (and of course also no mutants that can still bind DNA but can't take it up). At present, DNA binding must be measured indirectly, by giving cells radioactively labeled DNA and comparing cell-associated cpm with and without pretreatment with DNase, which removes DNA that has bound to cells but not been taken inside them.
The Neisseria paper (Aas et al, 2002) took advantage of the pilT gene, which H. influenzae and its relatives don't have. They showed that wildtype Neisseria cells associate with DNA, and that this does not require the DUS uptake sequence. Wildtype cells take up most of the DUS-containing DNA they bind - after a 30 minute incubation about 70% of the cell-associated DNA is DNase I-resistant - but take up less than 1% of DNA lacking a DUS.
PilT retracts the type 4 pili filaments thought to bind DNA into the cell (mutants that don't make pili don't associate with DNA), and they predicted that cells lacking PilT would bind DNA to their surface but be unable to bring it in. But very little DNA associated with their pilT mutants even when no DNase I was used. So the non-specific DNA binding seen on wildtype cells requires not just the machinery thought to assemble pili but also the protein thought to disassemble them.
Cells lacking the major pilus protein (pilin) encoded by pilE don't have visible pili and they didn't bind DNA or take up DNA, but cells lacking the minor pilin-like protein encoded by comP do have visible pili and they bound DNA normally, but they didn't take any of it up even if it had DUSs. Overexpressing the normally-scarce ComP protein increased by 20-fold the amount of DNA bound and taken up, and proportionately increased the transformation frequency. This increased uptake was specific for DNA containing a DUS, although a modest increase was also seen for DNA that lacked DUS.
Overexpressing ComP also caused increased and DUS-dependent DNA uptake by cells carrying a knockout of pilT, although the effect was small - these cells took up only about 5% as much DNA as wildtype cells with normal ComP expression.
What does all this mean? Because they couldn't show that ComP binds DNA, the authors concluded that it probably acts indirectly, by activating or stabilizing another component that does the DNA binding. They also thought that ComP may normally act as a minor component of pili composed mainly of PilE. They also suggested that pilus retraction by PilT may not be the main force pulling DNA into the cell; instead the DNA may be quickly handed over to a periplasmic receptor that does the bulk of the pulling-in.
Back to the DNA uptake proposal
The first goal is to get my mind clearly focused on them, so I'm starting by going through the draft proposal to CIHR, which at present is largely unchanged from its 2007 version. One (depressing) advance is that I now realize that the Introduction is a failure. It consists of three very nice paragraphs, clear and well-written. These paragraphs contain the following: 1) an introduction to natural competence; 2) the broad goals of my research program; 3) the medical relevance of DNA uptake and transformation. Unfortunately they're not the paragraphs the Introduction needs - most of the information they provide doesn't belong in the Introduction.
What should the Introduction introduce? The proposal is about the mechanism of DNA uptake, so it should probably start by framing the problem - that bacteria take up DNA but we don't know how the DNA gets across the membranes (maybe say why I think this is likely to be difficult). Maybe contrast this with uptake of simple molecules? And the medical relevance could be supplemented with some cell biology relevance - for understanding related processes, maybe getting a better picture of the roles and mechanism of type 4 pili. I should also reframe the broad goals of our research, with less emphasis on evolutionary function and more on the totality of competence.
I'm also reconsidering which committee I should direct this proposal to. The previous version went to the Microbiology and Infectious Diseases committee, partly because they had approved our previous proposal, but also because I was simultaneously sending a different proposal (on gene regulation) to the Biochemistry and Molecular Biology committee. That proposal was funded. This new proposal is quite biochemical, so I'm wondering if it should go the the BMB committee.
OK, now I need to reread the proposal, to once again remind myself of what it's all about.
Manuscript drags on
Analysis of uptake sequences by score
Here are the logos for the N. meningitidis and H. influenzae uptake sequences after sorting the occurrences by the scores that the Gibbs motif Sampler assigned them. (I'm pretty sure that each score is a measure of how well that occurrence's sequence matches the position weight matrix that Gibbs determined for this data set, but I don't know how the calculation is done.)The top set are the logos for 5381 N. meningitidis DUSs. The numbers are different than in yesterday's post because I realized I had been analyzing a N. gonorrhoeae data set. The overall picture is the same for N. meningitidis and N. gonorrhoeae - low-scoring DUS retain strong consensus for most of the central positions but have only very weak consensuses for the other positions. The drop-off is quite steep. The shapes of the logos are about the same for all the occurrences with scores lower than about 0.95.
The H. influenzae dataset is even more skewed; almost 60% of the USSs have perfect scores, and about 8% have zero scores. But the consensus decays fairly evenly across the positions, and even the zero-score occurrences have the full motif. Like the N. meningitidis DUS, the shapes of the USS logos are about the same for all occurrences with scores below 0.95.I think the question in my mind was whether there is a obvious place to draw a line between 'real uptake sequence' and 'degenerate sequence that doesn't deserve to be treated as an uptake sequence'. Unfortunately the analysis is complicated by the different sizes of the datasets - the N. meningitidis set has almost twice as many sites as the H. influenzae set.
OK, I've dug out another set of H. influenzae runs, done with a high 'expected' setting to maximize the number of sites found. This has 3466 USSs, with a lot more having zero scores than in the previous set. Now the first and last Gs in the core are seen to be weaker in USSs with low scores, though not in the larger set of USSs with zero scores. Overall the consensus still remains constant as the scores and consensus strengths decrease. Notably, the flanking AT-rich segments remain as important in poorly matched USSs as the core does.
Poster
While there I sat down with my former post-doc to discuss our manuscript on uptake sequence variation. We agreed that it needed major reorganization more urgently than it needed simple editing and polishing, so we worked out a new structure that we think will hold the ideas together much better. I made some new figures (cartoons of our explanation and of how our simulation model works) and rearranged the text into its now order, but I needed a paper copy to do any serious editing. Now I'm home I've printed out the rearranged draft and am hoping to do one quick pass through it and then set it aside till our grant proposals are done.
But of course I immediately got distracted by the data. I want to be able to say something about how our new analysis of uptake sequences as motifs gives us insight into their evolution. The most promising angle is the distribution of good and poor matches to the motif, which I can analyze because the Gibbs Motif Sampler assigns a score to each occurrence it finds.
I dug out a set of 4646 DUSs Gibbs found in the N. meningitidis genome, and sorted them by score. And now I've spent a lot of time trying to force Excel to draw a proper histogram. I get the histogram values using a math teachers' web site called Illuminations, and paste it into Excel, but Excel refuses to use the data ranges as the X-axis (instead using its line numbers). I've found a work-around - the graph is very ugly but here it is. The red bars are the numbers of DUS with each score range (0-.02, .02-.04, etc) and the lilac bars are the cumulative numbers with increasing scores. So about 0.5% of DUS have zero scores, very few have non-zero scores lower than 0.5, about 2% have scores in each category from 0.5 to 0.98, and more than 40% have scores greater than 0.98.
I made weblogos for the top-scoring 50% and bottom-scoring 50% of the DUS occurrences (my previous analysis had only looked at the high-scoring ones). Here they are; the bottom 50% logo isn't evenly weak at all positions, instead it's quite strong at some positions and very much weaker at others. I don't know what I'm going to do with this analysis... I guess I could make a range of logos, maybe for the 10 deciles (is that the word I want?), to see how the consensus decays. And I should probably do the same thing for the H. influenzae USSs.
Controlled nicking of supercoiled plasmid DNA

Last week I tried to introduce single-strand nicks into a closed circular plasmid (pUSS-R) by doing restriction digestion in the presence of ethidium bromide (background is here). I tested two different dilutions of HindIII, a restriction enzyme that should cut this plasmid once, with three different concentrations of ethidium bromide, to find conditions where a convenient range of incubation times gave a range of partially digested molecules.
I was hoping to see what's shown in the upper gel drawing - appearance of a novel band that migrated slower than the fully digested DNA in the rightmost lane and much slower than the supercoiled DNA in the leftmost lane. This new band would contain plasmid that had undergone a single-strand nick at its single HindIII site. I didn't know whether I might also see eventual appearance of linear DNA. But instead I saw what's shown in the middle gel drawing - gradual appearance of a linear-sized band as the supercoiled band disappeared.
I wondered if the HindIII now being sold no longer causes nicking (maybe New England Biolabs has 'improved' their HindIII clone...). So I did the control shown in the lower gel drawing, incubating pUSS-R with a very low concentration of DNase I, an enzyme widely used to create nicks in double-stranded DNA. Again I expected to see a novel, slow-migrating band, but instead saw only a linear-sized band. Damn!
I also ran samples from a few old plasmid preps. Most of these contained several bands, but I don't know if the slowest one is relaxed circles or dimers, because the plasmids may have been prepared from rec+ cells. I've now also checked the old literature, just in case I was mistaken in expecting nicked/relaxed DNA ("form II") to migrate slowly, but indeed it should (see for example the marker lanes in Fig. 2 on PNAS 86:1309).
Now what? Is the problem that the nicked DNA migrates at the same speed as linear DNA? Is this specific to this particular plasmid? Or do I not have any nicked DNA? Should I try another plasmid?
What to propose to NIH?
Goal: To fully characterize all of the biases and sequence specificities of transformational recombination in H. influenzae.
Specific questions to answer:
- Is DNA binding a distinct step that precedes the initiation of DNA uptake? If so, does it have the same sequence specificity as DNA uptake? Does it have any topological specificity?
- What is the complete sequence specificity of DNA uptake? How absolute is the requirement for a good match to the USS consensus at uptake (are non-USS DNAs taken up at lower frequency or not at all)? This will be investigated with plasmids or short fragments containing 12% degenerate USSs, and with ones containing completely random sequences (we could create these or just use fragments of unrelated DNA). Does uptake have any topological specificity?
- Does the translocation step impose any sequence specificity? How strict is the requirement for a pre-existing free end (does circular DNA sometimes get cut or nicked in the periplasm, or transported intact into the cytoplasm?
- Is there any sequence specificity to the DNA degradation that accompanies uptake and translocation?
- What proteins interact directly with DNA during binding, uptake and translocation?
- What recombination biases affect indels? (How efficiently are different indels transfered by recombination?
- What are the recombination biases along the full length of the chromosome?
- How does mismatch repair affect the outcome of transformational recombination? How much of the recombination bias found by #7 is due to mismatch repair?
Proposal priorities
Specific Aims:
- "What is the H. influenzae uptake specificity? A pool of USSs that have been intensively but randomly mutagenized and then selected for the ability to be taken up by competent cells will be sequenced to fully specify the uptake bias." I've been getting info about designing and ordering the degenerate oligos. The post-doc had already set up a spreadsheet that does the calculations I wanted to think about, and he and I agreed that we should start with 12% degeneracy rather than 9%. This will reduce the fraction of oligos that are strongly preferred, giving more sensitivity for detecting weaker effects. I've had replies from two custom-oligo companies, and the good news is that our degenerate oligo pool will be easier to get and much cheaper than I had expected. For the proposal we'll only need to do some conventional sequencing, and as our USSs will already be in plasmids we think we'll just sequence each plasmid insert separately. This will be wasteful but not very expensive if we do the DNA reactions and cleanups ourselves, and probably cheaper when we consider the time/money we'll save by not having to troubleshoot a more 'efficient' sequencing strategy.
- "What forces act on DNA during uptake? Laser-tweezer analysis of USS-dependent uptake by wild type and mutant cells will reveal the forces acting on the DNA at both the outer and inner membranes." My physicist collaborator is keen to have me back to get this working. The only think I could aim to get done before September 15 is to attach some chromosomal DNA to the styrene beads and show that Bacillus subtilis will bind to it and pull on it in the tweezers apparatus, and that H. influenzae doesn't stick nonspecifically to beads with no DNA on them (a concern raised by a reviewer). The biotin-linked chromosomal DNA I'll use for this will also be used in other preliminary experiments.
- "Does the USS polarize the direction of uptake? Using magnetic beads to block uptake of either end of a small DNA fragment will show whether DNA uptake is symmetric around the asymmetric USS." Our 1 micron paramagnetic streptavidin beads are on their way, and I'll use these to check for non-specific binding of cells to beads and for specific binding of competent cells to beads with DNA on them (using the same DNA prepared above). I found the source of the 50 micron beads and will order them tomorrow; maybe I can use them in the same way.
- "Does the USS increase DNA flexibility? Cyclization of short USS-containing fragments will reveal whether the USS causes DNA to be intrinsically bent or flexible, and whether ethylation or nicking can replace parts of the USS." I'm going to try the nicking protocol tomorrow.
- "Which proteins interact with incoming DNA? Cross-linking proteins to DNA tagged with magnetic beads, followed by HPLC-MS, will be used to isolate and identify proteins that directly contact DNA on the cell surface." We'll have the DNA on the big magnetic beads (see 3 above) and can use this to try out the formaldehyde cross-linking. It would be good to show that we can distinguish between non-specific proteins or peptides (independent of both DNA and competence) and peptides cross-linked only when cells are competent and the beads have DNA on them. One of the reviewers really liked this part, but thought we should be more ambitions and propose more diverse approaches to this problem, including in vitro cross-linking to purified secretin.
- "Which proteins determine USS specificity? Heterologous complementation with homologs from the related Actinobacillus pleuropneumoniae (which recognizes a variant USS) will identify the proteins responsible for this specificity." It would be good to have results of some complementation experiments with single gene plasmids, as the whole-operon plasmids seemed to cause growth problems.
Does EthBr turn restriction enzymes into nickases?
Restriction Enzyme Nicking Reactions. Covalently closed circular pBR322 DNA was nicked with restriction endonucleases HindIII, Cla I, or BamHI by incubating 10 pkg of plasmid DNA in a 100-Al solution of 20 mM Tris HCl, pH 7.8, 7 mM MgCl2, 7 mM 2-mercaptoethanol*, gelatin* (100 Ag/ml) and a concentration of ethidium bromide (50 ug/ml, 75 ug/ml, or 100 ug/ml, respectively) determined by titration to give an optimal level of nicking. An amount of restriction endonuclease was added sufficient to convert 50-90% of the input DNA to an open circular form on incubation at room temperature for 2-4 hr. The nicking reaction with the EcoRI enzyme consisted of 100 mM Tris'HCl, pH 7.6, 50 mM NaCL, 5 mM MgCl2, gelatin* (100 ug/ml), ethidium bromide (150 ug/ml). Reactions were stopped by addition of excess EDTA followed by phenol extraction and ethanol precipitation.I have lots of supercoiled plasmid that has one EcoRI or BamHI or HindIII site, and a big bottle of 1 mg/ml EthBr. So I'll just set up a series of digests with different concentrations of EthBr and restriction enzyme and the standard digestion buffers*, and run a gel to see what I get.
*In the old days we used to add mercaptoethanol and gelatin to stabilize our restriction enzymes, but I won't bother.
Email to potential suppliers of a degenerate USS oligonucleotide
I'm writing to inquire about ordering a large highly degenerate oligonucleotide.
We're hoping to obtain an oligonucleotide that will be about 50 nt long and 12% degenerate at every position (so this will really be a population of degenerate nucleotides). Below I've listed specifications.
Would you be able to synthesize such an oligonucleotide for us? If so, could you let me know the estimated cost and time frame, and any other issues you think we should consider?
Thanks very much,
Rosie Redfield
Sequence: Not yet finalized, but probably similar to: AAAGTGCGGTTAATTTTTACAGTATTTTTACTGGTACCATATGAATTCCA
Degeneracy: Every position should be degenerate, with 4% each of the three other bases. For example, every position specified as an A should be 88% A and 4% each of T, G and C.
Scale: Because synthesizing this oligo will probably require a dedicated run, we would like to obtain as much as can easily be prepared in a single run.
(Calculation: 1 micromole of a 50 nt fragment is 6.02 x 10^17 molecules, with mass = 15 mg.)Modifications: None.
Purification: For our experiments it will important to minimize the fraction of oligos that are not full length. What purification methods do you recommend for oligos of this length, and what contamination level should we expect?
Synthesis and sequencing of degenerate USS
- synthesizing the degenerate USS
- putting the degenerate USS in a plasmid vector
- having cells selectively take up plasmids from the pool
- reisolating the plasmids that have been taken up
- sequencing representative USS from the two plasmid pools (input and taken-up)
- interpreting the sequences (characterizing the sequence diversity)
First uptake experiment
Starting experiments again
The post-doc showed me the stock of the plasmid I'll use (pUSS-1, left by a previous post-doc). So this afternoon I'm just running a quick gel to check the state of this DNA. And pouring some plates to streak out the wildtype and rec-2 cells I'll be using. If the DNA is supercoiled I'll try it out tomorrow on some wildtype competent cells I have frozen.
We've also ordered the tester-kit of four kinds of streptavidin-coated magnetic beads with different surface properties. They're much bigger than I wanted, but I figure I can at least use them to find out whether H. influenzae cells stick to any of (or all of?) these beads in the absence of DNA. If they do, the planned experiments may not work.
The new post-doc's plan
tiny magnetic beads?
Now I can't find the supplier that had the 50 nm streptavidin-coated beads at all. The best I can find is a European company, Ademtech, that has 100 and 200 nm beads. I've sent them an email asking if they have a Canadian supplier, but I'm not optimistic. Maybe I should just try the tester kit of big beads, and then scale down...
Invitrogen charges more than $500 for a stand that holds 16 microfuge tubes against rare-earth magnets, to separate the beads from the liquid they're suspended in. The 6-tube one is only about $150. But, I can buy a sampler of 50 NdFeB magnets of assorted sizes (1/8 inch to 1/2 inch, rods and discs) from Lee Valley Tools for $16.50.
Real USS spacings are not random

The co-author only needed about 20 minutes to write the little perl script I'd asked for, and here are the results I generated with it. The blue bars show the spacing distribution of the real H. influenzae uptake sequence cores, and the red bars show the spacing distribution of the same number of random positions in a sequence of the same size (based on 10 replicates).
So the statistical advisor was correct; the real distributions have far fewer close spacings than expected for a random sequence. This is despite the the presence of two forces expected to increase close spacings - USS pairs serving as transcription terminators, and general clustering of USS in non-coding sequences. As hypothesized in yesterday's post, the discrepancy between real and random is stronger when we eliminate the terminator effect by only considering spacings between USS in the same orientation.
And the previous mathematical analysis was right, the real spacing distribution is more even than the random one - not only are there fewer close USS than expected, there are fewer far-apart USS too. (This is more obvious in the upper graph.)
Now I need to modify the perl script that locates uptake sequences so it will find the Neisseria DUS, and then do the same analysis for the N. meningitidis genome. The previously published analysis of these spacings simply compared the mean spacing of DUS with the mean length of apparent recombination tracts, concluding that the similar lengths was consistent with the hypothesis that DUS arise or spread by recombination.
Spacing of uptake sequences
I've been looking more at the spacings of the uptake sequences that arise in our model genomes and that are present in the real H. influenzae genome. First, in the panel on the left, is what we see in our simulated genomes. I only have data for simulations done with relatively small recombination fragments because the others take a very long time to run.The notable thing (initially I was surprised) is that almost all of the uptake sequences in these genomes are at least one fragment length apart. That is, if the fragments that recombined with the genome were 25 bp (top histogram), none of the resulting evolved uptake sequences are in the 0-10 and 10-20 bins. If the fragments were 500 bp (bottom histogram, not to the same scale), none are in any bin smaller than 500 bp.
For the 100 bp fragments I've plotted both the spacings between perfect US (middle green histogram) and the spacings between pooled perfect and singly mismatched US (blue histogram below the green one). When the mismatched US are included the distribution is tighter, but there are still very few closer than 500 bp.
This result makes sense to me, because the uptake bias only considers the one best-matched US in a fragment. So a fragment that contained 2 well-matched USs is no better off than a fragment with one, and selection will be very weak on USs that are within a fragment's length of another US.
I talked this result over with my statistics advisor, and we agreed that there's not much point in worrying about whether the distributions are more even than expected for random spacing, because these are so obviously not random spacing.
After getting these results I decided I should take my own look at the spacings of the real uptake sequences in the H. influenzae genome, and these are shown in the graphs in the lower panel. I first checked the spacings of all the uptake sequences that are in the same orientation; the upper graph shows the pooled data for spacings between 'forward' USS and spacings between 'reverse' USS. Let's consider this graph after the one below it.

The lower graph shows the spacings between all of the USSs, regardless of which orientation each is in. This is the data that has been previously analyzed and found to be non-random (in a mathematical analysis that's too sophisticated for me to follow).
We don't see the same abrupt drop-off at close spacings that's in the simulated genomes, but the statistics advisor thought the drop-off at close spacings might be significant. Rather than making a theoretical prediction, he suggested just generating randomly spaced positions (same number of occurrences as the real USS, and in the same length of sequence as the real genome) and comparing their distributions to the real ones in my graph. I'm asking my co-author to write me a little perl script to do this (yes, I probably could write it myself, but she can write it faster). I'll use it to generate 10 random pseudo-genomes worth of spacings, and compare their distributions to my real genome.
There are good reasons to expect the real distribution to not be random. First, uptake sequences are preferentially found outside of genes. In H. influenzae 30% of the USS are in the 10% of the genome that's non-coding, and this would be expected to cause some clustering, increasing the frequencies of close spacings. Second, USS often occur in closely spaced, oppositely oriented pairs that act as transcriptional terminators. This is also expected to increase the frequencies of close spacings. I can't easily correct for the first effect, but I can correct for the second one by only considering USS that are in the same orientation, and that's what I've done in the upper graph.
There we see a much stronger tendency to avoid close spacings. The numbers aren't really comparable because I'm only considering half as many USS in each orientation, but the overall shapes of the two distributions suggest fewer close spacings in the same-orientations graph.
So, two things to do: First, predict the spacings of random positions, using the perl script that isn't written yet. Second, I need to get better data for the simulated evolved genomes, because these were initiated with fake genomes consisting of pure strings of uptake sequences, and the uptake sequences in the resulting evolved genomes had not originated de novo but instead had been preserved from some of the originating ones. I need to redo this with evolved genomes that began as random sequences, but these are taking longer to simulate.
Uptake of closed circular DNA
How to do it? How was the original experiment done? (First step is to find the paper...) Would they have used radioactively labeled DNA, and if so how did they label it without compromising its supercoiled state? (Prep the plasmid from E. coli grown in the presence of massive amounts of 32-P? Yikes!) Or maybe they did a Southern blot to detect it? Find the paper...
Basic plan: Incubate competent cells (wild type or rec-2 mutant) with supercoiled plasmid containing a USS - pUSS1 isolated from E. coli using a kit should be fine, except for the labeling question. Treat cells with DNase I and wash them thoroughly to get rid of external DNA. Lyse cells and prep DNA using a plasmid-prep kit. Run in a gel. If the recovery is really high I might be able to see the unlabeled plasmid, especially if I use one of those ethidium alternatives that's much more sensitive. How high could recovery be? Say I have 10^9 cells, each takes up 10 plasmid molecules. That's only 10^10 plasmids; if they're 2kb long that's 2x10^10kb of DNA, which is about 0.02 micrograms. 20 ng can easily be seen in a gel, especially with a fancy dye, but I'd be out of luck if uptake is less or recovery is poor (for example, if the DNA is nicked the kit won't work).
Ideas from lab meeting
- "What is the H. influenzae uptake specificity? A pool of USSs that have been intensively but randomly mutagenized and then selected for the ability to be taken up by competent cells will be sequenced to fully specify the uptake bias."
- What forces act on DNA during uptake? Laser-tweezer analysis of USS-dependent uptake by wild type and mutant cells will reveal the forces acting on the DNA at both the outer and inner membranes.
- Does the USS polarize the direction of uptake? Using magnetic beads to block uptake of either end of a small DNA fragment will show whether DNA uptake is symmetric around the asymmetric USS.
- Does the USS increase DNA flexibility? Cyclization of short USS-containing fragments will reveal whether the USS causes DNA to be intrinsically bent or flexible, and whether ethylation or nicking can replace parts of the USS.
- Which proteins interact with incoming DNA? Cross-linking proteins to DNA tagged with magnetic beads, followed by HPLC-MS, will be used to isolate and identify proteins that directly contact DNA on the cell surface.
- Which proteins determine USS specificity? Heterologous complementation with homologs from the related Actinobacillus pleuropneumoniae (which recognizes a variant USS) will identify the proteins responsible for this specificity.
Before any of this, I or the post-doc should carefully test whether H. influenzae can indeed take up closed circular DNA intact. I've been citing an old report of this for years, but it's critical to my model of DNA uptake so I really need to test it myself. I'll write a separate post about this.
Preliminary experiments for our next grant proposals
Basically it's a lovely proposal with one massive flaw. The experiments are a fine mix of safe and risky, and the proposal is beautifully written and illustrated. BUT, there's no preliminary data to show that the experiments we propose are likely to work. (The reviewers also grumbled that I was being sneaky in submitting proposals to two different committees in the same competition. Such greed is, of course, un-Canadian, but that won't be an issue this time.)
So now it's time to think more strategically (i.e. about strategy) than I usually do. Which of the 6 lines of investigation I propose are easiest to generate preliminary data for? Which are the most severely weakened by their lack of preliminary data?
Here are the Specific Aims from that proposal, with brief notes about preliminary data in italics:
- "What is the H. influenzae uptake specificity? A pool of USSs that have been intensively but randomly mutagenized and then selected for the ability to be taken up by competent cells will be sequenced to fully specify the uptake bias." The new-post-doc is about to start doing preliminary experiments for this and related work, so we should at least be able to show that we can reisolate DNA that has been specifically bound by competent cells.
- "What forces act on DNA during uptake? Laser-tweezer analysis of USS-dependent uptake by wild type and mutant cells will reveal the forces acting on the DNA at both the outer and inner membranes." This work was begun by a physics grad student from the lab with the tweezers (across town). We have all the components ready, and I'm keen to work on this myself.
- "Does the USS polarize the direction of uptake? Using magnetic beads to block uptake of either end of a small DNA fragment will show whether DNA uptake is symmetric around the asymmetric USS." We've tried this unsuccessfully in the past with much bigger beads. The magnetic beads may also be valuable for the post-doc's DNA-reisolation work - I wonder what form preliminary data should take.
- "Does the USS increase DNA flexibility? Cyclization of short USS-containing fragments will reveal whether the USS causes DNA to be intrinsically bent or flexible, and whether ethylation or nicking can replace parts of the USS." We should at least try this cyclization assay - it shouldn't be a big deal.
- "Which proteins interact with incoming DNA? Cross-linking proteins to DNA tagged with magnetic beads, followed by HPLC-MS, will be used to isolate and identify proteins that directly contact DNA on the cell surface." It would be good to at least show that the cross-linking works...
- "Which proteins determine USS specificity? Heterologous complementation with homologs from the related Actinobacillus pleuropneumoniae (which recognizes a variant USS) will identify the proteins responsible for this specificity." I made the plasmids for this work several years ago, but they were not well-tolerated by the host cells. I should pull them out and see if I can get some results.
For the Discussion?
The Lewontin approach to confusion
I'm (at last) trying the approach I learned from Dick Lewontin when I was a beginning post-doc. When you're trying to understand a confusing situation where multiple factors are influencing the outcome, consider the extreme cases rather than trying to think realistically. Create imaginary super-simple situations where most of the variables have been pushed to their extremes (i.e. always in effect or never in effect). This lets you see the effect of each component in isolation much more clearly than if you just held the other variables constant at intermediate values. Then you can go back and apply the new insights to a more complex but more realistic situation.
So I've created versions of our model where the input DNA consists of just end-to-end uptake sequences, where the genome doesn't mutate (the fragments being recombined do have mutations), where the probability of recombination is effectively 1.0 for fragments with perfect uptake sequences and almost zero for fragments with less-than-perfect ones, and where this probability doesn't change during the run. The genome is small (2 kb) to speed up the run and the fragments are only 25 bp, so recombination doesn't bring in a lot of background mutation.
These runs reach equilibria with 45-50 uptake sequences in their 2 kb genomes. If they had one uptake sequence every 25 bp they would have 80 in the genome. Maybe this is as good as it can get - it's certainly way better than what we had previously.
Eliminating the background mutation in the genome makes a surprisingly large difference; a run with it present had only 16 perfect uptake sequences. I wonder if, with this eliminated, the amount of the genome recombined per generation now has much less effect on the final equilibrium? In the runs I've just done, fragments equivalent to 100% of the genome are recombined, replacing about 63% of the genome in each cycle. So I've now run one with half this recombination - it reaches the same equilibrium.
I had previously thought that the reason our runs reached equilibria that were independent of mutation rate was some complex effect of the weird bias-reduction component of our simulations. But in these new 'extreme' runs I had the bias reduction turned off, and I got the same equilibria with mutation rates of 0.01 and 0.001. I also tested a more-extreme versions of our position-weight matrix, where a singly mismatched uptake sequence was 1000-fold less likely to recombine rather than the usual 100-fold less, and this gave the same equilibrium as the usual matrix. So I tried using a weaker matrix, with only a 10-fold bias for each position. This gave about 30 uptake sequences in the 2 kb, a much less dramatic reduction than I had expected. Not surprisingly there were also more mismatched uptake sequences at the equilibrium.
This is excellent progress for two reasons. First, I now have a much better grasp of what factors determine the outcomes of our more realistic simulations. Second, I now know how to generate evolved genomes with a high density of uptake sequences, which we can analyze to determine how the uptake sequences are spaced around the chromosome.
The one big problem
The other variable I could change is how quickly the bias is reduced when fewer than the desired number of fragments have recombined. At present it's 0.75. Increasing it would make the early cycles go slower, but with such a high mutation rate this isn't a big problem. I'll try a couple of runs using 0.9 and 0.95.
Update: I've checked the runs my co-author did before she left. With 100 bp fragments, replacing 50% of the genome in each cycle, she always got less than one perfect US per kb. A run that started with 200 perfects seeded in a 200 kb synthetic sequence ended with 33, not 2000. A run that started with a random sequence ended with 18 perfects. These may not have been the final equilibria, but , in the last half of both runs the numbers of perfects were wandering around between 10 and 100, never more. In these runs the mutation rates for the genomes were very low (the rates for the fragments were 0.001), so the problem isn't background mutational decay of perfect US (the numbers of singly mismatched US at equilibrium were also low, about 100-150).
While I was traveling I did some runs with 500% of the genome replaced - in practice this means that most parts of the genome would have replaced several times before the cycle ended, and only a few bits would have not been replaced at all. These runs had less sensitivity because the genomes were only 20 kb, but they did give higher scores, with about 30-40 perfect US in the 20 kb genome. With this much recombination, the mutation rate of the genome itself becomes irrelevant as the almost the whole mutated genome is replaced by recombination in each cycle.
Should I repeat these runs, and do some with fragments only 50 bp long too? The only reason would be to increase my confidence that these are real equilibrium outcomes. I had terminated my runs when the scores seemed stable, but maybe they would creep up if left long enough (maybe using a seeded genome isn't a very good test of the true approach to equilibrium). The results were the same with mutation rates of 0.001 and 0.01, so I could use a rate of 0.01 and set it to run for 10,000 cycles, which is more than 10 times longer than my other runs went for.
More US-variation progress
US-variation progress
Getting organized
Unexpected perl-run success
Rob had tried to run our program and run into a problem - it wouldn't read the desired settings from the settings input file. This appeared to be due to some idiosyncracy with how our code was being interpreted by their perl compiler. When he tweaked the code to use a slightly different command that did the same thing, it worked. But I wasn't looking forward to having to go through our code and make these changes.
Karl had thoughtfully sent me some instructions when he sent me my password:
The cluster uses the `module` system for loading and unloading the environment variables required for the different applications, quick how to is as followsUnfortunately this was gibberish to me, as I don't know what a 'module' or a 'shell' is, or how to do a standard SGE install. I also couldn't understand the online documentation for remote use, which talked about things I'd never heard of, such as RSA auth KeyFiles and NX CLIENTs and Gnome.
Show the currently installed modules
`module avail`
Load the module for this shell only, not future shells
`module load mauve/2.2.0`
Load the module for all future shells
`module initadd mauve/2.2.0`
Remove the module
`module initrm mauve/2.2.0`
Beyond that pretty standard SGE install.
But wotthehell, I figured I might as well try doing the same things I do to install and run my program on our local WestGrid cluster. So I used Fugu to log on to Rob and Karl's cluster (that worked!) and to copy my program and its two input files into my cluster directory (that worked too!). Then I opened a Mac Terminal window, used ssh to log on to the cluster (that worked!) and told it to run my program. Based on Rob's experience I was expecting a string of error messages, but the program ran fine! In fact it ran about 60% faster than it does on my laptop, which is about five times faster than it runs on WestGrid.
Now I just have to decide what runs need to be done. This will require sorting through everything I did while I was away and what I learned from it, and then consulting with my former-post-doc co-author. Expect a post tomorrow describing this.





