Field of Science

Progress

I'm labeling with biotin-dUTP the restriction-digested MAP7 DNA that I want to attach to my streptavidin-coated beads. I wasn't sure what length range of DNA fragments would be best, so I've cut some with XhoI (mean length ~12 kb, ~4 µ) and some with EcoRI (mean length ~6 kb, ~ 2 µ).

After I've purified the biotin-tagged DNA (must get rid of ALL of the free biotin), I'll test its ability to bind to the beads by mixing some DNA with some beads and running the mix in an agarose gel. DNA that's bound to the beads should stay in the well rather than entering the gel, because the beads will be bigger than the gel mesh size. Well, the beads will be 1 µ or 2 µ in diameter (I'll try both, one with the XhoI-cut DNA and one with the EcoRI-cut DNA), and I'm quite sure that's much bigger than the mesh size of a 0.8% agarose gel. As controls I'll use the same DNA digests without the biotin tagging.

I'm also making more B. subtilis DNA for my transformation tests, because the first prep didn't have much DNA in it. I'll pool both preps and repurify by another round of phenol extraction and spooling.

Tomorrow I'm at SFU (the university across town), talking to their evolution group about why I'm working in a physics lab, and picking the brain of the Physics Dept's colloquium speaker.

B. subtilis progress

Most of the B. subtilis culture and transformation problems have been solved. What appeared to be growth of the trp- and met-requiring strain Y on minimal-glucose agar with no tryptophan or methionine turned out to be tiny crystals (of CaPO4?) that were seeded in the agar by the act of streaking over it. I now realize that I shouldn't have added CaCl2 to the agar at all.

The only problem remaining is that strain Y can grow moderately well on plates without added methionine, because the defined medium recipe says to add a small amount of 'casamino acids' to the agar. Casamin acids is the mix of amino acids obtained by breaking down the milk protein casein with acid; this destroys the tryptophan but the mixture is about 1% methionine by weight. The genotype of strain Y shows that it should only need methionine and tryptophan, and strain W shouldn't need any amino acid at all. My first test of medium without any casamino acids had no growth of either strain, but maybe I messed up. I've now tested using half and one sixth as much casamino acids as the recipe specifies (cells grow fine), and now I'm testing one tenth as much and none again.

I put strain Y through the B. subtilis competence ritual: Grow )shaking vigorously) in the minimal glucose medium plus casamino acids and a bit of yeast extract until growth has been slowing for 90 minutes, then dilute tenfold into more of the same medium with extra magnesium and calcium and shake for another hour, then mix cells with DNA of strain W. I definitely got transformants to Trp+, but the frequency was quite low. (I couldn't tell about transformation to Met+ because of the background growth due to the casamino acids.) The low frequency could be because the cells weren't very competent or because the strain W DNA was very old (it's been in the fridge for somewhere between 10 and 20 years ).

So now I'm going back into the lab to inoculate a culture of strain W so I can make some fresh DNA tomorrow.

The joys (?) of benchwork

So yesterday I streaked out the B. subtilis strain I'm planning to use as a control in my tweezers experiments, and its wildtype parent, first on rich medium (LB) and then (from that, after about 8 hours) on minimal plates with and without the supplements the control strain needs. This morning I found that the parent (call it W) had grown on all the plates, but the control strain (call it Y) had only grown on LB, when it should have grown on the plate with both supplements.

So I wondered whether my supplements (solutions of tryptophan and methionine) might have been harmed by being autoclaved rather than filter-sterilized. So I quickly made up more, this time filtering them, and streaked the right amounts onto the agar of supplemented plates like those that Y hadn't grown on, and then streaked both W and Y on them, from single colonies on the LB plates. Six hours later I could see tiny colonies of both strains on all the new plates. Not just the plate with both supplements, or just on the new-methionine plate, as might be expected if it was the original tryptophan stock that had been ruined by autoclaving. No, strain Y grew on the plate with the new methionine, and on the plate with the new tryptophan, and on the plate with both, as if it was the wildtype strain.

What control did I wish I had done - streaking both Y and W again on the original plate with both old supplements. Maybe I had accidentally streaked strain W in place of strain Y. So I streaked them on that. And then I made some new minimal plates, this time leaving out the trace amount of casamino acids the minimal recipe called for (maybe that had enough tryptophan and methionine???). I spread some of these plates with the new supplements, and streaked W and Y on plates with and without.

Then I wondered if I had somehow taken cells from the wrong vial of my 20-year-old- stocks, so I streaked from the vials again onto LB and old and new minimal plates. Could I have gotten the Y genotype wrong? No, a Google search for its genotype found a page specifying trp and met requirements.

Here's hoping the plates make more sense tomorrow morning.

Experiments begun

I'm starting the laser tweezers work with Bacillus subtilis cells rather than H. influenzae cells, for several reasons. B. subtilis cells are much bigger than H. influenzae cells, and also tougher. I've worked with them before and know how to make them competent. They take up DNA without any sequence specificity, so I can use them with the same bead-attached H. influenzae DNA that I'll use for the H. influenzae cells. And they've previously been used for optical tweezers measurement of DNA uptake. So they're an excellent positive control for the H. influenzae experiments I want to do.

Today I streaked out the wildtype and auxotroph (trp- met-) strains from some 20-year-old slants; they're growing up fine. And I made the defined media needed for competence induction, and poured some plates with and without tryptophan and methione, and streaked both strains on all the kinds of plates to test that I've made them correctly. I found a DNA stock from the wildtype strain in the fridge. It's about 19 years old but I expect it will still be fine. If the transformations don't work well I'll make fresh DNA.

Submitted!

Hooray, the CIHR proposal has been submitted. We think it looks good - we've incorporated just about all of the suggestions of our two internal reviewers, and we didn't have to cut any of our little illustrations to fit it into the 11 pages.

We took a break from proposal-polishing this afternoon to drag our visiting Chinese grad student to the pub to watch the hockey game, telling him that he'd otherwise be kicking himself for missing such an exceptional Canadian cultural experience. And it was!

Now I'll get into the lab, tomorrow morning, to start working on the preliminary techniques for the optical tweezers experiments. Things like figuring out how to get the cells to stick to coverslips without killing them, and how to tell that I haven't killed them.

Getting shorter...

The proposal, not me.

It's down to 11.7 pages. And that's with some additions to clarify previously obscure points, and with all the tiny illustrative figures retained. There's not much deadwood left, but I think there's still lots of sentences we can shorten and maybe even cut.

I've also been going through the figures for the Appendix, making them clearer and adding 'Conclusion' boxes.

Nearly done the CIHR proposal

I have to click 'Submit' on the CIHR proposal by 10:00 am on Monday. The signed forms are already at Research Services, who have managed to get UBC a blanket institutional-signoff extension from CIHR. (I suspect that they claimed that all the signing authorities would have fled the city, either to avoid the predicted Olympics-associated chaos or because they went to Hawaii on the money they made by renting out their house.).

It's looking pretty good now, thanks to valiant work by the postdoc and research associate, except that it's about 50% too long. It will get shorter when the reference citations ({name, year,Endnote #}) are replaced by simple numbers, and when the paragraph spacing is reduced from 6 pt to 3, but that won't be nearly enough to get it down from 17 pages to 11. So today's big challenge is to find and eliminate all the redundancies in the text, and all the irrelevant or dispensable information.

The research associate is off for the weekend - much deserved after some heroic work generating evidence of cross-species complementation of the pilin operons. I'm delighted with this result because I wasn't at all sure the operons could complement, and one of the best parts of the proposal depends on this complementation. I had started to test this several years ago but given up when my plasmids were misbehaving, but she's sorted out the plasmid mess (by throwing mine out and starting fresh) and gotten the result I had been hoping for.

Today and tomorrow the postdoc and I will cut and polish and cut and polish, with a break tomorrow afternoon to watch the hockey game (Canada vs US for Olympic gold!) at a campus pub. I'll also do the space-saving last minute editing that makes a paragraph a line shorter by eliminating one little word or nudging a paragraph margin by a millimeter. Then we'll assemble all the docs into one big pdf and I'll click 'Submit'.

And on Monday I'll get into the lab.

Explaining the gaps

OK, I'we worked out how to organize the first part of the Background section of our CIHR proposal: First a very brief overview of DNA uptake in all bacteria, with three sub-headings: (i) Regulation, (ii) Uptake and translocation, and (iii) Degradation and recombination. (There might also be a little figure showing the four stages, but this might not be needed now there's a detailed figure soon after.) This section serves also to emphasize the generality of the problem - I'm not just proposing to solve a H. influenzae problem.

The next short section is titled Current model of DNA uptake. It presents a new version of the figure I posted yesterday, with the gene names and the nucleotide import removed. It describes what we think happens in gram-negative bacteria, relying on evidence mainly from H. influenzae and the Neisserias. The basic steps are described, but the detailed evidence and the problems aren't pointed out until the Gaps sections.
  1. Initiation occurs internally on dsDNA fragments, preferentially at short sequence motifs called uptake sequences.
  2. The force for DNA uptake is produced by retraction of cylindrical protein multimers called pseudopili, short versions of long extracellular filaments called type IV pili.
  3. Transforming DNA enters the periplasm through outer membrane 'secretin' pores like those through whch type IV pili exit.
  4. DNA may bind nonspecifically to the pseudopilus, as it does to type IV pili, and this may be how the force of pseudopilus retraction is transmitted to the DNA.
  5. Once DNA is in the periplasm, a fragment end interacts with the translocation machinery. One DNA strand is degraded to nucleotides as the other enters the cytoplasm.
Evidence underlying this model comes from several kinds of analyses, in H. influenzae, the Neisserias and other bacteria: the phenotypes of mutants, the fate of radiolabelled or genetically marked DNA in wildtype and mutant cells, protein properties predicted from the sequences of genes induced in competent cells or otherwise implicated in competence, the activities of homologs that contribute to type IV pilus function.

This model is superficially satisfying, but a critical analysis reveals a number of problems. Next I describe the three serious gaps that will be filed by the experiments I propose.

Gap 1. What are the players? (What proteins contribute to DNA uptake? Which of these interact directly with DNA?)

The first paragraph explains that H. influenzae is particularly appropriate for this analysis because (i) we've identified all the members of its competence regulon (can't be done in Neisseria (no regulation), or the Gram-positives (no competence-specific regulation).

The next two paragraphs describe what's known and not known about the H. influenzae competence genes; it includes a little figure showing all the genes clustered under their likely roles. Only four genes can reasonably be excluded from roles in uptake. Four others have no known function, and most of the rest have only suggested functions. Only six are directly implicated in uptake by assays of labeled DNA and phenotypes of non-polar knockout mutations.

Another paragraph describes the need to identify the proteins that bind DNA, and the limited success to date (some evidence of non-specific binding).

This section ends with a paragraph describing the Specific Aim and the first two questions: which proteins contribute to DNA uptake (and translocation), and which proteins contact DNA during uptake and translocation. I think the paragraph should also summarize the strategy.

Gap 2. What is the uptake specificity? How does it act?

The first paragraph will give basic factoids about uptake specificity and uptake sequences. Both H. influenzae and N. gonorrheoeae have long been known to have very strong preferences for DNA from close relatives. We now know that this is due to an uptake bias favouring sequence motifs abundant in these DNAs; our measures show at least a 100-fold bias (published reports are inconsistent). The preferred H. influenzae sequence was initially identified as a 9 bp AAGTGCGGT (N. gonorrhoeae's is XXXXXXXXX). However once genome sequences became available the focus shifted to characterizing the many copies of these sequences in the respective genomes. Once my lab realized that the genome sequences are not replicative elements but motifs (like other protein-binding sites), we shifted to characterizing them as such (see attached manuscript).

Relevance of uptake sequences to the mechanism of uptake: 1. Uptake sequences are likely to result from general features of DNA uptake in Gram-negative bacteria. Even though their sequences are different the genomic H. influenzae and Neisseria uptake sequences share almost all other properties: frequency, spacing, frequent accessory role as transcription terminators, strong consensus (see attached manuscripts). These are the two best studied DNA Gram-negative uptake systems. Their uptake specificities are also shared by other members of the Neisseria, and by all Pasteurellaceae. There is also some evidence of uptake specificity in other systems (Campylobacter, others?), and most have not been examined. The absence of obvious uptake-sequence-like repeats in genomes doesn't mean that the species' DNA uptake machinery has no sequence specificity. The failure to easily find a single protein that binds specifically to uptake sequences (in either H. influenzae or N. meningtidis) suggests that uptake specificity is not a detachable 'add-on' to the mechanism (a kind of pre-screening) but is rather intrinsic to the process.

One big problem with the uptake model presented above is the need for sharp DNA kinking, and that may be resolved by uptake sequences. Double-stranded DNA is quite stiff on the scale needed for uptake, with a persistence length of >50 nm (for comparison, the secretin pore has a diameter of about 6-7 nm ). (Can this be indicated on the figure; is the figure drawn roughly to scale?) The model shows a slight bend at the uptake sequence (structural models predict a ??° bend at each of the two AT-rich segments in the H. influenzae uptake sequence, and we see the predicted gel retardation in a 200 bp model fragment), but clearly a sharp kink of nearly 180° is needed for passage through the pore. We know that uptake does not need to occur at the ends of fragments, because closed circular plasmids are taken up as efficiently as linear fragments. Current understanding of DNA structure is not good enough to predict whether uptake sequences are preferential kink sites , nor what kind of force might be needed to cause it. Because of DNA's high charge, bypassing the pore is not an option. (There's also the problem of fitting the DNA in the pore beside the pseudopilus, but it may be better to not even bring this up as I don't propose to solve it.)

Say that we can't just rely on the genomic uptake sequence motif as a surrogate for the uptake specificity. Evidence in the manuscript that uptake bias only imperfectly corresponds to the genomic motif, also evidence of strain-to-strain variation at uptake sequence positions doesn't correspond.

I propose to carry out a very high resolution analysis of uptake specificity, examining both sequences in the uptake sequence motif and effects of nearby sequences. The high resolution will allow investigation of effects on uptake by interactions between different positions in the motif (like the genomic covariation analysis in the manuscript). We will also investigate the physical properties of the newly defined uptake motif. Finally, we will exploit the variant specificity of the related A. pleuropneumoniae DNA uptake system in a cross-species complementation experiment to identify the genes responsible for their different uptake specificities.

Gap 3. What forces does DNA experience during uptake?

Force must act on DNA to create the kink needed to initiate uptake; force is also needed to continue uptake. The model shown above addresses this by using pseudopilus retraction to pull DNA across the outer membrane, but it overlooks two big problems, the need for a ratchet and the absence of the retraction protein PilT.

In principle, retraction using a type IV pilus mechanism is certainly able to generate a sufficiently strong force; measurements using optical tweezers (where a cell is stuck to a coverslip and its pilus is attached to a bead in the optical laser trap) have recorded forces in excess of 150 pN, the strongest molecular forces known. The details of pilus assembly and retraction are shown in the figure (??? maybe we'll have a figure here). Prepilin subunits in the inner membrane are freed from their leader sequence by a prepilin peptidase (pilD), and assembled into the base of the elongating pilus by the PilB ATPase. Retraction occurs by disassembly of the subunits by the reverse of this reaction, typically catalyzed by the related PilT ATPase; this is where the force is generated. The entire complex is thought to be restrained in the inner membrane by other proteins (name them??).

Although H. influenzae has good homologs of all the other Tfp proteins, it lacks any identifiable homolog of PilT, as do all of its Pasteurellacean relatives. Thus we do not know where the force comes from. (One strain has been shown to be capable of assembling and retracting pili under special conditions, so we know it has all the required genes; to date thee recognized ones are all int the competence regulon.) We will take two approaches to finding the source of the force. First, characterization of all proteins that interact with incoming DNA should identify it, even if it's not one of the known competence proteins. Second, using optical tweezers to characterize the forces acting on DNA during uptake will show whether the force has the properties expected of a Tfp mechanism.

The second problem with the current model of uptake is the need for a ratchet. This problem has been largely overlooked by the Neisseria researchers, perhaps because their main focus has been on the role of Nesseria's long type IV pili in pathogenesis. However, type IV pili have not been detected on H. influenzae cells under normal growth conditions, and even though competent cells dramatically upregulate all of the Tfp genes they do not have detectable pili. Neisseria cell also do not need long pili to take up DNA, as mutants defective in pilus assembly are proficient for uptake. Many other naturally competent Gram-negative bacteria also lack detectable type IV pili, despite possessing the same genes as H. influenzae (though they do have PilT).

Type IV pili can be several µ long, so a single pilus retraction could in principle pull in DNA fragments as long as 10 kb (if one end of the DNA bound to the proxinal end of the pilus) (see 'Not this' figure). But DNA is normally taken up not by pili but by pseudopili, which are thought to only span the distance between the inner and outer membranes (??? 20 nm???), and not to protrude significantly beyond the cell surface. Thus a single pseudopilus retraction can only be expected to pull in about 100 bp of DNA (see 'But this' figure).

One possible solution would be coupling of translocation to uptake, with the pseudopilus only needed to bring some part of the DNA into contact with the translocation machinery. But we know this is not the solution in H. influenzae, because circular plasmids can be fully taken into the periplasm even though they cannot be translocated, and because uptake proceeds normally in a rec2 mutant, which cannot translocate DNA.

The simplest solution (thought by no means simple) is for the pseudopilus to act as a ratchet, alternately elongating and binding DNA and then retracting and releasing the DNA. A detailed drawing of this mechanism is provided in the Appendix. The timescale of DNA uptake (a few minutes?) and the length of each retraction event makes this potentially detectable with optical tweezers.

In addition to their use for characterizing pilus retraction, optical tweezers have been used to measure forces on DNA during uptake by competent cells. This work shows that this is a good way to measure forces but was not informative about ratchet mechanisms because the bacteria used don't face this problem. The first measurements were done with B. subtilis, which needs its pseudopilus only to bring a part of the DNA to the cell surface where it is cut for translocation (mutants lacking the pseudopilus proteins and translocate DNA provided the cell wall barrier is removed). The only other tweezer measurements have been done with Campylobacter, which is (with its relative Helicobacter) the only uptake system that doesn't use Tfp machinery.

(Say more about forces here). (Also say that maybe the need for sequence specificity arises from the complications introduced by the outer membrane and the need for an uptake ratchet??)

Now describe what I propose to do. Need details here.

Other ideas to include somewhere else:

If we detect binding of a particular protein to DNA, we can then test whether mutations in other proteins affect this binding - this might help define the order of events within, e.g. DNA uptake. For example, if we see that pilin does contact DNA, we can test whether a secretin mutation prevents this.

If sequencing costs drop we will repeat the Q. 3 analysis of uptake specificity with A. pleuropneumoniae cells and their USS.

Background section for the CIHR proposal on DNA uptake

I've convinced myself that I need to reorganize the Background section of our upcoming proposal to CIHR (due March 1), but I'm not making much progress on paper so I thought I'd try to outline it here.

The plan is to first give a very brief overview of natural competence, saying what's generally true for all bacteria. This could also give some H. influenzae-specific information but I think it's better kept general.

Then I'll have a drawing of our current model of DNA uptake in gram-negative bacteria (applicable to both H. influenzae and N. meningitidis, and maybe to most Gram-negative bacteria). The figure below is one from a review we wrote - I'll modify it for the grant. I was originally (i.e. yesterday) planning to just briefly point out that this 'model' is really only a static picture of the known and hypothesized players. But I'm beginning to think I should give some description of the mechanisms illustrated the figure, also telling the reader that it's all hypotheses based on limited information (what does and doesn't happen to transforming DNA in wildtype and mutants, what properties proteins are predicted to have based on their sequences, what homologs of the proteins are thought to do in Tfp function).
  1. Initiation of uptake has a strong sequence bias towards a motif called the uptake sequence.
  2. Initiation occurs internally on DNA fragments.
  3. Force for uptake is produced by shortening of a protein multimer called the pseudopilus, which is closely related to long filaments called Type IV pili (more details below).
  4. Transforming DNA enters the periplasm through an outer membrane secretin pore like those through which type IV pili exit the periplasm.
  5. DNA can bind nonspecifically to type IV pili; this may be how the force is transmitted.
  6. Once DNA is in the periplasm, a fragment end interacts with translocation machinery.
  7. One strand is degraded (probably on the periplasm side of the membrane) and the other enters the cytoplasm.
  8. The competence protein Rec2 may form a pore in the membrane.
The model is not informative about many points. We don't know:
  1. What are the ultimate sources of the forces that pull DNA in (ATP? PMF?, other?).
  2. How the pseudopilus is disassembled.
  3. How the DNA fits through the pore?
  4. What role the sequence specificity plays .
  5. What is the full uptake specificity.
  6. Whether sequence specificity only matters at initiation.
  7. Whether uptake and translocation are usually coupled.
  8. What most of the genes in the H. influenzae competence regulon do
  9. What prevents backsliding.
The rest of the Background is headed by the three gaps I propose to fill.

Gap 1: Who (what?) are the players?
This section will describe what we know and don't know about the proteins that might contribute to DNA uptake. I need to also say what other proteins might do (process DNA in the cytoplasm)? End with an overview of what we'll do in Aim I.

Gap 2: What is the uptake specificity?
This section will describe why I think the uptake specificity is an important component of the mechanism (i.e why I don't think it's just a Haemophilus-specific artefact). Emphasize that properties (genomic and uptake) are shared by Neisseria (just not the actual sequence itself) and because these are the two best studied systems we have to take them as exemplars. Then I can describe what we know: crude uptake assays, detailed genomic analysis, two types in the pasteurellaceae, and give an overview of what we'll do in Aim II.

Gap 3: What are the forces?
This section will describe the unknowns about what forces act on the DNA, in the contexts of the B. subtilis and Helicobacter analyses. Here I'll describe the absence of PilT, and the apparent need for a ratchet mechanism. Also the backsliding problem. And the need for an uptake force that is independent of translocation.

Proposed dynamic model
The Background will end with my dynamic ratchet-based model of uptake.

Two submissions down

The NIH proposal got submitted (it just needed a 9-digit zip code), and yesterday the former post-doc submitted our manuscript on uptake sequence variation. It looks pretty good so we decided to submit to Genome Biology, even though BioMed Central is not my favourite journal publisher.

Now I'm back to working on the CIHR proposal about the mechanism of DNA uptake. I have comments from the two internal reviewers. One thought it was pretty good proposal, and had lots of comments about ways to increase the clarity and improve the explanations. The other thought the global organization was quite bad - this was depressing but his suggestions were very good so we're doing a major reordering of the material.

Yesterday the postdoc and RA gave me their revisions to our Encyclopedia of Genetics entry on transformation. I did a bit of polishing, and now it's just about ready to go. They're working at the bench, along with our visiting grad student from China, but I think I'll have to wait until the CIHR grant is in to get to my bench.

NIH's new forms won't accept zip codes!

OK, so I dotted every 't' and crossed every 'i' for the new NIH RO1 application forms. Everything done perfectly, to assemble the 16 form pages and dozen of so attachments into the multi-component Adobe-format application. But one ridiculous problem is preventing us (i.e. the UBC Research Services grants administrator) from submitting the application to NIH.

THE FORM DEMANDS BUT WON'T ACCEPT ZIP CODES!

My address, of course, doesn't have a zip code - I chose Country = Canada and entered the post code in the field above the zip code field, no problem. But my consultant is in the US, and I need to enter his address on the Key/Senior Personnel page. So I chose Country = United States and entered his zip code in the zip code field. But the text turned red and a pop-up window told me this isn't a valid zip code. I confirmed the zip code on his letter of support and on his university's web site, I tried other valid zip codes, no success. I tried to delete the Key/Senior Personnel section for the consultant, but that box is greyed out (apparently I can't delete it until I successfully complete it...). The 'Check application for errors' function told me that the application is perfect except for the invalid zip code.

So I took the all-but-complete form (on a memory stick) over to the grants administrator, hoping she could either submit the application as-is, or solve the problem. But the problem was the same on her computer, which told me that it's not a Mac-specific problem. She was too busy dealing with other people's messed-up NIH applications to contact NIH (or Grants.gov) for help.

Back to my office to try a few more things. Downloading a fresh copy of the application form took ages (mostly finding the right web page), but the new one had the same problem. In a way that's a relief, as I really don't want to have to redo the whole application on a fresh copy. As a test I tried changing my entry in the PI section, telling it that I was located in the United States, and giving myself a zip code. Same problem, so now I know it's a general problem with zip codes, not just with that one field.

While unsuccessfully searching Grants.gov for a technical support address or phone number, I found a tester version of the application form, provided to allow applicants to check that they had the right version of Adobe Reader. So I downloaded that form, and had no problem entering a zip code into it. So that tells me the problem isn't with my version of Adobe Reader.

I've tried Googling various combinations of terms (NIH, RO1, SF 424, "invalid zip code") but haven't found any evidence that other people are having this problem. The only solution I can think of now is to tell the form that the consultant is located in Canada (Charlottesville Virginia Canada?) and enter a fake post code.

Later: I was going to tell NIH that the University of Virginia is in the city of "Charlottesville VA USA 22908", in the country United Arab Emirates, but in the meantime the grants administrator discovered that the problem was simply a requirement for a 9-digit zip code!

Now I've also fixed the errors that were identified post-submission, which existed because I didn't realize that the 'Months' section of the Budget for salaries was to indicate the amount of their 'effort' each person would put into the project.

Last paragraph of the NIH proposal

I need to get it written in the next half hour, but my brain is jammed. Points it should make:
  • We're the best people to do this work. We have a unique combination of wonderful attributes.
  • The components of the work are well balanced. None is excessively risky, and later work is not dependent on the success of (or a particular outcome of) earlier work. Our preliminary results confirm that the basic strategy is robust.
  • The approach is cost-effective. Using genome sequencing to get answers about recombination is much cheaper than doing it with molecular biology, because of the breadth of information the sequences provide.
  • The results will give insights into the molecular mechanism of recombination.
  • The work is testing hidden assumptions about recombination. Over the past 60 years, studies of bacterial genetics have been forced to make assumptions about recombination events. These assumptions were reasonable, given the information available, but now we can finally test them.
  • The strains we will have sequenced are a resource for mapping clinically important phenotypes. They also provide a gold-standard control dataset for phylogenetic and epidemiological studies that must detect recombination.

Eureka!

Now I see how to give the Innovation section a narrative, and accomplish other things too. The key was to see it as a way to reinforce the rest of the proposal.

The proposal begins with a Specific Aims page, which provides a summary of what I propose to do and why. This is followed by several pages of Significance, which provide a more detailed explanation of what the problems are that this work will help address. After the Innovation section comes the Approach, where I spell out each Specific Aim in detail, emphasizing how it will be accomplished.

In an old-style proposal, the Specific Aims page would be called the Summary, the Significance section would be Background, the Approach would be Methods, and the place now occupied by Innovation would be where the Specific Aims were listed, connecting the problems raised in the Background with the solutions described in the Methods.

There's no reason that the Innovation section can't also accomplish what a traditional Specific Aims section accomplished. So mine will begin with "I am proposing three Specific Aims, each innovative in both concept and strategy." Then I will have a paragraph for each Aim in turn, explaining how our approach differs from previous approaches and why it is the best solution to its problem. In doing this I will also be giving the reader an overview of the Aims both in the context of the broad Significance they've just read and in the context of the detailed Approach sections they're about to read.

Are we innovative yet?

NIH wants its applicants to use somewhere between half a page and a full page of the 12-page proposal to explain how their proposals are innovative. Here's NIH's instructions:
  • Explain how the application challenges and seeks to shift current research or clinical practice paradigms.
  • Describe any novel theoretical concepts, approaches, methodologies, instrumentation, or intervention(s) to be developed or used and any advantage gained.
  • Explain any refinements, improvements, or new applications.
Other advice, from Jeffrey Benovic and Bruce Freeman):
  • Significance is why the work is important to do.
  • Innovation is why the work is different from (better than) what has been done before.
  • Definition of innovation: a new device or process resulting from study & experimentation; the act of introducing something new.
  • How will research in your field change as a result of your work?
  • Demonstrate the potential gains are not merely incremental.
  • Explain why concepts & methods are novel to one field or novel in a broad sense (or both).
  • Summarize (sans detailed data) novel findings to be presented as preliminary results in Approach
  • Focus on innovation in study design & outcomes
Morgan Giddings also emphasized describing specific ways the field will be different if our work is funded and successful.

The field is recombination (in its broad sense, everything from the molecular mechanisms to the evolutionary consequences), and we're going to fully characterize recombination between related strains. That is, we're going to find out (i) the properties of the recombination tracts and (ii) the probabilities of recombination for ~all differences between the strains. This will be done using deep sequencing of single and pooled recombinant genmes, and will give an extremely comprehensive picture of recombination across the ~40,000 snps and 300 indels and rearrangements that distinguish these two strains. Then we're going to use recombination to map genes responsible for the very low transformation frequency of one of the strains.

I've been making lists of things that are innovative about what we propose, but what's lacking is a coherent narrative that ties them together. So I'm going to just start listing them here and see if a narrative comes together...
  • This will be the first time that all of the recombination of a single recombinant genome (from a single transformation event) has been identified for any organism. And we're going to do it many times (50? 100?). Maybe include actual numbers - how many recombination tracts do we expect to characterize? how many breakpoints? Relate to how many actual recombination tracts have already been characterized (I'd have to dig into the literature...)?
  • By using deep sequencing to measure the frequency of recombination at tens of thousands of SNPs and indels we will characterize the full spectrum of sequence factors affecting recombination. The scale will be unprecedented. How many SNP-conversion events do we expect to detect? How many indel-conversion events?
  • We are breaking down the previously necessary tradeoff between high resolution and broad scope. Deep sequencing of tine genomes lets us have it all!
  • We will develop new analytical tools, to analyze recombinant genomes.
Paragraph about Aim III:
  • Nobody else is focusing on the causes of the poor transformability of many strains of 'transformable' species.
  • Will use genome sequencing of recombinants with altered transformability to identify recombination tracts carrying the responsible alleles. This may seem wasteful but is very efficient; one lane of sequencing (even without multiplexing) is likely to define a stretch of no more than 1 kb. (If the difference is a snp and not an indel - we should have done the phenotyping! This is a pitfall we need to write about.)
  • This will (we think) be the first time that the QTL mapping methods used for eukaryote genomes are applied to bacteria. (Is this right? We need to clarify the relationship between QTL mapping and sequencing. Is anyone sequencing recombinant genomes of yeast?) They were developed out of necessity for the very large eukaryote genomes, but, now that sequencing bacterial genomes is so cheap, are very efficient when applied to bacteria that lack the sophisticated genetic tools available for E. coli and B. subtilis. We're one of the first to start using deep sequencing to replace conventional molecular biology analysis (well, 'one of' is a cop-out as I don't really know...).
Maybe the narrative could reflect/reinforce the narrative of the Specific Aims and Significance sections: "Haemophilus is bad, recombination is devaluing our only weapons (vaccines and antibiotics), but finding out the ground truth about between-strain recombination can let us better predict and prevent it." Or at least emphasize that we're applying this innovation to a pathogen

about PilT

One issue we need to deal with better in our revised CIHR proposal is the identity of the H. influenzae protein that retracts its type 4 pili and/or pseudopili (short stb pili).

In the well characterized bacteria (Neisseria meningitidis, Pseudomonas aeruginosa), pilus retraction is done by the protein PilT, using energy it gets by hydrolyzing ATP. I'm just going to summarize the things I think are true, but once I've done that I'll need to read the latest papers to find the evidence for my statements, and to check what I may have gotten wrong.

The problem is that we (and others) haven't been able to identify a PilT homolog in H. influenzae, although everything we know about DNA uptake in other systems, and about the need for other proteins of the type 4 pilus system in H. influenzae, predicts that a PilT homolog should be needed to pull the DNA in (by pulling the pilus or pseudopilus in). The competence regulon (CRP-S regulon includes all of the other proteins with recognizable T4P-family signal sequences ('prepilin protease-dependent leader sequences'), but none of these are good homologs of the PilT proteins identified in other bacteria (nor of the related PilU). Nor are there recognizable PilT or PilU homologs among proteins that don't have this leader sequence.

The closest H. influenzae relative of PilT in other systems is a protein assigned as the PilB homolog. In other bacteria PilB is essential for assembly of the T4P, and we know that H. influenzae PilB is essential for DNA uptake. I'm pretty sure that PilB can't also do the job of PilT, because both proteins are ATPases. That is, in the bacteria where its function has been studied (mainly N. meningitidis and P. aeruginosa) PilB uses energy from hydrolyzing ATP to assemble pilin subunits into a pilus fiber, and PilT uses energy from hydrolyzing ATP to disassemble the pilus fiber into its subunits. From the perspective of the pilus, the PilT reaction is a reversal of the PilB reaction, but from the perspective of the ATP these are very different reactions - the energy requirement tells us that PilB will not be able to carry out pilus disassembly.

But some protein must do the work of pulling in the DNA. One possibility is that H. influenzae has a cryptic PilT homolog - maybe an ATPase that gets to the right place in the inner membrane without having a recognizable T4P targeting sequence. Another is that this function is done by an unrelated protein. I'd expect such a protein to be an ATPase,

(Here's a link to a couple of movies of P. aeruginosa cells whose pili have been made visible with fluorescent antibody. In one you can see the pili (~5 times longer than the cell) shortening, and in the other you can see the tip of an elongated pilus attaching to the slide surface at a point distant from the cell, and then shortening, pulling the cell to the attachment point.)

I can't think of any way to select or screen for a defect in pilus disassembly. This is partly because pili have not been detected on strain Rd - the group that showed pili only did this work in the clinical strain 86028-NP. That strain does detectable twitching motility under the alkaline conditions where it does produce visible pili, but it also doesn't have a PilT homolog. In fact (I think), none of the Pasteurellaceae have them. And I think that Pasteurellacean cells typically don't have type 4 pili at all. We do have a pilus-associated phenotype that we can screen for defects in - DNA uptake - and we expect a PilT mutant to be defective for this. But we have already identified lots of genes whose knockouts prevent DNA uptake and thus would be found in such a screen, so this is a lousy way to look for PilT mutants.

But let's think about this a bit more. Say there are about 25 genes needed for transformation. In principle we can do random knockouts using some well-behaved in vitro method and use transformation to put these into the chromosome (just like Gerry Barcak and Hanna Tomb did in Ham Smith's lab 20 years ago). Then we'd do a massive screen for strains that don't transform, and screen the clean nontransformers for DNA uptake, then check where the knockout was in each strain that didn't do uptake. Then we'd check out any new genes. (This strategy is looking more and more lousy with each additional screening step. If I had unlimited money I might do this, but I don't think I'd fund it over competing projects.)

An alternative plan is to start by screening the genome for proteins with ATPase motifs (Walker boxes) and then pick out the ones that have signals to target them to the inner membrane and don't have any other assigned function. If there are only a few candidates this would be reasonable, but there might be very many.

Another plan is to wait for someone else to solve the problem. Another bacterium lacking PilT is also able to retract type 4 pili - I thought it was Myxococcus, but there's a 2003 paper describing Myxococcus pilT gene - Oh, it's that Myxococcus pilT mutants can still retract their pili, though not as well as wildtype cells.

Articulating why us, why H. influenzae

The research associate has been going through the reviewers' comments on our unsuccessful CIHR research proposal. Much of the criticism was along the lines of 'Why should you do this when it's already been done in Neisseria*?", "Why should you do this in H. influenzae instead of Neisseria*?" and "Why should anyone try to do this, when other scientists have been unsuccessful?"

So yesterday she took the devil's advocate position, pushing me to defend our plans. One argument we will make is that we're in a much better position than the Neisseria researchers to use uptake sequences as a tool to study the uptake mechanism. For example, we have a resource that they lack - the availability of related species with variant uptake specificity. We also have done more investigation (in H. influenzae) into the details of the uptake specificity, measuring uptake with a series of uptake sequences altered at single or double positions. And Aim I of our proposal will give us an immensely detailed characterization of how strongly every base at every uptake sequence position contributes to uptake.

Another defense is that we have thought more deeply about how uptake could work than others have. Nobody else has considered that it's not enough to have a type four pilus or pseudopilus pull on the DNA (a ratchet is needed). Furthermore, nobody else has recognized the significance of the ability to take up circular molecules intact.

These issues are better addressed in H. influenzae. The need for a ratchet is not obvious in Neisseria, because it has long external pili. But most competent bacteria appear to pull DNA to the inner/cytoplasmic membrane with short stubby structures (pseudopili). Neisseria doesn't need long pili either, but this is only seen in mutants, whereas in H. influenzae we're studying the natural mechanism, not an aberration. We know that H. influenzae can take up circular DNAs that remain supercoiled in the periplasm, but we don't know that for Neisseria.

We will downplay testing whether competent H. influenzae have external pili. None are visible in the few published electron micrographs of competent cells, but nobody has ever specifically looked for them. We will begin with the reasonable assumption that competent cells lack external pili but will check this assumption using an anti-pilin antibody.

We may also downplay the search for the proteins that bind to DNA. It's intellectually messy work and we don't have any experience with mass-spec.

* The reviewers' emphasis on Neisseria made us wonder if one or both of them might have a Neisseria background. But I just checked the member lists for the previous two versions of this review committee, and none of them have any obvious connections to Neisseria or natural competence or pili. (These are previous committees, the membership list for our actual committee won't be released for months.) But whoever our reviewers were, they were surprisingly knowledgable!

If I'm on sabbatical why can't I get into the lab?

Well, the NIH proposal needs to get to Research Services in three weeks. I've received lots of very helpful advice, but it's mostly big-picture stuff that will take lots of work to implement. Lots of digging up papers (and reading them), lots of thinking about how stuff ties together, lots of trying to craft plausible descriptions of the significance for human health. Not to mention lots of struggling to understand what the post-doc is finding in all the wonderful sequencing data.

(I just realized that I should take advantage of the NIH database of funded proposals - I can search the list for 'recombination' + 'bacteria' to see the kinds of things people write. I also think I might focus the significance a bit more on H. influenzae, rather than just emphasizing the generality of the need to better understand recombination.)

The failed CIHR proposal also needs to be rewritten in the next few weeks, so we can get a draft to internal review a month before the due date. The research associate has offered to go through the reviewer's concerns, annotating the draft at each point that needs work. I'm afraid that this will be much of the proposal. She's working hard to get more preliminary data on our cross-species complementation plan. The concern is that her preliminary data could show that this isn't going to work as easily as we are hoping, so this is risky. Luckily the reviewers weren't concerned about the lack of preliminary data for the optical tweezers section - instead they were concerned about it's significance.

On Monday afternoon I'm giving a talk about DNA uptake to the biophysics group at the university across town (henceforth SFU) where I'll be doing my optical tweezers sabbatical project. This will be a chalk talk because I want to keep it informal and get lots of interaction with the audience, and because I don't want to take the time to prepare polished slides. Luckily the room's walls are covered with chalkboards, and the organizer has promised coloured chalk.

I'll prepare for this talk by rereading the CIHR proposal (killing two birds with one stone).

Bad news from CIHR

Yesterday I got the review and scores for our DNA uptake grant proposal - not nearly as good as I'd hoped. It only made the 51st percentile, which means it won't get funded.

The reviewers said some good things. They thought it was very well written (as it was). They liked the mix of risky and safe projects.

Some of the reviewers' concerns were very reasonable: the mass- spec bit lacked details that would make it credible (we've no experience); the basic cross-species complementation should have been tested; the search for a pilT homolog should have been completed; why continue to look for pili when they're not seen in EM. They also want more justification of the personnel budget. These problems can easily be addressed in the resubmission (due March 1).

Others mainly asked for better explanation of significance. Why study DNA uptake in H. influenzae when we already have information about uptake by Neisseria and Bacillus? What will we learn from the optical tweezers experiments? I thought we had done a really good job of explaining this, but I guess I thought wrong.

A bigger concern was the reviewers' general lack of enthusiasm, which I suspect may also account for several unfounded criticisms. One reviewer wondered why we were going to look at binding of Neisseria pili to DNA, when we had clearly stated that this was just a positive control. Another said that the optical tweezers experiments had already been done for Neisseria and Myxococcus, when in fact these had only looked at pilus retraction, not DNA uptake. If we had managed to create more enthusiasm, maybe the reviewers would have more carefully checked whether these criticisms were valid.

CIHR includes a 'community reviewer' for each proposal. These are interested members of the general public - they usually read only the lay summary. The community reviewer of the previous submission (3 years ago) complained that the lay summary hadn't included the name of the organism, so this time I made a point of naming Haemophilus influenzae, clearly explaining that it was a bacterium that causes respiratory diseases. But the reviewer of this proposal nevertheless complains that the name misled them into thinking this proposal was about influenza, and recommends that we don't give any species names!

We're going to try to get a revised version done by the end of January, so we can take advantage of UBC's internal review option.

The US-variation manuscript finally is coming together

(Yes, I know I've said this before...)

Abstract

Uptake signal sequences are DNA motifs that promote DNA uptake by competent bacteria in the family Pasteurellaceae and the genus Neisseria. The genomes of these bacteria contain many copies of their canonical uptake sequence (often >100-fold overrepresentation), causing the uptake machinery to prefer DNA derived from close relatives over DNAs from other sources. However the molecular and evolutionary forces responsible for both the uptake bias and the abundance of uptake sequences in these genomes are not well understood. Here we thoroughly evaluate the simplest explanation, that uptake sequences accumulate in genomes by a form of molecular drive, generated by biased DNA uptake and genetically neutral recombination. A computer simulation model shows that these simple assumptions are sufficient to drive uptake sequences to high densities, with the spacings, stabilities and strong consensuses typical of real uptake sequences. In the absence of strong evidence of selection for a recombination function, it may thus be more parsimonious to treat uptake sequences as an epiphenomenon of biased DNA uptake rather than as evidence for a sexual function of natural competence.


Resolving the hotspot paradox, at least partly)

A while back I described our previous work on the paradoxical activity of meiotic recombination hotspots (their mode of action is self-destructive). A new paper by Simon Myers and coauthors (Drive Against Hotspot Motifs in Primates Implicates the PRDM9 Gene in Meiotic Recombination.) now goes a long way towards resolving the paradox, though it doesn't explain how our recombination system got itself into this mess.

This group had previously identified a 13-nt sequence motif typical of human hotspots (thought not all of them); it's thought to be the sequence motif recognized by the process that initiates recombination by creating a double-strand break in the hotspot DNA. Previous work had suggested that chimpanzee hotspots are in different places than human hotspots, so the authors looked at the chimpanzee homologs of the human hotspots and found that although the sites did have this sequence (or variants of it), they didn't function as hotspots. The chimpanzee genome also had this motif at other sites, more than the human genome does. They concluded that many of the human occurrences of the motif had been lost from the human genome because their hotspot activity was self-destructive. They hypothesized that the motif was not ancestrally a hotspot but had become one in the human lineage, 1-2 million years ago.

They then decided to look for the protein that recognizes these sites. Based on their earlier work they had already hypothesized that it would be a zinc-finger protein with ≥12 'fingers' to bind the motif. So they used structural predictions to examine candidate zinc-finger proteins encoded by the human genome, and found five candidates, of which the best was a protein called PRDM9.

If PRDM9 is indeed the protein that, in humans, binds the 13 nt hotspot motif to initiate recombination, its human version should recognize this motif but its chimpanzee version should not. Consistent with this, PRDM9 was the only one of the five candidates that was different in chimpanzees (the other four had identical sequences in both species). Furthermore, it's sequence didn't just have random differences, but had many of its differences in the zinc-fingers that recognize DNA sequence,with hallmarks of positive selection for the changes. And independent work on this protein in mice and genetic mapping, both implicate it as playing a role in the initiation of recombination.

So what are the implications for the hotspot paradox? My simple view is that active hotspots do self-destruct over evolutionary time, as we predicted, and because we need recombination to hold chromosomes together in meiosis, this creates selection on the protein that recognizes them. Variant proteins that recognize new sequences are favoured (that's the positive selection) because they can cause more recombination and thus better prevent chromosome errors. So over long evolutionary periods, genomes may progress through a series of different hotspot motifs and locations.

Researchers have sometimes proposed that different kinds of genes evolve best with different amounts of recombination, and that the chromosomal locations of genes and/or hotspots have evolved to that optimize the amount of recombination. This new paper throws cold water on that idea.

Added later: Turns out the Myers paper was published together with two other papers about PRDM9's role in recombination, confirming and extending the conclusions. And, the same gene was identified last year as a locus very important in speciation, so maybe changing your hotspot specificity changes who you can reproduce with.