Field of Science

Reading frames vs USSs

Sorry for the dead air; I was away unexpectedly with no internet access.

We got the reviews back for our USS manuscript. Not too bad. Both reviewers asked for more analysis of the role of reading frames in USS locations. (See previous posts here and here.) This turned out to be both easy and fun to do.

Many USSs are in the protein-coding parts of the genome, and these can be sorted by which of the 6 possible reading frames the respective proteins are encoded. The first two figures show the relationships of the USSs to the reading frames.

The frames aren't equally used. Frames A, B and C have 49, 125 and 425 USSs respectively, and frames D, E and F have 474, 125 and 157. The differences are too large to be explained by chance alone, and we think they arise because USS in the less-used reading frames impose more severe constraints on protein function, so that many USSs arising in these frames are eliminated by natural selection.

The new analysis considers the two factors likely to contribute to this disparity. (Because the flanking segments exert only modest constraints on amino acid sequence, I've limited the new analysis to only the most frequent tripeptides encoded by the USS core in each frame.)

The first factor likely to affect USS reading frame usage is the differing biochemical properties of the tripeptides that USS cores in these reading frames will encode. Some combinations are intrinsically more versatile than others, useful at many different locations in a wide variety of proteins, whereas other tripeptide combinations will only be useful in particular contexts.

The second factor is 'codon bias'. Most amino acids are specifiable by at least 2 (often 4 and sometimes 6) different codons, some of which are more efficiently translated than others, and cells preferentially use the easiest-to-translate codons for proteins they need to make a lot of.

The new analysis evaluates the first factor by comparing the number of USSs in each frame to the total usage of the six tripeptides in all the proteins of the genome (the 'proteome'). If the differing versatilities of the tripeptides is responsible for their differing use in USSs, we should see a correlation between total number and USS-encoded number. The results are shown by the red symbols and line in the graph below. The predicted correlation exists. (The line fitted to the points has a confidence score of 0.86. I know enough statistics to know that this is a reasonably strong correlation.)

The blue symbols in the graph show the results of the other analysis. To look for a correlation between codon bias and USS reading frame usage, I needed a crude score indicating how easily each USS-encoded tripeptide would be translated. I was able to get a table of codon usage for the H. influenzae proteome from the TIGR website; this gave the percent usage of each codon. So for each tripeptide I calculated a 'USS-codons' score as the sum of the percentages of the codons specified by USSs in that reading frame, and a 'best-codons' score as the sum of the percentages of the most commonly used codons for its three amino acids. Then I calculated a 'codon cost' as the ratios of these scores, and multiplied it by 100 so it would fit neatly on the graph.

If codon cost contributes to the disparity in reading frame usage by USS, we expect the blue points to show an inverse relationship; the highest codon costs should be for the least-used reading frames, so the blue line should slope down to the right. Instead we see no correlation at all. This tells us that codon bias makes little or no contribution to persistence of USSs in different reading frames.

We're not surprised by this second result. (In fact one of the postdocs was so sure of it that she didn't think the analysis was worth doing.) Other analysis we've done has indicated that USS are rarely found in categories of genes known to be subject to strong codon bias. But that's for a different paper, and this new analysis will nicely address concerns raised by both reviewers of our present paper.

Next?

OK, so the induction of CRP-S genes experiments established clearly that the ppdA gene's CRP-S promoter isn't being induced by the growth conditions I tried. And that changes in beta-gal activity don't necessarily result from changes in promoter activity. What to do now?

To make progress in these experiments, I need better tools for manipulating genes. I want to put the ppdD::lacZ fusion into the chromosome (it's on a plasmid now) and I want to test the effect of knocking out various genes, and I want to try a 'wild' E. coli strain. All of these require recombining desired genes into the chromosome, so I need to get the recombineering technique working for me. One of the post-docs was working on this, so I just need to take up where she left off.

I also need to get back to the tests I was setting up of components of the laser-tweezers project. I was just about ready to test attaching DNA to beads (have the beads with the bound streptavidin, have the biotinylated nucleotides to put on the ends of the DNA, have the DNA prep). Separately, I need to work with the new post-doc on an antibody-based method to attach H. influenzae cells to other beads, so they can easily be pushed around under the microscope.

Not the result I was hoping for....

Italic
This is beta-galactosidase activity in cultures that have been transferred from LB (= rich medium) to M9 salts plus a tiny bit of casamino acids (starvation medium). Transfer was at time=0.

All cells carried the ppdA::lacZ fusion on the same plasmid. The black line is cells that are sxy+ and crp+. The blue line cells are sxy-, and the red line cells are crp-.

It's clear that beta-gal activity increased similarly in all three cultures. As seen for the cultures in rich medium (previous post), the crp- cells had slightly lower activity, while the sxy+ and sxy- cells had almost identical activity.

So I conclude that the treatments I've tried so far have not changed the Sxy-dependent activity of the ppdA promoter.

Disappointing result

The detailed time course of beta-galactosidase activity in the E. coli strain with the ppdA::lacZ fusion confirmed the results I posted a few days ago. But I'm not going to post these results, because they are superseded (made uninteresting?) by the results of the next experiment.

Having shown that beta-gal activity decreases as cells enter exponential growth, and increases once growth stops, I needed to find out whether these changes depended on the presence of the transcription activators Sxy and CRP. If the answer is yes, then they are due to changes in the activity of the ppdA gene's CRP-S promoter. If the answer is no, then the beta-gal changes are due to something much less interesting (from my present perspective of wanting to find out how E. coli CRP-S promoters are activated). Unfortunately the answer seems to be NO.

Here's the data. The top graph shows culture density as a function of time for the three strains I tested. The first points (t=0) are before the cells were diluted 300-fold. You can see the 'lag' for the first ~30 minutes, then the cells begin growing exponentially. (You can tell that they're doubling at a constant rate because the points fall on a straight line on this log scale.) After about 200 minutes growth slows. The lines aren't joined to the last points because there's a 700 minute gap separating them (this part of the graph isn't to scale). You can also see that one strain grows slower than the others (red line and points); this is the strain whose crp gene is knocked out.

The second graph shows the amount of beta-gal activity in the cultures. Some of the values for the crp- strain (again the red line and points) are a bit low, perhaps because of its slow growth. However the sxy+ and sxy- strains show almost identical patterns (black and blue lines respectively). This means that the changes in beta-gal activity do not depend on Sxy, and thus almost certainly do not reflect changing activity of the ppdA gene's CRP-S promoter. Again the last points are for samples taken after 1200 minutes, when the cultures had been in stationary phase for quite a while.

What could cause the changes in beta-gal activity? One possibility is that the number of copies of the ppdA::lacZ plasmid per cell might continue to increase after cell growth slows. Another is that the lacZ mRNA might be more stable in stationary phase than other mRNAs. There are probably other possibilities too. The only important possibility is that we're wrong about Sxy and CRP being transcriptional activators of this promoter.

What next? I'm right now testing whether the stronger induction seen when the crp+ sxy+ cells were transferred to starvation conditions depends on CRP and Sxy. I hope it does.

Why are we still using an archaic procedure to make H. influenzae competent?

Last week I did an experiment in parallel with one of the post-docs. Her preparations of competent H. influenzae had not been very competent at all, and we wanted to check if there was a problem with her procedure. We both produced cells that were reasonably (not excellently) competent, so we concluded that her past problems were not due to faults in her procedure.

In recent posts I've been describing my preliminary steps to finding a way to induce E. coli's 'cryptic' competence genes, and this and the above have started me thinking about the procedure we use for H. influenzae. It's a simple sprocedure: cells growing in a rich medium called sBHI are abruptly transferred to a starvation medium called M-IV. After about 90-100 minutes in this medium the cells are as competent as they are going to get. I deally this means that all ofthe cells are competent, each ready to take up a few hundred kb of chromosomal DNA.

This procedure was developed about 45 years ago, by more-or-less systematic fiddling with growth conditions. It hasn't been altered, simply because nobody has had a strong reason to bother trying. But now we know much more about what is going on when competence genes are induced, and we should be able to improve on this procedure. Improvements would certainly be useful, especially for the post-doc's ongoing survey of competence in 'wild' strains of H. influenzae. Being able to improve the induction procedure is also a good test of whether we really do understand how the genes are regulated.

What should we try? Adding cAMP should take the place of carbohydrate starvation, activating CRP to make its contribution to expression of CRP-S promoters. If we're right that nucleotide starvation is needed to induce sxy translation, we would like to try simpler ways to mimic this; I wonder if there's a specific inhibitor pf purine (or pyrimidine) synthesis? But this won't be much use if the cells are still in rich medium, as they are getting their nucleotide precursors from the medium not from de novo synthesis. Maybe we could simulate nucleotide starvation another way (slowing polymerase with an antibiotic resistance mutation? as we suggested in one of our grant proposals???).

Done and to do

Yesterday I did a big time-course experiment with the ppdA::lacZ fusion strain, to get a more detailed picture of the results in my last post. The information is now in my notebook: 50+ samples, each with sampling time and OD600 (cell density at that time) and volume assayed and assay time and OD420 (ONPG hydrolyzed by beta-galactosidase), but I haven't yet entered it into Excel and analyzed it.

The next important experiment is to demonstrate that the changes in expression are dependent on both CRP and Sxy. I will do this by doing simplified time-course analyses using cells carrying the ppdA fusion plasmid and a knockout of either crp or sxy. I grew up both these strains yesterday and hope to have time to test them this afternoon. If the changes are indeed due to increased activity of the ppdA CRP-S promoter, they shouldn't happen in the knockout backgrounds. I won't do this experiment until after I have analyzed yesterday's data, so I'll know which parts of the time course are most informative.

I didn't have time yesterday to repeat the transfer-to-starvation-medium analysis; I'll wait till I have the knockout results before doing this. That way I'll be sure that the expression changesI'm seeing are indeed due to the cause I want to investagate (changes in activity of CRP-S promoters).

Biocurious about tweezers

In between time courses I've been reading a M.Sc. thesis by the student who was developing the laser-tweezers analysis of DNA uptake. (You may know him as PhilipJ of Biocurious.)

I'm not an examiner of this thesis; I'm reading it to learn more about how laser-tweezers work. It's very well written, and pitched at exactly the right level for a neophyte like me. Any day now I should decide I've taken the time courses far enough for now, and get back to my attempt to follow in Philip's footsteps.

Time course of ppdA expression

On Saturday I did a time course examining whether the expression of the ppdA::lacZ fusion (and I hope thus of the sxy gene) changed as growth conditions changed while the cells were growing in rich medium (LB) and after transfer to a starvation medium.

The green line shows how the cells grew. First the density fell sharply because I diluted the cells 1/250 in fresh LB (+Amp to maintain the plasmid carrying the fusion). After a brief period of slow growth while the cells adjusted their metabolism to the improved conditions ("lag" phase), the cells grew exponentially, doubling about every 30 minutes ("log" phase). After about 220 minutes growth slowed down as the medium became depleted of nutrients (still growing so not yet in "stationary" phase).

The blue line in the second graph shows the expression of the lacZ fusion in LB. The first sample was from the dense culture (it had been left on my bench overnight). Once those cells were diluted into fresh LB the amount of beta-gal activity slowly fell, presumably because expression of the lacZ fusion decreased. Part of the reason the decrease is slow is that the cells still contain beta-galactosidase that will either be degraded or diluted out by cell growth. The amount of beta-gal activity remained low while the cells grew, until the cell density got quite high (OD600 about 1.0), when the activity began to increase. I know that the low activity reflects lacZ expression by cells in exponential growth, and not residual enzyme from earlier induction, because I diluted part of the log-phase culture 100-fold into fresh LB+Amp, and found the same activity after these cells had spent an additional 2 and 3 hours in exponential growth (blue circles).

But once the cells got so dense that growth slowed, the beta-gal activity increased, and by 300 minutes had become at least as high as that of the overnight cells from my bench. This is a weaker version of the gene induction we see as Haemophilus influenzae approaches stationary phase.

In H. influenzae, rapid transfer to a medium lacking most nutrients causes strong induction of CRP-S genes, including the homolog of ppdA. To mimic this I transfered the log-phase cells to a medium consisting of minimal salts (M9) plus the same small amount of amino acids we use for H. influenzae. The two red bars show that this transfer caused a more dramatic increase in beta-gal activity.

So, these are weak but nice results. They suggest that CRP-S genes in E. coli are regulated by similar factors to those regulating CRP-S genes in H. influenzae.

[Just to check, I also induced lacZ expression in wildtype (lac+) cells by adding IPTG. Even though these cells have only one copy of the lacZ gene (in the chromosome, rather than one on each of many copies of a plasmid), after 60 minutes they expressed 7 times as much beta-gal activity as the starved ppdA cells did after 120 minutes. This is slightly less than the ppdA cells expressed after Sxy was overexpressed from the pASKA plasmid, and suggests that the CRP-S promoter is being only partially induced by my starvation conditions.]

What next? I need to repeat this time course, starting with a proper overnight culture and taking more time points, especially after log phase. I need to get a longer log-phase series too. And a time course in the starvation medium.

Which E. coli is best?

I haven't yet tested whether different growth/non-growth conditions alter the expression of the ppdA fusion (though I did get the sxy manuscript resubmitted today), but a paper I came across reminded me of another issue I need to consider.

The paper is an opinion piece titled "Laboratory strains of Escherichia coli: model citizens or deceitful delinquents growing old disgracefully?"; it just came out in the journal Molecular Microbiology (Mol. Micb. 64:881-885). The authors argue that the standard K-12-derived strains of E. coli that microbiologists and molecular biologists typically use are not at all representative of the strains out in the natural environment. For one thing, lab strains have lost about 20% of their genomes. These are (by definition) non-essential genes, but their loss no doubt affects cellular metabolism, so that the metabolic interactions we see in K-12 strains may be quite different than those in natural strains. Furthermore, the culture conditions we use (rich broth, lots of oxygen, no competitors) are unlikely to ever occur in nature. Their long maintenance under these and other unnatural conditions means that the lab strains will have evolved by accumulating cryptic mutations that are beneficial under lab culture conditions but that may have very different and perhaps harmful effects in the natural environment.

I already knew this. It has important implications for my search for conditions that induce expression of competence genes in E. coli. Right now I'm working with standard lab strains, but it's all too possible that one of their ancestors lost the ability to express the genes encoding homologs of H. influenzae's competence genes, or to assemble the proteins into functional DNA uptake machinery.

So I should perform my tests on a less lab-adapted, more 'natural' strain as well as on the K-12 strain the ppdA::lacZ fusion is in. But this raises two problems. First, which strain should I use? I do have a fairly-ancestral K-12 strain, but really I should use a 'wild' strain recently isolated from the environment. The NCBI Microbial Genomes page lists 8 completely sequenced E. coli genomes; their sizes range from 4.6 million bp (K-12) to 5.6 million bp (O157:H7). And the various wild strains have genomes that are quite different from each other - does this mean I would need to test many strains before giving up? I'm also not meticulous enough to be trusted with a strain that's seriously pathogenic to humans, so the two O157:H7 strains are out, as is the uropathogenic strain*.

Second, I'll need to transfer the necessary genes into this strain; first a lacZ mutation making it Lac-, then the ppdA::lacZ fusion so I can assay induction by Sxy. This might mean that I need to get my P1 transductions working after all (I had been thinking I could let them slide now that I've found that the ppdA::lacZ fusion can be used as an indicator of CRP-S induction by Sxy). I don't need P1 to move in the ppdA fusion because it's on a plasmid, but transduction would certainly be the easiest way to move a lac- mutation in. If I'm lucky, maybe some of the wild strains are naturally Lac-. (Probably not; most screens for wild E. coli start by treating anything that's Lac- as not E. coli.)

* NCBI also lists 6 sequenced 'Shigella' genomes - we now know that the bacterial strains assigned to the genus Shigella are really variants of E. coli. But I have absolutely no intention of working with these very pathogenic bacteria.

Manuscripts progress

Last week I finally submitted our manuscript about how CRP acts at CRP-S sites in H. influenzae and E. coli. The most interesting result in it was obtained by the gsnpiw just before he left for Belize. He showed that CRP-S sites contain regulatory sequences just upstream of the FCRP-S site that are needed for transcriptional activation; these sequences aren't typically present in CRP-N promoters.
Still to resubmit is our manuscript about Sxy. It too has been enhanced by a new result from the gsnpiw. The experiments are described here). He only had time to do it once before he left, so we're not treating them as part of the paper's Results section, but the results are sufficiently good to be briefly described in the Discussion section. They show directly that mutations in sxy do alter the ability of the sxy mRNA to be translated. Two of the reviewers had suggested we do a different experiment ("toeprinting") which would have tested whether the mutant mRNAs differ in their ability to serve as templates for a polymerase. In my cover letter to the editor of our manuscript I'll explain why we think our new experiments are more informative than toeprinting would have been. I'll also explain that they're not ready to be part of the Results, and that, if the editor thinks that the experiments need to be completed and included in the Results, we'd like a two-month extension of our revisions deadline so they can be completed after the gsnpiw returns from Belize.

I spent much of yesterday making changes to the manuscript and figures and writing responses to the many points raised by the reviewers. A downside of getting four thorough reviews of a manuscript is the very large number of issues they raised. I'm not complaining, as almost all of of these issues lead to improvements; either the reviewer is right, and we make the suggested change, or the reviewer misunderstood what we meant, and we clarify our writing to prevent the misunderstanding.

Only a few points remain to be dealt with. We used lacZ fusions to examine the effect of sxy secondary structure, and one reviewer wants more background information about the behaviour of the reference fusion. This data is in the PhD thesis of a former grad student (an author on the paper), so we may be able to simply refer to it there rather than adding the data to the manuscript's Results.

Time course and IPTG dependence

I've finished the two experiments I described in the last post: (1) a time course to find out how long it takes for induction of sxy by IPTG to give production of beta-galactosidase by the ppdA::lacZ gene fusion, and (2) a 'dose-response curve' to see how much IPTG is needed to induce the fusion. And I'm posting the resulting graphs, along with a sketch of the E. coli cells I'm using for this experiment.

The sketch shows that the cells contain two plasmids. The one on the left is pASKAsxy; it carries the E. coli sxy gene (red) under the control of the Plac promoter (blue arrow). This promoter is normally OFF in these cells so the sxy gene is not expressed, but the promoter can be activated by adding a lactose-analog called IPTG. The plasmid on the right carries the lacZ gene, which codes for the enzyme beta-galactosidase. Here the lacZ gene has been placed under the control of the ppdA gene's promoter, so when ppdA would be expressed, the cells make beta-galactosidase. Beta-galactosidase normally digests lactose, but I can easily detect it because it also digests another lactose analog called ONPG, producing a bright yellow chemical (ONP).

So I give these cells IPTG, wait a while, give them ONPG, and measure how much yellow colour they have made. That's what I did yesterday. (Measuring the yellow colour was complicated by a problem with our spectrophotometer - it kept forgetting what the standard brightness was - but I managed to keep it working well enough to get my data.)

Here's the time course. I set up three cultures. One got no IPTG, one got the standard amount of IPTG (1.0 mM) and one got tenfold less (0.1 mM). It's easy to see that tenfold less IPTG gave just as strong a response as the standard amount, that neither dose produced a significant increase in beta-gal until at least 40 minutes had elapsed, and that there was a bit more beta-gal at 90 minutes than at 120 minutes.

So I have the most basic information I need. When I go on to test induction of the ppdA promoter by treating cells in ways that I think might induce expression of the normal chromosomal sxy gene, I now know that I should allow at least an hour to see the effect of my treatment.

There are no error bars because I only tested each combination of dose and time once. I'll need to repeat the experiment with more replicates to get publishable results. I'll probably also want to use more time points in the interval between 40 minutes and 90 minutes, and to try a denser culture as well as the quite dilute one I used this time.

It is worth repeating this experiment and getting some more details, because I find this long delay quite surprising. One control I didn't do is to examine IPTG induction of the normal lac operon (the lac promoter driving expression of the lacZ, lacY and lacA genes), but this is a very standard experiment done in most introductory biochemistry lab courses, and as I recall the ONPG starts to accumulate within a few minutes. A Google search just turned up a paper about mutations that affect transcription, which used IPTG induction of lacZ as an assay; their control showed that beta-gal began to accumulate after 2.3 minutes. If we assume that induction of the pASKAsxy promoter by IPTG starts producing Sxy by 2.3 minutes, why does it take so long for Sxy to induce the ppdA promoter and cause beta-gal to start accumulating? Is the sxy mRNA translated very inefficiently, so accumulation of enough Sxy protein takes a long time? Is the ppdA promoter activated only once a lot of Sxy has accumulated? The results may have implications for competence induction in H. influenzae too. We've found that cells need about 45 minutes to become competent, but I have been assuming (perhaps wrongly) that this time is mainly needed for assembly of the uptake machinery. We have some fusions of lacZ to H. influenzae competence gene promoters; I (or someone) should test the kinetics of induction of these.

Here are the results from the dose-response analysis. In this graph the X-axis is IPTG concentration, not time (and note the log scale on both axes). IPTG was added at time 0. We see that any IPTG concentration above 10 micromolar gives full induction (about 40-fold) after 120 minutes, and that even the highest concentration doesn't give any induction at 25 minutes. This may not be very useful information; because I don't have an easy to measure Sxy production directly, it may just tell me about the sensitivity of the lac repressor to IPTG.

So what next? I'm going to grow the cells containing the ppdA::lacZ plasmid (no pASKAsxy) and try the same changes to culture conditions I previously tried with the ppdD::lacZ fusion (which I now know couldn't have worked because it isn't inducible by Sxy).

What next (turning on Sxy in E. coli)?

OK, the fusion of the ppdA promoter to lacZ produces 100-fold more beta-gal when Sxy is overexpressed. What are the important things to do with it?

1. Find out how sensitive the ppdA promoter is to Sxy: The pASKAsxy plasmid produces a great deal of Sxy when its promoter is induced with IPTG, but we know that most of this Sxy aggregates into insoluble and probably nonfunctional 'inclusion bodies'. So I could try inducing with decreasing concentrations of IPTG, and measure the effect on beta-gal production. It may be that only very slight induction of Sxy will give a big induction of the ppdA promoter. Ideally the amount of Sxy would be measured too, but we don't (yet) have the antibody to E. coli Sxy that would let us easily make these measurements. The beta-gal vs IPTG assay will be easy. If I find that ppdA induction needs only a tiny bit of IPTG, I'll have more confidence that natural induction will need only modest induction of Sxy expression from its normal promoter.

2. I should also do a time course of induction, so I know how quickly the steps happen (IPTG causes sxy transcription, Sxy protein is made, Sxy causes transcription of ppdA::lacZ, beta-gal is made). I can do this with the standard (1 mM) IPTG concentration I used yesterday. Yesterday I waited 2 hours after adding IPTG before assaying beta-gal; now I'll try 1 min, 2 min, 5 min, 10 min, 30 min, 60 min, 90 min, 120 min. The results will let me optimize the assay conditions I'll use when looking for culture conditions that induce Sxy expression from its normal promoter.

What culture conditions should I test first? My previous finding that the baseline expression of this fusion is independent of Sxy and CRP means that I shouldn't bother looking for conditions that reduce expression, so I won't try adding glucose to the medium to turn off CRP. I will try adding cAMP to the medium to make sure CRP is fully active. I will do a time course, following growth of the culture in LB, looking especially at expression as the cells' growth slows at high cell density. I will transfer the cells from LB to minimal salts (with a small amount of amino acids) to mimic the inducing effect of transferring H. influenzae from rich medium to MIV. Maybe I'll try growing the E. coli in BHI before such a transfer, as BHI is a much richer medium than LB. Perhaps I should get an E. coli mutant that can't synthesize purines (or pyrimidines), to see if transferring it to minimal gives induction.

Other things I might do:

Should I try to get a strain carrying this fusion in the chromosome rather than on a plasmid? This would eliminate concern about possible variations in the number of copies of the plasmid, due to differences in culture conditions. And if the chromosomal insertion was stable I could eliminate the selective antibiotic from the test cultures. This sounds like a job for recombineering - I wonder if we can still get this working.

Should I sequence the ppdD::lacZ fusion and the hofM::lacZ fusion to see if they have mutations that would explain their failure to be induced by Sxy?

Results of the obvious next step

As I posted a few days ago, I planned to introduce a sxy-expression plasmid into my cells with CRP-S promoter fusions to lacZ, to see if the CRP-S promoters are indeed induced by Sxy. I did that yesterday, and I just finished calculating the amounts of beta-galactosidase activity these cells produced with and without induction of sxy expression.

Results: Of the three fusions I tested (all that I have), one produced almost 100-fold more beta-galactosidase when Sxy was present, one produced only slightly more, and one produced even less! The exclamation mark is because I expected them all to respond similarly.

The fusion that was induced by Sxy is to the ppdA promoter (really the promoter of a 4-gene operon). The H. influenzae homologs (not pilB but comNOPQ) are in a similar operon, which is very strongly induced in competent cells and strongly dependent on CRP and Sxy. The ComN and ComO proteins are known to be required for competence. This fusion is carried on a plasmid. The 100-fold induction when Sxy is overexpressed is what I was hoping to see. This means that the low expression when Sxy is not induced is indeed baseline, so I'm comfortable with this baseline expression not being dependent on Sxy or CRP.

The results with the other fusions are surprising. The one that produced only slightly more is a plasmid-borne fusion of lacZ to the hofM promoter (really the promoter of the hofMNOPQ operon). This is homologous to H. influenzae's comABCDE operon, and like comNOPQ it is required for competence and needs CRP and Sxy for induction. Its CRP-S promoter element doesn't look significantly worse than that of ppdA, but it was induced only three-fold by Sxy overexpression.

The fusion that produced even less beta-galactosidase when Sxy was induced is to the ppdD promoter; it's integrated into the chromosome rather than being carried on a plasmid. This is the most surprising result, as the grad-student-now-postdoc-in-waiting ('gsnpiw') used quantitative PCR to show that ppdD mRNA is induced about 100-fold by overexpression of Sxy from the same plasmid I just used. Induction of Sxy in the fusion-carrying cells not only failed to induce the ppdD::lacZ fusion, it actually made these cells quite sick; the cells with IPTG grew to less than half the density of the uninduced cells, and produced only about half as much beta-galactosidase per cell. Perhaps the chromosomal insertion in these cells has somehow made them sensitive to transcriptional activation of the ppdA gene.

Anyway, the good news is that I can use the ppdA fusion as an indicator of Sxy expression, in my search for conditions that induce Sxy and thus induce expression of CRP-S genes. I'll set aside (for now) the questions about why the other fusions don't behave similarly.

(Sorry I've been too lazy to include figures in my recent blog posts; I promise to improve.)

Cryptic effects of antibiotics and resistance alleles

This morning we had a meeting with members of another lab to discuss progress on a shared project. There isn't as much progress as I had hoped, but at least we now know what still needs to be done and who will do it.

The goal is to understand how gene expression changes when cells are exposed to antibiotics at concentrations so low they don't even slow growth of the cells, much less kill them ("sub-inhibitory concentrations"; abbreviated sub-MIC).

The first part of the project was to use microarray analysis to compare the amount of mRNA produced by each gene, when cells were grown with and without sub-MIC of the antibiotics rifampicin and erythromycin. Rifampicin inhibits production of mRNA by RNA polymerase, and erythromycin inhibits production of protein by ribosomes. This work was begun by a previous technician in our lab, and completed by an undergraduate working in the other lab. (The undergraduate just learned that she's been accepted into medical school here - Congratulations, Wendy!) I have a draft version of a conference poster with some results of this analysis, but we'll need to reanalyze the data in preparation for writing the paper we plan.

The complementary part of the project turns out to be still quite a long ways from completion. The plan is to compare the gene expression by normal (antibiotic-sensitive) cells growing without antibiotic to the expression by cells that carry a mutation making them resistant to the antibiotic (i.e. of a RifR strain and of an EryR strain). This will let us compare the set of gene-expression changes caused by resistance mutations to those caused by the antibiotic. We expect these two sets of changes to be quite different, because each represents a complex outcome of different adaptive and accidental responses to a different change.

Part of our lab's contribution was to isolate the necessary resistance mutations; that's done and we have sequenced the altered DNA so we know exactly what the changes are. The Rifampicin resistance mutation is in the rpoB gene; it creates a S->P amino acid substitution at position 509. An identical mutation is known to cause Rif resistance in Staphylococcus aureus. Isolating the erythromycin resistant strain was a lot more trouble, but it's sequenced too. It's in the L22 protein (part of the 'large' subunit of the ribosome; I forgot to note down the exact position).

Our collaborators have analyzed mRNA from the EryR strain, but unfortunately the cells were being grown in the presence of erythromycin rather than in antibiotic-free medium, so we can't use this analysis as we planned. I offered to make more mRNA, this time from cells grown without antibiotic, so the microarrays can be repeated.

Several other problems were discovered.

First, the array slides used for this analysis were quite old and had been stored in air at room temperature, which causes degradation of the DNA fragments spotted on them. When repeating the analysis we will need to use new arrays, and these must be ordered from the research group in London that makes them.

Second, the very expensive license (>$4000 per year) for the GeneSpring software used to analyze microarray results has expired. It belonged to another lab that had kindly let our collaborators use it. Luckily another lab in our group of labs is thought to have just purchased a new license, so we're going to approach the head of that lab to ask if we can use their software (on their computer, as it can't be copied to other computers) in exchange for a small financial consideration.

Alternatives to GeneSpring exist, and I think some are open access, but I suspect they require quite a bit more sophistication to use well. A Google search for "alternatives to GeneSpring" led me to BRB Array Tools, which is free from NIH and runs as an Excel add-in. It was developed by 'professional statisticians' (why do I not find this reassuring?). One strength of GeneSpring is its ability to integrate the array information with the genome sequence and metabolic pathways of the organism - this is especially valuable for simple compact genomes such as H. influenzae's.

A third problem is that neither research group (ours or our collaborators) has anyone with much experience with GeneSpring (much less any equivalent free software). I've used it (a few years ago), but I've forgotten most of what I learned. Luckily we don't need to use its more sophisticated abilities, just the basic analyses, but even so it's going to take a major investment of time to analyze the data once we get it all.

The obvious next step

I did a sloppy lab meeting presentation this afternoon on my inconclusive results so far with the lacZ fusion strains. We discussed whether I could just use the moderate constitutive expression from the CRP-S promoters as the baseline, and look for conditions that increase it (and test that the induction does depend on Sxy and CRP).

One of the post docs reminded me that we have a plasmid carrying an inducible version of the sxy gene, and made the excellent suggestion that I should introduce this plasmid into these strains and see if overexpressing sxy causes a dramatic increase in beta-galactosidase activity. If not, then the system is unlikely to be much use. But if it does (and it should) then I'll know the range of signal I can expect, and can get to work testing possibly-inducing conditions.

She already has a stock of the plasmid, so I just need to make the recipient cells competent and transform this plasmid in. I'll test both the strain with the chromosomal fusion and the two strains carrying the plasmid-borne fusions. She says that the origin of replication of the sxy-expression plasmid is compatible with (i.e. different from) the origin of the fusion plasmids.

Where'd the regulation go?

I wanted to find out why the plasmids with fusions of CRP-S promoters to lacZ (hofM::lacZ and ppdA::lacZ) give such high lacZ expression. Because CRP-S promoters require both CRP and Sxy for high activity, I transformed the plasmids into strains with crp or sxy knocked out. But I also needed crp+ sxy+ cells of the same genetic background (frozen as RR1321; I forget its original number).

So yesterday I made cells of this strain 'chemically competent' to take up plasmids by incubating them in a cold solution containing rubidium chloride. (The cells didn't become 'naturally competent' - that's the long-term goal of these experiments.) I scaled down the cumbersome procedure described in the methods manual. This made it much faster because the cells could be collected by filtration and further concentrated by microcentrifugation, rather than using a big slow centrifuge. I wasn't sure this would work well but it did - this morning I found about about 100,000 transformants on my ampicillin plates).

So now I had all the strains I needed to test whether expression of the fusion depends on CRP and Sxy. Here's the beta-galactosidase 'activity' produced by each strain (in 'Miller units'):
Parent with no plasmid: 3 units
Parent plus either plasmid: 194 and 230 units
crp knockout plus plasmid: 301 and 406 units
sxy knockout plus plasmids: 361 and 463 units.
The differences between the various plasmid-containing cultures are probably not significant, as in this quick experiment I didn't take pains to make sure the cultures were all at exactly the same growth stage. This is unlikely to be important because my previous analyses found no little dependence on growth stage.

The conclusion? Expression of these fusions is independent of both CRP and Sxy. This is a troublesome result. The most likely explanation is that the inserts that carry the CRP-S promoters also contain quite a bit of sequence upstream of the promoter, and these sequences may be responsible for the high basal expression.

I will first repeat the experiment, this time having all the cultures at the same growth stage. Then I'll email the researchers who kindly gave us these plasmids to ask if they found high basal expression too. And I'll start looking the sequences of the promoter-containing inserts, in preparation for engineering them to remove extraneous sequences.

Transformations (good) and transductions (bad)

My transformations worked. The crp knockout cells transformed with the ppd::lacZ or hofM::lacZ fusion plasmids gave me about 200 AmpR KanR colonies each, which should all contain the appropriate plasmid because the no-DNA control gave no colonies. The sxy knockout cells must have not been very competent, as they gave only 1 and 2 transformants with the two plasmids, but the no-DNA control gave none so these too are probably right.

I've streaked all the sxy knockout transformants and 3 of each of the crp knockout transformants. Before I go home (if the new colonies are big enough to see), I'll inoculate them into LB+Amp+Kan, so that tomorrow I can (1) do plasmid minipreps to check that they have the right plasmids, and (2) do beta-galactosidase assays to measure promoter activity in the mutant backgrounds.

I realized this morning that the proper comparison for these strains should not be with the plasmid-carrying strains I already tested but with plasmid-carrying strains that have the same genetic background as the knockouts. We do have the parental strain of the knockouts, so I've streaked it out and tomorrow will make it competent and transform it with the plasmids.

I was growing cells and pouring plates and making P1 lysates today, in preparation for various carefully done transductions, but discovered that the LB broth I was using for today's cultures is contaminated. So I'll need to start this over with fresh medium - probably not for a few days, as two manuscripts are urgently in need of attention. Both need to be completed within the week, as the senior author is about to leave for two months of ecological R&R in Belize. One is the sxy manuscript - data is still being generated and we have four reviewers' comments to address. The other is the manuscript describing what Sxy does at CRP-S promoters. Again data is still being generated, but it's otherwise almost ready to submit.

Final revisions of the sxy manuscript

The sxy manuscript we submitted in late February came back a month or so later with four (!) thorough and quite favourable reviews. (The exclamation point is because most papers get two reviews.) We're now doing the revisions, and hope to resubmit later this week.

It's taken us two months to get to this point because two of the reviewers asked that we improve the evidence for the role of secondary structure as a regulator of sxy mRNA translation by doing an analysis called "toeprinting". The name is by analogy with "footprinting", where a DNA-binding (or RNA-binding) protein protects a specific segment of DNA (or RNA) from cleavage by a nuclease. In toeprinting, an RNA-binding protein or other obstacle is detected as a position where a polymerase stalls while copying the mRNA.

The PhD student working on sxy regulation decided that a different technique would be more appropriate, so he's directly measured the translatability of wildtype and mutant sxy mRNAs in an in vitro system. The experiments were delayed while he finished and very capably defended his thesis, so now he's finishing the experiments as his first postdoctoral work. He's also revising the figures as requested by various of the reviewers, and has already done some of the rewriting as part of his thesis. So tomorrow we'll sit down and see how much we can finalize.

P1 lysates bad but alternatives good

Most of the old P1 lysates I had left on my bench were dead, though one was OK and the ones in the fridge were still pretty good. So I made fresh lysates yesterday on the wildtype strain W3110, but the titers are lower than I expected, and maybe not even high enough to give acceptable transduction.

So I'm going to do one more try. I found what looks to be an excellent protocol on the Open Wetware pages. I'll follow their instructions exactly (no more "I'm so smart I can safely cut corners"). I'll also try the lysate I generated yesterday, and make fresh lactose plates for the selection. Two days ago I streaked the donor and recipient strains on lactose plates to check how the donor and recipient grow. The donor grows well, giving small colonies after 24 hrs and big ones after 48 hrs. The recipient does grow a bit on lactose, giving tiny colonies after 48 hours, presumably because it lacks the lactose uptake protein but can grow very slowly on lactose that leaks in to the cell. These tiny colonies are what I saw in my last transduction.

Remember why I'm doing these transductions? I want to combine the ppdD::lacZ fusion (indicator of CRP-S promoter activity and thus of Sxy and CRP activity) with a sxy knockout and with a crp knockout to find out whether the high constitutive expression of the fusion depends on its CRP-S promoter. And I'm using the ppdD::lacZ fusion because I want to find out how to induce Sxy activity and CRP-S promoter activity in E. coli. If I find that the constitutive activity is dependent on Sxy and CRP, I will suspect that such promoters are normally moderately active. If the activity is not Sxy and CRP dependent, I will conclude that the fusion is not a good indicator and not work with it any more.

In Tuesday's post I described checking some plasmids, to see that they had the expected inserts. We asked for these plasmids (gift from other researchers) because they too contain fusions of CRP-S promoters to lacZ, and can be used in the same way as the ppdD chromosomal fusion. The two promoters are from the ppdA gene and the hofM gene, which are both homologs of CRP-S genes that H. influenzae requires for DNA uptake.

So yesterday I compared the amount of beta-galactosidase (product of lacZ) produced by cells carrying these plasmids to that produced by the ppdD fusion. Even though cells have many more copies of the plasmids than of a chromosomal fusion, they produced quite a bit less beta-galactosidase. I don't need to use P1 transduction to check whether this expression is dependent on Sxy and CRP. Because I already have preps of the plasmids, and I have strains carrying the sxy and crp knockouts, I can just transform the plasmids into the knockout strains, selecting for the ampicillin resistance of the plasmids and the kanamycin resistance of the knockouts. Even better, one of the post-docs has offered me some already-competent cells of the crp knockout.

P1

Yesterday I finally sat down at the bench and wrote up the results of my first round of P1 transductions. Three transductions, three failures.

Two of the transductions were to construct double-mutant E. coli strains carrying both the ppdD::lacZ fusion (reporter for activity of the CRP-S promoters, and thus of Sxy) and either a sxy::kan knockout or a crp::kan knockout. The goal is to test whether the high baseline expression of the reporter is due to Sxy-dependent CRP-dependent activity of the CRP-S promoter, or to baseline promoter-independent expression (e.g. from other sequences in the fusion).

Both of these transductions produced lots of colonies on the Amp+Kan selective plates. But one of the negative controls (same cells, but no P1) produced just as many, suggesting that something other than transduction was responsible for the colonies. As controls for the selection I had streaked the reporter and knockout cells onto the same plates, and onto Amp or Kan plates. These told me that the Amp+Kan plates were faulty, perhaps because the Amp was old.

I would have concluded that I should just do it again, with fresh plates, but my other transduction was a positive control for transduction (one that I knew should work), and it didn't work. This experiment tried to transduce the lacY gene from the wildtype strain W3110 into the lacY mutant C600, selecting for Lac+ by plating on minimal salts with lactose as the only sugar. The negative control was C600 with no P1. The positive control of plating the cells on minimal salts plus glucose worked well - everyone grew fine. I can't remember whether I also plated the donor cells on minimal lactose - this would have been a good control. Colonies on the lactose plates were initially very tiny, and the negative control (no phage) cells gave just as many as did the cells with phage (in fact more, probably because the phage killed some cells). So there's no evidence that this positive-control transduction worked either.

What to do next? Check growth of W3110 on the lactose plates? Repeat the lacY transduction? First I must recheck the titers of my supposedly-transducing lysates (plated last night). And go back over my notes, looking for any corners I might have cut. This is just another example of the truth of my favourite saying:
"Most scientists spend most of their time trying to figure out why their experiments won't work."

Plasmid checking

Today I made it to the bench, to do plasmid preps to check that some gift plasmids had the expected structure. I needed to check this before freezing stocks of them in our lab strain collection. The plasmids contain fusions of E. coli CRP-S promoters to the E. coli lacZYA genes. I'll be using them as reporters to test my attempts to induce expression of the E. coli sxy gene. The genes are hofM (the E. coli homolog of H. influenzae's comA) and ppdA (the E. coli homolog of H. influenzae's pilB)

So I did minipreps using our nice Sigma kit, and digested the plasmids with a pair of enzymes that would give one of two patterns. I didn't know which of three restriction sites the inserts were in, so I could only predict two possible patterns from my digests (either one 11kb fragment, one 1.2kb fragment and tiny fragments of either ~300 or ~600 (hofM or ppdA); or one 11kb fragment and fragments of ~1.5kb and 1.8kb (hofM or ppdA).

But for the first time in my research career, I absentmindedly plugged the gel box electrodes in backwards - black cable into red connection and red into black. When I went back to check 25 minutes later, the tracking dye was running backwards out the top end of the gel. Rather than starting over, I just reversed the electrodes and let it run back the way it was supposed to go for a couple of hours, hoping that some information would be usable. Much to my surprise, the gel turned out great (the migration process must be perfectly reversible even though the DNA travelled twice through the wells). If the inserts had given the two tiny fragments they would have been lost during the backwards excursion, but luckily the cloning had generated the 1.5 and 1.8kb fragments instead.

Might H. influenzae competence be phase-variable?

One of the postdocs just raised an issue I've never seriously considered. Many surface structures on bacterial cells undergo what's called "phase variation". That is, a key gene controlling the structure has evolved to have a high rate of mutations that switch it from an active allele to an inactive allele, and from the inactive allele to an active one.

By "high frequency" here I mean more often than one switch per million cell divisions. Switching is thus still a very rare event, but is much higher than the background mutation rate for normal DNA sequences. Such elevated frequencies are usually caused either by short sequence repeats that cause DNA polymerase to add or miss bases in critical positions, or by specific DNA-altering enzymes that recognize the gene.

That's the proximate cause of the variation. The ultimate cause (the evolutionary cause) is thought to be natural selection created by predators or host immune systems that recognize the surface structure and attack cells expressing it. Under such pressure, a cell that has turned the structure off will have an advantage, so cells with elevated mutation rates affecting the structure are favoured. Because the structure is strongly advantageous in the absence of external attack, selection favours cells that also have a high rate of reversion mutations that switch the structure back on. Such genes are often called "contingency loci".

Competence for DNA uptake requires expressing DNA uptake proteins on the cell surface, so it's a logical target for attack by the host immune system, and thus perhaps for phase variation. But how would we detect it? In Neisseria competence is known to be phase variable, but only because it depends on the phase-variable expression of type 4 pili, a phenotype that is easily assayed in the lab. Screening H. influenzae cells for phase variation of competence is likely to be very difficult, as our only assays are uptake of radioactive DNA and transformation to antibiotic resistance.

Rather than screening for variation, a more efficient approach is to examine the H. influenzae genome for sequences that could promote such variation, and check each for its ability to affect competence. These have been thoroughly investigated by Richard Moxon and his colleagues. They found no enzymatic switches but many short sequence repeats affecting production of complex carbohydrates on the cell surface. Now we need to carefully check whether any of these could also affect genes needed for DNA uptake.

Simplifying the cyclization experiments

One of the analyses I proposed in the DNA-uptake grant proposal is designed to find out whether uptake signal sequences are unusually flexible or bent, by testing whether presence of a USS helps short DNA sequences bend around into circles whose ends can be joined by DNA ligase.

Because competition for grants is very tight this time around, we want to be prepared to submit even better proposals in September if the ones we submitted two months ago aren't successful. This means we want to have lots more preliminary data showing that the experiments we propose will actually work. So one of the post-docs has been working to get preliminary results for this circularization experiment.

The DNA fragments should be about 200bp long; shorter ones usually can't circularize at all, and longer ones have enough flexibility that the ends readily bump into each other. Even for fragments in the right size range, exact length is critical because of the 'polarity' of the ends of the double-stranded DNA. Each 'end' is really the ends of two base-paired strands, one ending in a 3'-OH and the other in a 5'-P. When the ends do bump into each other, ligase can only join them if a 3'-OH is aligned with a 5'-P (the bond must be a 3'-5 connection, not a 3'-3' or 5'-5'). Because the two strands of the DNA wind around each other every 10.4 base pairs, the length of the DNA fragments must be approximately a multiple of this length so that the ends will meet in the right alignment. The strategy is to use PCR to synthesize fragments of the right length, and she designed two sets of primers, giving fragments of 208 and 260bp (my notes say 260 but I think I have it wrong as this seems too long, unless it's the positive control).

The ends created by the PCR process are not easy to ligate, because they have inconvenient incompatible tails, so her primers include sites for digestion by restriction enzymes. Cutting both ends of the PCR product with the same restriction enzyme will generate compatible 'sticky ends' that can base pair with each other. The base pairing will hold together any ends that do bump into each other in the right orientation until ligase can seal them together permanently.

Sounds good so far. But there's one more factor. The circularization reactions must be done at low DNA concentration to decrease the frequency of ends of different molecules bumping into each other. This intermolecular reaction has 'bimolecular' kinetics, meaning that its rate depends on the DNA concentration. In contrast, circularization has 'unimolecular' kinetics, and its rate is independent of the presence of other molecules and thus of DNA concentration. Doing the reaction at low concentration is easy (need less DNA), but detecting the results of the ligation is hard, because small amounts of DNA (linear or circular) are difficult to see in the gels used to separate the different conformations that result from the ligation.

The solution is to label the DNA fragments with 32-P, making even very small quantities easy to detect. The post-doc followed a published method for labeling DNA for these experiments, which should put a 32-P at each 3' end of each fragment. The fragments are first cut with the restriction enzyme, purified using little spin-columns, and then incubated first with a phosphatase, to remove the non-radioactive P from each 3' end, and then with the 32-P nucleotide and a kinase, to put the hot phosphate on. Then the fragments are purified again.

Initially I was concerned by the low recovery; only about 10% of the input DNA was recovered after all the intervening reactions and clean-up steps. After a bit more thought I became more concerned by the labeling reactions, mostly because I've always found phosphatases to be nasty treacherous enzymes that don't know when to stop. If the phosphatase removes more than the single terminal phosphate, the fragment will not be circularizable even if the kinase then does its job correctly. Furthermore, any fragment that the kinase misses will also not be circularizable, even if the phosphatase has behaved itself. Even if one end is processed correctly, any problem with the other end will still prevent circularization. In principle these problems can be controlled for, but any experiment-to-experiment variation will invalidate the conclusions we're hoping to achieve.

Luckily, once I started worrying about these issues I realized that we can eliminate both the recovery problem and the phosphatase/kinase problems by labeling the DNA internally rather than at its ends. So the new plan is to add a labeling step after the PCR reaction. This will be essentially one extra PCR cycle, this time with one radioactive precursor nucleotide added to the mix. The resulting DNA fragments will then only need to be digested with the restriction enzyme and cleaned up once.

When the post-doc gets back from her visit home, we'll still need to solve the problem of why the gels run so oddly, but at least we'll have enough labeled DNA to do lots of tests. The gel problem may be related to the high concentration of ligase needed in these reactions. The standard ligase stock is purchased in 50% glycerol at the low concentrations needed for cloning reactions, and the circularization reactions use so much ligase that they are about 25% glycerol. I'm hoping we'll be able to buy a high-concentration ligase stock, rather than having to make our own ligase....

no transductants??

I thought I had a lot of KanR AmpR transductants from infection of the ppdD::lacZ strain (AmpR) with the P1 lysate made on the sxy::kan strain, because lots of colonies grew overnight on the Kan+Amp plates from this infection but not from mock-infected cells. But I picked 4 colonies and streaked them on various plates (Kan, Amp, Kan+Amp, Maconkey-lac) and they mostly didn't grow. Now I suspect that some ofthe plates I used may have been too old (Amp is unstable) or otherwise problematic.

My test transduction of lacY from W3110 into C600 doesn't seem to have worked either (though I did find the TTC and get cute little red colonies). And the subsequent transduction of crp::kan into the ppdD::lacZ strain doesn't seem to have worked either.

After I carefully recheck the controls I'll redo it all with fresh plates.

Doing it right (P1)

I've been futzing around with phage P1, trying to get a high-titer lysate from a single clear (P1vir) plaque, with no success. Part of the problem is the high frequency of what I guess are revertant phage, giving larger and turbid plaques, and partly it's that I've been trying to squeeze the work in between other stuff (peer-reviewing manuscripts and a proposal, preparing my freshman biology final exam, etc.).

But today's clear, and I'm ready to roll. Last night I innoculated cultures from single colonies of the E. coli wildtype strain W3110 and three mutant strains: ppdD::lacZ, sxy::kan and crp::kan. Last night I poured lawns of W3110 with appropriate dilutions of two of my not-very-satisfactory lysates. When I could the plaques this morning I'll know the titer of these lysates, and can then use one or the other to infect the four E. coli strains. I'll do these infections in broth and maybe also on plates.

Then I'll collect the lysates and titer them all overnight on W3110 lawns. I may also do a test transduction overnight, using the W3110 lysate (before I know its titer) to transduce the lacY gene into the lacY mutant strain C600. This should give me all the information I'll need for tomorrow, when I want to infect the sxy and crp knockout mutant strains with the ppdD fusion lysate, and vice versa, and select for transductants that have both the ppdD fusion and the sxy or crp knockout. Then I can test whether the high 'baseline' expression of lacZ in the ppdD::lacZ fusion is due to high baseline activity of its CRP-S promoter, as such activity should drop dramatically if sxy or crp is knocked out.

I'll need to make some minimal-lactose plates to properly score the result of my test transduction, for which I'll need to make up and sterilize stock solutions of the amino acids threonine and leucine, the vitamin thiamine, and lactose, because C600 is thi- thr- leu- as well as lacY-. (I used this strain a lot in grad school, and its genotype is burned into my brain.) And ideally I should also make up some of the TTC we used to put into minimal plates that makes the colonies turn red and easy to see. Right now I can't remember what TTC stands for, but it will come back to me, and I think we have a big bottle of the stuff somewhere.

How to compare protein sequences?

In the last post I described an analysis that depends on comparing protein sequences. There are two different ways to do the comparison; and I need to decide which is more appropriate for our analysis.

Both methods rely on first aligning the amino acids in the proteins to be compared (here we'll only be comparing two proteins at a time), and then comparing the amino acids at each aligned position. The goal of the alignment is to align amino acids that are homologous - that is, those that are similar because of descent from the same position in the ancestral sequence both proteins evolved from. (If the proteins are not themselves homologous the analysis can't be done.)

In studies of evolutionary relationships, the usual method of comparison is to simply count how many of the positions have identical amino acids, giving a "% identity" score. In studies of protein function, scores based on the functional similarity of the aligned amino acids are often used. These rely on a matrix that gives similarity scores of all pairwise combinations of amino acids. These matrix scores are themselves derived from comparisons of large numbers of aligned amino acids, giving highest scores to amino acids which most often have evolved to perform the same role in a protein. For example, valine and leucine are commonly found in homologous positions, and matrices give this pair a high score. Wikipedia gives a good explanation of the use of matrices to compare protein sequences.

I'll post later about the issues raised by our analysis problem.

Analyzing the effect of USS on the coding function of genes

While I've been doing other things a collaborator has been working hard on a comparative genomics project that will tell us how much impact uptake signal sequences (USS) have on gene function.

Reminder: USS are short sequence motifs (the longest are ~30bp) present in many copies in the genomes of naturally transformable bacteria, probably because the cells preferentially take up DNA fragments containing the motif. Most of the USS in the Haemophilus influenzae genome are in coding sequences, and we want to find out whether their presence forces genes to specify sub-optimal amino acids at positions encoded by USS.

This analysis is testing the effect of USS by comparing the amino acid sequences of proteins with and without USS. For each H. influenzae gene with one or more USSs, we first find homologous protein sequences from at least three genomes with no USS. We compare these three protein sequences with each other (that's three no-USS comparison scores), to get a measure of how strongly selection acts on the protein, especially on the segment that in H. influenzae is specified by a USS. Then we compare each of the three with the H. influenzae sequence (that's three +USS comparison scores).

Then we compare the mean no-USS score with the mean +USS score; if the scores are similar then we conclude that the USS doesn't significantly constrain the protein's function. There's a lot of random variation, so we do this for every USS-encoded gene in the the genome and then plot each pair of scores as a point on a scatter-plot. Points that fall on a diagonal line represent genes whose USSs don't constrain them, and points that fall below the line represent genes whose USSs may be causing problems.

We're not interested in specific genes, but in the general picture - we want to know whether, on average, USSs cause problems or not. A preliminary analysis done years ago suggested they don't, but the answer from this new improved analysis will be interesting in any case.

Lysates!

Well, lysates aren't really that exciting, but it's been a while since I got to do anything with phage. I used both the turbid-plaque and the clear-plaque streaks to infect two cultures, one with the ppdD::lacZ reporter fusion and one with the sxy::kan knockout. They all grew up nicely and then promptly lysed. Titering showed that the clear-plaque phage gave about 2x10^10 pfu/ml, all the small clear plaques expected of P1vir. The other phage gave a mix of clear and turbid plaques, suggesting that the colleague who made the source lysate had neglected to start from a single plaque (naming no names).

Today I'm going to make a good stock lysate, by infecting the "wild type" E. coli strain W3110. (I put wild type in quotes because the strain does not carry the lambda phage present in the original "wild" K-12 isolate. And I'm going to do the transduction that was the point of getting a P1 lysate, to create a strain carrying both the ppdD::lacZ reporter and the sxy::kan knockout). I have lysates of both strains so I'll try the transduction in both directions, each selecting for AmpR KanR (the reporter carries AmpR). And a control transduction transferring lacY from W3110 into C600.

P1 (vir?)

I have some phage P1, courtesy of a colleague. But (as usual) complications have arisen. The colleague is out of town, and his P1 collection consisted of five different lysate tubes, of unknown ages as none of the tubes had a date. So I tested them all for viable phage by streaking a drop onto a plate that had a lawn of host cells on it, and looked this morning for the plaques (spots of lysis) that phage would make.

Two of the lysates gave no plaques, so they were probably very old. Two others (#1 and #4) gave lots of conspicuous plaques, and #2 gave tiny plaques. The colleague tells me (by email) that #1 is his newest stock, but the big plaques are not what I expected.

The phage should be not normal (wildtype) P1 but a mutant called P1vir (vir = virulent). Wildtype P1 can form lysogens, becoming dormant in cells that are then immune to being killed by other P1. The large plaques I see look like plaques made by wildtype P1, because they have cloudy centers (the plaques are "turbid"), where surviving cells are growing; these are likely to be lysogens. The virulent mutant should not form lysogens, so its plaques should be clear. Our manual of bacterial genetics methods says that P1vir plaques are tiny, so I think lysate #2 is more likely to be genuine P1vir. Lysates #1 and #4 would then have been grown from P1vir that had mutated, reverting to wildtype. (Perhaps the person who made the lysate mistakenly picked a "nice big plaque" instead of the more common tiny plaques.)

Wildtype P1 is not very suitable for doing transductions because many of the survivors are likely to be lysogens, whereas we usually want to work with phage-free cells. I'll do test transductions with lysates #1 and #2; if #2 works well I won't bother doing experiments on #1 to sort out the virulence issue, but just make a good fresh lysate from a single plaque of #2. Part of this lysate will be given to the colleague who supplied the original lysates.

Before making proper lysates and testing transduction I need to grow up appropriate E. coli host strains. I did the preliminary checking on the strain carrying the ppdD::lacZ fusion. Unfortunately the computer that we keep our strain list on is in the shop, but one of the post-docs has a recent backup.

Results of beta-galactosidase assays

So I did some beta-galactosidase assays on the strain with ppdD fused to lacZ. The immediate goal was to characterize baseline transcription of the ppdD gene, with the longer goal of finding ways to turn it up by inducing expression of sxy.

The researcher who kindly sent us the strain with this fusion said that its colonies are pale blue on plates with the beta-gal indicator X-gal, but for me they were quite a strong blue. So I wasn't too surprised that my assays showed moderate beta-gal activity in all the conditions I tested. I tested cells in exponential growth and after overnight culture in LB, in LB+glucose (which should prevent production of cAMP and thus expression of the ppdD gene's CRP-S-regulated promoter, and of cells in LB+glycerol, which should allow cAMP production like plain LB but provide the same amount of extra carbon source that glucose does. I also tested cells transferred from log-phase growth in LB+glucose to minimal salts with added amino acids ("M9+caa"); a treatment that might roughly approximate the competence-inducing effect of transferring H. influenzae cells from sBHI to MIV. None of these treatments made much difference; all the samples produced between 150 and 500 units of enzyme activity per ml of cells.

So the next test is to find out whether the transcription of ppdD depends on Sxy. If it does, then this suggests that Sxy is being produced at least a bit under the standard conditions I tested. If not, the expression is genuinely baseline, and maybe I should test other genes as indicators of sxy expression.

To do this test I need to introduce our sxy knockout into the ppdD::lacZ fusion strain. So tomorrow I'll go searching for the needed P1 lysate. The colleague who has it has unfortunately gone off to Europe for a couple of months, but he tells me that one of the people in his lab can help me find it.

Preparing for the E. coli sxy induction experiments

The strain and plasmids we've been waiting for arrived on Wednesday (thank you to the Pugsley lab), so now I can get started on my attempts to induce sxy expression in E. coli. And we do already have an E. coli sxy knockout from the Japanese group, so solving the recombineering problems isn't critical (though I should get to work on that anyway).

What to do first? The initial tests will use a reporter strain carrying a fusion of the ppdD gene to lacZ. The person who sent the strain says it should be pale blue on X-gal plates, indicating weak baseline expression of lacZ. I should first characterize its lacZ expression more precisely by growing it under defined conditions and measuring the amount of beta-galactosidase (the lacZ product) with the substrate ONPG. This is a simple classic assay, done in every introductory molecular biology lab. I'll grow cells in rich medium and minimal medium, doing the assay at various cell densities.

I need to find out whether this baseline expression is independent of the transcriptional regulators Sxy and CRP/cAMP, which we know activate the ppdD promoter. This requires transferring the ppdD::lacZ fusion into strains that carry knockouts of either sxy or crp (or cya, which encodes the adenylate cyclase that makes cAMP), or transferring the sxy and crp knockouts into the ppdD::lacZ strain. We have both these knockouts, but I don't yet have the P1 lysate I'll need to do the transfers. And I need to find out the antibiotic resistances associated with each of these knockouts, so I can plan the selections. And I need to get all the strains from the grad student who's been working with them.

I do have what I need to test whether baseline expression of the ppdD::lacZ fusion is affected by cAMP levels. The grad student tells me cAMP is normally high in the rich medium LB, presumably because LB lacks glucose, causing the phosphotransferase system to activate adenylate cyclase to make lots of cAMP. If this is correct, simply adding glucose to LB should reduce cAMP levels. So should growing cells in a minimal medium with glucose as the carbon source. If the baseline lacZ expression from the ppdD promoter is due to weak activation of the promoter by CRP + Sxy, these conditions should reduce it, and adding exogenous cAMP should restore it. If it's just constitutive activity of the promoter, these conditions should have no effect.

Negative control for E. coli sxy experiments

I'm still planning the experiments to identify conditions that induce sxy expression (and possibly competence) in E. coli.

One issue I hadn't considered previously is the appropriate negative controls. If I'm using the sxy-inducible ppdD::lacZ fusion to indicate sxy expression, the best control will be a sxy-knockout mutant carrying the same lacZ fusion. We might already have such a mutant (from the wonderful Japanese group that provides clones and knockouts of E. coli genes), but if not I'll need to make one. This would be best done using the genetic technique called recombineering, which one of the post-docs has been trying to get working for us. I'll start working with her on this.

How hot can it be?

I'm going to test biotin-tagging of the ends of my big DNA fragments by doing a preliminary tagging with a radioactive (33-P) nucleotide. I need this test because I don't have specific information about the short single-stranded overhangs I expect these fragments to have. I don't know whether most fragments have overhangs, whether 3' and 5' overhangs are equally common, or how long the average overhang is.

But I realized that the unless the overhangs are very long, they will constitute only a tiny fraction of the DNA in such long molecules. So I don't expect much of my 33-P or biotin to be incorporated. I do know the specific activity of the 33-P I'll use (2500 Curie/mmol* **), and this lets me do a back-of-the-envelope calculation of how hot the DNA can get if the tagging reaction works perfectly.

[This is a long but not difficult calculation, relying on the kind of 'dimensional analysis' I learned in Grade 11 Physics class. It has Avogadro's number and Rosie's universal constant (10^18 bp/g) and arithmetic-simplifying assumptions that the average fragment is 75kb long and that a single 33-P nucleotide gets incorporated at each end. This last assumption would be right if 3' and 5' ends were equally common, blunt ends insignificant, and the average overhang about 8 bases.]

The result of this calculation: a perfect labeling reaction could incorporate enough 33-P to give only about 240 dpm per microgram of DNA. This is so low that the 33-P labeling experiment may not be worth doing at all. Using 32-P won't make a big difference, as the standard specific activity of 32-P nucleotides is 3000 Ci/mmol.

- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -

*A Curie is 2.2 x 10^12 dpm (radioactive disintegrations per minute).

** This is close to the theoretical maximum specific activity, with a 33-P as the alpha phosphate of every nucleotide.