- Home
- Angry by Choice
- Catalogue of Organisms
- Chinleana
- Doc Madhattan
- Games with Words
- Genomics, Medicine, and Pseudoscience
- History of Geology
- Moss Plants and More
- Pleiotropy
- Plektix
- RRResearch
- Skeptic Wonder
- The Culture of Chemistry
- The Curious Wavefunction
- The Phytophactor
- The View from a Microbiologist
- Variety of Life
Field of Science
-
-
Change of address1 year ago in Variety of Life
-
Change of address1 year ago in Catalogue of Organisms
-
-
Earth Day: Pogo and our responsibility1 year ago in Doc Madhattan
-
What I Read 20241 year ago in Angry by Choice
-
I've moved to Substack. Come join me there.1 year ago in Genomics, Medicine, and Pseudoscience
-
-
-
-
Histological Evidence of Trauma in Dicynodont Tusks7 years ago in Chinleana
-
Posted: July 21, 2018 at 03:03PM8 years ago in Field Notes
-
Why doesn't all the GTA get taken up?8 years ago in RRResearch
-
-
Harnessing innate immunity to cure HIV10 years ago in Rule of 6ix
-
-
-
-
-
-
post doc job opportunity on ribosome biochemistry!11 years ago in Protein Evolution and Other Musings
-
Blogging Microbes- Communicating Microbiology to Netizens11 years ago in Memoirs of a Defective Brain
-
Re-Blog: June Was 6th Warmest Globally12 years ago in The View from a Microbiologist
-
-
-
The Lure of the Obscure? Guest Post by Frank Stahl14 years ago in Sex, Genes & Evolution
-
-
Lab Rat Moving House15 years ago in Life of a Lab Rat
-
Goodbye FoS, thanks for all the laughs15 years ago in Disease Prone
-
-
Slideshow of NASA's Stardust-NExT Mission Comet Tempel 1 Flyby15 years ago in The Large Picture Blog
-
in The Biology Files
Not your typical science blog, but an 'open science' research blog. Watch me fumbling my way towards understanding how and why bacteria take up DNA, and getting distracted by other cool questions.
DNA secretion by Neisseria
The first poster was about the genetic machinery responsible for secreting DNA, initially characterized in N. gonorrhoeae. It's a large 'island' of genes that strongly resemble the transfer genes of conjugative plasmids (like the E. coli F factor); they encode a type IV secretion system like that used by these plasmids to transfer a single DNA strand into a new host cell. The island appears to have been integrated into the genome by a single crossover with chromosomal DNA, likely mediated by the site-specific recombinase that normally functions to separate newly-replicated daughter DNAs when replication is finished. (This would mean that the island was originally a circular plasmid-like molecule.) Most but not all N. gonorrhoeae strains have some version of this island. Most N. meningitidis strains have it too, but usually with major deletions or other changes expected to make it non-functional.
So is this an adaptation to promote genetic exchange? The evidence looks good that this element does cause cells to release DNA into the medium. The student whose poster I was at said that the DNA is single-stranded and originates from a specific oriT-like site in the element, just as we would expect for a conjugation system.
Amount of DNA released: The first paper reports that cultures after 4 hr growth contained about 150-200 ng of DNA/ml; this was almost entirely eliminated by a mutation in a component of the secretion system. A later paper quantitates the DNA produced by different strains in log phase; producers had 0.2-0.6 ng DNA per µg of cell protein.
Transforming activity of the released DNA: Transformation frequencies are high, between 10^-4 and 10^-3, consistent with the presence of high concentrations of DNA in the medium. Transformation is reduced 250-fold by addition of DNase I, confirming that this is not cryptic conjugation.
Timing of DNA release: The DNA appears in the medium in log phase, whereas cell death is expected mainly when growth slows and stops. The secretion mutant doesn't produce much DNA until the end of growth.
Mechanism of DNA release: The genetic analysis supports the hypothesis that the secretion system causes the DNA release; cells with mutations in secretion genes don't release DNA. But it doesn't distinguish between active secretion and release caused by toxic effects of the element.
Strandedness of the released DNA: This is a critical point. If all the DNA that appears in the medium in log phase is released by secretion, it should all be single-stranded. But single-stranded DNA transforms poorly - a 1981 paper by Biswas and Sparling says it's 100-fold worse than double-stranded DNA - so perhaps much of the released DNA is double-stranded. If so, it's probably not secreted DNA. If it is secreted it should also all be from the same strand of the chromosome - I don't think this has been checked.
The actual state of the released DNA is not clear. The student whose poster this was said it was all single-stranded, and that the dye used to measure it was specific for single-stranded DNA. But the papers describe use of a dye, pico green, that is advertised for measuring double-stranded DNA, and although the authors say that in their hands it also detects single-stranded DNA, they appear to have used double-stranded DNA to calibrate their assay.
I've just sent an email to the senior author on these papers, asking all these quenstions.
Neisseria DNA secretion at the ASM GM
a secreted nuclease.
Sent from my iPhone; typed with my thumbs!
Open Science round table at the ASM AGM
- First, being as open as possible about what we've done. This means publishing in open-access journals such as PLoS where that's reasonable, and paying large sums for open-access publication of the papers we publish in subscription-based journals. It also means posting pdf copies of our papers on our web pages even where the copyright forms we signed say we mustn't.
- Second, being as open as possible about what we're presently doing. I have a research blog, where I write about the experiments and analyses I'm doing (what I'm going to do today, the results of what I did yesterday). I also write about the other aspects of doing science, but I try to stick to my own research-related experiences. The other members of my lab have research blogs too.
- Third, being as open as possible about what we're planning. So I blog about the grant proposals I'm writing, and when I'm finished writing them I post them on our web pages at the same time that I submit them.
Time to stop playing around with the model
The playing around was largely motivated by a worry that something was wrong with the program component that gradually reduces the uptake bias when not enough fragments have recombined. The bias seemed to be becoming very low very fast, and this could explain why the model couldn't maintain the uptake sequences in the genome segment I've been giving it. I've now convinced myself that the bias-reduction component is behaving as it should, so the low bias must mean that the program is going through MANY sets of fragments before it finds enough it likes. Which probably means it's recombining the same fragments over and over within a single evolutionary cycle, not at all what I want to happen.
Progress on the Gibbs and record-keeping fronts
The new N. meningitidis Gibbs searches specifying small numbers of 'expected' occurrences succeeded in finding the DUS motif. The searches expecting 200 found about 2500 occurrences, and those expecting 500 found about 2600. I've now queue'd up some searches expecting more (1000, 1500 and 2000), only because my analyses of other genomes1have used 1.5 times the number of perfect cores, and for N. meningitidis this needs 2809 occurrences.
The motifs Gibbs found from a genome sequence with the RS3 repeats removed are very similar to those from the full genome sequence. The 'residue frequencies', expressed as percent of each base at each position, are identical, and the 'motif probability models', expressed as probabilities to three decimal places, differ by no more than 0.001. This means that the RS3 repeats are not perturbing the results of whole-genome searches, but they may be perturbing the intergenic-sequence searches. I ran a bunch of intergenic searches last night, with and without RS3s removed. Some of these runs ran into 'segmentation fault' errors, which I remember dealing with a couple of years ago, and I haven't analyzed the outputs yet. I'll try to do that quickly because I really need to focus on the Perl simulation work today.
Update: I had two OK searches with the RS3 repeats and one without. They found nearly identical numbers of occurrences (1570, 1573 and 1572). I made logos of one of each type and they are indistinguishable, so I can conclude that RS3 repeats don't affect the searches of intergenic sequences, at least when the search is expecting a small number of repeats.
The analyses of N. meningitidis coding sequences also showed that it's better to expect an unreasonably low number. With exp=1500, the searches found about 6500, but most were not DUS and the motif was only strong for four of the 12 positions. But with exp=100 the replicate searches found 900 and 901 occurrences, and these gave a very strong DUS-type motif.
Keeping records of computer work

When I do benchwork I consistently keep pretty good notes. I write down everything I do as I do it, on numbered and dated sheets of paper that go into looseleaf binders, organized by experiment. If I make notes on a scrap of paper, I tape them in. I write a brief plan (a few sentences about the point and design of the experiment) and end with some sort of conclusion or summary.
Gibbs can't find the DUS
Progress on multiple fronts
On the Perl simulations front, I've got the program running and used it to do the control simulations. The first controls use random sequences the same lengths and base compositions as the concatenated H. influenzae or N. meningitidis intergenic sequences, run with matrices specifying the corresponding USS or DUS core but with no recombination. These controls tell us what the baseline USS or DUS score is for a genome that hasn't experienced any accumulation. The second controls use the real H. influenzae or N. meningitidis intergenic sequences instead of random sequences, and run for a long time to see how long the sequences take to degenerate to the predetermined baselines (i.e. to become randomized with respect to USS or DUS). The score isn't a very sensitive indicator for this degeneration, as the genome may still contain an excess of the imperfectly matched cores, but I'll be able to tell this from the final analysis done at the end of the run.
After screwing up the settings many times this afternoon (e.g. specifying the N. meningitidis sequence and matrix but forgetting to change to the corresponding base composition), I realized that I could save myself a lot of wasted time by making two versions of the program, each with its own matrix and sequence files and with a settings file that specifies the appropriate genome size, base composition, and matrix and sequence files. So I did. All of the analyses I've planned will be simulating the evolution of either USS in H. influenzae intergenic sequences or DUS in N. meningitidis intergenic sequences, so now I just need to open the right folder.
Gibbs progress, and on to the Perl simulations
I discovered that I don't need to repeat the leading/lagging strand analyses after all. I had forgotten that I'd already redone the N. meningitidis ones (showing that the original surprising result was a fluke), and I decided that the H. influenzae ones I've done don't need to be repeated.
I started analysis of the DUS in the N. meningitidis coding sequences. I ran 2x100 replicates overnight on Westgrid, but they didn't find the DUS even thought they used the prior that specified its sequence. Instead they found about 7000 instances of sequences that resemble it only in containing GCCG. I think the problem may be the low density of DUS in coding sequences (~650 perfect 10-mers in ~1.74 megabases; 0.37/kb); the whole genome has ~1900 in 2.2 megabases; 0.89/kb. So I've set up a couple more runs, this time telling the program to expect only about 100 occurrences (yesterday I told it to expect 1500).
Now I'm going to try to get some Perl simulations running, after at least skimming the copious notes and data the former post-doc left me.
Today's goal: Gibbs analyses
For today, I'm not going to try to do anything about the Perl simulation work. Instead I'll just focus on the Gibbs analyses. I did find my instructions-to-myself of how to do these. I haven't heard back from the RS3 expert (apparently because his spam filter didn't like the urls in my email signature), so the first step now is to find what I've done to analyze these repeats. In particular, I have a N. meningitidis genome sequence from which the RS3 repeats have been removed; maybe I can figure out how I did this, and then test whether it makes a difference to the Gibbs results. And I also need to figure out why I seem to have been using two slightly different counts of the number of perfect 10-mer DUSs in the N. meningitidis genome.
Checking a basic technique
- Do we always use a pipettor and disposable tips, or do we sometimes use glass pipettes?
- Do we use relatively large volumes, in culture tubes, or small volumes in microfuge tubes?
- Do we always make 1-in-10 dilutions, or sometimes make 1-in-100 or other proportions?
- When we use a pipettor, do we use a fresh pipette tip every time, or do we only change tips when we think it matters? Every time we sample from a different tube? Only if the new tube has a higher concentration of cells? Only if the new tube has a lower concentration of cells? Only if the volume we need to measure changes? What about when we're using glass pipettes - do we use a fresh one every time?
- Do we pipette liquid up and up and down in the tip or pipette before removing our sample? Do we pipette liquid up and down to rinse the tip or pipette out after putting the sample into the new tube?
- Do we always plate the same volume of dilution onto each agar plate, or do we use different volumes to refine our measurements?
What the uptake sequence variation manuscript needs
........
OK. Progress.
Genome analyses needed: I need to reanalyze the Neisseria meningitidis genome with the Gibbs motif sampler, but not until I've decided whether or not to first remove the copies of the RS3 repeat. I've emailed the person who discovered them, asking him whether he thinks they are insertions or arise in situ like uptake sequences. If the former, I'll use the genome sequence that I've already removed them from. I'll do the analysis on the whole genome, and then on the strands sorted by their direction of replication. I did this before and got weird results; if the same thing happens this time I'll investigate further.
I've already done the corresponding analyses for H. influenzae, though I should probably repeat the replication-direction analysis because that was done with a slightly different dataset.
I should also analyze both genome datasets for the numbers of one-off and two-off motifs (singly and doubly mismatched); that will be easy because we have a little Perl script (somewhere) to do that now.
I should look at the effect of coding constraints by doing Gibbs searches with the coding and intergenic subsets of both genomes. But I won't split up the coding subset by the different reading frames - this is messy and not very informative.
The analysis of covariation has been done for Neisseria. I can't remember whether the H. influenzae covariation analysis was only done with the old dataset and so should be redone. The control analysis for Neisseria showed an odd pattern of weak covariation between every third position of random sequence segments. I don't think it's due to coding effects because I see the same pattern, a bit weaker, in the noncoding dataset. Maybe it's those blasted RS3 elements, so perhaps I should redo the Neisseria analysis with the RS3-deleted dataset.
The analysis of within-species variation at uptake sequences in H. influenzae is done, and there's no N. meningitidis equivalent to do.
And finally, what needs to be done with the Perl simulation of uptake sequence evolution? The few paragraphs I've found in the manuscript (written by me last fall) say that I'm going to take 200kb of intergenic sequence (or maybe all the intergenic sequence) of H. influenzae and of N. meningitidis, and find out what combinations of mutation rates and uptake bias the simulation needs to maintain their present levels of uptake sequences. Sounds straightforward, though I bet it isn't really.
back to the uptake-sequence variation project
Clean negative results
What next? The RA is also going to repeat these experiments, using exactly the same method that gave apparent transformants previously. And I'm going to streak some of my KanR colonies onto MacConkey maltose (the recipeints are all Lac-) to see if any are crp- as true transformants should be.
The new post-doc's plans
One preliminary analysis we should do is comparison of the two genomes he'll use. Both are sequenced, and it would be good to provide table and/or a figure giving specific numbers of SNPs (is it called a polymorphism when you're only comparing two individuals?), numbers and lengths of indels, and information about specific differences relevant to the proposed analysis. This could be an appendix if such are allowed, or a small table in the text.
Another thing the proposal needs is a more explicit description of the calculations that underlie its claims that the scope of the sequencing is sufficient for the information desired. He'll be creating pools of DNA from different stages and sequencing these. Depending on the source of the pool of DNA, this will require a lot of sequencing, a ton of sequencing, and what until recently would have been an absurd amount of sequencing. We can afford to do some sequencing on our present budget, and the R01 proposal is mainly to get funding for the massive sequencing. I don't understand the sequencing methods he's proposing as well as I need to, and I haven't seen any of the calculations yet. We need to lay them out in enough detail that the reviewers will have confidence in us. Perhaps we can also create some simple diagrams illustrating how the different pools will be analyzed.
Follow-up frustrations
Because she's found plating anomalies depending on cell density, I'm planning to plate a wide range of dilutions, and I'm also going to use two kanamycin concentrations - the 10 µg/ml she's already used and also 20 µg/ml.
BUT... one of the recipient strains (RR3013, the DH5alpha derivative) refuses to grow in LB with 20 µg/ml chloramphenicol (the selection for the sxy plasmid), although the other strain, which carries the same plasmid, is growing just fine. Both were inoculated at the same time, each from a plump single colony on a LB Cm10 plate. And when I tried to look at the non-growing culture under our very expensive microscope, I discovered that the high-power lens has somehow become all crudded up, and that lens cleaning solution doesn't help.
So I guess I'll just do the transformations into the other strain (RR3015). Update: my reinoculation of RR3013 crept up to the appropriate density so I included it too.
* Our strain list says this strain is also called NK6051 (constructed by Nancy Kleckner), and the E. coli Genetic Stock Center says NK6051 has its Tn10 insertion in purK, not purE. That's fine, purK is the gene next door to purE, so the original mapping was probably an error.
Negative results (no transformants)
There's no evidence of transformation of either strain by any of the markers. As expected, the C600 strain produced Lac+ revertants (frequency 10^5 - 10^6) and Leu+ revertants (frequency 10^6 - 10^7). I know these are revertants because the control cells (no DNA or no sxy insert or no inducer) produced just as many Lac+ and Leu+ colonies as the cells given DNA, sxy and inducer. The thr-1 mutation didn't revert detectably (frequency <10^8) style="font-style: italic;">lacZ gene.
I know the plates used for the selection were OK because the donor strain (w3110) grows fine on all of them but the recipients don't. This also confirms that the donor strain carries the expected wildtype alleles of lac, leu and thr.
So what next? First the RA and I need to discuss her latest results checking out the crp::kan marker she's been using for her (apparently successful) transformations. Then I'll probably do another set of transformations and controls, this time with that marker.
Waiting for the colonies to grow
Experiment plans
The first step is to inoculate strain W3110 into, say, 10 ml of LB broth and grow it overnight, and then tomorrow morning do a DNA prep. This will be the donor DNA for my transformations.
I'll be transforming two strains, C600 and BW25113. C600 is a work-horse strain from my grad-school days; genetically it is lacY (it can't take up lactose), thi (it needs the vitamin thiamine in its medium), and thr leu (needs the amino acids threonine and leucine in its medium). Its full genotype is F- tonA21 thi-1 thr-1 leuB6 lacY1 glnV44 rfbC1 fhuA1 λ-. F- and λ- means it differs from the ancestral K-12 strain in lacking the conjugative plasmid F and the lambda prophage. The tonA21 mutation removes a protein used as a receptor by phage T1, the glnV44 mutation (also called supE44) changes a glutamine tRNA gene so it inserts its glutamines at what would be UAG stop codons, rfbC catalyzes a step in synthesis of the outer-membrane sugar rhamnose, and fhuA helps cells take up ferrichrome (an iron-scavenging siderophore). I've listed all these because I want to make sure I've thought about all the factors that might confound my experiments - I see none here.
The other strain, BW25113, comes from Barry Wanner by way of the fantastic Keio group in Japan who have generated many of the E. coli clones and cassette mutants we've used. Its full genotype is F- ∆(araD-araB)567, ∆lacZ4787(::rrnB-3), lambda-, rph-1, ∆(rhaD-rhaB)568, hsdR514. So it has deletions of the araBAD operon (can't use arabinose) and the rhaBAD operon (can't take up rhamnose). rph-1 helps processing of tRNAs, and hsdR is a restriction nuclease that cuts DNA lacking a specific methylation.
What controls will I want to do? No DNA, to confirm that I'm seeing transformants and not new mutants. Cells without the inducing sxy gene (with a no-insert version of the plasmid). Cells with the sxy plasmid but without IPTG induction.
What will I select for? In BW25113 I can only select for Lac+, so I'll need minimal plates with lactose as well as minimal plates with glucose. BW25113 doesn't need any supplements. In C600 I can select for Lac+ by providing lactose as the only sugar, but the plates need to be supplemented with thiamine, threonine and leucine. I can also select for Thr+ and for Leu+ by not adding these amino acids to the medium.
So I'll need stocks of glucose and lactose (20%, I think - it's been a very long time since I did simple genetics in E. coli) and of threonine and leucine (10 mg/ml is standard for amino acids, I think). Thiamine I might as well put into all the agar. And minimal salts, and autoclaved agar.
Because I'll be growing up cells with both the inducing plasmid and a no-insert plasmid, I can also follow their growth in enough detail to see whether the insert slows growth.
No real results yet
Today's experiment
Results of UV-sensitivity tests
This suggests that our DH5alpha strains are not really RecA-. That would be consistent with the RA's results in her transformation assays, but it seems unlikely that both our stocks are not what they're supposed to be.
Another weird result is that when DH5alpha is carrying a low-copy sxy expression plasmid it becomes as UV-sensitive as NM554. But the UV-sensitivity of the Rec+ strain BW25113 isn't altered by the same plasmid. (sxy expression in these cells wasn't induced with IPTG, but that may not matter because DH5alpha is deleted for lacZYA and maybe also lacI.)
I guess I should check other components of the genotype of our DH5 alpha strains before I start emailing the RecA experts for advice. It's supposed to be F-, φ80dlacZΔM15, Δ(lacZYA-argF) U169, deoR, recA1, endA1, hsdR17(rk-, mk+), phoA, supE44, λ-, thi-1, gyrA96, relA1. Easy to check for Lac- (the RA just made some MacConkey plates), but many E. coli strains are Lac- so that's not very diagnostic. Hmmm....
What I love about doing E. coli genetics
Photo documentation
This is an iphone snap of my test plate after overnight incubation. The first and third streaks are a recA+ control strain and the second and fourth are of a known recA- strain. Because I was working at an odd angle, the green lines I drew to guide the exposures aren't lined up with the actual exposures. The recA+ cells grew fine after 5 seconds of UV, and gave a few colonies even after 1 minute (probably cells that were shielded in some way from the UV). The recA- cells grew fine when they were not exposed to any UV, but didn't grow at all even after only 5 seconds UV.
So today I've streaked all the cells I want to test, each on at least 2 different plates. I also reduced the UV dose. Yesterday I used a high dose range (5, 15, 30 and 60 seconds) to make sure I had a dose that would kill the recA- cells; today I used 2, 5 10 and 15 seconds.
Doing an experiment at last!

What's done, what's not
- Analysis of the true consensus and variation in uptake sequence motifs in all the bacterial genomes that have uptake sequences (= the family Pasteurellaceae (USS) and the genus Neisseria (DUS)).
- Analysis of variation in DUS and USS motifs across different location categories (orientation wrt replication, in coding sequences, in non-coding sequences, in terminator positions).
- Analysis of covariation between the different positions of the DUS and USS uptake sequence motifs (e.g. does having a particular base at one position correlate with having a particular base at another position).
- Additional experimental data on how variation in uptake sequence affects uptake by H. influenzae. (This will just be a paragraph as it only modestly enriches a previously published dataset.)
- Development of a computer-simulation model of uptake sequence evolution, and use of it to investigate the roles of key factors in maintaining uptake sequences in the non-coding parts of genomes.
The Perl-model manuscript (and the data) still need lots of work
- Analysis of the true consensus and variation in uptake sequence motifs in all the bacterial genomes that have uptake sequences (= the family Pasteurellaceae (USS) and the genus Neisseria (DUS)).
- Analysis of variation in DUS and USS motifs across different location categories (in coding sequences, in non-coding sequences, in terminator positions).
- Analysis of covariation between the different positions of the DUS and USS uptake sequence motifs (e.g. does having a particular base at one position correlate with having a particular base at another position).
- Additional experimental data on how variation in uptake sequence affects uptake by H. influenzae. (This will just be a paragraph as it only modestly enriches a previously published dataset.)
- Development of a computer-simulation model of uptake sequence evolution, and use of it to investigate the roles of key factors in maintaining uptake sequences in the non-coding parts of genomes.
Excavated documents
I now remember that the stuff in the pile on the floor was there for a reason. It's all sources of important ideas that I keep forgetting about - either papers I've read that tell me things I really want to remember, or notes from previous research that should someday be followed up on. Filing them would have almost the same effect as just throwing them out. This way I periodically go through the pile hoping to tidy it away but instead discovering that I need to keep these things where I'll see them now and then.
So what did I find? (Maybe if I blog about them I can decide how to use some of them?) Starting from the top of the research notes pile:
- The table of contents of part of a former tech's lab notebook, indexing the DNA uptake experiments she had done.
- A table listing the H. influenzae strains in a 'tiling-path' collection that a colleague had given us, with a map of the large-insert plasmid they're in. These are cloned in E. coli, and I think I was planning to use them to test whether H. influenzae competence genes work in E. coli (maybe make E. coli competent). I should mention these to the RA when we sit down on Tuesday to discuss the E. coli transformation experiments.
- The abstract of a paper reporting that RadC (competence-induced in S. pneumoniae, H. influenzae and E. coli) does not contribute to transformation or DNA repair in S. pneumoniae. This belongs in the pile of articles, next to one showing that RadC contributes to replication-fork stabilization (highly relevant to our ideas about what else the 'competence' regulons control).
- A very old (c. 1992?) folder containing restriction maps of some plasmids we made with the H. influenzae cya gene. I think these could be filed away.
- A table summarizing 'Next-Generation Sequencing Informatics', printed four months ago and probably already out of date. But highly relevant to the new post-doc's research and our planned NIH proposal.
- A list of the research questions we hoped to answer by microarray analysis of H. influenzae gene expression (also c. 1992). We've certainly answered most of these, but perhaps not all. Certainly I can't remember the answers to some.
- An unpublished summary of research some colleagues did into the distribution of transformability in H. influenzae strains. I think it was presented as a poster about 4 years ago. They sent us the strains, and a former post-doc included some of them in her more detailed analysis of the distrobtuion of competence and transformability (now in press in Evolution).
- Notes from my analysis of PTS genes in Pasteurellaceae, particularly the glucose and fructose transporters. Of interest because the PTS regulates cAMP which regulates competence, and the H. influenzae PTS has only the fructose transporter.
- The reviewers' comments on a manuscript that the former post-doc is revising.
- A page of notes from last summer (or the summer before last) when I was planning to test conditions that might induce E. coli sxy by assaying for expression of two lacZ fusions (to comA and ppdA). I did this and saw no induction.
One manuscript accepted, another resubmitted
NIH programs
CMYK - the saga continues
Tne Sxy in E. coli manuscript
Instead I found lots of problems, so she and I have spent much of the last 10 days revising and re-polishing the sentences, and changing parts of the text, and completely replacing the Discussion, and making the figures clearer. Finally we have a version to send to the former post-doc for final approval. We think it's now quite good, and we certainly don't want to d any more writing, so he's been strictly warned that he's not allowed to make any substantive changes to the text, except maybe to the Discussion, which still needs a final paragraph.
Here's hoping we can submit this one on Monday.
CMYK!
Unfortunately the journal wants the files in CMYK format, but we made the figures with PowerPoint, which doesn't do CMYK. I spent much of yesterday evening looking for a way to convert them that didn't involve PhotoShop. I don't want to buy PhotoShop because it's far too sophisticated and complex for our needs, and it's very expensive.
So first I Googled the problem, but didn't find any easy free solutions. Then I tried my test version of the program Acorn, which costs less than 10% of what Photoshop costs but claims to do most of the same things, only easier. But Acorn can't convert files to CMYK. Then I Googled some more, and found that Macs come with a utility application called ColorSync that claims to do this. So I spent quite a while figuring out how, and doing it, only to discover that the CMYK files became their negatives every time they were saved. Everything goes black except the text, which goes white. The colours don't exactly go black, but they become very dark.
So then I emailed the former post-doc who had mastered these file conversions, and he said he'd had the same problem with ColorSync, but had done CMYK conversions using an old copy of Photoshop on one of our computers. I found our old copy of Photoshop Elements on that computer (it came free with a scanner), but it refused to open. The RA said she has Photoshop (no, wait, it's on the home computer), and that it's also on another old computer. That was again Photoshop Elements, and it did open. But when I tried to use it on my files, there was no CMYK option, and further Googling revealed that Photoshop Elements doesn't support CMYK at all.
I'd really like to get this done today, because the journal wants the resubmission by today so they can include it in their inaugural issue. So I think I'll ask around to see if someone else has a copy of Photoshop I can use for a little while.
Should I go to the SMBE meeting?
Starting to plan proposals
What I learned at the NIH workshop
What I learned this morning: (I'm collating my scattered notes here so I won't forget them)
Sign up for the weekly NIH Announcements emails. Most of the contents are of no interest to me, but I'm a fast reader so I should be skimming them anyway.
Find a Program Officer and talk to him/her repeatedly about the science. NIH is full of very helpful people, who measure their success by whether or not your grant gets funded! (Not like CIHR.) So the first step in grant preparation is finding the program officer whose interests best match yours. I have the url for a listing of these people (no, it's just the list of Institutes; I guess I should start by calling NIAID - they handle infectious diseases). Anyway, I'll do this on Wednesday. I can also ask US colleagues who their program officer is. And when I go to US meetings I should look for the NIH people and talk to them about my plans. I'm going to the big microbiology meeting in May -- NIH should have a substantial presence there -- and to two meetings on evolution (evolution of sex and molecular evolution) in June (NIH people might be there too).
Use the cover letter to direct your proposal to the right people and places: Name the names of the NIH people who are on your side.
Relationships matter as much as good science. Again, this is all about getting to know a program officer.
Brag! This is hard for Canadians, but being modest is a big msitake here.
NIH is flush right now. Obama has given them $10 billion that must be spent by Sept 2010 (on top of their annual budget of about $30 billion). Most of this can't go outside the country, but some can, as subcontracts and foreign components of domestic grants. And it will take the pressure off of the main grant stream, so hopefully getting funded will be easier for at least the next several years.
Focus on the 'R01' grants: Don't bother applying for the little 'R03" grants ($50k/year for only two years). But it might be worth applying for an 'R21' grant ($275K over two years) - you don't need as much preliminary data as you do for the regular 5-year R01 grants because they're intended to let you generate the preliminary data.
Money is always tight, so having US collaborators is a bonus. I wouldn't be a co-investigator or sub-contractor on someone else's grant (this is our own work), but I should be able to line up a couple of solid American collaborators, or at least have letters of support from Americans who will provide help and advice if needed. But the need for this varies with the NIH program, so I need to talk to 'my program officer'.
Spell out up front why this foreign grant deserves funding. There are three criteria. 1. opportunities not available in the US. In my case, it's that nobody in the US wants to do this, and only I have the combination of evolutionary and molecular expertise to see why it's so important and to carry it through. (Of course this argument will depend on how much evolution is in the grant.) 2. Augmentation of existing US resources. Can I claim this? 3. Potential for improving the health of US taxpayers (it's their money). I can argue that understanding why and how bacteria take up DNA will give is therapeutic targets.
Commit at least 15% of your 'effort' to this project. 20% is probably better. But be prepared to back this claim up with evidence - don't claim to be putting 50% of your effort into each of several projects.
The Budget section of NIH proposals is complex: I was looking forward to the new 'modular' budgeting, where you just need to say how many $25K modules you want per year (up to $250K/year), but this doesn't apply to foreign grants.
Ask for some salary. Canadian faculty have 12-month salaries and we're not allowed to ask for any salary support from the Canadian agencies. But salary support is a standard item on NIH grants so we're not seen as very serious if we don't ask for any. The legalities of doing this are a bit uncertain - I asked my Department Head whether the money is allowed to travel from the UBC Finance account into my bank account; he's going to ask the Dean. Someone said we are allowed to get the equivalent of two months' salary - I don't know if this only applies to consulting fees and running a company on the side, or also to salary from outside grants and contracts.
Some 'indirect costs' can become direct on foreign grants. Foreign grants get only 8% as indirect costs to the institution, and these are intended to defray only the costs of administering the grant, not to cover the indirect costs of doing the research. So some expenses that in the US would be considered indirect costs to be provided by the institution (e.g. phone, office supplies, secretarial support) on foreign grants can be included in direct costs.
Ask for more money than you need: My previous NIH grant was awarded the full amount I had asked for, but now across-the-board cuts of 10% or even 20% are common. So budget in some extra. So I guess I should ask for 20% of my salary?
Don't be overambitious: Don't propose to accomplish an unreasonable amount of science. The reviewers won't think you're exceptional, just naive.
Plan for the long term: When planning your proposal, think beyond the 5-year term. What will you want to do next? How will what you are proposing now affect that? What are your long-term goals and how does each project take you closer?
Be very careful not to write anything that might turn a reviewer against you. Don't be disparaging or smart-assed. Check the membership of the study sections your proposal is likely to be sent to, and be sure to cite all their relevant work.
The font matters? One speaker recommended using Ariel 11 font. I'll have to email him to ask why.
Don't leave town right after your proposal is submitted: The NIH system scans each proposal for technical errors (wrong kinds of information in wrong boxes), and gives you only two days after submission to fix these.
It's fine to apply to both NIH and CIHR for the same project. If both succeed, NIH is happy for the aims to be readjusted and the project split up into two parts, one funded by each agency.
Starting in 2010, a rejected proposal will only be allowed one resubmission. NIH is trying to clear out the deadwood of the grants that just won't die.
Once the project is funded, the budget allocation is flexible. I used to think that NIH budgets were quite rigid, but if I later decide that I need to spend the money on something other than what was originally budgeted I can do so. If I'm changing key personnel (not just a tech or student) or if it would change the 'scope of the work' I need to get NIH approval, but that's usually just an email. However there's no allowance for currency fluctuations.
Progress...
Minor manuscript submitted!
Minor manuscript progress (major progress on a minor manuscript)
ANOVA success
I first analyzed each group of tripeptides separately (the blue ones as one dataset, then the pink, then the yellow). The blue set had significant differences between the columns in the ANOVA (p=0.01). It also had significant differences between the bright-blue column and all other columns by Tukey's multiple comparison test. I used this rather then the Bonferroni test but I'm not sure which would have been more appropriate - I think this is less sensitive than the experiment deserves, because I had specific comparisons in mind from the start. The pink set had not-quite significant differences (p=0.058) in the ANOVA, and not-significant differences between any pairs of columns in the Tukey's test. The yellow data had very significant differences between the columns in the ANOVA (p<0.0001), and significant differences between the bright-yellow column and all other columns by the Tukey's test.
I then rearranged the data, putting the bright-colour data all in the same column (the 'cognate-proteome' column), and the pale-colour data in the other columns. This let me analyze all three colours together. The ANOVA found very significant differences between the columns (p<0.0001) and the Tukey's test found significant differences between the cognate-proteome column and all the other columns.
The control comparisons ( using reversed tripeptides) were never significant.
So now I can add a sentence to the manuscript, reporting that the effects shown in Figure 1 are statistically significant.
I should have paid more attention in stats class


One of the reviewers of the manuscript I'm revising for Genome Biology and Evolution asked if we could do some statistical analysis of the data we present in a graph. On the left I've put the graphs and the data . The lower graph panel and lower block of data are the controls; we can ignore them for now. I think we can also safely ignore what the data represent.
I'll describe the significance questions with respect to the top-panel graph (A):
We want to know the following:
In the left group (4 blocks of four bars, labels SAV, TAL, KEG, PHF/L), are the four blue bars significantly higher than the red, yellow and green bars beside them?
In the middle group (4 blocks of 4 bars, labels QAV, TAC, TSG, PLV), are the four red bars significantly higher than the blue, yellow and green bars beside them?
In the right group,(5 blocks of 4 bars, labels PSE, SDG, FRR, QTA, RLN/K), are the five yellow bars significantly higher than the blue, red and green bars beside them?
The actual numbers are in the upper part of the table, in the correspondingly coloured cells, and below I'll restate the above questions in terms of these numbers.
In the top four rows of the table (blue), are the numbers in the bright-blue cells significantly higher than the numbers in the light-blue cells in the same rows?
In the next four rows of the table (pink), are the numbers in the bright-pink cells significantly higher than the numbers in the light-pink cells in the same rows?
In the next four rows of the table (yellow), are the numbers in the bright-yellow cells significantly higher than the numbers in the light-yellow cells in the same rows?
I suspect this is an ANOVA (analysis of variance) type of problem. But I'm pretty sure it would require more complicated analysis than the simple ANOVA described the new statistics textbook my author-colleague kindly gave me (probably to get me off his back with dumb statistics questions). Hmmm, maybe it would be possible to do a separate ANOVA on each group -- i.e. one for the blue data, one for the red data, and one for the yellow data.
UPDATE:
My basic version of EXCEL doesn't have the statistics add-in needed for ANOVAs, and I can't even remember the name of the statistics/graphing package the lab owns (it's not installed on my computer). But I found an on-line applet to do two-way ANOVAs here ( I need two-way because I have two variables, the rows and the columns). So I pasted the data from the blue cells into the applet, with the following results.

"Conclusion on Treatments Effects: Very strong evidence against the null hypothesis." The null hypothesis is that all treatments (columns) gave the same results, so there are very significant differences between the data in the different columns (p=0.00058).
"Conclusion on Blocks Effects: Moderate evidence against the null hypothesis." The null hypothesis is that all blocks (rows) gave the same results, so there are moderately significant differences between the data in the different rows (p=0.011).
This is definitely the kind of information I want, so I guess I should find the lab's statistical/graphing package and find someone to show me how to use it to do ANOVAs properly.
But this analysis doesn't let me see whether it's only the bright-blue column that's significantly different from the others. I guess I could repeat the analysis, leaving out the bright-blue data, and see if the others are not significantly different, but I'm sure there's a better way to do this. After I play around with our statistical/graphing package for a bit, I might be knowledgeable enough to go ask my colleague for help without embarrassing myself too badly.
...despite claiming that one of their main goals was to determine whether uptake sequences had an effect on protein and organismal fitness, the authors did not look if these sites are under purifying/diversifying selection. It would be greatly relevant for their question of interest, which is currently only supported by indirect evidence.The reviewer is absolutely right. We didn't think of doing this analysis, but we should have (though of it, not necessarily done it).
I don't think our dataset is appropriate for anything more sophisticated than simply calculating dN/dS ratios, and I'm not at all sure it's even suitable for that. I had to start by pulling out my complimentary copy of Freeman and Herron's undergraduate textbook Evolutionary Analysis, which explains how dN/dS ratios and McDonald Kreitman tests are used to examine DNA sequences for evidence of purifying or diversifying selection on the amino acids they encode. For a pair of aligned DNA sequences, dN/dS is the ratio of the number of differences that change the encoded amino acid to the number of differences that don't change the encoded amino acid. There are lots of programs and web sites that will do this analysis, given pairs of aligned seuqences in the appropriate format.
I think that my bioinformatician coauthor has DNA sequences of hundreds of H. influenzae and N. meningitidis genes, each aligned with each of three 'standard' homologs from genomes that don't have uptake sequences. These alignments have been sorted into classes, based on how many uptake sequences the H. influenzae or N. meningitidis gene has (0, 1, 2, 3, >3). I think the appropriate analysis would be to score the dN and dS ratio for each alignment, calculate the mean score of the three standard alignments of each H. influenzae or N. meningitidis gene, and then calculate the grand mean score for all the genes in each class.
This analysis isn't hard to describe, but it might be harder for my coauthor to automate, depending on the details of how the alignments are fomatted and what the dN/dS programs will accept. I'm going to email my former post-doc who has a lot of sophisticated knowledge about these methods, asking for her advice.
Should Darwin be an 'ism'?
Open access at the American Society for Microbiology annual general meeting
In May I'll be part of a panel discussion on open-access publishing, at the big General Meeting of the American Society for Microbiology. The other participants are 'professional experts': Jon Eisen, Academic Editor in Chief, PLoS; Sam Kaplan, Chair of the ASM Publications Board (ASM publishes about a dozen journals and many books); and Joe Deken of the California Institute for Telecommunications and Information Technology. I guess what I'll bring to the table is the perspective of the ordinary scientist trying to do what's right.Writing
Here's an all-too-typical example: "In E. coli, the dramatic reduction in growth and eventual cell death caused by sxy overexpression made it impossible to test whether sxy induction produces the typical ‘natural competence’ phenotype of high-efficiency transformation with linear chromosomal DNA." It's a perfectly OK sentence, no grammar or syntax errors, but it's still a bit of an effort to read. Can I improve it by replacing some of the nominalizations (reduction, overexpression, induction, transformation) with verbs?