It looks like the unique sequences we found in the 2019-nCoV coronavirus were indeed significant!
In this article in last week's Science, the authors found key differences between this virus and SARS, focused most strongly on the N-terminal domain (NTD) and receptor binding domain (RBD) regions of the viruses spike glycoprotein. This is important to understand, because this protein is what the viruses uses to actually infect cells, and also a primary target for antibodies to identify or neutralize the virus.
These regions are also right where we pointed our spotlight in our bioRxiv paper, with the surface glyoprotein region of interest that we identified! In particular, we identified the region from amino acids 9 to 275 as the largest unique sequence, and found it was part of a cluster spanning from amino acids 9 to 883. In the Science paper, the key NTD sequence goes from amino acids 17 - 305, nearly a perfect match to our largest unique sequence, and the RBD sequence goes from amino acids 330 to 521, meaning that together the two cover the majority of our identified cluster!
Now, these folks went a lot deeper than we could (not being protein modelers ourselves), and I'm sure they didn't use our research, given they were likely starting their investigation at the same time we started ours. That said, it's a nice confirmation of our methods and their potential significance to have rapidly and independently identified these regions with our FAST-NA method.
My next question for other researchers, however, is this: what about the other two domains we found?
Wednesday, February 26, 2020
Sunday, February 09, 2020
Congratulations to Cassandra Overney!
Congratulations to my former intern Cassandra Overney, who is a finalist for the National Center for Women & Information Technology (NCWIT) Collegiate Award!
Cassandra is an undergrad at Olin College who first began working for me at BBN in the summer of 2018, contributing to the NSF Expeditions “Living Computing Project” by improving our TASBE Flow Analytics software package for calibrated flow cytometry (which you may remember from a post last year). Flow cytometry is a method for measuring the fluorescence of large numbers of cells, often used as a “logic probe” for genetic engineering projects, and TASBE Flow Analytics allows precise and replicable interpretation of the results of complex experiments, and is being used in a number of laboratories and large-scale projects.
Cassandra's recognition by NCWIT is based on the critical contributions that she made for this project, most notably developing an Excel-based user interface that has proven to be much simpler and more intuitive for most of its biologist users. In developing this software, Cassandra worked closely with the biologists who would become her users, prototyping, testing, and adjusting in multiple rounds in order to provide a workflow that has significantly increased the adoption of TASBE Flow Analytics by bench scientists. Better, though, why not learn about it from the video that Cassandra made for her NCWIT award entry?
Although her internship is long over, Cassandra has continued to work part-time on this project, further improving the user interface she designed and addressing other issues as raised by users. Wearing my selfish primary investigator hat, I'd hire her full time if I could, but wearing my mentor hat, I expect both she (and science) will be better served by instead continuing to explore her interests in different areas of potential research and going off to graduate school. This is the bittersweet joy of a mentor: the better the student you work with, the faster they are likely to leave the nest!
So congratulations again, Cassandra!
Cassandra is an undergrad at Olin College who first began working for me at BBN in the summer of 2018, contributing to the NSF Expeditions “Living Computing Project” by improving our TASBE Flow Analytics software package for calibrated flow cytometry (which you may remember from a post last year). Flow cytometry is a method for measuring the fluorescence of large numbers of cells, often used as a “logic probe” for genetic engineering projects, and TASBE Flow Analytics allows precise and replicable interpretation of the results of complex experiments, and is being used in a number of laboratories and large-scale projects.
Cassandra's recognition by NCWIT is based on the critical contributions that she made for this project, most notably developing an Excel-based user interface that has proven to be much simpler and more intuitive for most of its biologist users. In developing this software, Cassandra worked closely with the biologists who would become her users, prototyping, testing, and adjusting in multiple rounds in order to provide a workflow that has significantly increased the adoption of TASBE Flow Analytics by bench scientists. Better, though, why not learn about it from the video that Cassandra made for her NCWIT award entry?
Although her internship is long over, Cassandra has continued to work part-time on this project, further improving the user interface she designed and addressing other issues as raised by users. Wearing my selfish primary investigator hat, I'd hire her full time if I could, but wearing my mentor hat, I expect both she (and science) will be better served by instead continuing to explore her interests in different areas of potential research and going off to graduate school. This is the bittersweet joy of a mentor: the better the student you work with, the faster they are likely to leave the nest!
So congratulations again, Cassandra!
Tuesday, February 04, 2020
Organizing genome engineering for the gigabase scale
Just out in Nature Communications, our new paper on "Organizing Genome Engineering for the Gigabase Scale"!
This perspective piece, a companion to the technology perspective last fall, analyzes the trends in the growing size of organisms getting their genomes re-engineered, and concludes that, while impressive, it's growing more slowly than one might think: big, complex organisms like mammals and plants are only likely to become tractable around 2050. Moreover, the complexity of the projects has been growing exponentially as well, as measured by the number of authors per paper.
We look at this problem and see not just a genome technology issue, but a massive organizational challenge as well: these projects are going to be big, and in order to manage them effectively we're going to need a lot of friction-reducing software tooling automation. The bulk of the piece is then dedicated to looking at the design/build/test cycle and analyzing the sticking points and how to address them.
Bottom line: it's not going to be simple, but it looks quite tractable, and there are things that can be done right now that will likely have a significant impact on our ability to engineer ever-larger genomes.
This perspective piece, a companion to the technology perspective last fall, analyzes the trends in the growing size of organisms getting their genomes re-engineered, and concludes that, while impressive, it's growing more slowly than one might think: big, complex organisms like mammals and plants are only likely to become tractable around 2050. Moreover, the complexity of the projects has been growing exponentially as well, as measured by the number of authors per paper.
![]() |
| The largest engineered genomes have grown exponentially, doubling approximately every 3 years (a), but the number of authors credited on projetcs has been growing exponentially as well (b). |
Bottom line: it's not going to be simple, but it looks quite tractable, and there are things that can be done right now that will likely have a significant impact on our ability to engineer ever-larger genomes.
Sunday, February 02, 2020
Unique sequences found in Wuhan coronavirus
Like many people, I have some concerns about the emerging virus in Wuhan. I am also fortunate enough to have some tools that might turn out to be helpful. For the past two years, I've been leading a project on improving pathogen screening in DNA orders by applying cybersecurity tools, and was, in fact, in the midst of writing up a paper on our improved ability to detect small virus fragments with high precision.
So it just so happens that I've got software to hand that's very good at detecting the unique aspects of a viral pathogen, and a pre-existing collection of organized coronavirus data, and it looks like we may have found something interesting---some chunks of the virus that look unlike any of its known relatives. We've written this up in a quick manuscript that's now under review and up on bioRxiv:
Highly Distinguished Amino Acid Sequences of 2019-nCoV (Wuhan Coronavirus)
Using a method for pathogen screening in DNA synthesis orders, we have identified a number of amino acid sequences that distinguish 2019-nCoV (Wuhan Coronavirus) from all other known viruses in Coronaviridae. We find three main regions of unique sequence: two in the 1ab polyprotein QHO60603.1, one in surface glycoprotein QHO60594.1.
It's also been a fascinatingly fast project: we noticed the sequence and decided to evaluate it on Tuesday morning and got our first results that afternoon. On Wednesday, we refined and confirmed the results. Thursday, we checked with others that it might be interesting, and I wrote up the quick report. Friday was polishing and submission as a research letter to CDC Emerging Infectious Diseases and a bioRxiv preprint, and then it took 48 hours for bioRxiv to post it. At just under a week from project conception to submitted preprint with DOI, this is definitely my fastest experience with scientific publication, and it's been a strange experience.
I don't know just how important this might or might not be---I am definitely not a viral pathology specialist. And maybe the journal will just laugh at us and reject it all as naive. But I'm still happy that this is out there, no matter what, in case it may indeed be useful. More than anything else, I really hope that this gets in front of people who are, in fact, the right type of expert, so that they can evaluate it and see if they can put this information to effective use in helping diagnose, prevent, and mitigate this new disease.
Friday, October 18, 2019
Gigabase-scale genome engineering
Just out in today's Science: "Technological challenges and milestones for writing genomes." One of a pair of papers I've been working on with the GP-write consortium, both of which are asking the question: what, exactly, do we need in order to go from engineering millions of base-pairs of DNA in bacteria and yeast to the billions of base-pairs in complex organisms like mammals, plants, and people?
This paper focuses on the DNA-wrangling side of the problem, while its complement (on arXiv and under revision) focuses on the informational and coordination side of the problem. Both need to be addressed, and the complexity---while daunting---is tractable. Take a read-through and see our take on the matter!
This paper focuses on the DNA-wrangling side of the problem, while its complement (on arXiv and under revision) focuses on the informational and coordination side of the problem. Both need to be addressed, and the complexity---while daunting---is tractable. Take a read-through and see our take on the matter!
Monday, October 14, 2019
Getting plate readers right
If you ever used a plate reader to measure either OD or fluorescence, you'll want to check out the iGEM 2018 interlab preprint on bioRxiv!
We just submitted this manuscript, "Robust Estimation of Bacterial Cell Count from Optical Density," for review on Friday, but we think a lot of folks will want to make use of this information, and so we've gotten a preprint up early as well. The big deal of this study is that we've now got a good calibration process for both optical density (OD) measurement, which is commonly used for estimating cell count in a sample, and fluorescence measurement, which is commonly used as a "debugging probe" for estimating cellular activity. Both of these are usually reported in relative or arbitrary units right now, which causes lots of trouble interpreting what's even going on in your experiment, as well as greatly limiting how results can be shared and applied.
No more: we have protocols that are cheap (less than $0.10/run) and easy (reliably executed by high school students just getting started in a lab), and that this manuscript shows are also both precise and accurate. All you have to do is dilute little cell-sized silica beads and fluorescent dye, plug the measurements into a spreadsheet, and you're good to go.
And here's the most important result from our paper: a nearly perfect match between per-cell fluorescence estimate from plate reader measurements and the ground truth captured from single-cell measurements in flow cytometers.
In fact, this match is even better than we deserve: we know there are factors that should distort the plate reader measurements both up and down, but they're small and appear to be canceling one another out. The only device with a notable difference in measured value is the one that's got very low fluorescence---and even there it's not significant and conforms with our expectation that flow cytometers will be better able to measure extremely faint fluorescence than plate readers.
We just submitted this manuscript, "Robust Estimation of Bacterial Cell Count from Optical Density," for review on Friday, but we think a lot of folks will want to make use of this information, and so we've gotten a preprint up early as well. The big deal of this study is that we've now got a good calibration process for both optical density (OD) measurement, which is commonly used for estimating cell count in a sample, and fluorescence measurement, which is commonly used as a "debugging probe" for estimating cellular activity. Both of these are usually reported in relative or arbitrary units right now, which causes lots of trouble interpreting what's even going on in your experiment, as well as greatly limiting how results can be shared and applied.
No more: we have protocols that are cheap (less than $0.10/run) and easy (reliably executed by high school students just getting started in a lab), and that this manuscript shows are also both precise and accurate. All you have to do is dilute little cell-sized silica beads and fluorescent dye, plug the measurements into a spreadsheet, and you're good to go.
![]() |
| Serial dilution of fluorescein (from iGEM protocols page) |
And here's the most important result from our paper: a nearly perfect match between per-cell fluorescence estimate from plate reader measurements and the ground truth captured from single-cell measurements in flow cytometers.
![]() |
| Plate reader (calibrated with microsphere dilution) vs. flow cytometry showing a 1.07-fold mean difference over 6 test devices. |
This is new science, so there's lots of caveats, of course: this has only been validated for E. coli, and probably won't work well for murky cultures with a lot of background or for biofilms or long filamentous strands. Nevertheless, it's a big step forward, since a huge amount of what people use plate readers for is covered by this study already. We'll see what the reviewers think, but I expect this paper is going to have a big impact because it's addressing a problem that so many people are encountering.
The next key challenge, however, is this: can we get somebody manufacturing plate readers to make calibration plates so that people don't have to prepare their reference materials themselves?
Friday, September 27, 2019
New aggregate programming survey!
Just out, a new survey entitled "From distributed coordination to field calculus and aggregate
One of this nice things about this survey was that we also were able to spend some time tracing out the roots of this work in the past, including a something that I really like: a diagram of all the key different traces of past work coming together to form aggregate computing (not the one above, but something much more complicated). We also spent half a dozen pages laying out our view on key problems to be addressed and the likely roadmap for near-term progress in the area. If you're interested in either making use of this work or getting involved in research in this area yourself, this paper is a great place to start reading!
computing", which surveys aggregate programming work by my collaborators and myself. This paper expands on a conference version published last year, and gives a nice overview of how all of the different pieces of our work in this area fit together.
![]() |
| How the past, present, and future fit together in our view of aggregate programming. |
Tuesday, September 10, 2019
Damn you, asparagine!
Deep inside big public databases, you can find quite curious things, especially when biology is involved.
For example, I spent several hours today hunting down a mysterious bug in the DNA screening project that I've been leading. We're working on improving the ability to detect when somebody orders DNA that they shouldn't be ordering (e.g., smallpox, ebola), and so it's really important to not let anything get past. So while most classification projects might be fine with getting nearly everything right, our system has to catch every single problematic sequence every time.
That means I get to drill down and try to classify every miss our system makes, and I learn some strange and interesting things while doing it. For instance, these pseudo-fascinating trivia are amongst the things that I have recently learned:
With these discoveries and a few other tweaks, I was able to categorize and plan mitigations covering all of the classes of failures that our system was encountering. Almost.
There was just one miss that I just could not explain, a short little snippet from a virus coat. There were no related "safe" viruses that would cause us to overlook its sequence, nothing in the protein sequences and nothing that could even be mis-translated from other DNA sequences. And I thought, "that's funny..."
I dug down and dug down and eventually found something both embarrassing and wonderful. You see, in DNA sequences, there's often parts that are unknown, and so instead of the standard "A", "C", "T", and "G" DNA bases, these bits of missing information get marked as "N" for an unknown "any" base. These get used in ordering DNA too, to indicate places where you don't care what the sequence is. We've long been excluding these from matches, since it makes no sense to say, "Aha! Somebody once didn't know part of a virus, and you don't care what you get!" So our detector throws out potential matches that include an unknown.
Only thing is, when you're working with proteins, the missing information letter isn't "N". There are a lot more amino acids than nucleic acids, and so they use up more of the alphabet, including "N", which stands for the amino acid asparagine. With proteins, the missing information letter is "X" instead.
Most of our system knew that. Most of our system was doing the right thing. But one little part of one little script wasn't getting switched into protein mode at the right time.
We've been systematically excluding every protein pathogen signature with asparagine in it.
That's embarrassing. Easy to fix, but still embarrassing.
And yet...
Asparagine is a pretty common amino acid, so we've been accidentally throwing away around one third of the detection power of our system. And out of tens of thousands of tests, there was precisely one where this blatant and egregious error caused us to miss a detection.
The wonderful thing is that the system is still working almost perfectly, even while we've been unknowingly arbitrarily throwing away a vast amount of its ability to detect pathogens. That speaks to its resilience, and how many alternative routes it explores to achieve its goal. I can live with that, with a nice natural experiment accidentally conducted by a misbehaving script. We'll fix it, and move forward.
But such remarkable things you may find when you follow just one little thread of something funny in your data...
For example, I spent several hours today hunting down a mysterious bug in the DNA screening project that I've been leading. We're working on improving the ability to detect when somebody orders DNA that they shouldn't be ordering (e.g., smallpox, ebola), and so it's really important to not let anything get past. So while most classification projects might be fine with getting nearly everything right, our system has to catch every single problematic sequence every time.
That means I get to drill down and try to classify every miss our system makes, and I learn some strange and interesting things while doing it. For instance, these pseudo-fascinating trivia are amongst the things that I have recently learned:
- The same DNA sequence from the same publication is often uploaded twice and categorized differently each time.
- Fish in fish farms get sick with a virus related to rabies. It doesn't hurt humans, though.
- Somebody is running automated systems to infer the organisms that DNA sequences are associated with, and that produces a lot of "unknown member of [family/order]" entries.
- Somebody published a paper where they claimed to discover a bunch of new virus species by just sort of sequencing samples from healthy people and not actually checking in any way whether actual viruses were involved.
- When NCBI updates its taxonomy which organisms are related to which, the sequence records don't change to reflect their new taxonomy.
With these discoveries and a few other tweaks, I was able to categorize and plan mitigations covering all of the classes of failures that our system was encountering. Almost.
There was just one miss that I just could not explain, a short little snippet from a virus coat. There were no related "safe" viruses that would cause us to overlook its sequence, nothing in the protein sequences and nothing that could even be mis-translated from other DNA sequences. And I thought, "that's funny..."
I dug down and dug down and eventually found something both embarrassing and wonderful. You see, in DNA sequences, there's often parts that are unknown, and so instead of the standard "A", "C", "T", and "G" DNA bases, these bits of missing information get marked as "N" for an unknown "any" base. These get used in ordering DNA too, to indicate places where you don't care what the sequence is. We've long been excluding these from matches, since it makes no sense to say, "Aha! Somebody once didn't know part of a virus, and you don't care what you get!" So our detector throws out potential matches that include an unknown.
Only thing is, when you're working with proteins, the missing information letter isn't "N". There are a lot more amino acids than nucleic acids, and so they use up more of the alphabet, including "N", which stands for the amino acid asparagine. With proteins, the missing information letter is "X" instead.
Most of our system knew that. Most of our system was doing the right thing. But one little part of one little script wasn't getting switched into protein mode at the right time.
We've been systematically excluding every protein pathogen signature with asparagine in it.
![]() |
| Our system: "Damn you, asparagine! Get out of my house!" |
That's embarrassing. Easy to fix, but still embarrassing.
And yet...
Asparagine is a pretty common amino acid, so we've been accidentally throwing away around one third of the detection power of our system. And out of tens of thousands of tests, there was precisely one where this blatant and egregious error caused us to miss a detection.
The wonderful thing is that the system is still working almost perfectly, even while we've been unknowingly arbitrarily throwing away a vast amount of its ability to detect pathogens. That speaks to its resilience, and how many alternative routes it explores to achieve its goal. I can live with that, with a nice natural experiment accidentally conducted by a misbehaving script. We'll fix it, and move forward.
But such remarkable things you may find when you follow just one little thread of something funny in your data...
Sunday, August 11, 2019
Can we put an end to secret parenting?
Recently, one of my colleagues at BBN shared an article about "secret parenting," and the concept really struck a chord with me. The basic idea is that people often feel that they will be judged for choosing parenting over putting in more hours at the office, and so they end up hiding these choices, making excuses, and generally having their work-life balance (or lack thereof) degraded further.
It's unfortunately easy to simply brush away one's parenting, to pretend that it's not happening, to pretend it's not important. And it's not just parenting, of course: people have all sorts of other things outside of work. Parenting, however, is something that's particularly strong and gendered in its impact in American society, at least.
In my group at BBN, I think we do pretty well on not hiding our parenting. The group mailing list is always abuzz with notifications of people saying they're going to be out or working from home for personal or family reasons: taking the kids to the doctor, dealing with child-care failure, going to see a kid's baseball game, helping out with the grand-kids, fixing an air conditioner, keeping their new dog company, etc. Also, importantly, I see it coming very much from both men and women. I think that this visibility on the mailing list is really important, because it makes it much more comfortable to make those choices oneself, and to feel less pressure to engage in secret parenting. I definitely know that it matters for me.
With other colleagues outside of my home organization, however, I often do not feel such comfort. Whenever I make a choice that's driven by my desire to be a present and responsible parent (or other personal things, though parenting dominates in my life right now), I feel that I have to worry about things like:
This shows up in lots of little micro choices. Like, do I tell people I can't make it because I'm volunteering to drive for a field-trip at my daughter's school, or just say that I have a conflict? Do I say that I'm heading for the airport early because I want to see my kids in the morning, or just blame it on flight combinations to Iowa?
As I get to know somebody better, the barriers can come down, but in the world of science there are always new collaborators, new potential competitors, new program managers. I don't feel secure enough to expose myself in that way with people that I do not know well. And if I don't, as somebody who should probably be considered well established at this point in my career, how much more vulnerable my younger colleagues, my colleagues who are female or minorities?
On this blog, on my online persona, you get to see the highlights of my life. You don't get to see my times of burnout and depression. You don't get to see me struggle with imposter syndrome. I'm still not going to post these here, in full public record, for all to see, because I do not want to make myself that vulnerable to judgement. But dear reader, I would encourage you to count the posts that are not there.
Writing posts like these is a good sign, for me, because it's showing that I'm finding time enough to sit down and reflect and find the things I want to share. Posts show me operating at peak functionality in my life, and if I'm operating at Peak Jake, I'd probably post just about once a week. Thus, if you don't see a post, it doesn't necessarily mean that things are bad in my life---but it means that I don't feel I have the luxury to indulge in these delightful pseudo-conversations. Not without neglecting things that are more important to me, at least, like parenting and career.
But I do think that in my professional interactions, I'm going to try to shift my boundaries a bit more, indulge my trust a bit more freely in my outside-of-BBN colleagues. I don't like hiding my life from my work, parenting or otherwise, and since I am indeed in a somewhat secure and privileged position in my career, I think that one of my responsibilities is to help to shape my professional environment to be more of the sort in which I would like to live and work.
And with that, dear reader, let me sign off by informing you that this post appears in the midst of a two week vacation. My older daughter is between school and camp, and I've decided that I should spend that time with her, prioritizing parenting over work for at least a little while. I just hope that I don't pay too much for this choice in the state of my email and my projects at the time when I return.
It's unfortunately easy to simply brush away one's parenting, to pretend that it's not happening, to pretend it's not important. And it's not just parenting, of course: people have all sorts of other things outside of work. Parenting, however, is something that's particularly strong and gendered in its impact in American society, at least.
In my group at BBN, I think we do pretty well on not hiding our parenting. The group mailing list is always abuzz with notifications of people saying they're going to be out or working from home for personal or family reasons: taking the kids to the doctor, dealing with child-care failure, going to see a kid's baseball game, helping out with the grand-kids, fixing an air conditioner, keeping their new dog company, etc. Also, importantly, I see it coming very much from both men and women. I think that this visibility on the mailing list is really important, because it makes it much more comfortable to make those choices oneself, and to feel less pressure to engage in secret parenting. I definitely know that it matters for me.
With other colleagues outside of my home organization, however, I often do not feel such comfort. Whenever I make a choice that's driven by my desire to be a present and responsible parent (or other personal things, though parenting dominates in my life right now), I feel that I have to worry about things like:
- Will this person think less of me professionally?
- Will they worry I'm not sufficiently committed?
- Will they feel like I'm putting them at a lower priority?
This shows up in lots of little micro choices. Like, do I tell people I can't make it because I'm volunteering to drive for a field-trip at my daughter's school, or just say that I have a conflict? Do I say that I'm heading for the airport early because I want to see my kids in the morning, or just blame it on flight combinations to Iowa?
As I get to know somebody better, the barriers can come down, but in the world of science there are always new collaborators, new potential competitors, new program managers. I don't feel secure enough to expose myself in that way with people that I do not know well. And if I don't, as somebody who should probably be considered well established at this point in my career, how much more vulnerable my younger colleagues, my colleagues who are female or minorities?
On this blog, on my online persona, you get to see the highlights of my life. You don't get to see my times of burnout and depression. You don't get to see me struggle with imposter syndrome. I'm still not going to post these here, in full public record, for all to see, because I do not want to make myself that vulnerable to judgement. But dear reader, I would encourage you to count the posts that are not there.
Writing posts like these is a good sign, for me, because it's showing that I'm finding time enough to sit down and reflect and find the things I want to share. Posts show me operating at peak functionality in my life, and if I'm operating at Peak Jake, I'd probably post just about once a week. Thus, if you don't see a post, it doesn't necessarily mean that things are bad in my life---but it means that I don't feel I have the luxury to indulge in these delightful pseudo-conversations. Not without neglecting things that are more important to me, at least, like parenting and career.
But I do think that in my professional interactions, I'm going to try to shift my boundaries a bit more, indulge my trust a bit more freely in my outside-of-BBN colleagues. I don't like hiding my life from my work, parenting or otherwise, and since I am indeed in a somewhat secure and privileged position in my career, I think that one of my responsibilities is to help to shape my professional environment to be more of the sort in which I would like to live and work.
And with that, dear reader, let me sign off by informing you that this post appears in the midst of a two week vacation. My older daughter is between school and camp, and I've decided that I should spend that time with her, prioritizing parenting over work for at least a little while. I just hope that I don't pay too much for this choice in the state of my email and my projects at the time when I return.
Sunday, August 04, 2019
Two Maxims of Project Management
I hold these two maxims of project management to be unwaveringly true:
These two maxims come to me through long and painful experience, which I'd like to pass on to you, in hopes that your learning process will be less long and less painful.
The first maxim, "if it's not in the repository, it doesn't exist," is something that I first learned in writing LARPs but is just as true in scientific projects or any other form of collaboration. For any project I am working with people on, I always, always set up some sort of shared storage repository, whether it be DropBox, Google Drive, git, subversion, etc. If something matters, it needs to be in that repository, because if it isn't, there are oh-so-many ways for it to get accidentally deleted.
More importantly, however, anything in the repository can be seen by other people on the team, which means there's some accountability for its content. I can't count the number of times that somebody has said they're working on something, but it's just not checked in yet, and then it turns out that they weren't working on it at all, or they were working on it but it was terrible and wrong. Some of the worst experiences of my professional life, like nearly-quit-your-job level of painful, have involved somebody I was counting on failing me in this way. If somebody's reluctant to put their work in the team repository, well, that's a pretty good hint that they are embarrassed by it in some way, and thus that their work might as well not exist.
Share your work with your team. Even if it's "messy" and "not ready," insulate yourself from disaster and give people evidence that you are on the right track---or a chance to help you and correct you if you aren't.
The second maxim, "If it's not running under continuous integration, it's broken," appears on the surface to be more specific to software. Continuous integration is a type of software testing infrastructure, where on a regular basis a copy of your system gets checked out of the repository (see Maxim #1), and a batch of tests are run to see if it's still working or not. Typically, continuous integration gets run both every time something changes in the repository and also nightly (because something might have changed the external systems it depends on).
This makes a lot of sense to do for software, because software is complicated. When you improve one thing, it's easy to accidentally break another as a side effect. Building tests as you go is a way to make sure that you don't accidentally break anything (at least not anything you're testing for). If you don't test, it's a good bet that you will break things and not know it. Likewise, the environment is always changing too, as other people improve their software and hardware, so code tends to "rot" if left untouched and untested over time. So if you don't test, you won't know when it breaks, and if you don't automate the testing, you won't remember to run the tests, and then everything will break and it will be a pain.
Surprisingly, I find that this applies not just to software, but to pretty much anything where there's a chance to make a mistake and a chance to check your work. Whenever I analyze data, for example, I always make sure that I automate the calculation so that I can easily re-run the analysis from scratch---and then I add "idiot checks" that give me numbers and graphs that I can look at to make sure that the analysis is actually working properly. Things often go wrong, even in routine experiments and analyses, and if I put these tests in, then I can notice when things go wrong and re-run the analysis to make it right. I fear that I annoy my collaborators with these checks, sometimes, because they find embarrassing problems, but I'd much rather have a little bit of friction than a retraction due to easily avoidable mistakes in our interpretation of our experiments.
Even my personal finances use tests. In my spreadsheets, I always include check-sums that add things up two different ways so that I can make sure that they match. Otherwise I'm going to make some little cut-and-paste error or typo and then have some sort of unpleasant surprise when I figure out I've got ten thousand dollars less than I thought I did or something like that.
Check your work, and check it more than one way, and add a little bit of automation so that the checks run even when you don't think about them. It takes a bit of extra time and thought, and it's easy to neglect it because it's hard to measure disasters that don't happen. I promise you, though, investing in testing is worth it for the bigger mistakes that you'll avoid making and the crises that you'll avoid creating.
- If it's not in the repository, it doesn't exist.
- If it's not running under continuous integration, it's broken.
These two maxims come to me through long and painful experience, which I'd like to pass on to you, in hopes that your learning process will be less long and less painful.
If it's not in the repository, it doesn't exist
The first maxim, "if it's not in the repository, it doesn't exist," is something that I first learned in writing LARPs but is just as true in scientific projects or any other form of collaboration. For any project I am working with people on, I always, always set up some sort of shared storage repository, whether it be DropBox, Google Drive, git, subversion, etc. If something matters, it needs to be in that repository, because if it isn't, there are oh-so-many ways for it to get accidentally deleted.
More importantly, however, anything in the repository can be seen by other people on the team, which means there's some accountability for its content. I can't count the number of times that somebody has said they're working on something, but it's just not checked in yet, and then it turns out that they weren't working on it at all, or they were working on it but it was terrible and wrong. Some of the worst experiences of my professional life, like nearly-quit-your-job level of painful, have involved somebody I was counting on failing me in this way. If somebody's reluctant to put their work in the team repository, well, that's a pretty good hint that they are embarrassed by it in some way, and thus that their work might as well not exist.
Share your work with your team. Even if it's "messy" and "not ready," insulate yourself from disaster and give people evidence that you are on the right track---or a chance to help you and correct you if you aren't.
If it's not running under continuous integration, it's broken.
The second maxim, "If it's not running under continuous integration, it's broken," appears on the surface to be more specific to software. Continuous integration is a type of software testing infrastructure, where on a regular basis a copy of your system gets checked out of the repository (see Maxim #1), and a batch of tests are run to see if it's still working or not. Typically, continuous integration gets run both every time something changes in the repository and also nightly (because something might have changed the external systems it depends on).
This makes a lot of sense to do for software, because software is complicated. When you improve one thing, it's easy to accidentally break another as a side effect. Building tests as you go is a way to make sure that you don't accidentally break anything (at least not anything you're testing for). If you don't test, it's a good bet that you will break things and not know it. Likewise, the environment is always changing too, as other people improve their software and hardware, so code tends to "rot" if left untouched and untested over time. So if you don't test, you won't know when it breaks, and if you don't automate the testing, you won't remember to run the tests, and then everything will break and it will be a pain.
Surprisingly, I find that this applies not just to software, but to pretty much anything where there's a chance to make a mistake and a chance to check your work. Whenever I analyze data, for example, I always make sure that I automate the calculation so that I can easily re-run the analysis from scratch---and then I add "idiot checks" that give me numbers and graphs that I can look at to make sure that the analysis is actually working properly. Things often go wrong, even in routine experiments and analyses, and if I put these tests in, then I can notice when things go wrong and re-run the analysis to make it right. I fear that I annoy my collaborators with these checks, sometimes, because they find embarrassing problems, but I'd much rather have a little bit of friction than a retraction due to easily avoidable mistakes in our interpretation of our experiments.
Even my personal finances use tests. In my spreadsheets, I always include check-sums that add things up two different ways so that I can make sure that they match. Otherwise I'm going to make some little cut-and-paste error or typo and then have some sort of unpleasant surprise when I figure out I've got ten thousand dollars less than I thought I did or something like that.
Check your work, and check it more than one way, and add a little bit of automation so that the checks run even when you don't think about them. It takes a bit of extra time and thought, and it's easy to neglect it because it's hard to measure disasters that don't happen. I promise you, though, investing in testing is worth it for the bigger mistakes that you'll avoid making and the crises that you'll avoid creating.
Subscribe to:
Posts (Atom)







