Monday, May 06, 2024

Does your biosecurity screening work?

So, you've decided you don't want to help randos get ahold of smallpox, ebola, and ricin. That's a good start! So, you've decided to check if the DNA, RNA, or protein sequences that people are asking you to build or work with are dangerous, and you've either built your own biosecurity screening system or you've obtained a biosecurity screening system from us or one of the other screening tool providers. Great!

Does it work?

How would you even know?

That's the question that our new paper, Progress and Prospects for a Nucleic Acid Screening Test Set, is beginning to answer. 

We organized a collaboration of four tool providers and two synthesis companies to start with the simplest question of all: on what sequences do we agree if they're dangerous? Even that question wasn't straightforward, because all of the tools have qualitatively different ways of reporting their analysis and some approach the question of sequence danger as "innocent until proven guilty" while others are instead "guilty until proven innocent."

Still, we were able to come up with a way of making their results all sufficiently comparable, and we did a first test on three controlled threat organism groups - one viral, one bacterial, and one fungal:
  • The good news is that we all basically agreed on the viral sequences.
  • The bad news is that we couldn't agree on about 10% of the bacterial sequences.
  • The challenging news is that there wasn't enough information to decide one way or another for more than 30% of the bacterial sequences and more than 80% of the fungal sequences.


That last bit, about not being able to decide, is not actually as bad as it might sound: if nobody knows enough about a sequence to decide whether it's dangerous, then it's also not likely that anybody knows enough to actually do something bad with it. And being able to agree large numbers of sequences is a very good thing if you want to be able to do basic "competence tests" to make sure that a tool isn't making bad decisions.

This might sound very down in the weeds, but being able to answer this question is actually a big deal, and an important part of making the biosecurity screening framework just put out by the White House actually work. So now over the next few months we're scaling up our "bronze standard" test set effort to cover all of the regulated pathogens and toxins out there, and collaborating with EBRC and NIST to make sure what we're doing can be used to benefit the whole community.

So much of civilization depends on little details of measurement and standards... I just hope that we can work quickly and effectively enough to help ward off the threats that are coming over the next few years.

Tuesday, November 14, 2023

How do you describe genetic construction plans?

ACS Synthetic Biology has just published "Standardized Representation of Parts and Assembly for Build Planning", our new article on how to better communicate about building genetic constructs. The paper is basically a more friendly user manual for the best practice that we wrote up last year.

Fundamentally, this is all just about trying to reduce the confusion that commonly occurs when we're talking about build plans. If somebody shares a sequence, is it for the bit they want synthesized, what a vector will look like after the synthesized bit gets stuck in, what gets digested out of the vector, or what it looks like as part of the final construct after it gets ligated together with other constructs? 

When we were collaborating on building the new iGEM distribution, we ran into a lot of confusion amongst the many different participants along these lines, so we worked out a standard vocabulary for describing what we were talking about, with intuitive names for different stages in typical digestion/ligation assembly processes.


And once we humans were clear on what we wanted to say to one another, it was easy enough to take the next step and use SBOL3 to make a simple description to describe it to the machines as well, including the exact reactions one would want to run to actually execute the plan. This is one of the nice things about SBOL, which you can't do with formats like GenBank, FASTA, or GFF: describe not just a construct, but its relationship with other constructs and your whole plan for how to use it.

We're still using this vocabulary quite extensively in the iGEM Engineering Committee, as well as using the representations in our software, and we hope that others will find it useful for clarifying their discussions as well.

Monday, April 03, 2023

BLAST vs. custom tools for pathogen identification

Our analysis of issues with using BLAST vs. NCBI for pathogen identification is out today: “Studying pathogens degrades BLAST-based pathogen identification.”  This paper is the full published version of the preprint I posted about a few months ago, investigating an emergent dynamic, in which biological research and development ends up contaminating public databases with chimeric material that can confound biosecurity systems that trust those databases.

The most important addition between the preprint and this final version was to make direct head-to-head comparisons between BLAST vs. NCBI and two tools specifically designed for biosecurity analysis, our own FAST-NA Scanner and a free tool called SeqScreen (there are other tools we'd like to have compared with as well, but they were not available for comparison). 

As predicted, the actual biosecurity tools completely dominated over BLAST vs. NCBI, making more than an order of magnitude less mistakes---not a surprise, but nice to see experimentally validated. In fact, each biosecurity tool only made one mistake in judgement, and in both cases it was the same mistake that NCBI did, which is an important lesson: the big NCBI databases aren't bad, they're just dirty, and so they just need a lot of care and refinement when they're being put to a use (like biosecurity determinations) where mistakes can be costly and dangerous.

This is important for biosecurity, but I also think people need to be aware of this in the larger scientific world as well. In biology, curation quality really matters, and many people are far too blasé about the potential impact of dirty data on their applications. If you want to do biosecurity right, you need to use an actual biosecurity tool and not just trust the databases. I'm sure the same applies for many aspects of medicine, diagnostics, etc., and I fear that not enough people are taking these issues seriously.

Wednesday, July 27, 2022

Multicolor Plate Reader Fluorescence Calibration

Just out in OUP Synthetic Biology, "Multicolor Plate Reader Fluorescence Calibration" extends our prior work on calibrating green fluorescence and cell count to calibrate red and blue fluorescence as well. The results are no surprise (if we can use a green dye, we ought to be able to use other dyes too), but it's valuable to have specific recommendations for dyes to use and to have an interlab study validate that yes, they really do perform as well as the others. 

So everybody out there listening, please start using sulforhodamine-101 to calibrate your red fluorescence and Cascade Blue to calibrate your blue fluorescence! Everybody who uses your data will thank you for providing equivalent molecule/cell estimates rather than irreproductible arbitrary or relative units.

Red and blue fluorescence calibrants were just as precise as the prior green and cell-count calibrants 

The paper also reports on some of the travails we ran into making the study work: some of the fluorescent proteins we wanted to try out didn't work in our hands, and there were miscellaneous other problems: a promoter sequence got messed up,  some things wouldn't synthesize, one of the plasmids seemed problematic, and timing problems meant not all labs could run all constructs.

Problems like that are frustrating, but ultimately I'm happier reporting them than burying them. Remember: if you read a synthetic biology study with lab work and it doesn't talk about failures, it just means they either aren't aware of them or else they've pruned them from the narrative!  Calibration methods like these help us see better when things go wrong and understand what's happened.


Thursday, July 14, 2022

Studying Pathogens Degrades BLAST-based Pathogen Identification

Using the BLAST algorithm to search the NCBI databases is the typical way one goes about identifying a DNA sequence, so it's been the typical way biosecurity systems decide if something is potentially a dangerous pathogen or toxin too. Problem is, that's not what BLAST and those databases were designed for, and we've observed that they aren't working as well for that purpose as they used to, as we report in our new preprint: "Studying Pathogens Degrades BLAST-based Pathogen Identification"

Specifically, we've found an inherent problem that is growing in seriousness due to a non-obvious emergent dynamic. Now that sequencing and bioengineering tools are getting much more accessible, lots of sequences are being studied by modifying them with "tool" sequences like purification tags, fluorescent proteins, stabilizing sequences, etc. Those sequences get (appropriately) classified based on what's being studied, and now you've got chimeric material that includes both the subject of study and the bioengineering tool. Then when you run BLAST on a sequence with that tool, you start finding that tools are classified as what they're used to study.

Example of BLAST classification failure: using a purification tag to study an Ebola protein means that now a fluorescent protein plus a purification tag gets mis-identified as Ebola.

This doesn't seem to be much of a problem for most uses of BLAST against NCBI, but it's poisonous for making biosecurity decisions, since it can cause benign sequences to be classified as dangerous or vice versa. Moreover, the effect gets stronger the more problematic a pathogen is (since more sequences are recorded) and the more useful a tool is (since more chimeric material is produced), meaning that the problem is most likely to occur in the most important.  For example, over the last two years, quite a lot of stuff has started coming back as COVID-19, since everybody in the world is studying COVID-19 with all of the tools that they can get their hands on.

This is a serious problem, and it's not likely to get better, since NCBI and BLAST aren't doing the wrong thing: they're just getting less suitable to use as a short-cut for doing something that they were never designed to do. 

So how do we fix it? Switch to tools that are actually designed for pathogen identification. We've got one (FAST-NA Scanner), and a whole bunch of other folks worked on the same problem in the FunGCAT program. The solutions are there, we just have to help folks switch to them.

Wednesday, July 13, 2022

pySBOL3: SBOL3 for Python Programmers

Our Python library for the SBOL3 standard now has an official citable publication in ACS Synthetic Biology, called "pySBOL3: SBOL3 for Python Programmers." 

The article is a good short read, but for any Python programmers, out there I recommend just jumping straight in with the tutorial instead. Happy hacking, everyone!



Tuesday, July 05, 2022

Functional Synthetic Biology

Synthetic biology isn't about sequences. Don't agree? Tell me what this is without looking it up: atgcgtaaaggagaagaacttttcactggagttgtcccaattcttgttga

Tell you what, I'll give you a hint, make it easy. It's a coding sequence translating to MRKGEELFTGVVPILV. Everybody knows this one, right?

How about this instead?


That's right. That mystery sequence up top is the first 50 bases of BBa_E0040, the widely used iGEM part with a coding sequence for GFPmut3. Now that one, a great many folks working in synthetic biology know, have used in their work, and maybe even have strong opinions about.

Notice that this is a description of biological function: the important thing is that the coding sequence makes a protein that emits a lot of green light when you hit it with a blue laser. There's a sequence in there somewhere but that's not what gets put on the whiteboard or what gets discussed.

Don't get me wrong, sequences are important. But right now we're living with a mis-match in synthetic biology, where most of our discussions about design are about function, but nearly all of our tooling is heavily focused on sequences (e.g., GenBank format), with any information about function tacked on as an afterthought or else confined to specialized databases that each pose their own sui generis integration problem. 

We need a new focus on functional synthetic biology, and that's one of the things we've been working on in the iGEM Engineering Committee. We're trying to change how we do synthetic biology, so that we can pull together the work that lots of people have been doing on calibration, insulation, characterization, context effects, modeling, assembly, etc., in one place and make at least a small class of synthetic biology engineering really simple and predictable.

We aren't there yet, but we've gotten to the point where we think we've figured out some of the important shifts in thinking, representation, and tooling that need to happen in order to make functional synthetic biology possible. If you're interested in this too, I encourage you to read more in our newly available pre-print on Functional Synthetic Biology.

Thursday, May 05, 2022

AI for Synthetic Biology

Several of my colleagues have been organizing an series of "AI for SynBio" workshops over the last few years. I've been to some and they have been both stimulating and enjoyable. Now they have an article out in Communications of the ACM, along with a nice short video in which Aaron Adler introduces this increasingly important cross-disciplinary interaction for folks who aren't familiar with one or both of the subjects.

Friday, April 22, 2022

Talking measurement and standards with "The Living Revolution"

Yesterday I had an enjoyable conversation with Luke Roche and Sara Knurowska, who do a podcast called "The Living Revolution." They'd read some of my work on measurement, which led inevitably to a wide-ranging discussion including fundamental principles in engineering and science, when to standardize (or not), SBOL, etc.

Check out the podcast here (if it works for in your browser), or on Spotify or Apple Podcasts 


Friday, January 07, 2022

Two years of soap

Back in pre-pandemic times, I used to travel quite a lot, and like many other frequent travelers, I slowly accumulated a pile of little bars of complimentary soap from hotel rooms.  As a result, I hadn't actually purchased soap for myself for years. Today, however, I opened my last little leftover travel soap. A curious milestone and statistic: it appears that I'd had just under two years of soap in my little pile.

One of my daughter's stuffed animals traveling with me on my last pre-pandemic trip.

Monday, October 11, 2021

Meeting Measurement Precision Requirements for Effective Engineering of Genetic Regulatory Networks

We've got a new preprint up today, "Meeting Measurement Precision Requirements for Effective Engineering of Genetic Regulatory Networks", that is an unusual mixture of theoretical analysis and interlaboratory study. 

The work started out as an investigation of the replicability of flow cytometry measurements. Flow cytometry, as readers of this blog may know, is one of my favorite biological measurement tools, since it lets us obtain measurements from large numbers of individual cells. I've been involved in a number of projects that have put it to good use in engineering biological devices, and the calibration methods available let us put real, biologically-sensible units on the measurements. But just how good are these measurements and how reproducible?  That's what we set out to study with a consortium of collaborators and about two dozen flow cytometers.

Then we went to go write it up, and a rabbit hole opened beneath our feet, sucking us down into an unexpected set of theoretical questions. We had a number (~1.5-fold precision), but was that a good number? In fact, how do we even decide what a good number is? What do you even need to do good engineering? 

Maybe we should have just called that "future work" and published what we had. But we didn't. We followed that rabbit hole down and the manuscript went into limbo. But when it came out of limbo, the manuscript was standing on its head and had an answer. What started as an investigation of flow cytometry became an investigation of the general requirements for effective biological engineering, with the work on flow cytometry becoming one verified answer for how to meet those requirements.

Basically, you want to be on the left side of the red line.

We ended up with a (highly abstract, conservative) formula for estimating how well one needs to know values in order to engineer gene regulation. And for most state of the art work, it means you need to have a measurement precision somewhere in the range of 1.2-fold to 2.0-fold, with calibrated flow cytometry right smack in the middle.

I'm happy with these dual results, and I think they should be useful to help us move another couple of steps towards a world of reliable and predictable biological engineering.

Thursday, July 15, 2021

Predictable signal amplification with recombinases

New paper out today: "Quantitative characterization of recombinase-based digitizer circuits enables predictable amplification of biological signals." If we ever want to be able to make reliable controller in cells, we need to have well-separated control signals. Many of the biological sensors and other inputs that we work with, however, are really blurry, so we need devices that can clean them up. This paper demonstrates how this can be done in mammalian cells with a circuit that cleans up a poorly separated signal by nearly 3 decibels!

Blurry input (left) is predicted to be separated well by our recombinase device (middle), and that prediction is realized experimentally (right).

This work, part of the NSF Living Computing Project,  involved collaboration across several labs and a lot of work to connect the devices, analytics, and models. Making this work meant really getting down into what we wanted not just biologically but computationally, in terms of the signal properties of the device. The models and metrics guided adjustments in device design that ultimately feed back into a better performing system. I'm personally very happy with the result as an example of a getting really serious about the engineering approach biological systems.

Monday, June 07, 2021

From reproducibility failure to methodological success

Out today in PLOS ONE, "Comparative analysis of three studies measuring fluorescence from engineered bacterial genetic constructs" solves a mystery hiding in the iGEM interlaboratory studies for 2016, 2017, and 2018. You see, the publication of the 2017 interlab data was delayed, even after the publication of the iGEM 2016 study and iGEM 2018 study, because of a troubling mystery: the plate reader results from the 2016 and 2017 studies did not match.  This was a shock, because the 2017 study was intended to be a replication of the 2016 study, plus a few extensions and enhancements. But what we got was shockingly different, systematically off by a factor of more than 10. So which, if either of them, was right?

This is a terrible and unsettling place to find oneself in, but we couldn't actually answer the question until after we had run and analyzed the 2018 study. With that study, we finally had a way to put plate reader data on the same scale as flow cytometry data, so that we could assess accuracy through two independent measurements. So once we'd finally finished analyzing and publishing that data, we turned to comparing the three years to find out what had happened and how to understanding our failure to reproduce. And here is the story, finally, summed up in a single image:


It appears the 2017 plate reader results were right: they match both the 2018 results as well as the flow cytometry from 2016. There's a lot more detail in the paper, as well as additional confirmations, but the bottom line is that it looks like the calibrant that we prepared for the 2016 study did not have the concentration of fluorescein that it was intended to. 

Embarrassing, but actually, I think, good news in the end. Because we could tell! We are no longer held to the tyranny of uncertainty, unable to even know if our measurements have been reproduced. With multiple independent measures and a successful confirmation of values reproduced in three different studies (2016 flow, 2017 plate, 2018 both), we now have truly solid ground on which to stand, biologically. Every future study that we build can bootstrap off of these results, and know if the numbers that come out are reasonable or not.

But why are we still preparing our own fluorescent calibrants in the first place? We need metrological traceability and easily purchased commercial preparations with adequate quality control, just like we have for units of time and length. Calling all reagent suppliers: who will first start to sell a plate reader cellular quantification kit?

Monday, May 31, 2021

Analysis and Visualization of Gene Expression Data

A couple of days ago, I gave a seminar on analysis and visualization of gene expression data for After iGEM, which was recorded and made freely available online. The first half of the talk is focused on core issues on data analysis, covering unit calibration, use of geometric statistics, process controls, and relating measurements to biology. The second half is about how to make a good figure, applying lessons from my favorite instructor in the area, Edward Tufte, that are likely useful to anyone and everyone who makes a figure ever. For those interested, I'm embedding the video below, and have posted the slides on my website.

Monday, April 05, 2021

Sharing our ignorance

One of the both wonderful and challenging things about working in a highly interdisciplinary area like synthetic biology is that all of us who work there are painfully ignorant. 

No matter how much of an expert one is in some areas, there is simply too much complexity and too many things to know to allow one to be an expert in all of the relevant aspects of the field. Even an apparently simple task like measuring fluorescence from simple genetic constructs often contains quite a number of rabbit holes of complexity that one can go down. 

Working in a field like this, it's easy to feel insecure about how much one doesn't know. But ignorance can be a gift as well, providing an outside perspective and shedding light on unexamined assumptions. Moreover, what is collaboration if not constructive use of complementary ignorance? Indeed, this is what Joy's law is all about: tackling complex challenges is effectively impossible for "lone geniuses" and always involves expertise dispersed among many different people.

This is a key part of what we are trying to address with the Synthetic Biology StackExchange proposal. Knowledge flows slowly and noisily through person-to-person networking, but much more quickly through well-curated community Q&A like StackExchange supports. Instead of one person getting their question answered through oral tradition, we all get an answer that's confirmed by many peer reviewers and made easy to find for the next several hundred people who need to know.

All we need now to make this happen is another few dozen people to support the proposal and then come ask three good questions on one of the existing sites like Biology.SE or Bioinformatics.SE (thus hitting the "people able to use StackExchange" criteria for launch).  I've really been enjoying this myself, asking questions about simple laboratory information that's outside of my experience, like how hard it is to pipette right and the shelf-life of frozen bacteria, and receiving interesting and informative answers.

Come join us today and make a gift of your ignorance!

Friday, March 12, 2021

One year of lockdown

One year ago today, I was in England, nervously finishing up a standards meeting and hoping that I could make safely home without either getting infected with Covid or getting stuck on the wrong side of the border. 

On March 13th, 2020, I flew home via Chicago, on the last day before the border closed. My wife and I embraced, and we put our household into lockdown. We waited nervously for a week of potential incubation time, but I had apparently escaped infection. Spring break began, and we wondered if there would still be school at the end of it.

We were fortunate. I had noticed the potential trouble building and we had at least a month's supplies laid in our basement for our household. My wife and I could both keep doing our jobs online, and despite some friction with the kids and cats, we settled into an enclosed routine. I miss being able to get together with friends in person, but between my work and family, my life is full and over-full with social interactions, and in the evenings I usually just want a quiet place of solace. 

Some things, I am surprised that I do not miss, like restaurants. We've gotten much better at cooking, and the food we eat is healthier. I've lost fifteen pounds or so, thanks to my healthier lifestyle. More time with the kids is a silver lining too. But what a world of change we've been through, all of us, and it's not over yet.

Tomorrow will be the anniversary, one full year since the last time our house was open to the world. One full year of living our pandemic lives. 

There's a light at the end of the tunnel now, I think, but we're a long way yet from done. Our indoor cats perch on the windowsill, looking at the world outside denied to them to roam, and I think I can identify.



Thursday, March 11, 2021

Building the SynBio StackExchange community

We're getting closer to launching the Synthetic Biology StackExchange Q&A site, but building a new community on StackExchange is hard, and we need more people to come help!

There are more than 150 different topic sites in the StackExchange network, and over the course of launching them, StackExchange has learned a lot about what causes a community to succeed or fail. Most important, it seems, is having critical mass at the start, and they've set the thresholds for site launch accordingly.

What this means for SynBio StackExchange is that we'll launch as soon as we have:

  • 250 people committing to use the site when it launches, and
  • at least 100 of those having 200+ reputation on some other StackExchange site.

I've believe that both of these are quite reasonable and achievable. The 250 people threshold is pretty straightforward, and we're on track to hit that around the end of March. The second criteria, however, is harder since most SynBio folks aren't yet contributing to StackExchange, and will be the one that determines when we can actually launch the site.

Getting 200 reputation is pretty easy: it only takes about 4 questions or 3 answers. It's a big deal, though, since it means you've put a little skin in the game and figured out how to contribute on StackExchange. And once you learn to play the game, it's often pretty fun: helping people feels good, getting your own problems solved is great, and StackExchange is set up to really reward people who put in just a few minutes on a regular basis.

And so it is with building this community. Come join us, and bring one friend. Ask another one tomorrow, and ask your friend to do the same. We're doing pre-launch Q&A using the Biology StackExchange synthetic biology tag, so come ask one question there. Then ask another one tomorrow. And if you want more, ask me for an invitation to the community-building Slack where we're sharing tips and helping each other.

One step at a time, we're working to build this community, and when it's big enough, we'll get to have a dedicated watering hole where everybody who needs help in synthetic biology knows that they can come and find it.

Thursday, February 25, 2021

Getting close to a SynBio StackExchange!

Synthetic Biology StackExchange is one step closer to launching! We've now passed the definition phase for the site, and as soon as we have a critical mass of people committing to participate in the beta, the site will be officially launched. 

I'm very excited about getting this site going, since I think it will be an extremely valuable resource for many thousands of folks working in the area. I know even as an expert, I'm quite naive about a lot of details in the laboratory and issues outside of my own areas of expertise, and really look forward to asking as well as answering questions.

If you're interested too, please go to the site and commit to the beta today!


Monday, February 08, 2021

Come help define a Synthetic Biology StackExchange!

In the actual practice of working in synthetic biology, there are so many pragmatic details that don't get captured well in scientific papers. Right now, they are basically passed around by word of mouth, through oral tradition amongst students in a lab or hallway conversations at conferences. But there is a better way.

Pretty much everybody who programs makes use of StackOverflow, which provides well-curated answers for programming questions. The greater network of StackExchange sites spun off of it provides a great one-stop-shop for lots of other communities as well, from math, physics, and chemistry to travel, cooking, and personal finance. The iGEM Engineering Committee, as part of its educational mission, is trying to do the same for synthetic biology

Synthetic biology field is rapidly growing, highly cross-disciplinary (which means none of us can be experts at everything), and we think it could use a good, universal database of questions and answers. It will be good for students learning the field, and also good for professionals who need to know things outside of their personal expertise.  

Right now, we're in the "definition" phase at StackExchange, and need about 100 more people to add example questions and vote for example questions they like. Just follow this link or click on the imagebelow, make an account, and you can start asking and voting too!