Tuesday, July 05, 2022

Functional Synthetic Biology

Synthetic biology isn't about sequences. Don't agree? Tell me what this is without looking it up: atgcgtaaaggagaagaacttttcactggagttgtcccaattcttgttga

Tell you what, I'll give you a hint, make it easy. It's a coding sequence translating to MRKGEELFTGVVPILV. Everybody knows this one, right?

How about this instead?


That's right. That mystery sequence up top is the first 50 bases of BBa_E0040, the widely used iGEM part with a coding sequence for GFPmut3. Now that one, a great many folks working in synthetic biology know, have used in their work, and maybe even have strong opinions about.

Notice that this is a description of biological function: the important thing is that the coding sequence makes a protein that emits a lot of green light when you hit it with a blue laser. There's a sequence in there somewhere but that's not what gets put on the whiteboard or what gets discussed.

Don't get me wrong, sequences are important. But right now we're living with a mis-match in synthetic biology, where most of our discussions about design are about function, but nearly all of our tooling is heavily focused on sequences (e.g., GenBank format), with any information about function tacked on as an afterthought or else confined to specialized databases that each pose their own sui generis integration problem. 

We need a new focus on functional synthetic biology, and that's one of the things we've been working on in the iGEM Engineering Committee. We're trying to change how we do synthetic biology, so that we can pull together the work that lots of people have been doing on calibration, insulation, characterization, context effects, modeling, assembly, etc., in one place and make at least a small class of synthetic biology engineering really simple and predictable.

We aren't there yet, but we've gotten to the point where we think we've figured out some of the important shifts in thinking, representation, and tooling that need to happen in order to make functional synthetic biology possible. If you're interested in this too, I encourage you to read more in our newly available pre-print on Functional Synthetic Biology.

Thursday, May 05, 2022

AI for Synthetic Biology

Several of my colleagues have been organizing an series of "AI for SynBio" workshops over the last few years. I've been to some and they have been both stimulating and enjoyable. Now they have an article out in Communications of the ACM, along with a nice short video in which Aaron Adler introduces this increasingly important cross-disciplinary interaction for folks who aren't familiar with one or both of the subjects.

Friday, April 22, 2022

Talking measurement and standards with "The Living Revolution"

Yesterday I had an enjoyable conversation with Luke Roche and Sara Knurowska, who do a podcast called "The Living Revolution." They'd read some of my work on measurement, which led inevitably to a wide-ranging discussion including fundamental principles in engineering and science, when to standardize (or not), SBOL, etc.

Check out the podcast here (if it works for in your browser), or on Spotify or Apple Podcasts 


Friday, January 07, 2022

Two years of soap

Back in pre-pandemic times, I used to travel quite a lot, and like many other frequent travelers, I slowly accumulated a pile of little bars of complimentary soap from hotel rooms.  As a result, I hadn't actually purchased soap for myself for years. Today, however, I opened my last little leftover travel soap. A curious milestone and statistic: it appears that I'd had just under two years of soap in my little pile.

One of my daughter's stuffed animals traveling with me on my last pre-pandemic trip.

Monday, October 11, 2021

Meeting Measurement Precision Requirements for Effective Engineering of Genetic Regulatory Networks

We've got a new preprint up today, "Meeting Measurement Precision Requirements for Effective Engineering of Genetic Regulatory Networks", that is an unusual mixture of theoretical analysis and interlaboratory study. 

The work started out as an investigation of the replicability of flow cytometry measurements. Flow cytometry, as readers of this blog may know, is one of my favorite biological measurement tools, since it lets us obtain measurements from large numbers of individual cells. I've been involved in a number of projects that have put it to good use in engineering biological devices, and the calibration methods available let us put real, biologically-sensible units on the measurements. But just how good are these measurements and how reproducible?  That's what we set out to study with a consortium of collaborators and about two dozen flow cytometers.

Then we went to go write it up, and a rabbit hole opened beneath our feet, sucking us down into an unexpected set of theoretical questions. We had a number (~1.5-fold precision), but was that a good number? In fact, how do we even decide what a good number is? What do you even need to do good engineering? 

Maybe we should have just called that "future work" and published what we had. But we didn't. We followed that rabbit hole down and the manuscript went into limbo. But when it came out of limbo, the manuscript was standing on its head and had an answer. What started as an investigation of flow cytometry became an investigation of the general requirements for effective biological engineering, with the work on flow cytometry becoming one verified answer for how to meet those requirements.

Basically, you want to be on the left side of the red line.

We ended up with a (highly abstract, conservative) formula for estimating how well one needs to know values in order to engineer gene regulation. And for most state of the art work, it means you need to have a measurement precision somewhere in the range of 1.2-fold to 2.0-fold, with calibrated flow cytometry right smack in the middle.

I'm happy with these dual results, and I think they should be useful to help us move another couple of steps towards a world of reliable and predictable biological engineering.

Thursday, July 15, 2021

Predictable signal amplification with recombinases

New paper out today: "Quantitative characterization of recombinase-based digitizer circuits enables predictable amplification of biological signals." If we ever want to be able to make reliable controller in cells, we need to have well-separated control signals. Many of the biological sensors and other inputs that we work with, however, are really blurry, so we need devices that can clean them up. This paper demonstrates how this can be done in mammalian cells with a circuit that cleans up a poorly separated signal by nearly 3 decibels!

Blurry input (left) is predicted to be separated well by our recombinase device (middle), and that prediction is realized experimentally (right).

This work, part of the NSF Living Computing Project,  involved collaboration across several labs and a lot of work to connect the devices, analytics, and models. Making this work meant really getting down into what we wanted not just biologically but computationally, in terms of the signal properties of the device. The models and metrics guided adjustments in device design that ultimately feed back into a better performing system. I'm personally very happy with the result as an example of a getting really serious about the engineering approach biological systems.

Monday, June 07, 2021

From reproducibility failure to methodological success

Out today in PLOS ONE, "Comparative analysis of three studies measuring fluorescence from engineered bacterial genetic constructs" solves a mystery hiding in the iGEM interlaboratory studies for 2016, 2017, and 2018. You see, the publication of the 2017 interlab data was delayed, even after the publication of the iGEM 2016 study and iGEM 2018 study, because of a troubling mystery: the plate reader results from the 2016 and 2017 studies did not match.  This was a shock, because the 2017 study was intended to be a replication of the 2016 study, plus a few extensions and enhancements. But what we got was shockingly different, systematically off by a factor of more than 10. So which, if either of them, was right?

This is a terrible and unsettling place to find oneself in, but we couldn't actually answer the question until after we had run and analyzed the 2018 study. With that study, we finally had a way to put plate reader data on the same scale as flow cytometry data, so that we could assess accuracy through two independent measurements. So once we'd finally finished analyzing and publishing that data, we turned to comparing the three years to find out what had happened and how to understanding our failure to reproduce. And here is the story, finally, summed up in a single image:


It appears the 2017 plate reader results were right: they match both the 2018 results as well as the flow cytometry from 2016. There's a lot more detail in the paper, as well as additional confirmations, but the bottom line is that it looks like the calibrant that we prepared for the 2016 study did not have the concentration of fluorescein that it was intended to. 

Embarrassing, but actually, I think, good news in the end. Because we could tell! We are no longer held to the tyranny of uncertainty, unable to even know if our measurements have been reproduced. With multiple independent measures and a successful confirmation of values reproduced in three different studies (2016 flow, 2017 plate, 2018 both), we now have truly solid ground on which to stand, biologically. Every future study that we build can bootstrap off of these results, and know if the numbers that come out are reasonable or not.

But why are we still preparing our own fluorescent calibrants in the first place? We need metrological traceability and easily purchased commercial preparations with adequate quality control, just like we have for units of time and length. Calling all reagent suppliers: who will first start to sell a plate reader cellular quantification kit?

Monday, May 31, 2021

Analysis and Visualization of Gene Expression Data

A couple of days ago, I gave a seminar on analysis and visualization of gene expression data for After iGEM, which was recorded and made freely available online. The first half of the talk is focused on core issues on data analysis, covering unit calibration, use of geometric statistics, process controls, and relating measurements to biology. The second half is about how to make a good figure, applying lessons from my favorite instructor in the area, Edward Tufte, that are likely useful to anyone and everyone who makes a figure ever. For those interested, I'm embedding the video below, and have posted the slides on my website.

Monday, April 05, 2021

Sharing our ignorance

One of the both wonderful and challenging things about working in a highly interdisciplinary area like synthetic biology is that all of us who work there are painfully ignorant. 

No matter how much of an expert one is in some areas, there is simply too much complexity and too many things to know to allow one to be an expert in all of the relevant aspects of the field. Even an apparently simple task like measuring fluorescence from simple genetic constructs often contains quite a number of rabbit holes of complexity that one can go down. 

Working in a field like this, it's easy to feel insecure about how much one doesn't know. But ignorance can be a gift as well, providing an outside perspective and shedding light on unexamined assumptions. Moreover, what is collaboration if not constructive use of complementary ignorance? Indeed, this is what Joy's law is all about: tackling complex challenges is effectively impossible for "lone geniuses" and always involves expertise dispersed among many different people.

This is a key part of what we are trying to address with the Synthetic Biology StackExchange proposal. Knowledge flows slowly and noisily through person-to-person networking, but much more quickly through well-curated community Q&A like StackExchange supports. Instead of one person getting their question answered through oral tradition, we all get an answer that's confirmed by many peer reviewers and made easy to find for the next several hundred people who need to know.

All we need now to make this happen is another few dozen people to support the proposal and then come ask three good questions on one of the existing sites like Biology.SE or Bioinformatics.SE (thus hitting the "people able to use StackExchange" criteria for launch).  I've really been enjoying this myself, asking questions about simple laboratory information that's outside of my experience, like how hard it is to pipette right and the shelf-life of frozen bacteria, and receiving interesting and informative answers.

Come join us today and make a gift of your ignorance!