Saturday, January 23, 2016

How fresh are your bananas? An exploration of functionality in biological computing

A recent conversation got me thinking about how to explain quantification of "function" in biological circuits.  This is a critical issue and one of the key inhibitors to engineering complex biological systems, but it's not easy to explain because it involves a lot of technical electrical engineering / computer science concepts like signal-to-noise ratio, input/output transfer curves, non-linear amplification and threshold matching.  I think, however, that there may be nice biological metaphor that can make this concept easier to understand.

You see, a biological computing device is like a banana.

Let's say that I want something to eat, so I go into my kitchen and find a banana.  Is my banana "functional" as food? Well, it's very important for food to be fresh and healthy, so I'd better make sure my banana is fresh. But just how fresh is "fresh enough"?  Let us consider some different ways that my banana might look:
A spectrum of banana freshness (credits: yellowspottedbrownrotted)
The banana on the left is beautiful: no question that it's fresh, and indeed so perfect that I am sure that my three-year-old daughter would eat it without the slightest protest.

The second banana is getting on in age and starting to develop spots.  It might be a bit mushy inside, and my daughter will definitely not eat it, but I'd be happy enough to chow down.

The third banana probably won't taste good to eat on its own, but it's just perfect for making banana bread or other recipes that transform a banana from centerpiece to simply tasty flavoring.

And as for the fourth... I don't think I'd even want to feed that melted mess to livestock.

As you can probably see, the notion of a "fresh banana" is not a fixed concept, but a spectrum, and "fresh enough" depends entirely on what exactly we want to do with that banana.

Biological computing devices are the same way: a device has to be very high-performance in multiple dimensions (uniform yellow) to be safe and useful for complex circuits or applications like precision medical therapy, while simpler and less safety-critical circuits can tolerate some problems (spotted yellow), some applications just need a nudge or two in the right direction (brown), and some devices probably aren't good for anything at all (rotten).

Right now, the vast majority of our available biological devices are metaphorical brown bananas, with just a few spotted bananas available.  Understanding that fact is the first step, and getting on with building some fresher bananas is the second.

Monday, January 18, 2016

A week on the computer science / synthetic biology interface

Back in August, I spent a week in Seattle playing several different roles at the interface between computer science and synthetic biology.  It was all built around a conference that I've been involved with and participating in for a number of years now, the International Workshop on Bio-Design Automation (IWBDA) (amusingly, given my now geographic location, it has only recently displaced the Iowa Wholesale Beer Distributors Association as the top Google hit for that acronym).

This past year, I was the Publication Chair for IWBDA, which meant I ran herd on the paper submission and peer review process, making sure that we actually get clear judgement, as unbiased as possible, of which submissions are strong scientific contributions worthy of putting on stage as talks. In the end, I think we had a quite strong program, with some very exciting results from a lot of good groups (I gave a talk of my own as well, on circuit design using signal to noise ratio, which I will leave to others to judge), and I'm hoping the associated special journal issue will come out strongly as well.

In addition to IWBDA, two other events attached meetings, taking advantage of their overlap in interests with the IWBDA community.  Just before IWBDA (and, in fact, part of its "pre-conference" schedule) was the SBOL community meeting, aiming to disseminate information and support adoption of our data exchange standards for biological designs, as well as to get more input from more different groups into its development.  In that meeting, in my role as an editor of SBOL---one of the community's elected leadership---I presented some material and helped to facilitate discussion and organize plans for the community.  

Before that was a two-day meeting of the SemiSynBio Roadmap project, an effort sponsored by the semiconductor industry, which has gotten keenly interested in synthetic biology as a possible direction of expansion as Moore's law winds down, and is trying to identify the key directions of research that can set up a similar exponential expansion of capabilities and markets in its relationship with the biological world.  It's fascinating, unclear whether it will turn out meaningful or merely hopeful, and I sit on the Executive Committee of this project, trying to help ensure that we end up with a clear and productive vision out of the several working groups studying different aspects of that interface, an invited position that doesn't fit cleanly into any of my well-defined job responsibilities and yet is clearly a good use of my time and effort as a scientist.

In between, in the corners of my time, I pursued yet other pieces of my scientific life on the interface, including standards development with NIST, the 2015 iGEM interlab study, and various relationships and collaborations with other interesting characters who I enjoy and who live in similarly strange niches to myself.

Interesting things happen at interfaces, both in physics and in society, and scientific communities are no different.  It's an uncomfortable and delicate place to stand, when you're not really at the heart of any of the communities that you're trying to participate in and affect, but it also feels like home to me.

Wednesday, January 13, 2016

An Introduction to SBOL Visual

Following up on our recent publication on SBOL Visual, an important next step is to make it nice and easy for anybody and everybody who wants to illustrate a genetic construct to do so using SBOL Visual.

To that end, I've now posted "Introduction to SBOL Visual," a short set of slides intended to give all you need to know about making genetic construct diagrams in one simple and easy to digest package.

Please share, enjoy, and send feedback on adjustments that you think would improve the document!

Thursday, January 07, 2016

Beginning an ambitious (NSF) Expedition: Evolvable Living Computing

Today the official announcement is out and I can finally talk openly about what I've known for nearly two months now: we are beginning a major new synthetic biology project in partnership with Boston University and MIT, led by Doug Densmore over at BU.  "Evolvable Living Computing" is an NSF Expeditions in Computing project, a class of large-scale long-term basic research project that puts $10 million dollars into focused research on a deep challenge.

In the case of our project, that challenge is the computational side of synthetic biology.  Our goal, over the next few years, is to create the intellectual and scientific foundations for truly general computation in biological organisms.  Computation has long been an important goal in synthetic biology, dating back at least to the 1997 "Cellular Gate Technology" paper by Tom Knight and Gerry Sussman.  This vision has been elusive, however, partly because much of the funding available has been heavily focused on applications rather than foundations, and partly because we are only now beginning to overcome barriers to effective design design methods and develop metrics that accurately assess the computational power of biological devices.

In this project, we will be tackling those issues head-on, directly tackling the questions of performance, limitations, and scope of various biological computation models.  BBN's anticipated role in the project is foundations for the foundations, and I am excited to be funded on these works:
I expect that this will be an exciting several years indeed, and hope to build off this foundation in many new directions and collaborations.

Saturday, January 02, 2016

Science is an endless sequence of paths not taken

In the turning of the year, I have been going over old records, catching up on the neglected aspects of my scientific, family, and personal life.  One of the things I've spent time doing in this process was going over the past several years of my calendar, trips, and other records as I updated the list of invited talks I've given.  Yes, it's a bit of a mindless thing to do, and rather dry on record-keeping, but walking through all those dates and records produced an interesting set of observations as a side effect for me.

Month by month, year by year, my calendar is filled with paths not taken.  Collaboration discussions that were pleasant and interesting, but ultimately went nowhere.  Pilot projects begun, carried through to a useful starting point, but then never moving forward past that first beginning. Discussions about funding, white-papers and proposals written, most never leading to a funded project.  And yet, I did not feel a failure: woven around and within this set of paths not taken were the strains and threads that have indeed succeeded, grown and prospered and become the basis for the work I now am doing.

Years previous, I certainly did feel like a failure, when I was first starting out on my own to seek funding and hadn't landed any yet.  From the time that I finished my postdoc and joined BBN, my first time to seek funding truly under my own name and not with a professor serving as my PI and responsible adult supervision, it took nearly two years to land my first funding and more than two years to land the first piece of funding with myself as a primary investigator.  Two years of talking with program managers and pitching ideas, writing white-papers and proposals, building collaborations and having them shot down.  When I got those first contracts, I totaled up all of the attempts I'd made before and counted 11 misses before the first hit.  Some of those were cheap misses---a few emails, a couple of hours of research and preparation, a day in Washington to find a missed connection---while others were quite expensive, working with several others to put many weeks of effort into a big proposal that ultimately got turned down.  My record since, I think that I would count as little better, though I hope that I have improved my efficiency in terms of cost per try, if not on number.

Every one of those misses is a path not taken, a might-have-been that usually will never be.  A new and different twist building off my core work, to connect it to a particular opportunity, a particular set of interests of a funder and a place and time in where the science stands. Complementarily, every path that is taken bends the arc of research somewhat: the main themes at the core remain the same, but different applications build infrastructure and results and expertise in different directions, exercise different relationships and strengthen different collaborations, opening up new paths that might have never been able to exist before.

In my CV, I keep a list of projects that I have been funded on.  I do not keep a list of paths not taken. But perhaps I should.  It would be fascinating to know more about the might-have-beens of other scientists as well.  If you spend time talking to a scientist, you'll often start to hear about their paths not taken, especially if you are discussing some sort of possible work together or a new proposal: "We did a little bit of that, but then we ended up going in another direction...", "We started on that, but then the project ended...", "We wrote a proposal to do that, but it wasn't funded...", "We tried that once before and it didn't work, but now I think that the technology is better..."

I don't think that we teach this aspect of the life in graduate school enough.  The myths of science speak of only the core of the research, and only in retrospect, when the moral of the tale is clear. But I think false starts and paths not taken are the true tale, an endless sequence of paths not taken: so long as you are a creative and intelligent researcher, you will generate ideas and possibilities faster than you can find time to adequately pursue and to persuade others to sufficiently support them.  It's hard to learn to not take the failures-to-launch personally, however, to not suffer heartbreak from each glorious potential you describe that never comes to pass.  And yet, of course, to succeed you must take the failures seriously, and learn from them, and build your work yet stronger on the pieces that succeed.

Science is an endless sequence of paths not taken, and survival as a scientist means learning to delight in the paths that you can take and to let go of the paths that ultimately turn out not to be available.

Monday, December 28, 2015

Why aren't research grants centralized?

A recent question on Academia StackExchange asked something that looked simple to me at first, but turned out to be much deeper and more subtle that I had expected as I thought about it more.  The question is, in essence: "Why aren't research grants centralized?"  

In other words, why do countries generally have messy and complicated research funding systems like we do in the United States, where there are a bazillion different independent agencies and mechanisms for funding research, each with its own peculiar mechanisms and application rules? Wouldn't it make more sense to have some sort of unified science system where all of the science people can go and ask to get their science things funded?  

I enjoyed thinking about and exploring this question, and so I wish to share my answer with you here as well:

First, let us consider why there are many organizations that fund research, rather than a single research-funding organization. This is a matter of evolutionary organizational structure. In most countries, research has a non-trivial budget and applies to many different concerns of government. That means there has to be some (probably largely hierarchical) structure for organizing it. Now, let's consider two prototypical organizational structures for government-funded research. First, we might have a general research agency, which contains subdivisions addressing the research needs of various other governmental tasks: 

Alternatively, each government department might have its own research agency: 

Almost everywhere, we see organizations more like the second structure than the first---there might well be some countries in the world where research is so small or so controlled that is it organized in the first way, but if so, I am not aware of them. Why might that be?

Consider what happens if you are a leader in the department of agriculture, and you want to expand your agency's research work. Unless strong regulation prevents you from doing so, it's much easier to create or expand a research organization within the agriculture department than it is to get an independent research department to do it for you. A research sub-department within agriculture is also more likely to serve the peculiar needs, time scale, market structure, etc. as relates to agriculture. It's also easier and more rewarding to go to government leadership and fight to get resources for your own organization, where you can explain exactly how you plan to utilize them, than to fight to give them to somebody else.

Since both government structure and research needs evolve over time, we may thus expect research organizations to multiply, both across the government as a whole and also within individual sub-organizations. They are, in fact, occasionally reorganized and combined with the goal of making them simpler and more efficient to interact with, just as other government agencies are, but that will typically not reduce the number down to one, just to a smaller "many." Moreover, we've only discussed government funding, not industry funding or funding by foundations and NGOs, which all have their own separate needs and desires and further complicate the funding landscape.

Now, to the second aspect of the question: why is there no central database for applications? Sometimes there are, at least partially. For example, in the United States all government requests for proposals go through FedBizOpps. Most research solicitations can thus be found there (though not all, due to the diversity of mechanisms), along with requests for things like security guards for the US Embassy in Costa Rica. As you might guess, however, the sheer breadth means this often isn't a terribly efficient method of searching.

Likewise, every agency has different sorts of information it's looking for in research proposals. Again, taking the US as an example, the NSF really wants to know how its funds will support graduate student and postdoc education, since that's a key part of its mandate. AFRL, on the other hand, usually doesn't care much about supporting students, and has a mandate instead focusing on how its funds will affect current military concerns. As a result, a "universal" proposal would likely be quite cumbersome even if the bureaucracies were somehow reconciled.

Bottom line: "research" is too complex and pervasive a set of needs to readily stay contained within a single unified organization.

Sunday, December 06, 2015

A Publication Sea Change?

Now, in the closing of the year, is a time to start taking stock of my scientific progress over the last twelve months.  What in my professional life is going well, what things need care and focus, what things have changed on me a bit at a time without me noticing all of the accumulation?

As I've been looking back over my publications of the last year, I've noticed something that is unexpected: apparently, 2015 is the year my publications moved to journals.  During graduate school, I barely published in journals at all, being both a less mature writer and in the deepest depths of computer science workshop/conference culture.  By the time I graduated and came to BBN, my work was starting to mature and I began putting extended versions of my conference papers in journals, to a tune of about two journal papers per year, still only a fraction of my scholastic output.

This year is is different: this year I have twelve journal articles stamped with an official publication date of 2015 and three more that have appeared in "online early" editions.

In large part, this reflects the growth of the synthetic biology side of my research.  Synthetic biologists typically publish in journals rather than conferences, and so what might have been conference publications in computer science go to journals instead in synthetic biology.  I've been working seriously in synthetic biology for several years now, but my collaborations have been growing and maturing, and some of those articles reflect projects multiple years in the works that have had long and hard roads to publication.  There are also some that in practice were published online last year, but have only this year been officially assigned to a theoretical paper issue that no-one really reads that way any more.

Five of my publications, however, are from the aggregate programming / spatial computing side of my world, and that also reflects a major increase in activity.  Here, though, the time to publication is often a very much longer road indeed.  I have noticed that computer science journals are often much more comfortable with lengthy times in review and revision than biology journals are.  Once of my articles that's just come out, for example, was first submitted in February of 2014; another article was submitted in March of 2014 and should appear in mid-2016.  I think that this may be because of the conference culture in computer science: a journal can afford to be quite slow and dozy in review because the editors assume that everyone has already got access to the prior version of the paper, and the journal issue will simply be the extended remix.  I do not know if that's the case, but my experience has certainly showed a stark difference in the urgency that attends each culture's publications.

In fact, then, my surge of journal publications does not actually reflect a surge of writing in this year, but rather a more gradual increase over the last few years, first picking up on my computer science side, then rising on the biology side as well.  The slower computer science and faster biology waves then happen to coincide in 2015, creating this prominent spike in journal publications.

In fact, the rate at which I am writing publications does not appear to have changed all that much.  The actual numbers of publications that I am an author on that have been initiated in these last few years is:

  • 2012: 17 publications initiated
  • 2013: 17 publications initiated
  • 2014: 16 publications initiated
  • 2015: 20 publications initiated

The quality and intensity of those publications has risen, though, as has the degree of collaboration, which also no doubt leads to more publications per unit effort on my part.

So, what does this all mean?  In short: this means that I seem to be saying things that others are interested in scientifically, and working with more people to say more things more clearly, and overall I think that's good.

Saturday, December 05, 2015

SBOL Visual

We have just published what I believe is a very important paper, "SBOL Visual: A Graphical Language for Genetic Designs."  You can read it online for free, from PLOS Biology.  This is the culmination of a long and slow process (as most standards work tends to be) of looking at the different ways that people make diagrams explaining genetic designs and trying to boil it all down into a simple common language for communicating.

Diagram languages for communicating designs are practically universal, in any area of human endeavor where we need to talk about making complicated things.  Whether you're building electronic circuits or houses, writing software or maintaining a sewer system, sewing from patterns or folding origami, there are standard ways of drawing diagrams in order to communicate ideas and minimize confusion.  So of course we need them for engineering biological organisms as well, and thus, SBOL Visual.

The basic idea is quite simple, and can be captured in a simple image of genetic constructs organized along a DNA or RNA sequence "backbone":

The current set of icons covers a lot of the constructs that people engineer, though by no means all.  If you've got another thing to put on a diagram, you can use any icon you want, as long as it doesn't conflict with an existing SBOL Visual icon. To let SBOL Visual expand and become more universal, however, there's an open community process for adding more icons, with a number of icons slowly working their way through the process.
Current SBOL visual icons
Finally, although we provide "standard" icons, there is actually a great deal of flexibility in how you can style them, which makes it easy to use these icons in anything from scribbling on a whiteboard to computerized design software to figures in scientific publications.
All of these diagrams follow the SBOL visual standard.
In some ways, this is a very simple thing.  It is, however, extremely important to get these simple thing right, in order to reduce the amount of friction, frustration, and mistakes we make when we work together and communicate about the things we do.  SBOL Visual is an important step in getting that "less wrong" in the engineering of biology, and I'm glad that it's officially become published now as well.

Tuesday, November 24, 2015

Paying down organizational debt

When I first learned about the concept of technical debt (from a post on the excellent Coding Horror blog), it was like a light going on in a dark room where I had been tripping on things for years. Technical debt (also known as coding debt, when considering software) is a way of thinking about the cost of shortcuts.  When you are putting together a project and you need to get stuff to work now, we often choose an approach that's faster and easier to implement, but which we know isn't really the right way to do it.  That choice creates a piece of technical debt---an approximation, inelegance, or minor incorrectness.  When we do something else in the future, it may be harder because of the technical debt that we have created, either because we have to work around the poor implementation, because we have to fix the poor implementation, or simply because the poor implementation is confusing.

Technical debt is not necessarily bad, any more than financial debt is necessarily bad.  Taking on technical debt is a way of accomplishing something that you might not have been able to accomplish if you had to do everything "right" the first time.  It's also a way of deferring difficult decisions until we better understand which path is correct, in order to avoid doing something we think is "right" but that later turns out to have been wrong.  Like financial debt, however, technical debt can accumulate and can also lead to taking on additional technical debt, creating all sorts of havoc.  What's important is to keep track of your technical debt and to try to make wise decisions about when to allow it to accumulate and when to pay it off.

Lately, I've been realizing that this same idea can be extended more generally to organization of one's effort on across different projects.  As a scientist who leads my own investigative ventures, I have quite a number of different projects that are "live" at any given point, ranging in scope from an hour or two of effort here and there (e.g., service as an associate journal editor) to complex multi-year ventures (e.g., development of the aggregate programming stack).  When I make choices in managing these projects and my time between them, I often take on organizational debt.  This accumulates in lots of different ways, such as note cards accumulating on my desk, file directories growing large, open browser tabs, email messages on my "need to reply" list, etc.

Just like taking technical debt on in writing software, there are pluses and minuses in taking on organizational debt. If I spent lots of time trying to be hyper-organized and so that I avoid taking on organizational debt, then I will be much slower at actually accomplishing the things I'm trying to organize.  If my organizational debt causes me to overlook a deadline or to end up in a last-minute crisis trying to get things done, however, then my work and my life suffer in various ways. Some of these costs are quite signifiant and spill out of work into costs on the rest of my life, such as losing time with family and friends, lack of sleep, getting sick, and gaining weight.

My struggle of late, then, has been to recognize that paying down organizational debt is a real and legitimate part of my job, just as paying down technical debt is a real and legitimate task in executing particular projects in my job.  I wouldn't say that I've found a clear method of managing my organizational debt yet, but, as they say, the first step to solving a problem is to recognize clearly.  I have, however, made some steps forward that have helped immensely.

For example, I am now both publishing papers frequently and traveling frequently.  Both of those have major time lags involved.  For example, in publishing a paper I first submit a manuscript, then get reviews, send a revision, repeat until accepted or rejected, wait for publication, post and publicize; along the way I also need to obtain release permissions internally and sometimes also from funders.  This process can take years, and if I've got a bunch of papers in flight it's easy to lose track of a deadline and create a failure or crisis where none was needed.  Travel is similar: registration, booking flights, hotels, cars, sending in pre-expenses, actually taking the trip, waiting for expenses to register in the reimbursement system, submitting requests for reimbursement, and actually getting reimbursed often spans many months.  To address these problems, I've created a spreadsheet for each task which lets me have a "dashboard" view of what's going on overall with regards to that area of responsibility.  For example, here is part of the 2015 sheet from my publications spreadsheet:
Publications not yet available have their names tastefully redacted.
It's not a panacea.  Nothing is, for me, when it comes to technical or organizational debt, because I'm not willing to pay the additional cost of not taking on debt.  Moreover, I'm sure that my approaches will have to change periodically, as my career continues to evolve.  Solutions like these help, however, and recognizing the importance of (at least sometimes) tracking these moving targets is, I think, a useful step forward towards more improvements in my quality of life, both at work and also at home.

Sunday, November 15, 2015

Signs in the Snow

When I fly for business reasons, I always try to get a window seat.  I have always delighted in the view from the airplane window, at all the shifted perspectives one obtains when looking at the world on high.  As long as I still love these sights, I feel, I know that I have not become too jaded, and my sometimes-strained soul has not yet died (I speak in the metaphorical sense, of course). When I fly with family, it's different now: Harriet almost always claims the window, and despite my longing to displace her, I would rather share the joy than steal it.  When I am alone, however, even on my shortest and busiest flights, I will track the sights at least occasionally, and I often shutterbug my way through takeoff and landing, so happy that our minor electronics are once again officially allowed to be active at those time.

One of the things that always fascinates me most is the way the landscape radically transforms with seasons.  Moreover, despite what one might think, I find that it is winter that most brings out the texture of the land.  When snow is on the ground, its topography is highlighted, leaping out in dark lines on every vertical and slope.  With those thoughts in mind, I present to you dear readers, an album of interesting forms I've seen, which I think of by the title "Signs in the Snow":

Album: Signs in the Snow