Sunday, October 18, 2015

Racism, fond memories, and toddler education

As I was reading Harriet her bedtime stories tonight, I was struck once again by a thing that greatly pains me.  Many of my fondest childhood memories are laced with rather awful racism that I simply failed to be aware of.  Case in point, tonight one of the books we read was To Think that I Saw it on Mulberry Street. This book is a simple and delightful Dr. Seuss tale of a child's fantasies of what he saw while walking home, building from a simple horse and wagon to a fantastical parade.  And there, on the second to last page, is this:
"A Chinese man who eats with sticks" --Dr. Seuss
Apparently, Dr. Seuss thought that Chinese-Americans were just as unusual a freak-show as a man with a 10-foot beard, a magician pulling piles of rabbits from a hat, and two giraffes and an elephant towing a brass band down the street.  And so we get this image, on which I can count at least six blatant pieces of racism.  Worse yet, this is apparently the post-1978 revised edition in which the racism is toned way down: he's a "Chinese man" rather than "Chinaman" and he's no longer wearing a pigtail and painted bright yellow.

OK, I know that Dr. Seuss is well known to have done some awfully racist things over the years (e.g., this cartoon condemning Japanese-Americans during World War II).  I know this.  But it burns me up that I had no idea that this monstrosity was living inside a favorite childhood book.  In other words, it's not that Dr. Seuss was making racist drawings, but that I didn't remember the racism at all. We bought this book (well, I bought this book) for Harriet quite early on, on the strength of my fond memories, and I was shocked when I got to this point.  I also noticed that the police were Irish and was a little bit dubious about the Rajah riding the elephant.  Not being familiar enough with the subject matter, I wasn't sure if the Rajah was racist or just archaic (like a knight in shining armor or a lady in a wimple), so I asked my wife, who is South Asian.  Her answer? "Totally racist."

This leaves me with two dilemmas that I struggle with.  First, what does this say about me, to not have known I had such racism in my education?  Clearly there's at least a bit of "fish don't have a word for water" going on.  I did not have this racism called out to me, and thus I didn't realize that it was anything to notice.  It's there in many other things I loved as well, like If I Ran the Zoo (another Seuss), The Jungle Book, and Tintin (oh my goodness, Tintin).  I loved these things and, if I am honest with myself, still do.  My favorite Jungle Book story of all time is "Kaa's hunting," and now I cannot read its descriptions of the Bandar-Log monkeys without wondering if they are allegorical for Kipling's views of India.  Tintin in America is practically hallucinogenic in its kaleidoscope of stereotypes and disrespect for, well, everything, and I still would read it again if I had a copy here in front of me.

And that leads me to the second struggle: do I share these things with Harriet or do I censor them? Mostly, there's an obvious third path that avoids the issue: there are so many good things out there, that I can simply choose to select the ones that I find less problematic.  But what about the ones I find out afterward, like in Mulberry Street? Tonight, I didn't read the line.  I broke the rhyme and went straight to the big magician doing tricks.  Other times, I read it through.  Sometimes, I point things out to her and critique them ("this picture is being mean"), and sometimes I do not.  Mostly, I am uncomfortable and simply shift my strategies back and forth.  I find some of the advice out there about liking problematic media to be useful, but it's not the end of the story and I still have not found peace.

Saturday, October 17, 2015

SBOL 2.0, governance, and Jake's self-perception

This past summer, one of the most significant scientific milestones I've been involved with is the publication of the SBOL 2.0 standard for representation of biological designs.  What it's all about is being able to better describe and exchange information about the genetic constructs and similar such systems that people are trying to build.  Perhaps the best way to describe it is with this diagram I prepared for a talk, comparing SBOL 2.0 to previous standards:


FASTA is about as bare-bones as it comes: pretty much just listing out the DNA sequence that you want.  GenBank lets you annotate that sequence with descriptive information about what the different parts mean, and SBOL 1.0 lets you describe the structure of a design hierarchically in terms of annotated sequences that get combined together as "parts" to make bigger designs.  SBOL 2.0 lets you talk about function as well, describing the way that these parts interact with one another to create the overall behavior of a design.

Conceptually, it's fairly simple, but in practice it took several years to work out and the arguments are not yet over.  The document that we produced is more than 80 pages long, and we're still tinkering with bits and pieces as we try to understand all of the consequences of what we've built.

SBOL is heavy on my brain right now because for the past week, I've been at the COMBINE meeting, where the communities for SBOL and a number of other biological standards meet up to try to improve their systems, work on interoperability, etc.

This is still not something I ever thought I would be doing with my life.  Even now, in my prejudicial mind, standards design is still something done by grey little people who care passionately about trivial and boring things.  I struggle with this, because I look at my work in this area and simultaneously feel that it is highly important and mind-numbingly stultifying to anybody who isn't actually in the room arguing passionately about the potential long-term consequences of adding a single arrow to a diagram.

A case in point: one of the things that I'm most proud of this week was the updated governance document I drafted, and my mediation of discussion on this document, which helped tune it to become widely accepted; the updated version now appears well on its way to official approval by a formal community vote.  So, apparently I am proud of work I've done on adjusting the methods for making decisions regarding an experimental standard for interchange of information about biological designs that will allow faster prototyping of improved systems for biomedicine, biomanufacturing, etc.  That's at least five levels of separation from anything that really affects the larger world. Looked at in that light, this is clearly the very definition of obscurity.  And yet, let me spell it out in another way...
  • Good governance, which gets openness, power, and decision-making right, is critically important for the health of a community, and a number of little warning signs have indicated that the SBOL community needed to adjust its governance to match the way the group has developed and grown.
  • If the SBOL community governs itself effectively, then it will make better decisions that are more likely to lead to a useful and effective standard.
  • If the SBOL standard works well, it will make it a lot easier for people to develop good biological engineering tools.
  • Those biological engineering tools will make it a lot easier to safely and predictably engineer with and for living organisms.
  • Used responsibly, those capabilities can help make all of humanity healthier and safer, as well as improving our ability to manage our environmental impact on a global scale.


This nail I've driven in is very small and unimportant, almost certainly, and yet it matters.  It matters a lot, and not at all, all at the same time.  And I suppose that's just the way the world works, on a planet with seven billion interconnected and increasingly technologically powerful individuals.  Our civilization is remarkably strange and obscure in its operation, and I'm glad when I find satisfaction in the parts I play.

Thursday, October 08, 2015

Publication delays ARE aimed at manipulating impact factor!

A few months ago, I wrote a post with a question: Are publication delays aimed at manipulating impact factor?

Today, I have an answer to that question: yes.

A recently published article, "Editors’ JIF-boosting stratagems – Which are appropriate and which not?" (h/t RetractionWatch) investigates strategies that journal have been using to boost their impact factor and explicitly calls out what it calls the "online queue strategem."  The article is paywalled, so let me summarize here.  In addition to reviewing some of the better-known and clearly unethical practices used by some journals (e.g., forcing citations on authors, citation cartels), the paper carefully dissects the effects of having a long "online early" period of publication, finding four main effects:

  • Papers accumulate citations before "official" publication (multiplying by ~1.5 to 2)
  • Citation rates typically peak 3-4 years after publication, so shifting the time selects for a better citation date (adding another ~50%)
  • Queue order can be manipulated to publish the papers picking up the most citations earlier, (adding another ~30%)
  • Calendar-year boundaries mean that papers in early months count more than papers in later months, so strategic organization of early-month issues can further boost citations (adding another ~30%).

All of this adds up to around 5-fold potential distortion in impact factor.  Since the dynamic range of most journals is only around 0.5 to 10 anyway and even the very highest impact factor journals top out at ~50, this renders that most precious number completely useless.

Now, it's possible that many journals aren't deliberately and strategically manipulating their queues, meaning they'll only get about a 2x boost in impact factor from queuing.  So what?  It still means that impact factor is going to be highly distorted and basically only good for distinguishing journals into three categories: "glamour journal", "normal journal", and "ignored journal" (less than about 0.3).

Ironically, the article itself is dated February, 2016.

That's it: it's clearly time to adopt the wise strategy of my favorite satire journal, the Proceedings of the Natural Institute of Science.  Their current impact factor? "Leadership"

Wednesday, October 07, 2015

Scale-free distribution of payoffs in science

One of the things I've been enjoying these days has been answering questions on the Academia site on StackExchange.  This question-and-answer site is part of the vast network of Q&A sites that have flowered out of the wildly successful StackOverflow, which is pretty much the best source for coding help on the internet.  The model is that people ask question about the topic, e.g., academia, and other folks turn up and provide answers, and then you get or lose Fake Internet Points depending on whether the crowd thinks it's a good answer.  It's surprisingly effective and also, for me at least, pretty enjoyable and kinda addictive.

Anyway, I answered one this morning that made me think a lot, and I thought that I might share my thoughts here as well.  The question was simple, fundamental, and ill-posed: "What is the distribution of payoffs in research?"  Basically, the person is wondering whether every experiment is a roughly equivalent step forward, or whether some are much more valuable than others, and if so whether there's some sort of power-law relationship between topic, funding, and value of result.

This is ill-posed, because the whole notion of "payoff" is extremely vague and probably the wrong question to ask, but it really made me think.  My response, which I'd like to share with you, was this:

There's a vast amount of ill-definition and uncertainty wrapped up in your question... and yet despite that, the answer is almost certainly yes, there is a power-law distribution.
I'm going out on a limb a bit here, because I'm not building on any published analysis that I'm aware of. However, a little analysis of limit cases and fundamental principles can take us a long way here. Let us start with two simple and relatively uncontroversial statements:
  1. Better experimental design leads to better results. It seems self-evident that if you make a bad choice in designing and experiment, it's not going to get you the interesting results you want. At the micro-scale, some choices are clearly better than others, and some are clearly worse.
  2. Sub-fields appear, expand, shrink, and die. As I write this, CRISPR research is hot, and a lot of people are finding interesting results there, and accordingly that field is rapidly expanding. Nobody is doing research on the luminiferous aether because it's been discredited as an idea. Nobody is trying to prove that it's possible to generate machine code from high-level specifications because Grace Hopper did that in the 1950s, when she invented the compiler, thereby initiating what is now a fairly mature and stable research area.
So clearly, no matter how one defines "payoff," any sane definition will see a highly uneven distribution of payoffs both the micro-scale of individual experiments and at the fairly macro level of sub-fields.
Finally, we need to recognize that "significance" is a matter not only of objective value, but also of communication through human social networks. This means that the same result may have wildly different impacts depending on the methods and circumstances of its communication. The history of multiple discoveries in science is ample evidence of this fact; one nice illustrative example is the way in which Barbara McClintock's work on gene regulation was largely ignored until its later rediscovery by Jacob & Monod.
So, we have variation and we have interaction with human social networks, which tend to be rife with heavy-tailed distributions. All of this says to me that it would be remarkable if there were notsome sort of power-law distribution regarding pretty much any plausible of definition of impact, significance, and investment. For these same reasons, I think it would also be surprising if one can make any more than weak predictions using this information (e.g., "luminiferous aether research is unlikely to be productive", "CRISPR is pretty hot right now").
And the devil, of course, is in the details...

Saturday, October 03, 2015

Tribute to driving in Germany

I'm sitting in the Frankfurt airport right now, having just finished driving two hours on the Autobahn up from Schloss Dagstuhl, in Wadern near the French border.  As an American, I'm used to fairly titanic road networks, but driving on the Autobahn feels different to me: my impression is that while Americans use our roads, Germans really love their roads.

Out in the gently winding hills and valleys of Western Germany, forests and fields flash past traffic freely flowing at 100 miles per hour.  The gentle curves are well designed to encourage speed, and many people take good advantage of it.  Yes, it's true, on much of the Autobahn system there is simply no official speed limit (though you'll still get pulled over if the police think you're driving dangerously), and even where there is a limit it usually restricts you only down to 130 kph (a little over 80 mph).

In my little bitty economy rental car, I cruised along comfortably at 110 mph or so in 6th gear (don't even try asking for automatic in Germany), its happy German engineering not making the least complaint about the speed.  At home, my faithful Toyota starts to get very loud and quite unhappy by the time that I reach 85.  Even so, happy-looking people in bigger cars rocketed smoothly past me at significantly higher speed, bound for who knows where at the highest speed available.  And all I know is that my head has got this song on repeat, and I invite you to join along with me and sing:

Why I love iGEM

Last weekend was the annual iGEM jamboree---that is, the International Genetically Engineered Machine (iGEM) Competition. I poured in about 60 hours of my time into the contest over the course of 3.5 days, and by the end I was exhausted, both physically and mentally, but feeling absolutely elated and on top of the world, raring to go for another one in 2016.

What had I seen, and why was I so excited?  Well, iGEM is a magnificent and unique event, a gathering of students from every continent, from high school on up, all driven by a passion for biological engineering and simply overflowing with creativity.  Each team spends the summer working together on a project that they create, and in the fall they come together to have a big party, where everybody gives talks on what they've done and the best few are recognized in front of everybody for their superlative accomplishments.  There's lots of silly things, lots of over-ambitious ideas that don't get too far, and lots of nice little steps and learning by the students.

And in the middle of it all, some damned good science gets done as well.

Last year, I co-founded a new track at iGEM focused on measurement.  Yes, we're back to that again: my obsession with terribly unsexy rulers.  We had some very good teams last year, and this year again there were a bunch of excellent projects in the measurement track.  And this year, one of those projects stood out head and shoulders above all the rest.

The team from William & Mary, a small but long-standing and excellent public college in Virginia, chose to focus on an important but subtle problem: quantification of noise in gene expression.  Building on recent work in the area, they dug into the problem and ended up with a simple and easy to use kit for measuring this noise, then applied it to quantify noise for a few of the most widely used biological components in the iGEM parts registry.  Very deep and very geeky, but it matters a lot.  If we want to have safe and reliable genetic engineering, we need to be able to predict what will happen when we modify an organism, and this strikes right at that heart of that problem by measuring predictability.

But that wasn't all: they also worked with their county school system to develop a curriculum for synthetic biology.  It's magnificent, and you can get a copy for free online.  Inside this 80-page document, you can find 24 age-appropriate activities, from "DNA twizzlers" for 1st gradesr to Monster Genetics a couple years later (fire-breath is a dominant trait, but cyclopses are recessive), building all the way to adult-level work in high-school like PCR amplification of DNA and bioethics analysis.  The interactions with teachers really show, as the lessons are not only pretty but also give clear goals and a materials list and expected cost per student (usually a whole class can be supplied with a just a few dollars of groceries or arts & crafts supplies).  Even more remarkably, teachers have already begun enthusiastically adopting it, both throughout their county and in other states and nations.

The William & Mary team gave clear, understated presentations that simply let their work shine through, and the whole community recognized it, ultimately first giving them a chance to present as a finalist in front of all the thousands at the convention center, and finally awarding them the competition's top prize (along with a bunch of others as well).  This simple yet deep set of work comes from a team whose school doesn't even break the top 100 in US News' ranking for biology, and shows the power of careful and thoughtful work in science.  The Washington Post may have been too confused to even mention them, but their university is quite elated, and its staff took the time to understand and write a clear and accessible article about their project.

To me, all of this is a vindication not just of the work I've put in organizing and promoting measurement at iGEM, but of the entire scientific process.  Good things can come from unexpected places, and sharp minds thinking careful thoughts can be recognized and receive the recognition they deserve.  Yes, there are problems in the scientific world---quite many, in fact---but this is why we must fight to preserve and promote the scientific ideals, and to keep making that world more diverse, more inclusive, and more able to recognize and promote the potential to improve our world and make a difference.  This is why I love iGEM, why I'm proud to be involved and for what part I've had in helping to enable this, and why I'll be back again for more in 2016.

William & Mary, iGEM 2015 winners, with the Measurement Track committee

Wednesday, September 30, 2015

Aggregate Programming!

It's finally out: our IEEE Computer article, "Aggregate Programming for the Internet of Things" (free preprint) is up and online where all can get it for free.  Don't be fooled by the name: this isn't really just about the our increasingly networked possessions  ("the Internet of Things"), it's a much more general paper.  In fact, this is the first place we've really put all the pieces of our last few years' distributed systems research together, into a generally accessible article that clearly introduces a better framework for building distributed systems.  Please allow me to introduce the aggregate programming stack:

Aggregate programming stack, with examples from a crowd safety application.
In computer networks, the OSI model is a "stack" of abstraction layers that separate different aspects of computer communication.  The browser you are reading this on, for example, is probably obtaining it via HTTP at the Application Layer, routed to you via TCP/IP at the Transport and Network Layers, respectively, with the last link sent to you over something like 802.11 or Ethernet handling the Data Link and Physical layers.

Our aggregate programming model takes a similarly layered approach to the problems of designing networked systems, breaking these often extremely complex problems into five layers.  From bottom to top, these are:

  1. Device: this is the collection of actual electronic devices that comprise the system, with their various built-in sensors, actuators, ability to communicate with one another, etc.
  2. Field Calculus: this layer abstracts the devices into a simple (but universal) virtual model that can freely be mapped between an "aggregate" perspective in which the whole system acts like a single unified device, and a "local" perspective of individual device interactions that implement this model.
  3. Resilient Coordination: this layer consists of "building block" algorithms (implemented in field calculus) that provide guarantees that systems will be safe, resilient, and adaptive in various ways.
  4. Developer APIs: useful patterns and combinations of building blocks are then named and collected into application programming interfaces (APIs) that are easier to think about and program with.
  5. Application: Finally, distributed systems can be much more easily constructed, using the APIs just like one would any other single-machine library.
Constructing this stack factors the problems of distributed systems development into separable components, each much simpler than trying to tackle the whole complicated mess at once.  If you just want to build applications, you just need to learn the Developer API layer and work with that, just like web programmers learn about HTML and Javascript.  If you want to work on resilient algorithms, on the other hand, you get involved with the plumbing at the Resilient Coordination layer, and if you want to use the stack on a new device, you implement a copy of the interface required for a Device by an instance of the Field Calculus layer.

I'm very proud of this work, and think it's got a potential to really change the way that people deal with complex computer networks.  For the programmers amongst you, dear readers, I suggest you check out both this paper and our (still somewhat rough) implementation of field calculus in Protelis.

Saturday, September 19, 2015

A Golden Boston Sunset

Dear readers,

Some days, it's just good to be alive, and a thing comes out of nowhere unexpectedly to remind you of that fact.  Today, as my flight was gliding down into its final descent into Boston, the air was almost perfectly clear and the sun was just at that magic moment in its descent where everything begins to be golden and shadows stretch out just enough to give the third dimension of everything an extra bit of special emphasis.  As I gloried in the texture of the light, my camera came out and I snapped away---not blocking myself from an enjoyment of this sight, but finding that the aim to capture gave me extra focus and appreciation for the details.

My dear readers, I wish to share this joy with you, in the form of a few of the best moments of imagery I captured.  May this lighten your day as it has lightened mine.





Monday, September 14, 2015

A Tale of two CRISPRs

Last year, I had my first "glamour journal" publication, as second author of a Nature Methods paper on a new family of CRISPR-based synthetic regulatory devices.  Actually, I had two "glamour" publications---the other was a Nature Biotech paper on the SBOL language for communicating biological designs with 32 authors, the biggest collaborative publication I've been involved in to date.  That's a tale for another post, however---this one's all about CRISPR, CRISPR, CRISPR.

For those who haven't encountered the wonderful hype-storm around CRISPR, the acronym expands to the highly non-mellifluous "clustered regularly interspaced palindromic repeats," which tells you virtually nothing about why it's cool.  The reason it's cool is because one of the things this awkward acronym refers to a protein ("Cas9") that docks with fairly arbitrary "guide RNA" fragments in order to go act on DNA that matches those sequences.

Core CRISPR mechanism: Cas9 protein binds to gRNA, which targets the protein to a matching DNA sequence

Protein design is really hard, but DNA and RNA design has become reasonably straightforward, so CRISPR is an awesome mechanism: it lets us target a (fairly) predictable protein effect to pretty much any piece of DNA that we want.  People have used it for editing DNA, which has previously been done with lots of other mechanisms, but gets much easier with CRISPR (hence the recent controversies you may have seen in the news around human genetic engineering---the changes we can do aren't any different, they're just a lot cheaper, which is a meaningful difference of a different sort).

Our paper last year showed for the first time how to use the CRISPR mechanisms to make potentially large numbers of strong biological logic gates.  This is important because one of the big things that's been holding synthetic biology back is the difficulty in building reliable computation and control systems inside of cells.  We've known for a long time that biological computing is possible, but there's only been a handful of decent computational devices, and no good ways of making more.  Now, within the last few years, there have been several different families that have emerged, including TALE proteins, homolog mining, invertases, and now, with our paper, CRISPR repressors.  Our CRISPR repressors are nice because they can potentially easily generate thousands of high-performance devices and implement all sorts of complex computations, something that nobody currently has a clear approach for with any of the other families.

Diagram of one of our CRISPR repressors: a modified Cas9 protein (blue box) acts as "power supply" for an inverter logic gate implemented by having the gRNA (orange box) regulate a synthetic promoter (blue arrow). The important things to know are 1) the orange box and blue arrow are easy to design and we can make lots of them that don't interfere with one another, and 2) the blue box can potentially power lots of these gates at the same time.

So I was (and still am) very excited about this publication for two reasons: first because I think it's a big step forward scientifically, and second because it's in a big-name venue that lots of people are likely to pay attention to and where it's more likely to have a big impact on scientific practice.

Just last week, I had my second paper in Nature Methods, led by my same awesome collaborator, Samira Kiani, and following on the subject: this time, our paper shows how to use CRISPR devices to both compute and edit genes in the same circuit.  My reaction, however, has been much more mixed to this publication.  Don't get me wrong: I'm really happy to be published in a high-ranked journal again, and I really enjoy working with Samira (soon to upgrade from Dr. Kiani to Professor Kiani!), who I find an insightful and diligent collaborator and whose skills I think complement my own quite nicely.  Maybe it's just that I can't be so deliriously excited about getting published in a journal a second time?  I'm also not as excited about these results: it's a nice twist on previous results and a useful new capability, but in return we lose some of the device efficacy. Overall, though, I just don't feel like this paper is a game-changer in the way that our first paper might prove to be.

Still, it matters, and it's a step forward for all of us.  Soon, we will meet, celebrate this success with a toast and a fine dinner, and plan our next venture toward transformation of the world and toward posterity.

Sunday, September 06, 2015

Perhaps my least interesting publication ever

Just recently, I was listed as first author (out of five), on what is perhaps the least interesting scientific publication in my history as a researcher.  This includes even semi-embarassing old rants from when I was a young and arrogant graduate student---those at least give some sort of plausibly interesting perspective on what I was thinking about at the time.  Not to say this document isn't important: I think it was definitely worth the time and effort, and is useful.  That doesn't necessarily mean anybody will derive any particular joy or pleasure from encountering it.

So, what is this deadly dull publication that I've for some strange reason decided to advertise so loudly on the Internet?  Its formal name is: BBF RFC 107: Copyright and Licensing of BBF RFCs. This takes a little bit of explanation, so bear with me and please try not to fall asleep too quickly: one of ways that people in the synthetic biology community share their work is by posting open "Request For Comment" documents (RFCs)---essentially draft standards, following the main model used for developing the Internet.  These are cataloged by the BioBricks Foundation, hence BBF RFCs.  The first of these, BBF RFC 0 (yes, there were computer scientists involved, and we like to count starting with zero), sets out the process for how to submit a new RFC.  A few months ago, I noticed that the original handling of copyright had gotten out of date with respect to some current practices in accessing scientific documents online and current preferences for open standards development.  I raised these issues with the BBF RFC maintainers, and we figured out a legal "patch" for BBF RFC 0. The end result of all this is a 1.5 page document that makes two small changes in how new BBF RFC documents are handled:

  • The document is actually marked with a modern open copyright license, and
  • The authors share copyright with the BioBricks Foundation, rather than transferring it.

Now, unfortunately, the parts of BBF RFC 0 that we didn't replace weren't followed correctly in setting forth this RFC, which has caused some trouble with another BBF RFC that I'm involved in, but that story's even less interesting, and I'm sure it will all get sorted out eventually.

In case your eyes have well and truly glazed over, let me sum that all up more simply: I noticed a little thing about copyrighting certain scientific documents that needed tweaking.  By a quirk of process, doing so had the side effect of creating an archival scientific publication.

So, was it worth it?  Absolutely: it didn't take much time, and copyright is one of those things that it's often worth paying close attention to, because if you screw it up as a community, you can accidentally wind up poisoning all sorts of things down the line, if nasty people decide to try to take advantage of loopholes or cautious organizations get blocked from doing things by technicalities.  I'm just quite amused that this ends up in my list of publications as outwardly indistinguishable from BBF RFCs that took many people years of work and that gather lots of citations.  It's also kind of funny from a "what do scientists do all day" perspective.

But, for the love of all that you hold holy, don't read the document unless you actually need to.