Saturday, November 07, 2015

Jack Splat!

Last night the Iowa Children's Museum presented us with a wonderful surprise. We'd made a plan to go there for a Friday evening outing after school, and we walked unexpecting into the remarkable event known as Jack Splat! In the big "main street" open space at the heart of the museum, where normally one could access the music room and the post office, tarps were laid down and a cleanup crew stood below. Above on the balcony lurked a great swarm of over-aged pumpkins, a week past Halloween and ready to meet their end.

Our timing was perfect: Isaac Newton was just explaining to a rambunctious audience how his first law of motion meant that a falling pumpkin, once in motion, would keep moving until it encountered an opposing force---"The ground!" cried the children, and the pumpkin flew down to meet its fate.
Isaac Newton (resplendent in toilet-paper-roll wig) and lab assistant preparing to launch a pumpkin.
The physics lessons continued, accompanied by redolent meaty splats of demonstration. Before each pumpkin flew, the throwers read out its name and the name of its donor, as well as the specified method of execution (e.g., roll from the ledge face first, backflip up in the air).  The children cheered and chanted (though one little boy near us was quite upset, and asked why the people hated pumpkins so much), and the rain of gourds continued for nearly half an hour, quite challenging the cleanup crew to keep up with all the mess.
Getting ready for the next bombardment.
I much enjoyed the unexpected show, and was reminded of a similar but much smaller scale yearly event arranged by the undergraduates at MIT.  A good and rather cathartic end to a Friday evening.

Friday, November 06, 2015

The Power of Ignorance

When my friend and colleague called herself clueless this afternoon, it triggered something in my mind.  We've been working on a paper together, and since I'd taken the first pass I'd put everything in LaTeX, hoping to avoid having to figure out how to manage citations and such in Microsoft Word.  I hate writing in Word, and as a computer scientist I usually get to avoid it, opting instead for the nitpicking and precision control that LaTeX offers me. When I am collaborating with biologists, however, LaTeX might as well be Martian or Haskell to many of them, and we often default back to Word.  Selfishly, I'd avoided that, and just assumed I'd end up taking feedback notes or snippets of text and incorporating them myself.

But I had aimed too low, and here my friend surprised me.  Rather than complain or take the easy route, she asked me to put the document into Overleaf, an online LaTeX interface that she'd just learned about, and she did her editing there, including using the LaTeX todo notes package that I've been using for tracking commentary in the document.

And then she called herself a "clueless biologist," as she asked me questions about this fearsomely complex new technology she's been voluntarily educating herself on.

Those two words say something that is very important and also sad about the way that science often operates, and I think about our larger society as well.  My friend is largely ignorant about LaTeX, in the sense that she lacks knowledge, but "clueless" is a rather negative view of ignorance.  Adding "biologist" lumps her into a category, othering her and tying that to this "clueless" expectation---a stereotype that I'm quite familiar with.  I don't like it, though, because that phrase sets up an expectation of an "us vs. them," Men are From Mars/Women are from Venus, Computer Scientists are from LISP / Biologists are from S. cerevisiae sort of oppositional dichotomy.  That makes us all smaller, I feel, because it divides us and suggests our minds are alien to one another, and therefore that we ought not to attempt to learn from each other so much.

Instead, I think that we should be celebrating and embracing the power of our ignorance.

By this, I do not mean that we should avoid knowledge.  Knowledge is wonderful and empowering, and recognizing one's ignorance is a first step to doing something interesting involving others who are not ignorant in that same area. The Renaissance Man is dead---quite dead---and Joy's law is how we spend our lives, in a world where there are so many interesting things to know about, some important and some just fun for somebody.  We are all more ignorant than we can know, and not just in a snotty Socrates one-upmanship way of looking at it.

One of the hardest things I had to learn while I was a graduate student was how to say, "I don't understand" and "I don't know."  I learned it from Hal Abelson, the professor who always asked me the hardest simple questions I have ever heard.  Hal taught me that saying "I don't understand" did not have to be an admission of weakness.  It could simply be the truth, and then what happens next depends on why you don't know.  When I was talking to Hal, it was usually the case that Hal saying "I don't understand" was an indicator of some fundamental flawed or overlooked assumption in the thing that I was trying to explain to him.  I learned a lot about my own research from Hal's admissions of ignorance, and I also learned to stop being afraid of lacking knowledge.

I'm still learning that. It's easy in our competitive world to fear that admitting ignorance is the first step to losing out to other people who are better at putting up a front, and I still struggle with that. But fearing ignorance is almost as bad as being proud of it, and I prefer to avoid it when I'm not too panicked to consider the bigger picture that I'm living in.

The fear of ignorance is competition, but the power of ignorance is partnership and teamwork.  I'd much rather live in the second world, and I hope that I can be sufficiently wise to help encourage it for both myself and my compatriots.

Monday, November 02, 2015

Whole lotta ruttin' going on

That's what the highway sign said this morning:
WHOLE LOTTA RUTTIN GOING ON WATCH FOR DEER
It brought a smile to my face even as I was duly warned of the frightful danger that I face on the roads of Iowa.  This is one of the most intense states in the nation, when it comes to collision risk: currently third, with a 1 in 68 chance of hitting a deer each year.  Frankly, this number really blows my mind, especially coming from Massachusetts where the odds are an order of magnitude lower, and Boston in particular where you might as well not even bother thinking about the possibility.  I've had a recent close call of my own already, a couple weeks back on a date with my wife, when a deer darted across in front of us and I had to jam my breaks on to avoid an accident.

In Iowa, the distinction between countryside and city is much sharper and closer than New England, and it seems to me that it's a virtually ideal environment for deer.  Deer flourish on the edges of forests and in sparse woodlands, and those are things that Iowa has in great amount.  Drive through the countryside, and you see something that I have never really known before I moved out there into the Midwest: an entirely rural-industrial environment.

Growing up in New England, and with my father's stories of growing up in Colorado, I'm used to the idea of "rural" meaning lands where little or no people live.  Walk through the woods of Maine or New Hampshire or even Massachusetts, and you will first of all find that you've a very hard time walking through those dense pine woods at all, and second that they often extend in all directions for many miles, a fearsome wilderness broken only by the old stone walls of centuries-abandoned farms. The woods of Iowa, by contrast, are linear affairs, in which one would have a rather difficult time getting lost at all.  They curve along contours of land, following streams and rivers, in those narrow places where the land is too steep to be productively arable.  Elsewhere, the land is mostly farms, broken by an apparently arbitrary fractal dispersion of cities, towns, and factories.  Suburbs don't exist they way I knew them in New England, where the city slowly peters out into nothingness over the course of many miles: Iowa City stops about a mile East of our house on a straight-edged line, an instant transition from extremely dense developments to fields of corn and soy (rotating on a 3-year cycle).  Even in the most rural areas I've been, you're never more than a couple of miles from a dense aggregation of a few hundred people clustered together in a tight little town.

Biodiversity is low, in an environment like this, but species that do well on the boundaries with humanity, like deer and rabbits, flourish and expand.  All the halloween pumpkins in our neighborhood are attacked and eaten by marauding squirrels.  And this transplanted specimen still feels for roots, sorting out my place in this rich Midwestern soil.

Monday, October 26, 2015

An unknowing inheritance: BBN's stop and go history in genetic engineering

When I joined BBN back in 2008, one of the new things I brought with me to the company was my research in synthetic biology.  I was a starry-eyed and naive recent Ph.D. graduate, and it was one of my little exploratory sidelines, which would not expand into a full-scale line of research for another two years, blooming as my AI work slowly withered away into neglect.  Nobody at BBN was even thinking about synthetic biology at the time, and nobody I was working with had any institutional memory of such research being done at BBN before, and so I assumed that must be so, and indeed for most intents and purposes it was so.

In fact, however, BBN has been a significant player in the work of genetic engineering at least twice before that I now know about.  One of those times I learned about several years ago, and is not the brightest of episodes for the community.  The other, however, I only learned about a short time ago, as I prepared to give a talk on the new SBOL 2.0 standard for encoding genetic designs, and it makes me both proud of my institution's history and amazed that it has somehow dropped out of its memory as an institution.

The nearer and less proud episode was BioSPICE, and it haunts my every step as a non-lab-centric researcher.  As best I understand the project (I was still a grad student chasing strong AI), in the early and heady days of the word "synthetic biology," a bunch of the leading researchers in the field made a try on the big goal of predictable simulation and engineering of organisms, at that point thinking they already had a sufficient critical mass of good tools and knowledge to take something like a straight-up electrical-engineering-style approach to the problem.  Thus, BioSPICE, a big DARPA-funded project to try to build the equivalent of the SPICE tool for simulation and engineering of electronics, which started with much fanfare in the early 2000s and much more quietly folded up shop a few years later.  I had known about some of the academic side (quite distantly at the time), but only later came to learn that BBN had been significantly involved in some way---I'm still not quite sure how, since the few stories I've gathered don't seem to correspond to what searching online turns up.  Years later on, when I would open my mouth to talk about the promise of model-driven design, it was often BioSPICE that dogged my heels, and fueled cries of, "We know that doesn't work, just remember BioSPICE!"  I have fought that history hard, all the way to the last few years when we've been finally able to start producing evidence that we really can predict biological circuits from their component parts.

The other episode of BBN's involvement with genetic engineering is much older, quite long before my time.  Back in 1982, when I was only four years old and the genetic engineering revolution not so much older, BBN began development of GenBank, probably the most important repository of biological information in the world.  What is it, and why is it so important?  GenBank stores genetic sequence information: it's where pretty much all the important scientific information about genes and genomes gets stored, one way of another.  BBN, with subcontracting help from another apparently unlikely collaborator, Los Alamos National Laboratory, put it together and ran it for its first few years of existence, as it became established and started gathering information.  Eventually, as it became less a research project in and of itself and its contents became more and more important, it moved to curation by the NIH, who still manage it to this day, as an exponentially growing resource made publicly available to all of humanity.  Perhaps it's not quite as big a deal to work on as the internet or email, but pretty close, in my books.

Somehow, though, we seem to have almost entirely forgotten this history, as a company,  It's not trumpeted on the list of accomplishments on our front page, nor bragged about in the "history of BBN" materials that people pass around. The only name I've been able to find so far who was associated with the project from BBN is Howard Bilofsky, who apparently spent 17 years at BBN before leaving in 1990 for a long, distinguished, and apparently ongoing career in the biotech industry.  Someday, I would love to look him up and learn a bit more about the hidden corners of our corporate history.

GenkBan and BioSpice, triumph and failure.  And now a third wave of biology at BBN, with me, trying to navigate these waters once again, as best my limited scientific sight can guide me.

Friday, October 23, 2015

Is academia really just a huge competition?

Another question that really made me think was posed last night on the Academia site of StackExchange, and once again I'd like to share my answer with you, my dear readers.

The question was simple in its essense, yet deep and rather challenging:
Is academia really just a huge competition?
I started writing an answer several times before I finally ended up with a direction that I could really believe in what I was saying.  The result was this statement, that I think reflects some difficult passages of my own over the years, back and forth along the tension between cooperation and competition:
You've asked a question that is both very important and very difficult, as well as one that is likely to draw different answers from different people depending on their own experiences in academia. 
This is because there are both competitive and cooperative aspects to academia. Different people take different strategies with respect to the balance between these two, and that affects their communities as well, so that the mixture of competition and cooperation that you encounter will also radically differ between different academic communities. 
Some of the key factors for inducing cooperation are: 
  • Science is hard.
  • Working together, people can accomplish things that they cannot possibly accomplish alone.
  • Cooperation in a team gives you an advantage when competing with other teams.
  • Many people enjoy working together in teams, and this is just as true for science as it is for any other human endeavor.
  • Scientific discovery feels awesome and it can be really fun to share that feeling with other people.
Some of the key factors for inducing competition are: 
  • Inherent conflict of ideas: when theories compete, people often become polarized and begin competing based on the "team" they support intellectually.
  • Limited resources: you've got a good idea, but a lot of other people have good ideas too, and there is not enough funding to support all of them fully: some people will not get what they want. Likewise, the Hubble space telescope can only point at one thing at a time, and there are a lot more things people want to point at than time to point at them.
  • Explicit competition set up by external agencies. For example, DARPA will sometimes make scientists in the same program compete with one another, and the loser gets their funding cut off.
  • Many people are just plain competitive, and want to "win" over other people in various different ways, and this is just as true for science as it is for any other human endeavor.
Bottom line: just like everything else, academia can be a competition, and everyone faces some aspects of a competition. But it's not just a competition, and I feel sad for anyone who experiences it in that manner. 

Sunday, October 18, 2015

Racism, fond memories, and toddler education

As I was reading Harriet her bedtime stories tonight, I was struck once again by a thing that greatly pains me.  Many of my fondest childhood memories are laced with rather awful racism that I simply failed to be aware of.  Case in point, tonight one of the books we read was To Think that I Saw it on Mulberry Street. This book is a simple and delightful Dr. Seuss tale of a child's fantasies of what he saw while walking home, building from a simple horse and wagon to a fantastical parade.  And there, on the second to last page, is this:
"A Chinese man who eats with sticks" --Dr. Seuss
Apparently, Dr. Seuss thought that Chinese-Americans were just as unusual a freak-show as a man with a 10-foot beard, a magician pulling piles of rabbits from a hat, and two giraffes and an elephant towing a brass band down the street.  And so we get this image, on which I can count at least six blatant pieces of racism.  Worse yet, this is apparently the post-1978 revised edition in which the racism is toned way down: he's a "Chinese man" rather than "Chinaman" and he's no longer wearing a pigtail and painted bright yellow.

OK, I know that Dr. Seuss is well known to have done some awfully racist things over the years (e.g., this cartoon condemning Japanese-Americans during World War II).  I know this.  But it burns me up that I had no idea that this monstrosity was living inside a favorite childhood book.  In other words, it's not that Dr. Seuss was making racist drawings, but that I didn't remember the racism at all. We bought this book (well, I bought this book) for Harriet quite early on, on the strength of my fond memories, and I was shocked when I got to this point.  I also noticed that the police were Irish and was a little bit dubious about the Rajah riding the elephant.  Not being familiar enough with the subject matter, I wasn't sure if the Rajah was racist or just archaic (like a knight in shining armor or a lady in a wimple), so I asked my wife, who is South Asian.  Her answer? "Totally racist."

This leaves me with two dilemmas that I struggle with.  First, what does this say about me, to not have known I had such racism in my education?  Clearly there's at least a bit of "fish don't have a word for water" going on.  I did not have this racism called out to me, and thus I didn't realize that it was anything to notice.  It's there in many other things I loved as well, like If I Ran the Zoo (another Seuss), The Jungle Book, and Tintin (oh my goodness, Tintin).  I loved these things and, if I am honest with myself, still do.  My favorite Jungle Book story of all time is "Kaa's hunting," and now I cannot read its descriptions of the Bandar-Log monkeys without wondering if they are allegorical for Kipling's views of India.  Tintin in America is practically hallucinogenic in its kaleidoscope of stereotypes and disrespect for, well, everything, and I still would read it again if I had a copy here in front of me.

And that leads me to the second struggle: do I share these things with Harriet or do I censor them? Mostly, there's an obvious third path that avoids the issue: there are so many good things out there, that I can simply choose to select the ones that I find less problematic.  But what about the ones I find out afterward, like in Mulberry Street? Tonight, I didn't read the line.  I broke the rhyme and went straight to the big magician doing tricks.  Other times, I read it through.  Sometimes, I point things out to her and critique them ("this picture is being mean"), and sometimes I do not.  Mostly, I am uncomfortable and simply shift my strategies back and forth.  I find some of the advice out there about liking problematic media to be useful, but it's not the end of the story and I still have not found peace.

Saturday, October 17, 2015

SBOL 2.0, governance, and Jake's self-perception

This past summer, one of the most significant scientific milestones I've been involved with is the publication of the SBOL 2.0 standard for representation of biological designs.  What it's all about is being able to better describe and exchange information about the genetic constructs and similar such systems that people are trying to build.  Perhaps the best way to describe it is with this diagram I prepared for a talk, comparing SBOL 2.0 to previous standards:


FASTA is about as bare-bones as it comes: pretty much just listing out the DNA sequence that you want.  GenBank lets you annotate that sequence with descriptive information about what the different parts mean, and SBOL 1.0 lets you describe the structure of a design hierarchically in terms of annotated sequences that get combined together as "parts" to make bigger designs.  SBOL 2.0 lets you talk about function as well, describing the way that these parts interact with one another to create the overall behavior of a design.

Conceptually, it's fairly simple, but in practice it took several years to work out and the arguments are not yet over.  The document that we produced is more than 80 pages long, and we're still tinkering with bits and pieces as we try to understand all of the consequences of what we've built.

SBOL is heavy on my brain right now because for the past week, I've been at the COMBINE meeting, where the communities for SBOL and a number of other biological standards meet up to try to improve their systems, work on interoperability, etc.

This is still not something I ever thought I would be doing with my life.  Even now, in my prejudicial mind, standards design is still something done by grey little people who care passionately about trivial and boring things.  I struggle with this, because I look at my work in this area and simultaneously feel that it is highly important and mind-numbingly stultifying to anybody who isn't actually in the room arguing passionately about the potential long-term consequences of adding a single arrow to a diagram.

A case in point: one of the things that I'm most proud of this week was the updated governance document I drafted, and my mediation of discussion on this document, which helped tune it to become widely accepted; the updated version now appears well on its way to official approval by a formal community vote.  So, apparently I am proud of work I've done on adjusting the methods for making decisions regarding an experimental standard for interchange of information about biological designs that will allow faster prototyping of improved systems for biomedicine, biomanufacturing, etc.  That's at least five levels of separation from anything that really affects the larger world. Looked at in that light, this is clearly the very definition of obscurity.  And yet, let me spell it out in another way...
  • Good governance, which gets openness, power, and decision-making right, is critically important for the health of a community, and a number of little warning signs have indicated that the SBOL community needed to adjust its governance to match the way the group has developed and grown.
  • If the SBOL community governs itself effectively, then it will make better decisions that are more likely to lead to a useful and effective standard.
  • If the SBOL standard works well, it will make it a lot easier for people to develop good biological engineering tools.
  • Those biological engineering tools will make it a lot easier to safely and predictably engineer with and for living organisms.
  • Used responsibly, those capabilities can help make all of humanity healthier and safer, as well as improving our ability to manage our environmental impact on a global scale.


This nail I've driven in is very small and unimportant, almost certainly, and yet it matters.  It matters a lot, and not at all, all at the same time.  And I suppose that's just the way the world works, on a planet with seven billion interconnected and increasingly technologically powerful individuals.  Our civilization is remarkably strange and obscure in its operation, and I'm glad when I find satisfaction in the parts I play.

Thursday, October 08, 2015

Publication delays ARE aimed at manipulating impact factor!

A few months ago, I wrote a post with a question: Are publication delays aimed at manipulating impact factor?

Today, I have an answer to that question: yes.

A recently published article, "Editors’ JIF-boosting stratagems – Which are appropriate and which not?" (h/t RetractionWatch) investigates strategies that journal have been using to boost their impact factor and explicitly calls out what it calls the "online queue strategem."  The article is paywalled, so let me summarize here.  In addition to reviewing some of the better-known and clearly unethical practices used by some journals (e.g., forcing citations on authors, citation cartels), the paper carefully dissects the effects of having a long "online early" period of publication, finding four main effects:

  • Papers accumulate citations before "official" publication (multiplying by ~1.5 to 2)
  • Citation rates typically peak 3-4 years after publication, so shifting the time selects for a better citation date (adding another ~50%)
  • Queue order can be manipulated to publish the papers picking up the most citations earlier, (adding another ~30%)
  • Calendar-year boundaries mean that papers in early months count more than papers in later months, so strategic organization of early-month issues can further boost citations (adding another ~30%).

All of this adds up to around 5-fold potential distortion in impact factor.  Since the dynamic range of most journals is only around 0.5 to 10 anyway and even the very highest impact factor journals top out at ~50, this renders that most precious number completely useless.

Now, it's possible that many journals aren't deliberately and strategically manipulating their queues, meaning they'll only get about a 2x boost in impact factor from queuing.  So what?  It still means that impact factor is going to be highly distorted and basically only good for distinguishing journals into three categories: "glamour journal", "normal journal", and "ignored journal" (less than about 0.3).

Ironically, the article itself is dated February, 2016.

That's it: it's clearly time to adopt the wise strategy of my favorite satire journal, the Proceedings of the Natural Institute of Science.  Their current impact factor? "Leadership"

Wednesday, October 07, 2015

Scale-free distribution of payoffs in science

One of the things I've been enjoying these days has been answering questions on the Academia site on StackExchange.  This question-and-answer site is part of the vast network of Q&A sites that have flowered out of the wildly successful StackOverflow, which is pretty much the best source for coding help on the internet.  The model is that people ask question about the topic, e.g., academia, and other folks turn up and provide answers, and then you get or lose Fake Internet Points depending on whether the crowd thinks it's a good answer.  It's surprisingly effective and also, for me at least, pretty enjoyable and kinda addictive.

Anyway, I answered one this morning that made me think a lot, and I thought that I might share my thoughts here as well.  The question was simple, fundamental, and ill-posed: "What is the distribution of payoffs in research?"  Basically, the person is wondering whether every experiment is a roughly equivalent step forward, or whether some are much more valuable than others, and if so whether there's some sort of power-law relationship between topic, funding, and value of result.

This is ill-posed, because the whole notion of "payoff" is extremely vague and probably the wrong question to ask, but it really made me think.  My response, which I'd like to share with you, was this:

There's a vast amount of ill-definition and uncertainty wrapped up in your question... and yet despite that, the answer is almost certainly yes, there is a power-law distribution.
I'm going out on a limb a bit here, because I'm not building on any published analysis that I'm aware of. However, a little analysis of limit cases and fundamental principles can take us a long way here. Let us start with two simple and relatively uncontroversial statements:
  1. Better experimental design leads to better results. It seems self-evident that if you make a bad choice in designing and experiment, it's not going to get you the interesting results you want. At the micro-scale, some choices are clearly better than others, and some are clearly worse.
  2. Sub-fields appear, expand, shrink, and die. As I write this, CRISPR research is hot, and a lot of people are finding interesting results there, and accordingly that field is rapidly expanding. Nobody is doing research on the luminiferous aether because it's been discredited as an idea. Nobody is trying to prove that it's possible to generate machine code from high-level specifications because Grace Hopper did that in the 1950s, when she invented the compiler, thereby initiating what is now a fairly mature and stable research area.
So clearly, no matter how one defines "payoff," any sane definition will see a highly uneven distribution of payoffs both the micro-scale of individual experiments and at the fairly macro level of sub-fields.
Finally, we need to recognize that "significance" is a matter not only of objective value, but also of communication through human social networks. This means that the same result may have wildly different impacts depending on the methods and circumstances of its communication. The history of multiple discoveries in science is ample evidence of this fact; one nice illustrative example is the way in which Barbara McClintock's work on gene regulation was largely ignored until its later rediscovery by Jacob & Monod.
So, we have variation and we have interaction with human social networks, which tend to be rife with heavy-tailed distributions. All of this says to me that it would be remarkable if there were notsome sort of power-law distribution regarding pretty much any plausible of definition of impact, significance, and investment. For these same reasons, I think it would also be surprising if one can make any more than weak predictions using this information (e.g., "luminiferous aether research is unlikely to be productive", "CRISPR is pretty hot right now").
And the devil, of course, is in the details...

Saturday, October 03, 2015

Tribute to driving in Germany

I'm sitting in the Frankfurt airport right now, having just finished driving two hours on the Autobahn up from Schloss Dagstuhl, in Wadern near the French border.  As an American, I'm used to fairly titanic road networks, but driving on the Autobahn feels different to me: my impression is that while Americans use our roads, Germans really love their roads.

Out in the gently winding hills and valleys of Western Germany, forests and fields flash past traffic freely flowing at 100 miles per hour.  The gentle curves are well designed to encourage speed, and many people take good advantage of it.  Yes, it's true, on much of the Autobahn system there is simply no official speed limit (though you'll still get pulled over if the police think you're driving dangerously), and even where there is a limit it usually restricts you only down to 130 kph (a little over 80 mph).

In my little bitty economy rental car, I cruised along comfortably at 110 mph or so in 6th gear (don't even try asking for automatic in Germany), its happy German engineering not making the least complaint about the speed.  At home, my faithful Toyota starts to get very loud and quite unhappy by the time that I reach 85.  Even so, happy-looking people in bigger cars rocketed smoothly past me at significantly higher speed, bound for who knows where at the highest speed available.  And all I know is that my head has got this song on repeat, and I invite you to join along with me and sing: