Wednesday, October 07, 2015

Scale-free distribution of payoffs in science

One of the things I've been enjoying these days has been answering questions on the Academia site on StackExchange.  This question-and-answer site is part of the vast network of Q&A sites that have flowered out of the wildly successful StackOverflow, which is pretty much the best source for coding help on the internet.  The model is that people ask question about the topic, e.g., academia, and other folks turn up and provide answers, and then you get or lose Fake Internet Points depending on whether the crowd thinks it's a good answer.  It's surprisingly effective and also, for me at least, pretty enjoyable and kinda addictive.

Anyway, I answered one this morning that made me think a lot, and I thought that I might share my thoughts here as well.  The question was simple, fundamental, and ill-posed: "What is the distribution of payoffs in research?"  Basically, the person is wondering whether every experiment is a roughly equivalent step forward, or whether some are much more valuable than others, and if so whether there's some sort of power-law relationship between topic, funding, and value of result.

This is ill-posed, because the whole notion of "payoff" is extremely vague and probably the wrong question to ask, but it really made me think.  My response, which I'd like to share with you, was this:

There's a vast amount of ill-definition and uncertainty wrapped up in your question... and yet despite that, the answer is almost certainly yes, there is a power-law distribution.
I'm going out on a limb a bit here, because I'm not building on any published analysis that I'm aware of. However, a little analysis of limit cases and fundamental principles can take us a long way here. Let us start with two simple and relatively uncontroversial statements:
  1. Better experimental design leads to better results. It seems self-evident that if you make a bad choice in designing and experiment, it's not going to get you the interesting results you want. At the micro-scale, some choices are clearly better than others, and some are clearly worse.
  2. Sub-fields appear, expand, shrink, and die. As I write this, CRISPR research is hot, and a lot of people are finding interesting results there, and accordingly that field is rapidly expanding. Nobody is doing research on the luminiferous aether because it's been discredited as an idea. Nobody is trying to prove that it's possible to generate machine code from high-level specifications because Grace Hopper did that in the 1950s, when she invented the compiler, thereby initiating what is now a fairly mature and stable research area.
So clearly, no matter how one defines "payoff," any sane definition will see a highly uneven distribution of payoffs both the micro-scale of individual experiments and at the fairly macro level of sub-fields.
Finally, we need to recognize that "significance" is a matter not only of objective value, but also of communication through human social networks. This means that the same result may have wildly different impacts depending on the methods and circumstances of its communication. The history of multiple discoveries in science is ample evidence of this fact; one nice illustrative example is the way in which Barbara McClintock's work on gene regulation was largely ignored until its later rediscovery by Jacob & Monod.
So, we have variation and we have interaction with human social networks, which tend to be rife with heavy-tailed distributions. All of this says to me that it would be remarkable if there were notsome sort of power-law distribution regarding pretty much any plausible of definition of impact, significance, and investment. For these same reasons, I think it would also be surprising if one can make any more than weak predictions using this information (e.g., "luminiferous aether research is unlikely to be productive", "CRISPR is pretty hot right now").
And the devil, of course, is in the details...

Saturday, October 03, 2015

Tribute to driving in Germany

I'm sitting in the Frankfurt airport right now, having just finished driving two hours on the Autobahn up from Schloss Dagstuhl, in Wadern near the French border.  As an American, I'm used to fairly titanic road networks, but driving on the Autobahn feels different to me: my impression is that while Americans use our roads, Germans really love their roads.

Out in the gently winding hills and valleys of Western Germany, forests and fields flash past traffic freely flowing at 100 miles per hour.  The gentle curves are well designed to encourage speed, and many people take good advantage of it.  Yes, it's true, on much of the Autobahn system there is simply no official speed limit (though you'll still get pulled over if the police think you're driving dangerously), and even where there is a limit it usually restricts you only down to 130 kph (a little over 80 mph).

In my little bitty economy rental car, I cruised along comfortably at 110 mph or so in 6th gear (don't even try asking for automatic in Germany), its happy German engineering not making the least complaint about the speed.  At home, my faithful Toyota starts to get very loud and quite unhappy by the time that I reach 85.  Even so, happy-looking people in bigger cars rocketed smoothly past me at significantly higher speed, bound for who knows where at the highest speed available.  And all I know is that my head has got this song on repeat, and I invite you to join along with me and sing:

Why I love iGEM

Last weekend was the annual iGEM jamboree---that is, the International Genetically Engineered Machine (iGEM) Competition. I poured in about 60 hours of my time into the contest over the course of 3.5 days, and by the end I was exhausted, both physically and mentally, but feeling absolutely elated and on top of the world, raring to go for another one in 2016.

What had I seen, and why was I so excited?  Well, iGEM is a magnificent and unique event, a gathering of students from every continent, from high school on up, all driven by a passion for biological engineering and simply overflowing with creativity.  Each team spends the summer working together on a project that they create, and in the fall they come together to have a big party, where everybody gives talks on what they've done and the best few are recognized in front of everybody for their superlative accomplishments.  There's lots of silly things, lots of over-ambitious ideas that don't get too far, and lots of nice little steps and learning by the students.

And in the middle of it all, some damned good science gets done as well.

Last year, I co-founded a new track at iGEM focused on measurement.  Yes, we're back to that again: my obsession with terribly unsexy rulers.  We had some very good teams last year, and this year again there were a bunch of excellent projects in the measurement track.  And this year, one of those projects stood out head and shoulders above all the rest.

The team from William & Mary, a small but long-standing and excellent public college in Virginia, chose to focus on an important but subtle problem: quantification of noise in gene expression.  Building on recent work in the area, they dug into the problem and ended up with a simple and easy to use kit for measuring this noise, then applied it to quantify noise for a few of the most widely used biological components in the iGEM parts registry.  Very deep and very geeky, but it matters a lot.  If we want to have safe and reliable genetic engineering, we need to be able to predict what will happen when we modify an organism, and this strikes right at that heart of that problem by measuring predictability.

But that wasn't all: they also worked with their county school system to develop a curriculum for synthetic biology.  It's magnificent, and you can get a copy for free online.  Inside this 80-page document, you can find 24 age-appropriate activities, from "DNA twizzlers" for 1st gradesr to Monster Genetics a couple years later (fire-breath is a dominant trait, but cyclopses are recessive), building all the way to adult-level work in high-school like PCR amplification of DNA and bioethics analysis.  The interactions with teachers really show, as the lessons are not only pretty but also give clear goals and a materials list and expected cost per student (usually a whole class can be supplied with a just a few dollars of groceries or arts & crafts supplies).  Even more remarkably, teachers have already begun enthusiastically adopting it, both throughout their county and in other states and nations.

The William & Mary team gave clear, understated presentations that simply let their work shine through, and the whole community recognized it, ultimately first giving them a chance to present as a finalist in front of all the thousands at the convention center, and finally awarding them the competition's top prize (along with a bunch of others as well).  This simple yet deep set of work comes from a team whose school doesn't even break the top 100 in US News' ranking for biology, and shows the power of careful and thoughtful work in science.  The Washington Post may have been too confused to even mention them, but their university is quite elated, and its staff took the time to understand and write a clear and accessible article about their project.

To me, all of this is a vindication not just of the work I've put in organizing and promoting measurement at iGEM, but of the entire scientific process.  Good things can come from unexpected places, and sharp minds thinking careful thoughts can be recognized and receive the recognition they deserve.  Yes, there are problems in the scientific world---quite many, in fact---but this is why we must fight to preserve and promote the scientific ideals, and to keep making that world more diverse, more inclusive, and more able to recognize and promote the potential to improve our world and make a difference.  This is why I love iGEM, why I'm proud to be involved and for what part I've had in helping to enable this, and why I'll be back again for more in 2016.

William & Mary, iGEM 2015 winners, with the Measurement Track committee

Wednesday, September 30, 2015

Aggregate Programming!

It's finally out: our IEEE Computer article, "Aggregate Programming for the Internet of Things" (free preprint) is up and online where all can get it for free.  Don't be fooled by the name: this isn't really just about the our increasingly networked possessions  ("the Internet of Things"), it's a much more general paper.  In fact, this is the first place we've really put all the pieces of our last few years' distributed systems research together, into a generally accessible article that clearly introduces a better framework for building distributed systems.  Please allow me to introduce the aggregate programming stack:

Aggregate programming stack, with examples from a crowd safety application.
In computer networks, the OSI model is a "stack" of abstraction layers that separate different aspects of computer communication.  The browser you are reading this on, for example, is probably obtaining it via HTTP at the Application Layer, routed to you via TCP/IP at the Transport and Network Layers, respectively, with the last link sent to you over something like 802.11 or Ethernet handling the Data Link and Physical layers.

Our aggregate programming model takes a similarly layered approach to the problems of designing networked systems, breaking these often extremely complex problems into five layers.  From bottom to top, these are:

  1. Device: this is the collection of actual electronic devices that comprise the system, with their various built-in sensors, actuators, ability to communicate with one another, etc.
  2. Field Calculus: this layer abstracts the devices into a simple (but universal) virtual model that can freely be mapped between an "aggregate" perspective in which the whole system acts like a single unified device, and a "local" perspective of individual device interactions that implement this model.
  3. Resilient Coordination: this layer consists of "building block" algorithms (implemented in field calculus) that provide guarantees that systems will be safe, resilient, and adaptive in various ways.
  4. Developer APIs: useful patterns and combinations of building blocks are then named and collected into application programming interfaces (APIs) that are easier to think about and program with.
  5. Application: Finally, distributed systems can be much more easily constructed, using the APIs just like one would any other single-machine library.
Constructing this stack factors the problems of distributed systems development into separable components, each much simpler than trying to tackle the whole complicated mess at once.  If you just want to build applications, you just need to learn the Developer API layer and work with that, just like web programmers learn about HTML and Javascript.  If you want to work on resilient algorithms, on the other hand, you get involved with the plumbing at the Resilient Coordination layer, and if you want to use the stack on a new device, you implement a copy of the interface required for a Device by an instance of the Field Calculus layer.

I'm very proud of this work, and think it's got a potential to really change the way that people deal with complex computer networks.  For the programmers amongst you, dear readers, I suggest you check out both this paper and our (still somewhat rough) implementation of field calculus in Protelis.

Saturday, September 19, 2015

A Golden Boston Sunset

Dear readers,

Some days, it's just good to be alive, and a thing comes out of nowhere unexpectedly to remind you of that fact.  Today, as my flight was gliding down into its final descent into Boston, the air was almost perfectly clear and the sun was just at that magic moment in its descent where everything begins to be golden and shadows stretch out just enough to give the third dimension of everything an extra bit of special emphasis.  As I gloried in the texture of the light, my camera came out and I snapped away---not blocking myself from an enjoyment of this sight, but finding that the aim to capture gave me extra focus and appreciation for the details.

My dear readers, I wish to share this joy with you, in the form of a few of the best moments of imagery I captured.  May this lighten your day as it has lightened mine.





Monday, September 14, 2015

A Tale of two CRISPRs

Last year, I had my first "glamour journal" publication, as second author of a Nature Methods paper on a new family of CRISPR-based synthetic regulatory devices.  Actually, I had two "glamour" publications---the other was a Nature Biotech paper on the SBOL language for communicating biological designs with 32 authors, the biggest collaborative publication I've been involved in to date.  That's a tale for another post, however---this one's all about CRISPR, CRISPR, CRISPR.

For those who haven't encountered the wonderful hype-storm around CRISPR, the acronym expands to the highly non-mellifluous "clustered regularly interspaced palindromic repeats," which tells you virtually nothing about why it's cool.  The reason it's cool is because one of the things this awkward acronym refers to a protein ("Cas9") that docks with fairly arbitrary "guide RNA" fragments in order to go act on DNA that matches those sequences.

Core CRISPR mechanism: Cas9 protein binds to gRNA, which targets the protein to a matching DNA sequence

Protein design is really hard, but DNA and RNA design has become reasonably straightforward, so CRISPR is an awesome mechanism: it lets us target a (fairly) predictable protein effect to pretty much any piece of DNA that we want.  People have used it for editing DNA, which has previously been done with lots of other mechanisms, but gets much easier with CRISPR (hence the recent controversies you may have seen in the news around human genetic engineering---the changes we can do aren't any different, they're just a lot cheaper, which is a meaningful difference of a different sort).

Our paper last year showed for the first time how to use the CRISPR mechanisms to make potentially large numbers of strong biological logic gates.  This is important because one of the big things that's been holding synthetic biology back is the difficulty in building reliable computation and control systems inside of cells.  We've known for a long time that biological computing is possible, but there's only been a handful of decent computational devices, and no good ways of making more.  Now, within the last few years, there have been several different families that have emerged, including TALE proteins, homolog mining, invertases, and now, with our paper, CRISPR repressors.  Our CRISPR repressors are nice because they can potentially easily generate thousands of high-performance devices and implement all sorts of complex computations, something that nobody currently has a clear approach for with any of the other families.

Diagram of one of our CRISPR repressors: a modified Cas9 protein (blue box) acts as "power supply" for an inverter logic gate implemented by having the gRNA (orange box) regulate a synthetic promoter (blue arrow). The important things to know are 1) the orange box and blue arrow are easy to design and we can make lots of them that don't interfere with one another, and 2) the blue box can potentially power lots of these gates at the same time.

So I was (and still am) very excited about this publication for two reasons: first because I think it's a big step forward scientifically, and second because it's in a big-name venue that lots of people are likely to pay attention to and where it's more likely to have a big impact on scientific practice.

Just last week, I had my second paper in Nature Methods, led by my same awesome collaborator, Samira Kiani, and following on the subject: this time, our paper shows how to use CRISPR devices to both compute and edit genes in the same circuit.  My reaction, however, has been much more mixed to this publication.  Don't get me wrong: I'm really happy to be published in a high-ranked journal again, and I really enjoy working with Samira (soon to upgrade from Dr. Kiani to Professor Kiani!), who I find an insightful and diligent collaborator and whose skills I think complement my own quite nicely.  Maybe it's just that I can't be so deliriously excited about getting published in a journal a second time?  I'm also not as excited about these results: it's a nice twist on previous results and a useful new capability, but in return we lose some of the device efficacy. Overall, though, I just don't feel like this paper is a game-changer in the way that our first paper might prove to be.

Still, it matters, and it's a step forward for all of us.  Soon, we will meet, celebrate this success with a toast and a fine dinner, and plan our next venture toward transformation of the world and toward posterity.

Sunday, September 06, 2015

Perhaps my least interesting publication ever

Just recently, I was listed as first author (out of five), on what is perhaps the least interesting scientific publication in my history as a researcher.  This includes even semi-embarassing old rants from when I was a young and arrogant graduate student---those at least give some sort of plausibly interesting perspective on what I was thinking about at the time.  Not to say this document isn't important: I think it was definitely worth the time and effort, and is useful.  That doesn't necessarily mean anybody will derive any particular joy or pleasure from encountering it.

So, what is this deadly dull publication that I've for some strange reason decided to advertise so loudly on the Internet?  Its formal name is: BBF RFC 107: Copyright and Licensing of BBF RFCs. This takes a little bit of explanation, so bear with me and please try not to fall asleep too quickly: one of ways that people in the synthetic biology community share their work is by posting open "Request For Comment" documents (RFCs)---essentially draft standards, following the main model used for developing the Internet.  These are cataloged by the BioBricks Foundation, hence BBF RFCs.  The first of these, BBF RFC 0 (yes, there were computer scientists involved, and we like to count starting with zero), sets out the process for how to submit a new RFC.  A few months ago, I noticed that the original handling of copyright had gotten out of date with respect to some current practices in accessing scientific documents online and current preferences for open standards development.  I raised these issues with the BBF RFC maintainers, and we figured out a legal "patch" for BBF RFC 0. The end result of all this is a 1.5 page document that makes two small changes in how new BBF RFC documents are handled:

  • The document is actually marked with a modern open copyright license, and
  • The authors share copyright with the BioBricks Foundation, rather than transferring it.

Now, unfortunately, the parts of BBF RFC 0 that we didn't replace weren't followed correctly in setting forth this RFC, which has caused some trouble with another BBF RFC that I'm involved in, but that story's even less interesting, and I'm sure it will all get sorted out eventually.

In case your eyes have well and truly glazed over, let me sum that all up more simply: I noticed a little thing about copyrighting certain scientific documents that needed tweaking.  By a quirk of process, doing so had the side effect of creating an archival scientific publication.

So, was it worth it?  Absolutely: it didn't take much time, and copyright is one of those things that it's often worth paying close attention to, because if you screw it up as a community, you can accidentally wind up poisoning all sorts of things down the line, if nasty people decide to try to take advantage of loopholes or cautious organizations get blocked from doing things by technicalities.  I'm just quite amused that this ends up in my list of publications as outwardly indistinguishable from BBF RFCs that took many people years of work and that gather lots of citations.  It's also kind of funny from a "what do scientists do all day" perspective.

But, for the love of all that you hold holy, don't read the document unless you actually need to.

Sunday, July 19, 2015

Foul-Mouthed LARPing Advice

MIT Assassin's Guild dart-gun combat, March 2006
OK, dear readers, I'd like you to indulge me in one more trip down memory lane.  In cleaning my electronic life and the crud accumulated on my machine, I recently came once again across an ancient and delightful (to me at least) document.

You see, one of my primary indulgences back when I was an undergraduate was designing and playing live action roleplaying games (LARPs).  At MIT, this meant getting submerged into the very peculiar and "hard-core" culture of the MIT Assassin's Guild, where we would regularly play 10-day-long games with 60 people (as well as many shorter and smaller forms).  I quickly gravitated to the writing side, much more enjoying playing God and setting up the arena for others to contend within, rather than actually participating in the conflict myself (my personality is actually rather conflict-averse, although you wouldn't readily know it from my behavior as an arrogant young man). Even now, I still hold a record as one of the all-time most prolific authors in the 35-year history of the Guild, having written and run 25 different games, mostly over the period of 1998 through 2004, though I continued producing approximately one game per year all the way until 2013, when I moved to Iowa. Looking at the coincidence in timing between my decline of LARP output and my increase of scientific output, I suspect that actually it is no coincidence. Writing LARPs continued to be a very important both social and creative outlet for me, however, evolving with my life and circumstances and increasing interest in things outside the generic fantasy or science fiction boxes (the last game I co-authored was based on the Hindu legend of the churning of the ocean of milk).

All of this is preface and circumstance to explain the document that I intend to present to you.  It needs some explanation, because it's definitely got a Mature Content Advisory on it.  You see, back in 2000, right around the height of my youthful arrogance and joy in being transgressive, I wrote a long, flamey rant called the "Definitive Guide to Writing Guild Games for the Rest of Eternity," putting down my own extremely biased views about writing LARPs for the Assassin's Guild.  It was a foul-mouthed, irreverent, and joyful trip through my perspective of the time, and I circulated it privately amongst those who would likely not hold it against me too much, but it was never intended for wider dissemination.  Then, in 2005, my friend Joe Foley started putting together an attempt at a comprehensive guide to gamewriting, which he called the Mechanicomicon (transparently referring to H.P. Lovecraft), and he persuaded me to update and expand the document with my additional years of experience.  It didn't take much persuasion, and we decided to maintain the original tone of the document for publication, even if its content became a little more level-headed and we removed a couple of pieces of obvious and unnecessary slander.

And this, dear reader, is the document that I just rediscovered, still essentially in the same form as a decade ago (though it's had a few additional tweaks and bug-fixes since then).  I want to share it with you as a slice into another world, a world that I fondly remember the times I spent, and also a slice of my own mind.  LARPing prepared me for science in some surprising ways, especially in the challenges of organizing events and of managing large and complex projects across multi-year time-spans.  If you yourself are a gamer, you might find some interesting thoughts in there as well, though the Assassin's Guild is also a very peculiar culture of its own, and some of the references will be, I'm sure, completely and totally impenetrable.  All I ask, dear reader, is that you not judge me too harshly for things that were written long ago, and a side of me that was never intended for professional presentation.

In the continuing spirit of irreverence, however, I have tucked this document discretely at the very bottom of my professional webpage, formatted much like a scientific citation, and will delighted if I can get Google Scholar to pick it up, and even more delighted if it can pick up citations.  Just my little scientific prank, if not quite on the delightful scale of F.D.C. Willard.
Team Hufflepuff in a tense moment of planning during the Harry Potter 10-day game, January 2011.

Sunday, July 12, 2015

A fond goodbye to most of fiction

I've recently noticed a strange and surprising shift in myself: I appear to be losing much of my interest in fiction.

This is an odd and somewhat disquieting phenomenon for me to notice in myself, since I have been a committed fan of the stuff, and especially science fiction and fantasy, for such a long, long period of my life.  Like many a nerdy child before me, I got into Tolkien and Asimov at an early age, and most of my fiction reading growing up was science fiction.  Then I went away to college and discovered the MIT Science Fiction Society, and my cementation in the genre was complete.  For more than fifteen years, my need to buy fiction was almost entirely abrogated, and I gave back as a volunteer librarian, holding hours, processing part of the bookflow, and eventually writing reviews as well.
Processing books at MITSFS: here we were comparing donated books to current copies in order to decide which was in better condition and should be kept.
The first big shift in my fiction consumption was shifting from a walking/public-transit commute to bicycling or driving.  When I walked, I would read as I walked, and I had gotten quite good at it in several years of commuting while reading.  When I added a subway ride to my commute it stayed the same.  Towards the end of grad school, I started bicycling, however, and that meant that I could no longer even vaguely sanely desire to read while commuting.  When I moved from MIT to BBN, public transit become a much worse option for commuting, and it shifted to bicycling and car, and my commuting reading shifted entirely to podcasts and audiobooks.  There went at least six hours a week of reading, right there.

The next big shift in my fiction consumption was having a child.  When Harriet was born, my already scarce time dwindled even further, and I started needing to multi-task at home as well.  No longer were there good long times to read just curled up on a couch or in bed. Instead, I was generally doing something else, whether it be laundry or cooking or errands or something else, and more reading time converted from visual to auditory.

In the most recent shift, however, I've simply stopped consuming fiction at all for the most part.  Of the last 10 audiobooks that I have purchased, five have been non-fiction, while the others have been science fiction / fantasy.  That might not seem like a lot, but five years ago there might have been a single non-fiction in the lot.  Even more tellingly, of those ten books, I have finished all of the non-fiction and even listened to three of them twice.  Of the fiction, however, I have only even bothered to finish listening to two.  For me, this shift is frankly shocking.

I think that this most recent shift has something to do with the accompanying transitions in how I process fiction when it's auditory and not visual.  When reading a physical book, I read quickly, and move quite quickly through a story, more quickly when I'm not enjoying parts of it.  With an auditory book, on the other hand, the physical fact of the narration speed being so much slower than my usual reading speed means that I am given time to think about and process every sentence as it comes to me.  In turn, this leads to me thinking more deeply about the story I am listening to.  This ends up meaning several things: first, I need a much better quality of author to even keep me interested, since I am forced to consider quality of the prose more closely.  Second, it means that I am thinking in general more about the narrative and story structure, which means that I am generally learning more about not just this particular book that I am reading, but about storytelling as a craft and art.  That, in turn, has ended up bleeding over into other sorts of genres as well, and I am now having a much harder time enjoying anything in fiction (or fictionalized) in television and movies also.  I end up spending more time critiquing the story and looking at its structure than I do enjoying it (again, unless it is very well done indeed), and so a man who once would spend a whole day on the couch watching Cartoon Network is generally now bored by the second episode of something new, when I can see the signposts marking out the path for all the next ten episodes.

I think that I am finding non-fiction more interesting for exactly the same reasons that I am finding fiction less interesting.  Instead of giving me places to ponder the gaps in the world-building of the author and the hollowness behind the scenes, a good non-fiction leaves me wondering about the interconnections and larger implications of the things I'm thinking about.  For example, one of the non-fiction books that I have recently "read" multiple times is a lecture series called "The Barbarian Empires of the Steppes." This piece of history, despite its often degrading into a list of names of peoples who had fights at places, has really filled in a gap in the way that I have been educated about the world.  In my Euro-centric education, the world was anchored in the Mediterranean, with the rise of Islam off on one side and India and China as far away islands to talk about in very separate ways (mostly having to do with when Europeans turned up bent on conquest there).  In fact, however, the great land mass of Asia has always tied these lands together, as well as the worlds in the middle, such as Transoxiana, that were never even mentioned in my education.  This lecture series has shifted the way that I can understand our history of Eurasian civilization, and of many of the ways that I can understand the drivers backing current events, even as I have critiqued the scholarship and prejudices of the author in my head while listening.

In fiction, also, the things I still enjoy are works by authors such as C.J. Cherryh, who many find challenging because of the way that she submerges you so tightly in her characters' viewpoints and refuses to give much introduction to the worlds or technologies that they deal with.  A story by her is thus so densely packed with things to think about that I can read or listen to it many times and still be thinking about the implications and dynamics of the worlds that she has introduced, as well as just the general structure of the story.  A few words here and there can give a glimpse of an iceberg of relationships within a whole society, just as in our own world a single painful news story may become a touchstone for a much, much broader movement.  It's just that authors such as her are rare, and I treasure them on those occasions when I find them.

I suppose that in a way I might consider all of this just natural progress of maturity, but I don't know.  We don't necessarily get bored of all the things we love, not when those things can symbolize comfort and familiarity and home.  But when it comes to fiction, I think I'm moving on to wanting much more depth than once had been able to make me very satisfied.

Friday, July 03, 2015

The Scientific Richter Scale

Perhaps it's unhealthy statistical fixation, but I tend to mull over scientific citations, those being one of the important factors in the multi-currency economics of science.  From my thoughts, observations, and wanderings on Google Scholar, has gelled the following rough interpretation of scientific impact in terms of the rough order of magnitude of citations:
  • Order 100 citations: This paper is languishing in darkness and obscurity.  That doesn't mean it's a bad paper, just that nobody particularly cares about it.
  • Order 101 citations: Either some people have noticed this paper, or else it's part of a strong ongoing line of research and is picking up a lot of self-citations.
  • Order 102 citations: This paper is having a significant intellectual impact.  People have noticed it and, whether or not they are actually using the ideas within it, those ideas are having a noticeable effect on the scientific discourse.
  • Order 103 citations: This paper contains something that lots of people are actually finding that they need and are putting to use.  It is no coincidence that many of the "Top 100" papers identified by Nature are methodology papers (though the fact that the list omits Claude Shannon's exceedingly highly cited paper on information theory is another great example of just how shoddy citation databases are in their coverage of computing fields).
Of course, there's no sharp boundaries between levels, nothing to say that a paper with 300 citations (closer to magnitude 2) is really all that different than one with 350 citations (closer to magnitude 3). Perhaps it will be best to think of this as a scientific Richter scale, measuring the intellectual impact of individual scientific events [papers].  Right now, I don't think it makes much sense to look beyond three orders of magnitude because there are so few papers out there, but then, there aren't many magnitude 8.0 earthquakes either.

Thinking about this in terms of the Richter scale immediately sends the complexity theorist in my brain to think about scale-free distributions and log/log plots, and so I made a plot of my own current personal citational spectrum, according to Google Scholar.  It looks like this:
Yes, that certainly looks relatively linear on a log/log scale.  But then, so do so many things.  The the linear fit points out that there is definitely a big kink in the line, so it's not a straight scale-free distribution.  That means... I have no idea.  But this was certainly a fun way to fritter away half an hour while starting my vacation and watching my daughter sleep.