Prospecting for Scientific 'Gems' with Google

May 3rd, 2006 Google Page Rank -- PhysOrg.com

Science is all about quantitative measurement, so it should come as no surprise that scientists have a long tradition of measuring their influence on each other. Traditionally, the most important measure of scientific impact has been the number of citations an article receives -- but this method’s chief virtue is its simplicity. The authors of any given paper are much more likely to reference recent works, so a result whose importance is not immediately recognized can end up with a much lower citation index than it deserves.

A well-written and relevant paper is typically referenced by two or three dozen papers, mostly from researchers working on the same highly specialized problem. A particularly useful or innovative result will often receive a hundred or more citations. Seminal works can achieve over a thousand citations, although they usually take decades to reach that point. While the system works well overall, it has no way of distinguishing a highly relevant citation from the polite mention of a colleague’s work.

Many articles, for example, include an introduction section describing the history and current status of their specialized subject. This section can easily generate up to half of a paper’s references, even though few of the results given mention are actually used.

A database of scientific literature is very similar in structure to the World Wide Web. Just as individual web pages are connected to each other by one-way links, journal articles are connected to each other by one-way citations. The number of external sites linking to a given website, or “in-degree”, is equivalent to the citation index of a given article.

When Google tackled the problem of ranking websites by their influence, it didn’t consider the in-degree to be an appropriate measure. This would make it too easy to inflate the importance of a site by creating a host of useless linking pages. Instead they crafted a customized statistic, the Google PageRank (GPR) algorithm.

To illustrate this algorithm, consider the webpage PhysOrg.com. PageRank finds every other webpage with a link to PhysOrg.com, and divides each neighbor’s GPR by its total number of outgoing links. The GPR of Physorg.com is then calculated as the sum of all these factors. In other words, each site in the network can be thought of as evenly distributing its influence over all the sites that it links to. A page thus gains influence mainly by being associated with other influential pages. (The actual algorithm is a little more complicated, but this is its essential feature.) This method strikes a nice balance between content and connectivity, reducing the influence of high-traffic directories on the sites that they list.

So how can one calculate the GPR of any site when you first need to know the GPR of all its neighbors? It’s not a problem to be solved by pencil and paper! The answer is found through a recursive calculation: every website in the network is initialized with the same GPR, so that a new GPR can be calculated for each website simultaneously. This calculation is repeated until all the values stabilize.

Researchers Patrick Chen and Sidney Redner at Boston University, along with their colleagues Huafeng Xie and Sergei Maslov at Brookhaven National Lab, recently applied the PageRank algorithm to all 353,268 articles published by the Physical Review between 1893 and 2003. It comes as no surprise that on average, GPR correlates nicely with the citation index. More interesting are the outliers—those articles that somehow achieve a high ranking with relatively few incoming references.

After applying PageRank, Chen et al. sorted the papers in this network by their GPR values. Their recent article provides a sampling of famous papers from the top hundred results. Number 85, with only three citations, is a startling poster child of this new approach! The paper in question is a classic example of delayed influence. While it was the first to present a model which today sees widespread use, its result was refined and popularized by other researchers in a separate article. The “child paper” has accumulated 680 citations but makes only ten references to other works itself. The original paper thus collects a large share of its child’s impressive impact.

Nor is this the only example! Among the papers with over a hundred citations, most of the papers with an unusually high GPR are easily recognizable as seminal works. Such works compare favorably in overall influence with the very small population having over a thousand citations.

While “influence” may be easy to measure crudely, it is hard to measure reliably. These results show that although the two methods are comparable, Google’s PageRank algorithm seems to identify important scientific papers more reliably than a simple citation index.

If there is a lesson here, it is this: in giving due credit, one should not be short-cited!

Reference: Patrick Chen, Huafeng Xie, Sergei Maslov, & Sidney Redner 2006, “Finding Scientific Gems with Google”, http://xxx.lanl.gov/physics/0604130

By Ben Mathiesen, Copyright 2006 PhysOrg.com


print this article email this article download pdf blog this article bookmark this article     Digg this Stumble it share on Facebook share on Reddit add to delicious save to Yahoo! bookmarks
4.4/5 after 31 votes


May 3rd, 2006 all stories
Physics / General Physics

Comments: 0
Rank: 4.4/5 after 31 votes

  • Stumble this up

  • Digg this

  • Share it:
  • share on Facebook
  • share on MySpace
  • share on Slashdot
  • rss-newsfeed
  • share on Google
  • share on Reddit
  • add to delicious
  • save to Yahoo! bookmarks
  • share on Windows Live
  • Add to Mixx!
Rating: 4.4/5 after 31 votes


Tags


  • Physicists Demonstrate Quantum Memory with Matter Qubits
    Physicists Demonstrate Quantum Memory with Matter Qubits
    Physics / General Physics
    created Jul 03, 2009 | popularity 4.4 / 5 (17) | comments 1
  • 'Holey' Nanosheets for Wastewater Dye Removal
    Nanotechnology / Nanomaterials
    created Jul 01, 2009 | popularity 5 / 5 (5) | comments 1
  • Jellyfish Robot Swims Like its Biological Counterpart
    Jellyfish Robot Swims Like its Biological Counterpart
    Electronics / Robotics
    created Jun 26, 2009 | popularity 4.4 / 5 (8) | comments 1
  • Could Maxwell's Demon Exist in Nanoscale Systems?
    Could Maxwell's Demon Exist in Nanoscale Systems?
    Physics / General Physics
    created Jun 24, 2009 | popularity 4.4 / 5 (18) | comments 29
  • Living Safely with Robots, Beyond Asimov's Laws
    Living Safely with Robots, Beyond Asimov's Laws
    Electronics / Robotics
    created Jun 22, 2009 | popularity 4.6 / 5 (52) | comments 40
  • Other News

    UQ researchers break the law -- of physics

    Physics / General Physics

    created 51 minutes ago | popularity 4 / 5 (1) | comments 0

    (PhysOrg.com) -- Two UQ Science researchers have proved two famous physical laws that have been widely used for the past 25 years do not always work.


    Scientists create first electronic quantum processor

    Scientists create first electronic quantum processor

    Physics / General Physics

    created Jun 28, 2009 | popularity 4.8 / 5 (54) | comments 41

    A team led by Yale University researchers has created the first rudimentary solid-state quantum processor, taking another step toward the ultimate dream of building a quantum computer.


    Science journals

    How to Spot an Influential Paper Based on its Citations

    Physics / General Physics

    created Jul 04, 2009 | popularity 4 / 5 (9) | comments 6

    (PhysOrg.com) -- At first it may seem that the number of citations received by a published scientific paper is directly related to that paper's quality of content. The higher the quality, the more people read ...


    Fermilab's CDF observes Omega-sub-b baryon

    Fermilab's CDF observes Omega-sub-b baryon

    Physics / General Physics

    created Jun 29, 2009 | popularity 4.7 / 5 (17) | comments 7

    (PhysOrg.com) -- At a recent physics seminar at the Department of Energy’s Fermi National Accelerator Laboratory, Fermilab physicist Pat Lukens of the CDF experiment announced the observation of a new particle, ...


    New insights, and a new angle, on high-temperature superconductivity

    New insights, and a new angle, on high-temperature superconductivity

    Physics / Superconductivity

    created Jun 29, 2009 | popularity 4.8 / 5 (13) | comments 6

    (PhysOrg.com) -- A Princeton-led research team has revealed surprising information about how electron behavior influences the conduction of electricity in a class of high-temperature superconductors. An increased ...