Shrinking 'ridiculous' data sets to manageable size

May 14, 2009 By Bill Steele

Two decades ago a renowned statistician described a computer data set of 1 billion bytes as "huge" and 10 trillion bytes as "ridiculous."

Today, thanks to the use of computers to collect and generate data, such ridiculously large data sets are common, from genome databases to search engine logs to Wal-Mart sales data. But the ability to monitor and process the data has not kept up with the ability to create it.

With a new three-year, $551,508 Young Investigator Award from the U.S. Office of Naval Research (ONR), Ping Li, Cornell assistant professor of statistical science, is taking a new mathematical approach. His goal: to "shrink" massive data sets into manageable approximations that can be processed in a reasonable length of time to detect such anomalies as denial-of-service attacks on the Internet or to enable computers to learn from experience for such applications as natural language processing, Web searching and computer vision.

"Instead of storing the whole data, we compute and store a sketch of the data, which is small enough to fit in the memory and still contains enough information to recover crucial relationships of the data," Li explained.

From the resulting sketch, Li says that it is possible, for example, to compute a quantity known as the Shannon entropy, which is, roughly, a measure of the degree of uncertainty in a body of information. A change in this would warn engineers of an anomaly such as a network failure, a large transfer of money or perhaps terrorist chatter. Li also plans to develop and publicly distribute software that can be used as part of machine-learning applications on massive and high-dimensional data sets.

The ONR Young Investigator Program identifies and supports academic scientists and engineers who have received a doctorate or equivalent degrees within the past five years and who show exceptional promise for doing cutting-edge research.

Provided by Cornell University (news : web)


   
Rate this story - 4.6 /5 (5 votes)


May 14, 2009 all stories

Comments: 0

4.6 /5 (5 votes)

  • hide
  • Related Stories

  • New tool enables powerful data analysis
    created Jan 08, 2009 | popularity not rated yet | comments 0
  • New grant supports emerging field of massive data analysis and visual analytics
    created Aug 06, 2008 | popularity not rated yet | comments 0
  • Statistics Professor Hides Pictures, Messages in Problem Solutions
    created Apr 11, 2007 | popularity not rated yet | comments 0
  • Leading-edge data analytics and visualization enable breakthrough science
    created Apr 10, 2009 | popularity not rated yet | comments 0
  • Model helps computers sort data more like humans
    created Aug 25, 2008 | popularity not rated yet | comments 0



  • hide
  • Relevant PhysicsForums posts

  • Computer 5V or 0V output to Sensaphone Express II
    created Feb 04, 2010
  • Ti-89 ROM Image
    created Jan 29, 2010
  • TV ads
    created Jan 29, 2010
  • Apple introduces latest iNonsense
    created Jan 27, 2010
  • More from Physics Forums - Computing & Technology

Other News

The power of 'random'

The power of 'random': 'Seemingly loopy' technique could dramatically improve communications networks

Technology / Computer Sciences

created 19 hours ago | popularity 4.8 / 5 (9) | comments 5 | with audio podcast

A radical new approach to the design of communications networks, called "network coding," promises to make Internet file sharing faster, streaming video more reliable, and cell-phone reception better -- among ...


'Revolutionary' water treatment units on their way to Afghanistan

Technology / Engineering

created 13 hours ago | popularity 4.4 / 5 (7) | comments 5 | with audio podcast

The United States Army has taken delivery of the first two units of a "revolutionary" waste-water treatment system that will clean putrid water within 24 hours and leave no toxic by-products, according to scientists at Sam ...


Android

Google developing a translator for smartphones

Technology / Software

created 20 hours ago | popularity 4.8 / 5 (9) | comments 3 | with audio podcast report

(PhysOrg.com) -- Google is developing a translator for its Android smartphones that aims to almost instantly translate from one spoken language to another during phone calls.


Imec and Holst Centre achieve breakthrough in battery-less radios

Imec achieves breakthrough in battery-less radios

Technology / Semiconductors

created 14 hours ago | popularity 4.9 / 5 (14) | comments 1 | with audio podcast

At today's International Solid State Circuit Conference, Imec and Holst Centre report a 2.4GHz/915MHz wake-up receiver which consumes only 51µW power. This record low power achievement opens the door to battery-less ...


In Utah, company aims to store energy in air

Technology / Energy

created 23 hours ago | popularity 3.4 / 5 (10) | comments 2

A Utah company plans to dig a series of underground caverns that it hopes to one day fill with compressed air, releasing it to generate electricity by turning a turbine and solving one of the most vexing problems facing the ...