Showing posts with label data simplification. Show all posts
Showing posts with label data simplification. Show all posts

Tuesday, March 29, 2016

CLASS BLENDING: Simpson's Paradox

For the past two days, we've been posting on Class Blending. Simpson's paradox is a special case that demonstrates what may happen when classes of information are blended.


Simpson's paradox is a well-known problem for statisticians. The paradox is based on the observation that findings that apply to each of two data sets may be reversed when the two data sets are combined.

One of the most famous examples of Simpson's paradox was demonstrated in the 1973 Berkeley gender bias study (1). A preliminary review of admissions data indicated that women had a lower admissions rate than men:
Men    Number of applicants.. 8,442   Percent applicants admitted.. 44%
Women  Number of applicants.. 4,321   Percent applicants admitted.. 35%
A nearly 10% difference is highly significant, but what does it mean? Was the admissions office guilty of gender bias?

A closer look at admissions department-by-department showed a very different story. Women were being admitted at higher rates than men, in almost every department. The department-by-department data seemed incompatible with the combined data.

The explanation was simple. Women tended to apply to the most popular and oversubscribed departments, such as English and History, that had a high rate of admission denials. Men tended to apply to departments that the women of 1973 avoided, such as mathematics, engineering and physics. Men tended not to apply to the high occupancy departments that women preferred. Though women had an equal footing with men in departmental admissions, the high rate of women rejections in the large, high-rejection departments, accounted for an overall lower acceptance rate for women at Berkeley.

Simpson's paradox demonstrates that data is not additive. It also shows us that data is not transitive; you cannot make inferences based on subset comparisons. For example in randomized drug trials, you cannot assume that if drug A tests better than drug B, and drug B tests better than drug C, then drug A will test better than drug C (2). When drugs are tested, even in well-designed trials, the test populations are drawn from a general population specific for the trial. When you compare results from different trials, you can never be sure whether the different sets of subjects are comparable. Each set may contain individuals whose responses to a third drug are unpredictable. Transitive inferences (i.e., if A is better than B, and B is better than C, then A is better than C), are unreliable.

- Jules Berman (copyrighted material)

key words: data science, irreproducible results, complexity, classification, ontology, ontologies, classifications, data simplification, jules j berman

Reference:

1. Bickel PJ, Hammel EA, O'Connell JW. Sex Bias in Graduate Admissions: Data from Berkeley. Science 187:398-404, 1975.

2. Baker SG, Kramer BS. The transitive fallacy for randomized trials: If A bests B and B bests C in separate trials, is A better than C? BMC Medical Research Methodology 2:13, 2002

Monday, March 21, 2016

DATA SIMPLIFICATION: Published At Last!



Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.


On March 17, 2016, my book Data Simplification: Taming Information with Open Source Tools was published by Morgan Kaufmann, an imprint of Elsevier. [the Elsevier site indicates that the book is still on preorder, buy you can ignore that]. This past month, I've posted on topics relevant to data simplification. Beginning tomorrow, I'll be moving onto new subjects for this blog site, but I wanted to make one additional comment for anyone who might be on the fence about buying this book.

Most large data projects are total failures (1-21). Furthermore, in my humble opinion, most data projects that are deemed successes at the time of completion are actually failures of a kind, because the data that was collected during the project was abandoned when the project ended. Data shouldn't die. Data should be prepared in a manner that permits anyone (not just the people who planned the project) to confirm the conclusions, to reanalyze the data, to merge the data with other data sources, and to repurpose the data for future projects. To do so, the data must be prepared in a manner that is comprehensible and simplified. My book provides open source tools for creating data that can be used and repurposed, by generations of data scientists.

Enough said! Tomorrow, we move on.

- Jules Berman

key words: computer science, data analysis, data repurposing, data simplification, data science, information science, simplifying data, taming data, jules j berman

References:

[1] Kappelman LA, McKeeman R, Lixuan Zhang L. Early warning signs of IT project failure: the dominant dozen. Information Systems Management 23:31-36, 2006.

[2] Arquilla J. The Pentagon's biggest boondoggles. The New York Times (Opinion Pages) March 12, 2011.

[3] Lohr S. Lessons From Britain's Health Information Technology Fiasco. The New York Times Sept. 27, 2011.

[4] Dismantling the NHS national programme for IT. Department of Health Media Centre Press Release. September 22, 2011. Available from: http://mediacentre.dh.gov.uk/2011/09/22/dismantling-the-nhs-national-programme-for-it/ viewed June 12, 2012.

[5] Whittaker Z. UK's delayed national health IT programme officially scrapped. ZDNet September 22, 2011.

[6] Lohr S. Google to end health records service after it fails to attract users. The New York Times Jun 24, 2011.

[7] An assessment of the impact of the NCI cancer Biomedical Informatics Grid (caBIG). Report of the Board of Scientific Advisors Ad Hoc Working Group, National Cancer Institute, March, 2011.

[8] Heeks R, Mundy D, Salazar A. Why health care information systems succeed or fail. Institute for Development Policy and Management, University of Manchester, June 1999 Available from: http://www.sed.manchester.ac.uk/idpm/research/publications/wp/igovernment/igov_wp09.htm, viewed July 12, 2012.

[9] Brooks FP. No silver bullet: essence and accidents of software engineering. Computer 20:10-19, 1987.

[10] Unreliable research: Trouble at the lab. The Economist October 19, 2013.

[11] Kolata G. Cancer fight: unclear tests for new drug. The New York Times April 19, 2010.

[12] Ioannidis JP. Why most published research findings are false. PLoS Med 2:e124, 2005.

[13] Baker M. Reproducibility crisis: Blame it on the antibodies. Nature 521:274-276, 2015.

[14] Naik G. Scientists' Elusive Goal: Reproducing Study Results. Wall Street Journal December 2, 2011.

[15] Innovation or Stagnation: Challenge and Opportunity on the Critical Path to New Medical Products. U.S. Department of Health and Human Services, Food and Drug Administration, 2004.

[16] Hurley D. Why Are So Few Blockbuster Drugs Invented Today? The New York Times November 13, 2014.

[17] Ioannidis JP. Microarrays and molecular research: noise discovery? The Lancet 365:454-455, 2005.

[18] Vlasic B. Toyota's slow awakening to a deadly problem. The New York Times, February 1, 2010.

[19] Lanier J. The complexity ceiling. In: Brockman J, ed. The next fifty years: science in the first half of the twenty-first century. Vintage, New York, pp 216-229, 2002.

[20] Labos C. It Ain't Necessarily So: Why Much of the Medical Literature Is Wrong. Medscape News and Perspectives. September 09, 2014

[21] Gilbert E, Strohminger N. We found only one-third of published psychology research is reliable - now what? The Conversation. August 27, 2015. Available at: http://theconversation.com/we-found-only-one-third-of-published-psychology-research-is-reliable-now-what-46596, viewed on August 27,2015.

Saturday, March 19, 2016

DATA SIMPLIFICATION: Persistent Data


This is the last of my blogs related to topics selected from Data Simplification: Taming Information With Open Source Tools (released March, 2016). I hope that as you page back through my posts on Data Simplification topics, appearing throughout this month's blog, you'll find that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.

A file that big?
It might be very useful.
But now it is gone.

-Haiku by David J. Liszewski

Your scripts create data objects, and the data objects hold data. Sometimes, these data objects are transient, existing only during a block or subroutine. At other times, the data objects produced by scripts represent prodigious amounts of data, resulting from complex and time-consuming calculations. What happens to these data structures when the script finishes executing? Ordinarily, when a script stops, all the data produced by the script simply vanishes.

Persistence is the ability of data to outlive the program that produced it. The methods by which we create persistent data are sometimes referred to as marshalling or serializing. Some of the language specific methods are called by such colorful names as data dumping, pickling, freezing/thawing, and storable/retrieve.

Data persistence can be ranked by level of sophistication. At the bottom is the exportation of data to a simple flat-file, wherein records are each one line in length, and each line of the record consists of a record key, followed by a list of record attributes. The simple spreadsheet stores data as tab delimited or comma separated line records. Flat-files can contain a limitless number of line records, but spreadsheets are limited by the number of records they can import and manage. Scripts can be written that parse through flat-files line by line (i.e., record by record), selecting data as they go. Software programs that write data to flat-files achieve a crude but serviceable type of data persistence.

A middle-level technique for creating persistent data is the venerable database. If nothing else, databases are made to create, store, and retrieve data records. Scripts that have access to a database can achieve persistence by creating database records that accommodate data objects. When the script ends, the database persists, and the data objects can be fetched and reconstructed for use in future scripts.

Perhaps the highest level of data persistence is achieved when complex data objects are saved in toto. Flat-files and databases may not be suited to storing complex data objects, holding encapsulated data values. Most languages provide built-in methods for storing complex objects, and a number of languages designed to describe complex forms of data have been developed. Data description languages, such as YAML (Yet Another Markup Language) and JSON (JavaScript Object Notation) can be adopted by any programming language.

Data persistence is essential to data simplification. Without data persistence, all data created by scripts is volatile, obliging data scientists to waste time recreating data that has ceased to exist. Essential tasks such as script debugging and data verification become impossible. It is worthwhile reviewing some of the techniques for data persistence that are readily accessible to Perl, Python and Ruby programmers.

Perl will dump any data structure into a persistent, external file, for later use. Here, the Perl script, data_dump.pl, creates a complex associative array, "%hash", which nests within itself a string, an integer, an array, and another associative array. This complex data structure is dumped into a persistent structure (i.e., an external file named dump_struct).
#!/usr/local/bin/perl
use Data::Dump qw(dump);
%hash = (
    number => 42,
    string => 'This is a string',
    array  => [ 1 .. 10 ],
    hash   => { apple => 'red', banana => 'yellow'},);
open(OUT, ">dump_struct");
print OUT dump \%hash;
exit;
The Perl script, data_slurp.pl picks up the external file, "dump_struct", created by the data_dump.pl script, and loads it into a variable.
#!/usr/local/bin/perl
use Data::Dump qw(dump);
open(IN, "dump_struct");
undef($/);
$data = eval ;
close $in;
dump $data;
exit;
Here is the output of the data_slurp.pl script, in which the contents in the variable "$data" are dumped onto the output screen:
c:\ftp>data_slurp.pl
{
  array  => [1 .. 10],
  hash   => { apple => "red", banana => "yellow" },
  number => 42,
  string => "This is a string",
}
Python pickles its data. Here, the Python script, pickle_up.py, pickles a string variable
#!/usr/bin/python
import pickle
pumpkin_color = "orange"
pickle.dump( pumpkin_color, open( "save.p", "wb" ) )
exit
The Python script, pickle_down.py, loads the pickle file, "save.p" and prints it to the screen.
#!/usr/bin/python
import pickle
pumpkin_color = pickle.load( open( "save.p", "rb" ) )
print(pumpkin_color)
exit
The output of the pickle_down.py script is shown here:
c:\ftp\py>pickle_down.py
orange
Where Python pickles, Ruby marshalls. In Ruby, whole objects, with their encapsulated data, are marshalled into an external file and demarshalled at will. Here is a short Ruby script, object_marshal.rb, that creates a new class, "Shoestring", a new class object, "loafer", and marshalls the new object into a persistent file, "output_file.per".
#!/usr/bin/ruby

class Shoestring < String   
  def initialize 
    @object_uuid = (`c\:\\cygwin64\\bin\\uuidgen.exe`).chomp
  end
  def object_uuid
    print @object_uuid
  end
end

loafer = Shoestring.new
output = File.open("output_file.per", "wb")
output.write(Marshal::dump(loafer))
exit
The script produces no output other than the binary file, "output_file.per". Notice that when we created the object, loafer, we included a method that encapsulates within the object a full uuid identifier, courtesy of cygwin's bundled utility, "uuidgen.exe".

We can demarshal the persistent "output_file.per" file, using the ruby script, object_demarshal.rb:
#!/usr/bin/ruby

class Shoestring < String   
  def initialize 
    @object_uuid = `c\:\\cygwin64\\bin\\uuidgen.exe`.chomp
  end
  def object_uuid
    print @object_uuid
  end
end

array = []
$/="\n\n"
out = File.open("output_file.per", "rb").each do 
  |object|
  array << Marshal::load(object)
  array.each do
    |object|
    puts object.object_uuid
    puts object.class
    puts object.class.superclass
  end
end
exit
The Ruby script, object_demarshal.rb, pulls the data object from the persistent file, "output_file.per" and directs Ruby to list the uuid for the object, the class of the object, and the superclass of the object.
c:\ftp>object_demarshal.rb
c2ace515-534f-411c-9d7c-5aef60f8c72a
Shoestring
String
Perl, Python and Ruby all have access to external database modules that can build database objects that exist as external files that persist after the script has executed. These database objects can be called from any script, with the contained data accessed quickly, with a simple command syntax (1).

Here is a Perl script, lucy.pl, that creates an associative array and ties it to a external database file, using the SDBM_file (Simple Database Management File) module.
#!/usr/local/bin/perl
use Fcntl;
use SDBM_File;
tie %lucy_hash, "SDBM_File", 'lucy', O_RDWR|O_CREAT|O_EXCL, 0644;
$lucy_hash{"Fred Mertz"} = "Neighbor";
$lucy_hash{"Ethel Mertz"} = "Neighbor";
$lucy_hash{"Lucy Ricardo"} = "Star";
$lucy_hash{"Ricky Ricardo"} = "Band leader";
untie %lucy_hash;
exit;
The lucy.pl script produces a persistent, external file, from which any Perl script can access the associative array created in the prior script. If we look in the directory from which the lucy.pl script was launched, we will find two new SDBM (Simple DataBase Manager) files, lucy.dir and lucy.pag. These are the persistent files that will substitute for the %lucy_hash associative array when invoked within other Perl scripts.

Here is a short Perl script, lucy_untie.pl, that extracts the persistent %lucy_hash associative array from the SDBM file in which it is stored:
#!/usr/local/bin/perl
use Fcntl;
use SDBM_File;
tie %lucy_hash, "SDBM_File", 'lucy', O_RDWR, 0644;
while(($key, $value) = each (%lucy_hash))
  {
  print "$key => $value\n";
  }
untie %mesh_hash;
exit;
Here is the output of the lucy_untie.pl script:
c:\ftp>lucy_untie.pl
Fred Mertz => Neighbor
Ethel Mertz => Neighbor
Lucy Ricardo => Star
Ricky Ricardo => Band leader
Here is the Python script, lucy.py, that creates a tiny external database. [jb meta.txt]
#!/usr/local/bin/python
import dumbdbm
lucy_hash = dumbdbm.open('lucy', 'c')
lucy_hash["Fred Mertz"] = "Neighbor"
lucy_hash["Ethel Mertz"] = "Neighbor"
lucy_hash["Lucy Ricardo"] = "Star"
lucy_hash["Ricky Ricardo"] = "Band leader"
lucy_hash.close()
exit
Here is the Python script, lucy_untie.py, that reads all of the key,value pairs held in the persistent database created for the lucy_hash dictionary object.
#!/usr/local/bin/python
import dumbdbm
lucy_hash = dumbdbm.open('lucy')
for character in lucy_hash.keys():
  print character, lucy_hash[character]
lucy_hash.close()
exit
Here is the output produced by the Python script, lucy_untie.py script.
c:\ftp>lucy_untie.py
Fred Mertz Neighbor
Ethel Mertz Neighbor
Lucy Ricardo Star
Ricky Ricardo Band leader
Ruby can also hold data in a persistent database, using the gdbm module. If you do not have the gdbm (GNU database manager) module installed in your Ruby distribution, you can install it as a Ruby GEM, using the following command line, from the system prompt:
c:\>gem install gdbm
The Ruby script, lucy.rb, creates an external database file, lucy.db:
#!/usr/local/bin/ruby
require 'gdbm'
lucy_hash = GDBM.new("lucy.db")
lucy_hash["Fred Mertz"] = "Neighbor"
lucy_hash["Ethel Mertz"] = "Neighbor"
lucy_hash["Lucy Ricardo"] = "Star"
lucy_hash["Ricky Ricardo"] = "Band leader"
lucy_hash.close
exit
The Ruby script, ruby_untie.db, reads the associate array stored as the persistent database, lucy.db:
#!/usr/local/bin/ruby
require 'gdbm'
gdbm = GDBM.new("lucy.db")
gdbm.each_pair do |name, role|
  print "#{name}: #{role}\n"
end
gdbm.close
exit
The output from the lucy_untie.rb script is:
c:\ftp>lucy_untie.rb
Ethel Mertz: Neighbor
Lucy Ricardo: Star
Ricky Ricardo: Band leader
Fred Mertz: Neighbor
Persistence is a simple and fundamental process ensuring that data created in your scripts can be recalled by yourself or by others who need to verify your results. Regardless of the programming language you use, or the data structures you prefer, you will need to familiarize with at least one data persistence technique.


- Jules Berman (copyrighted material)

key words: computer science, data science, data analysis, data simplification, simplifying data, persistence, databases, jules j berman

References:

[1] Berman JJ. Methods in Medical Informatics: Fundamentals of Healthcare Programming in Perl, Python, and Ruby. Chapman and Hall, Boca Raton 2010.

Tuesday, March 15, 2016

DATA SIMPLIFICATION: The Many Uses of Random Number Generators


Over the next few weeks, I will be writing on topics related to my latest book, Data Simplification: Taming Information With Open Source Tools (release date March 17, 2016). I hope I can convince you that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.

If you are among the many students and professionals who are intimidated by statistics, then fear no more! With a little imagination, random number generators (to be accurate, pseudorandom number generators) can substitute for a wide range of statistical methods.

As it happens, modern computers can perform two simple processes, easily and very quickly. These two processes are: 1) generating random numbers, and 2) repeating sets of instructions thousands or millions of times. Using these two computational steps, we can accurately predict outcomes that would be intractable to any direct mathematical analysis. You are about to be rewarded with simple methods whereby every statistical test can be replicated and every probabilistic dilemma can be resolved; usually with a few lines of code (1-5).

To begin, let's perform a few very simple simulations that confirm what we already know, intuitively. Imagine that you have a pair of dice, and you would like to know how often you might expect each of the numbers (from one to six) to appear after you've thrown one die (5).

Let's simulate 600,000 throws of a die, using the Perl script, randtest.pl:
     #!/usr/bin/perl
     $count = 0;
     while ($count < 600000)
        {
        $count++;
        $one_of_six = (int(rand(6))+1);
        $hash{$one_of_six}++;
        }
      while(($key, $value) = each (%hash))
        {
      print "$key => $value\n";
        }
     exit;
The script, randtest.pl, begins by setting a loop that repeats 600,000 times, each repeat simulating the cast of a die. With each cast of the die, Perl generates a random integer, 1 through 6, simulating the outcome of a throw. The most important line of code is:
$one_of_six = (int(rand(6))+1);
The rand(6) command yields a pseudorandom number of value less than 6. We integerize the result using Perl's int() function, which truncates anything past the decimal point. This produces integer values of 0,1,2,3,4, or 5. We increment each value to produce 1,2,3,4,5 or 6. The script yields the total number die casts that would be expected for each of the possible outcomes.

Here is the output of randtest.pl.
C:\ftp>perl randtest.pl
1 => 100002
2 => 99902
3 => 99997
4 => 100103
5 => 99926
6 => 100070
As one might expect, each of the six equally likely outcomes of a thrown die occurred about 100,000 times, in our simulation.

Repeating the randtest.pl script produces a different set of outcome numbers, but the general result is the same. Each die outcome had about the same number of occurrences.
C:\ftp>perl randtest.pl
1 => 100766
2 => 99515
3 => 100157
4 => 99570
5 => 100092
6 => 99900
Let's get a little more practice with random number generators, before moving onto more challenging simulations. Occasionally in scripts, we need to create a new file, automatically, during the script's run time, and we want to be fairly certain that the file we create will not have the same filename as an existing file. An easy way of choosing a filename is to grab, at random, printable characters, concatenating them into an 11 character string suitable as a filename. The chance that you'll encounter two files with the same randomly chosen filename is very remote. In fact, the likelihood that any two selected filenames are identical exceeds to 2 to the 44th power.

Here is a Perl script, random_filenames.pl, that assigns a sequence of 11 randomly chosen uppercase alphabetic characters to a file name:
#!/usr/bin/perl
for ($count = 1; $count <= 12; $count++)
  {
  push(@listchar, chr(int(rand(26))+65));
  }
$listchar[8]= ".";
$randomfilename = join("",@listchar);
print "Your filename is $randomfilename\n";
exit;
Here is the output of the ranfile.pl script:
c:\ftp>random_filenames.pl
Your filename is OAOKSXAH.SIT
The key line of code in random_filenames.pl is:
push(@listchar, chr(int(rand(26))+65));
The rand(26) command yields a random value less than 26. The int() command converts the value to an integer. The number 65 is added to the value to produce a value ranging from 65 to 90, and the chr() command converts the numbers 65 through 90 to their ASCII equivalent; which just happen to be the uppercase alphabet from A to Z. The randomly chosen letter is pushed onto an array, and the process is repeated until a 12 character filename is generated.

Here is the equivalent Python script, random_filenames.py, that produces one random filename
#!/usr/bin/python
import random
filename = [0]*12
filename = map(lambda x: x is "" or chr(int(random.uniform(0,25) + 65)), filenam
e)
print ''.join(filename[0:8]) + "." + ''.join(filename[9:12])
exit
Here is the outcome of the random_filenames.py script:
c:\ftp>random_filenames.py
KYSDWKLF.RBA
In both these scripts, as in all of the scripts in this section, many outcomes may result from a small set of initial conditions. It's much easier to write these programs and observe their outcomes than to directly calculate all the possible outcomes of a set of governing equations.[jb outline.txt]

Let's use a random number generator to calculate the value of pi, without measuring anything, and without resorting to summing an infinite series of numbers. Here is a simple python script, pi.py, that does the job.
#!/usr/bin/python
import random
from math import sqrt
totr = 0
totsq = 0
for iterations in range(10000000):
  x= random.uniform(0,1)
  y= random.uniform(0,1)
  r= sqrt((x*x) + (y*y))
  if r < 1:
    totr = totr + 1
  totsq = totsq + 1
print float(totr)*4.0/float(totsq)
exit
The outcome of the pi.py script is:
output:
c:\ftp\py>pi.py
3.1414256
Here is an equivalent Ruby script. [jb]
#!/usr/local/bin/ruby
x = y = totr = totsq = 0.0
(1..100000).each do
  x = rand()
  y = rand()
  r = Math.sqrt((x*x) + (y*y))
  totr = totr + 1 if r < 1
  totsq = totsq + 1
end
puts (totr *4.0 / totsq)
exit
Here is an equivalent Perl script. [jb]
#!/usr/local/bin/perl
open(DATA,">pi.dat");
for (1..10000)
  {
  $x = rand();
  $y = rand();
  $r = sqrt(($x*$x) + ($y*$y));
  if ($r < 1)
    {
    $totr = $totr + 1;
    print DATA "$x\ $y\n";
    }
  $totsq = $totsq + 1;
  }
print eval(4 * ($totr / $totsq)); 
exit;
As one would hope, all three scripts produce approximately the same value of pi. The Perl script contains a few extra lines that produces an output file, named pi.dat, that will helps us visualize how these scripts work. The pi.dat script contains the x,y data points, generated by the random number generator, meeting the "if" statement's condition that the hypotenuse of the x,y coordinates must be less than one (i.e., less than a circle of radius 1). [jb]

We can plot the output of the script with a few lines of Gnuplot code:
gnuplot> set size square
gnuplot> unset key
gnuplot> plot 'c:\ftp\pi.dat'
The resulting graph is a quarter-circle within a square.


The data points produced by 10,000 random assignments
of x and y coordinates in a range of 0 to 1.
Randomly assigned data points whose hypotenuse
exceeds "1" are excluded.


The graph shows us, at a glance, how the ratio of the number of points in the quarter-circle, as a fraction of the total number of simulations, is related to the value of pi.


- Jules Berman (copyrighted material)

key words: computer science, data analysis, data repurposing, data simplification, simplifying data, random, pseudorandom, resampling, probability, simulations, Monte Carlo jules j berman

References:

[1] Simon JL. Resampling: The New Statistics. Second Edition, 1997. Available online at: http://www.resample.com/intro-text-online/, viewed on September 21, 2015.

[2] Efron B, Tibshirani RJ. An Introduction to the Bootstrap. CRC Press, Boca Raton, 1998.

[3] Diaconis P, Efron B. Computer-intensive methods in statistics. Scientific American, May, 116-130, 1983. Comment. Oft-cited explanation of resampling statistics, a field largely credited to Bradley Efron. The articles contains examples in Basic, Pascal, and Fortran source code.

[4] Anderson HL. Metropolis, Monte Carlo and the MANIAC. Los Alamos Science 14:96-108, 1986. Available at: http://library.lanl.gov/cgi-bin/getfile?00326886.pdf, viewed September 21, 2015.

[5] Berman JJ. Biomedical Informatics. Jones and Bartlett, Sudbury, MA, 2007.

Monday, March 14, 2016

DATA SIMPLIFICATION: Abbreviations and Acronyms


Over the next few weeks, I will be writing on topics related to my latest book, Data Simplification: Taming Information With Open Source Tools (release date March 17, 2016). I hope I can convince you that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.

"A synonym is a word you use when you can't spell the other one." -Baltasar Gracian

People confuse shortening with simplifying; a terrible mistake. In point of fact, next to reifying pronouns, abbreviations are the most vexing cause of complex and meaningless language. Before we tackle the complexities of abbreviations, let's define our terms. An abbreviation is a shortened form of a word or term. An acronym is a an abbreviation composed of letters extracted from the words composing a multi-word term. There are two major types of abbreviations: universal/permanent and local/ephemeral. The universal/permanent abbreviations are recognized everywhere and have been used for decades (e.g., USA, DNA, UK). Some of the universal/permanent abbreviations, ascend to the status of words whose long-forms have been abandoned. For example, we use laser as a word. Few who use the term know that "laser" is an acronym for "light amplification by stimulated emission of radiation". Local/ephemeral abbreviations are created for terms that are repeated within a particular document or a particular class of documents. Synonyms and plesionyms (i.e., near-synonyms) allow authors to represent a single concept using alternate terms (1).

Abbreviations make textual data complex, for three three principle reasons:

1. No rules exist with which abbreviations can be logically expanded to their full-length form.

2. A single abbreviation may mean different things to different individuals, or to the same individual at different times.

3. A single term may have multiple different abbreviations. (In medicine, Angioimmunoblastic lymphadenopathy can be abbreviated as ABL, AIL, or AIML.) These are the so-called polysemous abbreviations (See Glossary item, Polysemy). In the medical literature, a single abbreviations may have dozens of different expansions (1).

Some of the worst abbreviations fall into one of the following catagories:

Abbreviations that are neither acronyms nor shortened forms of expansions. For example, the short form of "diagnosis" is "dx", although no "x" is contained therein. The same applies to the "x" in "tx", the abbreviation for "therapy", but not the "X" in "TX" that stands for Texas. For that matter, the short form of "times" is an "x", relating to the notation for the multiplication operator. Roman numerals I, V, X, L and M are abbreviations for words assigned to numbers, but they are not characters included in the expanded words (e.g., there is no "I" in "one"). EKG is the abbreviation for electrocardiogram, a word totally bereft of any "K". The "K" comes from the German orthography. There is no letter "q" in subcutaneous, but the abbreviation for the word is sometimes "subq"; never "subc". What form of alchemy converts ethanol to its common abbreviation, "EtOH"?

Mixed-form abbreviations. In medical lingo "DSV" represents the Dermatome of the fifth (V) Sacral nerve. Here a preposition, an article, and a noun (of, the, nerve) have all been unceremoniously excluded from the abbreviation; the order or the acronym components have been transposed (dermatome sacral fifth); an ordinal has been changed to a cardinal (fifth changed to five), and the cardinal has been shortened to its roman numeral equivalent (V).

Prepositions and articles arbitrarily retained in an acronym. When creating an abbreviation, should we retain or abandon prepositions? Many acronyms exclude prepositions and articles. USA is the acronym for United States of America; the "of" is ignored. DOB (Date Of Birth) remembers the "of".

Single expansions with multiple abbreviations. Just as abbreviations can map to many different expansions, the reverse can occur. For instance, high-grade squamous intraepithelial lesion can be abbreviated as HGSIL or HSIL. Xanthogranulomatous pyelonephritis can be abbreviated as xgp or xgpn.

Recursive abbreviations. The following example exemplifies the horror of recursive abbreviations. The term SMETE is the abbreviation for the phrase "science, math, engineering, and technology education". NSDL is a real-life abbreviation, for "National SMETE digital Library community". To fully expand the term (i.e., to provide meaning to the abbreviation), you must recursively expand the embedded abbreviation, to produce "National science, math, engineering, and technology education digital Library community."

Stupid or purposefully unhelpful abbreviations. The term GNU (Gnu is not UNIX) is a recursive acronym. Fully expanded, this acronym is of infinite length. Although the N and the U expand to words ("Not Unix"), the letter G is simply inscrutable. Another example of an inexplicable abbreviation is PT-LPD (post-transplantation lymphoproliferative disorders). The only logical location for a hyphen would be smack between the letters p and t. Is the hyphen situated between the T and the L for the sole purpose of irritating us?

Abbreviations that change from place to place. Americans sometimes forget that most English-speaking countries use British English. For example an esophagus in New York is an oesophagus in London. Hence TOF makes no sense as an abbreviation of tracheo-esophageal fistula here in the U.S. but this abbreviation makes perfect sense to physicians in England, where a patients may have a Trancheo-Oesophageal Fistula. The term GERD (representing the phrase gastroesophageal reflux disease) makes perfect sense to Americans, but it must be confusing in Britain, where the esophagus is not an organ.

Abbreviations masquerading as words. Our greatest vitriol is reserved for abbreviations that look just like common words. Some of the worst offenders come from the medical lexicon: axillary node dissection (AND), acute lymphocytic leukemia (ALL), Bornholm Eye Disease (BED), and Expired Air Resuscitation (EAR). Such acronyms aggravate the computational task confidently translating common words. Acronyms commonly appear as uppercase strings, but a review of a text corpus of medical notes has shown that words could not be consistently distinguished from homonymous word-acronyms (2).

Fatal abbreviations. Fatal abbreviations are those which can kill individuals if they are interpreted incorrectly. They all seem to originate in the world of medicine:

MVR, which can be expanded to any of: mitral vale regurgitation, mitral valve repair, or mitral valve replacement;

LLL, which can be expanded to any of: left lower lid, left lower lip, or left lower lung;

DOA, dead on arrival, date of arrival, date of admission, drug of abuse.

Is a fear of abbreviations rational, or does this fear emanate from an overactive imagination? In 2004, the Joint Commission on Accreditation of Healthcare Organizations, a stalwart institution not known to be squeamish, issued an announced that, henceforth, a list of specified abbreviations should be excluded from medical records Rboodr.

Examples of Forbidden abbreviations are:

IU (International Unit), mistaken as IV (intravenous) or 10 (ten).

Q.D., Q.O.D. (Latin abbreviation for once daily and every other day), mistaken for each other.

Trailing zero (X.0 mg) or a lack of a leading zero (.X mg), in which cases the decimal point may be missed. Never write a zero by itself after a decimal point (X mg), and always use a zero before a decimal point (0.X mg).

MS, MSO4, MgSO4 all of which can be confused with one another and with morphine sulfate or magnesium sulfate. Write "morphine sulfate" or "magnesium sulfate."

Abbreviations on the hospital watch list were:

mg (for microgram), mistaken fir mg (milligrams), resulting in a 1000-fold dosing overdose.

h.s., which can mean either half-strength or the Latin abbreviation for bedtime or may be mistaken for q.h.s., taken every hour. All can result in a dosing error.

T.I.W. (for three times a week), mistaken for three times a day or twice weekly, resulting in an overdose.

The list of abbreviations that can kill, in the medical setting, is quite lengthy. Fatal abbreviations probably devolved through imprecise, inconsistent, or idiosyncratic uses of an abbreviation, by the busy hospital staff who enter notes and orders into patient charts. For any knowledge domain, the potentially fatal abbreviations is the most important to catch.

Nobody has ever found an accurate way of disambiguating and translating abbreviations (1). There are, however a few simple suggestions, based on years of exasperating experience, that might save you time and energy.

1. Disallow the use of abbreviations, whenever possible. Abbreviations never enhance the value of information. The time saved by using an abbreviation is far exceeded by the time spent attempting to deduce its correct meaning.

2. When writing software applications that find and expand abbreviations, the output should list every known expansion of the abbreviation. For example, the abbreviation, "ppp" appearing in a medical report, should have all these expansions inserted into the text, as annotations: pancreatic polypeptide, palatopharyngoplasty, palmoplantar pustulosis, pancreatic polypeptide, pentose phosphate pathway, platelet poor plasma, primary proliferative polycythaemia, primary proliferative polycythemia. Leave it up to the knowledge domain experts to disambiguate the results.


- Jules Berman (copyrighted material)

key words: computer science, data analysis, data repurposing, data simplification, simplifying data, abbreviations, acronyms, complexity jules j berman

References:

[1] Berman JJ. Pathology abbreviated: a long review of short terms. Arch Pathol Lab Med 128:347-352, 2004.

[2] Nadkarni P, Chen R, Brandt C. UMLS concept indexing for production databases. JAMIA 8:80-91, 2001.

Sunday, March 13, 2016

DATA SIMPLIFICATION: Doublet Lists


Over the next few weeks, I will be writing on topics related to my latest book, Data Simplification: Taming Information With Open Source Tools (release date March 17, 2016). I hope I can convince you that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.

Yesterday's blog covered lists of single words. Today we'll do doublets.

Doublet lists (lists of two-word terms that occur in common usage or in a body of text) are a highly underutilized resource. The special value of doublets is that single word terms tend to have multiple meanings, while doublets tend to have specific meaning.

Here are a few examples:

The word "rose" can mean the past tense of rise, or the flower. The doublet "rose garden" refers specifically to a place where the rose flower grows.

The word "lead" can mean a verb form of the infinitive, "to lead", or it can refer to the metal. The term "lead paint" has a different meaning than "lead violinist". Furthermore, every multiword term of length greater than two can be constructed with overlapping doublets, with each doublet having a specific meaning.

For example, "Lincoln Continental convertible" = "Lincoln Continental" + "Continental convertible". The three words, "Lincoln", "Continental", and "convertible" all have different meanings, under different circumstances. But the two doublets, "Lincoln Continental" and "Continental Convertible" would be unusual to encounter on their own, and produce a unique meaning, when combined.

Perusal of any nomenclature will reveal that most of the terms included in nomenclatures consist of two or more words. This is because single word terms often lack specificity. For example, in a nomenclature of recipes, you might expect to find, "Eggplant Parmesan" but you may be disappointed if you look for "Eggplant" or "Parmesan". In a taxonomy of neoplasms, available at: http://www.julesberman.info/figs/neocl_f.htm, containing over 120,000 terms, only a few hundred of those terms are single word terms (1).

Lists of doublets, collected from a corpus of text, or from a nomenclature, have a variety of uses in data simplification projects (1-3). We will show examples in Section 5.4, and in "On-the-fly indexing scripts" later in this chapter.

For now, you should know that compiling doublet lists, from any corpus of text, is extremely easy.

Here is a perl script, doublet_maker.pl, that creates a list of alphabetized doublets occurring in any text file of your choice (filename.txt in this example):
#!/usr/local/bin/perl
open(TEXT,"filename.txt")||die"cannot";
open(OUT,">doublets.txt")||die"cannot";
undef($/);
$var = ;
$var =~ s/\n/ /g;
$var =~ s/\'s//g;
$var =~ tr/a-zA-Z\'\- //cd;
@words = split(/ +/, $var);
foreach $thing (@words)
  {
  $doublet = "$oldthing $thing";
  if ($doublet =~ /^[a-z]+ [a-z]+$/)
    {
    $doublethash{$doublet}="";
    }
  $oldthing = $thing;
  }
close TEXT;
@wordarray = sort(keys(%doublethash));
print OUT join("\n",@wordarray);
close OUT;
exit;
Here is an equivalent Python script, doublet_maker.py:
#!/usr/local/bin/python
import anydbm, string, re
in_file = open('filename.txt', "r")
out_file = open('doubs.txt',"w")
doubhash = {}
for line in in_file:
  line = line.lower()
  line = re.sub('[.,<>?/;:"[]\{}|=+-_ ()*&^%$#@!`~1234567890]', ' ', line)
  hoparray = line.split()
  hoparray.append(" ")
  for i in range(len(hoparray)-1):
     doublet = hoparray[i] + " " + hoparray[i + 1]
     if doubhash.has_key(doublet):
          continue
     doubhash_match = re.search(r'[a-z]+ [a-z]+',  doublet)
     if doubhash_match:
         doubhash[doublet] = ""
for keys,values in sorted(doubhash.items()):
    out_file.write(keys + '\n')
exit
Here is an equivalent Ruby script, doublet_maker.rb that creates a doublet list from file filename.txt:
#!/usr/local/bin/ruby
intext = File.open("filename.txt", "r")
outtext = File.open("doubs.txt", "w")
doubhash = Hash.new(0)
line_array = Array.new(0)
while record = intext.gets
  oldword = ""
  line_array = record.chomp.strip.split(/\s+/)
  line_array.each do
    |word|
    doublet = [oldword, word].join(" ")
    oldword = word
    next unless (doublet =~ /^[a-z]+\s[a-z]+$/)
    doubhash[doublet] = ""
    end
end
doubhash.each {|k,v| outtext.puts k }
exit
I have deposited a public domain doublet list, available for download at:

http://www.julesberman.info/doublets.htm

The first few lines of the list are shown:
a bachelor
a background
a bacteremia
a bacteria
a bacterial
a bacterium
a bad
a balance
a balanced
a banana

- Jules Berman

key words: computer science, data analysis, data repurposing, data simplification, word lists, doublet lists, n-grams, complexity, open source tools, jules j berman

References:

[1] Berman JJ. Automatic extraction of candidate nomenclature terms using the doublet method. BMC Medical Informatics and Decision Making 5:35, 2005.

[2] Berman JJ. Doublet method for very fast autocoding. BMC Med Inform Decis Mak, 4:16, 2004.

[3] Berman JJ. Nomenclature-based data retrieval without prior annotation: facilitating biomedical data integration with fast doublet matching. In Silico Biol, 5:0029, 2005. Available at: http://www.bioinfo.de/isb/2005/05/0029/, viewed on September 6, 2015.

Saturday, March 12, 2016

DATA SIMPLIFICATION: Building Word Lists


Over the next few weeks, I will be writing on topics related to my latest book, Data Simplification: Taming Information With Open Source Tools (release date March 23, 2016). I hope I can convince you that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.


Word lists, for just about any written language for which there is an electronic literature, are easy to create. Here is a short Python script, words.py, that prompts the user to enter a line of text. The script drops the line to lowercase, removes the carriage return at the end of the line, parses the result into an alphabetized list, removes duplicate terms from the list, and prints out the list, with one term assigned to each line of output. This words.py script can be easily modified to create word lists from plain-text files (See Glossary item, Metasyntactic variable).
#!/usr/local/bin/python
import sys, re, string
print "Enter a line of text to be parsed into a word list"
line = sys.stdin.readline()
line = string.lower(line)
line = string.rstrip(line)
linearray = sorted(set(re.split(r' +', line)))
for i in range(0, len(linearray)):
   print(linearray[i])
exit
Here is some a sample of output, when the input is the first line of Joyce's Finegans Wake:
c:\ftp>words.py
Enter a line of text to be parsed into a word list

a way a lone a last a loved a long the riverrun, past Eve and Adam's, from 
swerve of shore to bend of bay, brings us by a commodius vicus

a
adam's,
and
bay,
bend
brings
by
commodius
eve
from
last
lone
long
loved
of
past
riverrun,
shore
swerve
the
to
us
vicus
way
Here is a nearly equivalent Perl script, words.pl, that creates a wordlist from a file. In this case, the chosen file happens to be "gettbysu.txt", containing the full-text of the Gettysburg address. We could have included the name of any plain-text file.
#!/usr/local/bin/perl
open(TEXT, "gettysbu.txt");
undef($/); 
$var = lc();
$var =~ s/\n/ /g;
$var =~ s/\'s//g;
$var =~ tr/a-zA-Z\'\- //cd;
@words = sort(split(/ +/, $var));
@words = grep($_ ne $prev && (($prev) = $_), @words);
print (join("\n",@words));
exit;
The words.pl script was designed for speed. You'll notice that it slurps the entire contents of a file into a string variable. If we were dealing with a very large file, that exceeded the functional RAM memory limits of our computer, we would need to modify the script to parse through the file line-by-line.

Aside from word lists you create for yourself, there are a wide variety of specialized knowledge domain nomenclatures that are available to the public (1), (2), (3), (4), (5), (6). Linux distributions often bundle a wordlist, under filename "words", that is useful for parsing and natural language processing applications. A copy of the linux wordlist is available at:

http://www.cs.duke.edu/~ola/ap/linuxwords

Curated lists of terms, either generalized, or restricted to a specific knowledge domain, are indispensable for a variety of applications (e.g., spell-checkers, natural language processors, machine translation, coding by term, indexing. Personally, I have spent an inexcusable amount of time creating my own lists, when no equivalent public domain resource was available.

- Jules Berman

key words: computer science, data analysis, data repurposing, data simplification, data wrangling, information science, simplifying data, taming data, complexity, system calls, Perl, Python, open source tools, utility, word lists, jules j berman

References:

[1] Medical Subject Headings. U.S. National Library of Medicine. Available at: https://www.nlm.nih.gov/mesh/filelist.html, viewed on July 29, 2015.

[2] Berman JJ. A Tool for Sharing Annotated Research Data: the "Category 0" UMLS (Unified Medical Language System) Vocabularies. BMC Med Inform Decis Mak, 3:6, 2003.

[3] Berman JJ Tumor taxonomy for the developmental lineage classification of neoplasms. BMC Cancer 4:88, 2004. http://www.biomedcentral.com/1471-2407/4/88, viewed Jan. 1, 2015.

[4] Hayes CF, O'Connor JC. English-Esperanto Dictionary. Review of Reviews Office, London, 1906. Availalable at: http://www.gutenberg.org/ebooks/16967 viewed on July 29, 2105.

[5] Sioutos N, de Coronado S, Haber MW, Hartel FW, Shaiu WL, Wright LW. NCI Thesaurus: a semantic model integrating cancer-related clinical and molecular information. J Biomed Inform 40:30-43, 2007.

[6] NCI Thesaurus. National Cancer Institute, U.S. National Institutes of Health, Bethesda, MD. Available at: ftp://ftp1.nci.nih.gov/pub/cacore/EVS/NCI_Thesaurus/ viewed on July 29, 2015.

Thursday, March 10, 2016

DATA SIMPLIFICATION: System Calls


Over the next few weeks, I will be writing on topics related to my latest book, Data Simplification: Taming Information With Open Source Tools (release date March 23, 2016). I hope I can convince you that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.


A system call is a command line, inserted into a software program, that interrupts the script while the operating system executes the command line. Immediately afterwords, the script resumes, at the next line. Any utility that runs from the command line can be embedded in any scripting language that supports system calls, and this includes all of the languages discussed in this book.

Here are the properties of system calls that make them useful to programmers:

1. System calls can be inserted into iterative loops (e.g., while loops, for loops), so that they can be repeated any number of times, on collections of files, or data elements.

2. Variables that are generated at run-time (i.e.,during the execution of the script) can be included as arguments added to the system call.

3. The results of the system call can be returned to the script, and used as variables.

4. System calls can utilize any operating system command and any program that would normally be invoked through a command line, including external scripts written in other programming languages. Hence, a system call can initiate an external script written in an alternate programming language, composed at run-time within the original script, using variables generated in the original script, and capturing the output from the external script for use in the original script!

System calls enhance the power of any programming language by providing access to a countless number of external methods and by participating in iterated actions using variables created at run-time.

How does the system call help with the task of data simplification? Data simplification is very often focused on uniformity and reproducibility. If you have 100,000 images, data simplification might involve calling ImageMagick to resize every image to the same height and width. If you need to convert spreadsheet data to a set of triples, than you might need to provide a UUID string (see prior blog) to every triple in the database, all at once. If you are working on a Ruby project, and you need to assert one of Python's numpy methods, on every data file in a large collection of data files, then you might want to create a short Python file that you can be accessed, via a system call, from your Ruby script.

Once you have gotten the hang of including system calls in your scripts, you will probably use them in most of your your data simplification tasks. It's important to know how system calls can be used to great advantage, in Perl, Python, and Ruby. A few examples follow.

The following short Perl script makes a system call, consisting of the DOS "dir" command:
#!/usr/bin/perl
system("dir");
exit; 
The "dir" command, launched as a system call, displays the files in the current directory. Here is the equivalent script, in Python:
#!/usr/local/bin/python
import os
os.system("dir")
exit
Notice that system calls in Python require the importation of the os (operating system) module into the script.

Here is an example of a Ruby system call, to ImageMagick's "Identify" utility [note: this only works if you have pre-installed ImageMagick]. The system call instructs the "Identify" utility to provide a verbose description of the image file3320_out.jpg, and to pipe the output into the text file, myimage.txt.
#!/usr/bin/ruby
system("Identify -verbose c:/ftp/3320_out.jpg >myimage.txt")
exit
Here is an example of a Perl system call, to ImageMagick's "convert" utility, that incorporates a Perl variable ($file, in this case) that is passed to the system call [note: this only works if you have pre-installed ImageMagick].
#!/usr/local/bin/perl
$file = "try2.gif";
system("convert -size 350x40 xc:lightgray -font Arial -pointsize 32 -fill black
-gravity north -annotate +0+0 \"Hello, World\" $file");
exit;
The following Python script opens the current directory and parses through every filename, looking for jpeg image files. When a jpeg file is encountered, the script makes a system call to imagemagick, instructing imagemagick's "convert" utility to copy the jpeg file to the thumb drive (designated as the f: drive), in the form of a grayscale image. If you try this script at home, be advised that it requires a mounted thumb drive, in the "f:" drive [note: this only works if you have pre-installed ImageMagick].
#!/usr/local/bin/python
import os, re, string
filelist = os.listdir(".")
for file in filelist:
  if ".jpg" in file:  
    img_in = file
    img_out = "f:/" + file 
    command = "convert " + img_in + " -set colorspace Gray -separate -average " + img_out
    os.system(command)
exit
Let's look at a Ruby script that calls a Perl script, a Python script, and another Ruby script, from within one Ruby script.

Here are the Perl, Python and Ruby scripts that will be called from within a Ruby script:
hi.py
#!/usr/local/bin/python
print("Hi, I'm a Python script")
exit

hi.pl
#!/usr/local/bin/perl
print "Hi, I'm a Perl script\n";
exit;

hi.rb
#!/usr/local/bin/ruby
puts "Hi, I'm a Ruby script"
exit
Here is the Ruby script, call_everyone.rb, that calls external scripts, written in Python, Perl and Ruby:
#!/usr/local/bin/ruby
system("python hi.py")
system("perl hi.pl")
system("ruby hi.rb")
exit
Here is the output of the Ruby script, call_everyone.rb:
c:\ftp>call_everyone.rb
Hi, I'm a Python script
Hi, I'm a Perl script
Hi, I'm a Ruby script
If you have some facility with a variety of language-specific methods and utilities, you can deploy them all from within your favorite scripting language.

- Jules Berman

key words: computer science, data analysis, data repurposing, data simplification, data wrangling, information science, simplifying data, taming data, complexity, system calls, Perl, Python, Ruby, jules j berman

Wednesday, March 9, 2016

DATA SIMPLIFICATION: Specifications to the Rescue!


Over the next few weeks, I will be writing on topics related to my latest book, Data Simplification: Taming Information With Open Source Tools (release date March 23, 2016). I hope I can convince you that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.


Today's blog continues yesterday's discussion of Standards and Specifications

Despite the problems inherent in standards, government committees cling to standards as the best way to share data. The perception is that in the absence of standards, the necessary activities of data sharing, data verification, data analysis, and any meaningful validation of the conclusions will be impossible to achieve (1). This long-held perception may not be true. Data standards, intended to simplify our ability to understand and share data, may have increased the complexity of data science. As each new standard is born, our ability to understand our data seems to diminish. Luckily, many of the problems produced by the proliferation of data standards can be avoided by switching to a data annotation technique broadly known as "specification." Although the terms "specification" and "standard" are used interchangeably, by the incognoscenti, the two terms are quite different from one another. A specification is a formal way of describing data. A standard is a set of requirements, created by an standards development organization, that comprise a pre-determined content and format for a set of data.

A specification is an accepted method for describing objects (physical objects such as nuts and bolts; or symbolic objects, such as numbers; or concepts expresed as text). In general, specifications do not require explicit items of information (i.e. they do not impose restrictions on the content that is included in or excluded from documents), and specifications do not impose any order of appearance of the data contained in the document (i.e., you can mix up and rearrange the data records in a specification if you like). Specifications are not generally certified by a standards organization. Examples of specifications are RDF (Resource Description Framework) produced by the W3C (WorldWide Web Consortium), and TCP/IP (Transfer Control Protocol/Internet Protocol), maintained by the Internet Engineering Task Force. The most widely implemented specifications are simple; thus, easily adopted.

Specifications proved a simple and uniform way of representing the information you choose to include in your reports, messages, and files. Some of the most useful and popular specifications are XML, RDF, Notation 3, and Turtle. In general, specifications do not require explicit items of information (i.e. they do not impose restrictions on the content that is included in or excluded from documents), and specifications do not impose any order of appearance of the data contained in the document (i.e., you can mix up and rearrange the data records in a specification if you like). Specifications are not typically certified by a standards organization, but they are developed by special interest groups. Their legitimacy depends on their popularity.

Files that comply with a specification can be parsed and manipulated by generalized software designed to parse the markup language of the specification (e.g., XML, RDF) and to organize the data into data structures defined within the file.

Specifications serve most of the purposes of a standard, plus providing many important functions that standards typically lack (e.g., full data description, data exchange across diverse types of data sets, data merging, and semantic logic). Data specifications spare us most of the heavy baggage that comes with a standard, which includes: limited flexibility to include changing data objects, locked-in data descriptors, licensing and other intellectual property issues, competition among standards that compete within the same data domain, and bureaucratic overhead (2).

Most importantly, specifications make standards fungible. A good specification can be ported into a data standard, and a reasonably good data standard can be ported into a specification. For example, there are dozens of image formats (e.g., jpeg, png, gif, tiff). Although these formats have not gone through a standards development process, they are used by billions of individuals and have achieved the status as de facto standards. For most of us, the selection of any particular image format is inconsequential. Data scientists have access to robust image software that will convert images from one format to another.

A common mistake committed by data scientists is to convert all their data (legacy data and newly acquired data) into a contemporary standard, and then relying on analytic software that is designed to operate exclusively upon the chosen standard. Doing so only serves to perpetuate their frustrations. You can be certain that your data standard and your software application will be unsuitable for the next generation of data scientists. It makes much more sense to port data into a general specification, from which data can be ported to any current or future data standard.

References:

[1] National Committee on Vital and Health Statistics. Report to the Secretary of the U.S. Department of Health and Human Services on Uniform Data Standards for Patient Medical Record Information. July 6, 2000. Available from: http://www.ncvhs.hhs.gov/hipaa000706.pdf

[2] Berman JJ. Repurposing Legacy Data: Innovative Case Studies. Morgan Kaufmann, Waltham, MA, 2015.

- Jules Berman

key words: computer science, data analysis, data repurposing, data simplification, data wrangling, information science, simplifying data, taming data, complexity, standards, specifications, semantic web, jules j berman

Tuesday, March 8, 2016

DATA SIMPLIFICATION: Substandard Standards


Over the next few weeks, I will be writing on topics related to my latest book, Data Simplification: Taming Information With Open Source Tools (release date March 23, 2016). I hope I can convince you that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.


"The nice thing about standards is that you have so many to choose from." -Andrew S. Tanenbaum

Data standards are the false gods of informatics. They promise miracles, but they can't deliver. The biggest drawback of standards is that they change all the time. If you take the time to read some of the computer literature from the 1970s or 1980s, you will come across the names of standards that have long-since fallen into well-deserved obscurity. You may find that the literature from the 1970s is nearly impossible to read with any level of comprehension, due to the large number of now-obsolete standards-related acronyms scattered through every page. Today's eternal standard is tomorrow's indecipherable gibberish (1).

The Open Systems Interconnection (OSI) was an internet protocol created in 1977 with approval from the International Organization for Standardization. It has been supplanted by TCP/IP, the protocol that everyone uses today. A hand full of programming languages have been recognized as standards by the American National Standards Institute. These include Basic, C,, Ada and Mumps. Basic and C are still popular languages. Ada, recommended by the Federal Government, back in 1995, as the recommended language for all high performance software applications, is virtually forgotten (2). Mumps is still in use, particularly in hospital information systems, but it changed its name to M, lost its allure to a new generation of programmers, and now comes in various implementations that may not strictly conform to the original standard.

In many cases, as a standard matures, it becomes hopelessly complex. As the complexity becomes unmanageable, those who profess to use the standard may develop their own idiosyncratic implementations. Organizations that produce standards seldom provide a mechanism to ensure that the standard is implemented correctly. Standards have long been plagued by non-compliance or (more frequently) under-compliance. Over time, so-called standard-compliant systems tend to become incompatible with one another. The net result is that legacy data, purported to conform to a standard format, is no longer understandable.

Regarding versioning, it is a very good rule of thumb that when you encounter a standard whose name includes a version number (e.g. International Classification of Diseases-10, Diagnostic and Statistical Manual of Mental Disorders-5), you can be certain that the standard is unstable, and must be continually revised. Some continuously revised standards cling tenaciously to life, when they really deserve to die. In some cases, a poor standard is kept alive indefinitely by influential leaders in their fields, or by entities who have an economic stake in perpetuating the standard.

Raymond Kammer, then Director of the U.S. National Institute of Standards and Technology, understood the down-side of standards. In a year 2000 government report, he wrote that "the consequences of standards can be negative. For example, companies-and nations-can use standards to disadvantage competitors. Embodied in national regulations, standards can be crafted to impede export access, sometimes necessitating excessive testing and even redesigns of products. A 1999 survey by the National Association of Manufacturers reported that about half of U.S. small manufacturers find international standards or product certification requirements to be barriers to trade. And according to the Transatlantic Business Dialogue, differing requirements add more than 10% to the cost of car design and development." (3)

As it happens, data standards are seldom, if ever, implemented properly. In some cases, the standards are simply too complex to comprehend. Try as they might, every implementation of a complex standard is somewhat idiosyncratic. Consequently, no two implementations of a complex data standard are equivalent to one another. In many cases, corporations and government agencies will purposefully veer from the standard to accommodate some local exigency. In some cases, a corporation may find it prudent to include non-standard embellishments to a standard to create products or functionalities that cannot be easily reproduced by their competitors. In such cases, customers accustomed to a particular manufacturer's rendition of a standard may find it impossible to switch providers).

The process of developing new standards is costly. Interested parties must send representatives to many meetings. In the case of international standards, meetings occur in locations throughout the planet. Someone must pay for the expertise required to develop the standard, improve drafts, and vet the final version. Standards development agencies become involved in the process, and the end-product must be shepherded through one of the agencies that confer final approval. After a standard is approved, it must be accepted by its intended community of users. Educating a community in the use of a standard is another expense. In some cases, an approved standard never gains traction. Because standards cost a great deal of money to develop, it is only natural that corporate sponsors play a major role in the development and deployment of new standards. Software vendors are clever and have learned to benefit from the standards-making process. In some cases, members of a standards committee may knowingly insert a fragment of their own patented property into the standard. After the standard is released and implemented, in many different vendor systems, the patent holder rises to assert the hidden patent. In this case, all those who implemented the standard may find themselves required to pay a royalty for the use of intellectual property sequestered within the standard (4).

Corporations can profit from standards by obtaining patents on the uses of the standard; not on the patent itself. For example, an open standard may have been created that can be obtained at no cost, and that is popular among its intended users, and that contains no hidden intellectual property. An interested corporation or individual may discover a use for the standard that is non-obvious, novel,and useful; these are the three criteria for awarding patents. The corporation or individual can patent the use of the standard, without needing to patent the standard itself. The patent holder will have the legal right to assert the patent over anyone who uses the standard for the purpose claimed by the patent. This patent protection will apply even when the standard is free and open (4).

Despite the problems inherent in standards, government committees cling to standards as the best way to share data. The perception is that in the absence of standards, the necessary activities of data sharing, data verification, data analysis, and any meaningful validation of the conclusions will be impossible to achieve (5). This long-held perception may not be true. Data standards, intended to simplify our ability to understand and share data, may have increased the complexity of data science. As each new standard is born, our ability to understand our data seems to diminish. Luckily, many of the problems produced by the proliferation of data standards can be avoided by switching to a data annotation technique broadly known as "specification." Although the terms "specification" and "standard" are used interchangeably, by the incognoscenti, the two terms are quite different from one another. A specification is a formal way of describing data. A standard is a set of requirements, created by an standards development organization, that comprise a pre-determined content and format for a set of data.

More on specifications in tomorrow's blog.

References:

[1] Berman JJ. Repurposing Legacy Data: Innovative Case Studies. Morgan Kaufmann, Waltham, MA, 2015.

[2] FIPS PUB 119-1. Supersedes FIPS PUB 119. 1985 November 8. Federal Information Processing Standards Publication 119-1 1995 March 13. Announcing the Standard for ADA. Available from: http://www.itl.nist.gov/fipspubs/fip119-1.htm, viewed August 26, 2012.

[3] Kammer RG. The Role of Standards in Today's Society and in the Future. Statement of Raymond G. Kammer, Director, National Institute of Standards and Technology, Technology Administration, Department of Commerce, Before the House Committee on Science Subcommittee on Technology, September 13, 2000.

[4] Berman JJ. Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information. Morgan Kaufmann, Waltham, MA, 2013.

[5] National Committee on Vital and Health Statistics. Report to the Secretary of the U.S. Department of Health and Human Services on Uniform Data Standards for Patient Medical Record Information. July 6, 2000. Available from: http://www.ncvhs.hhs.gov/hipaa000706.pdf


- Jules Berman

key words: computer science, data analysis, data repurposing, data simplification, data wrangling, information science, simplifying data, taming data, complexity, standards, specificationsjules j berman

Monday, March 7, 2016

DATA SIMPLIFICATION: Poor Identifiers, Horrific Consequences

Over the next few weeks, I will be writing on topics related to my latest book, Data Simplification: Taming Information With Open Source Tools (release date March 23, 2016). I hope I can convince you that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.



All information systems, all databases, and all good collections of data are best envisioned as identifier systems to which data (belonging to the identifier) can be added over time.

If the system is corrupted (e.g., multiple identifiers for the same object, data belonging to one object incorrectly attached to other objects), then the system has no value. You can't trust any of the individual records, and you can't trust any of the analyses performed on collections of records. Furthermore, if the data from a corrupted system is merged with the data from other systems, then all analyses performed on the aggregated data becomes unreliable and useless. This holds true even when every other contributor to the system shares reliable data.

Without proper identifiers, the following may occur: data values can be assigned to the wrong data objects; data objects can be replicated under different identifiers, with each replicant having an incomplete data record (i.e., an incomplete set of data values); the total number of data objects cannot be determined; data sets cannot be verified; and the results of data set analyses will not be valid.

In the past, individuals were identified by their name. When dealing with large numbers of names, it becomes obvious, almost immediately, that personal names are woefully inadequate. In a review of a population of 3.5 million, there occurred nearly 250,000 instances wherein individuals shared the same first and last name; there were 70,000 instances where two people shared the same first name, last name and birthdate (1)!

Aside from the obvious fact that they are not unique (e.g., surnames such as Smith, Zhang, Garcia, Lo, and given names such as John and Susan), one name can have multiple representations. The sources for these variations are many. Here is a partial listing (2):

1. Modifiers to the surname (du Bois, DuBois, Du Bois, Dubois, Laplace, La Place, van de Wilde, Van DeWilde, etc.).

3. Accents that may or may not be transcribed onto records (e.g., acute accent, cedilla, diacritical comma, palatalized mark, hyphen, diphthong, umlaut, circumflex, and a host of obscure markings).

4. Special typographic characters (the combined "ae").

5. Multiple "middle names" for an individual, that may not be transcribed onto records. Individuals who replace their first name with their middle name for common usage, while retaining the first name for legal documents.

6. Latinized and other versions of a single name (Carl Linnaeus, Carl von Linne, Carolus Linnaeus, Carolus a Linne).

7. Hyphenated names that are confused with first and middle names (e.g., Jean-Jacques Rousseau, or Jean Jacques Rousseau; Louis-Victor-Pierre-Raymond, 7th duc de Broglie, or Louis Victor Pierre Raymond Seventh duc deBroglie).

8. Cultural variations in name order that are mistakenly re-arranged when transcribed onto records. Many cultures do not adhere to the Western European name order (e.g., given name, middle name, surname).

9. Name changes; through marriage, legal action, aliasing, pseudonymous posing, or insouciant whim.

I have had numerous conversations with intelligent professionals who are tasked with the responsibility of assigning identifiers to individuals. At some point in every conversation, they will find it necessary to explain that although an individual's name cannot serve as an identifier, the combination of name plus date of birth provides accurate identification in almost every instance. They sometimes get carried away, insisting that the combination of name plus date of birth plus social security number provides perfect identification, as no two people will share all three identifiers: same name, same date of birth, same social security number. This is simply wrong. Let us see what happens when we create identifiers from the name plus birthdate.

Consider this example. Mary Jessica Meagher, born June 7, 1912 decided to open a separate bank account in each of 10 different banks. Some of the banks had application forms, which she filled out accurately. Other banks registered her account through a teller, who asked her a series of questions and immediately transcribed her answers directly into a computer terminal. Ms. Meagher could not see the computer screen and could not review the entries for accuracy.

Here are the entries for her name plus date of birth (1):

1. Marie Jessica Meagher, June 7, 1912 (the teller mistook Marie for Mary).

2. Mary J. Meagher, June 7, 1912 (the form requested a middle initial, not name).

3. Mary Jessica Magher, June 7, 1912 (the teller misspelled the surname).

4. Mary Jessica Meagher, Jan 7, 1912 (the birth month was constrained, on the form, to three letters; Jun, entered on the form, was transcribed as Jan).

5. Mary Jessica Meagher, 6/7/12 (the form provided spaces for the final two digits of the birth year. Through the miracle of bank registration, Mary, born in 1912, was re-born a century later).

6. Mary Jessica Meagher, 7/6/2012 (the form asked for day, month, year, in that order, as is common in Europe).

7. Mary Jessica Meagher, June 1, 1912 (on the form, a 7 was mistaken for a 1).

8. Mary Jessie Meagher, June 7, 1912 (Marie, as a child, was called by the informal form of her middle name, which she provided to the teller).

9. Mary Jessie Meagher, June 7, 1912 (Marie, as a child, was called by the informal form of her middle name, which she provided to the teller, and which the teller entered as the male variant of the name).

10. Marie Jesse Mahrer, 1/1/12 (an underzealous clerk combined all of the mistakes on the form and the computer transcript, and added a new orthographic variant of the surname).

For each of these ten examples, a unique individual (Mary Jessica Meagher) would be assigned a different identifier at each of 10 banks. Had Mary re-registered at one bank, ten times, the results may have been the same.

References:

[1] McCann E. The patient identifier debate: Will a national patient ID system ever materialize? Should it? Healthcare IT News February 18, 2013.

[2] Berman JJ. Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information. Morgan Kaufmann, Waltham, MA, 2013.


- Jules Berman

key words: computer science, data analysis, data repurposing, data simplification, data wrangling, information science, simplifying data, taming data, complexity, identifiers, jules j berman

Sunday, March 6, 2016

Data Simplification: Identifiers


Over the next few weeks, I will be writing on topics related to my latest book, Data Simplification: Taming Information With Open Source Tools (release date March 23, 2016). I hope I can convince you that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.


"I always wanted to be somebody, but now I realize I should have been more specific." -Lily Tomlin

An object identifier is anything associated with the object that persists throughout the life of the object and that is unique to the object (i.e., does not belong to any other object). Everyone is familiar with biometric identifiers, such as fingerprints, iris patterns, and genome sequences. In the case of data objects, the identifier usually refers to a randomly chosen long sequence of numbers and letters that is permanently assigned to the object and which is never assigned to any other data object.

An identifier system is a set of data-related protocols that satisfy the following conditions: 1) Completeness (i.e., every unique object has an identifier); 2) Uniqueness (i.e., each identifier is a unique sequence); 3. Exclusivity (i.e., each identifier is assigned to only one unique object, and to no other object, ever); 4) Authenticity (i.e., objects that receive identification can be verified as the objects that they are intended to be); 5) Aggregation (all information associated with an identifier can be collected); and 6) Permanence (i.e., an identifier is never deleted).

Uniqueness is a very strange concept, especially when applied to the realm of data. For example, if I refer to the number 1, then I am referring to a unique number among other numbers (i.e., there is only one number 1). Yet the number 1 may apply to many different things (i.e., 1 left shoe, 1 umbrella, 1 prime number between 2 and 5). The number 1 makes very little sense to us until we know something about what it measures (e.g., left shoe) and the object to which the measurement applies (e.g., shoe_id_#840354751) (1).

We refer to uniquely-assigned computer-generated character strings as "identifiers". As such, computer-generated identifiers are abstract constructs that do not need to embody any of the natural properties of the object. A long (e.g., 200 character length) character string, consisting of randomly chosen numeric and alphabetic characters is an excellent identifier, because the chances of two individuals being assigned the same string is essentially zero. When we need to establish the uniqueness of some object, such as a shoe or a data record, we bind the object to a contrived identifier.

Jumping ahead just a bit, if we say "part number 32027563 weighs 1 pound," then we are dealing with a meaningful assertion . The assertion tells us three things: 1) that there is a unique thing, known as part number 32027563, 2) the unique thing has a weight, and 3) the weight has a measurement of 1 pound. The phrase "weighs 1 pound" has no meaning until it is associated with a unique object (i.e., part number 32027563 weighs 1 pound). The assertion that "part number 32027563 weighs 1 pound" is a "triple," the embodiment of meaning in the field of computational semantics. A triple consists of a unique, identified object, matched to a pair of data and metadata (i.e., a data element and the description of the data element). Information experts use formal syntax to express triples as data structures.

Returning to the issue of object identification, there are various methods for generating and assigning unique identifiers to data objects (2), (3), (4), (5). Some identification systems assign a group prefix to an identifier sequence that is unique for the members of the group. For example, a prefix for a research institute may be attached to every data object generated within the institute. If the prefix is registered in a public repository, data from the institute can be merged with data from other institutes, and the institutional source of the data object can always be determined. The value of prefixes, and other reserved namespace designations, can be undermined when implemented thoughtlessly (1).

Identifiers are data simplifiers, when implemented properly, because they allow us to collect all of the data associated with a unique object, while ensuring that we exclude that data that should be associated with some other object.

UUID (Universally Unique IDentifier) is an example of one type of algorithm that creates collision-free identifiers that can be generated on command, at the moment when new objects are created (i.e., during the run-time of a software application). Linux systems have a built-in UUID utility, "uuidgen.exe", that can be called from the system prompt.

Here are a few examples of output values generated by the "uuidgen.exe" utility:
$ uuidgen.exe
312e60c9-3d00-4e3f-a013-0d6cb1c9a9fe

$ uuidgen.exe
822df73c-8e54-45b5-9632-e2676d178664

$ uuidgen.exe
8f8633e1-8161-4364-9e98-fdf37205df2f

$ uuidgen.exe
83951b71-1e5e-4c56-bd28-c0c45f52cb8a

$ uuidgen -t
e6325fb6-5c65-11e5-b0e1-0ceee6e0b993

$ uuidgen -r
5d74e36a-4ccb-42f7-9223-84eed03291f9
Data Simplification: Taming Information With Open Source Tools describes simple implementions of UUID utilities under Windows.

Notice that each of the final two examples have a parameter added to the "uuidgen" command (i.e., "-t" and "-r"). There are several versions of the UUID algorithm that are available. The "-t" parameter instructs the utility to produce a UUID based on the time (measured in seconds elapsed since the first second of October 15, 1582, the start of the Gregorian calendar). The "-r" parameter instructs the utility to produce a UUID based on the generation of a pseudorandom number. In any circumstance, the UUID utility produces a fixed length character string suitable as an object identifier. The UUID utility is trusted and widely used by computer scientists.

References:

[1] Berman JJ. Repurposing Legacy Data: Innovative Case Studies. Morgan Kaufmann, Waltham, MA, 2015.

[2] Leach P, Mealling M, Salz R. A Universally Unique IDentifier (UUID) URN Namespace. Network Working Group, Request for Comment 4122, Standards Track. Available from: http://www.ietf.org/rfc/rfc4122.txt, viewed Jan. 1, 2015.

[3] Mealling M. RFC 3061. A URN Namespace of Object Identifiers. Network Working Group, 2001. Available from: https://www.ietf.org/rfc/rfc3061.txt, view Jan. 1, 2015.

[4] Berman JJ. Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information. Morgan Kaufmann, Waltham, MA, 2013.

[5] Berman JJ. Methods in Medical Informatics: Fundamentals of Healthcare Programming in Perl, Python, and Ruby. Chapman and Hall, Boca Raton 2010.


- Jules Berman

key words: computer science, data analysis, data repurposing, data simplification, data wrangling, information science, simplifying data, taming data, complexity, identifiers, jules j berman

Saturday, March 5, 2016

Data Simplification: Hitting the Complexity Barrier


Over the next few weeks, I will be writing on topics related to my latest book, Data Simplification: Taming Information With Open Source Tools (release date March 23, 2016). I hope I can convince you that this is a book worth reading.


Blog readers can use the discount code: COMP315 for a 30% discount, at checkout.


"Nobody goes there anymore. It's too crowded." -Yogi Berra

It seems that many scientific findings, particularly those findings based on analyses of large and complex data sets, are yielding irreproducible results. We find that we can not depend on the data that we depend on. If you don't believe me, consider these shocking headlines:

1. "Unreliable research: Trouble at the lab." (1) The Economist, in 2013 ran an article examining flawed biomedical research. The magazine article referred to an NIH official who indicated that "researchers would find it hard to reproduce at least three-quarters of all published biomedical findings, the public part of the process seems to have failed." The article described a study conducted at the pharmaceutical company Amgen, wherein 53 landmark studies were repeated. The Amgen scientists were successful at reproducing the results of only 6 of the 53 studies. Another group, at Bayer HealthCare, repeated 63 studies. The Bayer group succeeded in reproducing the results of only one-fourth of the original studies.

2. "A decade of reversal: an analysis of 146 contradicted medical practices." (2) The authors reviewed 363 journal articles, reexamining established standards of medical care. Among these articles were 146 manuscripts (40.2%) claiming that an existing standard of care had no clinical value.

3."Cancer fight: unclear tests for new drug." (3). This New York Times article examined whether a common test performed on breast cancer tissue (Her2) was repeatable. It was shown that for patients who tested positive for Her2, a repeat test indicated that 20% of the original positive assays were actually negative (i.e., falsely positive on the initial test). (3).

4. "Reproducibility crisis: Blame it on the antibodies" (4). Biomarker developers are finding that they cannot rely on different batches of a reagent to react in a consistent manner, from test to test. Hence, laboratory analytic methods, developed using a controlled set of reagents, may not have any diagnostic value when applied by other laboratories, using different sets of the same analytes (4)

5. "Why most published research findings are false." (5). Modern scientists often search for small effect sizes, using a wide range of available analytic techniques, and a flexible interpretation of outcome results. The manuscript's author found that research conclusions are more likely to be false than true (5), (6).

6. "We found only one-third of published psychology research is reliable - now what?" (7). The manuscript authors suggest that the results of a first study should be considered preliminary and tentative. Conclusions have no value until they are independently validated.

Anyone who attempts to stay current in the sciences soon learns that much of the published literature is irreproducible (8); and that almost anything published today might be retracted tomorrow. This appalling truth applies to some of the most respected and trusted laboratories in the world (9), (10), (11), (12), (13), (14), (15), (16). Those of us who have been involved in assessing the rate of progress in disease research are painfully aware of the numerous reports indicating a general slowdown in medical progress (17), (18), (19), (20), (21), (22), (23), (24).

For the optimists, it is tempting to assume that the problems that we may be experiencing today are par for the course, and temporary. It is the nature of science to stall for a while and lurch forwards in sudden fits. Errors and retractions will always be with us so long as humans are involved in the scientific process.

For the pessimists, such as myself, there seems to be something going on that is really new and different; a game changer. This game changer is the "complexity barrier", a term credited to Boris Beizer, who used it to describe the impossibility of managing increasingly complex software products (25). The complexity barrier, known also as the complexity ceiling, reflects the intricacies of Big Science and applies to most of the data analysis efforts undertaken these days (26), (27).

Some of the mistakes that lead to erroneous conclusions in data-intensive research are well-known, and include the following:

1. Errors in sample selection, labeling, and measurement (28), (29), (30). For example, modern biomedical data is high-volume (e.g., gigabytes and larger), heterogeneous (i.e., derived from diverse sources), private (i.e., measured on human subjects), and multi-dimensional (e.g., containing thousands of different measurements for each data record). The complexities of handling such data correctly are daunting (31).

3. Misinterpretation of the data (32), (5), (33), (22), (34), (35), (36), (37)

4. Data hiding and data obfuscation (38), (39)

5. Unverified and unvalidated data (40), (41), (42), (43), (34), (44)

6. Outright fraud (39), (16), (45).

When errors occur in complex data analyses, they are notoriously difficult to discover (40).

Aside from human error, intrinsic properties of complex systems may thwart our best attempts at analysis. For example, when complex systems are perturbed from their normal, steady-state activities, the rules that govern the system's behavior become unpredictable (46). Much of the well-managed complexity of the world is found in machines built with precision parts having known functionality. For example, when an engineer designs a radio, she knows that she can assign names to components, and these components can be relied upon to behave in a manner that is characteristic of its type. A capacitor will behave like a capacitor, and a resistor will behave like a resistor. The engineer need not worry that the capacitor will behave like a semiconductor or an integrated circuit. The engineer knows that the function of a machine's component will never change; but the biologist operates in a world wherein components change their functions, from organism to organism, cell to cell and moment to moment. As an example, cancer researchers discovered an important protein that plays a role in the development of cancer. This protein, p53, was considered to be the primary cellular driver for human malignancy. When p53 mutated, cellular regulation was disrupted, and cells proceeded down a slippery path leading to cancer. In the past few decades, as more information was obtained, cancer researchers have learned that p53 is just one of many proteins that play some role in carcinogenesis, and that the role played by p53 changes depending on the species, tissue type, cellular microenvironment, genetic background of the cell, and many other factors. Under one set of circumstances, p53 may modify DNA repair; under another set of circumstances, p53 may cause cells to arrest the growth cycle (47), (48). It is difficult to predict the biological effect of a protein that changes its primary function based on prevailing cellular conditions.

At the heart of all data analysis is the assumption that systems have a behavior that can be described with a formula or a law, or that can lead to results that are repeatable and to conclusions that can be validated. We are now learning that our assumptions may have been wrong, and that our best efforts at data analysis may be irreproducible.

Science and society may have reached a complexity barrier beyond which nothing can be analyzed and understood with any confidence. In light of the irreproducibility of complex data analyses, it seems prudent to take the follow the following two recommendations:

1. Simplify your complex data, before you attempt analysis.

2. Assume that the first analysis of primary data is tentative and probably wrong. The most important purpose of data analysis is to lay the groundwork for data reanalysis.

References:

[1] Unreliable research: Trouble at the lab. The Economist October 19, 2013.

[2] Prasad V, Vandross A, Toomey C, Cheung M, Rho J, Quinn S, et al. A decade of reversal: an analysis of 146 contradicted medical practices. Mayo Clin Proc 88:790-8, 2013.

[3] Kolata G. Cancer fight: unclear tests for new drug. The New York Times April 19, 2010.

[4] Baker M. Reproducibility crisis: Blame it on the antibodies. Nature 521:274-276, 2015.

[5] Ioannidis JP. Why most published research findings are false. PLoS Med 2:e124, 2005.

[6] Labos C. It Ain't Necessarily So: Why Much of the Medical Literature Is Wrong. Medscape News and Perspectives. September 09, 2014

[7] Gilbert E, Strohminger N. We found only one-third of published psychology research is reliable - now what? The Conversation. August 27, 2015. Available at: http://theconversation.com/we-found-only-one-third-of-published-psychology-research-is-reliable-now-what-46596, viewed on August 27,2015.

[8] Naik G. Scientists' Elusive Goal: Reproducing Study Results. Wall Street Journal December 2, 2011.

[9] Zimmer C. A sharp rise in retractions prompts calls for reform. The New York Times April 16, 2012.

[10] Altman LK. Falsified data found in gene studies. The New York Times October 30, 1996.

[11] Weaver D, Albanese C, Costantini F, Baltimore D. Retraction: altered repertoire of endogenous immunoglobulin gene expression in transgenic mice containing a rearranged mu heavy chain gene. Cell 65:536 (inclusive), 1991.

[12] Chang K. Nobel winner in physiology retracts two papers. The New York Times September 23, 1010.

[13] Fourth paper retracted at Potti's request. The Chronicle March 3, 2011.

[14] Whoriskey P. Doubts about Johns Hopkins research have gone unanswered, scientist says. The Washington Post March 11, 2013.

[15] Lin YY, Kiihl S, Suhail Y, Liu SY, Chou YH, Kuang Z, et al. Retraction: Functional dissection of lysine deacetylases reveals that HDAC1 and p300 regulate AMPK. Nature 482:251-255, retracted November, 2013.

[16] Shafer SL. Letter: To our readers. Anesthesia and Analgesia. February 20, 2009.

[17] Innovation or Stagnation: Challenge and Opportunity on the Critical Path to New Medical Products. U.S. Department of Health and Human Services, Food and Drug Administration, 2004.

[18] Hurley D. Why Are So Few Blockbuster Drugs Invented Today? The New York Times November 13, 2014.

[19] Angell M. The Truth About the Drug Companies. The New York Review of Books Vol 51, July 15, 2004.

[20] Crossing the Quality Chasm: A New Health System for the 21st Century. Quality of Health Care in America Committee, editors. Institute of Medicine, Washington, DC., 2001.

[21] Wurtman RJ, Bettiker RL. The slowing of treatment discovery, 1965-1995. Nat Med 2:5-6, 1996.

[22] Ioannidis JP. Microarrays and molecular research: noise discovery? The Lancet 365:454-455, 2005.

[23] Weigelt B, Reis-Filho JS. Molecular profiling currently offers no more than tumour morphology and basic immunohistochemistry. Breast Cancer Research 12:S5, 2010.

[24] Personalised medicines: hopes and realities. The Royal Society, London, 2005.Available from: https://royalsociety.org/~/media/Royal_Society_Content/policy/publications/2005/9631.pdf, viewed Jan 1, 2015.

[25] Beizer B. Software Testing Techniques. Van Nostrand Reinhold; Hoboken, NJ 2 edition, 1990.

[26] Vlasic B. Toyota's slow awakening to a deadly problem. The New York Times, February 1, 2010.

[27] Lanier J. The complexity ceiling. In: Brockman J, ed. The next fifty years: science in the first half of the twenty-first century. Vintage, New York, pp 216-229, 2002.

[28] Bandelt H, Salas A. Contamination and sample mix-up can best explain some patterns of mtDNA instabilities in buccal cells and oral squamous cell carcinoma. BMC Cancer 9:113, 2009.

[29] Knight, J. Agony for researchers as mix-up forces retraction of ecstasy study. Nature 425:109, September 11, 2003.

[30] Gerlinger M, Rowan AJ, Horswell S, Larkin J, Endesfelder D, Gronroos E, et al. Intratumor heterogeneity and branched evolution revealed by multiregion sequencing. N Engl J Med 366:883-892, 2012.

[31] Berman JJ. Biomedical Informatics. Jones and Bartlett, Sudbury, MA, 2007.

[32] Ioannidis JP. Is molecular profiling ready for use in clinical decision making? The Oncologist 12:301-311, 2007.

[33] Ioannidis JP. Some main problems eroding the credibility and relevance of randomized trials. Bull NYU Hosp Jt Dis 66:135-139, 2008.

[34] Ioannidis JP, Panagiotou OA. Comparison of effect sizes associated with biomarkers reported in highly cited individual articles and in subsequent meta-analyses. JAMA 305:2200-2210, 2011.

[35] Ioannidis JPA, Panagiotou OA. "Comparison of effect sizes associated with biomarkers reported in highly cited individual articles and in subsequent meta-analyses. JAMA 305:2200-2210, 2011.

[36] Ioannidis JP: Excess significance bias in the literature on brain volume abnormalities. Arch Gen Psychiatry 68:773-780, 2011.

[37] Pocock SJ, Collier TJ, Dandreo KJ, deStavola BL, Goldman MB, Kalish LA, et al. Issues in the reporting of epidemiological studies: a survey of recent practice. BMJ 329:883, 2004.

[38] Harris G. Diabetes drug maker hid test data, files indicate. The New York Times July 12, 2010.

[39] Berman JJ. Machiavelli's Laboratory. Amazon Digital Services, Inc., 2010.

[40] Misconduct in science: an array of errors. The Economist. September 10, 2011.

[41] Begley S. In cancer science, many 'discoveries' don't hold up. Reuters Mar 28, 2012,

[42] Abu-Asab MS, Chaouchi M, Alesci S, Galli S, Laassri M, Cheema AK, et al. Biomarkers in the age of omics: time for a systems biology approach. OMICS 15:105-112, 2011.

[43] Moyer VA; on behalf of the U.S. Preventive Services Task Force. Screening for prostate cancer: U.S. Preventive Services Task Force recommendation statement. Ann Intern Med May 21, 2011

[44] How science goes wrong. The Economist Oct 19, 2013.

[45] Martin B. Scientific fraud and the power structure of science. Prometheus 10:83-98, 1992.

[46] Rosen JM, Jordan CT. The increasing complexity of the cancer stem cell paradigm. Science 324:1670-1673, 2009.

[47] Madar S, Goldstein I, Rotter V. Did experimental biology die? Lessons from 30 years of p53 research. Cancer Res 2009;69:6378-6380, 2009.

[48] Zilfou JT, Lowe SW. Tumor Suppressive Functions of p53. Cold Spring Harb Perspect Biol 00:a001883, 2009.


- Jules Berman

key words: computer science, data analysis, data repurposing, data simplification, data wrangling, information science, simplifying data, taming data, complexity, jules j berman