Showing posts with label ontology. Show all posts
Showing posts with label ontology. Show all posts

Tuesday, March 29, 2016

CLASS BLENDING: Simpson's Paradox

For the past two days, we've been posting on Class Blending. Simpson's paradox is a special case that demonstrates what may happen when classes of information are blended.


Simpson's paradox is a well-known problem for statisticians. The paradox is based on the observation that findings that apply to each of two data sets may be reversed when the two data sets are combined.

One of the most famous examples of Simpson's paradox was demonstrated in the 1973 Berkeley gender bias study (1). A preliminary review of admissions data indicated that women had a lower admissions rate than men:
Men    Number of applicants.. 8,442   Percent applicants admitted.. 44%
Women  Number of applicants.. 4,321   Percent applicants admitted.. 35%
A nearly 10% difference is highly significant, but what does it mean? Was the admissions office guilty of gender bias?

A closer look at admissions department-by-department showed a very different story. Women were being admitted at higher rates than men, in almost every department. The department-by-department data seemed incompatible with the combined data.

The explanation was simple. Women tended to apply to the most popular and oversubscribed departments, such as English and History, that had a high rate of admission denials. Men tended to apply to departments that the women of 1973 avoided, such as mathematics, engineering and physics. Men tended not to apply to the high occupancy departments that women preferred. Though women had an equal footing with men in departmental admissions, the high rate of women rejections in the large, high-rejection departments, accounted for an overall lower acceptance rate for women at Berkeley.

Simpson's paradox demonstrates that data is not additive. It also shows us that data is not transitive; you cannot make inferences based on subset comparisons. For example in randomized drug trials, you cannot assume that if drug A tests better than drug B, and drug B tests better than drug C, then drug A will test better than drug C (2). When drugs are tested, even in well-designed trials, the test populations are drawn from a general population specific for the trial. When you compare results from different trials, you can never be sure whether the different sets of subjects are comparable. Each set may contain individuals whose responses to a third drug are unpredictable. Transitive inferences (i.e., if A is better than B, and B is better than C, then A is better than C), are unreliable.

- Jules Berman (copyrighted material)

key words: data science, irreproducible results, complexity, classification, ontology, ontologies, classifications, data simplification, jules j berman

Reference:

1. Bickel PJ, Hammel EA, O'Connell JW. Sex Bias in Graduate Admissions: Data from Berkeley. Science 187:398-404, 1975.

2. Baker SG, Kramer BS. The transitive fallacy for randomized trials: If A bests B and B bests C in separate trials, is A better than C? BMC Medical Research Methodology 2:13, 2002

Sunday, March 27, 2016

Expunging a Blended Class: The Fall of Kingdom Protozoa

In yesterday's blog, we introduced and defined the term "Class blending". Today's blog extends this discussion by describing the most significant and most enduring class blending error to impact the natural sciences: the artifactual blending of all single cell organisms into the blended class, Protozoa.

For well over a century, biologists had a very simple way of organizing the eukaryotes (i.e., the organisms that were not bacteria, whose cells contained a nucleus) (1). Basically, the one-celled organisms were all lumped into one biological class, the protozoans (also called protists). With the exception of animals and plants, and some of the fungi (e.g., mushrooms), life on earth is unicellular. The idea of lumping every type of unicellular organism into one class, having shared properties, shared ancestry, and shared descendants, made no sense. What's more, the leading taxonomists of the nineteenth century, such as Ernst Haeckel (1834 - 1919), understood the class Protozoa was at best, a temporary grab-bag holding unrelated organisms that would eventually be split into their own classes. Well, a century passed, and complacent taxonomists preserved the Protozoan class. In the 1950s, Robert Whittaker elevated Class Protozoa as a kingdom in his broad new "Five Kingdom" classification of living organisms (2). This classification (more accurately, misclassification) persisted through the last five decades of the twentieth century.

Modern classifications, based on genetics, metabolic pathways, shared morphologic features, and evolutionary lineage, have dispensed with Class Protozoa, assigning each individual class of eukaryotes to its own hierarchical position. A simple schema demonstrates the modern classification of eukaryotes (3). Many modern taxonomists are busy improving this fluid list (vida infra), but, most significantly, Class Protozoa is nowhere to be found.
Eukaryota (organisms that have nucleated cells)
  Bikonta (2-flagella)
    Excavata
      Metamonada
      Discoba
        Euglenozoa
        Percolozoa
    Archaeplastida, from which Kingdom Plantae derives
    Chromalveolata
      Alveolata
        Apicomplexa
        Ciliophora
      Heterokontophyta
  Unikonta
    Amoebozoa
    Opisthokonta
      Choanozoa
      Animalia
      Fungi
Why is it important to expunge Class Protozoa from modern classifications of living organisms? Every class of living organism contains members that are pathogenic to other classes of organisms. To the point, most classes of organisms contain members that are pathogenic to humans, or to the organisms that humans depend on for their existence (e.g., other animals, food plants, beneficial organisms). There are way too many species of pathogens for us to develop specific drugs and techniques to control the growth of each disease-causing organism. Our only hope is to develop general treatments for classes of organisms, that share the same properties; hence the same weaknesses. For example, in theory, it's much easier to develop drugs that work on Apicomplexans that it is to develop separate drugs that work on each pathogenic species of Apicomplexan (3).

By lumping every single-celled organisms into one blended class, we have missed the opportunity to develop true class-based remedies for the most elusive disease-causing organisms on our planet. The past two decades have seen enormous progress in reclassifying the former protozoans. Unfortunately, the errors of the past are repeated in textbooks and dictionaries.

Here are three definitions of protozoa that I found on the web. Notice that these definitions don't even agree with one another. Notice that the first definition includes single celled organisms that may be free-living or parasitic. The second definition indicates that protozoans are obligate intracellular organisms. The third definition indicates that some protozoans are pathogenic in animals but omits mention of pathogenicity for other types of organisms. None of the definitions tell us that modern taxonomists have abandoned "protozoa" as a bona fide class of organisms.

from: http://www.dictionary.com/browse/protozoan
Protozoan: Any of a large group of one-celled organisms (called protists) that live in water or as parasites. Many protozoans move about by means of appendages known as cilia or flagella. Protozoans include the amoebas, flagellates, foraminiferans, and ciliates.

from: www.medicinenet.com/script/main/art.asp?articlekey=5091
Protozoa: A parasitic single-celled organism that can divide only within a host organism. For example, malaria is caused by the protozoa Plasmodium.

from: http://www.merriam-webster.com/dictionary/protozoan
Protozoan: any of a phylum or subkingdom (Protozoa) of chiefly motile and heterotrophic unicellular protists (as amoebas, trypanosomes, sporozoans, and paramecia) that are represented in almost every kind of habitat and include some pathogenic parasites of humans and domestic animals.


References:

[1] Scamardella JM. Not plants or animals: a brief history of the origin of Kingdoms Protozoa, Protista and Protoctista. Internatl Microbiol 2:207-216, 1999.

[2] Hagen JB. Five kingdoms, more or less: Robert Whittaker and the broad classification of organisms. BioScience 62:67-74, 2012.

[3] Berman JJ. Taxonomic Guide to Infectious Diseases: Understanding the Biologic Classes of Pathogenic Organisms. Academic Press, Waltham, 2012.


- Jules Berman (copyrighted material)

key words: data science, irreproducible results, complexity, classification, ontology, ontologies, protozoa, Apicomplexa, protists, protoctista,jules j berman

Saturday, March 26, 2016

Intro to Class Blending

I thought I'd devote the next few blogs to a concept that has gotten much less attention than it deserves: blended classes. Class blending lurks behind much of the irreproducibility in "Big Science" research, including clinical trials. It also is responsible for impeding progress in various disciplines of science, particularly the natural sciences, where classification is of utmost importance. We'll see that the scientific literature is rife with research of dubious quality, based on poorly designed classifications and blended classes.

For today, let's start with a definition and one example. We'll discuss many more specific examples in future blogs.

Blended class - Also known as class noise, subsumes the more familiar, but less precise term, "Labeling error." Blended class refers to inaccuracies (e.g., misleading results) introduced in the analysis of data due to errors in class assignments (i.e., assigning a data object to class A when the object should have been assigned to class B). If you are testing the effectiveness of an antibiotic on a class of people with bacterial pneumonia, the accuracy of your results will be forfeit when your study population includes subjects with viral pneumonia, or smoking-related lung damage. Errors induced by blending classes are often overlooked by data analysts who incorrectly assume that the experiment was designed to ensure that each data group is composed of a uniform and representative population. A common source of class blending occurs when the classification upon which the experiment is designed is itself blended. For example, imagine that you are a cancer researcher and you want to perform a study of patients with malignant fibrous histiocytomas (MFH), comparing the clinical course of these patients with the clinical course of patients who have other types of tumors. Let's imagine that the class of tumors known as MFH does not actually exist; that it is a grab-bag term erroneously assigned to a variety of other tumors that happened to look similar to one another. This being the case, it would be impossible to produce any valid results based on a study of patients diagnosed as MFH. The results would be a biased and irreproducible cacaphony of data collected across different, and undetermined, species of tumors. This specific example, of the blended MFH class of tumors, is selected from the real-life annals of tumor biology (1), (2).

References:

[1] Al-Agha OM, Igbokwe AA. Malignant fibrous histiocytoma: between the past and the present. Arch Pathol Lab Med 132:1030-1035, 2008.

[2] Nakayama R, Nemoto T, Takahashi H, Ohta T, Kawai A, Seki K, et al. Gene expression analysis of soft tissue sarcomas: characterization and reclassification of malignant fibrous histiocytoma. Modern Pathology 20:749-759, 2007.


- Jules Berman (copyrighted material)

key words: data science, irreproducible results, complexity, classification, ontology, ontologies, jules j berman

Thursday, February 4, 2016

A Species is a Biological Entity; Not a Mere Intellectual Abstraction

In the Disney retelling of a classic fairy tale, a human-made abstraction, a puppet named Pinocchio survives a series of perils and emerges as a real live boy. It seems farfetched that an abstraction could become a living biological organism, but it happens. In point of fact, the transformation of an abstract idea into a living entity is one of the most important scientific advancements of the past half century. For the most part, this miracle of science has gone unheralded. Nonetheless, if you think very deeply about the meaning of classifications, and if you can appreciate the role played by abstractions in the governance of our physical universe, you will appreciate the profound implications of the following story. We shall see that a human-made abstraction, that we name "species", has survived a series of perils, and has emerged as a real live biological entity.

In the classification of living terrestrial organisms, the bottom classes are known as "species". There is a species class for all the horses and another species class for all the squirrels, and so on. Speculation has it that there are 50 to 100 million different species of organisms on planet earth. We humans have assigned names to a few million species, a small fraction of the total.

It has been argued that nature produces individuals, not species; the concept of species being a mere figment of the human imagination, created for the convenience of taxonomists who need to group similar organisms. Biologists can collect feature data such as gene sequences, geographic habitat, diet, size, mating rituals, hair color, shape of skull and so on, for a variety of different animals. After some analysis, perhaps performed with the aid of a computer, we could cluster animals based on their similarities, and we could assign the clusters names, and the names of our clusters would be our species. The arbitrariness of species creation comes from the various ways we might select the features to be measured in our data sets, the choice of weights assigned to the the different features (e.g., should we give more weight to gene sequence than to length of gestation?), and to our choice of algorithm for assigning organisms to groups.

For myself, and for many other scientists who use classification, there can be no human arbitrariness in the assignment of species (1). A species is a fundamental building block of the natural world, no less substantial than the concept of a galaxy to astronomers or the number "e" to mathematicians.

The modern definition of species is "an evolving gene pool." As such, species have three properties that prove that they are biological entities.

1. Unique definition. Until recently, biologists could not agree on a definition of species. There were dozens of definitions to choose from, depending on which field of science you studied. Molecular biologists defined species by gene sequence. Zoologists defined species by mating exclusivity. Ecologists defined species by habitat constraints. The current definition equating species with an evolving gene pool serves as a great unifying theory for biologists.

2. The class "species" has a biological function that is not available to individual members of the species; namely, speciation. Species propagate, and when they do, they produce new species. Species are the only biological entities that can produce new species.

3. Species evolve. Individuals do not evolve. Evolution requires a gene pool; something that species have and individuals to not.


Species bear a biological relationship to individual organisms. Just as species are defined as evolving gene pools, individual organisms can be defined as set of propagating genes living within a cellular husk. Hence, the individual organism has a genome taken from the pool of genes available to his species.

The classification of living organisms has worked a true miracle, by breathing life into the concept of species, thus expanding reality.


[1] DeQueiroz K. Ernst Mayr and the modern concept of species. PNAS 102(suppl 1):6600-6607, 2005.

- Jules Berman (copyrighted material)

key words: classsification, ontology, species, speciation, jules j berman

Wednesday, February 3, 2016

Unclassifiable objects

Classifications create a class for every object and taxonomies assign each and every object to its correct class. This means that a classification is not permitted to contain unclassified objects; a condition that puts fussy taxonomists in an untenable position. Suppose you have an object, and you simply do not know enough about the object to confidently assign it to a class. Or, suppose you have an object that seems to fit more than one class, and you can't decide which class is the correct class. What do you do? Historically, scientists have resorted to creating a "miscellaneous" class into which otherwise unclassifiable objects are given a temporary home, until more suitable accommodations can be provided. I have spoken with numerous data managers, and everyone seems to be of a mind that "miscellaneous" classes, created as a stopgap measure, serve a useful purpose. Not so. Historically, the promiscuous application of "miscellaneous" classes have proven to be a huge impediment to the advancement of science. In the case of the classification of living organisms, the class of protozoans stands as a case in point. Ernst Haeckel, a leading biological taxonomist in his time, created the Kingdom Protista (i.e., protozoans), in 1866, to accommodate a wide variety of of simple organisms with superficial commonalities. Haeckel himself understood that the protists were a blended class that included unrelated organisms, but he believed that further study would resolve the confusion. In a sense, he was right, but the process took much longer than he had anticipated; occupying generations of taxonomists over the following 150 years. Today, Kingdom Protista no longer exists. Its members have been reassigned to various classes of unicellular eukaryotes. Nonetheless, textbooks of microbiology still describe the protozoans, just as though this name continued to occupy a legitimate place among terrestrial organisms. In the meantime, therapeutic opportunities for eradicating so-called protozoal infections, using class-targeted agents, have no doubt been missed (1). You might think that the creation of a class of living organisms, with no established scientific relation to the real world, was a rare and ancient event in the annals of biology, having little or no chance of being repeated. Not so. A special pseudoclass of fungi, deuteromyctetes (spelled with a lowercase "d", signifying its questionable validity as a true biologic class) has been created to hold fungi of indeterminate speciation. At present, there are several thousand such fungi, sitting in a taxonomic limbo, waiting to be placed into a definitive taxonomic class (2), (1).

[1] Berman JJ. Taxonomic Guide to Infectious Diseases: Understanding the Biologic Classes of Pathogenic Organisms. Academic Press, Waltham, 2012.

[2] Guarro J, Gene J, Stchigel AM. Developments in fungal taxonomy. Clinical Microbiology Reviews 12:454-500, 1999.

- Jules Berman (copyrighted material)

key words: classifications, ontology, classes, taxonomy, jules j berman

Friday, January 22, 2016

Signs of an Overly Complex Ontology

When modeling a complex system, you should always strive to design a model that is as simple as possible. There are a number of signs that tell the ontologist that her classification is just too complex.

1. Nobody, even the designers, fully understands the ontology model.

2. You realize that the ontology makes no sense. The solutions obtained by data analysts are impossible, or they contradict observations. Tinkering with the ontology doesn’t help matters.

3. For a given problem, no two data analysts seem able to formulate the query the same way, and no two query results are ever equivalent.

4. The ontology lacks modularity. It is impossible to remove a set of classes within the ontology without collapsing its structure. When anything goes wrong, the entire ontology must be fixed or redesigned, from scratch.

5. The ontology cannot fit under a higher level ontology or over a lower-level ontology.

6. The ontology cannot be debugged when errors are detected.

7. Errors occur without anyone knowing that the error has occurred.

8. You realize, to your horror, that your ontology has violated the cardinal rule of data simplification, by increasing the complexity of your data.

The practical ontologist may need to settle for a simplified approximation of the truth.

-Jules Berman (copyrighted material)

key words: ontology, classification, data simplification, jules j berman

Thursday, January 21, 2016

Bootstrapping Ontologies

Bootstrapping is the act of self-creation, from nothing. The term derives from the ludicrous stunt of pulling oneself up by one’s own bootstraps. Its shortened form, "booting," refers to the startup process in computers in which the operating system is somehow activated via its operating system, that has not been activated. The absurd and somewhat surrealistic quality of bootstrapping protocols serves as one of the most mysterious and fascinating areas of science. As it happens, bootstrapping processes lie at the heart of some of the most powerful techniques in data simplification.

It is worth taking the time to explore the philosophical and the pragmatic aspects of bootstrapping. Starting from the beginning, how was the universe created? For believers, the universe was created by an all-powerful deity. If this were so, then how was the all-powerful deity created? Was the deity self-created, or did the deity simply bypass the act of creation altogether? The answers to these questions are left as an exercise for the reader, but we can all agree that there had to be some kind of bootstrapping process, if something was created from nothing. Otherwise, there would be no universe, and this essay would be much shorter than it is.

Getting back to our computers, how is it possible for any computer to boot its operating system, when we know that the process of managing the startup process is one of the most important functions of the fully operational operating system? Basically, at startup, the operating system is nonfunctional. A few primitive instructions hardwired into the computer’s processors are sufficient to call forth a somewhat more complex process from memory, and this newly activated process calls forth other processes, until the operating system is eventually up and running. The cascading rebirth of active processes takes time, and explains why booting your computer may seem to be a ridiculously slow process.

What is the relationship between bootstrapping and classification? The ontologist creates a classification based on a worldview in which objects hold specific relationships with other objects. Hence, the ontologist’s perception of the world is based on preexisting knowledge of the classification of things; which presupposes that the classification already exists.

Paradoxically, you cannot build a classification without first having the classification. How does an ontologist bootstrap a classification into existence? She may begin with a small assumption that seems, to the best of her knowledge, unassailable. In the case of the classification of living organisms, she may assume that the first organisms were primitive, consisting of a few self-replicating molecules and some physiologic actions, confined to a small space, capable of a self-sustaining system. Primitive viruses and prokaryotes (ie, bacteria) may have started the ball rolling. This first assumption might lead to observations and deductions, which eventually yield the classification of living organisms that we know today.

Every thoughtful ontologist will admit that a classification is, at its best, a hypothesis-generating machine; not a factual representation of reality. We use the classification to create new hypotheses about the world and about the classification itself. The process of testing hypotheses may reveal that the classification is flawed; that our early assumptions were incorrect. More often, by testing hypotheses, we reassure ourselves that our assumptions were consistent with new observations, adding to our understanding of the relations between the classes and the instances within the classification.

- Jules Berman (copyrighted material)

key words: bootstrap, bootstrapping, classification, ontology, ontologist, data simplification, informatics, jules j berman

Wednesday, January 13, 2016

CLASSIFICATIONS, THE SIMPLEST OF ONTOLOGIES

Here is a short excerpt from my book Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information


CLASSIFICATIONS, THE SIMPLEST OF ONTOLOGIES

The human brain is constantly processing visual and other sensory information collected from the environment. When we walk down the street, we see images of concrete and asphalt and millions of blades of grass, and birds, and dogs, and other persons, and so on. Every step we take conveys a new world of sensory input. How can we process it all? The mathematician and philosopher Karl Pearson has (1857 - 1936) has likened the human mind to a "sorting machine". We take a stream of sensory information and sort it into objects; we then we collect the individual objects into general classes. The green stuff on the ground is classified as "grass," and the grass is subclassified under some larger grouping, such as "plants." A flat stretch of asphalt and concrete may be classified as a "road" and the road might be subclassified under "man-made constructions". If we lacked a culturally-determined classification of objects for our world, we would be overwhelmed by sensory input, and we would have no way to remember what we see, and no way to draw general inferences about anything. Simply put, without our ability to classify, we would not be human.

Every culture has some particular way to impose a uniform way of perceiving the environment. In English-speaking cultures, the term "hat" denotes a universally recognized object. Hats may be composed of many different types of materials, and they may vary greatly in size, weight, and shape. Nonetheless, we can almost always identify a hat when we see one, and we can distinguish a hat from all other types of objects. An object is not classified as a hat simply because it shares a few structural similarities with other hats. A hat is classified as a hat because it has a class relationship; all hats are items of clothing that fit over the head. Likewise, all biological classifications are built by relationships, not by similarities.

Aristotle was one of the first experts in classification. His greatest insight came when he correctly identified a dolphin as a mammal. Through observation, he knew that a large group of animals was distinguished by a gestational period in which a developing embryo is nourished by a placenta, and the offspring are delivered into the world as formed, but small versions of the adult animals (i.e., not as eggs or larvae), and the newborn animals feed from milk excreted from nipples, overlying specialized glandular organs (mammae). Aristotle knew that these features, characteristic of mammals, were absent in all other types of animals. He also knew that dolphins had all these features; fish did not. He correctly reasoned that dolphins were a type of mammal, not a type of fish. Aristotle was ridiculed by his contemporaries for whom it was obvious that dolphins were a type of fish. Unlike Aristotle, they based their classification on similarities, not on relationships. They saw that dolphins looked like fish and dolphins swam in the ocean like fish, and this was all the proof they needed to conclude that dolphins were indeed fish. For about two thousand years following the death of Aristotle, biologists persisted in their belief that dolphins were a type of fish. For the past several hundred years, biologists have acknowledged that Aristotle was correct after all; dolphins are mammals. Aristotle discovered and taught the most important principle of classification; that classes are built on relationships among class members; not by counting similarities. We will see in later chapters, that methods of grouping data objects by similarity can be very misleading, and should not be used as the basis for constructing a classification or an ontology.

A classification is a very simple form of ontology, in which each class is limited to one parent class. To build a classification, the ontologist must do the following: 1) define classes (i.e., find the properties that define a class and extend to the subclasses of the class); 2) assign instances to classes; 3) position classes within the hierarchy; and 4) test and validate all the above.

The constructed classification becomes a hierarchy of data objects conforming to a set of principles:

1. The classes (groups with members) of the hierarchy have a set of properties or rules that extend to every member of the class and to all of the subclasses of the class, to the exclusion of unrelated classes . A subclass is itself a type of class wherein the members have the defining class properties of the parent class plus some additional property(ies) specific for the subclass.

2. In a hierarchical classification, each subclass may have no more than one parent class. The root (top) class has no parent class. The biological classification of living organisms is a hierarchical classification.

3. At the bottom of the hierarchy is the class instance. For example, your copy of this book is an instance of the class of objects known as "books".

4. Every instance belongs to exactly one class.

5. Instances and classes do not change their positions in the classification. As examples, a horse never transforms into a sheep, and a book never transforms into a harpsichord.

6. The members of classes may be highly similar to one another, but their similarities result from their membership in the same class (i.e., conforming to class properties), and not the other way around (i.e., similarity alone cannot define class inclusion).

Classifications are always simple; the parental classes of any instance of the classification can be traced as a simple, non-branched list, ascending through the class hierarchy. As an example, here is the lineage for the domestic horse (Equus caballus), from the classification of living organisms:

Equus caballus
Equus subg. Equus
Equus
Equidae
Perissodactyla
Laurasiatheria
Eutheria
Theria
Mammalia
Amniota
Tetrapoda
Sarcopterygii
Euteleostomi
Teleostomi
Gnathostomata
Vertebrata
Craniata
Chordata
Deuterostomia
Coelomata
Bilateria
Eumetazoa
Metazoa
Fungi/Metazoa group
Eukaryota
cellular organisms
The words in this zoologic lineage may seem strange to laypersons, but taxonomists who view this lineage instantly grasp the place of domestic horses in the classification of all living organisms.

A classification is a list of every member class, along with their relationships to other classes. Because each class can have only one parent class, a complete classification can be provided when we list all the classes, adding the name of the parent class for each class on the list. For example, a few lines of the classification of living organisms might be:
Craniata, subclass of Chordata
Chordata, subclass of Duterostomia
Deuterostomia, subclass of Coelomata
Coelomata, subclass of Bilateria
Bilateria, sublcass of Eumetazoa
Given the name of any class, a programmer can compute (with a few lines of code), the complete ancestral lineage for the class, by iteratively finding the parent class assigned to each ascending class.

A taxonomy is a classification with the instances "filled in." This means that for each class in a taxonomy, all the known instances (i.e., member objects) are explicitly listed. For the taxonomy of living organisms, the instances are named species. Currently, there are several million named species of living organisms, and each of these several million species is listed under the name of some class included in the full classification.

Classifications drive down the complexity of their data domain, because every instance in the domain is assigned to a single class, and every class is related to the other classes through a simple hierarchy.

It is important to distinguish a classification system from an identification system. An identification system puts a data object into its correct slot within the classification. For example, a fingerprint matching system may look for a set of features that puts a fingerprint into a special subclass of all fingerprint, but the primary goal of fingerprint matching is to establish the identity of an instance (i.e., to show that two sets of fingerprints belong to the same person). In the realm of medicine, when a doctor renders a diagnosis on a patient's diseases, she is not classifying the disease; she is finding the correct slot within the pre-existing classification of diseases that holds her patient's diagnosis.

-Jules Berman (Copyrighted material)

key words: classification, ontology, taxonomy, jules j berman

Friday, April 18, 2008

MeSH (Medical Subject Headings) more complex than simple "Trees"

MeSH (Medical Subject Headings) is a wonderful nomenclature of medical terms available from the U.S. National Library of Medicine.

The download site is:

http://www.nlm.nih.gov/mesh/filelist.html


MeSH is one of the greatest gifts provided by the U.S. National Library of Medicine and can be used freely for a variety of projects involving indexing, tagging, searching, retrieving, coding, analyzing, merging, and sharing biomedical text. In my opinion, there are many projects that rely on commercial and legally encumbered nomenclatures that would be better served by MeSH.

My only quibble with MeSH is that it is incorrectly described as a Tree structure.

Here is the official word (from the NLM website) on MeSH Trees from: http://www.nlm.nih.gov/mesh/intro_trees2007.html

"Because of the branching structure of the hierarchies, these lists are sometimes referred to as "trees". Each MeSH descriptor appears in at least one place in the trees, and may appear in as many additional places as may be appropriate. Those who index articles or catalog books are instructed to find and use the most specific MeSH descriptor that is available to represent each indexable concept."

When you look at individual entries in MeSH, you find that a single entry may be assigned multiple MeSH numbers.

For example, the MeSH term, "Family" is assigned two MeSH numbers,

MN = F01.829.263
MN = I01.880.225

The parent "number" for any MeSH entry is found by removing the last set of decimal demarcated digits.

For example:
F01.829.263 MeSH name, Family
F01.829 MeSH name, Plychology, Social
F01 MeSH Name, Behavior and Behavior Mechanisms

For each MeSH number, there is a separate hierarchy.

It is tempting to think of each hierarchy for each number as a tree (then MeSH could be envisioned as a dense forest), but each parent term could be assigned multiple MeSH numbers, each producing a multi-branching hierarchy.

Because each MeSH term (including the ancestral terms for a MeSh term) may be assigned multiple MeSH numbers, each with its own hierarchy, the MeSH data structure is more accurately thought of as a complex ontology, with terms existing in multiple classes, with specified relationships among any class and its parent classes.

The tree metaphor breaks down because branches and nodes within a branch can be connected to other branches and to other nodes. Trees do not do this kind of thing.

It is possible to write a script that parses through every MeSH entry, finds all of the MeSH numbers for the entry, determines the parent terms for the MeSH numbers, determines all of the alternate MeSH numbers for the parent terms, then finds all of the grandparent terms for all of the parent terms, etc., until all of the hierarchical terms for the term are found.

Here is the Perl script. This Perl script is provided "as is", by its creator, Jules J. Berman, without warranty of any kind, express or implied, including but not limited to the warranties of merchantability, fitness for a particular purpose and noninfringement. in no event shall the authors or copyright holders be liable for any claim, damages or other liability, whether in an action of contract, tort or otherwise, arising from, out of or in connection with the software or the use or other dealings in the software.

#!/usr/local/bin/perl
open(MESH, "D2007.BIN"); #the file name for the raw ascii MeSH version
open(OUT, ">mesh.out");
$/ = "\n\n\*NEWRECORD\n";
$line = " ";
@cumlist;
%numberhash;
%namehash;
while ($line ne "")
{
my $numbers = "";
$line = <MESH>;
$line =~ /\nMH = ([^\n]+)\n/;
$name = $1;
while ($line =~ m/\nMN ?= ?([^\n]+)(?=\n)/mg)
{
$number = $1;
$number =~ s/^ *//o;
$number =~ s/ *$//o;
$number =~ s/ +/ /;
$numberhash{$number} = $name;
$numbers = $numbers . " " . $number;
}
$numbers =~ s/^ *//o;
$numbers =~ s/ *$//o;
$numbers =~ s/ +/ /o;
$namehash{$name} = $numbers;
}
close(MESH);
while((my $key, my $value) = each (%namehash))
{
@cumlist = ("");
print OUT "\nTERM LINEAGE FOR " . uc($key) . "\n";
my @valuelist = split(/ /,$value);
@cumlist = (@cumlist, @valuelist);
&splitlist(@cumlist);
for(1..30)
{
@cumlist = grep { $marked{$_}++; $marked{$_} == 1; }
@cumlist;
undef(%marked);
&allmeshnums(@cumlist);
@cumlist = grep { $marked{$_}++; $marked{$_} == 1; }
@cumlist;
undef(%marked);
&splitlist(@cumlist);
}
@cumlist = grep { $marked{$_}++; $marked{$_} == 1; }
@cumlist;
undef(%marked);
foreach my $thing (@cumlist)
{
print OUT "$thing $numberhash{$thing}\n";
}
}
sub splitlist()
{
@valuelist = @_;
foreach my $meshno (@valuelist)
{
for(1..30)
{
if ($meshno =~ /\.[0-9]+$/)
{
$meshno = $`;
push(@cumlist, $meshno);
}
else
{
last;
}
}
}
}
sub allmeshnums()
{
@meshnumber = @_;
foreach my $thing (@meshnumber)
{
my $name = $numberhash{$thing};
my $value = $namehash{$name};
my @valuelist = split(/ /,$value);
@cumlist = (@valuelist, @cumlist);
}
}
exit;

The output file, mesh.out is over 9 megabytes in length.

Here is an example of one entry, in the output file, mesh.out

TERM LINEAGE FOR GIANT CELLS, FOREIGN-BODY
A11.118.637 Leukocytes
A15.145.229.637 Leukocytes
A15.382.490 Leukocytes
A11.118.637.555 Leukocytes, Mononuclear
A15.145.229.637.555 Leukocytes, Mononuclear
A15.382.490.555 Leukocytes, Mononuclear
A15.378 Hematopoietic System
A11.148 Bone Marrow Cells
A15.378.316 Bone Marrow Cells
A12.207.152 Blood
A15.145 Blood
A11.118 Blood Cells
A15.145.229 Blood Cells
A11.329.372.376 Giant Cells, Foreign-Body
A11.502.376 Giant Cells, Foreign-Body
A11.627.624.480.376 Giant Cells, Foreign-Body
A11.733.397.376 Giant Cells, Foreign-Body
A15.382.680.397.376 Giant Cells, Foreign-Body
A15.382.812.522.376 Giant Cells, Foreign-Body
A11.329 Connective Tissue Cells
A11 Cells
A11.502 Giant Cells
A11.118.637.555.652 Monocytes
A11.148.580 Monocytes
A11.627.624 Monocytes
A11.733.547 Monocytes
A15.145.229.637.555.652 Monocytes
A15.378.316.580 Monocytes
A15.382.490.555.652 Monocytes
A15.382.680.547 Monocytes
A15.382.812.547 Monocytes
A11.627 Myeloid Cells
A11.733 Phagocytes
A15.382.680 Phagocytes
A15.382 Immune System
A15 Hemic and Immune Systems
A11.329.372 Macrophages
A11.627.624.480 Macrophages
A11.733.397 Macrophages
A15.382.680.397 Macrophages
A15.382.812.522 Macrophages
A15.382.812 Reticuloendothelial System
A12.207 Body Fluids
A12 Fluids and Secretions

When we examine the multi-lineage ancestry of "Foreign body giant cells" we see that MeSH is not a tree hierarchy. This means that the MeSH data structure is highly complex and requires some computational know-how to fully explore all the term relationships.

Jules Berman

tags: Perl programming for medicine and biology, nomenclature, thesaurus, nlm, medical subject headings, open source, medical indexing, medical data retrieval, medical informatics, biomedical informatics, national library of medicine
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Friday, March 14, 2008

Creating a directed graph from an RDF schema

In yesterday's blog , I announced the newest version of the Neoplasm Classification. The Neoplasm Classification is an open source document available as a plain-text flat-file, as an XML file, and as and RDF document. The top of the RDF document contains the complete RDF Schema for the Classification. The remainder of the file (>99% of the file) is devoted to the entries for the individual neoplasm terms (over 135,000 of them).

I have a small schematic, that represents the organization of the Neoplasm Classification.



In addition to this small schematic, I have a large schematic that represents the complete hierarchy of the Classification. Though it is too large to fit in this blog, each part of the schematic is quite simple.

It took a few seconds to create the complete diagram of the classification, in digraph (directed graph) form, using a short Perl script that parsed the Classification's RDF Schema + one command-line instruction to the free and open source GraphViz application.

Here's the script. As per usual, the following statement applies. The software is provided "as is", without warranty of any kind, express or implied, including but not limited to the warranties of merchantability, fitness for a particular purpose and noninfringement. in no event shall the authors or copyright holders be liable for any claim, damages or other liability, whether in an action of contract, tort or otherwise, arising from, out of or in connection with the software or the use or
other dealings in the software.

#!/usr/bin/perl
open (TEXT, "schema.txt");
open (OUT, ">schema.dot");
$/ = "\<\/rdfs\:Class>";
print OUT "digraph G \{\n";
print OUT "size\=\"15\,15\"\;\n";
print OUT "ranksep\=\"2\.00\"\;\n";
$line = " ";
while ($line ne "")
{
$line = <TEXT>;
last if ($line !~ /\<rdfs\:/);
if ($line =~
/\:resource\=\"[a-z0-9\:\/\_\.\-]*\#([a-z\_]+)\"/i)
{
$father = $1;
}
if ($line =~ /rdf\:ID\=\"([a-z\_]+)\"/i)
{
$child = $1;
}
print OUT "$father \-\> $child\;\n";
print "$father \-\> $child\;\n";
}
print OUT "\}";
exit;

If you work with RDF (and every biomedical professional should understand how RDF is used to specify data), you will want a method that can instantaneously render a schematic of your RDF Schema (ontology) or of any descendant section of your Schema.

Tomorrow, I'll go into some detail to describe just how the Perl script produces a GraphViz script that can render an RDF Schema as a visual digraph.

- Jules Berman


My book, Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information was published in 2013 year by Morgan Kaufmann.



I urge you to read more about my book. Google books has prepared a generous preview of the book contents. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.
tags: big data, metadata, data preparation, data analytics, data repurposing, datamining, data mining, medical autocoding, medical data scrubbing, medical data scrubber, medical record scrubbing, medical record scrubber, medical text parsing, medical autocoder, nomenclature, terminology, ontology, VizGraph, Neoplasm Classification, semantic web, ontologies, digraph, directed graph, tumor, cancer, tumour, neoplasia

Sunday, October 28, 2007

Developmental Classification of Neoplasms now an RDF Ontology

I am publishing today the first ontology version of the Developmental Lineage Classification and Taxonomy of Neoplasms. It is available for download in several file versions.

The full ontology is a 10 Megabyte RDF file. Note that the file is so large that some browsers may not be able to open the entire file. On my computer, I had no trouble opening the file in my Internet Explorer browser, but the file was too large for my Mozilla browser.
http://www.julesberman.info/neordf.xml


The file was validated using the w3c validator service at http://www.w3.org/rdf/validator/, with a caveat. The full ontology file (10+ Mbytes) was too large for the validator, so I truncated the ontology, validated the truncated file (that contained all of the classes, subclasses, properties), and left out the repetitive list of terms. Then I took the entire file and validated it with an XML parser to verify that the file was well-formed. That really covers everything (RDF logic and XML structure).

The gzipped version of the RDF file (under 1 Megabyte).
http://www.julesberman.info/neorxml.gz


The flat file version, listing each term followed by its lineage (gzipped file).
http://www.julesberman.info/neoself.gz


The plain old XML version, with no RDF semantics (gzipped file). http://www.julesberman.info/neoclxml.gz

The ontology contains several parts:

1. The neoplasm classification proper (as illustrated in the schematic)



2. A listing of cancer terms that will probably never be entered into the proper classification (more about this later)

3. A listing of hyperplasias or hamartomas, some of which will be entered into the proper classification and others of which will remain in class Hyperplasia

4. A listing of precancer terms

5. A listing of syndromes associated with increased risk for cancer.

In this version, there are 5841 classified types of neoplasms and 130,503 terms representing the 5,841 types of neoplasms.

This represents the largest nomenclature of neoplasms in existence and, with today's publication, the largest formal ontology (in RDF syntax) of neoplasm names.

Over the next few weeks, I'll post additional blogs to further explain the RDF ontology files.

- Jules Berman
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Thursday, October 25, 2007

New schema for the Neoplasm Classification

The Developmental Lineage Classification and Taxonomy of Neoplasms first came out in 2003, and I've been making revisions and updates since then. Most of the work has involved adding new names of neoplasms, but this month I've made a change to the basic organization of the classification.

The schema is summarized here:



There are now six major classes, under the root class, Neoplasm


Endoderm/Ectoderm
Mesoderm
Neuroectoderm
Neural Crest
Germ cell
Trophectoderm


Every neoplasm falls under one of the six major classes.

The rationale of the classification is that tumors inherit key cellular pathways through their developmental lineages. This assertion is supported by decades of morphologic evaluations of tumors. More recently, molecular biological observations have shown that genetic markers and pathways are carried through cell lineage. Tumors grouped by cell lineage may share responses to new chemotherapeutic and chemopreventive agents targeted to specific pathways. If this is true, we can start to develop agents (and combinations of agents) that are effective against groups of neoplasms that share a common developmental lineage.

This theme has been developed in several of my early papers. The most popular paper, which has had many thousands of downloads, is:

Tumor classification: molecular analysis meets Aristotle

The complete classified taxonomy is available as a gzipped XML file at:

http://www.julesberman.info/neoclxml.gz

The taxonomy contains the names of over 5,000 different neoplasms, and about 130,000 synonymous terms. It is the most comprehensive listing of neoplasms in the world, and it is distributed under the GNU Free Documentation License.

-Jules J. Berman

In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

tags: common disease, orphan disease, orphan drugs, rare disease, subsets of disease, disease genetics, genetics of complex disease, genetics of common diseases, cryptic disease

Friday, October 19, 2007

Specifying a classification with GraphViz

I made this graph with GraphViz (a freely available software application).

In my prior post, I used this same figure, which displays the entire Neoplasm Classification Schema, in a graphic form.

Here's how to use GraphViz to display a classification or ontology.

The GraphViz download site is:

http://www.graphviz.org/Download.php

Windows users can download graphviz-2.14.1.exe (5,614,329 bytes).

You can install the software by running the .exe file.

GraphViz has its own language, in which you list the relationships among the different classes, and then it builds a graphic view of the classification.

GraphViz has sub-applications:dot, fdp, twopi, neato, and circo. The twopi application, which I used, creates graphs that have a radial layout.



digraph G {
size="10,16";
ranksep="1.75";
node [style=filled color=gray65];
Neoplasm [label="Neoplasm"];
node [style=filled color=lightgray];
EndodermEctoderm
[label="Endoderm\/\nEctoderm"];
NeuralCrest [label="Neural Crest"];
GermCell [label="Germ cell"];
Neoplasm -> EndodermEctoderm;
Neoplasm -> Mesoderm;
Neoplasm -> GermCell;
Neoplasm -> Trophectoderm;
Neoplasm -> Neuroectoderm;
Neoplasm -> NeuralCrest;
node [style=filled color=gray95];
Trophectoderm -> Molar;
Trophectoderm -> Trophoblast;
EndodermEctoderm -> Odontogenic;
EndodermEctodermPrimitive
[label="Endoderm\/Ectoderm\nPrimitive"];
EndodermEctoderm -> EndodermEctodermPrimitive;
Endocrine
[label="Endoderm/Ectoderm\nEndocrine"];
EndodermEctoderm -> Endocrine;
EndodermEctoderm -> Parenchymal;
Odontogenic
[label="Endoderm/Ectoderm\nOdontogenic"];
EndodermEctoderm -> Surface;
MesodermPrimitive
[label="Mesoderm\nPrimitive"];
Mesoderm -> MesodermPrimitive;
Mesoderm -> Subcoelomic;
Mesoderm -> Coelomic;
NeuroectodermPrimitive
[label="Neuroectoderm\nPrimitive"];
NeuroectodermNeuralTube
[label="Central Nervous\nSystem"];
Neuroectoderm -> NeuroectodermPrimitive;
Neuroectoderm -> NeuroectodermNeuralTube;
NeuralCrestMelanocytic
[label="Melanocytic"];
NeuralCrestPrimitive
[label="Neural Crest\nPrimitive"];
NeuralCrestEndocrine
[label="Neural Crest\nEndocrine"];
PeripheralNervousSystem
[label="Peripheral\nNervous System"];
NeuralCrestOdontogenic
[label="Neural Crest\nOdontogenic"];
NeuralCrest -> NeuralCrestPrimitive;
NeuralCrest -> PeripheralNervousSystem;
NeuralCrest -> NeuralCrestEndocrine;
NeuralCrest -> NeuralCrestMelanocytic;
NeuralCrest -> NeuralCrestOdontogenic;
GermCell -> Differentiated;
GermCell -> Primordial;
}

You create your the image neo2.png, from the neo2.dot specification (above) by invoking the twopi subapplication on a command line.

c:\ftp>twopi -Tpng neo2.dot -o neo2.png

-Jules Berman tags: classification, directed graph, graph, ontology,


Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.


Sunday, September 16, 2007

Latest update of Neoplasm Classification available

The latest version of the Developmental Lineage Classification and Taxonomy of Neoplasms is now available as a gzipped file at:

http://www.julesberman.info/neoclxml.gz

This Neoplasm Classification has been described at:

http://www.biomedcentral.com/1471-2407/4/10

It contains 5,827 neoplasm classified concepts and 130,283 different terms (codes beginning with "C"). It is more than ten times larger than any other neoplasm classification.

In addition to specific neoplasm concepts, it also contains general neoplastic terms (coded as C0000000), inherited conditions associated with neoplasms (codes beginning with "S") and terms related to the stage or anatomic location of neoplasms (codes beginning with "ST").


In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D.