Showing posts with label informatics. Show all posts
Showing posts with label informatics. Show all posts

Monday, January 25, 2016

A Data Science Definition of Semantics

Semantics is the study of meaning (Greek root, semantikos, signficant meaning). In the context of data science, semantics is the technique of creating meaningful assertions about data objects. A meaningful assertion, as used here, is a triple consisting of an identified data object, a data value, and a descriptor for the data value. In practical terms, semantics involves making assertions about data objects (ie, making triples), combining assertions about data objects (ie, merging triples), and assigning data objects to classes; hence relating triples to other triples. As a word of warning, few informaticians would define semantics in these terms, but most definitions for semantics are functionally equivalent to the definition offered here.

Most language is unstructured and meaningless. Consider the assertion: Sam is tired. This is an adequately structured sentence with a subject verb and object. But what is the meaning of the sentence? There are a lot of people named Sam. Which Sam is being referred to in this sentence? What does it mean to say that Sam is tired? Is "tiredness" a constitutive property of Sam, or does it only apply to specific moments in Sam's life? If so, for what moment in time is the assertion, "Sam is tired" actually true? To a computer, meaning comes from assertions that have a specific, identified subject associated with some sensible piece of fully described data (metadata coupled with the data it describes).

- Jules Berman (copyrighted material)

key words: meaning, data science, informatics, triple, semantics, jules j berman

Thursday, January 21, 2016

Bootstrapping Ontologies

Bootstrapping is the act of self-creation, from nothing. The term derives from the ludicrous stunt of pulling oneself up by one’s own bootstraps. Its shortened form, "booting," refers to the startup process in computers in which the operating system is somehow activated via its operating system, that has not been activated. The absurd and somewhat surrealistic quality of bootstrapping protocols serves as one of the most mysterious and fascinating areas of science. As it happens, bootstrapping processes lie at the heart of some of the most powerful techniques in data simplification.

It is worth taking the time to explore the philosophical and the pragmatic aspects of bootstrapping. Starting from the beginning, how was the universe created? For believers, the universe was created by an all-powerful deity. If this were so, then how was the all-powerful deity created? Was the deity self-created, or did the deity simply bypass the act of creation altogether? The answers to these questions are left as an exercise for the reader, but we can all agree that there had to be some kind of bootstrapping process, if something was created from nothing. Otherwise, there would be no universe, and this essay would be much shorter than it is.

Getting back to our computers, how is it possible for any computer to boot its operating system, when we know that the process of managing the startup process is one of the most important functions of the fully operational operating system? Basically, at startup, the operating system is nonfunctional. A few primitive instructions hardwired into the computer’s processors are sufficient to call forth a somewhat more complex process from memory, and this newly activated process calls forth other processes, until the operating system is eventually up and running. The cascading rebirth of active processes takes time, and explains why booting your computer may seem to be a ridiculously slow process.

What is the relationship between bootstrapping and classification? The ontologist creates a classification based on a worldview in which objects hold specific relationships with other objects. Hence, the ontologist’s perception of the world is based on preexisting knowledge of the classification of things; which presupposes that the classification already exists.

Paradoxically, you cannot build a classification without first having the classification. How does an ontologist bootstrap a classification into existence? She may begin with a small assumption that seems, to the best of her knowledge, unassailable. In the case of the classification of living organisms, she may assume that the first organisms were primitive, consisting of a few self-replicating molecules and some physiologic actions, confined to a small space, capable of a self-sustaining system. Primitive viruses and prokaryotes (ie, bacteria) may have started the ball rolling. This first assumption might lead to observations and deductions, which eventually yield the classification of living organisms that we know today.

Every thoughtful ontologist will admit that a classification is, at its best, a hypothesis-generating machine; not a factual representation of reality. We use the classification to create new hypotheses about the world and about the classification itself. The process of testing hypotheses may reveal that the classification is flawed; that our early assumptions were incorrect. More often, by testing hypotheses, we reassure ourselves that our assumptions were consistent with new observations, adding to our understanding of the relations between the classes and the instances within the classification.

- Jules Berman (copyrighted material)

key words: bootstrap, bootstrapping, classification, ontology, ontologist, data simplification, informatics, jules j berman

Wednesday, January 20, 2016

TRIPLES: THE BASIC UNIT OF MEANING IN DATA SCIENCE

Data, by itself, has no meaning. It is the job of the data scientist to assign meaning to data, and this is done with data objects, triples, and classifications. Our most familiar data constructions (eg, spreadsheets, relational databases, flat-file records) convey meaning through triples and data objects; we just don’t perceive them as such.

The three conditions for a meaningful assertion are:
1. There is a specific data object about which the statement is made.
2. There is data that pertains to the specified object.
3. There is metadata that describes the data
Simply put, assertions have meaning whenever a pair of metadata and data (the descriptor for the data and the data itself) is assigned to a specific object. In the informatics field, assertions come in the form of so-called triples, consisting of the object, then the metadata, and then the data.

Here are some examples of triples, as they might occur in a medical dataset:
"Jules Berman" "blood glucose level" "85"
"Mary Smith" "blood glucose level" "90"
"Samuel Rice" "blood glucose level" "200"
"Jules Berman" "eye color" "brown"
"Mary Smith" "eye color" "blue"
"Samuel Rice" "eye color" "green"
Here are a few triples, as the might occur in a haberdasher’s dataset
"Juan Valdez" "hat size" "8"
"Jules Berman" "hat size" "9"
"Homer Simpson" "hat size" "9"
"Homer Simpson" "hat_type" "bowler"
We can combine the triples from a medical dataset and a habderdasher’s data set that apply to a common object:
"Jules Berman" "blood glucose level" "85"
"Jules Berman" "eye color" "brown"
"Jules Berman" "hat size" "9"
Triples can port their meaning between different databases because they bind described data to an object. The portability of triples permits us to achieve data integration of heterogeneous data, and facilitates the design of software agents. Data integration involves merging related data objects, across diverse data sets. As it happens, if data supports introspection and data is organized as meaningful assertions (ie, as identified triples), then data integration is implicit (ie, an intrinsic property of the data). In essence, data integration is awarded to data scientists who apply data simplification techniques.

- Jules Berman (copyrighted material)

key words: triple, meaning, informatics, computer science, data integration, heterogeneous data, jules j berman

Tuesday, January 12, 2016

REALLY BAD IDENTIFIER METHODS

Here is a short excerpt from my book Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information


"I always wanted to be somebody, but now I realize I should have been more specific." -Lily Tomlin

Names are poor identifiers. Aside from the obvious fact that they are not unique (e.g., surnames such as Smith, Zhang, Garcia, Lo, and given names such as John and Susan), a single name can have many different representations. The sources for these variations are many. Here is a partial listing:

1. Modifiers to the surname (du Bois, DuBois, Du Bois, Dubois, Laplace, La Place, van de Wilde, Van DeWilde, etc.).

3. Accents that may or may not be transcribed onto records (e.g., acute accent, cedilla, diacritical comma, palatalized mark, hyphen, diphthong, umlaut, circumflex, and a host of obscure markings).

4. Special typographic characters (the combined "ae").

5. Multiple "middle names" for an individual, that may not be transcribed onto records. Individuals who replace their first name with their middle name for common usage, while retaining the first name for legal documents.

6. Latinized and other versions of a single name (Carl Linnaeus, Carl von Linne, Carolus Linnaeus, Carolus a Linne).

7. Hyphenated names that are confused with first and middle names (e.g., Jean-Jacques Rousseau, or Jean Jacques Rousseau; Louis-Victor-Pierre-Raymond, 7th duc de Broglie, or Louis Victor Pierre Raymond Seventh duc deBroglie).

8. Cultural variations in name order that are mistakenly re-arranged when transcribed onto records. Many cultures do not adhere to the Western European name order (e.g., Given name, middle name, surname).

9. Name changes; through legal action, aliasing, pseudonymous posing, or insouciant whim.

Aside from the obvious consequences of using names as record identifiers (e.g., corrupt database records, impossible merges between data resources, impossibility of reconciling legacy record), there are non-obvious consequences that are worth considering. Take, for example, accented characters in names. These word decorations wreak havoc on orthography and on alphabetization. Where do you put a name that contains an umlauted character? Do you pretend the umlaut isn't there, and put it in alphabetic order with the plain characters? Do you order based on the ASCII-numeric assignment for the character, in which the umlauted letter may appear nowhere near the plain-lettered words in an alphabetized list. The same problem applies to every special character.

A similar problem exists for surnames with modifiers. Do you alphabetize de Broglie under "D" or under "d" or under "B"? If you choose B, then what do you do with the concatenated form of the name, "deBroglie"?

When it comes down to it, is is impossible to satisfactorily alphabetize a list of names. This means that searches based on proximity in the alphabet will always be prone to errors.

I have had numerous conversations with intelligent professionals who are tasked with the responsibility of assigning identifiers to individuals. At some point in every conversation, they will find it necessary to explain that although an individual's name cannot serve as an identifier, the combination of name plus date of birth provides accurate identification in almost every instance. They sometimes get carried away, insisting that the combination of name plus date of birth plus social security number provides perfect identification, as no two people will share all three identifiers: same name, same date of birth, same social security number. This argument, rises to the height of folly, and completely misses the point of identification. As we will see, it is relatively easy to assign unique identifiers to individuals and to any data object, for that matter. For managers of Big Data resources, the larger problem is ensuring that each unique individual has only one identifier (i.e., denying one object multiple identifiers).

Let us see what happens when we create identifiers from the name plus birthdate. We will examine name + birthdate + social security number later in this section.

Consider this example. Mary Jessica Meagher, born June 7, 1912 decided to open a separate bank account in each of 10 different banks. Some of the banks had application forms, which she filled out accurately. Other banks registered her account through a teller, who asked her a series of questions and immediately transcribed her answers directly into a computer terminal. Ms. Meagher could not see the computer screen and could not review the entries for accuracy.

Here are the entries for her name plus date of birth:

1. Marie Jessica Meagher, June 7, 1912 (the teller mistook Marie for Mary).

2. Mary J. Meagher, June 7, 1912 (the form requested a middle initial, not name).

3. Mary Jessica Magher, June 7, 1912 (the teller misspelled the surname).

4. Mary Jessica Meagher, Jan 7, 1912 (the birth month was constrained, on the form, to three letters; Jun, entered on the form, was transcribed as Jan).

5. Mary Jessica Meagher, 6/7/12 (the form provided spaces for the final two digits of the birth year. Through the miracle of bank registration, Mary, born in 1912, was re-born a century later).

6. Mary Jessica Meagher, 7/6/2012 (the form asked for day, month, year, in that order, as is common in Europe).

7. Mary Jessica Meagher, June 1, 1912 (on the form, a 7 was mistaken for a 1).

8. Mary Jessie Meagher, June 7, 1912 (Marie, as a child, was called by the informal form of her middle name, which she provided to the teller).

9. Mary Jessie Meagher, June 7, 1912 (Marie, as a child, was called by the informal form of her middle name, which she provided to the teller, and which the teller entered as the male variant of the name).

10. Marie Jesse Mahrer, 1/1/12 (an underzealous clerk combined all of the mistakes on the form and the computer transcript, and added a new orthographic variant of the surname).

For each of these ten examples, a unique individual (Mary Jessica Meagher) would be assigned a different identifier at each of 10 banks. Had Mary re-registered at one bank, ten times, the results may have been 10 different registration identifiers, for one person.

If you toss the social security number into the mix (name + birth date + social security number) the problem is compounded. The social security number for an individual is anything but unique. Few of us carry our original social security cards. Our number changes due to false memory ("You mean I've been wrong all these years?"), data entry errors ("Character trasnpositoins, I mean transpositions, are very common"), intention to deceive ("I don't want to give those people my real number), or desperation ("I don't have a number, so I'll invent one"), or impersonation ("I don't have health insurance, so I'll use my friend's social security number"). Efforts to reduce errors by requiring patients to produce their social security cards have not been entirely beneficial.

Beginning in the late 1930s, the E. H. Ferree Company, a manufacturer of wallets, promoted their product's card pocket by including a sample social security card with each wallet sold. The display card had the social security number of one of their employees. Many people found it convenient to use the card as their own social security number. Over time, the wallet display number was claimed by over 40,000 people. Today, few institutions require individuals to prove their identity by showing their original social security card. Doing so puts an unreasonable burden on the honest patient (who does not happen to carry his/her card) and provides an advantage to criminals (who can easily forge a card).

Entities that compel individuals to provide a social security number have dubious legal standing. The social security number was originally intended as a device for validating a person's standing in the social security system. More recently, the purpose of the social security number has been expanded to track taxable transactions (i.e., bank accounts, salaries). Other uses of the social security number are not protected by law. The Social Security Act (Section 208 of Title 42 U.S. Code 408) prohibits most entities from compelling anyone to divulge his/her social security number.

Considering the unreliability of social security numbers in most transactional settings, and considering the tenuous legitimacy of requiring individuals to divulge their social security numbers, a prudently designed medical identifier system will limit its reliance on these numbers. The thought of combining the social security number with name and date of birth will virtually guarantee that the identifier system will violate the strict one-to-a-customer rule.

-Jules Berman (Copyrighted material)

key words: big data, complex data, data identification, data identifiers, jules berman, jules j. berman, informatics, ehr, emr, precision data, electronic health record, electronic medical record

Thursday, March 31, 2011

Post-Informatics Pathology

For those who have been reading my blogs sequentially, I apologize for my lapse in the google ngram series. I've been preoccupied with other projects, but I hope to pick up where I left off, soon.

In the meantime, the Journal of Pathology Informatics has just published my article on "Post-Informatics Pathology." It is available at:

http://www.jpathinformatics.org/text.asp?2011/2/1/18/78499


- Jules Berman