Showing posts with label nomenclature. Show all posts
Showing posts with label nomenclature. Show all posts

Wednesday, September 10, 2014

Pitfalls in Medical Terminology

Back in 2008, I posted a list of medical terms that are easily confused, such as ileum (part of small intestine), and ilium (a pelvic bone). Medical transcriptionists and healthcare workers who input chart data (i.e., just about everybody), should be aware of medical term-pairs that have nearly the same orthography, are often pronounced identically, and have completely different meanings. These words are not picked up by spell checkers (because they are not misspelled). You can avoid such errors if you know what to look for.

Since 2008, there have been many updates to the list:
acinic, actinic
anisakiasis, anisokaryosis
aptotic, apoptotic
arboreal, aboriginal
arteritis, arthritis
auxilliary, axillary
brachial, brachium, branchial
callous, callus
causality, casualty
chlorpropamide, chlorpromazine
chondroid, chordoid
chondroma, chordoma
chorionic, chronic
cingula, singular
coitus, colitis
colic, colonic
colitis, coitus
costal, coastal
cryptogam, cryptogram
cygnet, signet
decease, disease
deceased, desist
digitate, digitize
dioecious, deciduous
diploic, diploid
disc, disk
disease, decease
diseased, deceased
dyskaryosis, dyskeratosis
dysphasia, dysphagia
ectatic, ecstatic
endochondral, enchondral (these are synonyms)
engram, n-gram, ngram
epistasis, epistaxis, epitaxis (the last is a misspelling of the second)
exxon, exon
facial, fascial
facies, faeces
fetal, fatal
firearm, forearm
foreword, forward
hallux, helicis
helicis, hallux
herpetic, herpangina
hydatid, hydatidiform
ileitis, iliitis
ileum, ilium
insular, insulin
intercostal, intercoastal
intubation, incubation
isotope, isotrope
kerasin, kerosene, keratin
keratotic, keratinic, actinic
keratinocytic, keratinolytic
keratosis, ketosis
lipoma, lymphoma
lumbar, lumber
malleolus, malleus
metachronous, metacrinus
milia, milium
miotic, mitotic, meiotic
mitosis, meiosis, myosis, myiasis
monogenic, monogenetic, and Monogenetic (last, related to class Monogenea)
mucous, mucus
myelofibrosis, myofibrosis
myofibroma, myelofibroma
neuroplastic, neoplastic
nucleus, nucleolus
oncocyte, onychocyte
oncology, ontology, ontogeny
organic, organoid
palatal, palatial
paleodontology, paleontology
palette, palate
palpation, palpitation
parasite, pericyte
parental, parenteral
pathogen, parthenogen
pathogenesis, parthenogenesis
pathogenic, pathogenetic (these two are synonyms)
penal, penile, pineal, panel
penicillamine, penicillin
perineal, peroneal, perianal
pleiotropic, pleiotrophic, pleiotypic (the first two are synonyms)
plural, pleural
porphyria, porphyruria
proptosis, ptosis
prostrate, prostate
protuberant, protruberant (the second term is simply a common misspelling)
quinine, quinidine
rachischisis, rachitis, rachischitic, rachitic
radial, radical
relics, relicts
reticle, reticule, radical
rett, ret
rosacea, rosea
semantic, somatic
serous, serious
silicon, silicone
singleton, singultus
sinusitis, synositis
somatic, semantic
sonography, stenography
taenia, tinea
takoma, trachoma
thecoma, thekeoma
torsion, distortion
trachoma, trachea
trichina, trachoma, trichura
trichinosis, trichosis, trichuriasis
trichrome, trichome
trochlear, tracheal
troglobite, troglodyte, trilobite
tuberous sclerosis, tuberculosis
tunicate, tourniquet
urethral, ureteral
vagitis, vaginitis
venous, venus
If you know the meaning of half of the terms in this list, you have a good grasp of medical terminology; but please don't settle for half measures. Physicians, nurses, chart reviewers, and medical transcriptionists should be aware of the correct meaning of each alternate word in these listed pairs.

In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D. tags: common disease, orphan disease, orphan drugs, rare disease, medical terminology, medical errors, malaprop, malapropism, definition, confusing terms, confused medical terms, medical definitions, medical transcription, nomenclature, terminology, transcription errors, transcription mistakes, EMR, EHR, electronic medical record, electronic chart, electronic health record, avoidable errors, avoidable mistakes, sources of confusion, sources of error, common mistakes, common sources of confusion

Wednesday, September 15, 2010

Naming rocks, minerals, and gems

Rocks, minerals, and gems have a rich vocabulary. There seems to be only one naming rule: no uppercase letters. This might be intended to simplify the nomenclature, but it can be confusing when you encounter a name such as "childrenite." You assume that name was inspired by a child, but the name comes from an adult; J.G. Children. The practice of using lowercase for eponyms is different from anatomic nomenclatures, wherein capitalization is preserved (e.g., Eustachian tube, not eustachian tube, after Bartolomeo Eustachio).

Similarly, you would think that "greenockite" must be a green rock (it's most often yellow, not green). It gets its name from Lord Greenock.

Likewise, biotite does not derive from a biologic precursor. It's named after Jean Biot.

Contrarily, rocks are sometimes named for lowercase (common) nouns.

Sepiolite is named for the cuttlefish bone, sepia.

Serpentine is named after the snake.

Where rocks get names from people, it's usually the surname. But not always.

Torbernite is named for Torbern Olaf Bergmann (the given name).

But don't get carried away. Bruceite was not named after someone whose given name was Bruce. It was named for a surname (Archibald Bruce). Likewise, Vivianite was not named after a woman named Vivian. It also came from the surname: J.G. Vivian

Perhaps the final solution for naming rocks after mineralogists was solved with the naming of frankhawthorneite after Professor Frank Hawthorne, of the University of Manitoba.

Here are a few more surprises:

If a gem is given a name, you'd think that it must be distinguishable from other gems with different names. No. Sapphire and ruby are the same gem, with different colorations (due to impurities in the stone). They're color-variants of corundum. Similarly, amethyst is just a color variant (purple) of common quartz.

Color in a mineral's name can be highly misleading. Glaucodot (greek for "blue"), is a gray to white mineral; never blue. Glaucodot is, however, used in the manufacture of blue glass, but you'd never know that by looking at the mineral.

Some rocks are named after the place where it was discovered or mined.

For example bytownite is named for Bytown, the former name for what is now called Ottawa.

This can be confusing, as franklinite is not named for Ben Franklin. It's named for Franklin, New Jersey. The city was named for Ben Franklin, but not the rock.

Consider Trona, California, where trona (sodium bicarbonate, and variously called tron) is mined. The mineral was not named for the city. The city was named for the mineral, which took it's name from tron, a shortened form of natron, the Arabic word for sodium.

Some rocks are named for their taste:

Calomel (probably from ancient Greek, meaning honey-taste)

or odor:

Scorodite (garlic-like, in Greek)

Some rocks are named after their included elements.

Bismuthinite contains bismuth, as you would expect. Zincote (a zinc dispersion) contains zinc, as does zincite (a true mineral).

But

Zinkenite contains no zinc (named after JKL Zinken, a German mineralogist).

Similarly, selenite, a clear crystal form of gypsum, contains no selenium.

If you're interested in the names of rocks, you must really know your Latin. Septarian concretions have complex internal structures, with multiple branches. You might think that the term comes from the latin septem (seven) referring to the number of branches. You'd be wrong. The name comes from the latin septum (partition).

It's also good to know your mythology. Pollucite was named for Pollux, the twin of Castor. The name has a certain inevitability. Pollucite is often found alongside petalite, previously known as castorite.

Some rocks are named for their geometry. There's tetrahedrite (tetrahedral crystals), triplite, clinoclase, microcline, and anorthoclase. But you can never generalize in mineralogy. Anglesite is not named for the angles in the crystal. It's named for Anglesey, Wales, where it is mined.

Sometimes, the relationship between a rock and it's name can be the opposite of what you might imagine. Fluorite is a rock that fluoresces. You might imagine that it was named because it had the property of fluorescence. Wrong. Fluorite was named for fluorine, from which it is composed (CaF2). The mineral fluorite was found to change color under UV light. The phenomenon was called fluorescence, after the first mineral shown to produce the effect. Today, every mineral that changes it's emission color under UV light is said to be fluorescent, whether it contains fluorine or not.

Sometimes you're sure a rock's name has been misspelled. Surely goethite should be geothite. Alas, no. Goethite is named for the polymath Wolfgang van Goethe.

In summary, if you're interested in the semiotics of rocks, you (unlike the rocks) will need to be flexible.

- © 2010 Jules Berman tags: nomenclature, specification, geology, rocks and minerals, hobbyists, semiotics, logophiles

Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Monday, June 14, 2010

Anatomic adjectival forms

In anatomy, most nouns have a corresponding adjective. In medicine, the adjectival forms are not always derived from the same root as the noun form.

Doctors don't usually refer to stomach flu, when they can use a term like gastric flu.

Likewise, the adjective for finger is not finger-like or fingy; it's digital. Speaking of "digital," medical software developers who work in natural language processing, need to have lists of the adjectival forms of anatomic nouns. I prepared the following computer-parsable list for my own use. I thought that others working in the field of medical informatics may have a use for these terms. If you have additional terms to add, please submit them as a comment to this posting.

As with all the documents and software I provide, the following
disclaimer holds:

The data list is provided "as is", without warranty of any kind,
express or implied, including but not limited to the warranties
of merchantability, fitness for a particular purpose and
noninfringement. in no event shall the authors or copyright
holders be liable for any claim, damages or other liability,
whether in an action of contract, tort or otherwise, arising
from, out of or in connection with the data list or the use or
other dealings in the data list.

The data list, created by Jules J. Berman, on June 5, 2010, is
donated to the Public Domain.

abdomen <=> abdominal
adenohypophysis <=> adenohypophyseal
adnexa <=> adnexal
adnexae <=> adnexal
alveolus <=> alveolar
amygdala <=> amygdaloid
anatomy <=> anatomic or anatomical
antecubitus <=> antecubital
antrum <=> antral
anus <=> anal
aorta <=> aortic or aortal
appendix <=> appendiceal
arm <=> brachial
artery <=> arterial
aryepiglottis <=> aryepiglottic
atrium <=> atrial
bladder (urinary bladder or vesica) <=> vesical
bone <=> osseous
brachium <=> brachial
bronchus <=> bronchial
caecum <=> caecal
calf <=> sural
callosum <=> callosal
cecum <=> cecal
cerebrum <=> cerebral
cervix <=> cervical
clitoris <=> clitoral
cloaca <=> cloacal
coelom <=> coelomic
colon <=> colonic
commissure <=> commissural
cranium <=> cranial
cuticle <=> cuticular
cutis <=> cutaneous
cytology <=> cytologic or cytological
decidua <=> decidual
dermis <=> dermal
diaphragm <=> diaphragmatic
digit <=> digital
dorsum <=> dorsal
duct <=> ductal
duodenum <=> duodenal
dura <=> dural
ear <=> aural
embryo <=> embryonic or embryonal
endocervix <=> endocervical
endometrium <=> endometrial or endometrioid
endothelium <=> endothelial
epicanthus <=> epicanthal or epicanthic
epicardium <=> epicardial
epidermis <=> epidermal or epidermic
epiglottis <=> epiglottal or epiglottic
epithelium <=> epithelial
esophagus <=> esophageal
ethmoid <=> ethmoidal
eye <=> ocular
face <=> facial
faeces <=> faecal
fascia <=> fascial
feces <=> fecal
fetus <=> fetal
finger or toe<=> digital
fibula <=> fibular
focus <=> focal
foetus <=> foetal
foot <=> pedal
forearm or elbow <=> cubital
front <=> frontal
gestation <=> gestational
gland <=> glandular
globe <=> global
glottis <=> glottal or glottic
gluteus <=> gluteal
gonad <=> gonadal
gyrus<=> gyral
haemorrhoid <=> haemorrhoidal
hallux (big toe) <=> hallucal
ham (back of knee) <=> popliteal
heart <=> cardiac or myocardial
hemorrhoid <=> hemorrhoidal
hepatocyte <=> hepatocellular
hernia <=> hernial
hiatus <=> hiatal
histiocyte <=> histiocytic
histology <=> histologic or histological
hypophysis <=> hypophyseal
ileum <=> ileac or ileal (intestine)
ilium <=> iliac or ilial (bone)
intestine <=> intestinal or enteric or enteral
ischium <=> ischial
jejunum <=> jejunal
kidney <=> renal or nephric or nephroid
labium <=> labial
larynx <=> laryngeal
leg <=> crural
leukocyte <=> leukocytic
liver <=> hepatic
lumen <=> luminal
lung <=> pulmonary or pulmonic
lymph <=> lymphatic or lymphoid
megakaryocyte <=> megakaryocytic
meninges <=> meningeal
metaphysis <=> metaphyseal
monocyte <=> monocytic or monocytoid
mouth or os<=> oral
mucus <=> mucous
myocardium <=> myocardial
myometrium <=> myometrial
neck <=> nuchal or cervical
node <=> nodal
nose <=> nasal
oesophagus <=> oesophageal
omentum <=> omental
orbit <=> orbital
os <=> oral
ovaries <=> ovarian
ovary <=> ovarian
pallidus <=> pallidal
pancreas <=> pancreatic
pathologic <=> pathologic or pathological
pelvis <=> pelvic
penis <=> penile
pericardium <=> pericardial
peritoneum <=> peritoneal
peroneum <=> peroneal
phalanges <=> phalangeal
phallus <=> phallic
pharynx <=> pharyngeal
pia <=> pial
pollex (thumb) <=> pollical
prostate <=> prostatic
rectum <=> rectal
retina <=> retinal
scrotum <=> scrotal
skin <=> cutaneous or dermal
sphincter <=> sphincteric
spine <=> spinal or vertebral
spleen <=> splenic or lienal or splanchnic or splenial
stomach <=> gastric
subcutis <=> subcutaneous
synovium <=> synovial
tear <=> lacrimal
testis or testicle <=> testicular
thalamus <=> thalamic
thorax <=> thoracic
throat <=> glottal
thymus <=> thymic
tongue <=> lingual or glossal or glottic
tooth or teeth <=> dental
trachea <=> tracheal
ureter <=> ureteral or ureteric
uterus <=> uterine
vagina <=> vaginal
vein <=> venous
ventricle <=> ventricular
vesicle (organelle, not to be confused with vesica or bladder) <=> vesicular
vessel <=> vascular
vulva <=> vulvar or vulval


In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D. tags: common disease, orphan disease, orphan drugs, rare disease, subsets of disease, disease genetics, genetics of complex disease, genetics of common diseases, cryptic disease, terminology pitfalls and confusing terminology, adjectives, anatomic adjectival forms, medical terminology, nomenclature, pathology, synonyms, plesionyms, public domain, listing, list

Monday, December 15, 2008

CDC Mortality Data: 6

This is the sixth in a series of posts on the CDC's (Centers for Disease Control and Prevention) public use mortality data sets.

Yesterday, we showed how to create a dictionary of ICD code/term pairs that could be used to assign disease terms to the death certificate record codes occurring in the CDC Mortality data sets. This morning I prepared a , web page that contains output data, from yesterday's blog post, that could not fit in the blog page.

I also mentioned, yesterday, that I would explain how the CDC data could be used in mashup projects. So, today, we'll begin a series of blogs that explain how mashup technology can integrate the CDC mortality data sets, and answer biomedical hypotheses.

Data mashups combine and integrate different data sources to produce a graphical representation of data that could not be achieved with any single available data source. Many people apply the term "mashup" to Web-based applications that employ two or more web services or that use two or more web-based applications that have web-accessible APIs (Application Progam Interfaces) that permit their data to be integrated into a derivative application. Because I am a biomedical information specialist, I apply "mashup" to any application that integrates available biomedical data sources, to answer questions with a graphic output (with or without Web involvement).

The classic medical mashup was done by Dr. John Snow, in London, in 1854. Wikimedia has an excellent essay on the subject. The story goes that a major outbreak of cholera occurred in late-August and early September of 1854, in the Soho district of London. By the end of the outbreak, 616 people died.

At the time, nobody understood the biological cause of cholera. At the height of the outbreak, Dr. Snow conducted a rapid, interview-based survey of the site of occurrences of new cases of cholera, producing a case-density map (hand-drawn by the doctor himself).



This map is now in the public domain. A higher-resolution version of the map is available from Wikimedia.

Examination of the map revealed that the epidemic expanded from a water source, the Broad Street pump. The pump was quickly shut. Dr. Snow's historic mashup is sometimes credited with ending the cholera epidemic and heralding a new age in scientific biomedical investigation.

To create a map mashup, we will need a data source that lists occurrences of disease and the localities in which they occur; a data source that provides the latitude and longitude of localities, and a map whose East, West, North, and South boundaries have known latitudes and longitudes. We will also need a programming language that can transform data to graphics and transfer graphics to a a map. We'll use Ruby because I like the Ruby interface to Image Magick, but Perl or Python would work equally well.

Much more importantly, we will need to have a question or hypothesis, whose solution requires a mashup. Much of computational medicine can be described as a solution in search of a question. We have many ways of analyzing data, but we often lack important questions. In the next several blogs, we will show how the CDC mortality data files can be used to test medical hypotheses. Through examples, we will introduce concepts and tools used in mashups, and we will end this series with several mashups, of increasing complexity.

If you are new to this blog, you might want to review the prior 5 blog posts, in the series, sequentially.

As I remind readers in almost every blog post, if you want to do your own creative data mining, you will need to learn a little about computer programming.

For Perl and Ruby programmers, methods and scripts for using a wide range of publicly available biomedical databases, are described in detail in my prior books:

Perl Programming for Medicine and Biology

Ruby Programming for Medicine and Biology

An overview of the many uses of biomedical information is available in my book,
Biomedical Informatics.

More information on cancer is available in my recently published book, Neoplasms: Principles of Development and Diversity.

© 2008 Jules Berman

As with all of my scripts, lists, web sites, and blog entries, the following disclaimer applies. This material is provided by its creator, Jules J. Berman, "as is", without warranty of any kind, expressed or implied, including but not limited to the warranties of merchantability, fitness for a particular purpose and noninfringement. in no event shall the author or copyright holder be liable for any claim, damages or other liability, whether in an action of contract, tort or otherwise, arising from, out of or in connection with the material or the use or other dealings in the material.

Tuesday, March 4, 2008

Medical Linguistics, Part 5

The past few blogs have been a series devoted to Medical Linguistics. Yesterday's blog discussed the fundamental linguistic principles underlying the doublet method.

In today's blog, I'm posting a Perl script that extracts, from a large nomenclature, the terms that cannot be composed from doublets contained in other terms (i.e., the terms that must include unique doublets).

The key lines in the script are (that create the list of doublets) are shown here:

foreach $thing (@words)
{
$doublet = "$oldthing $thing";
if ($doublet =~ /^[a-z]+ [a-z]+$/)
{
$doublethash{$doublet} =
$doublethash{$doublet} + 1;
}
$oldthing = $thing;
}


These lines are used in virtually every Perl script that uses the doublet method. Basically, they move through an array consisting of the consecutive words in a nomenclature term, two words at a time, creating a new doublet and a new member of a doublet hash structure, with each loop. If you know Perl, this little piece of code should be easy to understand.

The entire Perl script follows here. As with all my posted scripts, the software is provided "as is", without warranty of any kind, express or implied, including but not limited to the warranties of merchantability, fitness for a particular purpose and noninfringement. in no event shall the authors or copyright holders be liable for any claim, damages or other liability, whether in an action of contract, tort or otherwise, arising from, out of or in connection with the software or the use or other dealings in the software.

#!/usr/local/bin/perl
open(TEXT,"neocl.xml")||die"cannot";
open(OUT,">dubuniq.txt")||die"cannot";
$line = " ";
while ($line ne "")
{
$line = <TEXT>;
next if ($line !~ /\"(C[0-9]{7})\"/);
next if ($line !~ /\"\> ?(.+) ?\<\//);
$line =~ /\"\> ?(.+) ?\<\//;
$phrase = $1;
@words = split(/ /, $phrase);
foreach $thing (@words)
{
$doublet = "$oldthing $thing";
if ($doublet =~ /^[a-z]+ [a-z]+$/)
{
$doublethash{$doublet} =
$doublethash{$doublet} + 1;
}
$oldthing = $thing;
}
}
close TEXT;
open(TEXT,"neocl.xml")||die"cannot";
$phrase = "";
$line = " ";
$count = 0;
while ($line ne "")
{
$oldthing = "";
$rightflank = "";
$leftflank = "";
$line = <TEXT>;
next if ($line !~ /\"(C[0-9]{7})\"/);
next if ($line =~ /\"C0000000\"/);
next if ($line =~ /\"C0000001\"/);
next if ($line !~ /\"\> ?(.+) ?\<\//);
$line =~ /\"\> ?(.+) ?\<\//;
$phrase = $1;
@words = split(/ /, $phrase);
next if (scalar(@words) < 3);
foreach $thing (@words)
{
$newdoublet = "$oldthing $thing";
if ($newdoublet =~ /^[a-z]+ [a-z]+$/)
{
if (exists($doublethash{$newdoublet}))
{
if ($doublethash{$newdoublet} == 1)
{
if ($phrase =~ /[a-z]+ $oldthing/)
{
$leftflank = $&;
}
if ($phrase =~ /$thing [a-z]+/)
{
$rightflank = $&;
}
unless ($doublethash{$leftflank} > 1
&& $doublethash{$rightflank} > 1)
{
$uniqphrase{$phrase} = "";
}
}
}
}
$oldthing = $thing;
}
}

while ((my $key, my $value) = each(%uniqphrase))
{
$count++;
print OUT "$count $key\n";
}
exit;

The output consists of a file composed of a list of terms, one line per term, that cannot be constructed from doublets found in other terms. This output was discussed in a prior blog. A sample list of doublets is available for download.

- Jules Berman

My book, Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information was published in 2013 by Morgan Kaufmann.



I urge you to read more about my book. Google books has prepared a generous preview of the book contents. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

tags: big data, metadata, data preparation, data analytics, data repurposing, datamining, data mining, doublet method, medical linguistics, medical algorithm, nomenclature

Sunday, March 2, 2008

Medical Linguistics, Part 3

In yesterday's blog, I wrote that medical terms are composed of doublets, each of which convey a very specific meaning. Individual words seldom have a single meaning.

Most terms in a medical nomenclature are composed of doublets found elsewhere in the terminology. In other words, unique terms are composed of common doublets, with very few exceptions.

The Neoplasm Classification contains over 130,000 names of neoplasms. Among these large numbers of terms, there are about 1,500 terms that contain a doublet that is uniquely found in the term (i.e., not found in one or more additional terms in the nomenclature). This represents about 1% of the total number of terms in the nomenclature. (The entire Neoplasms Classification is available as a gzipped file from my web site.

The Perl script that produces the list of terms that cannot be constructed from doublets found in other terms, is discussed in a later blog.

- Jules Berman

key words: doublet method, neoplasm classification, nomenclature,

Saturday, March 1, 2008

Medical linguistics, Part 2

This is a continuation of yesterday's blog.

One of the many challenges in the field of machine translation is that expressions (multi-word terms) convey ideas that transcend the meanings of the individual words in the expression. Consider the following sentence:

"The ciliary body produces aqueous humor."

The example sentence has unambiguous meaning to anatomists, but each word in the sentence can have many different meanings. "Ciliary" is a common medical word, and usually refers to the action of cilia. Cilia are found throughout the respiratory and GI tract and have an important role locomoting particulate matter. The word "body" almost always refers to the human body. The term "ciliary body" should (but does not) refer to the action of cilia that move human bodies from place to place. The word "aqueous" always refers to water. Humor relates to something being funny. The term "aqueous humor" should (but does not) relate to something that is funny by virtue of its use of water (as in squirting someone in the face with a trick flower). Actually, "ciliary body" and "aqueous humor" are each examples of medical doublets whose meanings are specific and contextually constant (i.e. always mean one thing). Furthermore, the meanings of the doublets cannot be reliably determined from the individual words that constitute the doublet, because the individual words have several different meanings. Basically, you either know the correct meaning of the doublet, or you don't.

Any sentence can be examined by parsing it into an array of intercalated doublets:

"The ciliary, ciliary body, body produces, produces aqueous, aqueous humor."

The important concepts in the sentence are contained in two doublets (ciliary body and aqueous humor). A nomenclature containing these doublets would allow us to extract and index these two medical concepts. A nomenclature consisting of single words might miss the contextual meaning of the doublets.

What if the term were larger than a doublet? Consider the tumor "orbital alveolar rhabdomyosarcoma." The individual words can be misleading. This orbital tumor is not from outer space, and the alveolar tumor is not from the lung. The 3-word term describes a sarcoma arising from the orbit of the eye that has a morphology characterized by tiny spaces of a size and shape as may occur in glands (alveoli). The term "orbital alveolar rhabdomyosarcoma" can be parsed as "orbital alveolar, alveolar rhabdomyosarcoma" Why is this any better than parsing the term into individual words, as in "orbital, alveolar, rhabdomyosarcoma"? The doublets, unlike the single words, are highly specific terms that are unlikely to occur in association with more than a few specific concepts.

Very few medical terms are single words. In the Neoplasm classification, there are over 135,000 terms and only about 500 are single words. The doublet method uses the multi-word feature of medical terms to extract meaning from text.

This topic is covered in detail in my book, Biomedical Informatics.
To be continued.

- Jules Berman


My book, Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information was published in 2013 by Morgan Kaufmann.



I urge you to read more about my book. Google books has prepared a generous preview of the book contents. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

tags: big data, metadata, data preparation, data analytics, data repurposing, datamining, data mining, medical autocoding, medical data scrubbing, medical data scrubber, medical record scrubbing, medical record scrubber, medical text parsing, medical autocoder, nomenclature, terminology

Saturday, February 23, 2008

Confused medical terms

There are many medical terms that have nearly the same orthography, are often pronounced identically, and have completely different meanings. These words are not picked up by spell checkers (because they are not misspelled), and occasionally appear as erroneous text within medical records.
Examples are:
acinic, actinic
anisakiasis, anisokaryosis
Apert syndrome, Alport syndrome (Apert syndrome 
    is a rare disorder characterized by early 
    fusion of skull bones. Alport syndrome is a 
    rare disorder characterized by kidney disease,
    hearing loss, and eye abnormalities.)
aptotic, apoptotic
arboreal, aboriginal
arteritis, arthritis
aural, oral
auxilliary, axillary
brachial, brachium, branchial
callous, callus
Carney triad, Carney complex (Carney Triad
    is gastric leiomyosarcoma, pulmonary chondroma
    extraadrenal paraganglioma, occurring 
    mainly in young women. Carney complex is
    myxoma, spotty pigmentation, and
    endocrinopathy.  Both Carney triad and Carney
    complex are technically Carney syndromes.
causality, casualty
chlorpropamide, chlorpromazine
chondroid, chordoid
chondroma, chordoma
chorionic, chronic
cingula, singular
coitus, colitis
colic, colonic
colitis, coitus
costal, coastal
cryptogam, cryptogram
cygnet, signet
decease, disease
deceased, desist
digitalize, digitize
digitate, digitize
dioecious, deciduous
diploic, diploid
disc, disk
disease, decease
diseased, deceased
disseminated sclerosis, systemic sclerosis (the first is 
        multiple sclerosis, and the second is scleroderma)
dyskaryosis, dyskeratosis
dysphasia, dysphagia
E coli (the Amoebozoa), E coli (the Enterobacteriaceae)
ectatic, ecstatic
endochondral, enchondral (these are synonyms)
engram, n-gram, ngram
epistasis, epistaxis, epitaxis 
        (the last is a misspelling of the second)
exxon, exon
facial, fascial
facies, faeces
falx, false
fetal, fatal
fibrinous, fibrous
fibrosis, fibrositis
firearm, forearm
foreword, forward
fossa, phossy
Gnathostoma, Gnathostomata (Gnathostoma genus of helminths, 
   Gnathostomata class of jawed vertebrates)
hallux, helicis
helicis, hallux
herpetic, herpangina
hydatid, hydatidiform
hypochondrium, hypochondria
ileitis, iliitis
ileum, ilium
insular, insulin
intercostal, intercoastal
intubation, incubation
isotope, isotrope
kerasin, kerosene, keratin
keratotic, keratinic, ketotic
keratinocytic, keratinolytic
keratosis, ketosis
lipoma, lymphoma
lumbar, lumber
malleolus, malleus
meniere disease, menetrier disease (former is an inner ear disorder, 
      latter is a hypertrophic gastropathy)
metachronous, metacrinus
milia, milium
miotic, mitotic, meiotic
mitogenic, mitogenomic
mitosis, meiosis, myosis, myiasis
monogenic, monogenetic, and Monogenetic (last, 
      related to class Monogenea)
mucous, mucus
myelofibrosis, myofibrosis
myelogenous, myelopathy (the former refers 
      to blood forming cells, the latter to spinal cord disease)
myelopathy, myotilinopathy (the former is 
      spinal cord disease, the latter a type of myofibrillar myopathy)
myofibroma, myelofibroma
neuroplastic, neoplastic
nucleus, nucleolus
oncocyte, onychocyte
oncology, ontology, ontogeny
organic, organoid
ornithine, ornithurine
otic, optic (otic ear, optic eye)
palatal, palatial
paleodontology, paleontology
palette, palate
palpation, palpitation
panacea, placebo (one cures all, the other cures none)
parasite, pericyte
parental, parenteral
pathogen, parthenogen
pathogenesis, parthenogenesis
pathogenic, pathogenetic (these two are synonyms)
pediculated, pedunculated
penal, penile, pineal, panel
penicillamine, penicillin
perineal, peroneal, perianal
phyllodes, phylloides
pigmentosa, pigmentosum (retinitis pigmentosa and xeroderma pigmentosum)
pleiotropic, pleiotrophic, pleiotypic (the first two are synonyms)
plural, pleural
polypoid, polyploid
porphyria, porphyruria
proptosis, ptosis
prostrate, prostate
protuberant, protruberant (the second term is simply a common misspelling)
pyelonephritis, pyonephritis
quinine, quinidine, quinone
rachischisis, rachitis, rachischitic, rachitic
radial, radical
relics, relicts
reticle, reticule, radical
rett syndrome, RET gene, 
rett syndrome, Tourette syndrome
rosacea, rosea
semantic, somatic
serous, serious
silicon, silicone
singleton, singultus
sinusitis, synositis
somatic, semantic, semitic
sonography, somnography, stenography
taenia, tinea
takoma, trachoma
thecoma, thekeoma
Tietze syndrome, Tietz syndrome (Tietze syndrome 
   is chondropathia tuberosa or costochondral junction 
   syndrome; Tietz syndrome is albinism-deafness
   syndrome, an autosomal dominant congenital disorder)  
torsion, distortion
trachoma, trachea
trichina, trachoma, trichura
trichinosis, trichosis, trichuriasis
trichrome, trichome
trochlear, tracheal
troglobite, troglodyte, trilobite
tuberous sclerosis, tuberculosis
tunicate, tourniquet, turbinate
typhoid, typhus
urethral, ureteral
vagitis, vaginitis
venous, venus
viscous, viscus
Medical transcriptionists and other healthcare professionals should be aware of the correct meaning of each alternate word in these listed pairs and groups.

In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site.

- Jules J. Berman, Ph.D., M.D. tags: common disease, orphan disease, orphan drugs, rare disease, medical terminology, medical errors, malaprop, malapropism, definition, confusing terms, confused medical terms, confusing medical terms, medical definitions, medical transcription, nomenclature, orphan drugs, rare diseases, terminology, medical transcription, common mistakes, common errors, typographical errors, sources of confusion, sources of error, medical dictionary, tricky medical terms, medical informatics

Wednesday, February 13, 2008

Ruby, Perl and Python medical autocoders

In the past two days on this blog, I've provided very short, fast, and accurate medical autocoders in Ruby and Perl. I thought I might as well offer the equivalent Python script. The Python script runs about twice as fast as either the Ruby or the Perl script.

The Ruby, Perl and Python scripts and their equivalent output are provided at:

http://www.julesberman.info/coded.htm

They are distributed under a GNU license.

All three scripts use a public domain file of 20,000 PubMed Citations, available at:

http://www.julesberman.info/tumorabs.txt

They all use an external tumor nomenclature contained within the Neoplasm Classification and available as a gzipped XML file distributed under a GNU license at:

http://www.julesberman.info/neoclxml.gz

- Jules Berman
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Wednesday, January 30, 2008

Arcane "iform" words in UMLS

Anonymous commented, on my Jan 7 blog Possessive forms of eponymous neoplasms:

"Interesting post. As a pathologist, I must interject that I have never used the term "rubriform" and don't know of any pathologists in the US that use it either. Perhaps it is an old term from the literature? Many of the terms for skin diseases, used by both dermatologists and dermatopathologists, are famously baroque - I can imagine a "rubriform" in that arena somewhere..."

Anonymous is correct. Most pathologists would never use "rubriform."

To satisfy my own curiosity, I extracted all of the English "iform" words found in UMLS.

Here they are:

acneiform
ansiform
apoplectiform
bacilliform
canaliform
cerebriform
chancriform
choreiform
chyliform
coliform
coralliform
cribiform
cribriform
cruciform
cuneiform
disciform
emboliform
epileptiform
falciform
filariform
filiform
flagelliform
fundiform
fungiform
fusiform
gigantiform
herpetiform
hydatidiform
hydatiform
ichthyosiform
intercuneiform
juxtarestiform
lentiform
morbilliform
multiform
neuralgiform
nonhydatidiform
pampiniform
piriform
pisiform
plexiform
prepyriform
proteiform
psoriasiform
punctiform
pyriform
reniform
restiform
retiform
retrolentiform
rubelliform
sacciform
scarlatiniform
schizophreniform
spongiform
storiform
subcuneiform
sublentiform
unciform
uniform (unintended)
varicelliform
varioliform
vermiform
verruciform
vitelliform
zosteriform

The UMLS seems to be missing a few that I have seen:

morpheiform
moniliform

In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D. tags: common disease, orphan disease, orphan drugs, rare disease, disease genetics, nomenclature, terminology

Monday, October 29, 2007

The high level classes in the Developmental Neoplasm ontology

The past several blogs have been devoted to the Developmental Lineage Classification and Taxonony of Neoplasms.

The rationale of the classification is that tumors inherit key cellular pathways through their developmental lineages. This assertion is supported by decades of morphologic evaluations of tumors. More recently, molecular biological observations have shown that genetic markers and pathways are carried through cell lineage. Tumors grouped by cell lineage may share responses to new chemotherapeutic and chemopreventive agents targeted to specific pathways. If this is true, we can start to develop agents (and combinations of agents) that are effective against groups of neoplasms that share a common developmental lineage.

The taxonomy contains the names of over 5,000 different neoplasms, and about 130,000 synonymous terms. It is the most comprehensive listing of neoplasms in the world, and it is distributed under the GNU Free Documentation License. Download information is available from my website home page.

Here are the top level classes in the cancer ontology:


‹rdfs:Class rdf:ID="Neural_tube_parenchyma"›
‹rdfs:subClassOf
neo:resource="#Neural_tube"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Sub_coelomic"›
‹rdfs:subClassOf
neo:resource="#Mesoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Endoderm_or_ectoderm"›
‹rdfs:subClassOf
neo:resource="#Neoplasm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Syndrome"›
‹rdfs:subClassOf
neo:resource="#Unclassified"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Neural_crest"›
‹rdfs:subClassOf
neo:resource="#Neoplasm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Germ_cell"›
‹rdfs:subClassOf
neo:resource="#Neoplasm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Sub_coelomic_gonadal"›
‹rdfs:subClassOf
neo:resource="#Sub_coelomic"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Molar"›
‹rdfs:subClassOf
neo:resource="#Trophectoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Neural_crest_endocrine"›
‹rdfs:subClassOf
neo:resource="#Neural_crest"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Sub_coelomic_endocrine"›
‹rdfs:subClassOf
neo:resource="#Sub_coelomic"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Fibrous_tissue"›
‹rdfs:subClassOf
neo:resource="#Connective_tissue"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Mesoderm_primitive"›
‹rdfs:subClassOf
neo:resource="#Mesoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Trophectoderm"›
‹rdfs:subClassOf
neo:resource="#Neoplasm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Sub_coelomic_nephric"›
‹rdfs:subClassOf
neo:resource="#Sub_coelomic"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Tumor_classification"›
‹rdfs:subClassOf
rdfs:resource=
"http://www.w3.org/2000/01/rdf-schema#Class"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Neoplasm"›
‹rdfs:subClassOf
rdfs:resource="#Tumor_classification"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Vascular"›
‹rdfs:subClassOf
neo:resource="#Connective_tissue"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Germ_cell_differentiated"›
‹rdfs:subClassOf
neo:resource="#Germ_cell"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Endoderm_or_ectoderm_parenchymal"›
‹rdfs:subClassOf
neo:resource="#Endoderm_or_ectoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Peripheral_nervous_system"›
‹rdfs:subClassOf
neo:resource="#Neural_crest"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Coelomic_ductal"›
‹rdfs:subClassOf
neo:resource="#Coelomic"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Coelomic_gonadal"›
‹rdfs:subClassOf
neo:resource="#Coelomic"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Trophoblast"›
‹rdfs:subClassOf
neo:resource="#Trophectoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Muscle"›
‹rdfs:subClassOf
neo:resource="#Connective_tissue"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Mesenchyme"›
‹rdfs:subClassOf
neo:resource="#Mesoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Neural_crest_primitive"›
‹rdfs:subClassOf
neo:resource="#Neural_crest"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Neural_crest_ectomesenchymal"›
‹rdfs:subClassOf
neo:resource="#Neural_crest"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Stage"›
‹rdfs:subClassOf
neo:resource="#Unclassified"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Neural_tube"›
‹rdfs:subClassOf
neo:resource="#Neuroectoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Neuroectoderm"›
‹rdfs:subClassOf
neo:resource="#Neoplasm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Connective_tissue"›
‹rdfs:subClassOf
neo:resource="#Mesenchyme"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Neural_crest_melanocytic"›
‹rdfs:subClassOf
neo:resource="#Neural_crest"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Neural_tube_lining"›
‹rdfs:subClassOf
neo:resource="#Neural_tube"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Mesoderm"›
‹rdfs:subClassOf
neo:resource="#Neoplasm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Unclassified_precancer"›
‹rdfs:subClassOf
neo:resource="#Unclassified"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Unclassified_cancer"›
‹rdfs:subClassOf
neo:resource="#Unclassified"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Coelomic"›
‹rdfs:subClassOf
neo:resource="#Mesoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Bone_cartilage"›
‹rdfs:subClassOf
neo:resource="#Connective_tissue"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Coelomic_cavity"›
‹rdfs:subClassOf
neo:resource="#Coelomic"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Heme_lymphoid"›
‹rdfs:subClassOf
neo:resource="#Mesenchyme"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Adipose_tissue"›
‹rdfs:subClassOf
neo:resource="#Connective_tissue"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Neuroectoderm_primitive"›
‹rdfs:subClassOf
neo:resource="#Neuroectoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Endoderm_or_ectoderm_primitive"›
‹rdfs:subClassOf
neo:resource="#Endoderm_or_ectoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Unclassified"›
‹rdfs:subClassOf
neo:resource="#Tumor_classification"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Primordial"›
‹rdfs:subClassOf
neo:resource="#Germ_cell"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Endoderm_or_ectoderm_surface"›
‹rdfs:subClassOf
neo:resource="#Endoderm_or_ectoderm"/›
‹/rdfs:Class›

‹rdfs:Class rdf:ID="Endoderm_or_ectoderm_endocrine"›
‹rdfs:subClassOf
neo:resource="#Endoderm_or_ectoderm"/›
‹/rdfs:Class›



In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D.

Sunday, October 28, 2007

Developmental Classification of Neoplasms now an RDF Ontology

I am publishing today the first ontology version of the Developmental Lineage Classification and Taxonomy of Neoplasms. It is available for download in several file versions.

The full ontology is a 10 Megabyte RDF file. Note that the file is so large that some browsers may not be able to open the entire file. On my computer, I had no trouble opening the file in my Internet Explorer browser, but the file was too large for my Mozilla browser.
http://www.julesberman.info/neordf.xml


The file was validated using the w3c validator service at http://www.w3.org/rdf/validator/, with a caveat. The full ontology file (10+ Mbytes) was too large for the validator, so I truncated the ontology, validated the truncated file (that contained all of the classes, subclasses, properties), and left out the repetitive list of terms. Then I took the entire file and validated it with an XML parser to verify that the file was well-formed. That really covers everything (RDF logic and XML structure).

The gzipped version of the RDF file (under 1 Megabyte).
http://www.julesberman.info/neorxml.gz


The flat file version, listing each term followed by its lineage (gzipped file).
http://www.julesberman.info/neoself.gz


The plain old XML version, with no RDF semantics (gzipped file). http://www.julesberman.info/neoclxml.gz

The ontology contains several parts:

1. The neoplasm classification proper (as illustrated in the schematic)



2. A listing of cancer terms that will probably never be entered into the proper classification (more about this later)

3. A listing of hyperplasias or hamartomas, some of which will be entered into the proper classification and others of which will remain in class Hyperplasia

4. A listing of precancer terms

5. A listing of syndromes associated with increased risk for cancer.

In this version, there are 5841 classified types of neoplasms and 130,503 terms representing the 5,841 types of neoplasms.

This represents the largest nomenclature of neoplasms in existence and, with today's publication, the largest formal ontology (in RDF syntax) of neoplasm names.

Over the next few weeks, I'll post additional blogs to further explain the RDF ontology files.

- Jules Berman
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Saturday, October 27, 2007

National Cancer Institute Thesaurus

The National Cancer Institute (NCI) Thesaurus is a free medical vocabulary available in OWL format from:

ftp://ftp1.nci.nih.gov/pub/cacore/EVS/NCI_Thesaurus/


It's really quite an impressive document, and there are very few standardized vocabularies that have been prepared as formal ontologies. The creators wisely used the semantics of OWL (Web Ontology Language), a dialect of RDF.

The NCI thesaurus contains terms related to the interests of the NCI and contains the names of many neoplasms.

This vocabulary has been curated for over a decade by in-house ontologists (NCI employees), contractors, and through the use of domain consultants (including some pathologists). It is updated monthly. A lot of money has gone into the development of the NCI Thesaurus, and it is one of the most worked-on vocabularies in the medical field.

The NCI Thesaurus has been reviewed by Barry Smith and colleagues, who found it somewhat lacking.

http://ontology.buffalo.edu/medo/NCIT.pdf


"RESULTS: We found many mistakes and inconsistencies
with respect to the term-formation principles used,
the underlying knowledge representation system,
and missing or inappropriately assigned verbal and
formal definitions.."
Ceusters W, Smith B, Goldberg L.
A terminological and ontological analysis of the
NCI Thesaurus. Methods Inf Med. 2005;44(4):498-507.

My question is, "If the Thesaurus contains many different knowledge domains (medications, general diseases, neoplasms, etc.) how can it adequately cover all of its constituent domains?" In the neoplasm domain, it is missing many thousands of names of neoplasms. The terminology may be sufficient for its intended purpose (meeting the needs of the NCI community), but because the terminology is not comprehensive, the NCI Thesaurus will not necessarily serve those who want a thesaurus that comes close to including the names of ALL neoplasms.

Also, there doesn't seem to be any single organizing principle for the neoplasm domain. Some neoplasms are subclassed by their anatomic site (e.g. urinary tract neoplasm). Others are subclassed by their tissue type (e.g. soft tissue neoplasm). And so on. This is allowable under an ontology, so long as the ontology maintains consistency and competence (ability to answer questions about the members of classes). But I wonder if this is the best way of organizing tumors. Of course, I'm deeply biased. The Developmental Lineage Classification and Taxonomy of Neoplasms has a single organizing principle.

The NCI Thesaurus is an impressive piece of work and definitely worth looking over.

tags: biomedical informatics, cancer, classification, nomenclature, thesaurus, vocabulary, ontology, rare diseases, orphan drugs, genetics of disease, pathology, common diseases, complex diseases

In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D.

Sunday, September 16, 2007

Latest update of Neoplasm Classification available

The latest version of the Developmental Lineage Classification and Taxonomy of Neoplasms is now available as a gzipped file at:

http://www.julesberman.info/neoclxml.gz

This Neoplasm Classification has been described at:

http://www.biomedcentral.com/1471-2407/4/10

It contains 5,827 neoplasm classified concepts and 130,283 different terms (codes beginning with "C"). It is more than ten times larger than any other neoplasm classification.

In addition to specific neoplasm concepts, it also contains general neoplastic terms (coded as C0000000), inherited conditions associated with neoplasms (codes beginning with "S") and terms related to the stage or anatomic location of neoplasms (codes beginning with "ST").


In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D.

Wednesday, July 4, 2007

Unclassified terms in the Neoplasm Classification

The gzipped version (de-compress with gunzip utility) of the Developmental Lineage Classification and Taxonomy of Neoplasms is available for public download.

The total number of included cancer-related terms exceeds 146,000.

In addition to (and following within the file) the list of classified neoplasm terms is a list of unclassified cancer related terms (all identified by the same identifier, "C0000000").

This list of unclassified terms consists of general cancer terms that do not specify any particular neoplasm; overly specific terms that provide so-call pre-coordinated annotations related to terms contained elsewere in the Classification; and valid terms that have not been added (yet) to the list of classified neoplasm terms.

Examples of non-specific cancer-related terms are:

-borderline tumor
-mucinous tumor
-blast crisis
-preinvasive carcinoma
-dysplasia

Examples of overly specific terms are:

-squamous carcinoma of the nasal vestibule
-gastric non-hodgkin lymphoma of mucosa-associated lymphoid tissue
-primary primitive neuroectodermal tumor of the kidney

The terms that are currently unclassified and are awaiting inclusion in the classified section were added by putting curated candidate terms in an external file and parsing these candidate terms with a Perl script that checks to see if they are already in the Classification and that automatically assigns them a "C0000000" code if they are new. The Perl script, addterm.pl, is one of many "helper" scripts that the curator uses to facilitate growth of the classification. It is shown here:


#!/usr/local/bin/perl
#addterm.pl
#
#This Perl script was created by Jules J. Berman and is entered
#into the Public Domain
#
#The software is provided "as is", without warranty of any kind,
#express or implied, including but not limited to the warranties
#of merchantability, fitness for a particular purpose and
#noninfringement. in no event shall the authors or copyright
#holders be liable for any claim, damages or other liability,
#whether in an action of contract, tort or otherwise, arising
#from, out of or in connection with the software or the use or
#other dealings in the software.
#
open (TEXT,"neocl.xml")||die"Cannot";
my $line = " ";
my %doubhash;
while ($line ne "")
{
$line = <TEXT>;
next if ($line !~ /C[0-9]{7}/);
$line =~ /\"\> ?(.+) ?\<\//;
$phrase = $1;
$doubhash{$phrase}="";
}
close TEXT;
open (TEXT,"newneocl.txt")||die"Cannot";
open (OUT,">new.out")||die"Cannot";
my $key = " ";
while ($key ne "")
{
$key = <TEXT>;
$key =~ s/\n//;
next if ($key eq "");
if (exists $doubhash{$key})
{
print "$key already exists\n";
}
else
{
print OUT "\ print OUT "\= \"C0000000\"\>";
print OUT "$key\<\/name\>\n";
}
}
exit;

-Jules J. Berman
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Tuesday, June 19, 2007

"Precancer" versus "early cancer"

Precancers are the lesions from which cancers grow. Some people question why we need to specify some lesions as precancers when we know that carcinogenesis is a multistep process and that every cancer traverses many un-named biological states as it develops into a fully malignant lesion. Why can't we recognize that precancers are just an early form of cancer and refer to the precancers by the name of its developed cancer? Can't we just use adjectives like "early stage" squamous carcinoma or "non-invasive" pancreatic carcinoma? Wouldn't that make life a lot easier than inventing names for the pre-invasive stage of every cancer?

Much as I like data simplification, it just can't be done in the case of the precancers. Precancers have specific, characteristic properties that separate them from cancers. Because of these properties, the treatment of precancers may be very different from the treatment of cancers. In fact, if we take full advantage of the biologic features that separate the precancers from the cancers, we may actually find that we can eliminate deaths from cancer.

What are these special properties of the precancers?

1. Precancers, unlike cancers, tend to regress. Cancers tend to grow and only rarely regress. Furthermore, it some cases, we can influence the rate of regression of the precancers with relatively non-toxic drugs. Understanding the biology of regression is something that we can only learn from the precancers.

2. When a precancer progresses, it progresses to cancer. But not all precancers progress. Many precancers just stay precancers indefinitely, as far as we can tell. Why should we think of a precancer as an early stage of a cancer if it never becomes a cancer?

3. Precancers that progress to cancer can apparently progress into more than one type of cancer. Consequently, there are more types of cancers than there are types of precancers. For instance, in the lung, squamous metaplasia/dysplasia of bronchial epithelium may give rise to bronchogenic squamous cell carcinoma, bronchogenic adenocarcinoma, bronchogenic small cell carcinoma, or bronchogenic mixed carcinoma. If a lesion can progress into any of several different lesions, it is impossible to pretend that the lesion is just an early form of one named cancer.

4. Precancers can be cured. When a precancer is cured, the cancer never develops. The treatments that we use for precancers are likely to be different from (and much less toxic than) the treatments that we use for cancers.

Because the biology of precancer is distinguishable from the biology of cancer, and because there are clinically useful reasons (i.e., treatment and prevention of cancer) to make these distinctions, the precancers should be curated as designated entities.

Jules Berman

Friday, March 2, 2007

New version of neoplasm classification available

The latest version of the Developmental Lineage Classification and Taxonomy of Neoplasms is now available at:

NEOCLXML.GZ 716,963 bytes and
NEOSELF.GZ 1,086,677 bytes

Neoclxml.gz expands to over 10 Megabytes and is an XML file.
Neoself.gz expands to over 20 Megabytes and is a flat-file.

Each file contains over 145,000 neoplasm terms grouped in >6,000 concepts, and classified according to embryonic lineage. This is, by far, the most extensive nomenclature and classification of neoplasms in existence. It is copyrighted to Jules J. Berman and distributed under a GNU document license.

Detailed information on the classification is available in my article:
Tumor classification: molecular analysis meets Aristotle

tags: cancer, medical terminology, nomenclature, open access, open source
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.