Showing posts with label neoplasm classification. Show all posts
Showing posts with label neoplasm classification. Show all posts

Wednesday, January 28, 2009

Update of Neoplasm Classification is now available

I'm interrupting my series of blogs on bimodal cancer age distributions to announce the release of the most recent version of the Developmental Lineage Classification and Taxonomy of Neoplasms.

The current classification contains 6083 neoplasm concepts (types of neoplasms) classified under 122,698 terms. It also contains a large number of unclassified neoplasm terms as addendum items. It is, by far and away, the world's largest neoplasm nomenclature.

The classification is available in XML, RDF and flat-file formats. Here is the preface text distributed with each formatted version:

"This file was prepared by Jules J. Berman. The first version of this file was created November 15, 2003. The current version was created on January 27, 2009.

Copyright © 2003-2009 Jules J. Berman

Permission is granted to copy, distribute and/or modify this document under the terms of the GNU Free Documentation License, Version 1.2 or any later version published by the Free Software Foundation; with no Invariant Sections, no Front-Cover Texts, and no Back-Cover Texts. A copy of the license is available.

The neoclxml file is provided "as is", without warranty of any kind, expressed or implied, including but not limited to the warranties of merchantability, fitness for a particular purpose and noninfringement. in no event shall the author or copyright holder be liable for any claim, damages or other liability, whether in an action of contract, tort or otherwise, arising from, out of or in connection with the software or the use or other dealings in the software.

An explanation of the classification can be found in the following two publications, which should be cited in any publication or work that may result from any use of this file.

Berman JJ. Tumor classification: molecular analysis meets Aristotle. BMC Cancer 4:8, 2004.

Berman JJ. Neoplasms: Principles of Development and Diversity. Jones and Bartlett Publishers, Sudbury, MA, 2009.

In the Neoplasm Classification, all classified names of neoplasms are coded with a "C" followed by a 7 digit number other than 0000000.

For example, "C9168000" = rectal signet ring adenocarcinoma

In addition to classified terms, there are three groups of unclassified terms that are provided special items that follow the list of classified terms in this file.

"C0000000"
"S" followed by 7 digits
"ST" followed by 7 digits

This list of unclassified terms coded as "C0000000" consists of general cancer terms that do not specify any particular neoplasm; overly specific terms that provide so-call pre-coordinated annotations related to terms contained elsewere in the Classification; and valid terms that have not been added (yet) to the list of classified neoplasm terms.

Examples of overly specific terms are:

squamous carcinoma of the nasal vestibule, gastric non-hodgkin lymphoma of mucosa-associated lymphoid tissue, primary primitive neuroectodermal tumor of the kidney

The terms that are coded "S" followed by 7 digits are inherited syndromes that have a neoplastic component (i.e., the occasional or frequent appearance of neoplasms in the syndrome).

The terms that are coded "ST" followed by 7 digits are staging terms used by oncologists.

The classification is meant for informatics projects that use computer parsing techniques. Programmers should simply insert statements that filter the unclassified terms included in the file."

Additional information may be available from the author's web site:
http://www.julesberman.info/devclass.htm

The Neoplasm Classification is available as a zipped XML file at:
http://www.julesberman.info/neoclxml.zip

The Neoplasm Classification is available as a zipped flat file at:
http://www.julesberman.info/neoself.zip

The Neoplasm Classification is available as a zipped RDF file at:
http://www.julesberman.info/neordf.zip

© 2009 Jules Berman

key words: medical nomenclature, classification, rdf, xml, ontology, data mining, ontology, science
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Sunday, August 31, 2008

New version of the Developmental Lineage Classification now available

The latest update to the Developmental Lineage Classification and Taxonomy of Neoplasms is now available at:

http://www.julesberman.info/devclass.htm

The Neoplasms Classification contains about 135,000 neoplasm names collected under about 6,000 concepts. It is the largest neoplasm nomenclature in existence and is available as an open source document in several different file formats.

-Jules Berman

key words: nomenclature, ontology, cancer, tumors, terminology

Tuesday, August 26, 2008

Neoplasm synonym look-up

I just created a web site that permits anyone to enter the name of a neoplasm and retrieve all of the synonyms and all of the related terms for the entry term.

The search engine is available at:

http://www.julesberman.info/neoget.htm

It uses the Developmental Lineage Classification and Taxonomy of Neoplasms, which consists of over 135,000 neoplasm terms. This open source dataset is available as compressed RDF, text or XML files:

The gzipped version of the RDF file (under 1 Megabyte)

http://www.julesberman.info/neorxml.gz


The flat file version, listing each term followed by its lineage (gzipped file).

http://www.julesberman.info/neoself.gz


The plain old XML version, with no RDF semantics (compressed gzip file).

http://www.julesberman.info/neoclxml.gz


The plain old XML version, with no RDF semantics (compressed zip file).

http://www.julesberman.info/neoclxml.zip


More information on uses for the Developmental Lineage Classification is available from my home page.

- Jules J. Berman

key words: ontology, tumors, cancers, neoplasia, medical terminology, nomenclature, dictionary, concept unique identifier, medical terminology, concept identifier, neoplasm classification, taxonomy

Friday, March 14, 2008

Creating a directed graph from an RDF schema

In yesterday's blog , I announced the newest version of the Neoplasm Classification. The Neoplasm Classification is an open source document available as a plain-text flat-file, as an XML file, and as and RDF document. The top of the RDF document contains the complete RDF Schema for the Classification. The remainder of the file (>99% of the file) is devoted to the entries for the individual neoplasm terms (over 135,000 of them).

I have a small schematic, that represents the organization of the Neoplasm Classification.



In addition to this small schematic, I have a large schematic that represents the complete hierarchy of the Classification. Though it is too large to fit in this blog, each part of the schematic is quite simple.

It took a few seconds to create the complete diagram of the classification, in digraph (directed graph) form, using a short Perl script that parsed the Classification's RDF Schema + one command-line instruction to the free and open source GraphViz application.

Here's the script. As per usual, the following statement applies. The software is provided "as is", without warranty of any kind, express or implied, including but not limited to the warranties of merchantability, fitness for a particular purpose and noninfringement. in no event shall the authors or copyright holders be liable for any claim, damages or other liability, whether in an action of contract, tort or otherwise, arising from, out of or in connection with the software or the use or
other dealings in the software.

#!/usr/bin/perl
open (TEXT, "schema.txt");
open (OUT, ">schema.dot");
$/ = "\<\/rdfs\:Class>";
print OUT "digraph G \{\n";
print OUT "size\=\"15\,15\"\;\n";
print OUT "ranksep\=\"2\.00\"\;\n";
$line = " ";
while ($line ne "")
{
$line = <TEXT>;
last if ($line !~ /\<rdfs\:/);
if ($line =~
/\:resource\=\"[a-z0-9\:\/\_\.\-]*\#([a-z\_]+)\"/i)
{
$father = $1;
}
if ($line =~ /rdf\:ID\=\"([a-z\_]+)\"/i)
{
$child = $1;
}
print OUT "$father \-\> $child\;\n";
print "$father \-\> $child\;\n";
}
print OUT "\}";
exit;

If you work with RDF (and every biomedical professional should understand how RDF is used to specify data), you will want a method that can instantaneously render a schematic of your RDF Schema (ontology) or of any descendant section of your Schema.

Tomorrow, I'll go into some detail to describe just how the Perl script produces a GraphViz script that can render an RDF Schema as a visual digraph.

- Jules Berman


My book, Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information was published in 2013 year by Morgan Kaufmann.



I urge you to read more about my book. Google books has prepared a generous preview of the book contents. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.
tags: big data, metadata, data preparation, data analytics, data repurposing, datamining, data mining, medical autocoding, medical data scrubbing, medical data scrubber, medical record scrubbing, medical record scrubber, medical text parsing, medical autocoder, nomenclature, terminology, ontology, VizGraph, Neoplasm Classification, semantic web, ontologies, digraph, directed graph, tumor, cancer, tumour, neoplasia

Wednesday, February 13, 2008

Ruby, Perl and Python medical autocoders

In the past two days on this blog, I've provided very short, fast, and accurate medical autocoders in Ruby and Perl. I thought I might as well offer the equivalent Python script. The Python script runs about twice as fast as either the Ruby or the Perl script.

The Ruby, Perl and Python scripts and their equivalent output are provided at:

http://www.julesberman.info/coded.htm

They are distributed under a GNU license.

All three scripts use a public domain file of 20,000 PubMed Citations, available at:

http://www.julesberman.info/tumorabs.txt

They all use an external tumor nomenclature contained within the Neoplasm Classification and available as a gzipped XML file distributed under a GNU license at:

http://www.julesberman.info/neoclxml.gz

- Jules Berman
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Tuesday, January 8, 2008

Rhabdoid tumors, a sui generis class in the Neoplasm Classification

Rhabdoid tumors occur in just a few dozen children each year in the U.S. These aggressive tumors arise from brain (class Neuroectoderm in the Neoplasm Classification) and from kidney (class Mesoderm in the Neoplasm Classification) and contain morphologically distinctive cells (so-called rhabdoid cells, named for their superficial similarity to muscle cells).

Not long ago, the rhabdoid tumors that arose in the kidney were thought to be different from the rhabdoid tumors that arose from the brain. The common rhabdoid cell was considered to be a peculiar morphologic variant that just happened to occur in both types of tumors. The kidney tumor was called MRT (malignant rhabdoid tumor). MRT was considered, by many pathologists, to be a variant form of nephroblastoma. The rhabdoid brain tumor was called AT/RT (Atypical teratoid rhabdoid tumor), and some pathologists classified AT/RT among the PNETs.

The perceived distinction between MRT and AT/RT started to disappear apart when it was found that about 10% of patients with MRT also developed AT/RT or so-called primitive neuroectodermal neoplasm of brain. Recently, a characteristic genetic abnormality has been found in rhabdoid tumors of CNS or renal origin: bi-allelic loss of INI1 gene expression. Immunostaining for the protein produced by the INI1 gene is absent in almost all reported cases of rhabdoid tumor cells.

Rhabdoid cells are large cells with an eosinophilic cytoplasm. Ultrastructural examination of rhabdoid cells shows characteristic whorls of intermediate filaments. Intermediate filaments are proteins contribute to the the structure of cells and provide resistance to deformity. Normal cells contain intermediate filaments that are specific for their developmental lineage. Cells of endodermal or ectodermal origin contain cytokeratin filaments. Cell of Neuroectodermal origin contain neurofilaments filaments. Cells of mesenchymal origin contain desmin filaments. Rhabdoid cells contain all these different types of intermediate filaments.

Rhabdoid tumor apparently disobey some of the rules of neoplastic development.

-Rhabdoid tumors are lineage non-specific and can arise from neuroectodermal cells or mesodermal cells. All other tumors of somatic cells (non-germ cells) arise from a single germ lineage.

-Cells within a single rhabdoid tumor seem to have differentiated along several developmental lineages (ectodermal, endodermal, neuroectodermal and mesodermal). This phenomenon is otherwise restricted to totipotent germ cell tumors.

-Cells within a single rhabdoid tumor may include primitive cells indistinguishable from PNET tumors (primitive neuroectodermal tumors). PNET tumors are typically monomorphic tumors.

-Rhabdoid tumors are all associated with a specific phenotypic cell (the rhabdoid cell) regardless of the developmental origin of the tumor (Neuroectoderm or Mesoderm).

-The rhabdoid cells contain several different intermediate filaments. In normal cells, only one type of intermediate filaments is found in any single cell, and that filament is specific for the lineage of the cell.

-Almost all tumors arise from cells that resemble an observable normal cell. For example, a squamous cell carcinoma is composed of cells that resemble normal squamous cells biochemically, ultrastructurally, and by light microscopic examination. The rhabdoid cell has no known counterpart in any adult tissue or in any stage of development.

-Rhabdoid tumors seem to be caused by bi-allelic loss of the tumor suppressor gene, INI1. Tumor suppressor genes are normally associated with different types of tumors. INI1 gene loss seems to be exclusively associated with rhabdoid tumors or with rhabdoid tumor cell subpopulations developing within other types of tumors. Currently, there is no other known genetic mutation that produces a specific phenotype akin to the association between INI1 and rhabdoid cells.

-INI1 loss produces rhabdoid tumors in mice. Eight of 125 mice with germline haploid complement of INI (Snf5+/- mice) developed INI1-negative tumors of soft tissue origin and rhabdoid cell morphology RrobbR. haploid. The mouse tumor is morphologically and genetically identical to the human tumor. Despite the phenotype and genotypic similarities between murine and human rhabdoid tumors, the mouse tumor arises from the branchial arch soft tissue, a tissue of origin not observed in human rhabdoid tumors.

How is this possible? How can a tumor suppressor gene mutation in a non-germ cell produce tumors that arise from tissues of different developmental lineage and contain a common tumor cell that contains intermediate filaments specific for multiple cell lineages?

The INI1 gene codes a subunit of the SWI/SNF chromatin remodelling complex. The study of SWI/SNF complexes in mammals, flies and plants suggests that they strongly influence many developmental pathways RkwoaR. This suggests two possible mechanisms for the action of the INI1 tumor suppressor gene in rhabdoid tumors

-1. The cell of origin of rhabdoid tumors is a primitive, pluripotent somatic stem cell with a lineage position that precedes the development of germ layers.

-2. Loss of INI1 gene expression produces a pluripotent stem cell when it occurs in cells of several different lineages (e.g., neuroectodermal or endodermal or mesodermal cells).

Of these two possibilities, the first seems unlikely because subpopulations of INI1-negative rhabdoid cells occur secondarily within adult tumors, and this would not be expected to occur if the INI1 mutation only produced rhabdoid tumors derived from primitive cells. Also, if there were a very primitive cell (with a lineage that precedes the development of the germ layers), why would the INI1 mutation be the only mutation that could produce tumors of these cells?

The second possibility, if true, might explain the odd biological features of rhabdoid tumors. The idea of mutations in differentiated cells conferring pluripotentiality is not unprecedented. In recent work, Yamanaka and coworkers successfully created pluripotent stem cells from differentiated fibroblasts by introducing (with retroviruses) several chosen genes (Oct3/4, Sox2, c-Myc and Klf4) and subsequent selection for Fbx15 RokiaR.

See: Okita K, Ichisaka T, Yamanaka S. Generation of germline-competent induced pluripotent stem cells. Nature 448:313-317 2007.

These experiments demonstrated that differentiated cells could be altered to become stem cells and provides researchers with a source of stem cells other than embryos.

It would seem that regardless of the lineage of origin, INI1 biallelic loss creates tumors of a unique phenotype.

As a result, I'm changing the class for rhabdoid tumors within the Neoplasms Classification. It is now a one-of-a-kind tumor in its own class, placed under the Neoplasm superclass.

The updated Neoplasm Classification is available as a gzipped XML file at:

http://www.julesberman.info/neoclxml.gz


An early paper on the Neoplasm Clasification is available as an open source document.

In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D. tags: common disease, orphan disease, orphan drugs, rare disease, disease genetics, medical terminology, rare cancers,

Saturday, July 7, 2007

Synonymy in the Neoplasm Classification

The Developmental Lineage Classification and Taxonomy of Neoplasms contains (on 7/7/07), 130,359 classified neoplasm terms distributed over 5,826 concepts, yielding an average exceeding 20 terms per concept. An example are the synonymous terms for adenocarcinoma of the prostate.

prostate with adenoca
adenoca arising in prostate
adenoca involving prostate
adenoca arising from prostate
adenoca of prostate
adenoca of the prostate
prostate with adenocarcinoma
adenocarcinoma arising in prostate
adenocarcinoma involving prostate
adenocarcinoma arising from prostate
adenocarcinoma of prostate
adenocarcinoma of the prostate
adenocarcinoma arising in the prostate
adenocarcinoma involving the prostate
adenocarcinoma arising from the prostate
prostate with ca
ca arising in prostate
ca involving prostate
ca arising from prostate
ca of prostate
ca of the prostate
prostate with cancer
cancer arising in prostate
cancer involving prostate
cancer arising from prostate
cancer of prostate
cancer of the prostate
cancer arising in the prostate
cancer involving the prostate
cancer arising from the prostate
prostate with carcinoma
carcinoma arising in prostate
carcinoma involving prostate
carcinoma arising from prostate
carcinoma of prostate
carcinoma of the prostate
carcinoma arising in the prostate
carcinoma involving the prostate
carcinoma arising from the prostate
prostate adenoca
prostate adenocarcinoma
prostate ca
prostate cancer
prostate carcinoma
prostatic cancer
prostatic carcinoma
prostatic adenocarcinoma
prostate gland adenocarcinoma
adenocarcinoma of the prostate gland
adenocarcinoma of prostate gland
prostate gland carcinoma
carcinoma of the prostate gland
carcinoma of prostate gland

This kind of synonymy is needed if you want to implement autocoding software that will successfully capture all of the cancer terms included in a sampled text.


In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D. tags: common disease, orphan disease, orphan drugs, rare disease, medical terminology, medical transcription, nomenclature, terminology, pitfalls in medical terminology

Wednesday, July 4, 2007

Unclassified terms in the Neoplasm Classification

The gzipped version (de-compress with gunzip utility) of the Developmental Lineage Classification and Taxonomy of Neoplasms is available for public download.

The total number of included cancer-related terms exceeds 146,000.

In addition to (and following within the file) the list of classified neoplasm terms is a list of unclassified cancer related terms (all identified by the same identifier, "C0000000").

This list of unclassified terms consists of general cancer terms that do not specify any particular neoplasm; overly specific terms that provide so-call pre-coordinated annotations related to terms contained elsewere in the Classification; and valid terms that have not been added (yet) to the list of classified neoplasm terms.

Examples of non-specific cancer-related terms are:

-borderline tumor
-mucinous tumor
-blast crisis
-preinvasive carcinoma
-dysplasia

Examples of overly specific terms are:

-squamous carcinoma of the nasal vestibule
-gastric non-hodgkin lymphoma of mucosa-associated lymphoid tissue
-primary primitive neuroectodermal tumor of the kidney

The terms that are currently unclassified and are awaiting inclusion in the classified section were added by putting curated candidate terms in an external file and parsing these candidate terms with a Perl script that checks to see if they are already in the Classification and that automatically assigns them a "C0000000" code if they are new. The Perl script, addterm.pl, is one of many "helper" scripts that the curator uses to facilitate growth of the classification. It is shown here:


#!/usr/local/bin/perl
#addterm.pl
#
#This Perl script was created by Jules J. Berman and is entered
#into the Public Domain
#
#The software is provided "as is", without warranty of any kind,
#express or implied, including but not limited to the warranties
#of merchantability, fitness for a particular purpose and
#noninfringement. in no event shall the authors or copyright
#holders be liable for any claim, damages or other liability,
#whether in an action of contract, tort or otherwise, arising
#from, out of or in connection with the software or the use or
#other dealings in the software.
#
open (TEXT,"neocl.xml")||die"Cannot";
my $line = " ";
my %doubhash;
while ($line ne "")
{
$line = <TEXT>;
next if ($line !~ /C[0-9]{7}/);
$line =~ /\"\> ?(.+) ?\<\//;
$phrase = $1;
$doubhash{$phrase}="";
}
close TEXT;
open (TEXT,"newneocl.txt")||die"Cannot";
open (OUT,">new.out")||die"Cannot";
my $key = " ";
while ($key ne "")
{
$key = <TEXT>;
$key =~ s/\n//;
next if ($key eq "");
if (exists $doubhash{$key})
{
print "$key already exists\n";
}
else
{
print OUT "\ print OUT "\= \"C0000000\"\>";
print OUT "$key\<\/name\>\n";
}
}
exit;

-Jules J. Berman
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Sunday, July 1, 2007

Neoplasm classification structural validation

In yesterday's post, I announced the newest version of the Developmental Lineage Classification and Taxonomy of Neoplasms (also called the Neoplasm Classification).

When you have a nomenclature that contains hundreds of thousands of terms, and when new versions of the nomenclature are regularly released, you need computational methods to check the internal consistency of the nomenclature. The classification is in XML, and this makes it easy to write a multi-purpose parsing script.

The Perl script (below) has three purposes:

1. It checks that neocl.xml is well-formed xml

2. It checks that a concept identifying code in one class is not repeated in any
other class within neocl.xml

3. It checks that no term in neocl.xml is ever repeated

On my 2.5 GHz computer, the xmlvocab.pl Perl script takes about 4 seconds to parse and check the 10+ Megabyte neocl.xml file. The script provides messages indicating any problem terms in the nomenclature.

#!/usr/bin/perl
#xmlvocab.pl
#
#This Perl script was created by Jules J. Berman and
#updated on 5/19/2005
#
#Copyright (c) 2005 Jules J. Berman
#
#Permission is granted to copy, distribute and/or
#modify this document
#under the terms of the GNU Free Documentation
#License, Version 1.2
#or any later version published by the Free
#Software Foundation;
#with no Invariant Sections, no Front-Cover Texts,
#and no Back-Cover Texts.
#
#The software is provided "as is", without warranty
#of any kind, express or implied, including but not
#limited to the warranties of merchantability,
#fitness for a particular purpose and
#noninfringement. in no event shall the authors
#or copyright holders be liable for any claim, damages
#or other liability, whether in an action of contract,
#tort or otherwise, arising from, out of or in connection
#with the software or the use or other dealings in the
#software.
#
#An explanation of the classification can be found in
#the following two publications, which should be cited
#in any publication or work that may result from any
#use of this file.
#
#Berman JJ. Tumor classification: molecular analysis
#meets Aristotle. BMC Cancer 4:8, 2004.
#
#neocl.xml is the classification of all neoplastic lesions.
#
use XML::Parser;
my $parser = XML::Parser->new( Handlers => {
Init => \&handle_doc_start,
Final => \&handle_doc_end,
});
$file = "neocl.xml";
$parser -> parsefile($file);

sub handle_doc_start
{
print "\nBeginning to parse $file now\n";
}

sub handle_doc_end
{
print "\nFinished. $file is a well-formed XML File.\n";
}

open (TEXT, $file);
#open (OUT,">neocl.out");
my $countcode = 0;
my $line = " ";
my %code;
my $classname;
my $phrasecount = 0;
while ($line ne "")
{
$line = <TEXT>;
next unless ($line =~ /\<.+\>/);
if ($line =~ /^\<([a-z\_]+)\>/)
}
$classname = $1;
next;
}
if ($line =~ /[CS]([0-9]{7})/)
{
$phrasecount++;
if (exists $code{$&})
{
if ($code{$&} ne $classname)
{
print "$& is a problem\n";
}
}
else
{
$code{$&} = $classname;
$countcode++;
}
}
}
close TEXT;
print "The total number of concepts is $countcode\n";
print "The total number of phrases is $phrasecount\n";
open (TEXT, $file);
undef %code;
$line = " ";
my %item;
while ($line ne "")
{
$line = <TEXT>;
if ($line =~ /([CS])([0-9]{7})/)
{
$prefix = $1;
$line =~ /\"\> ?(.+) ?\<\//;
$phrase = $prefix . $1;
if (exists $item{$phrase})
{
print $. . " More than one occurrence of \"$phrase\"\n";
}
$item{$phrase}="";
}
}
close TEXT;
exit;

-Jules J. Berman

In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

tags: common disease, orphan disease, orphan drugs, rare disease, subsets of disease, disease genetics, genetics of complex disease, genetics of common diseases, cryptic disease