Showing posts with label biomedical informatics. Show all posts
Showing posts with label biomedical informatics. Show all posts

Thursday, January 1, 2009

Updated and new files on neoplasm occurrences, by age

Happy New Year!

I've just uploaded a new version of my previously published file on the age distribution of occurrences for 626 different types of cancers.

http://www.julesberman.info/seerdist.pdf

This file is intended to be a resource for pathologists, epidemiologists and cancer researchers.

I've also uploaded a new file on cancers with multimodal age distributions (i.e., more than one peak in the age distribution for the neoplasm).

http://www.julesberman.info/bimode.pdf

I'll be discussing this file in the next several blog posts.

-© 2009 Jules Berman
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Saturday, July 5, 2008

Bomedicine in the Post-Information Age: 7 of 7

This is the seventh of 7 blogs on biomedicine in the post-information age.

In the post-information age, there is universal access to information, computational power, and the world-wide communications infrastructure.

The first point I've been trying to make in this series of 7 blogs is that the "work" of the post-information age is to derive meaning from our ubiquitous information. The second point I've been trying to make is that individuals, rather than institutions, are in the best position to make the most rapid and the most startling advances in this new age.

Throughout this series of blogs, I've mentioned data annotation, without explaining what I meant. In its simplest form, data annotation is adding metadata (data descriptors) to the data in a document. The purpose of adding metadata is to make the data specific. So a date can be an item on a calendar, or a social event, or a type of fruit. Metadata allows you to specify your intent.

If a date is a fruit, then it is a type of organism:

ID : 42345
PARENT ID : 4719
RANK : species
GC ID : 1
MGC ID : 1
SCIENTIFIC NAME : Phoenix dactylifera
GENBANK COMMON NAME : date palm
SYNONYM : Phoenix dactylifera L.
HIERARCHY
Phoenix dactylifera
Phoenix
Phoeniceae
Coryphoideae
Arecaceae
Arecales
commelinids
Liliopsida
Magnoliophyta
Spermatophyta
Euphyllophyta
Tracheophyta
Embryophyta
Streptophytina
Streptophyta
Viridiplantae
Eukaryota
cellular organisms

The date grows on a date tree (Phoenix dactylifera) and inherits the properties of its ancestors. The organism ancestry (phylogeny) of the data was obtained at my web page, by entering Phoenix dactylifera in the query box.

http://www.julesberman.info/post.htm

By using metadata that is specified in a classification or an ontology, we can use annotated data to draw inferences that are beyond the intent of the original document. By merging annotated documents (a product of the information age), and applying post-information age data analysis tools, we can achieve a great deal.

The new age starts with data specification.

- Copyright (C) 2008 Jules J. Berman

key words: semantics, semantic web, RDF, biomedical informatics
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Thursday, July 3, 2008

Biomedical Informatics "Search Inside"

Readers of this blog know that I published Biomedical Informatics (Copyright 2007) with Jones & Bartlett Publishers. This book has never been scanned by Google Books (for sample chapters), and was never given a "Search Inside" by Amazon. A "Search Inside" is a web site that features excerpts from the book, including the full table of contents.

Amazon has, at long last, produced a "Search Inside," for Biomedical Informatics. I hope that readers of this blog will visit the Amazon site.

- Jules Berman

Monday, June 30, 2008

Biomedicine in the Post-Information Age: 6

This is part six of a multi-part blog on biomedicine in the post-information age.

In the post-information age, solo experts will use three tools (universal access to information, computational power, and the world-wide communications infrastructure) to be innovative and productive, without being employed by bricks-and-mortar institutions.

What is an example of a post-information age innovation? I'll give you an example from my personal experience.

Governments advocate the development of a standard biomedical vocabulary that will put an end to the profusion of non-standard vocabularies that are used to annotate biomedical tests. Text annotation (sometimes called text coding) is a necessary step for data retrieval, indexing, classification, integration, etc.

So, instead of having lots of separate nomenclatures, the governments prefer a single , standard nomenclature that everyone uses. In the U.S., England and much of Europe, there is a push to use SNOMED-CT as the standard medical vocabulary.

There is one problem with this. It has proven impossible to build a single nomenclature that includes all of the terminology used in specialty domains. A specialist in the domain of dermatologic diseases (in which there are many thousands of obscure diseases and multiple synonyms for individual diseases, and very little biological research to relate these diseases with other skin diseases or with systemic diseases) is unlikely to be satisfied with a general disease nomenclature.

In my area of specialty (tumor biology), this is also true. I found the standard nomenclatures (ICD-O, SNOMED-CT, NCI Thesaurus, UMLS metathesaurus) to have only a small number of the neoplasms that can be found in the biomedical literature. In addition, the relationships among the different neoplasms were, in my opinion, not adequately expressed in these standard nomenclatures.

So, I built my own specialty nomenclature, the Developmental Classification and Taxonomy of Neoplasms (usally called the Neoplasm Classification), with includes its own biological hierarchy of neoplasms and which has about ten-fold the number of neoplasm terms as the standard nomenclatures.

Anyone who wants to use a comprehensive, biologically classified list of neoplasms, is welcome to use the nomenclature that I, as a post-information age solo expert, developed. This is an open source document available in gzipped XML format at:

http://www.julesberman.info/neoclxml.gz

Or in zipped XML format at:

http://www.julesberman.info/neoclxml.zip

Do you need to abandon the standard nomenclature? No. Use both. I have written extensively on autocoding, double autocoding (autocoding with two or more nomenclatures), re-coding (autocoding again and again to satisfy the requirements of a particular project), and on-the-fly coding. My papers are linked to full-text articles on my publications page.

This is just one example showing how an post-information age individual can contribute in areas that large groups and institutions have ignored.

- Copyright (C) 2008 Jules J. Berman

key words: biomedical informatics, medical informatics, health coverage, health insurance, medical insurance
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Sunday, June 29, 2008

Biomedicine in the Post-Information Age: 5

This is part five of a multi-part blog on biomedicine in the post-information age.

As the prior blogs in this series emphasized, the distinctive feature of the post-information age is that everyone has personal access to compuational power, communications, and information. In the post-information age, individuals will use these empowering tools to be innovative and productive, without being employed by bricks-and-mortar institutions.

Who gets to be a player in the post-information age?

In the U.S., the lingering impediment to being a solo information expert is medical insurance. Here, health insurance coverage is something that is usually received through employment. Individuals who are not part of an empoyer's group can be denied health insurance by the insurance agencies, for almost any reason. Because medical care is extravagantly expensive in the U.S., it is very important to have a health insurance provider. Many people hold onto unrewarding jobs, just for the available health insurance (for themselves and their families).

It's really an enormous waste of potential talent, because many of the opportunities for innovation are best accomplished by small groups of experts (maybe 1, 2 or 3 people) who might be geographically dispersed. Contries that guarantee health care to their citizens (e.g., the EU), will have an enormous advantage in the new post-information age.

- Copyright (C) 2008 Jules J. Berman

key words: biomedical informatics, medical informatics, health coverage, health insurance, medical insurance
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Saturday, June 28, 2008

Biomedicine in the Post-Information Age: 4

This is part four of a multi-part blog on biomedicine in the post-information age.

To continue from the prior posts, in the post-information age (when everyone has instant access to enormous amounts of information), many services will be rendered by solo experts, who are not employees of bricks-and-mortar institutions, but who are contracted, as needed, to work on specific projects.

These solo contractors will write on-demand software utilities (very short turn-around time), annotate large databases (so that they can be integrated with other databases), check work done by the staff of bricks-and-mortar institutions (for mistakes and weaknesses, for conformity to some specialized standard), munge non-standard data into any of many standards, do literature/data research, assist in writing grants and proposals, etc.

I predict that there will be big changes in the book publishing industry as solo experts begin writing/illustrating/publishing/marketing/distributing their own books. Basically, we're just waiting for someone to market an inexpensive, convenient high-quality ebook reader. Once we get a good ebook reader, individuals will find that they have the expertise to manage every facet of the book industry. Traditional publishers will have a major role in this post-information age enterprise, only if they are willing to change with the times and use their established marketing and production skills to create innovative books that utilize all of the available facilities of the ebook medium (including seamless links between books and other media). If all goes well, the public will benefit from wonderful books that offer a remarkably exciting (and educational) reading experience.

In the post-information age, turn-around for all sorts of services will be shortened. Users will not tolerate a procrastinating work-force. Solo contractors will be valued for their rapid turn-around and high accuracy.

The skill sets of the solo contractor will change. They will be information integrators, and they will be expected to master ontologies (and the syntax for ontologies, RDF), several high-level programming languages (e.g., Perl, Python, Ruby), collective intelligence tools, web services, at least one highly specialized data domain (such as molecular biology, biomedical imaging, or genetics), and the legal/ethical aspects of their services.

- Copyright (C) 2008 Jules J. Berman

key words: biomedical informatics, medical informatics, big data, metadata, data preparation, data analytics, data repurposing, datamining, data mining
My book, Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information was published in 2013 by Morgan Kaufmann.



I urge you to explore my book. Google books has prepared a generous preview of the book contents.

Friday, June 27, 2008

Biomedicine in the Post-Information Age: 3

This is part three of a multi-part blog on biomedicine in the post-information age.

I confess that my core concept of the post-information age comes from Blade Runner. In that movie, there were large bricks-and-mortar corporations that controlled much of the world's industry (constructing dangerous androids, conducting off-world mining operations, etc.), and there were solo contractors who worked in the back rooms of antique stores and noodle shops and who made the android skin or nano-bots or psionic circuits and what-not for the replicants and the off-world mining operations, etc. The idea is that the large companies dealt with the macro-economics but that the little guys had the specialized expertise that was used by the large corporations.

This is how I think of the post-information age. Big corporations like Microsoft and Google will dominate the information world, but highly trained free-lancers will do some of their most specialized work. In the biomedical world, large academic universities and federal and private funding agencies will spearhead huge initiatives (hundreds of millions of dollars), but the most fastidious work will be done by free-lancers.

Why is this? Why won't specialized work be done in-house? When large corporations hire, they are looking for people with a generalized skill-set that is appropriate for the activities of a department. So a Department of Surgery hires lots of surgeons. They may even hire an information officer or two. But they will never be in a position to hire (and keep) someone with all of the computational skills needed for a complex project that collects clinical data and integrates it with biomedical data from heterogeneous sources. It just makes sense to identify one of the few people in the world with the needed skills and have that person help out, when the need arises, for a negotiated fee.

In the post-information age, everyone has access to computers and software and lots of people have access to information. In the case of biomedicine, this information would be public biological databases, and de-identified medical databases, and associated ontologies, nomenclatures and classifications that help integrate all the data. The free-lancers would be hired to add value to or make sense of the data or write software to handle some specific purpose, that sort of thing. In the next few blogs I'll provide some examples.

- Copyright (C) 2008 Jules J. Berman

key words: biomedical informatics, medical informatics, common disease, orphan disease, orphan drugs, genetics of disease, disease genetics, rules of disease biology, rare disease, pathology
In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site.

Thursday, June 26, 2008

Biomedicine in the Post-Information Age: 2

In yesterday's blog , I began a series on the post-information age of biomedicine.

In the post-information age, everyone is empowered with lots of information, as well as the hardware and software tools to use the information.

This means that there will be less dependence on bricks-and-mortar institutions to carry on research, development, and entrepreneurial ventures. People can do an awful lot from their homes, or from nearly any location on the planet.

My guess is that we will see a growing workforce of talented, free-lance technologists who make enormous contributions to biomedical research in the post-information age. These individuals will come from two groups:

1. The recent college grads, who are technology-enabled and who developed a group of collaborators through social networking sites while they were in college.

2. Retirees, who bring their technical expertise with them into their retirements and who are fully capable of leading technologically productive lives from their homes.

Just about everyone one else (i.e., age 30 to 60) is fully invested in the bricks-and-mortar paradigm. They're dependent on regular pay checks, and on the family health coverage provided by their employers. They are not sufficiently secure, financially or medically, to leave their jobs to begin a new life as free-lancers.

Tomorrow, I'll discuss the kinds of projects that can be done "from home" in the post-information age.

- Copyright (C) 2008 Jules J. Berman

key words: biomedical informatics, medical informatics
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Wednesday, June 25, 2008

Biomedicine in the Post-Information Age: 1

Apparently, we have entered the post-information age. I've been doing a little research on "post-information", and I'm not sure that a good definition exists.

My impression is that you enter a "post-fill in the blank" age when the basic tools and principles of an age are completed and available. In the "post" age, the world develops new and useful advances from pre-existing tools.

So, for example, the industrial revolution involved developing machinery and engines that could perform some of the difficult chores that humans were struggling with: ginning cotton, moving goods from the Eastern states to the Western territories. The post-industrial age involved using the principles of machine design to clever things, like visiting the moon.

Another example is the genomic and post-genomic ages. The age of the genome was devoted to sequencing the bases in human DNA. Once the human genome was sequenced, the post-genomic age began. Now, we're expected to use our knowledge of genes to cure cancer and halt the aging process.

The information age was focused on building powerful, fast, and affordable computers and to develop computational strategies for collecting, storing, accessing, exchanging, annotating, and analyzing huge amounts of information. We've done that, and we've used computers and software to do many of the tedious tasks that were once done "by hand." So now we're in the post-information age. Now, we're expected to use all that information to do completely new things; things that were not envsioned during the information age.

Over the next few days or weeks, I hope to write a few blogs about what we might expect from the post-information age. As usual, I will tie everything to my favorite subjects (data annotation, classification, new methods of data analysis, and biomedical progress).

- Copyright (C) 2008 Jules J. Berman

key words: biomedical informatics, post-genomic
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Thursday, June 19, 2008

Biomedical Informatics Book

I just visited (6/19/08) my book's Amazon page to see how sales were doing, and I found that Amazon has reduced its price. They're currently selling it for $50.91 (a 32% savings) and free shipping. This is pretty good, because they usually sell it with no reduction, or with a negligible reduction.

If you're curious about Biomedical Informatics, here is a list of contents.

0. Preface.

1. What is biomedical data, and what do we do with it?

1.1. Background.

1.2. The challenge of translational research.

1.3. Disasters in translational research.

1.4. The role of biomedical data in translational research.

1.5. Expertise in biomedical informatics.

1.6. The good news: no-cost tools.

1.7. The bad news: the high-cost of human cooperation.

1.8. Realistic opportunities for biomedical informaticians.

2. The data of biomedical informatics.

2.1. Background.

2.2. Data files and databases.

2.3. Medical databases and hospital information systems.

2.4. Every patient must be uniquely identified within the system.

2.5. All data entered should be retrievable.

2.6. Entered data should only be modified with great caution.

2.7. The government as a source of biomedical data.

2.8. Your right to obtain government data - freedom of information act.

2.9. Access to research data discovered under u.s. grants.

2.10. Grantees strike back: the u.s. bayh-dole act.

2.11. Intellectual property.

2.12. Fair use and other academic privileges.

2.13. Madey v duke and the erosion of academic privilege.

2.14. Further cautions on the use of proprietary software and data.

2.15. The often misunderstood concept of patient data "ownership".

2.16. Sharing data.

2.17. Legacy data.

2.18. Free, open source and proprietary software and data.

2.19. Undifferentiated software.

2.20. What are some of the open access biomedical databases?

2.21. Open access medical terminologies.

2.22. Mesh (the national library of medicine's medical subject headings).

2.23. Taxonomy.

2.24. Disease data and epidemiologic data.

2.25. The impact of free and open source data and software on biomedical informatics.

3. Confidential biomedical data.

3.1. Background.

3.2. Human subject risks.

3.3. The risk to life and health as a direct result of a medical intervention.

3.4. The risk of loss of database functionality.

3.5. The differences between confidentiality and privacy.

3.6. Example: loss of privacy resulting from participation in a medical study.

3.7. Loss of confidentiality.

3.8. The responsibilities of biomedical informaticians to human subjects.

3.9. Patient record anonymization.

3.10. Patient record de-identification.

3.11. An example of the law of unintended consequences.

3.12. Violations against the common rule (in u.s.).

3.13. Violations against hipaa (in u.s.).

3.14. Tort and violations against individuals.

3.15. What consents does the patient have on record?

3.16. Consented versus unconsented human subject research.

4. Standards for biomedical data.

4.1. Background.

4.2. The criticality of common standards.

4.3. The non-role of government (in the u.s.) in standards-making.

4.4. The hazards of creating a new standard.

4.5. Overview of standards development.

4.6. How are standards developed, approved and adopted?

4.7. The utility of non-standards.

4.8. The non-standard present - specifications and unique objects.

4.9. The non-standard future - data semantics.

4.10. Unique object identifiers.

4.11. Life science unique identifiers.

4.12. Hl7 unique identifiers.

4.13. Unique problems associated with uniqueness.

4.14. Specifying information: do you have the time?

4.15. Introduction to meaning.

5. Just enough programming.

5.1. Background.

5.2. Why you should learn some fundamental programming.

5.3. Just enough perl.

5.4. Downloading perl.

5.5. File operations.

5.6. Perl script basics.

5.7. The directory path to perl.

5.8. Accessing files.

5.9. The open1.pl script, line by line.

5.10. An 8-line perl word processor.

5.11. Don't panic! perl will forgive you.

5.12. Pseudocode for a general biomedical informatics program.

5.13. Interactively reading lines from a file.

5.14. Scanning enormous files quickly.

5.15. Getting just what you want with perl regular expressions.

5.16. Pseudocode for common uses of regex (regular expression pattern matching).

5.17. Regular expression syntax.

5.18. Removing periods that do not delineate sentences.

5.19. Counting all the words in a text file.

5.20. Finding the frequency of occurrence of each word in a text file (zipf distribution).

5.21. Creating a persistent database object.

5.22. Retrieving information from a persistent database object.

5.23. Validating xml tags using regular expressions.

5.24. What have we learned?

6. Programming common biomedical informatics tasks.

6.1. Background.

6.2. Computing a one-way hash for a word, phrase or file.

6.3. Simple statistics.

6.4. Invoking statistical tests through perl modules.

6.5. Avoiding type 4 errors with resampling.

6.6. Using random numbers.

6.7. Resampling and monte carlo statistics.

6.8. How often can i have a bad day?

6.9. Rough test of the built-in random number generator.

6.10. The monty hall problem: solving what we cannot grasp.

6.11. Internal and external math modules for perl.

6.12. Using external modules - fast fourier transform.

6.13. Indexing text.

6.14. Searching large text files.

6.15. Finding needles fast using a binary-tree search of the haystack.

6.16. Clustering: algorithms that group similar objects.

6.17. Retrieving information from the internet.

6.18. Gene sequence parsing: finding palindromes in a gene database.

6.19. Why counting is non-trivial and important.

6.20. Why you should write your own counting programs.

6.21. Software utilities versus software applications.

6.22. Software evaluation.

7. Biomedical nomenclatures.

7.1. Background.

7.2. Big nomenclatures and small nomenclatures.

7.3. Curating nomenclatures.

7.4. Automatic expansion of a medical nomenclature.

8. Misbehaving text: dealing with poorly written medical text.

8.1. Background.

8.2. Spelling errors.

8.3. Homonymous terms.

8.4. Abbreviations that are sometimes both acronyms and shortened forms.

8.5. Prepositions and articles retained in an acronym.

8.6. Single expansions with multiple abbreviations.

8.7. Nonsense abbreviations.

8.8. Common usage that confounds meaning.

8.10. Pejorative abbreviations.

8.11. Locale-dependent abbreviations.

8.12. Classifying abbreviations by their expansion algorithms.

8.13. Ephemeral abbreviations.

8.14. Hyponymous abbreviations.

8.15. Polysemous abbreviations.

8.16. Abbreviations masquerading as words.

8.17. Fatal abbreviations: innocent victims of abbreviation drift.

8.18. Forbidden abbreviations.

9. Autocoding unstructured data (narrative ext).

9.1. Background.

9.2. Machine translation.

9.3. Autocoding.

9.4. Human fallibility and the limitations of human-collected data.

9.5. A fast lexical autocoder.

9.6. Evaluating autocoders: dealing with precision and recall.

9.7. Other performance issues.

9.8. On-the-fly coded data retrieval without pre-coding.

9.9. Different philosophical approaches to term-based data retrieval.

9.10. Why it is important to have fast autocoding software.

10. Computational methods for de-identification and data scrubbing.

10.1. Background.

10.2. Anonymization, de-identification, data scrubbing.

10.3. Identifiers.

10.4. Stripping identifiers.

10.5. How good is good enough?

10.6. Scrubbing data.

10.7. De-identification algorithms.

10.8. Feasibility of de-identification.

10.9. Non-uniqueness and de-identification.

10.10. Leveraging some confidential information to learn more confidential information.

10.11. Performance considerations for de-identification software.

10.12. De-identification and data sharing patents.

11. Cryptography in biomedical informatics.

11.1. Background.

11.2. One-way hashing algorithms.

11.3. One-way hash weaknesses: dictionary attacks and collisions.

11.4. Zero-knowledge patient reconciliation.

11.5. Threshold protocol.

11.6. Electronic signatures.

12. Describing data with metadata.

12.1. Background.

12.2. Metadata, xml (extensible markup language) and rdf (resource description framework).

12.3. Enforced and defined structure (xml rules and schemas).

12.4. Formal metadata (through the iso11179 specification).

12.5. Namespaces (sharing metadata).

12.6. Linking data via the internet.

12.7. Logic and meaning.

12.8. Self-awareness (embedded protocols and commands).

12.9. Integrating heterogeneous data with rdf.

12.10. Meaning requires a fully-specified subject.

12.11. Meaningfully biomedical description with notation 3.

12.12. The daml extension of rdf .

12.13. Owl extension of daml.

13. Simplifying complex data with classifications and ontologies.

13.1. Background.

13.2. The value of hospital information technology.

13.3. Understanding complexity.

13.4. The importance of data simplification.

13.5. Example case: a molecular classification of cancer.

13.6. Cancer nomenclatures, taxonomies, classifications and ontologies.

13.7. Practical limitations of classifications.

13.8. Ontologies: multi-class inheritance and logical inferences.

13.9. Go, the gene ontology that is not an ontology.

14. Clinical trials: the informatician lives in a statistical world.

14.1. Background.

14.2. Do we need clinical trials?

14.3. The length and expense of clinical trials.

14.4. An imaginary clinical trial.

14.5. Modeling a clinical trial.

14.6. What do models tell us?

14.7. The informatics of clinical trials.
14.8. Clinical trials need to be validated by post-trial experience.

15. Distributed computing.

15.1. Background.

15.2. Remote procedure calls, soap, web services and grid computing.

15.3. Data utopia.

15.4. Data dystopia.

16. Grantsmanship for biomedical informaticians.

16.1. Background.

16.2. Institutional risks from biomedical informatics research.

16.3. Funders' risks from biomedical informatics research.

16.4. Suggestions for biomedical informaticians who write grant applications.

17. A practical approach to ethics for biomedical informaticians.

17.1. Background.

17.2. Is it ever ok to lie?

17.3. When can you use unconsented identified medical records?

17.4. When can you use proprietary software and standards?

17.5. When is it ok to have conflicts of interest?

17.6. When is it ok to refuse consent?

17.7. Is it ethical to patent biomedical discoveries?

17.8. The etiquette of free software usage.

17.9. Hoarding research data.

17.10. Are there ethical alternates to hipaa's safe harbor de-identification method?

17.11. Can you use consented data for unconsented research?

17.12. When is it ethical to enforce copyright medical research publications?

17.13. Is it ok to profit from tissue banking services?

17.14. How likely is a hipaa lawsuit?

17.15. Being fair to the outraged patient.

17.16. When can i be wrong?

17.17. Closing platitudes.

18. References (commented).

19. Appendix.

19.1. The c programming language.

19.2. The java programming language.

19.3. Perl, open source programming language.

19.4. Python, open source programming language.

19.5. Ruby, open source object oriented programming language.

19.6. Swig, open source glue tool.

19.7. Open microscopy environment (ome).

19.8. R open source statistical programming language and bioconductor.

19.9. Open source bioperl, biopython, bioruby.

19.10. Open source electronic laboratory notebook, neurosys.

19.11. Open source gimp image software.

19.12. Open source nih image.

19.13. Pov-ray image rendering open source software.

19.14. Open source compression and archiving utilities (gzip, gunzip, tar, 7-zip, bunzip).

19.15. Cygwin, open source unix/linux emulator.

19.16. Gnupg, open source encryption tool.

19.17. Wget web site mirroring software.

19.18. Open source indexing software (swishe-e and lucene).

19.19. Open source wordprocessing software (abiword and openoffice writer).

19.20. Open source emacs text editor.

19.21. Open source spreadsheet software.

19.22. Open source presentation software.

19.23. Mumps, an ansi standard programming language for medical informatics.

19.24. MySQL, open source database software.

19.25. Protege, open source ontology editor.

19.26. Vista, a free hospital information system courtesy of the u.s. government.

19.27. CWM, a closed world machine for rdf (in python).

19.28. Pubmed and pubmed central.

19.29. Resources from the national center for biotechnology information.

19.30. Database issue of nucleic acids research.

19.31. Locuslink and its successor, entrez gene.

19.32. Time stamping.

19.33. Google, as if you didn't already know.

19.34. Sourceforge.

19.35. CVS, concurrent versions system.

19.36. Cpan, the comprehensive perl archive network.

19.37. Requests for comment.

19.38. Omim - online mendelian inheritance in man.

19.39. Loinc, logical observations identifiers, names, and codes.

19.40. HL7 - health level 7.

19.41. Seer.

19.42. U,LS metathesaurus.

19.43. Medical subject headings - mesh.

19.44. Gene ontology - GO .

19.45. OBO (open biology ontologies).

19.46. Ushik metadata registry.

19.47. Neoplasm classification.

19.48. US census.

20. Glossary.

21. List of lists.

22. Index.

23. Author biography.

More book information is available from the publisher's web site.

-Jules Berman

key words: medical informatics, bioinformatics, Perl programming, biomedical data, medical confidentiality, medical privacy, hipaa, big data, metadata, data preparation, data analytics, data repurposing, datamining, data mining
My book, Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information was published in 2013 by Morgan Kaufmann.



I urge you to explore my book. Google books has prepared a generous preview of the book contents.

Thursday, January 31, 2008

More on "iform" words in UMLS

Here is the Perl script that extracted the "iform" words from UMLS:


#!/usr/local/bin/perl
open (TEXT, "MRCONSO");
$line = " ";
while ($line ne "")
{
$line = <TEXT>;
next if ($line !~ /ENG/);
if ($line =~ /\b[a-z]+iform\b/i)
{
$term = lc($&);
$subhash{$term}++;
}
}
foreach $key (sort keys %subhash)
{
print "$key\n";
}
exit;


The MRCONSO file (previously called the Mr. Con file) is the large (greater than 800 Megabyte) UMLS Metathesaurus file that contains all of the metathesaurus terms. It is available free from the U.S. National Library of Medicine, but you need to register and complete an online license agreement before they will release the metathesaurus files to you.

The Perl script (above) can be easily modified for simple extraction projects. If you're interested in learning Perl to help you with biomedical projects, you might want to read my book, Methods in Medical Informatics: Fundamentals of Healthcare Programming in Perl, Python, and Ruby (Chapman & Hall/CRC Mathematical and Computational Biology).

Here is the complete list of "iform" words:

acneiform
acniform
ansiform
apoplectiform
arciform
bacilliform
canaliform
cerebriform
chancriform
choreiform
chyliform
coliform
coralliform
cordiform
cribiform
cribriform
cruciform
cuciform
cuneiform
cupuliform
curariform
dendriform
dermiform
disciform
emboliform
epileptiform
equiform
falciform
fetiform
filariform
filiform
filliform
flagelliform
framboesiform
fundiform
fungiform
fusiform
gadiform
gelatiniform
gigantiform
gyriform
herpetiform
hydatidiform
ichthyosiform
intercuneiform
juxtarestiform
kaposiform
lentiform
licheniform
moniliform
morbilliform
multiform
myrtiform
naviculocuneiform
neuralgiform
nonhydatidiform
noviform
pampiniform
pectiniform
perciform
piriform
pisiform
pityriasiform
plexiform
prepiriform
prepyriform
proteiform
psoriasiform
punctiform
pyriform
rediform
reniform
restiform
retiform
retrolentiform
rhabditiform
rubelliform
sacciform
scarlatiniform
scarletiniform
schizophreniform
sclerodermiform
scorpaeniform
spongiform
storiform
subcuneiform
sublentiform
unciform
uniciform
uniform
valpiform
varicelliform
varioliform
vermiform
verruciform
vitelliform
zosteriform

In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D. tags: common disease, orphan disease, orphan drugs, rare disease, disease genetics, biomedical informatics, perl programming

Monday, January 14, 2008

Parsable Doublets List now available in public domain

Word doublets are two-word phrases that appear in text (i.e., they are not randomly chosen two-word sequences.

Doublets can be used in a variety of informatics projects: indexing, data scrubbing, nomenclature curation, etc. Over the next few days, I will provide examples of doublet-based informatics projects.

A list of over 200,000 word doublets is available for download.

The list was generated from a large narrative pathology text. Thus, the doublets included here would be particularly suitable for informatics projects involving surgical pathology reports, autopsy reports, pathology papers and books, and so on.

The Perl script that generated the list of doublets by parsing through a text file ("pathold.txt"), is shown:


#!/usr/local/bin/perl
open(TEXT,"pathold.txt")||die"cannot";
open(OUT,">doublets.txt")||die"cannot";
undef($/);
$var = <TEXT>;
$var =~ s/\n/ /g;
$var =~ s/\'s//g;
$var =~ tr/a-zA-Z\'\- //cd;
@words = split(/ +/, $var);
foreach $thing (@words)
{
$doublet = "$oldthing $thing";
if ($doublet =~ /^[a-z]+ [a-z]+$/)
{
$doublethash{$doublet}="";
}
$oldthing = $thing;
}
close TEXT;
@wordarray = sort(keys(%doublethash));
print OUT join("\n",@wordarray);
close OUT;
exit;


You can generate your own list by substituting any text file you like for "pathold.txt". Keep in mind that the Perl script slurps the entire text file into a string variable, so the script won't work if you use a file that exceeds the memory of the computer. For most computers (with RAM memories that exceed 256 MBytes) this will not be a problem. On my computer (about 2.8 GHz and 512 Mbyte RAM) the script takes about 5 seconds to parse a 9 Megabyte text file).

Since the doublet list below consists of a non-narrative collection of words, it cannot be copyrighted (i.e., it is distributed as a public domain file).

-Jules Berman
My book, Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information was published in 2013 by Morgan Kaufmann.



I urge you to explore my book. Google books has prepared a generous preview of the book contents.

tags: big data, metadata, data preparation, data analytics, data repurposing, datamining, data mining, biomedical informatics, curation, data scrubbing, deidentification, medical nomenclature, Perl script, public domain, doublets list

Friday, January 4, 2008

Google preview of Perl Programming for Medicine and Biology

"Google Books" has chosen my work, Perl Programming for Medicine and Biology (2007), for limited preview.

You can browse the table of contents and read selected excerpts from the book. This comes about as close as in in-store browse as anyone might want, and I'm very grateful that Google provides this service. I have no idea how they choose which books get previewed, but I wish they would do it for every book-in-print.

-Jules Berman

Tuesday, November 6, 2007

Ruby climbs up a notch on TIOBE index

On my September 14 blog, I wrote that Ruby had risen to a rank of 10 on the TIOBE index of programming languages.

On the November, 2007 TIOBE index, Ruby climbed another notch to number 9.

According to their web site, the TIOBE index ratings "are based on the world-wide availability of skilled engineers, courses and third party vendors."

This rating system seems to overlook one of the strongest features of Ruby: its appeal to non-programmers. Ruby is a language that unskilled engineers can use productively. One of the themes of my Specified Life blog site is that biomedical research requires a little bit of programming (usually less than a dozen lines of Ruby code) to perform common computational tasks related to data organization, data sharing and data analysis. Ruby empowers non-programmers (people who do not identify themselves as programmers) to perform the common computational tasks in their fields.

Still, it is re-assuring to know that Ruby is moving up as a language used by professional programmers.

- Jules Berman


Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.


Saturday, October 27, 2007

National Cancer Institute Thesaurus

The National Cancer Institute (NCI) Thesaurus is a free medical vocabulary available in OWL format from:

ftp://ftp1.nci.nih.gov/pub/cacore/EVS/NCI_Thesaurus/


It's really quite an impressive document, and there are very few standardized vocabularies that have been prepared as formal ontologies. The creators wisely used the semantics of OWL (Web Ontology Language), a dialect of RDF.

The NCI thesaurus contains terms related to the interests of the NCI and contains the names of many neoplasms.

This vocabulary has been curated for over a decade by in-house ontologists (NCI employees), contractors, and through the use of domain consultants (including some pathologists). It is updated monthly. A lot of money has gone into the development of the NCI Thesaurus, and it is one of the most worked-on vocabularies in the medical field.

The NCI Thesaurus has been reviewed by Barry Smith and colleagues, who found it somewhat lacking.

http://ontology.buffalo.edu/medo/NCIT.pdf


"RESULTS: We found many mistakes and inconsistencies
with respect to the term-formation principles used,
the underlying knowledge representation system,
and missing or inappropriately assigned verbal and
formal definitions.."
Ceusters W, Smith B, Goldberg L.
A terminological and ontological analysis of the
NCI Thesaurus. Methods Inf Med. 2005;44(4):498-507.

My question is, "If the Thesaurus contains many different knowledge domains (medications, general diseases, neoplasms, etc.) how can it adequately cover all of its constituent domains?" In the neoplasm domain, it is missing many thousands of names of neoplasms. The terminology may be sufficient for its intended purpose (meeting the needs of the NCI community), but because the terminology is not comprehensive, the NCI Thesaurus will not necessarily serve those who want a thesaurus that comes close to including the names of ALL neoplasms.

Also, there doesn't seem to be any single organizing principle for the neoplasm domain. Some neoplasms are subclassed by their anatomic site (e.g. urinary tract neoplasm). Others are subclassed by their tissue type (e.g. soft tissue neoplasm). And so on. This is allowable under an ontology, so long as the ontology maintains consistency and competence (ability to answer questions about the members of classes). But I wonder if this is the best way of organizing tumors. Of course, I'm deeply biased. The Developmental Lineage Classification and Taxonomy of Neoplasms has a single organizing principle.

The NCI Thesaurus is an impressive piece of work and definitely worth looking over.

tags: biomedical informatics, cancer, classification, nomenclature, thesaurus, vocabulary, ontology, rare diseases, orphan drugs, genetics of disease, pathology, common diseases, complex diseases

In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

- Jules J. Berman, Ph.D., M.D.

Wednesday, October 17, 2007

Updates to my web site

I've recently made updates to several old files on my website.

They are:

2003 Letter to Human Pathology regarding precancers

2004 Letter to Human Pathology regarding precancers

2004 Editorial to Am J Clinical Pathology regarding data sharing in pathology

2001 List of 12,000+ medical abbreviations

2003 Neoplasm Classification gzipped file

- Jules Berman
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Tuesday, October 9, 2007

Ruby scripts from Ruby Programming for Medicine and Biology

For those interested, here is a list of Ruby scripts that are included and described in Ruby programming for Medicine and Biology, my book that was published last month (September, 2007).

SCRIPTS IN RUBY PROGRAMMING FOR BIOLOGY AND MEDICINE

1.3.1. Getinput.rb retrieves a line of keyboarded text.

1.4.1. Grow.rb simulates six generations of bacterial growth.

1.4.3. Grow2.rb uses explicit Ruby objects and statements.

1.5.1. Lowclass.rb provides syntax for user-created classes.

1.5.3. Noclass.rb requires an external Person class definition.

1.5.4. Person_class_file.rb, a class library for script no_class.rb.

2.7.1. Combo.rb parses an array into all possible ordered subarrays.

2.8.1. Hash.rb creates and displays key/value pairs for a Hash instance object.

2.8.4. Neohash.rb creates three hash instances for the Neoplasm Classification.

2.10.1. Glob.rb demonstrated the Dir class glob method.

2.10.3. Dirlist.rb lists the files in the current directory.

2.12.1. Time.rb measures the length of time for any process.

3.4.2. Mod1.rb defines and includes a simple module.

3.4.3. Mod2.rb calls a module with the scope operator.

3.4.4. Mod3.rb embeds a module within a class.

6.2.1. Readsome.rb reads the first 20 lines of the MRCONSO file.

6.3.2. Zipf.rb prints the number of occurrences of words in a string.

6.3.4. Zipf2.rb creates a Zipf distribution of the words in OMIM.

6.4.1. Snom_get.rb extracts SNOMED-CT terms from UMLS.

6.5.1. Disease.rb collects SNOMED-CT diseases from UMLS.

6.6.1. Neosdbm.rb creates three persistent database objects.

6.7.1. Sdbmget.rb retrieves data from persistent object.

7.3.1. Sentence.rb, a simple sentence parser.

7.6.1. Search.rb searches through any file for lines matching a Regex expression.

7.7.1. Pubemail.rb extracts e-mail addresses from a PubMed search.

8.3.1. Base64.rb encodes strings in Base64 notation.

8.4.1. Dircopy.rb copies files from one directory to another.

8.5.1. Dcm2jpg.rb converts a DICOM file into a jpeg file.

8.5.3. Dcmsplit.rb converts a DICOM file into a jpeg and a text file.

8.6.1. Jpeg_add.rb inserts textual information into a jpeg image.

9.2.1. Concord.rb creates a concordance.

9.3.1. Indexer.rb creates an index.

10.2.1. Haystack.rb performs a binary search on a file.

10.3.2. Tinysort.rb sorts the lines of a file.

10.3.4. Bigsort.rb sorts large files quickly.

10.4.1. Anatomy.rb extracts SQL data from the Functional Model of Anatomy.

10.5.2. Alldata.rb sums the census districts to yield the total U.S. population.

11.2.1. Scrubit.rb scrubs and deidentifies any input line.

12.2.2. Autocode.rb provides nomenclature terms and codes for an input sentence.

12.3.1. Fastcode.rb improves performance compared with aucode.rb.

12.4.3. Icd.rb collects ICD10AM codes from the UMLS Metathesaurus.

12.5.1. Seer.rb determines the occurrences in the U.S. of tumor types found in the SEER public-use data files.

13.5.1. Fibo.rb computes the first twenty elements of the Fibonacci series.

13.6.1. Mean.rb computes the mean from an array of numbers.

13.7.1. Std_dev.rb computes the standard deviation for an array of numbers.

13.8.1. Randtest.rb simulates 600,000 casts of the die.

13.10.1. Error.rb uses resampling to simulate runs of errors.

14.6.2. Thresh.rb divides a text file into two threshold files.

14.6.5. Threshrv.rb computes original file from two threshold files.

15.2.2. Neopull.rb searches a server file for a web client query.

15.3.1. Neosafe.rb improves the security of neopull.rb.

17.3.1. Biohack.rb, converts a gene sequence into a protein sequence.

18.14.2. Rdf3.rb extracts triples from an RDF document.

18.19.2. Jpg2b64 inserts a Base64 jpeg image into an RDF file.

19.2.2. Ancestor.rb determines the ancestor lineage for organisms.

19.2.4. Class_lineage.rb determines ancestral lineage.

19.3.1. Neoself.rb provides tag hierarchy for XML file.

20.5.1. Conflict.rb overrides a class assignment.

-Jules Berman tags: bioinformatics, biomedical informatics, medical informatics, Ruby, ruby language, Ruby programming, scripts

Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.


Friday, June 15, 2007

Ruby Programming

I've recently written a book, "Ruby Programming for Medicine and Biology." The Table of Contents is available at the Jones and Bartlett Web Site.

In my opinion, it's important that biomedical informaticians become self-sufficient and less reliant on vendor-supplied applications. The simple act of writing your own programs is an empowering experience and permits us to develop and try new ideas, something that would not be feasible with commercial software.

For a long time, I've been an advocate of Perl (see my book).

Perl is very good when you want to do imperative programming (sometimes called procedural programming). Basically, in imperative programming, the program consists of the implementation of an algorithm in the syntax of the programming language. Each line of the program is another command that executes a step in the algorithm. Procedural programming is virtually the same as imperative programming. The only difference is that in procedural programming, a step in the algorithm may involve calling an external method (i.e., another algorithm). You can think of procedural programs as imperative programs with subroutines. This is what Perl does very well. Because it's easy to learn Perl syntax and because the built-in Perl commands and the available Perl modules provide most of the functionality that anyone would need in the biomedical field, Perl has become a very popular language among bioinformaticians.

The problem with Perl is that it is not well suited as a language that models and integrates biomedical classifications and ontologies. This last jargon-heavy sentence deserves a little explanation, but you probably don't need the standard essay on the data-intensive aspects of modern biomedicine. Suffice it to say that when you have lots and lots of complex data, you need some way to simplify the data and to relate one kind of data to other kinds of data. The best way to simplify data is with classifications or ontologies that can annotate data in a manner that everyone can understand and exchange. When you talk about classifications and ontologies, you're talking about data objects, object (instance) methods, class methods, inheritance, metadata descriptions, specifications, on and on. These are the things that object oriented languages provide.

Ruby is a great object oriented language because it is free, open source, has a very simple and logical syntax, and gracefully models existing biomedical classifications and ontologies. I tried using object-oriented Perl for my work with classifications and ontologies, but it just was not a good fit. I dabbled in Python (an excellent object-oriented programming language that has many of the features I was seeking), but it lacked a few things that I wanted.

Let's not get into an endless argument over Ruby v Python v Java. Let me just say that Python is fine (I won't get into my peeves regarding Java), but I chose Ruby because 1) its syntax was beautiful and simple, and I had no trouble learning the language; 2) it enforces single lineage inheritance (which greatly simplifies the language and fits well with the biological classifications that I work with), and 3) it uses the so-called open world paradigm for evaluating assertions, returning true, false or nil (rather than the true/false dichotomy of Perl and Python). I really need Ruby's "nil".

When do you use Perl, and when do you use Ruby? I use Perl whenever I want to create simple utility scripts (transforming one file into another file of a different structure, performing a single algorithm on an input, and so on). In the past, most of my work was this sort of thing. I don't use Ruby to create short utilities because Ruby is slower than Perl. A Ruby script will execute in about twice the time as a Perl script for the same algorithm. This is true of all object-oriented languages. The primary reason they run slowly is because they need to traverse their object libraries when methods are sent to objects.

I use Ruby for modeling biomedical domains. This usually means that if I'm using RDF, ontologies, classifications, objects, object libraries, I use Ruby.

Some of you may have heard of Ruby on Rails (RoR). This is a web server programming environment for creating simple, quick, elegant, object-oriented Web applications. It is wildly popular at the moment. It's just one more perk to learning Ruby.


-Jules Berman
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Friday, March 9, 2007

The uses of free standards can be patented

In a prior post, I described several ways in which data standards can become encumbered with intellectual property. One of these involves patenting the way that a standard is used. Even when a patent is free, there is nothing to stop an inventor from patenting uses for the standard.

DICOM (Digital Imaging and Communications in Medicine) exemplifies a standard that has a patented use. DICOM is a widely used image standard for radiologic images. Currently, there is an effort to have all medical specialties adopt DICOM as the exclusive format for all medical images.

U.S. Patent 6725231 , issued Apr 20, 2004, to Jingkun Hu and Kwok Pun Lee and assigned to Koninklijke Philips Electronics N.V., has the following claim.

"1. A method for mapping a DICOM specification into an XML document, comprising: mapping each entry of a DICOM table of the DICOM specification into a corresponding XML element of a plurality of XML elements,outputting each XML element of the plurality of XML elements to the XML document, in an output format that conforms to at least one of: an XML document-type-definition and an XML Schema."

A similar patent by the same parties sits at the European Patent Office (EPO).

Informaticians will note that teasing the data elements from a data object and porting them into XML is the bread-and-butter of modern informatics. A patent claim that covers this basic use of DICOM may be highly problematic.

SDOs(Standards Development Organizations) cannot stop inventors from patenting new and useful applications of their standards. However, there are easy ways for SDOs to reduce the risk of inventors patenting the common, expected uses of their standards. These will be described in a future post.

-Jules Berman tags: biomedical informatics, converting to xml, data standards, DICOM, embedded patents, european patent office, medical images, patent claims, radiology images, sdo, uspto, xml, science
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.

Thursday, March 8, 2007

Specifications versus Standards

In a prior blog I suggested 16 ways that SDOs (Standards Development Organizations) can protect their standards from embedded patents. Suggestion 12 was "Make specifications, not standards."

This suggestion, I'm sure, is cryptic to most people. A major theme of this blog site is that specifications are different from standards and have a number of features that make them more suitable than standards for describing and exchanging many types of biomedical information.

Though informaticians often use the terms "specification" and "standard" interchangeably, a specification is just a formal way (usually employing RDF) of describing any data object. A data standard is a set of requirements, created by an SDO, that comprise a pre-determined content and format for a set of data related to a very specific kind of data object.

Features of a "specified" object:

1. Anyone can understand the composition and construction of the object

2. If the object is unique, anyone can distinguish the object from all other objects.

3. If the object falls into a known class of objects, anyone can determine, from the specification, the class of the object.


A specification serves most of the purposes of a standard, and much more (data description, data exchange, data merging, data interoperability, semantic logic). Data specifications spare us most of the heavy baggage that comes with a standard (limited flexibility to include changing data objects, locked-in data descriptors, licensing and other intellectual property issues, competing standards for the same domain resulting in limited interoperability, bureaucratic overhead, etc.).

Readers of this blog might want to read an introduction to RDF data specifications written by myself and Dr. G. William Moore. I believe that standards are important, but that specifications are even more important. There are instances in the field of biomedical informatics where specifications could serve better than standards. This was a developed theme in my book, Biomedical Informatics.I hope to provide many examples of specifications (how they are created and used) in future blogs here.

-Jules Berman
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.