Showing posts with label data sharing. Show all posts
Showing posts with label data sharing. Show all posts

Saturday, August 4, 2018

Second Edition of Principles and Practice of Big Data now on Science Direct

The Second edition of my book Principles and Practice of Big Data has just been released and is available for purchase at many sites, including Amazon.

For those of you fortunate enough to have access to Science Direct, you can download chapters of my book at:

https://www.sciencedirect.com/science/book/9780128156094



TABLE OF CONTENTS

  Author's Preface to Second Edition 

  Author's Preface to First Edition 

  Chapter 1. Introduction
    Section 1.  Definition of Big Data
    Section 2.  Big Data Versus small data
    Section 3.  Whence Comest Big Data?
    Section 4.  The Most Common Purpose of Big Data is to Produce small data
    Section 5.  Big Data Sits at the Center of the Research Universe
    Section 6.  Case Study: From the Press: Big Claims for Big Data

  Chapter 2. Providing Structure to Unstructured Data
    Section 1.  Nearly all Data is Unstructured and Unusable in its Raw Form
    Section 2.  Term Extraction
    Section 3.  Autocoding
    Section 4.  Concordances
    Section 5.  Indexing
    Section 6.  Machine Translation
    Section 7.  Case Study: Sorted Lists (Why and Why Not)
    Section 8.  Case Study: Doublet Lists 
    Section 9.  Case Study: Ngram Lists 
    Section 10.  Case Study: Proximity Searches Using Only a Concordance  
    Section 11.  Case Study (Advanced): Burrows Wheeler Transform (BWT) 

  Chapter 3. Identification, Deidentification, and Reidentification
    Section 1.  What are Identifiers?
    Section 2.  Difference Between an Identifier and an Identifier System
    Section 3.  Generating Identifiers
    Section 4.  Really Bad Identifier Methods
    Section 5.  Registered Unique Object Identifiers
    Section 6.  Deidentification
    Section 7.  Reidentification
    Section 8.  Case Study: Data Scrubbing
    Section 9.  Case Study: Identifiers in Image Headers
    Section 10.  Case Study: Hospital Registration
    Section 11.  Case Study: One-Way Hashes

  Chapter 4. Metadata, Semantics, and Triples
    Section 1.  Metadata
    Section 2.  eXtensible Markup Language
    Section 3.  Namespaces
    Section 4.  Semantics and Triples
    Section 5.  Case Study: Syntax for Triples 
    Section 6.  Case Study: RDF Schema
    Section 7.  Case Study: RDF Parsers and the Fungibility of Triples
    Section 8.  Case Study: Dublin Core 

  Chapter 5. Classifications and Ontologies
    Section 1.  It's All About Object Relationships 
    Section 2.  The Difference Between Object Relationships and Object Similarities
    Section 3.  Classifications, the Simplest of Ontologies
    Section 4.  Ontologies, Classes with Multiple Parents
    Section 5.  Choosing a Class Model
    Section 6.  Paradoxes
    Section 7.  Class Blending
    Section 8.  Common Pitfalls in Ontology Development
    Section 9.  Case Study: An Upper Level Ontology 
    Section 10.  Case Study: Visualizing Class Relationships 
    Section 11.  Case Study: Bringing Order from Chaos with the Classification of Living Organisms

  Chapter 6. Introspection
    Section 1.  Knowledge of Self
    Section 2.  Data Objects
    Section 3.  How Big Data Uses Introspection 
    Section 4.  Case Study: Timestamping Data 
    Section 5.  Case Study: A Visit to the TripleStore 

  Chapter 7. Data Integration and Software Interoperability
    Section 1.  Another Big Problem for Big Data
    Section 2.  The Standard for Standards
    Section 3.  Standard Trajectories
    Section 4.  Specifications and Standards
    Section 5.  Versioning
    Section 6.  Compliance Issues
    Section 7.  Interfaces to Big Data Resources
    Section 8.  Case Study: Standardizing the Chocolate Teapot

  Chapter 8. Immutability and Immortality
    Section 1.  The Importance of Data that Cannot Change  
    Section 2.  Immutability and Identifiers
    Section 3.  Persistent Data Objects
    Section 4.  Coping with the Data that Data Creates
    Section 5.  Reconciling Identifiers Across Institutions
    Section 6.  Case Study: The Trusted Timestamp
    Section 7.  Case Study: Blockchains and Distributed Ledgers
    Section 8.  Case Study: Zero-Knowledge Reconciliation   

  Chapter 9. Assessing the Adequacy of a Big Data Resource
    Section 1.  Looking at the Data 
    Section 2.  The Minimal Necessary Properties of Big Data 
    Section 3.  Case Study: Utilities for Viewing and Manipulating Very Large Files
    Section 4.  Case Study: Flattened Data 
    Section 5.  Case Study: Data that Comes with Conditions 

  Chapter 10. Measurement
    Section 1.  Accuracy and Precision
    Section 2.  Data Range
    Section 3.  Counting
    Section 4.  Normalizing, and Transforming Your Data
    Section 5.  Reducing Your Data
    Section 6.  Understanding Your Control
    Section 7.  Practical Significance of Measurements
    Section 8.  Case Study: Gene Counting
    Section 9.  Case Study: The Significance of Narrow Data Ranges
    Section 10.  Case Study (Advanced): Fast Fourier Transform
    Section 11.  Case Study (Advanced): Principal Component Analysis

  Chapter 11. Indispensable Tips for Fast and Simple Big Data Analysis
    Section 1.  Speed and Scalability
    Section 2.  Fast Operations, Suitable for Big Data, that Every Computer Supports
    Section 3.  Fast Correlation Methods
    Section 4.  Clustering 
    Section 5.  Methods for Data Persistence (Without Using a Database)
    Section 6.  Back_of_Envelope Computations for Big Data
    Section 7.  Fast Data Retrieval for Lists of any Size 
    Section 8.  Case Study: One-Pass Mean and Standard Deviation
    Section 9.  Case Study: Climbing a Classification
    Section 10.  Pre-computing lookup lists: Google's PageRank
    Section 11.  Case Study: A Database Example 
    Section 12.  NoSQL and other Non-Relational Big Data Databases

  Chapter 12. Finding the Clues in Large Collections of Data
    Section 1.  Denominators 
    Section 2.  Frequency Distributions
    Section 3.  Multimodality
    Section 4.  Outliers and Anomalies
    Section 5.  Case Study: Discarding the Noisiest Frequencies in a Data Signal
    Section 6.  Case Study: Predicting User Preferences
    Section 7.  Case Study: Multimodality in Legacy Data
    Section 8.  Case Study: Big and Small Black Holes

  Chapter 13. Using Random Numbers to Your Big Data Analytic Problems Down to Size
    Section 1.  The Remarkable Utility of (Pseudo)Random Numbers 
    Section 2.  Resampling and Permutating 
    Section 3.  Case Study: Sample Size and Power Estimates
    Section 4.  Monte Carlo Simulations
    Section 5.  Case Study: Monty Hall Problem: Solving What We Cannot Grasp
    Section 6.  Case Study: Frequency of Unlikely String of Occurrences 
    Section 7.  Case Study: The Infamous Birthday Problem
    Section 8.  Case Study: A Bayesian Analysis of Insurance Costs 

  Chapter 14. Special Considerations in Big Data Analysis
    Section 1.  Theory in Search of Data 
    Section 2.  Data in Search of Theory
    Section 3.  Overfitting
    Section 4.  Bigness Bias
    Section 5.  Too Much Data
    Section 6.  Fixing Data
    Section 7.  Data Subsets in Big Data: Neither Additive nor Transitive
    Section 8.  Additional Big Data Pitfalls
    Section 9.  Case Study: Curse of Dimensionality

  Chapter 15. Big Data Failures and How to Avoid (Some of) Them
    Section 1.  Failure is Common
    Section 2.  Failed Standards
    Section 3.  Blaming Complexity
    Section 4.  Perils of Redundancy
    Section 5.  Save Time and Money; Don’t Protect Data that Does not Need Protection
    Section 6.  An Approach to Big Data that May Work For You
    Section 7.  After Failure
    Section 8.  Case Study: Cancer Biomedical Informatics Grid, a Bridge too Far
    Section 9.  Case Study: The Gaussian Copula Function

  Chapter 16. Legalities
    Section 1.  Responsibility for the Accuracy and Legitimacy of Data
    Section 2.  Rights to Create, Use, and Share the Resource
    Section 3.  Copyright and Patent Infringements Incurred by Using Standards
    Section 4.  Protections for Individuals
    Section 5.  Consent
    Section 6.  Unconsented Data
    Section 7.  Good Policies are a Good Policy
    Section 8.  Case Study: The "Inconclusive" Data Analysis
    Section 9.  Case Study: The Havasupai Story
    Section 10.  Case Study: Double-edged Sword of the U.S. Data Quality Act 

  Chapter 17. Data Sharing 
    Section 1.  What Is Data Sharing, and Why Don't We Do More of It?
    Section 2.  Common Complaints
    Section 3.  Case Study: Life on Mars
    Section 4.  Case Study: Who Shares Their Data 
    Section 5.  Case Study: National Patient Identifier

  Chapter 18. Data Reanalysis: Much More Important than Analysis
    Section 1.  First Analysis (Nearly) Always Wrong 
    Section 2.  Why Reanalysis is More Important than Analysis
    Section 3.  Case Study: Reanalysis of Old JADE Collider Data 
    Section 4.  Case Study: Vindication Through Reanalysis 
    Section 5.  Case Study: Finding New Planets from Old Data 

  Chapter 19. Repurposing Big Data
    Section 1.  What is Data Repurposing? 
    Section 2.  Dark Data, Abandoned Data, and Legacy Data 
    Section 3.  Case Study: From Postal Code to Demographic Keystone 
    Section 4.  Case Study: Fingerprints and Data-driven Forensics
    Section 5.  Scientific Inferencing from a Databases of Genetic Sequences
    Section 6.  Case Study: Linking global warming to high-intensity hurricanes
    Section 7.  Case Study: Inferring climate trends with geologic data
    Section 8.  Case Study: Old tidal data, and the iceberg that sank the Titanic
    Section 9.  Case Study: Lunar Orbiter Image Recovery Project
    Section 10.  Case Study: The Cornucopia of the Natural Sciences

  Chapter 20. Societal Issues
    Section 1.  How Big Data Is Perceived by the Public
    Section 2.  Reducing Costs and Increasing Productivity with Big Data
    Section 3.  Public Mistrust
    Section 4.  Saving Us from Ourselves 
    Section 5.  Who is Big Data?
    Section 6.  Hubris and Hyperbole
    Section 7.  Case Study: The Citizen Scientists
    Section 8.  Case Study: 1984, by George Orwell

  




- Jules Berman

Thursday, May 30, 2013

Big Data Book Contents

I've taken a hiatus from the Specified Life blog while I wrote my latest book, entitled, Principles of Big Data: Preparing, Sharing, and Analyzing Complex Information.



The Kindle edition is available now, and Amazon has a "look-inside" option on their book page. The print version will be available in a week or two, and Amazon is taking pre-orders Here is the complete Table of Contents:
Acknowledgments xi
Author Biography xiii
Preface xv
Introduction xix

1. Providing Structure to Unstructured Data
  Background 1
  Machine Translation 2
  Autocoding 4
  Indexing 9
  Term Extraction 11

2. Identification, Deidentification, and Reidentification
  Background 15
  Features of an Identifier System 17
  Registered Unique Object Identifiers 18
  Really Bad Identifier Methods 22
  Embedding Information in an Identifier: Not Recommended 24
  One-Way Hashes 25
  Use Case: Hospital Registration 26
  Deidentification 28
  Data Scrubbing 30
  Reidentification 31
  Lessons Learned 32

3. Ontologies and Semantics
  Background 35
  Classifications, the Simplest of Ontologies 36
  Ontologies, Classes with Multiple Parents 39
  Choosing a Class Model 40
  Introduction to Resource Description Framework Schema 44
  Common Pitfalls in Ontology Development 46

4. Introspection
  Background 49
  Knowledge of Self 50
  eXtensible Markup Language 52
  Introduction to Meaning 54
  Namespaces and the Aggregation of Meaningful Assertions 55
  Resource Description Framework Triples 56
  Reflection 59
  Use Case: Trusted Time Stamp 59
  Summary 60

5. Data Integration and Software Interoperability
  Background 63
  The Committee to Survey Standards 64
  Standard Trajectory 65
  Specifications and Standards 69
  Versioning 71
  Compliance Issues 73
  Interfaces to Big Data Resources 74

6. Immutability and Immortality
  Background 77
  Immutability and Identifiers 78
  Data Objects 80
  Legacy Data 82
  Data Born from Data 83
  Reconciling Identifiers across Institutions 84
  Zero-Knowledge Reconciliation 86
  The Curator’s Burden 87

7. Measurement
  Background 89
  Counting 90
  Gene Counting 93
  Dealing with Negations 93
  Understanding Your Control 95
  Practical Significance of Measurements 96
  Obsessive-Compulsive Disorder: The Mark of a Great Data Manager 97

8. Simple but Powerful Big Data Techniques
  Background 99
  Look at the Data 100
  Data Range 110
  Denominator 112
  Frequency Distributions 115
  Mean and Standard Deviation 119
  Estimation-Only Analyses 122
  Use Case: Watching Data Trends with Google Ngrams 123
  Use Case: Estimating Movie Preferences 126

9. Analysis
  Background 129
  Analytic Tasks 130
  Clustering, Classifying, Recommending, and Modeling 130
  Data Reduction 134
  Normalizing and Adjusting Data 137
  Big Data Software: Speed and Scalability 139
  Find Relationships, Not Similarities 141

10. Special Considerations in Big Data Analysis
  Background 145
  Theory in Search of Data 146
  Data in Search of a Theory 146
  Overfitting 148
  Bigness Bias 148
  Too Much Data 151
  Fixing Data 152
  Data Subsets in Big Data: Neither Additive nor Transitive 153
  Additional Big Data Pitfalls 154

11. Stepwise Approach to Big Data Analysis
  Background 157
  Step 1. A Question Is Formulated 158
  Step 2. Resource Evaluation 158
  Step 3. A Question Is Reformulated 159
  Step 4. Query Output Adequacy 160
  Step 5. Data Description 161
  Step 6. Data Reduction 161
  Step 7. Algorithms Are Selected, If Absolutely Necessary 162
  Step 8. Results Are Reviewed and Conclusions Are Asserted 164
  Step 9. Conclusions Are Examined and Subjected to Validation 164

12. Failure
  Background 167
  Failure Is Common 168
  Failed Standards 169
  Complexity 172
  When Does Complexity Help? 173
  When Redundancy Fails 174
  Save Money; Don’t Protect Harmless Information 176
  After Failure 177
  Use Case: Cancer Biomedical Informatics Grid, a Bridge Too Far 178

13. Legalities
  Background 183
  Responsibility for the Accuracy and Legitimacy of Contained Data 184
  Rights to Create, Use, and Share the Resource 185
  Copyright and Patent Infringements Incurred by Using Standards 187
  Protections for Individuals 188
  Consent 190
  Unconsented Data 194
  Good Policies Are a Good Policy 197
  Use Case: The Havasupai Story 198

14. Societal Issues
  Background 201
  How Big Data Is Perceived 201
  The Necessity of Data Sharing, Even When It Seems Irrelevant 204
  Reducing Costs and Increasing Productivity with Big Data 208
  Public Mistrust 210
  Saving Us from Ourselves 211
  Hubris and Hyperbole 213

15. The Future
  Background 217
  Last Words 226

Glossary 229

References 247

Index 257
In the next few days, I'll be posting short excerpts from the book, along with commentary. Best,
Jules Berman

key words: big data, heterogeneous data, complex datasets, Jules J. Berman, Ph.D., M.D., immutability, introspection, identifiers, de-identification, deidentification, confidentiality, privacy, massive data, lotsa data

Tuesday, March 11, 2008

MISFISHIE Specification (for in situ hybridization and immunohistochemistry) now available

Under the leadership of Eric Deutsch, a specification for annotating In Situ Hybridization and Immunohistochemistry Experiments (MISFISHIE) has just been published in Nature Biotechnology.

Deutsch EW, Ball CA, Berman JJ, Bova GS, Brazma A, Bumgarner RE, Campbell D, Causton HC, Christiansen JH, Daian F, Dauga D, Davidson DR, Gimenez G, Goo YA, Grimmond S, Henrich T, Herrmann BG, Johnson MH, Korb M, Mills JC, Oudes AJ, Parkinson HE, Pascal LE, Pollet N, Quackenbush J, Ramialison M, Ringwald M, Salgado D, Sansone SA, Sherlock G, Stoeckert CJ Jr, Swedlow J, Taylor RC, Walashek L, Warford A, Wilkinson DG, Zhou Y, Zon LI, Liu AY, True LD. Minimum information specification for in situ hybridization and immunohistochemistry experiments (MISFISHIE). Nat Biotechnol. 2008 Mar;26(3):305-12.

MISFISHIE is modelled after the MIAME (Minimum Information About a Microarray Experiment) specification for microarray experiments.

It has been a constant theme in this blog that data specifications are, in many instances, much better than data standards. Data specifications, like MIAME and MISFISHIE specify the information content without dictating a format for encoding that information.

Nature Biotechnology put up the entire MISFISHIE specification for public comment, and it is currently available at:

http://www.nature.com/nbt/consult/pdf/Deutsch.pdf


- Jules Berman
In June, 2014, my book, entitled Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases was published by Elsevier. The book builds the argument that our best chance of curing the common diseases will come from studying and curing the rare diseases.



I urge you to read more about my book. There's a generous preview of the book at the Google Books site. If you like the book, please request your librarian to purchase a copy of this book for your library or reading room.

tags: common disease, orphan disease, orphan drugs, genetics of disease, disease genetics, rules of disease biology, rare disease, pathology,annotation, data sharing, data specifications

Friday, June 1, 2007

Funding opportunity in precancer research

The U.S. National Cancer Institute (NCI) has put out an innovative Request for Applications (RFA) for precancer research. As you know, I support the idea that attacking precancers is the best way to eliminate human cancer. There's every reason to think that precancers can be treated successfully with low-toxicity agents that interfere with the pathways of precancer growth and progression or that enhance the pathways of precancer death. Most of the research in this field will be data-intensive. Those who know how to specify their data will probably welcome the data sharing provisions in the RFA.

The RFA, focused on breast precancers, just came out, and can be viewed at:

http://grants.nih.gov/grants/guide/rfa-files/RFA-CA-07-047.html

Release/Posted Date: May 30, 2007
Opening Date: September 14, 2007
Letters of Intent Receipt Date: October 14, 2007

The RFA uses an R01 funding mechanism (that's good).

The RFA cites our November 2004 conference on precancers that was co-sponsored by George Washington University.

From the RFA: "The NCI as well as experts in the extramural scientific community recommend further research related to the biology of the pre-malignant state in human breast cancer. An expert panel convened at the November 2004 NCI Workshop on Pre-Cancers identified delineation of the biological, genetic, and functional characteristics of pre-cancers as major scientific needs (Cancer Detect Prev. 2006;30(5):387-94). The distinctive early lesions that occur have characteristic properties that should permit them to be detected, diagnosed, and prevented from progressing to invasive cancer. The Workshop participants noted a number of impediments to conducting research on pre-cancers, including:

* insufficient understanding of normal and pre-cancer biology;
* limited access to appropriate specimens;
* a highly subjective, histology-based classification scheme; and
* the lack of strategic partnerships among research communities."


-Jules Berman tags: cancer research, data sharing, funding, nci, precancer, science
Science is not a collection of facts. Science is what facts teach us; what we can learn about our universe, and ourselves, by deductive thinking. From observations of the night sky, made without the aid of telescopes, we can deduce that the universe is expanding, that the universe is not infinitely old, and why black holes exist. Without resorting to experimentation or mathematical analysis, we can deduce that gravity is a curvature in space-time, that the particles that compose light have no mass, that there is a theoretical limit to the number of different elements in the universe, and that the earth is billions of years old. Likewise, simple observations on animals tell us much about the migration of continents, the evolutionary relationships among classes of animals, why the nuclei of cells contain our genetic material, why certain animals are long-lived, why the gestation period of humans is 9 months, and why some diseases are rare and other diseases are common. In “Armchair Science”, the reader is confronted with 129 scientific mysteries, in cosmology, particle physics, chemistry, biology, and medicine. Beginning with simple observations, step-by-step analyses guide the reader toward solutions that are sometimes startling, and always entertaining. “Armchair Science” is written for general readers who are curious about science, and who want to sharpen their deductive skills.