Skip to main content

Biomedical and Electrical Engineer with interests in information theory, evolution, genetics, abstract mathematics, microbiology, big history, IndieWeb, mnemonics, and the entertainment industry including: finance, distribution, representation

boffosocko.com

chrisaldrich

chrisaldrich

+13107510548

chris@boffosocko.com

stream.boffosocko.com

www.boffosockobooks.com

chrisaldrich

mastodon.social/@chrisaldrich

micro.blog/chrisaldrich

 

17w5131: Statistical & Computational Challenges in Large Scale Molecular Biology Workshop @BIRS_Math 3/2017 #ITBio

Arriving in Banff, Alberta Sunday, March 26 and departing Friday March 31, 2017

Organizers

  • Barbara Engelhardt (Princeton University)
  • Anna Goldenberg (University of Toronto)
  • Manolis Kellis (Massachusetts Institute of Technology)
  • Jacob Laurent (Centre national de la recherche scientifique)
  • Jeff Leek (John Hopkins University)
  • Stephen Montgomery (Stanford University)

Objectives

Over the past few years, an increasing number of large scale data sets have been made available in molecular biology. GTEx, for example, produced more than 18,000 RNA-Seq assays for multiple tissues in 900 individuals, Mindact generated gene expression data from about 7000 breast tumors in a single study, and 23andMe claims to have sequenced about 900,000 genomes. This growth in the available genomic data is expected to increase our capacity to identify cancer subtypes, regulatory genes, SNPs associated with phenotypes of interest, and biomarkers for many human traits. It also suggests exploring more complex feature representations when analyzing these datasets.

However, increasing the number of samples and features leads to a set of \textbf{interrelated statistical and computational problems}. Accordingly, the objectives of our workshop will be to:

Systematically identify the statistical and computational

problems arising during the analysis of large scale data in molecular biology;

Bring together experts in computational biology, molecular

biology, computer science, and statistics to propose innovative solutions to these problems, by leveraging recent advances in each of these fields.

Relevance, importance and timeliness

A number of studies generating high throughput molecular data for a large number of biological samples have been completed over the past five years. \textbf{Our workshop is important because the availability of these datasets holds great promises in terms of health improvement and understanding of molecular biology}. First of all, if exploited correctly, larger sample sizes should improve our ability to predict phenotypes of interest from molecular data. This entails very important applications such as improving the survival of cancer patients by better predicting which treatment they should receive, or decreasing bacterial resistances by predicting which antibiotic is efficient against a new strain. Correctly exploiting large scale datasets should also allow us to \textbf{better identify genetic and epigenetic determinants of these phenotypes, yielding a better understanding of human diseases and potentially guiding the development of new treatments and prevention policies}. In particular, more samples should allow the detection of less frequent variants in the human genome, or more complex features involving several modalities (copy number, expression, methylation, etc) associated with diseases. Finally, larger sample sizes should help with essential unsupervised tasks such as the \textbf{inference of regulation networks, or the identification of cancer subtypes}.

Our workshop is relevant because \textbf{all of these promises are conditioned on our solving of new statistical and computational challenges}. First (Challenge 1), we need to build new feature spaces and estimators whose complexity is adapted to these larger sample sizes, which involves designing novel, potentially more complex descriptors of the samples but still controlling the bias/variance trade-off. Second (Challenge 2), we need to build models which correctly integrate different modalities, such as copy number variation and gene expression. Third (Challenge 3), larger scale studies are more prone to unwanted variations, because they typically involve different labs and technical changes which can affect the measurements and become confounders in retrospective analyses. Similar or worse problems arise when trying to combine several existing datasets. We need methods which take this unwanted variation into account. Finally, (Challenge 4), we need new algorithms that make existing statistical tools scalable to the new sample sizes, and make estimation over the larger and more complex features of Challenge 1 tractable.

We also believe our workshop is very timely because \textbf{some of these statistical and computational challenges are starting to be addressed in other application fields} of statistics. It is crucial to recognize that the orders of magnitude are still very different in molecular biology and other data science application fields because of the cost and complexity of the data generation process: current large scale high throughput sequencing data sets typically contain a few thousand of samples but millions of features while computer vision, web, or astronomy datasets can involve billions or trillions of samples and relatively fewer features. A first consequence is that not all recent developments in machine learning are immediately transferable to computational biology. For example, so called deep learning methods have gained a lot of popularity and now represent the state of the art in computer vision but may not be the most appropriate tool for prediction of cancer outcome from molecular data. However, the fact that other fields already have much larger sample sizes also means that they had to develop efficient and scalable algorithms for basic tasks like feature selection, classification or clustering. \textbf{These recent developments are a great source of inspiration for computational biology, where large scale computation is still an emerging challenge}.

We believe \textbf{having a small scale workshop involving international experts in machine learning, statistics, computational biology and molecular biology is of utmost importance} for three main reasons. The first reason is that the technical advances we are referring to are very recent, often unknown to computational biologists and involve paradigms such as online optimization, accelerated gradient methods and network flow optimization, with which they are sometimes unfamiliar. The second reason is that it is not always obvious to non-statisticians which novel methods are appropriate given the current n/p regime. Conversely, the third reason is that statisticians do not know what the recent challenges are in molecular biology. Having them work on abstract versions of the problems is often not satisfactory as it is necessary to be aware of technical realities and of the underlying biology of the problem to come up with useful solutions.

 
 

The Evolution of Information Gathering: Operational Constraints by Cynthia F. Kurtz

The Evolution of Information Gathering: Operational Constraints
Cynthia F. Kurtz
1991 Master's Thesis, SUNY Stony Brook, Ecology & Evolution

Abstract: I present two new approaches to the study of information in foraging theory. First, rather than determine the cost a forager should pay to obtain information, I concentrate on the consequences of information use in an interacting population. I describe a density-dependent model which tracks genotypes with high and low information access through evolutionary time. Stable polymorphisms result. I suggest that the value of information is not monotonically increasing. Second, I present a scheme for partitioning the information used in the decision making process. Three types of information are recognized: internal information, or an individual's internal state; external information, or environmental factors; and relational information, or rules for predicting transformations of internal state. Interactions between the three types are examined in an extension of the basic model.

 

‘Sleeping beauty’ papers slumber for decades #complexity

Interesting statistics on sleeping papers. There are some cool things going on here...

 

"A few exciting words": information and entropy revisited | Lyn Robinson and David Bawden - Academia.edu

A review is presented of the relation between information and entropy, focusing on two main issues: the similarity of the formal definitions of physical entropy, according to statistical mechanics, and of information, according to information theory; and the possible subjectivity of entropy considered as missing information. The paper updates the 1983 analysis of Shaw and Davis. The difference in the interpretations of information given respectively by Shannon and by Wiener, significant for the information sciences, receives particular consideration. Analysis of a range of material, from literary theory to thermodynamics, is used to draw out the issues. Emphasis is placed on recourse to the original sources, and on direct quotation, to attempt to overcome some of the misunderstandings and oversimplifications that have occurred with these topics. While it is strongly related to entropy, information is neither identical with it, nor its opposite. Information is related to order and pattern, but also to disorder and randomness. The relations between information and the “interesting
complexity,” which embodies both patterns and randomness, are worthy of attention.

 

"Waiting for Carnot": information and complexity | Lyn Robinson and David Bawden - Academia.edu

Abstract: The relationship between information and complexity is analysed, by way of a detailed literature analysis. Complexity is a multi‐faceted concept, with no single agreed definition. There are numerous approaches to defining and measuring complexity and organisation, all involving the idea of information. Conceptions of complexity, order, organization and ‘interesting order’ are inextricably intertwined with those of information. Shannon’s formalism captures information’s unpredictable creative contributions to organized complexity; a full understanding of information’s relation to structure and order is still lacking. Conceptual investigations of this topic should enrich the theoretical basis of the information science discipline, and create fruitful links with other disciplines which study the concepts of information and complexity.