Arriving in Banff, Alberta Sunday, March 26 and departing Friday March 31, 2017
Over the past few years, an increasing number of large scale data sets have been made available in molecular biology. GTEx, for example, produced more than 18,000 RNA-Seq assays for multiple tissues in 900 individuals, Mindact generated gene expression data from about 7000 breast tumors in a single study, and 23andMe claims to have sequenced about 900,000 genomes. This growth in the available genomic data is expected to increase our capacity to identify cancer subtypes, regulatory genes, SNPs associated with phenotypes of interest, and biomarkers for many human traits. It also suggests exploring more complex feature representations when analyzing these datasets.
However, increasing the number of samples and features leads to a set of \textbf{interrelated statistical and computational problems}. Accordingly, the objectives of our workshop will be to:
Systematically identify the statistical and computational
problems arising during the analysis of large scale data in molecular biology;
Bring together experts in computational biology, molecular
biology, computer science, and statistics to propose innovative solutions to these problems, by leveraging recent advances in each of these fields.
A number of studies generating high throughput molecular data for a large number of biological samples have been completed over the past five years. \textbf{Our workshop is important because the availability of these datasets holds great promises in terms of health improvement and understanding of molecular biology}. First of all, if exploited correctly, larger sample sizes should improve our ability to predict phenotypes of interest from molecular data. This entails very important applications such as improving the survival of cancer patients by better predicting which treatment they should receive, or decreasing bacterial resistances by predicting which antibiotic is efficient against a new strain. Correctly exploiting large scale datasets should also allow us to \textbf{better identify genetic and epigenetic determinants of these phenotypes, yielding a better understanding of human diseases and potentially guiding the development of new treatments and prevention policies}. In particular, more samples should allow the detection of less frequent variants in the human genome, or more complex features involving several modalities (copy number, expression, methylation, etc) associated with diseases. Finally, larger sample sizes should help with essential unsupervised tasks such as the \textbf{inference of regulation networks, or the identification of cancer subtypes}.
Our workshop is relevant because \textbf{all of these promises are conditioned on our solving of new statistical and computational challenges}. First (Challenge 1), we need to build new feature spaces and estimators whose complexity is adapted to these larger sample sizes, which involves designing novel, potentially more complex descriptors of the samples but still controlling the bias/variance trade-off. Second (Challenge 2), we need to build models which correctly integrate different modalities, such as copy number variation and gene expression. Third (Challenge 3), larger scale studies are more prone to unwanted variations, because they typically involve different labs and technical changes which can affect the measurements and become confounders in retrospective analyses. Similar or worse problems arise when trying to combine several existing datasets. We need methods which take this unwanted variation into account. Finally, (Challenge 4), we need new algorithms that make existing statistical tools scalable to the new sample sizes, and make estimation over the larger and more complex features of Challenge 1 tractable.
We also believe our workshop is very timely because \textbf{some of these statistical and computational challenges are starting to be addressed in other application fields} of statistics. It is crucial to recognize that the orders of magnitude are still very different in molecular biology and other data science application fields because of the cost and complexity of the data generation process: current large scale high throughput sequencing data sets typically contain a few thousand of samples but millions of features while computer vision, web, or astronomy datasets can involve billions or trillions of samples and relatively fewer features. A first consequence is that not all recent developments in machine learning are immediately transferable to computational biology. For example, so called deep learning methods have gained a lot of popularity and now represent the state of the art in computer vision but may not be the most appropriate tool for prediction of cancer outcome from molecular data. However, the fact that other fields already have much larger sample sizes also means that they had to develop efficient and scalable algorithms for basic tasks like feature selection, classification or clustering. \textbf{These recent developments are a great source of inspiration for computational biology, where large scale computation is still an emerging challenge}.
We believe \textbf{having a small scale workshop involving international experts in machine learning, statistics, computational biology and molecular biology is of utmost importance} for three main reasons. The first reason is that the technical advances we are referring to are very recent, often unknown to computational biologists and involve paradigms such as online optimization, accelerated gradient methods and network flow optimization, with which they are sometimes unfamiliar. The second reason is that it is not always obvious to non-statisticians which novel methods are appropriate given the current n/p regime. Conversely, the third reason is that statisticians do not know what the recent challenges are in molecular biology. Having them work on abstract versions of the problems is often not satisfactory as it is necessary to be aware of technical realities and of the underlying biology of the problem to come up with useful solutions.
Great conversation with @Cesifoti on economic complexity: Why Information Grows | @TheEconomist http:/
And minutes later, I've got an advance copy of @seanmcarroll's book The Big Picture to read for the weekend!
#cantwait #complexity #ITBio
László Babai posted #complexity related paper to arXiv: Algorithm Solves Graph Isomorphism in Record Time https:/
#Conference Complex Networks: From theory to interdisciplinary applications http:/
Relative Entropy in Biological Systems by John Baez and Blake Pollard https:/
László Babai has posted the first of his two talks on complexity on his webpage:
http:/
Some additional thoughts on information theory and complexity for models in ecology.
In addition to some of my commentary on <a href="http://biology.stackexchange.com/a/40555/6629">StackExchange</a>, I'll include a few additional thoughts which may be of help/assistance.
First for those who don't have the background, I <i>highly</i> recommend reading the two seminal papers on information theory and statistical mechanics by ET Jaynes and the standard text on information theory <i><a href="http://amzn.to/1NThEjd">Elements of Information Theory</a></i> by Thomas M. Cover and Joy A. Thomas.
In addition to the information theoretic related areas, you might want to take a look at the discipline of complexity theory, which has primarily grown out of the Santa Fe Institute over the past several decades and which includes information theory as part of its disciplines. If you're unfamiliar with the broader topic, Melanie Mitchell has an excellent overview with her book <i><a href="http://amzn.to/1YbcUZ4">Complexity: A Guided Tour</a></i>. Also related to complexity is the area of cellular automata which one could view as a very base model of more complex ecological systems. Here, perhaps Stephen Wolfram's <i><a href="http://amzn.to/1kwIAtM">A New Kind of Science</a></i> or <i><a href="http://amzn.to/1Ybd3Me">Cellular Automata and Complexity</a></i> will be enlightening. The broader theories coming out of these primarily mathematical areas may be useful to you.
In particular, given the types of models in ecosystems, I might suggest taking a look at some of the mathematical modeling going on at the intersection of complexity and economics. For a relatively simple introduction to this area, one could look at the relatively introductory text <i><a href="http://amzn.to/1kwJ6Ic">Complexity and the Economy</a></i> by W. Brian Arthur which is very interesting. The economy is essentially a very specific type of ecology dealing with human beings, assets, and the monetary system.
Another area which I've seen a lot of literature over the last few years is applicable to the ideas of resiliency and complexity in cities, for assisting in designing more robust city planning. This really isn't that far from the naturally evolving systems being looked at in ecology settings.
For those looking for researchers in the area of complexity, I have a <a href="https://twitter.com/ChrisAldrich/lists/complexity/members">list of many who are on twitter</a> in a variety of sub-areas. In addition to individuals, it also includes a number of institutes and related organizations as well.
I'd also suggest, that for the broadest theoretical setting, one could actually start with the topic known as "Big History" which takes the broadest approach of looking at history and the evolution of the cosmos over 13.7 billion years since the big bang. This conceptualization includes ideas like evolution, complexity, and emergence on the biggest scales, a set of theories that could be similarly applied to ecologies both large and small. For this viewpoint, I would suggest two works by David Christian including <i><a href="http://amzn.to/1kwJPJy">Maps of Time: An Introduction to Big History</a></i> and <i><a href="http://amzn.to/1YbdExx">Big History: The Big Bang, Life on Earth, and the Rise of Humanity</a></i>.
In essence, with many of these topics and viewpoints, one is treating individual animals or even entire species as elementary particles and then using the mathematical models of statistical thermodynamics to tease out specific types of data or trends. As layers of overlapping "particles" interact with each other, they cause emergent types properties, and then these resultant emergent properties combine to create further layers of emergent properties, none of which might have been necessarily deduced from the initial conditions. Within Big History, these types of emergence go from the big bang and basic particles in the early universe to the ultimate evolution of humankind by way of a variety of stages.
Mathematician László Babai claims breakthrough in complexity theory http:/
My reply to @PeterCoffee "Information should be more ‘force’ than ‘mass’" | Diginomica
You're certainly asking the right types of questions here, but doing so in a somewhat flimsy framework in what is already a fairly solid mathematical theory. Mathematicians would call this a "hand-waving argument." Some of your conceptualization is clouded by mistaking the semantics of the commonly accepted definition of information with Claude Shannon's formal mathematical definition, something which he admonishes the reader about in the opening paragraphs of his seminal paper.
I would suggest that you take a look at Cesar Hidalgo's recent and very accessible book Why Information Grows: The Evolution of Order, from Atoms to Economies (MIT Press, 2015) [http:/
For those who want to take the economics piece a step further, one can then delve into the broader topics at the intersection of fields like complexity theory and economics similar to those framed by the Santa Fe Institute over the past two decades. (W. Brian Arthur's Complexity and the Economy, (Oxford, 2014) [http:/
From a business perspective, these theories underpin many of the ideas expounded by Judith Rodin's The Resilience Dividend: Being Strong in a World Where Things Go Wrong (Public Affairs, 2014) or many of the probabilistic arguments made in Nassim Nicholas Taleb's various works.
Wish I could have gone to Urban Design and Complexity talk. Looking forward to the full paper. #Complexity
Automata course starting soon via @Stanford on @Coursera https:/
Paul, thanks for the provocative piece, though the state of the art is certain much further along that your piece intimates. For the general reader, I would suggest reading MIT professor Cesar Hidalgo's recent book Why Information Grows (MIT Press, 2015) for some general structure and philosophy.
One of the best definitions and frameworks I've seen thus far has to be that of Christoph Adami. To start, and depending on your level of sophistication, take a look at his recent arXiv paper (Information-theoretic considerations concerning the origin of life - http:/
For further references, I maintain a nice list of resources at Information Theory and Biology Resources [http:/
For those who like to watch video material, I'll refer them to some videos from the NIMBioS Workshop on Information and Entropy in Biological Systems [http:/