Skip to main content

Biomedical and Electrical Engineer with interests in information theory, evolution, genetics, abstract mathematics, microbiology, big history, IndieWeb, mnemonics, and the entertainment industry including: finance, distribution, representation

boffosocko.com

chrisaldrich

chrisaldrich

+13107510548

chris@boffosocko.com

stream.boffosocko.com

www.boffosockobooks.com

chrisaldrich

mastodon.social/@chrisaldrich

micro.blog/chrisaldrich

 

Could I volunteer some web development time/energy to help any interested @RadioRookies journalists build their own websites/platforms for empowering/owning their voices, stories, and reporting portfolios? https://indieweb.org/Indieweb_for_Journalism

 

I think a lot of the problem comes down to all of the siloed walls out there which are causing most of the friction. We're still relatively early days yet and only a tiny few are using the concept of salmention which would help keep running threads working properly. Admittedly having the context live somewhere and then having proper threaded communications isn't easy, so many do what they're able to for the moment.

I'm usually attempting to manually accomplish salmention as best as I'm able, but I may not hit every syndicated target unless you're displaying it directly. Additionally some targets just don't make sense--I'll webmention your original, for example, but this lengthy reply just won't look right at micro.blog if you syndicated a simple headline and URL there, so why bother since you'll see it at the original anyway? Others who are on micro.blog may miss out on part of the conversation, but presumably if they're looking at your copy on micro.blog, they'll be able to see the original as it was intended.

Colin, you mention that not all of your content needs to go to micro.blog. Perhaps, but to think so in my mind is part of the older silo way of thinking. The only reason you're syndicating there is as a stopgap to reach the people who don't currently have the time or luxury to be doing things the way you are. Otherwise they could subscribe to you directly at the source (and potentially even circumscribe the types of posts, keywords, or content to get exactly what they want from your site.) In some sense you syndicate there to reach and communicate with the non-IndieWeb crowd. Perhaps some of your content doesn't make as much sense there as Micro.blog is limited in what it is able to do, but that is its limitation, not yours. Eventually in a fully IndieWeb-ified world, everyone would have their own domain, their own data, and syndication of any sort won't have a real need to exist at all.

As to Jack's comment, syndicating things out to multiple places is often difficult as is getting all the responses back. (Fortunately services like Brid.gy make things far easier though they don't cover all the bases.) I do it in large part because while I prefer to own all of my content and have all the conversation take place on my personal site, I can't necessarily make that choice for everyone else. My mom is likely to never have her own domain much less a site. The only way she'll see my content (whether it's meant for her or not) is to syndicate it to Facebook. For those who aren't yet aware of the IndieWeb or using it, they're still reading and interacting on other platforms, which, for me is fine since I can still have my cake and eat it too. Eventually there will be inexpensive platforms that will let people who don't want to deal with the development cost and overhead that allow much of the IndieWeb-types of functionalities they're not currently getting from their silo platforms for free. I suspect that these will be easier and easier (as well as cheaper) to use over time. I suspect more people will use them for their freedom, flexibility, and increased control. Until then, I have the privilege of using my site much the way I would Facebook, Twitter, Instagram, Google+, Flickr, GoodReads, etc., etc., I just get to do it through a more unified experience instead of having to juggle dozens of accounts and only being able to interact with my fractions of friends, family, and colleagues who coincidentally happen to be spending the time and effort to interact on those websites. As an example, I have dozens of friends who interact with me on Facebook about things I'm currently reading or finished reading, but if I was only doing this on GoodReads, they'd never have a chance to see it as they don't have accounts there or even know it exists. (Coincidentally, this is also the reason that GoodReads and most other silos allow one to syndicate their accounts to Twitter, Facebook, etc.)

Not all social sites are as lucky as Facebook to have such massive adoption. This creates a value imbalance with respect to the classic "network effect" (see: https://en.wikipedia.org/wiki/Network_effect). Philosophically I think that an decentralized and distributed version of IndieWeb philosophies add far more value than having hundreds or even thousands of individual silos.

Invariably some people are likely to stick with Twitter, Facebook, or others because they don't have the same values I do. (Currently I suspect the majority do it because it's frictionless and easy.) But this doesn't change the fact that one can't have a "universal" conversation if one prefers. When I look at various platforms, some of them have different personalities and types of conversations because of the (possibly) self-selecting group of people on them. I loved Twitter more in the early days because of it's smaller and more engaged community--things have naturally changed drastically since those early days. Micro.blog is a bit more like it now, but to me it's not so much a social media replacement for Twitter. To me it's really a social reader that I use to quickly follow a subset of interesting and thoughtful people until I have a better social reader built into my own site. The conversation would be somewhat different if these silos were working on niche audience content like knitting or quantum mechanics, but typically they're not. Most social silos are geared toward mass adoption and broad topic discussion or content posting for everyone/everywhere. Their goal is for their site to be the proverbial "phone number and dial-tone of the web". Why do this when I already have a connection and a "phone number" that is my own site URL?

Jack, while some bloggers have turned off comments, they've often done so saying "Post your reply on your own site, or on Twitter, Facebook, etc." This has just pushed the conversation on their ideas off somewhere else which is disconnected and not as easily searchable or discoverable. And without some kind of notification mechanism, the author of the original post has no idea it exists. (I'll elide a conversation about blocking trolls and abuse here.) I've written some thoughts on comments sections in reply to such a blogger who recently re-enabled comments which also links to several interesting articles about the pros/cons of having comments at all. To me, between webmentions, spam filtering, and even moderation, we're lightyears ahead of where we were in the early 2000s.

One other thing I do find interesting is that the way all of this is set up is allowing us the ability to write extended thoughts and extended multiple replies (with civility as relative strangers). I don't think there are many sites on the web that allow this type of interaction, and they certainly don't do it with anything remotely close to the open architecture we're using. While at times it can be a headache for maintenance and problems, I find far more value in it than using anything else.

 

Replied to a post on github.com :

@Tamaracks There's no need to explicitly set a Post Format in WordPress as it will automatically be set based on the Post Kind setting in Post Kinds on initial publish. (You can also tweak them by hand if necessary, so for example when using the `note` kind it automatically sets the Post Format to `aside`, but I prefer to use `status`, so I replaced `aside` with `status` just under 'note' => array( in the code here: https://github.com/dshanske/indieweb-post-kinds/blob/master/includes/class-kind-taxonomy.php

I also think @dshanske fixed the proper kind settings on micropub in the latest development branch of post kinds as well. https://github.com/dshanske/indieweb-post-kinds/issues/87 (Hopefully this should take care of checkins as well.)

 
 

Replied to a post on github.com :

I think the biggest hurdle to wider adoption is simply the fact that there are so many individual plugins and this takes up far more mental space for the user than it should.

So, another option which I'd like to suggest and advocate for is to **bundle all the plugins into one big single plugin** instead of sub-plugins. You could almost sell it as "the part of WordPress core you always wished you had" and now you can with two clicks: download and activate. (That's got to sound good, even to your mom who's still figuring out how to upload her profile picture.)

From the user's standpoint, this wouldn't require much more than some slightly better UI/descriptions. (And I'm more than happy to write them.) This could consist of a single main settings page with on/off toggles for Post Kinds, Syndication Links, Webactions(?), Micropub, Hum, and IndieAuth. A tabbed interface on this same page with tabs primarily for settings/set up and usage description for all of these (except for maybe Webactions?) would complete the cycle.

Most of the sub plugins don't have many (if any) actual settings other than installation/activation right now which is creating the biggest part of the (mostly mental) hurdle for every day users. I think the average WordPress user probably wouldn't know that they had Webmentions, Semantic Linkbacks, or Webmentions for Threaded Comments installed because they "just work", require no configuration, but are far prettier than any of their predecessors. Why make them carry the mental overhead of what they are and what they do aside from a few subtle lines that they exist? In fact, treating them as if they should have been in WordPress core all along may actually make it more likely to happen.

Additionally things like Micropub which would only have an on/off toggle wouldn't be noticed or used by many unless they had interest in alternate posting interfaces. (And based on the popularity and growth in Twitter interfaces/apps a few years ago, I'm surprised WordPress didn't do this, though perhaps it's part of the reason they're adding a more robust API over the past few years?)

It also means having slightly better or more intuitive explanations of what the individual pieces are (mostly Syndication Links and Post Kinds) near their on/off toggles to better explain what is being activated. Much of this can be taken from the current interface or from the WordPress wiki pages, or added on the individual tabs for the settings for these portions.

I would suggest that doing this would not only make it easier on end users who then wouldn't have to spend the mental space and capacity to keep track of what 10 individual plugins are doing (in addition to the space these take up on the plugin admin page and the fact that, once activated, they disappear from the IndieWeb plugin's list of plugins), but that it would actually dramatically increase the uptake of the single big plugin and its functionality and simultaneously the use of the all the sub plugins individually.

I'd argue that bigger plugins like Yoast SEO or something like PressForward have huge numbers of options and settings and could have been done as separate sub-plugins (the way IndieWeb Plugin is now), but that their value proposition is such that it's well worth spending the handful of minutes reading through the interface to know what the options are, what they mean, and using them to their fullest advantage. I think that Indieweb (and the suite of tools offered on WordPress) is at this tipping point in terms of offering must-have functionality for the future web and that having a simpler integrated set up would help to push it over the edge to broader adoption. (Certainly simpler than the old WP-Social, which users have indicated that they thought was far simpler than Indieweb plugin, though Social actually required more set up.) Additionally all of the seemingly dense text in the "getting started" page could be moved into smaller bit-sized chunks relating to individual portions on a tabbed-interface, for example.

I come to this in part after having spent part of the weekend revamping a bit of the IndieWeb.org documentation on getting started with WordPress and setting up Bridgy with WordPress. A lot of the description is "get this plugin, install, and activate" which takes up a big piece of mental space for the user as well--particularly for the Gen2, 3, 4 users who want a plug and play experience. Far better would be to install one plugin and then modify these handful of settings.

If this is done, then the only remaining (small) hurdle is making sure that the underpinning rel-me data input required of the user is done in a more explicit manner, because this seems to be the lynch-pin holding a lot of it together and making it work. As a result, I'd recommend unbundling the reliance on the User Profile page and put all the rel-me URL fields on their own page in the settings interface for such a single plugin (with all important links just underneath them to encourage users to visit, for example, Twitter's edit profile page to include their website URL in either the website field or in the bio field to enable the bi-directional rel-me.)

Finally a "Tools" tab in the settings page could provide pointer links to additional things like the H-Card Widget or the IndieWeb-PressThis bookmarklets.

When all of this is done, it could also be a simple manner of adding another settings tab to the interface to set up Bridgy with one button links from the plugin to the set up pages for each of the main backfeed services there. Bridgy then automatically checks for the webmention endpoint and checks for rel-me to do it's work, so that part is already automated and relatively user friendly too.

The one caveat I can imagine is that making it all into one big plugin potentially means some small added overhead in development with maintaining some of them as stand alone pieces. I'd recommend keeping them as standalone objects as I honestly believe that pieces like webmentions and micropub are so fundamental to the web, that they should be part of WordPress core and maintaining them separately could help speed this along.

 

Jeremy, Like you, I had some of the same issues and questions when I first started. I'd had a primary website for a while that was a bit more blog-ish on WordPress with a few other subsidiary sites for work related things. When I got into IndieWeb, WithKnown had a great plug-and play set up for almost everything, so for me, it made a nice easy place to start. I also wanted to play with Known and use it to get my feet wet. I was particularly interested in owning a lot of the shorter form posting/microblogging like Twitter without overwhelming my prior subscribers on my WordPress blog with a lot of shorter "fluff". Once I'd gotten a few things working, Known was also incredibly good at quick short posts using bookmarklets (or this mobile solution I'd come up with: http://stream.boffosocko.com/2016/sharing-from-the-indieweb-on-mobile-android-with-apps-and ). Initially the big difference between my sites was where I stored longer form content and shorter form.

Slowly over time, as I've been adding bits and pieces here and there to my WordPress set up, it has been owning more and more while somewhat less lives on my Known Install. There's also been a huge amount of community development on the WordPress side, so that it's tremendously better now than it was when I started and continues to grow. I think I hit a bigger turning point after IndieWebCamp LA when I was able to work out a bit more about how I want to own and syndicate things from my primary hub (on WordPress). (See also: http://boffosocko.com/2016/12/18/rss-feeds-a-follow-up-on-my-indieweb-commitment-2017/ which has some thoughts about adding RSS feeds from one site to cover multiple other sites). The nice part about owning your own site is that you get to figure out what to do with it and you can literally do almost anything. For example, very shortly I'm about to attempt to own all of my checkins on my WordPress site instead of on Known as I've got almost all the moving pieces to do it in a way I like.

Ultimately, I'd recommend doing what works best for you at the moment as there are literally hundreds/thousands of ways to pull it off. Do things one at a time in small bite-sized chunks until you have it the way you want it. There's no reason to do everything all at once. Sometimes I've found that creating posts literally by hand in raw html has been a great way to start to see both how it looks/works, and then use that experience to figure out how to best/most easily automate it going forward.

I'll also note that my coding skills were old and rusty and have been slowly getting better over time as each piece evolves. Another nice part of the process is that you'll have a chance to see how others are doing things with examples from the wiki to help you figure out what might work best for you. Over the process I've also seen others stop, change gears, and even platforms, and try out something totally different. The key seems to be to start with something you know and begin working from there based on your particular "itches" and needs [http://indieweb.org/itches]. Document in the wiki pages what works for you and what doesn't and why to help others along the way. What works for you won't necessarily be the best solution for others and vice-versa, but there's lots of examples out there to work off of.

Your post was certainly a good start for taking stock of what you've got and where you might like to go. Now just make a list of your itches and do things a step at a time. I know there's an IndieWeb Cat [https://indiewebcat.com/] and an IndieWeb Library [http://boffosocko.com/2016/07/16/the-indieweb-ified-library/], so I'm glad to think that there might be some IndieWeb bread out there soon too. (This reminds me that I want to get back to microformats for recipes again in relation to JetPack's shortcode work: https://jetpack.com/support/shortcode-embeds/).

Good luck!

 

@cswordpress I hadn't run across the keyboard issue with the bookmarklet yet, but I'm glad you're taking a look at fixing it. If you get some time, feel free to add your fix and links to the code at http://indieweb.org/accessibility so that others can benefit from your expertise. I know there was some significant discussion about accessibility at one of the early IndieWebCamps in 2016, but I don't remember which.

Several of us in the WordPress IndieWeb "camp" have been pushing to have more microformats2 compatible themes available to increase the selection. If you can push the Genesis framework project to include them, that would be phenomenal! Some of the notes on this work can be found and added to at https://indieweb.org/WordPress/Development#Brainstorming

 

@Chronotope I remember back in the day when Josh Kesselman had that in development...

 

Real-time MRI for precise and predictable intra-arterial stem cell delivery to the central nervous system

An MRI shows stem cells labeled with iron oxide nanoparticles being injected into an animal’s brain. Click to view video. (Credit: Piotr Walczak/Johns Hopkins Medicine)

Working with animals, a team of scientists reports it has delivered stem cells to the brain with unprecedented precision by threading a catheter through an artery and infusing the cells under real-time MRI guidance.

In a description of the work, published online Sept. 12 in the Journal of Cerebral Blood Flow and Metabolism, they express hope that the tests in anesthetized dogs and pigs are a step toward human trials of a technique to treat Parkinson’s disease, stroke, and other brain damaging disorders.

“Although stem cell-based therapies seem very promising, we’ve seen many clinical trials fail. In our view, what’s needed are tools to precisely target and deliver stem cells to larger areas of the brain,” says Piotr Walczak, M.D., Ph.D., associate professor of radiology at the Johns Hopkins University School of Medicine’s Institute for Cell Engineering. The therapeutic promise of human stem cells is derived from their ability to develop into any kind of cell and, in theory, regenerate injured or diseased tissues ranging from the insulin-making islet cells of the pancreas that are lost in type 1 diabetes to the dopamine-producing brain cells that die off in Parkinson’s disease.

Ten years ago, Shinya Yamanaka’s research group in Japan raised hopes further when it developed a technique for “resetting” mature cells, such as skin cells, to become so-called induced pluripotent stem cells. That gave researchers an alternative to embryonic stem cells that could allow the creation of therapeutic stem cells that matched the genetic makeup of each patient, greatly reducing the chances of cell rejection after they were infused or transplanted. But while induced pluripotent stem cells have enabled great strides forward in research, Walczak says they are not yet approved for any treatment, and barriers to success remain.

In a bid to address once such barrier – how to get the cells exactly where needed and no place else – Walczak and his colleague Miroslaw Janowski, M.D., Ph.D., assistant professor of radiology, sought a way around strategies that require physicians to puncture patients’ skulls or inject them intravenously. The former, Walczak says, is not only unpleasant, but also only allows delivery of stem cells to one limited place in the brain. In contrast, injecting cells intravenously scatters the cells throughout the body, with few likely to land where they’re most needed, says Walczak.

“Our idea was to do something in between,” says Janowski, using intra-arterial injection, which involves threading a catheter, or hollow tube, into a blood vessel, usually in a leg, and guiding it to a vessel in a hard-to-reach spot like the brain. The technique currently is used mainly to repair large vessels in the brain, says Janowski, but the research team hoped it might also be used to get stem cells to the exact place where they were needed. To do that, they would need a way of monitoring the catheter placement and movement of implanted cells in real time.

Walczak and Janowski teamed with colleagues including Monica Pearl, M.D., an associate professor of radiology practicing in the Division of Interventional Neuroradiology, who specializes in intra-arterial procedures. Usually the procedure is performed using an X-ray image as a guide, but that approach ruled out watching injected stem cells’ movements and making adjustments in real time.

In their experiments, after placing the catheter under X-ray guidance, they transferred anesthetized dog and pig subjects to an MRI machine, where images were taken every few seconds throughout the procedure. Once the catheter was in the brain, Pearl pre-injected small amounts of a harmless contrast agent that included iron oxide and could be detected on the MRI. “By using MRI to see in real time where the contrast agent went, we could predict where injected stem cells would go and make adjustments to the catheter placement, if needed,” says Janowski. Adds Jeff Bulte, Ph.D., a professor of radiology who participated in the study, “It’s like having GPS guidance in your car to help you stay on the right route, instead of only finding out you’re lost when you arrive at the wrong place.”

The team then injected both small stem cells (glial progenitor cells from the brain) and large mesenchymal stem cells from bone marrow into the animals under MRI, and found that in both cases, the pre-injected contrast agent and MRI allowed them to accurately predict where the cells would end up. They could also tell whether clumps of cells were forming in arteries and, if so, possibly intervene to avoid letting the clumps grow large enough to cut off blood flow and pose a danger. “If further research confirms our progress, we think this procedure could be a big step forward in precision medicine, allowing doctors to deliver stem cells or medications exactly where they’re needed for each patient,” says Walczak. The research team is planning to test the procedure in animals as a treatment for stroke and cancer, delivering both medications and stem cells while the catheter is in place.

Other authors on the paper are Joanna Wojtkiewicz, Aleksandra Habich, Piotr Holak, Zbigniew Adamiak and Wojciech Maksymowicz of the University of Warmia and Mazury in Poland; Adam Nowakowski and Barbara Lukomska of the Mossakowski Medical Research Center in Poland; Jiadi Xu of the Kennedy Krieger Institute; and Moussa Chehade and Philippe Gailloud of the Johns Hopkins University.

The study was funded by the National Institute of Neurological Disorders and Stroke (grant numbers NS076573, NS045062, NS081544), the Maryland Stem Cell Research Fund, the Department of Defense (grant number PT120368), the Polish National Science Centre (grant number NCN 2012/07/B/NZ4/01427), the National Centre for Research and Development, and a Mobility Plus Fellowship from the Polish Ministry of Science and Higher Education.

 

Attending: Improving your Drupal 8 development workflow using Composer | DrupalCamp LA 2016 https://2016.drupalcampla.com/sessions/improving-your-drupal-8-development-workflow-using-composer

 

Replied to a post on github.com :

Potential future features · Issue #2 · kshaffer/hypothesis_aggregator https://github.com/kshaffer/hypothesis_aggregator/issues/2

This plugin is spectacular by the way... Here are some thoughts on future development:

It would be great if the displayed text output from the shortcode included a permalink to the original annotation on Hypothesis. As a UI suggestion, perhaps you could place it after the title of the annotated page thusly:
URL WRAPPED TITLE | < a href="permalinkhere">#</a>

Being able to circumscribe dates for displayed annotations could be useful, particularly in cases where one wants to highlight only a handful of annotations from a particular user with a particular tag but specific ones within certain time parameters. Additionally pages with embeds could, over time, require importing huge numbers of annotations which may not be ideal, or the context may change with significant future annotations that aren't relevant to what could be a static post that's trying to highlight just a few particular annotations. Time parameters could also prevent future possible spam on unwatched embeds with only "tag" parameters that could be targeted by bad actors.

 

Thinking about how @WordPress might drive development by adding additional categories to their Feature Filter https://wordpress.org/themes

 

Replied to a post on github.com :

You could start with http://indiewebcamp.com/Getting_Started_on_WordPress which will walk you through most of the process. You should start with the Webmention Plugin and Semantic Linkbacks to send/receive webmention support and to help comments look better respectively. Brid.gy connects your blog to responses from sites such as Facebook, Twitter, et al.

Keep in mind that some of the other plugins you'll run across may have small bugs or quirks as they're being actively developed. Most of the plugins in the WordPress.org repository are usually more stable than some you'll run across on GitHub, but those also are dated in terms of development work, so you'll have a better idea how current they are. You can usually find one or more of us via IRC, or in social media if you need help.

 

Replied to a post on github.com :

It looks like no one is maintaining this actively and the last patch was in 2014. Just keeping up with Twitter and Facebook API changes can be painful.

You might find similar functionality using a group of IndieWeb related plugins https://wordpress.org/plugins/indieweb/ much of which is in active development on GitHub as well. Documentation and additional information can be found at http://indiewebcamp.com/wordpress.

Some of the rest of the autoposting portions can be had through JetPack's Sharing functionality or alternately through NextScript's SNAP plugin: https://wordpress.org/plugins/social-networks-auto-poster-facebook-twitter-g/

 

17w5131: Statistical & Computational Challenges in Large Scale Molecular Biology Workshop @BIRS_Math 3/2017 #ITBio

Arriving in Banff, Alberta Sunday, March 26 and departing Friday March 31, 2017

Organizers

  • Barbara Engelhardt (Princeton University)
  • Anna Goldenberg (University of Toronto)
  • Manolis Kellis (Massachusetts Institute of Technology)
  • Jacob Laurent (Centre national de la recherche scientifique)
  • Jeff Leek (John Hopkins University)
  • Stephen Montgomery (Stanford University)

Objectives

Over the past few years, an increasing number of large scale data sets have been made available in molecular biology. GTEx, for example, produced more than 18,000 RNA-Seq assays for multiple tissues in 900 individuals, Mindact generated gene expression data from about 7000 breast tumors in a single study, and 23andMe claims to have sequenced about 900,000 genomes. This growth in the available genomic data is expected to increase our capacity to identify cancer subtypes, regulatory genes, SNPs associated with phenotypes of interest, and biomarkers for many human traits. It also suggests exploring more complex feature representations when analyzing these datasets.

However, increasing the number of samples and features leads to a set of \textbf{interrelated statistical and computational problems}. Accordingly, the objectives of our workshop will be to:

Systematically identify the statistical and computational

problems arising during the analysis of large scale data in molecular biology;

Bring together experts in computational biology, molecular

biology, computer science, and statistics to propose innovative solutions to these problems, by leveraging recent advances in each of these fields.

Relevance, importance and timeliness

A number of studies generating high throughput molecular data for a large number of biological samples have been completed over the past five years. \textbf{Our workshop is important because the availability of these datasets holds great promises in terms of health improvement and understanding of molecular biology}. First of all, if exploited correctly, larger sample sizes should improve our ability to predict phenotypes of interest from molecular data. This entails very important applications such as improving the survival of cancer patients by better predicting which treatment they should receive, or decreasing bacterial resistances by predicting which antibiotic is efficient against a new strain. Correctly exploiting large scale datasets should also allow us to \textbf{better identify genetic and epigenetic determinants of these phenotypes, yielding a better understanding of human diseases and potentially guiding the development of new treatments and prevention policies}. In particular, more samples should allow the detection of less frequent variants in the human genome, or more complex features involving several modalities (copy number, expression, methylation, etc) associated with diseases. Finally, larger sample sizes should help with essential unsupervised tasks such as the \textbf{inference of regulation networks, or the identification of cancer subtypes}.

Our workshop is relevant because \textbf{all of these promises are conditioned on our solving of new statistical and computational challenges}. First (Challenge 1), we need to build new feature spaces and estimators whose complexity is adapted to these larger sample sizes, which involves designing novel, potentially more complex descriptors of the samples but still controlling the bias/variance trade-off. Second (Challenge 2), we need to build models which correctly integrate different modalities, such as copy number variation and gene expression. Third (Challenge 3), larger scale studies are more prone to unwanted variations, because they typically involve different labs and technical changes which can affect the measurements and become confounders in retrospective analyses. Similar or worse problems arise when trying to combine several existing datasets. We need methods which take this unwanted variation into account. Finally, (Challenge 4), we need new algorithms that make existing statistical tools scalable to the new sample sizes, and make estimation over the larger and more complex features of Challenge 1 tractable.

We also believe our workshop is very timely because \textbf{some of these statistical and computational challenges are starting to be addressed in other application fields} of statistics. It is crucial to recognize that the orders of magnitude are still very different in molecular biology and other data science application fields because of the cost and complexity of the data generation process: current large scale high throughput sequencing data sets typically contain a few thousand of samples but millions of features while computer vision, web, or astronomy datasets can involve billions or trillions of samples and relatively fewer features. A first consequence is that not all recent developments in machine learning are immediately transferable to computational biology. For example, so called deep learning methods have gained a lot of popularity and now represent the state of the art in computer vision but may not be the most appropriate tool for prediction of cancer outcome from molecular data. However, the fact that other fields already have much larger sample sizes also means that they had to develop efficient and scalable algorithms for basic tasks like feature selection, classification or clustering. \textbf{These recent developments are a great source of inspiration for computational biology, where large scale computation is still an emerging challenge}.

We believe \textbf{having a small scale workshop involving international experts in machine learning, statistics, computational biology and molecular biology is of utmost importance} for three main reasons. The first reason is that the technical advances we are referring to are very recent, often unknown to computational biologists and involve paradigms such as online optimization, accelerated gradient methods and network flow optimization, with which they are sometimes unfamiliar. The second reason is that it is not always obvious to non-statisticians which novel methods are appropriate given the current n/p regime. Conversely, the third reason is that statisticians do not know what the recent challenges are in molecular biology. Having them work on abstract versions of the problems is often not satisfactory as it is necessary to be aware of technical realities and of the underlying biology of the problem to come up with useful solutions.