Skip to main content

Biomedical and Electrical Engineer with interests in information theory, evolution, genetics, abstract mathematics, microbiology, big history, IndieWeb, mnemonics, and the entertainment industry including: finance, distribution, representation

boffosocko.com

chrisaldrich

chrisaldrich

+13107510548

chris@boffosocko.com

stream.boffosocko.com

www.boffosockobooks.com

chrisaldrich

mastodon.social/@chrisaldrich

micro.blog/chrisaldrich

 

@kristenhare A growing group of journalists are joining the IndieWeb movement to better own their work and data within their own personal archives as well as using tools like Ben's to archive their work on larger institutional repositories. There's a stub page on the group's wiki dedicated to ideas like this for journalists at https://indieweb.org/Indieweb_for_Journalism

Coincident with these particular sites disappearing, there's now also news today that Peter Thiel may purchase Gawker in a bid to make it disappear from the internet, which makes these tools all the more relevant to the thousands who wrote for that outlet over the past decade.

For journalists and technologists who are deeply committed to these ideas, I'd recommend visiting the Reynolds Journalism Institute. They just finished a two day conference entitled "Dodging the Memory Hole" at the Internet Archive last week focused on saving/archiving digital news in various forms. (https://www.rjionline.org/events/dodging-the-memory-hole-2017) Naturally most of the conference was streamed and is available on YouTube (as well as archived.) Keep your eyes peeled for next year's conference which typically occurs in November.

 

The more I think about archiving the web this week, the more value and stability I think that the W3C's Webmentions spec could be adding to the internet and copies of it.
https://indieweb.org/Webmention

 
 

, @palewire has built a nice tool for archiving your stories
https://twitter.com/palewire/status/930812507316297728

 

Are you archiving your own Tweets? Try following @LinkArchiver
https://twitter.com/LinkArchiver/status/883426532282249216

 

@andreweckford I'll wait to see what the "webmaster" posts later. :) Thanks for the live tweeting you're doing, particularly the linked stream. I know you're most of the way through, but next time you might give http://www.noterlive.com/ a try for live tweeting. It allows you to quickly search for and change speakers (and their twitter handles), does auto-threading, as well as does auto hashtagging for following the conversation. This'll free you up for simply typing and hitting enter and not fiddling with all the other bits along the way.

It will also save your entire tweet stream from the conference in a simple format so you can cut and paste your work into a blog post quickly after-the-fact for archiving into the conference website (or your own).

It's one of the best conference tools I've come across for this type of thing.

Instruction manual if you need it: https://github.com/kevinmarks/noterlive/wiki/Noter-Live-Instruction-Manual

 

@kevinmarks I like sidebar with stream for the hashtag that allows you to pull tweets into your notes for archiving for later.

 
 

For getting more value in following/interacting with colleagues on Twitter, particularly in relation to conference/topic hashtags, I recommend setting up an IFTTT.com "filter" to add people who tweet about a particular topic(s) to a Twitter list. This can save a LOT of time hunting, searching, and doing this by hand. Details and an example of this can be found at "Some thoughts on creating conference lists, live tweeting and archiving events" [http://boffosocko.com/2016/10/12/twitter-list-for-dtmh2016-participants-dodging-the-memory-hole-2016-saving-online-news/ ]

 

The nice part too is that at the end of the day you own and host your own data. In a similar vein, one might also look at Calibre [https://calibre-ebook.com/], which is free software that is somewhat like iTunes for books/e-books. It's highly configurable and in addition to also archiving and storing actual ebook versions of your books (and magazines, newspapers, or other digital files), it allows for additional common metadata or even custom fields. It also has a variety of plugins including one for GoodReads which allows you to sync your book data from GoodReads to your own computer/server.

 

Replied to a post on github.com :

Cross referencing #889 which may be related,

The items that won't delete are from feeds from mid-December (the feeds of which have since been deleted). They not only exist in the database, but their `post status` is still marked as `publish` after being "deleted".

It appears that for currently extant/live feeds that the delete button "works" properly for them in that `post status` has been changed to `removed_pf_feed_item` and they don't re-appear in the interface upon refresh. I am surprised that the entire post wasn't deleted from the database however--is this a bug or a feature? I might expect that archiving a post would hide it while still keeping it in the database, but not delete.

Currently using v4.2.1 on WP 4.7.1.

 

Feature request: Archive internal page/post links to Internet Archive on publish/update · Issue #23 · ManageWP/broken-link-checker https://github.com/ManageWP/broken-link-checker/issues/23

I might suggest the following functionality could fit in well with the plugin's general purpose, particularly since the Internet Archive recommends the plugin. It could also help to "close the loop" in the plugin's overall functionality for helping to maintain data integrity and working links for WordPress sites on the web.

**Suggested Functionality**
When one initially publishes (or possibly updates) a post/page, it would be awesome if all of the URLs referenced on the page as well as that of the page itself were pinged for archiving to the Internet Archive's Wayback Machine at the day and time of their being referenced in the post. (Or perhaps within a day or two of the post so as not to overwhelm Archive.org's servers with multiple subsequent updates for typos/tweaks which invariably happen post-publication.)

With this functionality, then in the future, if (when) resources change, move, etc. one could use the primary functionality in Broken Link Checker not only to restore a link to a close or reasonable copy of the original, but restore it to the _same_ day it was originally referenced.

As I'm sure you're all too aware, this can be very handy as the average web page has a lifespan of 100 days or less. I can see this being very useful to not only the general public, but particularly for bloggers, linkbloggers, journalists, and academics.

**Implementation**
To my knowledge, there are no plugins within the WordPress repository that manage this type of functionality as a standalone plugin, though there is the heavily underused [Post Archival in the Internet Archive](https://wordpress.org/plugins/post-archival/) plugin which essentially adds one's individual post/page permalink URL to the Internet Archive as it's published, though it doesn't include the archival of any links (references) within that same post.

I'm sure that Archive.org/The Wayback Machine may provide some additional documentation for implementation, though I suspect the code in the above-referenced pluign is a very good examplar. I also recently came across this snippet: https://indieweb.org/Internet_Archive#Trigger_an_Archive after a recent conference on Saving/Archiving News Sites which may be beneficial as well. It certainly exhibits at least an interest/demand for such a functionality.

I haven't dug into WordPress core, but I'm guessing some of the functionality for parsing URLs within pages/posts for sending Trackbacks/Pingbacks would have a filter or hook for providing all of the URLs in a post/page necessary for such processing and archiving while the_permalink() or get_permalink() gives the last.

Given the popularity of this spectacular plugin, it could also potentially become one of the largest forces for archiving vast swaths of the internet to the Internet Archive, short of WordPress adding such functionality directly into core.

 

@OpenStudy Are you going to provide a mechanism for data export? How about archiving the data to @InternetArchive before you turn off system?

 

PubSub (previously PubSubHubbub) is now a Public Working Draft from the W3C https://www.w3.org/TR/pubsub @MarkGraham