facebook, social media, media, social, internet, network, blog, seo, web, marketing, business, website, design, symbol, icon, online, search, optimization, communication, strategy,. Internet culture archives: a practical reference
Photo by Firmbee on Pixabay

Reviews

Internet culture archives: a practical reference

Internet culture archives come in four kinds, each with a characteristic gap, plus how to cite from them and the ethics of keeping what was meant to pass.

The internet is often described as a place where nothing is ever deleted. The opposite is closer to true. Most of what has been posted is gone, most of what remains is unfindable, and the surviving fraction was selected by accidents that have nothing to do with what mattered. The archives that exist are heroic and partial, and knowing which kind of archive you are holding is the difference between using it and being misled by it.

This page describes the kinds, what each keeps and loses, how to cite from them without overclaiming, and the ethical problem of preserving what was meant to pass.

What to take away

  • There are four kinds of archive, made by different people for different reasons, and each has a characteristic gap.
  • Text survives; everything else decays. Anything that needed specific software to display is probably unopenable, and anything behind a login was probably never captured.
  • An archived copy proves that something existed at an address by a date. It does not prove it existed nowhere earlier, and it does not prove the page looked like that to anyone.
  • Preservation is not neutral. Saving what people expected to vanish is a decision made on their behalf, and the field has not settled how to make it.

The four kinds

Kind Who made it and why What it keeps well Its characteristic gap
Crawl archives Institutions or projects running automated capture across the public web Public, static, popular pages, captured repeatedly Anything behind a login, anything built by script in the reader's browser, anything obscure enough to be crawled rarely or never
Community wikis and catalogs Enthusiasts documenting a subject they care about Context, explanation, the connections between items, the vocabulary Accuracy; the entries inherit each other's origin claims, and a wiki is a consensus, not a record
Personal collections Individuals saving what they liked Material that was never public, or was public briefly; the intimate and the ephemeral Coverage and provenance; a saved file has no record of where it came from, and the collector's taste is the selection rule
Institutional deposits Libraries, museums and universities accepting donated material Stability, cataloging, long-term access Lag and formality; material arrives years late and only from donors who thought to donate

The gaps are complementary, which is why serious work uses all four and trusts none alone. A crawl archive will have a capture of a page but not the conversation around it; a community wiki will have the conversation but assert an origin it cannot support; a personal collection will have the one file nobody else saved and no way to date it.

What survives, and why it is not random

The digital preservation problem is usually described as technical, and it partly is: formats become unreadable, media degrades, software stops running. But the pattern of loss has a shape, and the shape biases everything built from the remains.

Text survives best. It is small, it is easy to copy, and it migrates between formats with almost no loss. A plain message from decades ago is as readable now as the day it was written.

Images survive moderately, with degradation. Each copy is re-encoded, cropped or captioned, and the version that survives is usually several generations from the original, with metadata stripped along the way.

Anything interactive or dependent on a specific runtime survives worst. A page assembled by script, a piece built for a plugin nobody runs, a game, an animated format: these are often captured as shells, and the shell is preserved while the thing it held is gone. The fear of a digital dark age, a period whose record is unreadable to its successors, is a fear about this layer specifically.

Popular things were copied more and therefore survived more. This is the bias that matters most for this site's subject: the surviving record over-represents whatever was already widely seen, so a history built from it concludes that the period was dominated by the things that happened to be saved, which proves nothing. The same circularity affects any account of how a format spread, described under memes and virality.

And private things left nothing. The communities that were closed by design, which includes most small ones, produced conversation no crawler reached. A history of online community written from archives is a history of the public and the large.

Citing an archived item without overclaiming

An archived copy supports a narrow claim, and most citations stretch it. What a capture actually establishes:

  • That content existed at that address on the capture date. Not earlier, not elsewhere, and not that the address was the original one.
  • That the capture tool saw that content. A page can be served differently to a crawler than to a person, and a capture of a script-heavy page may show an empty frame that no human ever saw.
  • That the captured version is what it is. It does not establish that the version was the first, the last or the widely seen one.

A citation that respects these limits names the address, the capture date, the archive, and the date the citation was checked, and says which of the four kinds of archive it came from. That is more than most citations carry and it is the minimum that lets a later reader reproduce the check. The reasons the check has to be reproducible, and the ways archived material fails when built into a chronology, are set out under platform timelines.

Provenance in a personal collection

The most valuable material is often in the fourth kind, the personal collection: the file somebody saved that no crawler captured. It is also the material with the least provenance. A saved image carries no record of where it was saved from. Its file date is the date it was written to that disk, which may be a copy of a copy. Its metadata may be original, or stripped, or rewritten by whatever service it passed through.

What can be done: record what the collector remembers, marked as recollection; record what the file itself says, marked as file metadata; and look for the same item in one of the other three kinds to anchor a date. What cannot be done is treat the file as dated. The early social networks page describes why so much of that period's record exists only in this form, and why its dates are the least reliable thing about it.

The ethics of keeping what was meant to pass

Much of what archives hold was posted by people who expected it to vanish, in places they thought were small, under names they have since abandoned. Preserving it makes a decision on their behalf, and the decision is not obviously right.

The argument for preservation is that the record of a period belongs to everyone who comes after, that the people who posted are part of the history, and that the alternative is a period with no record at all. The argument against is that a message to a room of thirty was not a publication, that the person who wrote it did not consent to a permanent public copy, and that an archive can turn a discarded persona into a permanent identity for someone who moved on.

The field has not resolved this, and this site does not pretend to. What it does is decline to reproduce material from small or closed spaces, describe patterns rather than people, and treat visible as different from published. An archive is a tool, and the people inside it were not consulted about being in it.

Common questions

If a page is not in the crawl archives, did it exist?

Very possibly. Coverage is uneven, skewed toward the public and popular, and the absence of a capture is the normal condition for most of what was ever posted.

Which kind of archive is most reliable?

For dates, the crawl archives, within their coverage. For context, the community catalogs, with their origin claims discounted. For material nobody else has, personal collections, with no reliable dates. Institutional deposits for stability. None for everything.

Can an archived copy be wrong?

It can be a capture of a page as served to a crawler rather than as seen by a person, it can be a capture of a shell with the content missing, and it can be a capture of a version that was live for an hour. It is accurate about what it captured and silent about what it did not.

Should I archive material from a small community I belong to?

Ask the community. The answer will tell you something about the community, and it is their record.

More in Reviews

Reviews

How to make sense of memes and virality platforms

Memes and virality on platforms that bind them: four dependencies that stop a format traveling, and what actually crosses the boundary instead.

Latest from Reporting Desk