When the Web Becomes Evidence
My first Digital Humanities reflection
One of the first ideas I am carrying into Digital Humanities is that the internet is not the same thing as a permanent record. It feels permanent because information is easy to copy, search, and share, but the readings for Weeks 1 and 2 show how quickly digital evidence can disappear. A page can vanish when a company closes, a platform changes its policies, a server is not renewed, or a legal dispute changes what is allowed to remain online.
A history that can disappear
Chris Stokel-Walker’s BBC article begins with a statistic that made the problem concrete: research found that 25 percent of web pages posted between 2013 and 2023 had vanished, with older pages disappearing at an even higher rate. The article presents the Internet Archive and its Wayback Machine as an important response. Volunteers and archivists have built an enormous collection of snapshots, books, videos, and other digital materials so that researchers can still encounter pieces of the web that no longer exist in their original form.
What surprised me was how fragile that solution is. The Internet Archive is doing public-facing historical work, but it still depends on money, storage, technical infrastructure, volunteers, and legal permission. Cyberattacks and lawsuits can threaten the archive itself. That tension changed the way I think about “the cloud.” Digital information may look weightless, but it depends on physical data centers, institutions, and people willing to maintain it.
The article also made me think about what future historians will not be able to see. Much of our everyday life now happens through websites, social media, online communities, and digital services. If those spaces disappear without being preserved, the loss is not only technical. It changes the evidence available for understanding how people lived, communicated, organized, and made culture.
From scarcity to abundance
Roy Rosenzweig’s “Scarcity or Abundance? Preserving the Past in a Digital Era” approaches the same problem through history. His example of the “Bert Is Evil” website is memorable because it shows that even a strange piece of internet culture can become historical evidence. The site disappeared after its creator deleted it, yet copies and mirrors survived in other places. That story raises a difficult question: what counts as worth preserving, and who gets to decide?
Rosenzweig describes a shift from an older age of scarcity to a digital age of abundance. In the past, historians often worried that too few documents survived. In the digital era, we may face the opposite problem: too much information is produced too quickly, and much of it is poorly organized, undocumented, or dependent on software that may not exist in the future. Abundance does not automatically create better history. Researchers still need ways to select, describe, preserve, interpret, and teach from the record.
Rosenzweig’s argument also makes preservation feel less like a purely technical assignment. File formats and storage systems matter, but so do copyright, funding, public policy, institutional responsibility, and access. The Internet Archive is powerful precisely because it makes digital history available to the public, but its private and legally uncertain position shows why preservation cannot depend on one organization or one passionate founder alone.
Why this matters to a Data Science major
These readings connected directly to the way I think about datasets. A dataset can appear complete while still reflecting decisions about what was collected, labeled, excluded, or lost. A digital archive works the same way. Its search box may make a collection feel neutral, but the collection is shaped by crawl schedules, metadata, file formats, language support, permissions, and the choices of the people who built it.
The biggest takeaway for me is that preservation is also a question of representation. If a community’s websites, images, posts, or local histories are not archived, future researchers may mistake absence for unimportance. That is why Digital Humanities asks both technical and humanistic questions: How do we store a record, and whose record are we storing?
Alongside the readings, I began learning how web archiving works by exploring the Wayback Machine. The basic process is simple: enter a web address, choose an available date, and compare an archived snapshot with the live or missing page. The skill is not only clicking through old pages. It is learning to treat an archived page as evidence with a date, a capture history, and limitations.
The challenge is that an archive is never a perfect copy of the web. Images, scripts, links, and interactive features may be missing, and a page may have been captured only once. I managed that uncertainty by paying attention to the capture date and by treating the snapshot as one source rather than the entire story. If I taught this skill to another student, I would begin with a familiar website that has changed over time, then ask them to record what survived, what disappeared, and what questions the archive cannot answer.
What I am carrying forward
Weeks 1 and 2 taught me that the digital past is threatened by both scarcity and abundance. Some records disappear before anyone saves them, while other records accumulate so quickly that future researchers may struggle to understand them. My responsibility as a student is not to assume that digital information will take care of itself. I need to ask how a source was created, how it was preserved, what context it lost, and who can still access it.
That question will stay with me as I continue studying Digital Humanities: when I find a digital source, am I only using it, or am I also helping understand its history?