Sarvix Codes
DHUM field notes · Rowan University

Data has a story before it becomes a dataset.

I’m Naman Parmar, a Data Science major exploring Digital Humanities, digital archives, and the human questions inside technical systems.

01 Weekly reflection Tool Hypothesis annotation Focus Visibility & power
What I’m carrying forward

Three ideas from Week 3.

Digital Humanities is not only about using tools. It is about understanding how tools shape knowledge, access, and memory.

01 / RECORD

Digital records are constructed.

Archives and datasets inherit decisions about what is collected, described, preserved, and made discoverable.

02 / SEARCH

Speed changes the question.

Full-text search expands discovery across borders, but the easiest sources to find are not automatically the most representative.

03 / PRACTICE

Technical skill is ethical skill.

Building, mapping, annotating, and analyzing can become forms of public work when they are done with context and care.

When the Web Becomes Evidence

My first Digital Humanities reflection

One of the first ideas I am carrying into Digital Humanities is that the internet is not the same thing as a permanent record. It feels permanent because information is easy to copy, search, and share, but the readings for Weeks 1 and 2 show how quickly digital evidence can disappear. A page can vanish when a company closes, a platform changes its policies, a server is not renewed, or a legal dispute changes what is allowed to remain online.

A history that can disappear

Chris Stokel-Walker’s BBC article begins with a statistic that made the problem concrete: research found that 25 percent of web pages posted between 2013 and 2023 had vanished, with older pages disappearing at an even higher rate. The article presents the Internet Archive and its Wayback Machine as an important response. Volunteers and archivists have built an enormous collection of snapshots, books, videos, and other digital materials so that researchers can still encounter pieces of the web that no longer exist in their original form.

What surprised me was how fragile that solution is. The Internet Archive is doing public-facing historical work, but it still depends on money, storage, technical infrastructure, volunteers, and legal permission. Cyberattacks and lawsuits can threaten the archive itself. That tension changed the way I think about “the cloud.” Digital information may look weightless, but it depends on physical data centers, institutions, and people willing to maintain it.

The article also made me think about what future historians will not be able to see. Much of our everyday life now happens through websites, social media, online communities, and digital services. If those spaces disappear without being preserved, the loss is not only technical. It changes the evidence available for understanding how people lived, communicated, organized, and made culture.

From scarcity to abundance

Roy Rosenzweig’s “Scarcity or Abundance? Preserving the Past in a Digital Era” approaches the same problem through history. His example of the “Bert Is Evil” website is memorable because it shows that even a strange piece of internet culture can become historical evidence. The site disappeared after its creator deleted it, yet copies and mirrors survived in other places. That story raises a difficult question: what counts as worth preserving, and who gets to decide?

Rosenzweig describes a shift from an older age of scarcity to a digital age of abundance. In the past, historians often worried that too few documents survived. In the digital era, we may face the opposite problem: too much information is produced too quickly, and much of it is poorly organized, undocumented, or dependent on software that may not exist in the future. Abundance does not automatically create better history. Researchers still need ways to select, describe, preserve, interpret, and teach from the record.

Rosenzweig’s argument also makes preservation feel less like a purely technical assignment. File formats and storage systems matter, but so do copyright, funding, public policy, institutional responsibility, and access. The Internet Archive is powerful precisely because it makes digital history available to the public, but its private and legally uncertain position shows why preservation cannot depend on one organization or one passionate founder alone.

Why this matters to a Data Science major

These readings connected directly to the way I think about datasets. A dataset can appear complete while still reflecting decisions about what was collected, labeled, excluded, or lost. A digital archive works the same way. Its search box may make a collection feel neutral, but the collection is shaped by crawl schedules, metadata, file formats, language support, permissions, and the choices of the people who built it.

The biggest takeaway for me is that preservation is also a question of representation. If a community’s websites, images, posts, or local histories are not archived, future researchers may mistake absence for unimportance. That is why Digital Humanities asks both technical and humanistic questions: How do we store a record, and whose record are we storing?

Digital practice I began: exploring the Wayback Machine

Alongside the readings, I began learning how web archiving works by exploring the Wayback Machine. The basic process is simple: enter a web address, choose an available date, and compare an archived snapshot with the live or missing page. The skill is not only clicking through old pages. It is learning to treat an archived page as evidence with a date, a capture history, and limitations.

The challenge is that an archive is never a perfect copy of the web. Images, scripts, links, and interactive features may be missing, and a page may have been captured only once. I managed that uncertainty by paying attention to the capture date and by treating the snapshot as one source rather than the entire story. If I taught this skill to another student, I would begin with a familiar website that has changed over time, then ask them to record what survived, what disappeared, and what questions the archive cannot answer.

What I am carrying forward

Weeks 1 and 2 taught me that the digital past is threatened by both scarcity and abundance. Some records disappear before anyone saves them, while other records accumulate so quickly that future researchers may struggle to understand them. My responsibility as a student is not to assume that digital information will take care of itself. I need to ask how a source was created, how it was preserved, what context it lost, and who can still access it.

That question will stay with me as I continue studying Digital Humanities: when I find a digital source, am I only using it, or am I also helping understand its history?

The Search Box Is Not Neutral

My first Digital Humanities reflection

Before this week, I usually thought of digital tools as ways to make research faster. Search engines find information, databases organize it, maps display it, and code helps us process it. As a Data Science major, I am used to thinking about tools in terms of efficiency: How much information can I collect? How quickly can I find a pattern? How can I make a result easier to interpret?

The Week 3 readings made me slow down and question that mindset. Digital Humanities is not just about applying technology to humanities material. It is also about asking what technology changes, what it hides, and who gets represented when information becomes digital.

Digital Humanities as intervention

In the introduction to New Digital Worlds, Roopika Risam describes postcolonial Digital Humanities as an intervention in the “digital cultural record.” Her example of the Puerto Rico Mapathon made this idea concrete. After Hurricane Maria, students, university staff, and community members worked with OpenStreetMap to improve geographic information that could support disaster relief. This was not technology for technology’s sake. Mapping became a form of public service, collaboration, and teaching.

What stood out to me was the connection between theory, praxis, and pedagogy. Theory helps us understand the historical and political problems inside digital systems. Praxis means acting on those problems by building or improving tools, archives, maps, and workflows. Pedagogy means teaching other people how to participate. The Mapathon brought all three together: participants thought critically, performed technical work, and helped others learn the process.

Risam also argues that the digital cultural record contains gaps, omissions, and older colonial patterns. Digital projects can preserve and share knowledge, but they can also repeat the same inequalities found in older archives. If some languages, communities, and histories were underrepresented in the original cultural record, digitizing that record does not automatically make it fair. The process of choosing what to scan, how to describe it, how to organize it, and who can access it all matters.

That point connected strongly to my Data Science background. A dataset is never simply “the truth.” It reflects decisions about collection, categories, missing values, labels, and access. A digital archive has similar decisions built into it. The database may look objective because it is organized and searchable, but its structure still reflects human choices and institutional power.

The opportunities and blind spots of search

Lara Putnam’s article, “The Transnational and the Text-Searchable,” focuses on how digitized sources and full-text search have changed historical research. Search makes it possible to discover people, places, words, and connections across large collections without physically traveling to every archive. Putnam calls attention to the value of “side-glancing”: looking beyond the boundaries of a familiar topic or location to notice unexpected connections.

This is exciting to me as someone who works with data. Search increases the scale and speed of discovery. It can reveal patterns that would have been difficult or impossible to find through older research methods. Putnam explains that digital search has helped historians think across national borders because information is no longer as tightly connected to one physical place.

At the same time, the article warns that easier discovery can create new forms of ignorance. Search results are shaped by what has been digitized, what can be recognized by optical character recognition, what language the system handles well, and what terms a researcher already knows to enter. The sources that are easiest to find can begin to look like the most important sources, even when they are only the most searchable ones.

Digital search does not simply open a window onto the past. It creates a particular window, with a particular frame.

Putnam’s argument changed the way I think about a search box. It can help us see connections, but it can also make us overlook local context, people who left fewer written records, and communities whose materials were never digitized. In data terms, search is powerful, but it still has coverage problems, bias, and blind spots.

Digital skill I practiced: Hypothesis annotation

The new tool I practiced this week was Hypothesis, a web-based annotation tool. Instead of reading a PDF silently and keeping all of my reactions separate from the text, I could select a passage, highlight it, and attach a note directly to the relevant sentence. I could also read other annotations and see how classmates interpreted the same passage.

The basic workflow was simple: select text, choose Annotate, write a response, and then return to the page to read the conversation around it. The harder part was deciding what deserved an annotation. I became more intentional when I used three questions: What is the author claiming? What does this connect to? What do I want to question?

The biggest challenge was moving between close reading and the interface. Annotation adds another layer to reading because I was not only trying to understand the author’s argument; I was also trying to explain my own reaction clearly enough for someone else to follow. Reading classmates’ notes showed me that a passage can generate several legitimate questions. Annotation became less like finding one correct answer and more like making my thinking visible.

If I taught Hypothesis to another student, I would begin with a short passage and ask them to write one observation, one connection, and one question. I would also explain that a useful annotation does not have to be long. It should make a specific part of the reading easier to discuss.

What I am taking forward

My main takeaway from this week is that Digital Humanities is both technical and ethical. It asks us to build, search, map, annotate, and analyze, but it also asks us to examine the assumptions inside those actions. The goal is not simply to add more information to the internet. The goal is to create digital knowledge more carefully, with attention to context, representation, access, and power.

As a Data Science major, I am interested in the possibilities of computation. This week reminded me that computational skill is strongest when it is paired with humanistic questions. Before trusting a dataset, archive, search result, or visualization, I should ask: Who created it? What was left out? Which communities are represented? What would I fail to see if I only followed the easiest results?

That is the question I want to carry into the rest of this course—and into the rest of my work with data: not only “What can this tool show me?” but also “What is this tool making harder to see?”