The notebook is a school exercise book with a blue cover, the kind you buy at any stationery shop in Bengaluru for twelve rupees. Its pages are ruled, soft, already curling at the edges from the humidity of a monsoon that has barely begun. On the first page, in neat Kannada handwriting, is a list of house numbers. Not addresses—these houses do not have addresses in any postal sense—but identifying numbers that the ASHA worker, Lakshmamma, has assigned herself based on the layout of the lanes she walks every week. House 14 is the one with the blue tarpaulin roof visible from the main drain. House 22 is the household where the grandmother speaks Telugu, not Kannada, and where Lakshmamma has learned to ask about fevers using a different phrase. House 31 is the family that moved here from Tumakuru six months ago, whose two children she has been watching for signs of underweight since their first visit.
This notebook is not a supplementary artifact. By any functional definition, it is a disease surveillance system. It just does not look like one to anyone trained to read databases.
Lakshmamma has been an Accredited Social Health Activist for eleven years in a settlement off Sarjapur Road that officially does not appear on most city maps. The settlement has approximately 340 households, a number that fluctuates weekly because demolition notices arrive and disappear, because construction work in nearby apartment complexes draws and releases migrant labor families, and because the boundary of what counts as “the settlement” shifts depending on which municipal department is doing the counting. She knows these 340 households by name. By symptom history. By the layout of their kitchens. By which ones have taps that work and which ones have taps that exist but have not produced water in two years.
The notebook contains four kinds of entries. The first is what she calls her fever list: every household where someone has had a fever in the past two weeks, with the date it started, whether it came with chills, whether it broke and returned, and whether the person went to the primary health center or took something from the local pharmacy. The second is her water list: notes on which houses reported the water tasting different, looking cloudy, or running from the tap with a smell she describes as gehun—a word that in this context means something between “earthy” and “wrong.” The third is her breath list: names of children and elderly residents coughing for more than two weeks, with a shorthand notation she has developed herself, using symbols to mark whether the cough is dry, productive, worse at night, or accompanied by the particular sound she associates with past tuberculosis cases in the area. The fourth list has no formal name but functions as a cross-reference: which houses had a wedding, a funeral, or a religious gathering in the past month, because she has learned that these events predict fever clusters two to three weeks later.
This is not nostalgia for paper. The point is not that handwritten is better than digital. The point is that the act of naming and recording what counts as a health event—what Lakshmamma does every day with a ballpoint pen—is a form of epidemiological intelligence structurally different from what standardized digital surveillance systems capture. And that difference matters most in the communities whose health is already worst.
What Digitization Sees and What It Erases
Over the past several years, state health departments across India have been migrating ASHA worker reporting to digital platforms. The logic is reasonable on its surface. Paper records are difficult to aggregate, impossible to search across regions, and slow to surface in district-level dashboards. A digital form can be submitted from a phone, routed to a supervisor in real time, and counted in a state-level summary by end of day. For a system that has historically struggled to know what is happening in its own facilities, the appeal of real-time data is enormous.
But the digital form Lakshmamma is now required to fill has fields that do not match the information she collects. The form asks for the number of fever cases in the past week, categorized as “fever with rash,” “fever with bleeding,” “fever with altered consciousness,” or “fever other.” It asks whether the household has a toilet. It asks whether the pregnant woman in the household has received her iron-folic acid tablets. It does not ask about the water tasting wrong, because there is no field for taste. It does not ask about the funeral three weeks ago, because there is no field for social gatherings. It does not ask about the cough that sounds like the one that preceded a tuberculosis diagnosis in the adjacent lane last year, because there is no field for sound.
The form is not malicious. It is the product of a design process that began with what the health system wants to count, not with what the health worker already knows. And that starting point determines everything that follows.
Consider what happens to the water list. In the digital system, water quality is tracked through a separate vertical. Samples are collected at predetermined points—the treatment plant, a distribution reservoir, a public tap identified by the municipal water board—and tested for pH, turbidity, and bacterial counts. These readings are then aggregated into a ward-level water quality score. The system is designed to answer the question: “Is the water in this ward safe?” It is not designed to answer the question: “Is the water at House 14, the one with the blue tarpaulin roof, safe today?” That question requires a different kind of data, collected at a different spatial and temporal resolution, by someone already walking past that house.
The water list in Lakshmamma’s notebook captures something the municipal sampling protocol cannot: intra-ward variation in water quality that correlates with the geometry of the distribution network, the timing of tanker deliveries, and the specific storage practices of individual households. When she notes that House 14’s water “tastes wrong” on a Thursday, she is recording a data point that, if aggregated across weeks and across her peers’ notebooks, would reveal a pattern. The water in the lower-elevation lanes of the settlement is contaminated after the tanker refill on Wednesdays, because the tanker’s hose touches the open drain at the lane entrance. No municipal sampling protocol captures this, because no municipal sampling protocol samples at House 14 on a Thursday morning. The sample is taken at the tanker, at the distribution point, or at a tap selected when the ward was first mapped—and that tap may have been disconnected two years ago.
The Politics of Categories
Every surveillance system depends on acts of naming. Before you can count a case, you have to define what counts as a case. Before you can define what counts as a case, you have to decide which symptoms, in which combination, at what threshold of severity, constitute an event worth recording. These decisions are embedded in the architecture of the form, the structure of the database, and the logic of the dashboard. They are not neutral.
That same discipline applies to title and framing decisions: before publishing, editors need a way to test a heading promises the same thing the article actually delivers, which is where a novel title generator that fits the project can function as a planning aid rather than a substitute for domain evidence.
The same principle—systems can only see what they are designed to name—appears in engineering and cybersecurity frameworks, but it is most consequential when applied to communities whose health risks do not fit predefined categories. The Google SRE framework makes this explicit: what a monitoring system classifies as an event is what it can learn from, and everything outside that classification is invisible. Similarly, the NIST Cybersecurity Framework encodes assumptions about what kinds of threats are worth detecting, and its standardization—designed so that one organization can recognize what another has classified—comes at the cost of excluding risks that do not fit its vocabulary. In public health, this tradeoff is not abstract. When the system has no slot for a cough that sounds like last year’s TB, or for water that tastes wrong, the information does not get lost in transit. It never enters the system at all.
The ASHA Worker as an Epidemiological System
What Lakshmamma’s notebook reveals is not that informal systems are superior to formal ones. It is that they are doing a different kind of work, and that work is irreplaceable. The notebook is a situated surveillance system. It records health events in the specific context of a specific community, using categories developed through years of observation in that community. The categories are not generalizable—they would not work in a settlement in Mumbai or a village in Odisha without adaptation. But they are precise for this settlement, and that precision is what makes them useful.
Consider the breath list. Lakshmamma’s shorthand for cough types is not a clinical classification. It is a local taxonomy developed from eleven years of listening to coughs in this specific community, cross-referenced with subsequent diagnoses at the primary health center. When she marks a cough with a particular symbol, she is not diagnosing tuberculosis. She is flagging a household for follow-up, based on a pattern she has learned to recognize. In the digital system, this cough becomes “fever other” or nothing at all, because the form does not have a field for “cough that sounds like the one that preceded a TB diagnosis in the adjacent lane last year.” The information is not lost because Lakshmamma stops noticing it. It is lost because the system she reports to has no place to receive it.
Or consider the cross-reference list—the one that tracks weddings, funerals, and religious gatherings. This is, functionally, a social mixing pattern database. It records the events that bring households into contact with each other and with visitors from outside the settlement, creating the conditions for disease transmission. Formal contact tracing systems attempt to reconstruct these networks after a case is identified, by asking patients where they have been and whom they have met. Lakshmamma’s notebook records these networks prospectively, before anyone gets sick, because she is already present in the community and already knows what is happening. The difference between prospective and retrospective social mixing data is the difference between an early warning system and a confirmation system. The digital form, as currently designed, can only function as the latter.
What Gets Lost in the Gap Between Knowing and Reporting
The gap between what Lakshmamma knows and what the system receives is not just data loss. It is a structural blind spot that reproduces health inequity. When the digital system reports that there are three fever cases in the settlement this week, and the notebook records seventeen, the discrepancy is not a rounding error. It is a difference in what counts as a case. The digital system counts cases that fit its categories. The notebook counts cases that Lakshmamma has learned to recognize as significant, using categories she has developed through sustained observation. The three cases that make it into the digital system are the ones that match the form’s predefined fields. The fourteen that do not are the ones that would, in a different system, trigger an early response.
This is not a hypothetical concern. In the settlement where Lakshmamma works, there was a dengue outbreak in 2022 that the formal surveillance system detected three weeks after the first cases appeared. Lakshmamma’s notebook showed a spike in “fever with chills and eye pain” beginning in the second week of June. She reported this to her supervisor, but the digital form’s categories did not include “eye pain” as a symptom field, and the supervisor had no mechanism to flag a pattern that did not match the system’s case definition. The outbreak was confirmed when three children were admitted to the district hospital with platelet counts below 50,000. By that point, Lakshmamma’s notebook already showed twenty-three households affected.
The system did not fail because it was digital. It failed because the digital form was designed to capture what the system already knew to look for, not what the community was already experiencing. The form’s categories were derived from a national surveillance protocol that defines dengue by a specific combination of symptoms, laboratory confirmation, and clinical severity. This definition is appropriate for aggregate-level surveillance and for comparing trends across districts. It is not appropriate for detecting the beginning of an outbreak in a settlement of 340 households where the first signal is not a laboratory test but a pattern of fevers that a trained community health worker has learned to recognize.
The Cost of Legibility
The political scientist James Scott, in Seeing Like a State, argued that modern states have a recurring impulse to make populations “legible”—to standardize naming, categorization, and measurement so that citizens can be counted, taxed, and governed at scale. This impulse is not inherently destructive. Standardized categories make it possible to compare a ward in Bengaluru with a ward in Delhi, to allocate resources based on population, and to detect outbreaks that cross administrative boundaries. Without standardization, there is no aggregate, and without aggregate, there is no system.
But the cost of legibility is the erasure of what does not fit. And in health surveillance, what does not fit is often the earliest signal of an emerging problem—the fever that does not match the case definition yet, the cough that sounds like last year’s TB but has not been confirmed, the water that tastes wrong but tests clean at the sampling point. These signals exist in the notebook and nowhere else. When the notebook is replaced by the form, they do not migrate. They disappear.
Designing for Coexistence, Not Replacement
The question is not whether to digitize. The question is whether the digital system can be designed to receive what the notebook already contains—whether it can be built to accept the categories that community health workers have developed through sustained observation, rather than only the categories that the central system has predefined. This is a design problem, not a technology problem. It requires starting from the community’s epistemology, not the institution’s ontology.
This principle extends to every domain where naming determines outcomes. The ASHA workers I have spoken with across Karnataka do not think of their notebook entries as “data”—they think of them as observations, the way a clinician thinks of a patient history. But when those observations need to travel upward into a system that speaks a different language, the act of titling and framing them becomes political. The same dynamic appears wherever people must compress lived knowledge into categories they did not design—whether that is a community health worker labeling a symptom cluster or a researcher using a novel title generator to find the framing that makes a story legible to an audience that does not share its context. What you call a thing determines what you can say about it, and what you cannot.
There are practical steps that follow from this. First, digital health tools designed for ASHA workers should include free-text or narrative fields alongside structured categories, so that observations that do not fit predefined slots are not lost but are available for pattern recognition by supervisors and epidemiologists. Second, the categories themselves should be reviewed periodically against the notebooks, not against the dashboard, to identify what the system is failing to capture. Third, the design process should include ASHA workers as co-designers, not as end-users, because they are the only people who know what the form is missing.
The Next Outbreak
The next outbreak in Lakshmamma’s settlement will not announce itself through a laboratory confirmation or a hospital admission. It will announce itself through a cluster of fevers that do not match the form, a change in the water that no one is sampling, a cough that sounds like one she has heard before. The notebook will record it. The digital system will not. The question is whether the health system will have the humility to read the notebook before the outbreak becomes the kind of event it is finally designed to detect.
This is not a call to abandon digital surveillance. It is a call to recognize that the most sophisticated epidemiological intelligence in this settlement is written in ballpoint pen on a twelve-rupee notebook, and that the person holding the pen has been doing this work for eleven years without being asked what she knows. The next outbreak is already visible. The question is who is willing to look where it appears.