Comparing Geographic Data: British Colonial and South Asian Texts

Background for this study

Pathways in Kim. (Just for illustration)


I’ve spent a good part of the summer working with a graduate student collaborator on a Scalar project that puts Colonial British writing next to South Asian texts from the same period. (I outlined the basics of the new "Literature of Colonial South Asia" project in my previous post.)

The idea for the present project came out of an earlier project called The Kiplings and India. In that project, I found myself growing frustrated with the writings of the primary focus of the digital archive, namely the Kipling family, and wanted to explore a broader conversation about colonial India that would include more Indian/South Asian voices as well as other British voices. 

The current project is an attempt at modeling the broader conversation. Historically, it has been easier to think of these as two parallel but separate bodies of work – the world view and experience of writers like the Hindi/Urdu progressive Munshi Premchand or the Bengali writer Saratchandra Chattopadhyay was so different from that of a Rudyard Kipling that they aren’t really of a piece. However, here, I was interested in putting them together in a single corpus, in order to look for connections between them as well as points of divergence. 

To do that, one can take traditional qualitative (thematic and formal) as well as quantitative approaches. Thematically, a typical line of inquiry might be: British writers in India were famously fond of ghost stories – what about their Indian peers? (Quick answer: yes! Tagore, for example, wrote quite few ghost stories, though they were a little different from the kind written by writers like Kipling or Alice Perrin; see ghostly reincarnation.) Also, British writers tended to write a fair bit about their summer escapes to hill stations like Simla and Darjeeling. What about their Indian peers? What was their imaginative geography? The Tags on the Scalar site might provide an opening into some of these qualitative points of connection between the British and the colonial South Asian texts. 

Below, I also want to open up some preliminary quantitative inquiries, specifically on geographic references in the corpus. 


Analyzing Geography in Colonial South Asian Fiction: Experiments with SpaCy and Stanza


I’ve learned a lot from reading scholarship by Matthew Wilkens and Elizabeth Evans, particularly their co-authored book, Gender and Literary Geography (2025), and their co-written articles, especially “Nation, Ethnicity, and the Geography of British Fiction” (Journal of Cultural Analytics, 2018), and I wanted to briefly explore their methods and conclusions before giving my own analysis from the corpus described above. 

For non-DH readers, this type of work relies on software that extracts geographic information from texts using Named Entity Recognition (NER). Essentially, software can figure out from context that if you say, “We are going to Simla this summer to escape the heat of the plains,” that “Simla” is a geographic place entity (GPE). (It is generally not simply looking up words from a list [i.e., dictionary] of previously recognized places.) There are several software packages that do this kind of analysis; the two I have been exploring are SpaCy and Stanford University’s Stanza. Both are relatively easy to use if you can run a Python environment like Google Colab. 

Are we making a map? (Not so much) It's probably important to note that what we're aiming for is not quite a map, but geographic data. Maps, as Wilkens and Evans point out, can be helpful at visualizing geographic information in a single novel or small group of novels. But if we are looking at a corpus of 100 or more texts, there will be so much data that any maps we produce will look like noise. So what we are actually trying to do here is collect and analyze geographic data

In Gender and Literary Geography, Wilkens and Evans devised analytics to study “geographic mobility,” “international attention,” and “spatial range” in modern British fiction. First, generally speaking, modern British literature was remarkably mobile and international – Wilkens and Evans find that roughly two-thirds of all geographic references in their corpus were to places beyond British borders (this is much higher than the number found in American fiction). This makes sense given the proximity to Europe and the importance of European travel to British people in the upper classes. It also makes sense given the scale and scope of the British Empire throughout the 19th and early 20th centuries. 

They are also interested in the gender breakdown – it’s been traditionally assumed that writers who are men are more mobile than women, so there should be greater evidence of mobility and spatial range. However, Wilkens and Evans find that character space (the gender of characters in a novel) is actually a bigger predictor of mobility than author gender

One analytic they use that I find especially helpful and relevant in the South Asian context is “geographic intensity.” Geographic intensity speaks to the frequency of mentions of geographic place-names; certain writers tend to constantly mention place names when their characters are traveling; others not so much. One finding Wilkens and Evans have right off the bat is that nonfiction texts tend to show much higher geographic intensity than do works of fiction – so there’s no point comparing them. (Nonfiction books are much more likely to be instructional in some way, or structured as reference texts.) My own experiments with the present corpus lend support to that distinction – a book like Cornelia Sorabji’s nonfiction memoir India Calling has a much higher geographic intensity than do the same author’s fictional texts, so we should really bracket off the nonfiction works. 

My hunch going into this would be that British writers would have a much higher degree of geographic intensity in their accounts of colonial India than their Indian peers. To begin with, Anglo-Indian writers traveled back and forth quite a bit (between England and India), and with their assumed English readership they might be inclined to announce all of the exotic places their characters might visit in the course of a novel. But what will the data show? 

Measuring Geographic Intensity. Wilkens and Evans have a sophisticated way of measuring geographic intensity, though in my own experiments so far, I take a relatively straightforward approach – I ask NER software like Spacy or Stanza to extract a list of place-names (including duplicates), and then create a ratio based on the total number of words in a text. A geographic intensity score of 0.72% -- what I found for Rudyard Kipling’s Kim using SpaCy -- means that just under 1% of all the words in the novel are place-names. (Note: another software package, Stanza, gives a rather different result for Kim here – just 0.38%. I’ll talk more about what might be going on below)

In Kim, the SpaCy software finds a list of place-names that looks like you might expect. Here is a selection: 

Afghanistan, Aminabad, Amritsar, Bengal, Bhotiyal, Bikaner, Bombay, Cabul, Dacca, Delhi, Lahore, England, Ferozepore, Gunga, Hindustan, India, Indus, Karachi, Kotgarh, Kurdistan, Ladakh, Lahore, Lucknow… 


Most of these are familiar names, though many use archaic or idiosyncratic spelling (Dacca instead of Dhaka; Cabul instead of Kabul; Gunga instead of Ganga/Ganges). Idiosyncratic spelling could be challenging, especially if one plans to feed the output list into software that might produce a map. That said, at least one platform I have worked with -- Pelagios's Recogito -- does do a good job of standardizing archaic place names. When I ran Kim through their system, it successfully located "Dacca" in "Dhaka." (See Recogito's version of Kim here).

Another map of Kim from Voyant Tools


However, SpaCy also finds a fair number of false positives – for some reason, SpaCy thinks “Mahbub” is a place-name (though readers of the novel would know that this is part of a character name, Mahbub Ali). Those false positives throw off the score – for some books, I found the false positive effect to be as high as 20% – so the final number should not be taken as an accurate one. Along with false positives, I noted a fair number of false negatives – locations that I see as clearly evident in the text, but which Spacy missed. These could be as high as 10%.

SpaCy vs. Stanza. Stanza is better. If one compares the SpaCy results against the Stanza results, the Stanza results seem much better. The false positives are much lower (<5%) and the false negatives are also much better (also <5%). Based on those results, ideally one would just use Stanza. However, the downside of Stanza is that it is considerably slower and more resource intensive than Spacy – so running a small corpus of 20 novels on a basic Windows laptop could take two hours or more to complete. 

As of now, I have not run all of the fictional texts I have in my corpus through the system, though I hope to do that sometime when I have access to a GPU. For now, I have just done two small tests with ten texts each.

Here is my geographic intensity data for ten relatively random texts by mostly British writers set in India (Sara Jeannette Duncan was Canadian), with both Stanza and Spacy results.  






There is not a lot we need to say about this chart, though it is worth noting that the writers who are women here (B.M. Croker, Sara Jeannette Duncan) are ranked at essentially the same level as the men. Kim seems like an outlier using Spacy, but with Stanza its results are roughly comparable to others in this small corpus. 

And here is comparable data for geographic intensity in ten random texts by Indian authors from this period




Preliminary observation. One notable result: the numbers for these Indian writers – again, a fairly random selection of ten works of fiction – are slightly lower than the percentages for their English peers, though not by very much. It is also worth noting that the books by Tagore and Bankim Chandra Chatterjee were translated from the Bengali, while the books by S.M. Mitra and Cornelia Sorabji were originally published in English by British publishing houses. From looking at the texts themselves, it is clear that these two books were primarily written with English readers in mind. 

An Anomaly. In the chart above, there is a rather anomalous result for a relatively obscure book called Hindupore by S.M. Mitra. This is a travel narrative and political commentary, only thinly fictionalized, with an Irish Member of Parliament traveling to India with an Indian Maharajah. Mitra on the whole is a bit of an odd bird – he was essentially a supporter of British rule, who made a career of explaining Indian culture for English readers in often simplistic ways. I am keeping the text on the list for the present, though in terms of geographic intensity it behaves more like a nonfiction travel narrative than a typical work of fiction. 

A case could be made that for Indian writers who lived permanently abroad, like Mitra (or, in the U.S. context, Dhan Gopal Mukerji), their texts should be in a different category than texts by Indian authors who lived in India itself -- and maybe excluded from a corpus of "Indian fiction" in future tests. 
Global interconnections in S M Mitra's Hindupore


Preliminary Observations: In short, at least with these small test corpora, there does appear to be a slightly higher geographic intensity result for the British colonial texts than for the Indian ones. That said, the difference is not major -- I now want to go back to basics and re-read (with my eyes, not with a computer) texts like The Home and the World with an eye to how places are described. Also, as Wilkens and Evans concluded with a different corpus, there is also not a vast difference between men and women authors with respect to geographic intensity (on the British side at least). 

In the next few weeks I will also aim to expand the study to the full corpus to confirm these results, while also studying the granular results from various software packages. (What are the false positives and false negatives in each set of results?) 




New Project: Literature of Colonial South Asia

I've been working on a new project, and it seems time to announce it publicly for a kind of soft-launch. Last summer, I was fortunate to work with a graduate student collaborator, Srishti Raj, on a project on Adivasi Writers. This summer, I've been working with another student, Sana Asifriyaz, on a project called Literature of Colonial South Asia. (Both projects were supported by internal English department grants that allowed the student to be paid a small stipend while working on the project.)

This is a rather different kind of project -- several years ago, I had already assembled a text corpus with some of these materials as a plain text folder on Google Drive. It seemed like something someone might need, though I myself didn't do very much with the corpus, and I didn't hear back from many scholars who appeared to be using it. 

The present collection takes the basic idea of that corpus -- a collection of materials from the colonial period that includes both South Asian writers and Western writers who lived in India -- and presents it in a way that I hope will make the value of collecting these materials in this way a little clearer. The version we put online is not just a corpus anymore; it's also now a kind of Anthology (with editorial commentary and context) as well as a Digital Archive. 

Some of the work we have been doing as we've been creating plain text versions of primary texts on the site include creating contextual introductions that focus on the following topics: 

The idea of these various pages is to enable unpacking of the different constituencies in the collection.  As with other Scalar projects I've worked on, one can also use thematic and contextual tags to see links and connections that might not typically be visible. (For instance, the "Calcutta" tag is connected to a number of texts by Tagore and other Bengal Renaissance writers, but also texts by Western writers in India like Edward Money, W.D. Arnold, and Sara Jeannette Duncan, who also lived in or wrote about Calcutta in their works.) 

The Heart of the Collection: Tagore's Short Stories. In some ways, the heart of the current collection might be Tagore's short stories. Putting together this collection of nearly 50 short stories in translation started out as what we were calling a "mini-project," but soon became something bigger. Only around 35 or so of Tagore's nearly 100 published short stories could be found in translations that were out of copyright. Others were available in some form in old magazines like The Modern Review. And a few we decided to translate ourselves using Generative AI (while comparing the versions created by the LLM against published translations by translators like Carolyn Brown or William Radice). The tag system we developed to show the relationships between these stories became a kind of "ontology" that shaped our thinking about other texts on the site. (See the Tag Cloud and Network Visualization here.)

Classroom use: I would hope that the pages listed above might be helpful to folks in the classroom. For instance, I could imagine a really effective unit in classes on writing by women that could use the "South Asian Women Writers" collection as a starting point for students. Similarly, I could imagine a really profitable exploration of Tagore's writings through his short stories (which are world-class -- and generally underappreciated in Tagore's body of work). 

Some of the materials on the site are admittedly carried over from previous projects (such as The Kiplings and India, and The Collected Writings of Henry Louis Vivian Derozio). 

Centering South Asian perspectives: As mentioned, here we tried as much as possible to ground the 'ontology' of the tag system we developed for this project in writings by South Asian writers, not so much Kipling and his peers. If you center writers like Kipling and B.M. Croker in your account of colonial South Asia, you have a sense of Indian life dominated by summers in the Hill Stations, lots of big-game hunting ("Shikar"), and occasional encounters with Indians in service roles ("Syce," "Mahout," "Khitmutgar," "Ayah," etc.). On the other hand, if you center the experience of writers like Tagore or Pandita Ramabai, you get a very different picture -- and you might make different choices regarding site architecture, conceptual and thematic categories, and so on. 

For example, both Western writers and Tagore wrote "Indian ghost stories" that are included on the site, but Tagore also has a category of stories that invoke something I call "ghostly reincarnation," where a deceased loved one related to the protagonist reappears in ghostly form -- slightly different than a conventional Western ghost story... (Though admittedly, Kipling does have at least one story that fits that tag.) Again, see the Tag cloud and network visualization to get a sense of how our ontology shaped up in developing this project. 

Decentering Kipling/Interracial romance: Even within the body of texts by Western writers in India, there are certain advantages to decentering Kipling. For example, Kipling was famously dismissive of interracial romance and uninterested in mixed-race ("Eurasian") communities in India; if you just read Kipling, you would think that was the generally prevalent view of all British colonials in India. However, if you center women who were his peers in many respects -- writers like Maud Diver, F.E. Penny, and B.M. Croker -- you see a pretty extensive array of texts that explore interracial romance and mixed-race characters in interesting ways.  

Summaries and the Great Unread: The summaries also seemed like an important innovation, especially for researchers looking for a way into new materials and authors. There is a pretty vast "great unread" of material from this period, whether it's the novels of F.E. Penny and B.M. Croker or Indian writers like Peary Chand Mitra or Sarat Chandra Chatterjee (Chattopadhyay). The summaries we developed, sometimes using help from NotebookLM (Gemini Notebook), might help researchers get a sense of some of these texts before diving in to read them in toto. (The tags and the pages listed above might also help readers find constellations of texts on certain topics of interest.) 

There is more to come with this project -- for instance, I have some quantitative comparisons I want to work on -- but for now I just wanted to let people know the project is there & available to use. 

Statement of Purpose Tips (English Graduate programs -- M.A. and Ph.D. applications)

I served as the Director of Graduate Studies in the English department at Lehigh for three years; before that, I served on the graduate admissions committee for more than a decade. Over that time, I read hundreds, if not thousands, of applications, and advised dozens of my own students. Here are some tips for the Statement of Purpose, one of the most important documents in the graduate school application. This document emerged out of committee work, so some of my colleagues in the English department may have contributed ideas or language to the document below, though the vast majority of the prose is my own.

Applicants will also want to be aware of the extreme challenges of the current academic job market for English Ph.D.s. These tips assume you already know how bad things are, and still want to give it a go. 


General hint: this is a very challenging document to write for most applicants, and you should plan to do several drafts & versions. Start early! Preferably in the summer before the fall when you will be applying. Before you start to write, take some time to think about how you'd like to present yourself professionally and intellectually. 

As you work on the document, you should also plan to get feedback from trusted friends (especially those who know how English graduate programs work) and faculty mentors. It is especially important that faculty mentors who will be writing your recommendation letters see your SOP, as they will inevitably be using that to shape their own letters on your behalf.   

What is the Statement of Purpose? 

Fundamentally: The SOP should:


1) describe your research interests, including traditional period and region interests,
2) give an overview of previous experience and background (courses taken, sampling of topics covered), and
3) identify theoretical investments.

Slides for ACH (2026):

 I'm giving a talk at the Association for Computing in the Humanities conference, which is being held virtually June 24-26 this year. My talk will be on June 26 in the morning. The talk describes a newer digital collection, Adivasi Writers.

Slides for "The Space Between" (2026) : Publications of the Ghadar Party 1913-1930

I'm giving a talk at the Space Between / FIMA conference in Greensboro, NC, this week. This is a new project exploring writings from the Ghadar party. There is also a thread following the American writer Agnes Smedley, who worked closely with both Ghadarites and the Young India activists in New York in the late 1910s, and later published an autobiographical novel, Daughter of Earth. 

Slides for ALA 2026: African American Periodical Poetry: A Data-Driven Approach

 I'm giving a talk at the ALA this year, based on the periodical poetry dataset I developed for the  "Responsible Datasets in Context" project. Here I describe what that dataset is and why I wanted to construct it. I am also sharing some of my conclusions and observations about what the dataset shows, and how it might be used.

Gendered Pronouns in Early 20th Century Fiction: A Simple Quantitative Study

Gendered Pronouns in Early 20th Century Fiction: A Simple Quantitative Study

The following short essay is a work in progress -- I am exploring the uses of a corpus of early 20th century literature I have been developing for a few months. The study below represents an attempt to make use of that corpus to query a topic that has been of interest in quantitative DH in recent years. 


I have long been fascinated by a DH paper published in 2018, “The Transformation of Gender in English-Language Fiction” (link here; authors were Ted Underwood, David Bamman and Sabrina Lee) that has suggested strong statistical evidence that men were increasingly dominating the world of fiction in late 19th and early 20th centuries – that between 1850 and 1950 the percentage of published novels that were authored by women dropped dramatically (from near parity to more like a third or a quarter). Thus, at the exact period when we might have expected women to be gaining visibility and influence – associated with the early 20th-century suffrage movement and the appearance of important feminist voices like Virginia Woolf – they were actually losing position on the whole in the publishing world


According to the authors, the pattern only started to reverse in the second half of the twentieth century (and today, the publishing industry would of course look very different). Also, within their fiction, “The Transformation of Gender” authors indicate that men writers tend to write more about men, while writers who are women might be closer to gender parity in the amount of time given men and women in the social world represented in the story. The authors suggest that particular tendency hasn’t improved or changed as much.


Source: Underwood, Bamman, and Lee (2018)

Incidentally, the concern with the growing marginalization of writers who were women alluded to above is not a new one. The authors of “The Transformation of Gender” cite a 1989 study, Edging Women Out: Victorian Novelists, Publishers, and Social Change (Gaye Tuchman and Nina Fortin), where the authors did quantitative (but not digital!) scholarship with similar findings. Tuchman and Fortin counted and classified entries in Leslie Stephen’s Dictionary of National Biography to compare how women writers were talked about versus men writers. They found that while books by men were reviewed more frequently on the whole, the gender disparity in the more recent authors (late 19th century) became especially sharp with respect to works of nonfiction.  The authors of “The Transformation of Gender” used a very large corpus of tens of thousands of novels from HathiTrust (and checked against the smaller University of Chicago novel corpus) as well as sophisticated modeling techniques built around Natural Language Processing (NLP) to infer gender within a text and derive percentages. Some years ago, I finally gained enough confidence in basic Python to explore some of these methods on my own, using David Bamman’s BookNLP software (sadly, that software does not appear to be working at present, so I will not be using it for the results below).


One other bit of background: in the revised version of the essay published in his book, Distant Horizons, Ted Underwood mentions the Gendered Language Visualizer, a simple but deceptively powerful tool that tracks the association between non-gendered words and gendered pronouns in works of fiction. The technique behind that led to the beautifully illustrative image below (from the jointly written 2018 essay)


Source: Underwood, Bamman, and Lee (2018)

What it shows: women in fiction tend to "smile" and "laugh"; men tend to "grin" and "chuckle." (Though note that the divergence diminishes over time -- so in contemporary fiction that 'gendering of mirth' would be much less pronounced than it was at the peak of the divergence, around 1950.) 


Earlier studies: I should say that this is a more complex version of a type of analysis scholars have been doing in stylistics for many years; there are studies that go back to the 1990s that aimed to predict the gender of a writer based on characteristics of function words and articles. Koppel et al. (2002) used sophisticated statistical techniques with a fairly straightforward counting to find that writers who are men tend to use a higher proportion of noun specifiers (a, the, that), and numbers in their fiction. They also claim women tend to use more pronouns (she, herself), negation (not), and certain prepositions (for, with) and conjunctions (and). By lining up counts of these various parts of speech, the authors claim to be able to predict the gender of an author of an anonymized text with 80% accuracy. (Note: for what it’s worth, I tried to replicate their results with my own small, early 20th-century corpus, and failed. The only place where I saw a clear correlation was with gendered pronouns -- which might explain how I got to the design of the present study below.)


Moving past binarized gender thinking: Admittedly, I am not so interested in this particular application for my own research – it’s almost never the case with 20th-century fiction that the gender identity of an author is unknown. I also tend to be interested in writers who pushed against conventional gender roles and expectations in any case, many of whom might be understood as LGBTQIA+ today – writers like Virginia Woolf, E.M. Forster, D.H. Lawrence, Radclyffe Hall, or Wallace Thurman. Today, most scholars would find the "predict the gender" type of analysis overly restrictive and as essentially reinforcing binarized gender thinking. If E.M. Forster, for example, breaks with the expected pattern in novels that feature women protagonists (spoiler: he does!), that would be a more interesting finding than simply that reconfirming that 80% of men are from Mars, as it were.  


A simplified method for the present study: What if we drastically simplified the query with a corpus of early 20th-century fiction? As a starting point for thinking about patterns with respect to gendered socialization, why not simply look at gendered pronouns: he/him/his and she/her/hers? If the conclusions by Underwood et al. are correct, we should expect to see a lopsided homosocial tendency in fiction by men (men mostly talking to and about other men, and only occasionally mentioning a woman), and maybe a more balanced gender representation in fiction by women. We might also see some interesting anomalies in the patterns that might be worth exploring.  


Before doing this at a mid-range scale, I was curious to see how authors I know would shake out. Over the past few months, I’ve been developing a custom corpus of early 20th-century texts. I have described the basic design of the corpus here; it contains about 1000 total texts, including about 100 texts that might be thought of as canonical high modernist texts, 130 texts by African American authors, and about 90 texts associated with colonial South Asia. It also contains a substantial amount of genre fiction. The results below only reference works of fiction, though there are works of poetry, drama, and nonfiction in the corpus. 


With a little help from generative AI coding assistants, I devised a simple bit of code to count the use of gendered pronouns (he, him, his vs. she, her, hers), first, in a single novel, then in a batch of files, and then derive a percentage from the total word length of the file. I then took those gendered pronoun percentages, and compared them to one another to get a ratio. Rather than overwhelm the reader with a vast array of raw data, I’ll start with some smaller findings, initially focused on gendered pronoun ratios in a small set of ‘high modernist’ works of fiction, mainly by white British and American authors. I’ll then expand the conversation to other authors and consider broadly why any of this might be significant. 


From my limited high modernist collection, what are some texts that are especially lopsided towards men? (If you expected to see Ernest Hemingway on this list, you would be right!)




Text Ratio of masculine to feminine pronouns
Ernest Hemingway: Men Without Women 11.4 to 1
Hemingway: In Our Time 9.4 to 1
James Joyce: Portrait of the Artist as a Young Man 9.2 to 1
John Dos Passos: Three Soldiers 7.3 to 1
D.H. Lawrence: Kangaroo 4.2 to 1
Hemingway: The Sun Also Rises 3.5 to 1
James Joyce: Ulysses 3.0 to 1
James Joyce: Dubliners 2.2 to 1
E.M. Forster: A Passage to India 2.2 to 1
F. Scott Fitzgerald: The Great Gatsby 2.0 to 1

What to make of the lopsided nature of some of these texts? I should say, off the bat, that I don’t think the lopsidedness necessarily serves as an indictment of someone like Hemingway. The relative absence of women in his various short stories is partly due to their settings (several in Men Without Women deal with soldiers and World War I, and “The Undefeated,” about an aging Spanish bullfighter out for a last hurrah, is a pretty marvelous critique of dysfunctional masculinity). Moreover, A Portrait of the Artist as a Young Man is a coming-of-age narrative for Stephen Dedalus at schools that only admit boys and men with teachers who are also only men, so it’s not a huge surprise that the social world represented in the text is also pretty lopsided. (The imbalance might have been less if Joyce had kept in more of the love interest/romantic sections that were in the original Stephen Hero version of his manuscript.) The lopsidedness of other writers (and other Joyce texts) is less extreme, though it’s striking to see novels by D.H. Lawrence and E.M. Forster here (especially since Forster, with Howards End, is also on my second list below). 


Again, I don’t see it as an indictment per se, or as a reason to drop Hemingway or Joyce from my syllabus, though it is still worth knowing. (Do readers want or need to see characters that match their own gender identity or expression in order to connect with a text? Probably not, but my hunch is that it might help...) Still, the pattern does appear to show that there is a pretty limited role for women in the social worlds we find in these texts. It is not as if the authors don’t know it, either: the title Men Without Women can be read as self-critique of a symptomatic nature. These are men without women, and perhaps that’s why they are so broken.


And what about woman-centered texts by writers of literary fiction typically associated with high modernism? 



Text Ratio of feminine to masculine pronouns
Dorothy Richardson: Pilgrimage 1Pointed Roofs 13.1 to 1
Bryher: Development 9.9 to 1
Richardson: Pilgrimage (other volumes) varies between 5 to 1 and 2 to 1
Nella Larsen: Passing 4.9 to 1
Radclyffe Hall: The Unlit Lamp 3.9 to 1
Radclyffe Hall: The Well of Loneliness 2.8 to 1
Wallace Thurman: The Blacker the Berry 2.8 to 1
Gertrude Stein: Three Lives 1.6 to 1
Virginia Woolf: Mrs. Dalloway 1.6 to 1
Katherine Mansfield: The Garden Party And Other Stories 1.6 to 1
Mansfield: Bliss and Other Stories 1.5 to 1
Virginia Woolf: The Voyage Out 1.5 to 1
Woolf: Night and Day 1.4 to 1
Woolf: Orlando 1.4 to 1
Woolf: To the Lighthouse 1.3 to 1
Forster: A Room With a View 1.2 to 1
Forster: Howards End 1.2 to 1
It was not hugely surprising to see Pilgrimage: Pointed Roofs as the most lopsided she/her centered text in the high modernist selection from my text corpus. Pointed Roofs is the story of a young woman teaching at a girls’ boarding school, so, as with Portrait of the Artist above it is not surprising that it reflects a homosocial world with largely girls and women as characters. 


Also, anyone who has read Passing recently would not be surprised to see how prevalent she/her/hers pronouns are there: it really is a novel focused on the relationship between two women. (If anything, this finding only reconfirms readings that have stressed the homoerotic subtexts of that relationship.)


I was intrigued to see a book by a man, Wallace Thurman, come out fairly high on this list (2.8 to 1). I am not entirely sure what to make of it; the novel in question is a thoughtful and often bitter account of colorism within the Black community with a woman protagonist. 


The bigger takeaway might be that the pattern described by Underwood et al. appears to be in evidence with this small group of high modernist writers – writers who were women were, on the whole, less lopsided than were their peers who were men. Instead of a ratio of 10 to 1 or 4 to 1 or even 2 to 1, the median here for writers like Woolf and Mansfield – two of the core authors in the modern feminist canon – is closer to 1.5 to 1. 



Expanding the Range of Authors: Genre Fiction Writers


Now, let’s move to the broader dataset. The first discovery might be that the gendered pronoun disparity can be wildly lopsided in adventure fiction and westerns: 


Text Ratio of masculine to feminine pronouns


Zane Grey, The Young Pitcher 630 to 1

Zane Grey, Ken Ward in the Jungle 433 to 1

G. K. Chesterton, The Man Who Was 

Thursday 139 to 1

H. G. Wells, The First Men in the Moon 104 to 1

John Buchan, Prester John 101 to 1

Lord Dunsany, The Gods of Pegana 57 to 1

L. Frank Baum, The Master Key 47 to 1

Dhan Gopal Mukerji, Kari the Elephant 44 to 1

G. K. Chesterton, The Man Who Knew 

Too Much 43 to 1

Jack London, The Call of the Wild 18 to 1

John Buchan, The Thirty-Nine Steps 15 to 1

Dorothy Sayers, Lord Peter Views 

The Body 5.8 to 1

Agatha Christie, The Big Four 4.0 to 1


The scale of lopsidedness is pretty vast – and consistent – with early 20th century men who wrote westerns, science fiction, and detective fiction all showing a highly lopsided, man-centered social world. (I ran hundreds of titles for this study, and am only including a few noteworthy titles on these tables; readers who want to see the raw data can find it here; note that it contains texts that are not works of fiction--I've been disregarding those in the present study) Even women who wrote detective fiction tended to show a version of it, though Dorothy Sayers’ Lord Peter Views the Body (at 5.8 to 1) is still much less imbalanced than something like The Young Pitcher (another narrative of a young man at school, with no girls or women about). 


And what about woman-centered genre fiction / popular fiction? 


Text Ratio of feminine to masculine pronouns


Rokeya Hossain, Sultana’s Dream 10.8 to 1

Vita Sackville-West, The King’s Daughter 5.7 to 1

Edith Wharton The Old Maid 5.4 to 1

Elinor Glyn, Man and Maid 5.0 to 1

Gertrude Atherton, The Living Present 3.8 to 1

L. M. Montgomery, Anne of Green 

Gables 3.2 to 1

Somerset Maugham, Liza of Lambeth 2.9 to 1

Louis Bromfield, The Green Bay Tree 2.4 to 1

Edith Wharton, The House of Mirth 2.4 to 1

Zane Grey, The Call of the Canyon 2.4 to 1

Temple Bailey, Judy 2.0 to 1

H.G. Wells, Ann Veronica 1.8 to 1



Again, while there are some texts that are highly woman-centered (Sultana’s Dream is, famously, a feminist utopia with men kept in enclosures, while women run the world), the imbalance for romance fiction writers like Elinor Glyn or girl-oriented children’s fiction writers like L.M. Montgomery (Anne of Green Gables) is considerably less pronounced than with their counterparts who were men. 


Given how lopsided Zane Grey generally is, it is interesting to see one of his novels here (a shell-shocked World War I veteran moves to Arizona and has to choose between two different women). It’s also noteworthy to see an instance of H.G. Wells’ “new woman” fiction here. (Again, if anyone would like to see the full / raw data, it is here.) 


Quick conclusions; Next steps in the analysis?

Admittedly, this is a fairly crude method. At most, it shows some general patterns and trends, and confirms (albeit with a very small sample of texts) what Underwood/Bamman/Lee claimed using a much larger statistical model. Here is Underwood in Distant Horizons

“It turns out that women are consistently under-represented in books by men. On average, only a third of the words men use in characterization are used to describe feminine characters. Women writers, on the other hand, spend equal time on fictional men and fictional women. This difference remains depressingly constant across two centuries, and it may help explain why books by men tend to have more stereotyped gender roles.” (Distant Horizons, 127)

For me, the next steps might not be more quantitative queries. Rather, I am curious to look at the anomalies and exceptions in the early 20th-century corpus to try and learn more about what might have been going on, perhaps the old-fashioned way (i.e., actually reading the novels in question). For instance, for a writer who was so dramatically lopsided towards men otherwise, how did Zane Grey's  The Call of the Canyon feature women's voices in a 2:1 ratio? What was he doing differently here? (Especially curious since for someone like Hemingway, fiction responding to the psychic effects of World War I was often overwhelmingly oriented to men.)

Also, for writers like E.M. Forster and Wallace Thurman, both writers of literary fiction who lived their lives as closeted gay men, it is intriguing to see they both wrote novels with women protagonists who scored fairly high on the second table above. It might be interesting to gather together other novels written by cis-identified men with women as protagonists. Are there any patterns that can be gleaned from them? 

Finally, I'm curious about the representations of animals in the corpus. It's striking that a big part of the reason Hemingway's "The Undefeated" and Jack London's The Call of the Wild appear so lopsided in terms of gendered pronouns is that the vast majority of the animals are gendered male in both texts (alongside the human protagonists of those stories, of course). It might be interesting to make a small corpus of animal-oriented fiction from this fiction and study how animals are gendered (perhaps adding in more emphasis on non-gendered pronouns...).